Speech Generation Method, Apparatus, Electronic Device, and Storage Medium

By obtaining and processing the Mel spectrum features of the target user's voice, combining the characteristics of the text to be synthesized, and generating voices corresponding to the target identity characteristics and content characteristics, the problem of single tone in the existing technology is solved and personalized voice generation is achieved.

CN115116426BActive Publication Date: 2025-06-13BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210654618.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-06-13
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

Existing voice synthesis technology is difficult to generate voices that meet users' personalized needs, and the tone is single and cannot be flexibly adjusted.

Method used

By obtaining the Mel spectrum features of the text to be synthesized and the target user's voice, input it to the identity encoder and content encoder of the pre-trained speech generation model to generate speech corresponding to the target identity and content features.

Benefits of technology

It realizes that the identity and text characteristics are set flexibly according to the different target voice and the text to be synthesized, to meet the personalized needs of users and to generate diverse voices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116426B_ABST
    Figure CN115116426B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a voice generation method, apparatus, electronic device, and storage medium, and relates to the technical field of voice signal processing. The present disclosure aims to at least solve the problem in the related art that it is impossible to generate voice that meets the personalized needs of users. The method includes: obtaining the text to be synthesized and the target user's voice; determining the Mel spectrogram features of the target user's voice, and inputting the Mel spectrogram features of the target user's voice into the identity encoder of a pre-trained voice generation model to obtain target identity features; determining the Mel spectrogram features of the text to be synthesized, and inputting the Mel spectrogram features of the text to be synthesized into the content encoder of the voice generation model to obtain content features; inputting the target identity features and the content features into the decoder of the voice generation model to obtain the target voice; the target voice is the voice corresponding to the target identity features and the content features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of speech signal processing, and in particular, to a speech generation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the continuous development of artificial intelligence (AI), speech synthesis technology has been widely used, such as intelligent customer service, chatbots, etc. Speech synthesis technology can convert text into natural human voices. Specifically, speech synthesis technology collects multiple segments of speech of a natural person as training data, trains a speech synthesis model, and then synthesizes speech with the same timbre as this natural person according to the speech synthesis model.

[0003] However, the speech generated by using the above speech synthesis technology has a single timbre, that is, only one timbre of speech can be generated by one speech synthesis model in the above speech synthesis technology. It can be seen that although the current speech synthesis method can generate speech of various sentences, its timbre is fixed and it is difficult to meet the personalized needs of users. Summary of the Invention

[0004] The present disclosure provides a speech generation method, apparatus, electronic device, and storage medium to at least solve the problem in the related art that speech that meets the personalized needs of users cannot be generated. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a speech generation method is provided, including: obtaining a text to be synthesized and a target user's speech; determining the Mel spectrum feature of the target user's speech, and inputting the Mel spectrum feature of the target user's speech into an identity encoder of a pre-trained speech generation model to obtain a target identity feature; determining the Mel spectrum feature of the text to be synthesized, and inputting the Mel spectrum feature of the text to be synthesized into a content encoder of the speech generation model to obtain a content feature; inputting the target identity feature and the content feature into a decoder of the speech generation model to obtain a target speech; the target speech is a speech corresponding to the target identity feature and the content feature.

[0006] Optionally, determining the Mel spectrum feature of the text to be synthesized includes: using a preset speech synthesis model to obtain the text speech corresponding to the text to be synthesized, and determining the Mel spectrum feature of the text speech as the Mel spectrum feature of the text to be synthesized.

[0007] Optionally, the method further includes: obtaining multiple groups of first voice samples, each group of first voice samples including a first voice and a second voice; determining a first input sample for each group of first voice samples, the first input sample including the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice; respectively inputting the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice into a preset first neural network to obtain a first predicted identity feature of the first voice and a second predicted identity feature of the second voice; for each group of first voice samples, determining the identity feature difference degree between the first predicted identity feature and the second predicted identity feature to obtain the identity feature difference degrees of multiple groups of first voice samples; training the first neural network according to the identity feature difference degrees of multiple groups of first voice samples to obtain an identity encoder.

[0008] Optionally, training the first neural network according to the identity feature difference degrees of multiple groups of first voice samples to obtain an identity encoder includes: when the first voice and the second voice correspond to the same user, if the identity feature difference degrees of multiple groups of first voice samples are all less than or equal to a first preset threshold, then determining that the identity encoder is obtained; when the first voice and the second voice correspond to different users, if the identity feature difference degrees of multiple groups of first voice samples are all greater than or equal to a second preset threshold, then determining that the identity encoder is obtained; wherein, the second preset threshold is greater than the first preset threshold.

[0009] Optionally, the above method further includes: obtaining multiple groups of second voice samples, each group of second voice samples including a third voice and a fourth voice; determining a second input sample for each group of second voice samples, the second input sample including the Mel spectrogram features of the third voice and the Mel spectrogram features of the fourth voice; respectively inputting the Mel spectrogram features of the third voice and the Mel spectrogram features of the fourth voice into a preset second neural network to obtain a first predicted content feature of the third voice and a second predicted content feature of the fourth voice; for each group of second voice samples, determining the content feature difference degree between the first predicted content feature and the second predicted content feature to obtain the content feature difference degrees of multiple groups of second voice samples; training the second neural network according to the content feature difference degrees of multiple groups of second voice samples to obtain a content encoder.

[0010] Optionally, the second neural network is trained according to the content feature difference degree of multiple groups of second voice samples to obtain a content encoder, including: when the third voice and the fourth voice correspond to the same text and the third voice and the fourth voice correspond to different users, if the content feature difference degrees of multiple groups of second voice samples are all less than or equal to a third preset threshold, then the content encoder is determined; when the third voice and the fourth voice correspond to different texts and the third voice and the fourth voice correspond to different users, if the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to a fourth preset threshold, then the content encoder is determined; wherein, the fourth preset threshold is greater than the third preset threshold.

[0011] Optionally, the above method further includes: obtaining multiple sample voices, and determining the sample Mel spectrogram features of each sample voice; inputting the sample Mel spectrogram features of each sample voice into an identity encoder to obtain the sample identity features corresponding to each sample voice; inputting the sample Mel spectrogram features of each sample voice into the content encoder to obtain the sample content features corresponding to each sample voice; inputting the sample identity features and the sample content features into a preset third neural network to obtain the predicted Mel spectrogram features of each sample voice; for each sample voice, determining the Mel spectrogram feature difference degree between the sample Mel spectrogram features and the predicted Mel spectrogram features to obtain the Mel spectrogram feature difference degrees of multiple sample voices; training the third neural network according to the Mel spectrogram feature difference degrees of multiple sample voices to obtain a decoder.

[0012] According to a second aspect of the embodiments of the present disclosure, there is provided a voice generation device, including an acquisition unit, a determination unit, and a generation unit; the acquisition unit is configured to acquire a text to be synthesized and a target user voice; the determination unit is configured to determine the Mel spectrogram features of the target user voice, and input the Mel spectrogram features of the target user voice into an identity encoder of a pre-trained voice generation model to obtain target identity features; the determination unit is further configured to determine the Mel spectrogram features of the text to be synthesized, and input the Mel spectrogram features of the text to be synthesized into a content encoder of the voice generation model to obtain content features; the generation unit is configured to input the target identity features and the content features into a decoder of the voice generation model to obtain a target voice; the target voice is a voice corresponding to the target identity features and the content features.

[0013] Optionally, the determination unit is specifically configured to: use a preset voice synthesis model to obtain a text voice corresponding to the text to be synthesized, and determine the Mel spectrogram features of the text voice as the Mel spectrogram features of the text to be synthesized.

[0014] Optionally, the voice generation device further includes a training unit; the training unit is configured to obtain multiple groups of first voice samples, each group of first voice samples including a first voice and a second voice; the training unit is further configured to determine a first input sample for each group of first voice samples, the first input sample including the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice; the training unit is further configured to input the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice into a preset first neural network respectively, to obtain a first predicted identity feature of the first voice and a second predicted identity feature of the second voice; the training unit is further configured to, for each group of first voice samples, determine the identity feature difference degree between the first predicted identity feature and the second predicted identity feature, to obtain the identity feature difference degrees of multiple groups of first voice samples; the training unit is further configured to train the first neural network according to the identity feature difference degrees of multiple groups of first voice samples, to obtain an identity encoder.

[0015] Optionally, the training unit is specifically configured to: when the first voice and the second voice correspond to the same user, and when the identity feature difference degrees of multiple groups of first voice samples are all less than or equal to a first preset threshold, determine to obtain the identity encoder; when the first voice and the second voice correspond to different users, and when the identity feature difference degrees of multiple groups of first voice samples are all greater than or equal to a second preset threshold, determine to obtain the identity encoder; wherein, the second preset threshold is greater than the first preset threshold.

[0016] Optionally, the training unit is further configured to: obtain multiple groups of second voice samples, each group of second voice samples including a third voice and a fourth voice; determine a second input sample for each group of second voice samples, the second input sample including the Mel spectrogram features of the third voice and the Mel spectrogram features of the fourth voice; input the Mel spectrogram features of the third voice and the Mel spectrogram features of the fourth voice into a preset second neural network respectively, to obtain a first predicted content feature of the third voice and a second predicted content feature of the fourth voice; for each group of second voice samples, determine the content feature difference degree between the first predicted content feature and the second predicted content feature, to obtain the content feature difference degrees of multiple groups of second voice samples; train the second neural network according to the content feature difference degrees of multiple groups of second voice samples, to obtain a content encoder.

[0017] Optionally, the training unit is specifically configured to: when the third voice and the fourth voice correspond to the same text and the third voice and the fourth voice correspond to different users, if the content feature difference degrees of multiple groups of second voice samples are all less than or equal to a third preset threshold, then determine the content encoder; when the third voice and the fourth voice correspond to different texts and the third voice and the fourth voice correspond to different users, if the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to a fourth preset threshold, then determine the content encoder; wherein, the fourth preset threshold is greater than the third preset threshold.

[0018] Optionally, the training unit is further configured to: obtain multiple sample voices, and determine the sample Mel spectrogram features of each sample voice; input the sample Mel spectrogram features of each sample voice into the identity encoder to obtain the sample identity features corresponding to each sample voice; input the sample Mel spectrogram features of each sample voice into the content encoder to obtain the sample content features corresponding to each sample voice; input the sample identity features and the sample content features into a preset third neural network to obtain the predicted Mel spectrogram features of each sample voice; for each sample voice, determine the Mel spectrogram feature difference degree between the sample Mel spectrogram features and the predicted Mel spectrogram features to obtain the Mel spectrogram feature difference degrees of multiple sample voices; train the third neural network according to the Mel spectrogram feature difference degrees of multiple sample voices to obtain the decoder.

[0019] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor, and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the voice generation method of the first aspect above.

[0020] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which instructions are stored, and when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the voice generation method as described in the first aspect above.

[0021] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes computer instructions, and when the computer instructions are executed by a processor, the voice generation method as described in the first aspect above is implemented.

[0022] The technical solutions provided by the present disclosure at least bring the following beneficial effects: First, in the present disclosure, the speech generation device obtains the text to be synthesized and the target user's speech. Compared with the related art where a large number of users' speeches need to be obtained and trained based on them, the present disclosure only needs to obtain a small amount of speech (i.e., the target user's speech), and there is no need to train the target user's speech. Further, the speech generation device determines the Mel spectrogram features of the target user's speech, and inputs the Mel spectrogram features of the target user's speech into the identity encoder of the speech generation model that has been pre-trained to obtain the target identity features; the speech generation device determines the Mel spectrogram features of the text to be synthesized, and inputs the Mel spectrogram features of the text to be synthesized into the content encoder of the speech generation model to obtain the content features. Compared with the related art where one model can only generate speech with a fixed timbre or identity feature and cannot flexibly adjust the timbre or identity feature of the speech, and cannot generate speech that meets the personalized needs of users, the present disclosure can input the target identity features and content features into the decoder of the speech generation model according to the differences between the target speech and the text to be synthesized, and generate speech corresponding to the target identity features and content features, realizing flexible setting of the identity features and the text to meet the personalized needs of users.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.

[0025] Figure 1 is a schematic structural diagram of a speech generation system shown according to an exemplary embodiment;

[0026] Figure 2 is a schematic flowchart of a speech generation method shown according to an exemplary embodiment;

[0027] Figure 3 is a flowchart of the use of an identity encoder shown according to an exemplary embodiment;

[0028] Figure 4 is a flowchart of the use of a content encoder shown according to an exemplary embodiment;

[0029] Figure 5 is a flowchart of the use of a decoder shown according to an exemplary embodiment;

[0030] Figure 6 is a schematic flowchart of a speech generation method shown according to an exemplary embodiment;

[0031] Figure 7 It is the third flowchart diagram of a voice generation method shown according to an exemplary embodiment;

[0032] Figure 8 It is the fourth flowchart diagram of a voice generation method shown according to an exemplary embodiment;

[0033] Figure 9 It is the fifth flowchart diagram of a voice generation method shown according to an exemplary embodiment;

[0034] Figure 10 It is the sixth flowchart diagram of a voice generation method shown according to an exemplary embodiment;

[0035] Figure 11 It is the seventh flowchart diagram of a voice generation method shown according to an exemplary embodiment;

[0036] Figure 12 It is a flowchart diagram of the training process of a decoder shown according to an exemplary embodiment;

[0037] Figure 13 It is a structural diagram of a voice generation device shown according to an exemplary embodiment;

[0038] Figure 14 It is a structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0039] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order different from those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0041] In addition, in the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. The "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present disclosure, "a plurality of" means two or more than two.

[0042] It should be noted that the user information (including but not limited to user device information, user personal information, user behavior information, etc.) and data (including but not limited to program code, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0043] The voice generation method provided by the embodiments of the present disclosure can be applied to a voice generation system, which is used to solve the problem that in the related art, a voice that meets the personalized needs of users cannot be generated. Figure 1 A schematic structural diagram of the voice generation system is shown. As Figure 1 shown, the voice generation system 10 includes a voice generation device 11 and an electronic device 12. The voice generation device 11 is connected to the electronic device 12. The voice generation device 11 and the electronic device 12 can be connected in a wired manner or in a wireless manner. The embodiments of the present invention do not limit this.

[0044] The voice generation device 11 is used to obtain the text to be synthesized and the target user voice. The voice generation device 11 is also used to determine the Mel spectrum feature of the target user voice, and input the Mel spectrum feature of the target user voice into the identity encoder of the voice generation model that has been pre-trained to obtain the target identity feature. The voice generation device 11 is also used to determine the Mel spectrum feature of the text to be synthesized, and input the Mel spectrum feature of the text to be synthesized into the content encoder of the voice generation model to obtain the content feature. The voice generation device 11 is also used to input the target identity feature and the content feature into the decoder of the voice generation model to obtain the target voice; the target voice is the voice corresponding to the target identity feature and the content feature.

[0045] The voice generation device 11 can be implemented in various electronic devices 12 that can process voice data. The electronic device 12 at least has a sound collection device, a transmission device, and a voice playback device, such as a television, a smart phone, a portable terminal, a computer, a notebook computer, a tablet computer, and so on.

[0046] In different application scenarios, the voice generation device 11 and the electronic device 12 can be independent devices or integrated into the same device. The embodiments of the present invention do not specifically limit this.

[0047] When the voice generation device 11 and the electronic device 12 are integrated into the same device, the data transmission method between the voice generation device 11 and the electronic device 12 is the data transmission between internal modules of the device. In this case, the data transmission process between the two is the same as the "data transmission process between the voice generation device 11 and the electronic device 12 when they are independent of each other".

[0048] In the following embodiments provided by the embodiments of the present invention, the embodiments of the present invention will be described by taking the voice generation device 11 and the electronic device 12 being independently provided as an example.

[0049] Figure 2 It is a flowchart of a voice generation method shown according to some exemplary embodiments. In some embodiments, the above voice generation method can be applied to a voice generation device and an electronic device as shown in Figure 1 and can also be applied to other similar devices.

[0050] Such as Figure 2 shown, the voice generation method provided by the embodiments of the present invention includes the following S201-S206.

[0051] S201. The voice generation device acquires the text to be synthesized and the target user voice.

[0052] As a possible implementation manner, the voice generation device acquires the text to be synthesized and the target user voice from the electronic device.

[0053] It should be noted that both the text to be synthesized and the target user voice are obtained by the electronic device in response to the user's input operation. For example, the text to be synthesized can be the text input by the user into the electronic device according to the input operation, or the text selected by the user from multiple texts preset in the electronic device according to the input operation. The target user voice can be the voice input by the user into the electronic device according to the input operation, or the voice selected by the user from multiple voices preset in the electronic device according to the input operation. Among them, the form of the input operation can be text input, voice collection, target collection, etc., and the present disclosure embodiments do not limit the specific form of the input operation.

[0054] S202. The voice generation device determines the Mel spectrum feature of the target user voice.

[0055] As a possible implementation, the voice generation device performs analog-to-digital conversion (converting an analog signal to a digital signal) on the acquired target user voice to obtain target audio data. Further, the voice generation device performs a Fourier transform on the target audio data to obtain the target spectrum corresponding to the target user voice. The voice generation device inputs the target spectrum into a preset Mel filter function to obtain the Mel spectrum of the target user voice, and determines the Mel spectrum feature of the target user voice as this Mel spectrum.

[0056] It should be noted that the Mel filter function is preset by the operation and maintenance personnel in the voice generation device and is used to convert ordinary spectrum features into Mel-scale spectra (i.e., Mel spectra). The Mel spectrum is used to simulate the sensitivity of the human ear's hearing to actual frequencies, that is, the Mel spectrum is closer to the human ear's perception of the spectrum.

[0057] S203. The voice generation device inputs the Mel spectrum feature of the target user voice into the identity encoder of the pre-trained voice generation model to obtain the target identity feature.

[0058] As a possible implementation, the voice generation device inputs the Mel spectrum feature of the target user voice into the identity encoder of the pre-trained voice generation model and outputs the target identity feature.

[0059] It should be noted that the voice generation model includes an identity encoder, a content encoder, and a decoder, which are preset by the operation and maintenance personnel in the voice generation device. Among them, the identity encoder is used to analyze the input Mel spectrum feature and output the identity feature. The identity feature is essentially a voice feature used to reflect the speaker's identity. For example, the identity feature includes at least one of the frequency feature, amplitude feature, and timbre feature of the voice. The present disclosure does not make specific limitations on this.

[0060] Exemplarily, as Figure 3 shown, a flowchart of the use of an identity encoder is shown. Among them, the Mel spectrum feature of the target user voice is a, and the identity encoder is E id . The voice generation device inputs the Mel spectrum feature a of the target user voice into the identity encoder E id and outputs the target identity feature f id . Among them, f id = E id (a).

[0061] S204. The voice generation device determines the Mel spectrum feature of the text to be synthesized.

[0062] As a possible implementation, the speech generation device converts the text to be synthesized into text-to-speech and determines the Mel spectrogram of the text-to-speech. Further, the speech generation device determines the Mel spectrogram of the text-to-speech as the Mel spectrogram feature of the text to be synthesized. The specific implementation of determining the Mel spectrogram of the text-to-speech can refer to the above step S202, with the difference that the target user speech is replaced by the text-to-speech.

[0063] S205. The speech generation device inputs the Mel spectrogram feature of the text to be synthesized into the content encoder of the speech generation model to obtain a content feature.

[0064] As a possible implementation, the speech generation device inputs the Mel spectrogram feature of the text to be synthesized into the content encoder of the pre-trained speech generation model and outputs a content feature.

[0065] It should be noted that the content encoder is used to analyze the input Mel spectrogram feature and output a content feature. The content feature is used to reflect the content of the speech. For example, the content feature includes at least one of the language type, character length, and word meaning of the speech, and the present disclosure does not make specific limitations thereon.

[0066] Exemplarily, as Figure 4 shown, a flowchart of the use of a content encoder is shown, where the Mel spectrogram feature of the text to be synthesized is a', and the pre-trained content encoder is E con . The speech generation device inputs a' into E con and outputs the content feature f of the text to be synthesized con . Among them, f con = E con (a').

[0067] S206. The speech generation device inputs the target identity feature and the content feature into the decoder of the speech generation model to obtain a target speech.

[0068] Among them, the target speech is the speech corresponding to the target identity feature and the content feature.

[0069] As a possible implementation, after determining the target identity feature and the content feature, the speech generation device inputs the target identity feature and the content feature into the decoder and outputs the target speech corresponding to the target identity feature and the content feature.

[0070] Exemplarily, the decoder is D m , the speech feature is f id , the content feature is f con , and the speech generation device inputs f id and f con into D m to obtain the target speech

[0071] As another possible implementation, after determining the target identity feature and content feature, the voice generation device inputs the target identity feature and content feature into a decoder to output the Mel spectrum feature of the target voice. Further, the voice generation device inputs the Mel spectrum feature of the target voice into a vocoder to obtain the target voice.

[0072] It should be noted that the decoder is used to fuse the input identity feature and content feature and output a voice or the corresponding Mel spectrum feature.

[0073] The vocoder is pre-set in the voice generation device by the operation and maintenance personnel and can convert digital signals into analog signals. For example, it can convert the Mel spectrum feature into a voice.

[0074] Exemplarily, as Figure 5 shown, a flowchart of the use of a decoder is shown. Among them, the voice generation device inputs the identity feature and content feature into the decoder respectively to output the Mel spectrum feature. Further, the voice generation device inputs the Mel spectrum feature output by the decoder into the vocoder to generate a voice.

[0075] The technical solutions provided in the above embodiments at least bring the following beneficial effects: First, in the present disclosure, the voice generation device obtains the text to be synthesized and the target user voice. Compared with the related art that needs to obtain the voices of a large number of users and train according to the voices of a large number of users, the present disclosure only needs to obtain a small amount of voices (i.e., the target user voice) and does not need to train the target user voice. Further, the voice generation device determines the Mel spectrum feature of the target user voice and inputs the Mel spectrum feature of the target user voice into the identity encoder of the pre-trained voice generation model to obtain the target identity feature; the voice generation device determines the Mel spectrum feature of the text to be synthesized and inputs the Mel spectrum feature of the text to be synthesized into the content encoder of the voice generation model to obtain the content feature. Compared with the related art that a model can only generate a voice with a fixed timbre or identity feature and cannot flexibly adjust the timbre or identity feature of the voice and cannot generate a voice that meets the personalized needs of users, the present disclosure can input the target identity feature and content feature into the decoder of the voice generation model according to the differences between the target voice and the text to be synthesized, and generate a voice corresponding to the target identity feature and content feature, realizing the flexible setting of the identity feature and the text to meet the personalized needs of users.

[0076] In one design, in order to be able to determine the Mel spectrum feature of the text to be synthesized, as Figure 6 shown, the above S204 provided by the embodiments of the present disclosure specifically includes the following S2041-S2042:

[0077] S2041. The voice generation device uses a preset voice synthesis model to obtain the text voice corresponding to the text to be synthesized.

[0078] As a possible implementation, the voice generation device inputs the text to be synthesized into the preset voice synthesis model to obtain the text voice corresponding to the text to be synthesized.

[0079] It should be noted that the voice synthesis model is used to convert text into voice. Generally, a voice synthesis model can only generate a voice with a fixed timbre. The voice synthesis model can be any open-source voice synthesis model or the voice generation model in the embodiments of the present disclosure. The embodiments of the present disclosure do not limit the specific voice synthesis model.

[0080] Exemplarily, the voice generation device inputs the text t to be synthesized into the voice synthesis model to obtain the voice corresponding to the text t to be synthesized, and obtains the voice data a' corresponding to the voice.

[0081] S2042. The voice generation device determines the Mel spectrum feature of the text voice as the Mel spectrum feature of the text to be synthesized.

[0082] As a possible implementation, the voice generation device determines the Mel spectrum feature of the text voice and determines the Mel spectrum feature of the text voice as the Mel spectrum feature of the text to be synthesized. The specific implementation of the voice generation device for determining the Mel spectrum feature of the text voice can refer to S204 above and will not be elaborated here.

[0083] It can be understood that the embodiments of the present disclosure use a preset voice synthesis model to convert the text to be synthesized into text voice and determine the Mel spectrum feature of the text voice as the Mel spectrum feature of the text to be synthesized, unifying the data format of the text to be synthesized and making it easier to extract the content features of the text to be synthesized.

[0084] In one design, in order to obtain an identity encoder, as Figure 7 shown, the voice generation method provided by the embodiments of the present disclosure further includes the following S301 - S305 before the above S203:

[0085] S301. The voice generation device obtains multiple groups of first voice samples.

[0086] Each group of first voice samples includes a first voice and a second voice.

[0087] As a possible implementation, the voice generation device obtains multiple groups of first voice samples from the first dataset of the electronic device.

[0088] It should be noted that the first dataset is pre-stored in the electronic device by the operation and maintenance personnel. The first dataset includes multiple pre-collected voices. For example, the operation and maintenance personnel collect the voices of n users to obtain the first dataset S 1 , and store the first dataset S 1 in the electronic device.

[0089]

[0090] Among them, represents the k i th voice corresponding to the i-th user, that is, the same row in the dataset S 1 represents multiple voices of the same user.

[0091] The first voice and the second voice in the first voice sample pair are any two voices in the dataset S 1 .

[0092] In practical applications, the voice generation device can first collect the voice of any row from the dataset (multiple voices corresponding to the same user, such as the k i th voice data ) of the i-th user. Further, the voice generation device collects any two voices from the collected voice of any row to obtain a set of first voice samples, also called the first sample pair, and this first sample pair is a positive sample pair.

[0093] The voice generation device can also first collect the voices of any two rows from the dataset (multiple voices corresponding to different users, such as the k i th voice of the i-th user and the k w th voice of the w-th user). Further, the voice generation device collects one voice from each of the two collected rows of voices (for example ), to obtain a set of first voice samples, also called the first sample pair, and this first sample pair is a positive sample pair.

[0094] S302. The voice generation device determines the first input sample of each group of first voice samples.

[0095] Among them, the first input sample includes the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice.

[0096] As a possible implementation manner, the voice generation device determines the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice in each group of first voice samples, and uses the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice as the first input sample.

[0097] For the implementation of the voice generation device to specifically determine the Mel spectrum features of the first voice and the Mel spectrum features of the second voice, reference may be made to S202 above. The difference is that the target user voice can be replaced with the first voice or the second voice, and details will not be elaborated here.

[0098] S303. The voice generation device inputs the Mel spectrum features of the first voice and the Mel spectrum features of the second voice into a preset first neural network respectively, and obtains a first predicted identity feature of the first voice and a second predicted identity feature of the second voice.

[0099] As a possible implementation, the voice generation device inputs the Mel spectrum features of the first voice into a preset first neural network to obtain a first predicted identity feature of the first voice. Further, the voice generation device inputs the Mel spectrum features of the second voice into the preset first neural network to obtain a second predicted identity feature of the second voice.

[0100] It should be noted that the first neural network is pre-set in the voice generation device by the operation and maintenance personnel, and the first neural network can be a convolutional neural network.

[0101] Exemplarily, for the first voice sample The voice generation device inputs the Mel spectrum features of the first voice and the Mel spectrum features of the second voice into the convolutional neural network respectively, and obtains the predicted identity feature f 1 and the predicted identity feature f 2 .

[0102] Another exemplarily, for the first voice sample The voice generation device inputs the Mel spectrum features of the first voice and the Mel spectrum features of the second voice into the convolutional neural network respectively, and obtains the predicted identity feature f n1 and the predicted identity feature f n2 .

[0103] S304. For each group of first voice samples, the voice generation device determines the identity feature difference degree between the first predicted identity feature and the second predicted identity feature, and obtains the identity feature difference degrees of multiple groups of first voice samples.

[0104] As a possible implementation, for each group of first voice samples, the voice generation device calculates the identity feature difference degree between the first predicted identity feature and the second predicted identity feature according to a preset distance function, so as to obtain the identity feature difference degrees of multiple groups of first voice samples.

[0105] It should be noted that the distance function is pre-set in the voice generation device by the operation and maintenance personnel. The distance function can be the cosine distance or the Euclidean distance.

[0106] Exemplarily, D() is the preset distance function, f 1 is the first predicted identity feature, f 2 is the second predicted identity feature, then D(f 1 , f 2 ) is used to calculate the cosine distance or the Euclidean distance between f 1 and f 2 , and the voice generation device determines the calculation result as the identity feature difference degree between the first predicted identity feature and the second predicted identity feature.

[0107] S305. The voice generation device trains the first neural network according to the identity feature difference degrees of multiple groups of first voice samples to obtain an identity encoder.

[0108] As a possible implementation, for each group of first voice samples, the voice generation device determines the identity feature difference degree condition corresponding to the first voice sample, uses the identity feature difference degree condition as an expectation, and adjusts the parameters of the first neural network in combination with the identity feature difference degree of the first voice sample. Repeat the above actions to train the first neural network to obtain an identity encoder.

[0109] The technical solutions provided in the above embodiments at least bring the following beneficial effects: After the voice generation device obtains multiple groups of first voice samples including the first voice and the second voice, it inputs the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice into a preset first neural network respectively, and obtains the first predicted identity feature of the first voice and the second predicted identity feature of the second voice, so as to clarify the identity features corresponding to the two voices respectively. Further, the voice generation device determines the identity feature difference degrees of multiple groups of first voice samples, and trains the first neural network according to the identity feature difference degrees of multiple groups of first voice samples to obtain an identity encoder. In this way, in the subsequent process, the voice generation device can directly use the identity encoder to determine the identity feature of any voice.

[0110] In one design, in order to obtain an identity encoder, as Figure 8 shown, the above S305 provided by the embodiments of the present disclosure specifically includes the following S3051-S3055:

[0111] S3051. The voice generation device determines whether the first voice and the second voice correspond to the same user.

[0112] As a possible implementation, the voice generation device determines whether the first voice and the second voice correspond to the same user according to the user identifiers of the first voice and the second voice. When the user identifiers are the same, the voice generation device determines that the first voice and the second voice correspond to the same user; when the user identifiers are different, the voice generation device determines that the first voice and the second voice correspond to different users.

[0113] It should be noted that when the voice generation device obtains the first voice and the second voice from the first dataset, the voice of the same user has the same user identifier. For example, refer to S1 in the above step S301. Among them, represents the k i th voice corresponding to the i 1 th user, that is, the same row in the dataset S

[0114] In practical applications, the sample types of each group of first voice samples are usually the same, that is, usually all positive sample pairs (the first voice and the second voice correspond to the same user) or all negative sample pairs (the first voice and the second voice correspond to different users). The sample types of each group of first voice samples can also be different. Among multiple groups of first voice samples, there are both positive sample pairs and negative sample pairs. The embodiments of the present disclosure do not limit this. For the convenience of introduction, the embodiments of the present disclosure introduce with the sample types of each group of first voice samples being the same among multiple groups of first voice samples.

[0115] S3052. When the first voice and the second voice correspond to the same user, the voice generation device determines whether the identity feature difference degrees of multiple groups of first voice samples are all less than or equal to a first preset threshold.

[0116] As a possible implementation, when the first voice and the second voice correspond to the same user, the voice generation device compares the identity feature difference degrees of each group of first voice samples with the first preset threshold to determine whether the identity feature difference degrees of multiple groups of first voice samples are all less than or equal to the first preset threshold.

[0117] It should be noted that the first preset threshold is set in the voice generation device by the operation and maintenance personnel in advance. The first preset threshold should be set as small as possible.

[0118] S3053. When the identity feature difference degrees of multiple groups of first voice samples are all less than or equal to the first preset threshold, the voice generation device determines to obtain an identity encoder.

[0119] As a possible implementation, after the speech generation device trains the first neural network a number of times, if the identity feature difference degrees of multiple groups of first speech samples are all less than or equal to the first preset threshold, the identity encoder is determined. Otherwise, the speech generation device continues to train the first neural network (constantly adjusting the parameters) until the identity feature difference degrees of multiple groups of first speech samples are all less than or equal to the first preset threshold.

[0120] It can be understood that since multiple groups of first speech samples are all positive sample pairs, that is, the first speech and the second speech correspond to the same user, the identity feature difference degrees of multiple groups of first speech samples should be as small as possible to ensure the accuracy of the identity encoder.

[0121] S3054. When the first speech and the second speech correspond to different users, the speech generation device determines whether the identity feature difference degrees of multiple groups of first speech samples are all greater than the second preset threshold.

[0122] Among them, the second preset threshold is greater than the first preset threshold.

[0123] As a possible implementation, when the first speech and the second speech correspond to different users, the speech generation device compares the identity feature difference degrees of each group of first speech samples with the second preset threshold to determine whether the identity feature difference degrees of multiple groups of first speech samples are all greater than or equal to the second preset threshold.

[0124] It should be noted that the second preset threshold is set by the operation and maintenance personnel in the speech generation device in advance. The second preset threshold should be set as large as possible.

[0125] S3055. When the identity feature difference degrees of multiple groups of first speech samples are all greater than the second preset threshold, the speech generation device determines the identity encoder.

[0126] As a possible implementation, after the speech generation device trains the first neural network a number of times, if the identity feature difference degrees of multiple groups of first speech samples are all greater than the second preset threshold, the identity encoder is determined. Otherwise, the speech generation device continues to train the first neural network (constantly adjusting the parameters) until the identity feature difference degrees of multiple groups of first speech samples are all greater than the first preset threshold.

[0127] It can be understood that since multiple groups of first speech samples are all negative sample pairs, that is, the first speech and the second speech correspond to different users, the identity feature difference degrees of multiple groups of first speech samples should be as large as possible to ensure the accuracy of the identity encoder.

[0128] In some embodiments, the voice generation device may also train a first neural network based on the predicted identity features of the first voice, the predicted identity features of the second voice, and a first constraint condition to obtain an identity encoder.

[0129] Wherein, when the first voice and the second voice correspond to the same user, the first constraint condition includes: the difference degree between the predicted identity features of the first voice and the predicted identity features of the second voice data is less than a first preset threshold.

[0130] When the first voice and the second voice correspond to different users, the first constraint condition includes: the difference degree between the predicted identity features of the first voice and the predicted identity features of the second voice is greater than a second preset threshold.

[0131] As a possible implementation manner, the voice generation device uses the predicted identity features of the first voice and the predicted identity features of the second voice as sample features, and uses the first constraint condition as a label. When the predicted identity features of the first voice and the predicted identity features of the second voice satisfy the first constraint condition, the voice generation device trains to obtain an identity encoder. When the predicted identity features of the first voice and the predicted identity features of the second voice do not satisfy the first constraint condition, the voice generation device iteratively trains the first neural network with a new first voice sample until the predicted identity features of the first voice and the predicted identity features of the second voice satisfy the first constraint condition.

[0132] Exemplarily, when the first voice is and the second voice is , at this time the first voice and the second voice correspond to the same user (at this time the first voice samples corresponding to the first voice and the second voice are the first positive sample pair), then the first preset threshold can be set to minD(f 1 , f 2 ), where f 1 is 's predicted identity feature, f 2 is 's predicted identity feature, D() is a preset distance function used to calculate the cosine distance or Euclidean distance between f 1 and f 2 . When the distance between f 1 and f 2 is less than or equal to the first threshold, the voice generation device trains to obtain an identity encoder. When the distance between f 1 and f 2 is greater than the first threshold, the voice generation device iteratively trains the first neural network with a new first positive sample pair until the predicted f 1 and f2 until the distance therebetween is less than or equal to a first threshold value.

[0133] In another exemplary case, when the first voice is and the second voice is , at this time, the first voice and the second voice correspond to two different users (at this time, the first voice and the second voice correspond to the first voice sample as the first negative sample pair), then the second preset threshold value can be set as maxD(f n1 , f n2 ), where f n1 is 's predicted identity feature, f n2 is 's predicted identity feature, D() is a preset distance function for calculating the cosine distance or Euclidean distance between f n1 and f n2 . When the distance between f n1 and f n2 is greater than or equal to the second threshold value, the voice generation device trains to obtain an identity encoder. When the distance between f n1 and f n2 is less than the second threshold value, the voice generation device uses a new first negative sample pair to iteratively train the first neural network until the distance between the predicted f n1 and f n2 is greater than or equal to the second threshold value.

[0134] In one design, in order to obtain a content encoder, as Figure 9 shown, the voice generation method provided by the embodiments of the present disclosure further includes the following S401-S405 before S205 above:

[0135] S401. The voice generation device obtains multiple groups of second voice samples.

[0136] Each group of second voice samples includes a third voice and a fourth voice.

[0137] As a possible implementation manner, the voice generation device obtains multiple groups of second voice samples from the second data set of the electronic device.

[0138] It should be noted that the second data set is pre-stored in the electronic device by the operation and maintenance personnel, and the second data set includes multiple pre-collected voices. For example, the operation and maintenance personnel collect the voices of m users to obtain the second data set S 2 , and store the second data set S 2 in the electronic device.

[0139]

[0140] Among them, represents m voice data synthesized by the i-th open-source speech synthesis model, that is, the data set S 2 In the same row in represents the voice data synthesized by the same open-source speech synthesis model, and the same column represents the voice data synthesized by different open-source speech synthesis models according to the same text (text data set T).

[0141] The third voice and the fourth voice in the second voice sample are any two voice data in the data set S 2 among them.

[0142] In practical applications, the voice generation device can first collect any column of voices from the data set S 2 Among them, the voices in the same column are the voices synthesized by different open-source speech synthesis models for the same text, such as the d voices in the first column Furthermore, the voice generation device collects any two voices from the collected voices in any column (for example ), to obtain a second voice sample, and this second voice sample is a positive sample pair.

[0143] The voice generation device can also first collect any two columns of voices from the data set S 2 Among them, the voices in different columns are the voices synthesized by different open-source speech synthesis models for different texts, such as the d voices in the first column and the d voices in the second column Furthermore, the voice generation device collects one voice from each of the two columns of collected voices (for example ), to obtain a second voice sample, and this second voice sample is a negative sample pair.

[0144] S402. The voice generation device determines the second input sample of each group of second voice samples.

[0145] Among them, the second input sample includes the Mel spectrum features of the third voice and the Mel spectrum features of the fourth voice.

[0146] As a possible implementation manner, the voice generation device determines the Mel spectrum features of the third voice and the Mel spectrum features of the fourth voice in each group of second voice samples, and uses the Mel spectrum features of the third voice and the Mel spectrum features of the fourth voice as the second input sample.

[0147] The implementation manner for the voice generation device to specifically determine the Mel spectrum features of the third voice and the Mel spectrum features of the fourth voice can refer to the above S202, the difference is that the target user voice is replaced by the third voice or the fourth voice, which will not be elaborated here.

[0148] S403. The speech generation device inputs the Mel spectrum features of the third speech and the Mel spectrum features of the fourth speech into a preset second neural network respectively, and obtains the first predicted content feature of the third speech and the second predicted content feature of the fourth speech.

[0149] As a possible implementation manner, the speech generation device inputs the Mel spectrum features of the third speech into a preset second neural network, and obtains the first predicted content feature of the third speech. Further, the speech generation device inputs the Mel spectrum features of the fourth speech into a preset second neural network, and obtains the second predicted content feature of the fourth speech.

[0150] It should be noted that the second neural network is pre-set in the speech generation device by the operation and maintenance personnel, and the second neural network can be a convolutional neural network.

[0151] Exemplarily, for the second speech sample The speech generation device inputs the third speech data the fourth speech data into the convolutional neural network respectively, and obtains the predicted content feature p1 of and

[0152] Another exemplarily, the speech generation device collects a second speech sample from the second dataset inputs the third speech the fourth speech into the convolutional neural network respectively, and obtains the predicted content feature p n1 and the predicted content feature p n2 .

[0153] S404. For each group of second speech samples, the speech generation device determines the content feature difference degree between the first predicted content feature and the second predicted content feature, and obtains the content feature difference degrees of multiple groups of second speech samples.

[0154] As a possible implementation manner, for each group of second speech samples, the speech generation device calculates the content feature difference degree between the first predicted content feature and the second predicted content feature according to a preset distance function, so as to obtain the identity feature difference degrees of multiple groups of second speech samples.

[0155] It should be noted that the distance function is pre-set in the speech generation device by the operation and maintenance personnel. The distance function can be a cosine distance or an Euclidean distance.

[0156] Exemplarily, D() is a preset distance function, p1 is the first predicted content feature, and p2 is the second predicted content feature. Then D(p1, p2) is used to calculate the cosine distance or Euclidean distance between p1 and p2, and the speech generation device determines the content feature difference degree between the first predicted content feature and the second predicted content feature based on the calculation result.

[0157] S405. The speech generation device trains the second neural network according to the content feature difference degrees of multiple groups of second speech samples to obtain a content encoder.

[0158] As a possible implementation manner, for each group of second speech samples, the speech generation device determines the content feature difference degree condition corresponding to the second speech sample, uses the content feature difference degree condition as an expectation, and adjusts the parameters of the second neural network in combination with the content feature difference degree of the second speech sample. Repeat the above operations to train the second neural network to obtain a content encoder.

[0159] The technical solutions provided in the above embodiments at least bring the following beneficial effects: After the speech generation device obtains multiple groups of second speech samples including the third speech and the fourth speech, it inputs the Mel spectrum features of the third speech and the Mel spectrum features of the fourth speech into a preset second neural network respectively, to obtain the first predicted content feature of the third speech and the second predicted content feature of the fourth speech, so as to clarify the content features corresponding to the two speeches respectively. Further, the speech generation device determines the content feature difference degrees of multiple groups of second speech samples, and trains the second neural network according to the content feature difference degrees of multiple groups of second speech samples to obtain a content encoder. In this way, in the subsequent process, the speech generation device can directly use the content encoder to determine the content feature of any speech.

[0160] In one design, in order to obtain a content encoder, as Figure 10 shown, the above S405 provided in the embodiments of the present disclosure specifically includes the following S4051 - S4055:

[0161] S4051. The speech generation device determines whether the third speech and the fourth speech correspond to the same text.

[0162] As a possible implementation manner, the speech generation device determines whether the third speech and the fourth speech correspond to the same text according to the text identifiers of the third speech and the fourth speech. When the text identifiers are the same, the speech generation device determines that the third speech and the fourth speech correspond to the same text; when the text identifiers are different, the speech generation device determines that the third speech and the fourth speech correspond to different users.

[0163] It should be noted that when the voice generation device obtains the third voice and the fourth voice from the second dataset, the voice of the same text has the same text identifier. For example, refer to S2 in the above step S401. The same column represents the voices synthesized by different open-source voice synthesis models based on the same text, and the text identifiers of each voice are the same.

[0164] In practical applications, the sample types of each group of second voice samples are usually the same, that is, they are usually all positive sample pairs (the third voice and the fourth voice correspond to the same text) or all negative sample pairs (the third voice and the fourth voice correspond to different texts). The sample types of each group of second voice samples can also be different. Among multiple groups of second voice samples, there are both positive sample pairs and negative sample pairs. The embodiments of the present disclosure do not limit this. For the convenience of introduction, the embodiments of the present disclosure introduce with the sample types of each group of second voice samples being the same among multiple groups of second voice samples.

[0165] S4052. When the third voice and the fourth voice correspond to the same text, the voice generation device determines whether the identity feature difference degrees of multiple groups of second voice samples are all less than or equal to a third preset threshold.

[0166] As a possible implementation manner, when the third voice and the fourth voice correspond to the same text, the voice generation device compares the content feature difference degrees of each group of second voice samples with the third preset threshold to determine whether the content feature difference degrees of multiple groups of second voice samples are all less than or equal to the third preset threshold.

[0167] It should be noted that the third preset threshold is pre-set in the voice generation device by the operation and maintenance personnel. The third preset threshold should be set as small as possible.

[0168] S4053. When the content feature difference degrees of multiple groups of second voice samples are all less than or equal to the third preset threshold, the voice generation device determines to obtain the content encoder.

[0169] As a possible implementation manner, after the voice generation device trains the second neural network for several times, if the content feature difference degrees of multiple groups of second voice samples are all less than or equal to the third preset threshold, it determines to obtain the content encoder. Otherwise, the voice generation device continues to train the second neural network (constantly adjusting parameters) until the content feature difference degrees of multiple groups of second voice samples are all less than or equal to the third preset threshold.

[0170] It can be understood that since multiple groups of second voice samples are all positive sample pairs, that is, the third voice and the fourth voice correspond to the same text, the content feature difference degrees of multiple groups of second voice samples should be as small as possible to ensure the accuracy of the content encoder.

[0171] S4054. When the third voice and the fourth voice correspond to different texts, the voice generation device determines whether the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to a fourth preset threshold.

[0172] Among them, the fourth preset threshold is greater than the third preset threshold.

[0173] As a possible implementation manner, when the third voice and the fourth voice correspond to different texts, the voice generation device compares the content feature difference degrees of each group of second voice samples with the fourth preset threshold to determine whether the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to the fourth preset threshold.

[0174] It should be noted that the fourth preset threshold is set by the operation and maintenance personnel in the voice generation device in advance. The fourth preset threshold should be set as large as possible.

[0175] S4055. When the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to the fourth preset threshold, the voice generation device determines to obtain a content encoder.

[0176] As a possible implementation manner, after the voice generation device trains the second neural network for several times, if the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to the fourth preset threshold, it determines to obtain a content encoder. Otherwise, the voice generation device continues to train the second neural network (constantly adjusting parameters) until the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to the fourth preset threshold.

[0177] It can be understood that since multiple groups of second voice samples are all negative sample pairs, that is, the third voice and the fourth voice correspond to different texts, the content feature difference degrees of multiple groups of second voice samples should be as large as possible to ensure the accuracy of the content encoder.

[0178] In some embodiments, the voice generation device can also train the second neural network based on the predicted content features of the third voice, the predicted content features of the fourth voice, and the second constraint condition to obtain a content encoder.

[0179] Among them, when the third voice and the fourth voice correspond to the same text, the second constraint condition includes: the difference degree between the predicted content features of the third voice and the predicted content features of the fourth voice is less than or equal to the third preset threshold.

[0180] When the third voice and the fourth voice correspond to different texts, the second constraint condition includes: the difference degree between the predicted content features of the third voice and the predicted content features of the fourth voice is greater than or equal to the fourth preset threshold.

[0181] As a possible implementation, the speech generation device uses the predicted content features of the third speech and the predicted content features of the fourth speech as sample features, and uses the second constraint condition as a label. When the predicted content features of the third speech and the predicted content features of the fourth speech satisfy the second constraint condition, the speech generation device trains to obtain a content encoder. When the predicted content features of the third speech and the predicted content features of the fourth speech do not satisfy the second constraint condition, the speech generation device iteratively trains the second neural network using a new second speech sample until the predicted content features of the third speech and the predicted content features of the fourth speech satisfy the second constraint condition.

[0182] Exemplarily, when the third speech is and the fourth speech is , at this time the third speech and the fourth speech correspond to the same text (at this time the third speech and the fourth speech corresponding to the second speech sample are the second positive sample pair), then the third preset threshold can be set to minD(p 1 , p 2 ), where p 1 is the predicted content feature of , p 2 is the predicted content feature of , D() is a preset distance function used to calculate the cosine distance or Euclidean distance between p 1 and p 2 . When the distance between p 1 and p 2 is less than or equal to the third threshold, the speech generation device trains to obtain a content encoder. When the distance between p 1 and p 2 is greater than the third threshold, the speech generation device iteratively trains the second neural network using a new second positive sample pair until the predicted distance between p 1 and p 2 is less than or equal to the third threshold.

[0183] Another exemplarily, when the third speech is and the fourth speech is , at this time the third speech and the fourth speech correspond to different texts (at this time the third speech and the fourth speech corresponding to the second speech sample are the second negative sample pair), then the fourth preset threshold can be set to maxD(p n1 , p n2 ), where p n1 is the predicted content feature of , p n2 is the predicted content feature of , D() is a preset distance function used to calculate pn1 The cosine distance or Euclidean distance with p n2 When the distance between p n1 and p n2 is greater than or equal to the fourth threshold, the speech generation device trains to obtain a content encoder. When the distance between p n1 and p n2 is less than the fourth threshold, the speech generation device uses a new pair of second negative samples to iteratively train the second neural network until the predicted p n1 and p n2 The distance between them is greater than or equal to the fourth threshold.

[0184] In one design, in order to obtain a decoder, as Figure 11 shown, the speech generation method provided by the embodiments of the present disclosure further includes the following S501-S507 before S206 above:

[0185] S501. The speech generation device obtains a plurality of sample voices.

[0186] As a possible implementation manner, the speech generation device obtains sample voices from the sample data set of the electronic device.

[0187] It should be noted that the sample data set includes a plurality of voices. For example, the sample data set can be the first data set S 1 , or it can be the second data set S 2 , and the embodiments of the present disclosure do not limit this.

[0188] Exemplarily, the speech generation device obtains sample voices from the first data set S 1

[0189] S502. The speech generation device determines the sample Mel spectrum features of each sample voice.

[0190] The implementation manner for the speech generation device to specifically determine the sample Mel spectrum features of each sample voice can refer to S202 above, the difference being that the target user voice is replaced with the sample voice, which will not be elaborated here.

[0191] S503. The speech generation device inputs the sample Mel spectrum features of each sample voice into the identity encoder to obtain the sample identity features corresponding to each sample voice.

[0192] As a possible implementation manner, the speech generation device inputs the sample Mel spectrum features of each sample voice into the trained identity encoder respectively to obtain the sample identity features corresponding to each sample voice.

[0193] S504. The voice generation device inputs the sample Mel spectrogram features of each sample voice into the content encoder to obtain the sample content features corresponding to each sample voice.

[0194] As a possible implementation, the voice generation device inputs the sample Mel spectrogram features of each sample voice into the trained content encoder respectively to obtain the sample content features corresponding to each sample voice.

[0195] S505. The voice generation device inputs the sample identity features and the sample content features into a preset third neural network to obtain the predicted Mel spectrogram features of each sample voice.

[0196] It should be noted that the third neural network is pre-set in the voice generation device by the operation and maintenance personnel, and the third neural network can be a convolutional neural network.

[0197] S506. For each sample voice, the voice generation device determines the Mel spectrogram feature difference degree between the sample Mel spectrogram features and the predicted Mel spectrogram features to obtain the Mel spectrogram feature difference degrees of multiple sample voices.

[0198] As a possible implementation, for each sample voice, the voice generation device calculates the Mel spectrogram feature difference degree between the sample Mel spectrogram features and the predicted Mel spectrogram features according to a preset distance function, so as to obtain the Mel spectrogram feature difference degrees of multiple sample voices.

[0199] It should be noted that the distance function is pre-set in the voice generation device by the operation and maintenance personnel. The distance function can be a cosine distance or an Euclidean distance.

[0200] S507. The voice generation device trains the third neural network according to the Mel spectrogram feature difference degrees of multiple sample voices to obtain a decoder.

[0201] As a possible implementation, for each sample voice, the voice generation device takes the sample Mel spectrogram features as expectations and adjusts the parameters of the third neural network in combination with the predicted Mel spectrogram features. Repeat the above actions to train the third neural network to obtain a decoder.

[0202] In some embodiments, the voice generation device takes the sample Mel spectrogram features of the sample voice as labels. When the predicted Mel spectrogram features and the sample Mel spectrogram features satisfy the third constraint condition, the voice generation device trains to obtain a decoder. When the predicted Mel spectrogram features and the sample Mel spectrogram features do not satisfy the third constraint condition, the voice generation device uses new sample voices to perform iterative training on the third neural network until the predicted Mel spectrogram features and the sample Mel spectrogram features satisfy the third constraint condition.

[0203] The third constraint condition may be that the difference degree between the predicted Mel spectrum feature and the sample Mel spectrum feature is less than or equal to a fifth preset threshold value.

[0204] Exemplarily, the identity feature of sample speech 1 (the corresponding sample Mel spectrum feature is a) is f a , and the content feature is f c . The speech generation device inputs f a , f c into the third neural network and outputs If the distance D between and a is D = minD 2 it indicates that the third neural network is trained to obtain a decoder. Among them, D2() is a distance function used to calculate the cosine distance or Euclidean distance between and a.

[0205] As Figure 12 shown, a training flow chart of a decoder is shown. Among them, the speech generation device inputs the sample speech into the identity encoder and the content encoder respectively to obtain the identity feature and the content feature; furthermore, the speech generation device inputs the identity feature and the content feature into the third neural network to obtain a predicted Mel spectrum feature that satisfies the third constraint condition with the sample speech data.

[0206] The technical solutions provided in the above embodiments at least bring the following beneficial effects: Through the above training process, a decoder is obtained. In the subsequent process, the speech setting device can directly use the decoder to determine a speech corresponding to any identity feature and any content feature according to any identity feature and any content feature.

[0207] The above embodiments mainly introduce the solutions provided in the embodiments of the present disclosure from the perspective of the device (equipment). It can be understood that in order to implement the above methods, the device or equipment includes the corresponding hardware structure and / or software module for executing each method process, and these corresponding hardware structures and / or software modules for executing each method process can constitute a device for determining material information. Those skilled in the art should easily realize that, combined with the algorithm steps of each example described in the embodiments disclosed in this article, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the present disclosure.

[0208] In the embodiments of the present disclosure, the device or equipment can be divided into functional modules according to the above method examples. For example, the device or equipment can correspond to each function and divide each functional module, or integrate two or more functions into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present disclosure is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0209] Figure 13 is a schematic structural diagram of a voice generation device shown according to an exemplary embodiment. Refer to Figure 13 As shown, the voice generation device 60 provided in the embodiments of the present disclosure includes an acquisition unit 601, a determination unit 602, and a generation unit 603.

[0210] The acquisition unit 601 is configured to acquire the text to be synthesized and the target user voice; the determination unit 602 is configured to determine the Mel spectrum feature of the target user voice, and input the Mel spectrum feature of the target user voice into the identity encoder of the voice generation model pre-trained to obtain the target identity feature; the determination unit 602 is further configured to determine the Mel spectrum feature of the text to be synthesized, and input the Mel spectrum feature of the text to be synthesized into the content encoder of the voice generation model to obtain the content feature; the generation unit 603 is configured to input the target identity feature and the content feature into the decoder of the voice generation model to obtain the target voice; the target voice is the voice corresponding to the target identity feature and the content feature.

[0211] Optionally, the determination unit 602 is specifically configured to: use a preset voice synthesis model to obtain the text voice corresponding to the text to be synthesized, and determine the Mel spectrum feature of the text voice as the Mel spectrum feature of the text to be synthesized.

[0212] Optionally, the voice generation device 60 further includes a training unit 604; the training unit 604 is configured to obtain multiple groups of first voice samples, each group of first voice samples including a first voice and a second voice; the training unit 604 is further configured to determine a first input sample for each group of first voice samples, the first input sample including the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice; the training unit 604 is further configured to input the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice into a preset first neural network respectively, to obtain a first predicted identity feature of the first voice and a second predicted identity feature of the second voice; the training unit 604 is further configured to, for each group of first voice samples, determine the identity feature difference degree between the first predicted identity feature and the second predicted identity feature, to obtain the identity feature difference degrees of multiple groups of first voice samples; the training unit 604 is further configured to train the first neural network according to the identity feature difference degrees of multiple groups of first voice samples, to obtain an identity encoder.

[0213] Optionally, the training unit 604 is specifically configured to: when the first voice and the second voice correspond to the same user, and when the identity feature difference degrees of multiple groups of first voice samples are all less than or equal to a first preset threshold, determine to obtain the identity encoder; when the first voice and the second voice correspond to different users, and when the identity feature difference degrees of multiple groups of first voice samples are all greater than or equal to a second preset threshold, determine to obtain the identity encoder; wherein, the second preset threshold is greater than the first preset threshold.

[0214] Optionally, the training unit 604 is further configured to: obtain multiple groups of second voice samples, each group of second voice samples including a third voice and a fourth voice; determine a second input sample for each group of second voice samples, the second input sample including the Mel spectrogram features of the third voice and the Mel spectrogram features of the fourth voice; input the Mel spectrogram features of the third voice and the Mel spectrogram features of the fourth voice into a preset second neural network respectively, to obtain a first predicted content feature of the third voice and a second predicted content feature of the fourth voice; for each group of second voice samples, determine the content feature difference degree between the first predicted content feature and the second predicted content feature, to obtain the content feature difference degrees of multiple groups of second voice samples; train the second neural network according to the content feature difference degrees of multiple groups of second voice samples, to obtain a content encoder.

[0215] Optionally, the training unit 604 is specifically configured to: when the third voice and the fourth voice correspond to the same text and the third voice and the fourth voice correspond to different users, if the content feature difference degrees of multiple groups of second voice samples are all less than or equal to a third preset threshold, then determine to obtain a content encoder; when the third voice and the fourth voice correspond to different texts and the third voice and the fourth voice correspond to different users, if the content feature difference degrees of multiple groups of second voice samples are all greater than or equal to a fourth preset threshold, then determine to obtain a content encoder; wherein, the fourth preset threshold is greater than the third preset threshold.

[0216] Optionally, the training unit 604 is further configured to: obtain multiple sample voices, and determine the sample Mel spectrogram features of each sample voice; input the sample Mel spectrogram features of each sample voice into an identity encoder to obtain the sample identity features corresponding to each sample voice; input the sample Mel spectrogram features of each sample voice into a content encoder to obtain the sample content features corresponding to each sample voice; input the sample identity features and the sample content features into a preset third neural network to obtain the predicted Mel spectrogram features of each sample voice; for each sample voice, determine the Mel spectrogram feature difference degree between the sample Mel spectrogram features and the predicted Mel spectrogram features to obtain the Mel spectrogram feature difference degrees of multiple sample voices; train the third neural network according to the Mel spectrogram feature difference degrees of multiple sample voices to obtain a decoder.

[0217] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0218] Figure 14 is a schematic structural diagram of an electronic device provided by the present disclosure. As Figure 14 , the electronic device 70 may include at least one processor 701 and a memory 702 for storing instructions executable by the processor. Among them, the processor 701 is configured to execute the instructions in the memory 702 to implement the voice generation method in the above embodiments.

[0219] In addition, the electronic device 70 may further include a communication bus 703 and at least one communication interface 704.

[0220] The processor 701 may be a central processing unit (CPU), a microprocessing unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present disclosure solution.

[0221] The communication bus 703 may include a path for transmitting information between the above components.

[0222] A communication interface 704, using any transceiver-like device, is used to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0223] The memory 702 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited to this. The memory can exist independently and be connected to the processing unit through a bus. The memory can also be integrated with the processing unit.

[0224] Among them, the memory 702 is used to store the instructions for implementing the solution of the present disclosure and is controlled by the processor 701 to execute. The processor 701 is used to execute the instructions stored in the memory 702, thereby implementing the functions in the method of the present disclosure.

[0225] As an example, in combination with Figure 14 , the functions implemented by the acquisition unit 601, the determination unit 602, the generation unit 603, and the training unit 604 in the voice generation device 60 are the same as the functions of the processor 701 in Figure 14 .

[0226] In a specific implementation, as an embodiment, the processor 701 can include one or more CPUs, such as Figure 14 the CPU0 and CPU1 in

[0227] In a specific implementation, as an embodiment, the electronic device 70 can include multiple processors, such as Figure 14The processors 701 and 707 therein. Each of these processors can be a single-CPU processor or a multi-CPU processor. The processors here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0228] In a specific implementation, as an example, the electronic device 70 may further include an output device 705 and an input device 706. The output device 705 communicates with the processor 701 and can display information in various ways. For example, the output device 705 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 706 communicates with the processor 701 and can accept user input in various ways. For example, the input device 706 can be a mouse, a keyboard, a touch screen device, or a sensing device, etc.

[0229] Those skilled in the art can understand that Figure 14 the structure shown in does not constitute a limitation on the electronic device 70, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0230] In addition, the present disclosure also provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the voice generation method provided in the above embodiments.

[0231] In addition, the present disclosure also provides a computer program product, including computer instructions. When the computer instructions run on the electronic device, the electronic device executes the voice generation method provided in the above embodiments.

[0232] Those skilled in the art will readily think of other implementations of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. A voice generation method, characterized in that, comprising: obtaining the text to be synthesized and the target user's voice; determining the Mel spectrogram features of the target user's voice, and inputting the Mel spectrogram features of the target user's voice into the identity encoder of a pre-trained voice generation model to obtain target identity features; the identity encoder is trained as follows: obtaining multiple groups of first voice samples, each group of first voice samples including a first voice and a second voice; determining the first input samples of each group of first voice samples, the first input samples including the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice; respectively inputting the Mel spectrogram features of the first voice and the Mel spectrogram features of the second voice into a preset first neural network to obtain the first predicted identity features of the first voice and the second predicted identity features of the second voice; for each group of first voice samples, determining the identity feature difference degree between the first predicted identity features and the second predicted identity features to obtain the identity feature difference degrees of the multiple groups of first voice samples; training the first neural network according to the identity feature difference degrees of the multiple groups of first voice samples to obtain the identity encoder; determining the Mel spectrogram features of the text to be synthesized, and inputting the Mel spectrogram features of the text to be synthesized into the content encoder of the voice generation model to obtain content features; inputting the target identity features and the content features into the decoder of the voice generation model to obtain a target voice; the target voice is the voice corresponding to the target identity features and the content features.

2. The voice generation method according to claim 1, characterized in that, the determining the Mel spectrogram features of the text to be synthesized includes: using a preset voice synthesis model to obtain the text voice corresponding to the text to be synthesized, and determining the Mel spectrogram features of the text voice as the Mel spectrogram features of the text to be synthesized.

3. The voice generation method according to claim 1, characterized in that, the training the first neural network according to the identity feature difference degrees of the multiple groups of first voice samples to obtain the identity encoder includes: when the first voice and the second voice correspond to the same user, if the identity feature difference degrees of the multiple groups of first voice samples are all less than or equal to a first preset threshold, then the identity encoder is determined to be obtained; when the first voice and the second voice correspond to different users, if the identity feature difference degrees of the multiple groups of first voice samples are all greater than or equal to a second preset threshold, then the identity encoder is determined to be obtained; wherein, the second preset threshold is greater than the first preset threshold.

4. The voice generation method according to claim 1, characterized in that, the method further includes: obtaining multiple groups of second voice samples, each group of second voice samples including a third voice and a fourth voice; determining the second input samples of each group of second voice samples, the second input samples including the Mel spectrogram features of the third voice and the Mel spectrogram features of the fourth voice; Input the Mel spectrogram features of the third speech and the Mel spectrogram features of the fourth speech into a preset second neural network respectively, to obtain the first predicted content feature of the third speech and the second predicted content feature of the fourth speech; For each group of second speech samples, determine the content feature difference degree between the first predicted content feature and the second predicted content feature, to obtain the content feature difference degrees of the multiple groups of second speech samples; Train the second neural network according to the content feature difference degrees of the multiple groups of second speech samples, to obtain the content encoder.

5. The speech generation method according to claim 4, wherein, the training the second neural network according to the content feature difference degrees of the multiple groups of second speech samples to obtain the content encoder includes: in the case that the third speech and the fourth speech correspond to the same text and the third speech and the fourth speech correspond to different users, when the content feature difference degrees of the multiple groups of second speech samples are all less than or equal to a third preset threshold, then determine to obtain the content encoder; in the case that the third speech and the fourth speech correspond to different texts and the third speech and the fourth speech correspond to different users, when the content feature difference degrees of the multiple groups of second speech samples are all greater than or equal to a fourth preset threshold, then determine to obtain the content encoder; wherein, the fourth preset threshold is greater than the third preset threshold.

6. The speech generation method according to claim 1, wherein, the method further includes: acquire a plurality of sample speeches, and determine the sample Mel spectrogram features of each sample speech; input the sample Mel spectrogram features of each sample speech into the identity encoder, to obtain the sample identity feature corresponding to each sample speech; input the sample Mel spectrogram features of each sample speech into the content encoder, to obtain the sample content feature corresponding to each sample speech; input the sample identity feature and the sample content feature into a preset third neural network, to obtain the predicted Mel spectrogram feature of each sample speech; for each sample speech, determine the Mel spectrogram feature difference degree between the sample Mel spectrogram feature and the predicted Mel spectrogram feature, to obtain the Mel spectrogram feature difference degrees of the plurality of sample speeches; train the third neural network according to the Mel spectrogram feature difference degrees of the plurality of sample speeches, to obtain the decoder.

7. A speech generation device, wherein, it includes an acquisition unit, a determination unit, a generation unit and a training unit; the acquisition unit is configured to acquire a text to be synthesized and a target user speech; the determination unit is configured to determine the Mel spectrogram features of the target user speech, and input the Mel spectrogram features of the target user speech into the identity encoder of a pre-trained speech generation model, to obtain a target identity feature; the training unit is configured to acquire multiple groups of first speech samples, and each group of first speech samples includes a first speech and a second speech; The training unit is further configured to determine a first input sample for each group of first speech samples, where the first input sample includes the Mel spectrogram features of the first speech and the Mel spectrogram features of the second speech; The training unit is further configured to input the Mel spectrogram features of the first speech and the Mel spectrogram features of the second speech into a preset first neural network respectively, to obtain a first predicted identity feature of the first speech and a second predicted identity feature of the second speech; The training unit is further configured to, for each group of first speech samples, determine an identity feature difference degree between the first predicted identity feature and the second predicted identity feature, to obtain the identity feature difference degrees of the multiple groups of first speech samples; The training unit is further configured to train the first neural network according to the identity feature difference degrees of the multiple groups of first speech samples, to obtain the identity encoder; The determining unit is further configured to determine the Mel spectrogram features of the text to be synthesized, and input the Mel spectrogram features of the text to be synthesized into the content encoder of the speech generation model, to obtain content features; The generating unit is configured to input the target identity feature and the content features into the decoder of the speech generation model, to obtain a target speech; the target speech is a speech corresponding to the target identity feature and the content features.

8. The speech generation device according to claim 7, wherein, The determining unit is specifically configured to: Obtain the text speech corresponding to the text to be synthesized by using a preset speech synthesis model, and determine the Mel spectrogram features of the text speech as the Mel spectrogram features of the text to be synthesized.

9. The speech generation device according to claim 7, wherein, The training unit is specifically configured to: In the case where the first speech and the second speech correspond to the same user, when the identity feature difference degrees of the multiple groups of first speech samples are all less than or equal to a first preset threshold, then determine to obtain the identity encoder; In the case where the first speech and the second speech correspond to different users, when the identity feature difference degrees of the multiple groups of first speech samples are all greater than or equal to a second preset threshold, then determine to obtain the identity encoder; wherein, the second preset threshold is greater than the first preset threshold.

10. The speech generation device according to claim 7, wherein, The training unit is further configured to: Obtain multiple groups of second speech samples, each group of second speech samples including a third speech and a fourth speech; Determine a second input sample for each group of second speech samples, where the second input sample includes the Mel spectrogram features of the third speech and the Mel spectrogram features of the fourth speech; Input the Mel spectrogram features of the third speech and the Mel spectrogram features of the fourth speech into a preset second neural network respectively, to obtain a first predicted content feature of the third speech and a second predicted content feature of the fourth speech; For each group of the second speech samples, determine the content feature difference degree between the first predicted content feature and the second predicted content feature, and obtain the content feature difference degrees of the multiple groups of the second speech samples; Train the second neural network according to the content feature difference degrees of the multiple groups of the second speech samples to obtain the content encoder.

11. The speech generation device according to claim 10, wherein, the training unit is specifically configured to: in the case that the third speech and the fourth speech correspond to the same text and the third speech and the fourth speech correspond to different users, when the content feature difference degrees of the multiple groups of the second speech samples are all less than or equal to a third preset threshold, determine to obtain the content encoder; in the case that the third speech and the fourth speech correspond to different texts and the third speech and the fourth speech correspond to different users, when the content feature difference degrees of the multiple groups of the second speech samples are all greater than or equal to a fourth preset threshold, determine to obtain the content encoder; wherein, the fourth preset threshold is greater than the third preset threshold.

12. The speech generation device according to claim 7, wherein, the training unit is further configured to: obtain a plurality of sample speeches, and determine the sample Mel spectrogram features of each sample speech; input the sample Mel spectrogram features of each sample speech into the identity encoder to obtain the sample identity features corresponding to each sample speech; input the sample Mel spectrogram features of each sample speech into the content encoder to obtain the sample content features corresponding to each sample speech; input the sample identity features and the sample content features into a preset third neural network to obtain the predicted Mel spectrogram features of each sample speech; for each sample speech, determine the Mel spectrogram feature difference degree between the sample Mel spectrogram feature and the predicted Mel spectrogram feature, and obtain the Mel spectrogram feature difference degrees of the plurality of sample speeches; train the third neural network according to the Mel spectrogram feature difference degrees of the plurality of sample speeches to obtain the decoder.

13. An electronic device, wherein, comprising: a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the speech generation method according to any one of claims 1-6.

14. A computer-readable storage medium, on which instructions are stored, wherein, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the speech generation method according to any one of claims 1-6.

15. A computer program product, wherein, the computer program product includes computer instructions, and when the computer instructions are executed by a processor, the speech generation method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Voice real-time cloning method and device based on small sample, equipment and medium

    CN111681635A