Voice cloning method, device, training method, electronic device and storage medium
Through the method of decoupling and reverse model coupling of multi-layer neural network models, the imitation effect of speech cloning technology is improved, the problems of weak decoupling capabilities and dependence on data diversity in the existing technology are solved, and more efficient speech cloning is achieved.
Patent Information
- Application Number
- CN202111676414.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The existing voice cloning technology has weak decoupling ability to treat cloned speaker characteristics, resulting in poor imitation effect and high quantity and diversity of relying on training data, which increases development costs and low efficiency.
The multi-layer neural network model is used to decouple the features of cloned speech, obtaining multiple, multi-grained and multi-level speaker characteristics, and coupling the speaker characteristics with the text content characteristics through the reverse model to generate cloned speech.
It improves the imitation effect of speech cloning, reduces dependence on the quantity and diversity of training data, reduces development costs and improves efficiency.
Smart Images

Figure CN114333847B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a voice cloning method, apparatus, training method, electronic device, and storage medium. Background Art
[0002] Voice cloning technology is a technology that uses a reference voice signal to synthesize any text, but the target voice signal with speaker characteristics such as timbre, rhythm, style, etc. similar to the reference voice signal. It can meet the needs of personalized customization of voice or speaking style and is applied to various mobile phone assistants, e-books, intelligent telephone customer service, audio and video dubbing, intelligent interactive robots, etc. Benefiting from the rapid development of deep learning technology, neural network-based speech synthesis technology has achieved great success, and its synthesized speech has approached the effect of real human voice quality, making it difficult to distinguish between true and false. However, with the rapid increase in the demand for personalized customization of speech synthesis, the traditional method of collecting a large amount of training data and separately modeling a certain voice will not only increase the development cost but also reduce the development efficiency. With the open source and sharing of more and more multi-speaker and multi-voice style data, relying on the principle of transfer learning in deep learning, fine-tuning and transferring the target voice or style to the average model trained on this data has achieved good results, which will significantly reduce the company's development cost and improve efficiency.
[0003] However, the inventors of the present invention have found that in the existing voice cloning technology, due to the weak decoupling ability of the features of the speaker to be cloned, its imitation ability is poor, the imitation effect is poor, and there is a disadvantage of high dependence on the quantity and diversity of training data. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a voice cloning method, apparatus, and model training method for a voice cloning apparatus, which can improve the imitation effect of voice cloning.
[0005] To solve the above technical problems, an embodiment of the present invention provides a voice cloning method, including the following steps: using a first neural network model to decouple the features of the voice to be cloned to obtain the speaker features of the voice to be cloned, where the speaker features are the features unrelated to the text content in the voice to be cloned, and the first neural network model is a multi-layer neural network model; encoding the text to be synthesized to obtain the text content features of the text to be synthesized; using a second neural network model to couple the speaker features of the voice to be cloned and the text content features of the text to be synthesized to generate a cloned voice.
[0006] Embodiments of the present invention also provide a voice cloning device, including: a content encoder, which is used to encode the text to be synthesized and output the text content features of the text to be synthesized; a spectrogram encoder, which is used to decouple the features of the voice to be cloned to obtain the speaker features of the voice to be cloned, where the speaker features are features unrelated to the text content in the voice to be cloned, and the first neural network model running in the spectrogram encoder is a multi-layer neural network model; a spectrogram decoder, which is used to couple the text content features of the text to be synthesized and the speaker features of the voice to be cloned to generate a cloned voice.
[0007] Embodiments of the present invention also provide a model training method for a voice cloning device, including: obtaining a plurality of sample voices and sample texts corresponding to each of the sample voices; and training the model of the aforementioned voice cloning device according to the sample voices and the sample texts.
[0008] Embodiments of the present invention also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can execute the aforementioned voice cloning method or the model training method of the aforementioned voice cloning device.
[0009] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the aforementioned voice cloning method or the model training method of the aforementioned voice cloning device.
[0010] Compared with the prior art, in the embodiments of the present invention, the first neural network model used for decoupling the voice to be cloned is a multi-layer neural network model. Therefore, the speaker features decoupled by the first neural network model are multiple, multi-granularity, and multi-level speaker features, so that the speaker features can better represent the speaker features and improve the cloning effect of voice cloning.
[0011] In addition, using the first neural network model to decouple the features of the voice to be cloned to obtain the speaker features of the voice to be cloned includes: using each network layer of the first neural network model to perform encoding operations on the voice to be cloned, taking the hidden variables obtained by the operations of each network layer as the speaker features of the voice to be cloned, and taking the encoding result output by the first neural network model as the text content features of the voice to be cloned.
[0012] In addition, the coupling of the speaker feature of the speech to be cloned and the text content feature of the text to be synthesized by using the second neural network model includes: using the inverse model of the first neural network model to couple the speaker feature of the speech to be cloned and the text content feature of the text to be synthesized. The first neural network model for decoupling the speech to be cloned is the inverse model of the second neural network model for splicing to form the cloned speech, that is, the network layer structures of the first neural network model and the second neural network model are the same. Therefore, when decoupling the speech to be cloned, the system parameters generated in the first neural network model can all be applied in the second neural network model when synthesizing the cloned speech, reducing the loss of parameters. And these parameters generally include speaker features such as the speaking style and speaking voice information of the speaker in the speech to be cloned. Reducing the loss of parameters can make the simulation effect of the subsequently synthesized cloned speech better, thus improving the imitation effect of speech cloning.
[0013] In addition, the coupling of the speaker feature of the speech to be cloned and the text content feature of the text to be synthesized by using the second neural network model includes: respectively inputting the latent variables obtained by the operations of the respective network layers into the respective network layers identical to the respective network layers of the second neural network model, and coupling the text content feature of the text to be synthesized according to the latent variables of the respective network layers.
[0014] In addition, the second neural network model running in the spectrogram decoder is the inverse model of the first neural network model.
[0015] In addition, input the sample text into the content encoder to obtain the text content features of the sample text; input the sample speech into the spectrogram encoder to obtain the speaker features of the sample speech and the text content features of the sample speech, where the speaker features are the features unrelated to the text content in the sample speech; input the speaker features of the sample speech and the text content features of the sample text into the spectrogram decoder to obtain the cloned speech; establish a first loss function between the cloned speech and the sample speech, and establish a second loss function between the text content features of the sample speech and the text content features of the sample text; perform model training on the spectrogram encoder and the spectrogram decoder according to the multiple sample speeches and the sample texts corresponding to each of the sample speeches until both the first loss function and the second loss function converge. During the process of model training for the voice cloning device, while establishing the first loss function between the cloned speech and the sample speech and establishing the second loss function between the text content features of the sample speech and the text content features of the sample text, when performing model training on the spectrogram encoder and the spectrogram decoder, the training ends only when both the first loss function and the second loss function converge, which can make the text content features of the sample speech decoupled by the spectrogram encoder closer to the text content features of the sample text obtained by the content encoder. Since the neural network models running in the spectrogram encoder and the spectrogram decoder are the same, the closer the text content features of the sample speech decoupled by the spectrogram encoder are to the text content features of the sample text obtained by the content encoder, the better the imitation effect of the cloned speech spliced by the spectrogram decoder, thereby further improving the voice cloning effect of the trained voice cloning device. Description of the Drawings
[0016] Figure 1 is a schematic flowchart of the voice cloning method provided by the first embodiment of the present invention;
[0017] Figure 2 is a schematic flowchart of the voice cloning method provided by the second embodiment of the present invention;
[0018] Figure 3 is a schematic structural diagram of the voice cloning device provided by the third embodiment of the present invention;
[0019] Figure 4 is a schematic structural diagram of the voice cloning device provided by the fourth embodiment of the present invention;
[0020] Figure 5 is a schematic flowchart of the model training method of the voice cloning device provided by the fifth embodiment of the present invention;
[0021] Figure 6It is a schematic flowchart of the model training method of the voice cloning device provided by the sixth embodiment of the present invention;
[0022] Figure 7 It is a schematic structural diagram of the electronic device provided by the seventh embodiment of the present invention. Specific embodiments
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following will elaborate on each embodiment of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present invention, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0024] The first embodiment of the present invention relates to a voice cloning method, and the specific process is as Figure 1 shown, including:
[0025] Step S101: Encode the text to be synthesized to obtain the text content features of the text to be synthesized.
[0026] Specifically, in this step, the text to be synthesized is encoded to obtain an encoded sequence containing a context relationship. Specifically, a rule-based encoding method expressing the context relationship can be used, such as the method of the speech synthesis system based on the hidden Markov model used in traditional speech synthesis systems. The calculated encoded sequence is the text content feature of the text to be synthesized.
[0027] Step S102: Use the first neural network model to decouple the features of the voice to be cloned to obtain the speaker features of the voice to be cloned.
[0028] Specifically, in this step, the voice to be cloned is a segment of speech spoken by the speaker to be cloned. In actual use, the voice to be cloned can be the voice of the speaker collected in real time, or the recording of the speaker to be cloned stored in advance, and can be flexibly used according to actual needs.
[0029] The voice to be cloned contains the text content spoken by the speaker and the text content features corresponding to the text content, that is, the pronunciation and pauses of the text content that conform to human language habits described in the previous step. This part of the text content features is the common features of the speaker in the current language environment and is replaceable; in addition, the voice to be cloned also includes speaker features such as the speaker's intonation, timbre, and speaking style, and this part of the speaker features is the speaker's own timbre, speaking style and other features and is not replaceable. Use the first neural network model to decouple the voice to be cloned, that is, separate the speaker features belonging to the speaker in the voice to be cloned and the text content features related to the speaking content in the voice to be cloned, and extract the speaker features such as the unique intonation, timbre, and speaking style of the speaker.
[0030] Specifically, in this step, use the first neural network model to encode the voice to be cloned, and use the encoding result as the text content features of the voice to be cloned, and use the hidden variables of each network layer in the first neural network model (that is, the mean and variance calculated in each network layer) as the speaker features of the voice to be cloned. Among them, in this embodiment, the first neural network model can be composed of a one-dimensional residual CNN (Convolutional Neural Networks) and an Instance Normalization module. Among them, the speaker features of the voice to be cloned are calculated by the Instance Normalization module of each network layer in the "first neural network model", that is, the speaker features of the voice to be cloned are the hidden variables in the Instance Normalization module of the network layer.
[0031] Step S103: Use the second neural network model to couple the speaker features of the voice to be cloned and the text content features of the text to be synthesized to form a cloned voice.
[0032] Specifically, in this step, use the inverse model of the first neural network model to couple the speaker features of the voice to be cloned and the text content features of the text to be synthesized. That is to say, the second neural network model is the inverse model of the first neural network model, that is, the model structure of the second neural network model is the same as that of the first neural network model and the operation direction is opposite. For example, the second neural network model sequentially includes a total of N network layers from the first to the Nth, where the first network layer is the input network layer and the Nth network layer is the output network layer. Then the first neural network model also includes a total of N network layers from the first to the Nth, and the structures of each network layer are the same. The difference is that the Nth network layer in the first neural network model is the input network layer and the first network layer is the output network layer.
[0033] Specifically, in this step, the second neural network model is used to couple the speaker features of the speech to be cloned and the text content features of the text to be synthesized. Since the speaker features of the speech to be cloned are decoupled from the speech to be cloned by the first neural network model, and the second neural network model is the same as the first neural network model but only has the opposite operation direction, when using the second neural network model to splice the speaker features of the speech to be cloned and the text content features of the text to be synthesized, the loss of the speaker features of the speech to be cloned is small, and the cloned speech obtained by splicing has better reducibility of the speaker features such as the intonation, timbre, and speaking style of the speaker of the speech to be cloned.
[0034] Specifically, in this embodiment, the second neural network model can also be composed of a one-dimensional residual CNN (Convolutional Neural Networks) and an Instance Normalization module. In step S102, the speaker features of the speech to be cloned are calculated by the Instance Normalization module of each network layer in the "first neural network model", that is, the mean and variance of the latent variables in the Instance Normalization module of the network layer are used as the speaker features of the speech to be cloned; in this step, the mean and variance calculated in step S102 can be respectively input into the respective network layers in the second neural network model that are the same as the network layers where the mean and variance are calculated. For example, in step S102, the first mean and the first variance are calculated by the first network layer, and the second mean and the second variance are calculated by the second network layer. Then, in this step, the first mean and the first variance are input into the network layer that is the same as the first network layer, and the second mean and the second variance are input into the network layer that is the same as the second network layer. Inputting the mean and variance calculated by the first neural network model into the respective network layers in the second neural network model that are the same as the network layers where the mean and variance are calculated can reduce the loss of the speaker features of the speech to be cloned during the calculation process and improve the cloning effect.
[0035] Compared with the prior art, in the speech cloning method provided by the first embodiment of the present invention, the first neural network model used to decouple the speech to be cloned is a multi-layer neural network model. Therefore, the speaker features decoupled by the first neural network model are multiple, multi-granularity, and multi-level speaker features, so that the speaker features can better represent the speaker features and improve the cloning effect of speech cloning.
[0036] The second embodiment of the present invention relates to a speech cloning method. The second embodiment is substantially the same as the first embodiment, and the specific steps are as Figure 2 shown, including:
[0037] Step S201: Perform noise reduction on the voice to be cloned.
[0038] Specifically, in this step, environmental noise or interference noise caused during storage or transmission may exist in the voice to be cloned. Before cloning the voice to be cloned, noise reduction is performed on the voice to be cloned, that is, the noise in the voice to be cloned is reduced through a noise reduction algorithm.
[0039] Step S202: Perform format conversion on the text to be synthesized, and convert the text to be synthesized into a preset standard format.
[0040] Specifically, in this step, for the input text information, operations such as "illegal symbol filtering", "special symbol conversion", "number to Chinese character", "text regularization", and "pinyin prediction" are required to convert it into a "consonant and vowel" sequence in general writing habits. At the same time, relevant "prosody" marking information such as word segmentation boundaries, pause boundaries, stress, and emotion can also be added.
[0041] Taking Chinese as an example, the pinyin identifier of each character in the text to be synthesized is obtained in sequence, and the pinyin identifiers of all characters in the text to be synthesized are summarized as the text content feature of the text to be synthesized. Among them, the pinyin identifier includes at least one of the initial consonant, final vowel, and tone of each character. Specifically, if the flat tone, rising tone, falling tone, entering tone, and light tone in the tone are represented by the numbers 1, 2, 3, 4, and 5 respectively, then for the text "The Double Ninth Festival is an important festival" to be cloned, the pinyin identifier of the first "chong" obtained is "chong2", and the pinyin identifier of the second "chong" obtained is "zhong4". By analogy, the pinyin identifiers of each character in the text to be cloned are obtained and summarized to obtain the corresponding text "chong2 yang2 jie2 shi4 yi2 ge4 zhong4 yao4 de5 jie2 ri4" to be cloned in pinyin.
[0042] It can be understood that the foregoing is only an illustrative example in this embodiment taking Chinese characters as an example, and does not constitute a limitation. In actual use, the input text encoding type is not restricted in any way. Chinese can also directly model Chinese characters, or syllable pinyin, etc.; for an English voice cloning system, phonetic symbols or English letters can also be selected as the standard form of text input.
[0043] Step S203: Encode the text to be synthesized to obtain the text content feature of the text to be synthesized.
[0044] Step S204: Use the first neural network model to decouple the features of the voice to be cloned to obtain the speaker feature of the voice to be cloned.
[0045] Step S205: Use the second neural network model to couple the speaker features of the speech to be cloned and the text content features of the text to be synthesized to form cloned speech.
[0046] It can be understood that steps S203 to S205 in this embodiment are substantially the same as steps S101 to S103 in the first embodiment, and will not be elaborated here. For specific details, reference can be made to the specific description of the first embodiment.
[0047] Compared with the prior art, in the speech cloning method provided by the second embodiment of the present invention, before performing speech cloning on the speech to be cloned, noise reduction processing is performed on the speech to be cloned, which can reduce the influence of noise on the cloning effect of the speech to be cloned and further improve the cloning effect; in this embodiment, before performing speech cloning on the speech to be cloned, format conversion is also performed on the text to be synthesized, and the text to be synthesized is converted into a preset standard format, which can speed up the speech cloning speed and improve the cloning efficiency. In addition, this embodiment also has the technical effects of the first embodiment and will not be elaborated here.
[0048] It can be understood that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this patent; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process are all within the protection scope of this patent.
[0049] The third embodiment of the present invention relates to a speech cloning device, as Figure 3 shown, including: a content encoder 10, the content encoder 10 is used to encode the text to be synthesized and output the text content features of the text to be synthesized; a spectrogram encoder 20, the spectrogram encoder 20 is used to decouple the features of the speech to be cloned to obtain the speaker features of the speech to be cloned, and the speaker features are the features irrelevant to the text content in the speech to be cloned. The first neural network model running in the spectrogram encoder is a multi-layer neural network model; a spectrogram decoder 30, the spectrogram decoder 30 is used to couple the text content features of the text to be synthesized and the speaker features of the speech to be cloned to generate cloned speech.
[0050] It is not difficult to find that this embodiment is a device embodiment corresponding to the first embodiment, and this embodiment can be implemented in cooperation with the first embodiment. The relevant technical details and technical effects mentioned in the first embodiment are still valid in this embodiment and will not be elaborated here to avoid repetition. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.
[0051] In an implementation of the present invention, the neural network model running in the spectrogram encoder 20 is the inverse model of the neural network model running in the spectrogram decoder 30. The neural network model running in the spectrogram encoder 20 being the inverse model of the neural network model running in the spectrogram decoder 30 means that the model structures of the neural network models running in the spectrogram encoder 20 and the spectrogram decoder 30 are the same. Therefore, when decoupling the voice to be cloned, all the system parameters generated in the spectrogram encoder 20 can be applied in the spectrogram decoder 30 when synthesizing the cloned voice, reducing the loss of parameters. And these parameters generally include speaker characteristics such as the speaking style and speaking voice information of the speaker in the voice to be cloned. Reducing the loss of parameters can make the simulation effect of the subsequently synthesized cloned voice better, thereby improving the imitation effect of voice cloning.
[0052] It is worth mentioning that each module involved in this implementation is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present invention, units not closely related to solving the technical problems proposed by the present invention are not introduced in this implementation, but this does not mean that there are no other units in this implementation.
[0053] The fourth implementation of the present invention relates to a voice cloning device. The fourth implementation is substantially the same as the third implementation and also includes a content encoder 10, a spectrogram encoder 20, and a spectrogram decoder 30, specifically as Figure 4 shown. The voice cloning device further includes: a noise suppressor 40 connected to the spectrogram encoder 20. The noise suppressor is used to receive the target audio and perform noise suppression on the target audio to remove the noise in the target audio to obtain the voice to be cloned. In addition, the voice cloning device further includes: a format converter 50 connected to the content encoder 10; the format converter 50 is used to perform format conversion on the text to be synthesized and convert the text to be synthesized into a standard format that the content encoder 10 can recognize.
[0054] Since the second implementation corresponds to this implementation, this implementation can be implemented in cooperation with the second implementation. The relevant technical details and technical effects mentioned in the second implementation are still valid in this implementation, and the technical effects achievable in the second implementation can also be achieved in this implementation. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this implementation can also be applied in the second implementation.
[0055] The fifth embodiment of the present invention relates to a model training method for a voice cloning device, which is applied to model training of the voice cloning device provided in the foregoing embodiment. The specific steps are as follows Figure 5 shown and include:
[0056] Step S301: Obtain a plurality of sample voices and sample texts corresponding to each sample voice.
[0057] Specifically, in this step, the content of the sample voice and the sample text is the same, that is, the text content spoken by the speaker in the sample voice is the same as the text content in the sample text.
[0058] Step S302: Perform model training on the voice cloning device according to the sample voice and the sample text.
[0059] Specifically, in this step, a first loss function between the cloned voice and the sample voice is established in advance, and a second loss function between the text content features of the sample voice and the text content features of the sample text is established. Among them, the sample text is input into the content encoder, and the output of the content encoder is the text content feature of the sample text; the sample voice is input into the spectrogram encoder, and after the spectrogram encoder decouples the sample voice, the speaker feature of the sample voice and the text content feature of the sample voice can be obtained. The speaker feature is the feature unrelated to the text content in the sample voice; the speaker feature of the sample voice and the text content feature of the sample text are input into the spectrogram decoder, and after the spectrogram decoder couples the speaker feature of the sample voice and the text content feature of the sample text, the cloned voice can be obtained.
[0060] It can be understood that step S202 in this embodiment is substantially the same as steps S101 to S103 in the first embodiment. For specific details, reference can be made to the specific description of the foregoing embodiment, and details will not be repeated here.
[0061] Step S303: Use a plurality of sample voices and sample texts corresponding to each sample voice to perform model training on the spectrogram encoder and the spectrogram decoder until both the first loss function and the second loss function converge.
[0062] Compared with the prior art, before training the spectrogram encoder and the spectrogram decoder using multiple sample voices and the sample texts corresponding to the respective sample voices, a first loss function is established between the cloned voice and the sample voice, and a second loss function is established between the text content features of the sample voice and the text content features of the sample text. When training the spectrogram encoder and the spectrogram decoder, the training is completed only when both the first loss function and the second loss function converge. This can make the text content features of the sample voice decoupled by the spectrogram encoder closer to the text content features of the sample text obtained by the content encoder. Since the neural network models running in the spectrogram encoder and the neural network models running in the spectrogram decoder are the same, the closer the text content features of the sample voice decoupled by the spectrogram encoder are to the text content features of the sample text obtained by the content encoder, the better the imitation effect of the cloned voice spliced by the spectrogram decoder, thereby further improving the voice cloning effect of the trained voice cloning device.
[0063] The sixth embodiment of the present invention relates to a model training method for a voice cloning device, and the specific steps are as Figure 6 shown, including:
[0064] Step S401: Train the model of the content encoder.
[0065] Specifically, in this step, the model of the content encoder can be trained by a multi-speaker, multi-style corpus speech synthesis system.
[0066] Step S402: Obtain multiple sample voices and the sample texts corresponding to the respective sample voices.
[0067] Specifically, in this step, the content of the sample voice and the sample text is the same, that is, the text content spoken by the speaker in the sample voice is the same as the text content in the sample text.
[0068] Step S403: Train the model of the voice cloning device according to the sample voice and the sample text.
[0069] Step S404: Use multiple sample voices and the sample texts corresponding to the respective sample voices to train the spectrogram encoder and the spectrogram decoder until both the first loss function and the second loss function converge.
[0070] It can be understood that steps S402 to S404 in this embodiment are substantially the same as steps S301 to S303 in the fifth embodiment. For specific details, reference can be made to the specific description of the foregoing embodiment, and details will not be repeated here.
[0071] Compared with the prior art, in the model training method of the voice cloning device provided by the sixth embodiment of the present invention, before training the model of the spectrogram encoder and the spectrogram decoder, the content encoder is pre-trained. After obtaining the trained content encoder, during the training of the spectrogram encoder and the spectrogram decoder, the content encoder is no longer trained, that is, during the training of the spectrogram encoder and the spectrogram decoder, the model parameters of the content encoder are locked and no longer changed. Thereby reducing the influence of the text content on the voice cloning result during the training process, and further reducing the influence of the text content on the cloning result during the voice cloning process of the voice cloning device, and improving the cloning effect of the voice cloning device after training.
[0072] The seventh embodiment of the present invention relates to an electronic device, such as Figure 7 shown, including: at least one processor 701; and a memory 702 communicatively connected to the at least one processor 701; wherein, the memory 702 stores instructions executable by the at least one processor 701, and the instructions are executed by the at least one processor 701 to enable the at least one processor 701 to execute the voice cloning method provided by the foregoing embodiment or the model training method of the voice cloning device provided by the foregoing embodiment.
[0073] Wherein, the memory 702 and the processor 701 are connected by a bus. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 701 and the memory 702 together. The bus may also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be one element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 701 is transmitted on the wireless medium through the antenna. Further, the antenna also receives the data and transmits the data to the processor 701.
[0074] The processor 701 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. And the memory 702 can be used to store the data used by the processor 701 when executing operations.
[0075] The eighth embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.
[0076] That is, those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
Claims
1. A voice cloning method, characterized in that, it includes: using a first neural network model to decouple the features of the voice to be cloned, obtaining the speaker features of the voice to be cloned, where the speaker features are the features irrelevant to the text content in the voice to be cloned, and the first neural network model is a multi-layer neural network model; encoding the text to be synthesized to obtain the text content features of the text to be synthesized; using a second neural network model to couple the speaker features of the voice to be cloned and the text content features of the text to be synthesized to generate a cloned voice; wherein, the second neural network model is an inverse model with the same model structure as the first neural network model and the opposite operation direction.
2. The voice cloning method according to claim 1, characterized in that, the step of using a first neural network model to decouple the features of the voice to be cloned and obtain the speaker features of the voice to be cloned includes: using each network layer of the first neural network model to perform encoding operations on the voice to be cloned, taking the hidden variables obtained by the operations of each network layer as the speaker features of the voice to be cloned, and taking the encoding result output by the first neural network model as the text content features of the voice to be cloned.
3. The voice cloning method according to claim 2, characterized in that, the step of using a second neural network model to couple the speaker features of the voice to be cloned and the text content features of the text to be synthesized includes: using the inverse model of the first neural network model to couple the speaker features of the voice to be cloned and the text content features of the text to be synthesized.
4. The voice cloning method according to claim 3, characterized in that, the step of using a second neural network model to couple the speaker features of the voice to be cloned and the text content features of the text to be synthesized includes: respectively inputting the hidden variables obtained by the operations of each network layer into the same network layers as each network layer in the second neural network model, and coupling the text content features of the text to be synthesized according to the hidden variables of each network layer.
5. A voice cloning device, characterized in that, it includes: a content encoder for encoding the text to be synthesized and outputting the text content features of the text to be synthesized; a spectrogram encoder for decoupling the features of the voice to be cloned to obtain the speaker features of the voice to be cloned, where the speaker features are the features irrelevant to the text content in the voice to be cloned, and the first neural network model running in the spectrogram encoder is a multi-layer neural network model; a spectrogram decoder for coupling the text content features of the text to be synthesized and the speaker features of the voice to be cloned to generate a cloned voice; the second neural network model running in the spectrogram decoder is an inverse model with the same model structure as the first neural network model and the opposite operation direction.
6. A model training method for a voice cloning device, characterized in that, it includes: Obtain multiple sample voices and sample texts corresponding to each of the sample voices; Perform model training on the voice cloning device provided in claim 5 according to the sample voices and the sample texts.
7. The model training method of the voice cloning device according to claim 6, characterized in that, the performing model training on the voice cloning device provided in claim 5 according to the sample voices and the sample texts includes: Input the sample text into the content encoder to obtain the text content features of the sample text; Input the sample voice into the spectrogram encoder to obtain the speaker features of the sample voice and the text content features of the sample voice, where the speaker features are the features unrelated to the text content in the sample voice; Input the speaker features of the sample voice and the text content features of the sample text into the spectrogram decoder to obtain a cloned voice; Establish a first loss function between the cloned voice and the sample voice, and establish a second loss function between the text content features of the sample voice and the text content features of the sample text; Perform model training on the spectrogram encoder and the spectrogram decoder according to the multiple sample voices and the sample texts corresponding to each of the sample voices until both the first loss function and the second loss function converge.
8. An electronic device, characterized in that, it includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the voice cloning method according to any one of claims 1 to 4 or the model training method of the voice cloning device according to any one of claims 6 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the voice cloning method according to any one of claims 1 to 4 or the model training method of the voice cloning device according to any one of claims 6 to 7.