Voice cloning method, training method, device and medium

By training voiceprint features to build an acoustic model, the problems of high audio data volume and high device performance requirements in the existing voice cloning technology are solved, and efficient voice cloning is achieved, suitable for clients and servers.

CN114049873BActive Publication Date: 2025-07-08BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD

Patent Information

Application Number
CN202111277694.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-07-08
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

The existing voice cloning technology has high requirements for the amount of audio data of cloned objects, long adaptive training time, and high equipment performance requirements, which limits its usage range and processing efficiency.

Method used

By training voiceprint features, acoustic models are built, and the mapping relationship between voiceprint features and acoustic features is used to reduce the audio data volume requirement, simplify the adaptive training process, and is suitable for clients and servers.

Benefits of technology

It reduces the audio data volume requirement of cloned objects, improves processing efficiency, expands the scope of application, improves user experience, and reduces device performance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114049873B_ABST
    Figure CN114049873B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a voice cloning method, a training method, a device and a medium. The voice cloning method specifically includes: receiving text and the original audio of a cloning object; determining the voiceprint feature corresponding to the original audio; inputting the text and the voiceprint feature into an acoustic model to obtain a corresponding acoustic feature, where the acoustic model is obtained according to the voiceprint features corresponding to training samples; and determining a corresponding target audio according to the acoustic feature. The embodiment of the present invention can reduce the audio data volume of the cloning object, and can improve the processing efficiency and application scope of voice cloning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of speech processing, and particularly to a speech cloning method, a training method, a device, and a medium. Background Art

[0002] Speech cloning technology refers to using a small amount of audio of a cloning object to complete the cloning of the voice of the cloning object. Generally, speech cloning technology can generate a target audio that approximates the voice of the cloning object according to any input text.

[0003] Traditional speech cloning methods usually include: First, training a speech cloning model for multiple people; Second, collecting the audio of the cloning object; performing a series of operations such as noise reduction, feature extraction, and duration segmentation on the audio of the cloning object to obtain corresponding processing results; then using the above processing results to perform adaptive training on the speech cloning model for multiple people to adjust the speech cloning model for multiple people and obtain a speech cloning model of the cloning object, and this speech cloning model of the cloning object is used to clone the voice of the cloning object.

[0004] In practical applications, the above-mentioned adaptive training has certain requirements for the audio data volume of the cloning object, usually requiring the audio of the cloning object to be dozens to hundreds of sentences, which increases the difficulty of obtaining the audio of the cloning object. Moreover, adaptive training requires additional training time, which affects the processing efficiency. In addition, adaptive training has certain requirements for device performance, which affects the scope of use of the speech cloning method. For example, currently, the speech cloning method can only be applied to the server side. Summary of the Invention

[0005] How to reduce the audio data volume of the cloning object and how to improve the processing efficiency and scope of use of speech cloning are technical problems that need to be solved by those skilled in the art. In view of the above problems, embodiments of the present invention are proposed to provide a speech cloning method, a device, and a medium that overcome the above problems or at least partially solve the above problems.

[0006] To solve the above problems, the present invention discloses a training method, including:

[0007] Determine the voiceprint feature corresponding to the training sample;

[0008] Train an acoustic model according to the voiceprint feature corresponding to the training sample.

[0009] To solve the above problems, the present invention discloses a speech cloning method, including:

[0010] Receive the text and the original audio of the cloning object;

[0011] Determine the voiceprint feature corresponding to the original audio;

[0012] Input the text and the voiceprint feature into an acoustic model to obtain corresponding acoustic features; wherein, the acoustic model is obtained based on the voiceprint features corresponding to training samples.

[0013] Determine a corresponding target audio according to the acoustic features.

[0014] On the other hand, an embodiment of the present invention discloses a training device, including:

[0015] A voiceprint determination module, configured to determine the voiceprint features corresponding to training samples;

[0016] An acoustic training module, configured to train an acoustic model according to the voiceprint features corresponding to the training samples.

[0017] On the other hand, an embodiment of the present invention discloses a voice cloning device, including:

[0018] A receiving module, configured to receive text and the original audio of a cloning object;

[0019] A voiceprint determination module, configured to determine the voiceprint features corresponding to the original audio;

[0020] An acoustic determination module, configured to input the text and the voiceprint features into an acoustic model to obtain corresponding acoustic features; wherein, the acoustic model is obtained based on the voiceprint features corresponding to training samples.

[0021] An audio determination module, configured to determine a corresponding target audio according to the acoustic features.

[0022] On yet another aspect, an embodiment of the present invention discloses a device for training a voice cloning model, including a memory, and one or more programs, wherein the one or more programs are stored in the memory, and when the one or more programs are executed by one or more processors, the steps of the foregoing method are implemented.

[0023] On yet another aspect, an embodiment of the present invention discloses a device for voice cloning, including a memory, and one or more programs, wherein the one or more programs are stored in the memory, and when the one or more programs are executed by one or more processors, the steps of the foregoing method are implemented.

[0024] An embodiment of the present invention also discloses one or more machine-readable media, characterized in that instructions are stored thereon, and when executed by one or more processors, the device is caused to execute the foregoing method.

[0025] An embodiment of the present invention also discloses a computer program product, which includes computer instructions stored in a computer-readable storage medium and adapted to be read and executed by a processor, so that a computer device having the processor executes the foregoing method.

[0026] The embodiments of the present invention have the following advantages:

[0027] In the embodiments of the present invention, an acoustic model is trained according to the voiceprint features corresponding to the training samples. Among them, the acoustic model can represent the mapping relationship between the input (such as text and voiceprint features) and the output (acoustic features), and can obtain an output that matches the input. Since the input of the acoustic model includes voiceprint features, the embodiments of the present invention can obtain acoustic features that match the voiceprint features based on the acoustic model.

[0028] In the voice cloning process of the embodiments of the present invention, the original audio of the cloning object is used to determine the voiceprint features. Since the determination of voiceprint features has a low requirement for the amount of audio data, the embodiments of the present invention can reduce the requirement for the original audio of the cloning object, that is, can reduce the amount of audio data of the cloning object. In practical applications, the cloning object can record one or more sentences to achieve voice cloning, so the user experience can be improved.

[0029] Moreover, the principle of the voice cloning process in the embodiments of the present invention is specifically as follows: Voice cloning is achieved according to the voiceprint features corresponding to the original audio and the mapping relationship between the input (including voiceprint features) and the acoustic features represented by the acoustic model. Since the adaptive training corresponding to the original audio can be saved, the voice cloning process in the embodiments of the present invention can save the time of adaptive training and improve the processing efficiency; moreover, the voice cloning process in the embodiments of the present invention can reduce the requirement for device performance, and can be applicable to both the server side and the client side, so the applicable range can be increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flowchart of the steps of a training method according to an embodiment of the present invention;

[0031] Figure 2 is a schematic diagram of the training process of an acoustic model according to an embodiment of the present invention;

[0032] Figure 3 is a schematic structural diagram of a voice cloning model according to an embodiment of the present invention;

[0033] Figure 4 is a flowchart of the steps of a voice cloning method according to an embodiment of the present invention;

[0034] Figure 5It is a schematic diagram of the usage process of an acoustic model according to an embodiment of the present invention;

[0035] Figure 6 It is a block diagram of the structure of a voice cloning device according to an embodiment of the present invention;

[0036] Figure 7 It is a block diagram of the structure of a training device according to an embodiment of the present invention;

[0037] Figure 8 It is a block diagram of a device 1300 for voice cloning according to an embodiment of the present invention; and

[0038] Figure 9 It is a schematic diagram of the structure of a server according to an embodiment of the present invention. Detailed implementation manners

[0039] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0040] The embodiments of the present invention can be applied to the voice cloning scenario. The voice cloning scenario can be used to generate a target audio that approximates the voice of the cloning object according to any input text and a small amount of audio of the cloning object.

[0041] In view of the technical problems of how to reduce the audio data volume of the cloning object, how to improve the processing efficiency and application scope of voice cloning, the embodiments of the present invention provide a voice cloning method, which specifically includes: receiving text and the original audio of the cloning object; determining the voiceprint feature corresponding to the original audio; inputting the text and the voiceprint feature into an acoustic model to obtain the corresponding acoustic feature; wherein, the acoustic model can be obtained according to the voiceprint feature corresponding to the training sample; and determining the corresponding target audio according to the acoustic feature.

[0042] The embodiments of the present invention use voiceprint features in the training process and voice cloning process of the acoustic model. A voiceprint is the sound wave spectrum carrying speech information displayed by electroacoustic instruments. Modern scientific research shows that a voiceprint not only has specificity but also has the characteristic of relative stability, so it can represent the identity of a user.

[0043] The embodiments of the present invention train an acoustic model according to the voiceprint feature corresponding to the training sample. Among them, the acoustic model can represent the mapping relationship between the input (such as text and voiceprint feature) and the output (acoustic feature), and can obtain an output that matches the input. Since the input of the acoustic model includes the voiceprint feature, the embodiments of the present invention can obtain an acoustic feature that matches the voiceprint feature based on the acoustic model.

[0044] During the voice cloning process of the embodiments of the present invention, the original audio of the cloning object is used to determine the voiceprint features. Since the determination of voiceprint features has relatively low requirements for the amount of audio data, the embodiments of the present invention can reduce the requirements for the original audio of the cloning object, that is, can reduce the amount of audio data of the cloning object. In practical applications, the cloning object can record one or more sentences to achieve voice cloning, thus improving the user experience.

[0045] Moreover, the principle of the voice cloning process of the embodiments of the present invention is specifically as follows: Voice cloning is achieved based on the voiceprint features corresponding to the original audio and the mapping relationship between the input (including voiceprint features) characterized by the acoustic model and the acoustic features. Since the adaptive training corresponding to the original audio can be saved, the voice cloning process of the embodiments of the present invention can save the time of adaptive training, improve the processing efficiency; and the voice cloning process of the embodiments of the present invention can reduce the requirements for device performance, and can be applied to both the server side and the client side, so the applicable range can be increased.

[0046] The voice cloning method provided by the embodiments of the present invention can be applied to the application environment corresponding to the client and the server. The client and the server are located in a wired or wireless network, and through this wired or wireless network, the client and the server perform data interaction.

[0047] Optionally, the client can run on a terminal. The above terminal specifically includes but is not limited to: smart phones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, in-vehicle computers, desktop computers, set-top boxes, smart televisions, wearable devices, and so on.

[0048] The client can correspond to a website or an APP (Application). For example, the client can correspond to application programs such as a voice processing APP and a voice cloning APP.

[0049] In the training stage, the server can execute the training method so that the trained acoustic model can characterize the mapping relationship between the input (text, voiceprint features, etc.) and the output (acoustic features).

[0050] In the voice cloning stage, the server can receive the text and the original audio sent by the client, use the voice cloning method to obtain the target audio, and return the target audio to the client. Alternatively, the client can obtain the trained acoustic model from the server, use the voice cloning method to obtain the target audio, and generalize the target audio to the user.

[0051] Method Embodiment 1

[0052] This embodiment describes the process of determining the voiceprint feature.

[0053] In this embodiment, according to a preset voice library, a voiceprint model is trained; the voiceprint model is used to obtain the voiceprint feature based on the acoustic feature.

[0054] In a specific implementation, the preset voice library can be the first voice library. The first voice library can include voice samples of N (N can be a natural number) speakers, so that the trained voiceprint model has the ability to extract voiceprint features. The specific value of N in the embodiments of the present invention is not limited. For example, the order of magnitude of N can be hundreds, thousands, ten thousands, etc.

[0055] The environments of the voice samples in the first voice library can be various. For example, the environments of the voice samples in the voice library can include: indoor environment and outdoor environment. Another example is that the environments of the voice samples in the voice library can include: noisy environment and noise-free environment, etc. The voice samples in multiple environments can improve the robustness of the voiceprint model.

[0056] In an alternative embodiment of the present invention, a mathematical model can be trained based on the voice samples to obtain a voiceprint model, and the voiceprint model can represent the mapping relationship between the input data (the first acoustic feature) and the output data (the voiceprint feature).

[0057] A mathematical model is a scientific or engineering model constructed using mathematical logic methods and mathematical language. A mathematical model is a mathematical structure that generally or approximately expresses, in mathematical language, the characteristics or quantitative dependency relationships of a certain thing system. This mathematical structure is a relational structure depicted by means of mathematical symbols. A mathematical model can be one or a set of algebraic equations, differential equations, difference equations, integral equations, or statistical equations and their combinations, which quantitatively or qualitatively describe the mutual relationships or causal relationships among the variables of the system. In addition to the mathematical models described by equations, there are also models described by other mathematical tools, such as algebra, geometry, topology, mathematical logic, etc. Among them, a mathematical model describes the behavior and characteristics of a system rather than the actual structure of the system. Among them, methods such as machine learning and deep learning methods can be used to train the mathematical model. Machine learning methods can include: linear regression, decision tree, random forest, etc. Deep learning methods can include: Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), etc.

[0058] As the input of the voiceprint model, the first acoustic feature can be an acoustic feature that has not been processed by the acoustic model; there can be a weak correlation between the first acoustic feature and the speaker's voiceprint feature. The acoustic feature output by the acoustic model is an acoustic feature that can reflect the speaker's voiceprint feature and is strongly correlated with the speaker's voiceprint feature. For the convenience of distinction, the acoustic feature output by the acoustic model is denoted as the second acoustic feature. The correlation between the first acoustic feature and the speaker's voiceprint feature can be weaker than the correlation between the second acoustic feature and the speaker's voiceprint feature.

[0059] In practical applications, examples of the first acoustic feature can include: Linear Prediction Cepstral Coefficients (LPCC), Mel Frequency Cepstrum Coefficients (MFCC), etc.

[0060] After the voiceprint model training is completed, the trained voiceprint model can be used to determine the voiceprint feature corresponding to any audio. Taking the determination of the voiceprint feature corresponding to the original audio as an example, the first acoustic feature of the original audio can be extracted and input into the voiceprint model to obtain the voiceprint feature output by the voiceprint model.

[0061] Method Embodiment Two

[0062] This embodiment describes the training process of the acoustic model.

[0063] Reference Figure 1 , a step flowchart of a training method according to an embodiment of the present invention is shown. The method may specifically include the following steps:

[0064] Step 101, determine the voiceprint feature corresponding to the training sample;

[0065] Step 102, train an acoustic model according to the voiceprint feature corresponding to the training sample.

[0066] Figure 1 The method embodiment shown can be executed by the server. It can be understood that the specific execution entity of the method embodiment in the present invention is not limited.

[0067] In step 101, the training sample may be from the first speech database, or may be from a second speech database different from the first speech database. The order of magnitude of the training samples in the second speech database can also be hundreds, thousands, tens of thousands, etc. Optionally, the training samples in the second speech database can be speech in a noise-free environment to improve the performance of the acoustic model.

[0068] In specific implementation, the first acoustic feature of the training sample can be extracted, and the first acoustic feature of the training sample is input into the voiceprint model to obtain the voiceprint feature corresponding to the training sample output by the voiceprint model.

[0069] In step 102, the mathematical model can be trained based on the training sample to obtain an acoustic model. The acoustic model can represent the mapping relationship between the input (such as text and voiceprint feature) and the output (acoustic feature).

[0070] During the training process, in addition to using the text and voiceprint feature as the input of the acoustic model, the actual duration feature, the first acoustic feature, the actual prosody feature, etc. can also be used as the input of the acoustic model.

[0071] The text can be the text corresponding to the speech represented by the training sample. The text corresponding to the training sample can be determined based on speech recognition technology or manual annotation technology.

[0072] The duration feature can be used to represent the duration of the phoneme corresponding to the text. The duration feature can depict the cadence, stress and rhythm in the speech. The actual duration feature can be used to determine the error of the acoustic model in duration prediction.

[0073] The prosody feature can include information such as emotion, speech rate, speech quality level, etc., and can make the speech more natural and emotional. The actual prosody feature can be used to determine the error of the acoustic model in prosody prediction. In practical applications, vae (Variational Auto-Encoder) can be used to determine the corresponding actual prosody feature according to the text and voiceprint feature.

[0074] The first acoustic feature can be used to determine the prediction of the acoustic model in acoustic prediction.

[0075] In one implementation of the present invention, the acoustic model may specifically include: a duration prediction module, a prosody prediction module, and an acoustic prediction module.

[0076] Among them, the duration prediction model is used to predict the duration feature corresponding to the text and the voiceprint feature. The input of the duration prediction model may include: the text and the voiceprint feature, and the output may include: the predicted duration feature.

[0077] The prosody prediction module is used to predict the prosody feature corresponding to the text and the voiceprint feature. The input of the prosody prediction module may include: the text and the voiceprint feature, and the output may include: the predicted prosody feature.

[0078] The acoustic prediction module is used to predict the acoustic feature corresponding to the text, the voiceprint feature, the duration feature, and the prosody feature. The input of the acoustic prediction module may include: the text, the voiceprint feature, the duration feature, and the prosody feature, and the output may include: the second acoustic feature.

[0079] In another implementation of the present invention, in addition to including: a duration prediction module, a prosody prediction module, and an acoustic prediction module, the acoustic model may further include: a prosody extraction module, which is used to extract the actual prosody feature corresponding to the training sample. It can be understood that it is feasible to set the prosody extraction module inside or outside the acoustic model, and the embodiments of the present invention do not limit the specific position of the prosody extraction module.

[0080] During the process of training the acoustic model, the error can be determined, so that during the backpropagation process, the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module are updated according to the error, so that the training of the acoustic model converges, and further the trained acoustic model can represent the mapping relationship between the input (such as text and voiceprint feature) and the output (acoustic feature). The training convergence condition of the acoustic model may include: the error is less than a preset error value, etc. It can be understood that the embodiments of the present invention do not limit the specific training convergence condition.

[0081] In one implementation, the above-mentioned training of the acoustic model may specifically include:

[0082] According to the voiceprint feature corresponding to the training sample, determine the first error corresponding to the duration prediction module, determine the second error corresponding to the prosody prediction module, and determine the third error corresponding to the acoustic prediction module;

[0083] Fuse the first error, the second error, and the third error to obtain the corresponding first fused error, so as to update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the first fused error during the backpropagation process.

[0084] Among them, the first error can be the error between the output of the duration prediction model and the actual duration feature. The second error can be the error between the output of the prosody prediction model and the actual prosody feature. The third error can be the error between the output of the acoustic prediction module and the first acoustic feature.

[0085] In an alternative implementation, the third error can be an adversarial error. Specifically, a GAN (Generative Adversarial Networks) can be used to determine the adversarial error based on the output of the acoustic prediction module and the first acoustic feature. GAN has the advantages of simple process and low hardware pressure.

[0086] The embodiment of the present invention fuses the first error, the second error, and the third error, and uses the obtained first fused error for the backpropagation process of the three modules, which can make the parameters of the three modules meet the requirements in terms of duration prediction, prosody prediction, and acoustic prediction.

[0087] In one implementation, the above training of the acoustic model may specifically include:

[0088] Determine the first error corresponding to the duration prediction module, the second error corresponding to the prosody prediction module, and the third error corresponding to the acoustic prediction module according to the voiceprint feature corresponding to the training sample;

[0089] For the predicted acoustic feature output by the acoustic prediction module, determine the corresponding predicted voiceprint feature;

[0090] Determine the fourth error according to the voiceprint feature and the predicted voiceprint feature;

[0091] Fuse the first error, the second error, the third error, and the fourth error to obtain the corresponding second fused error, so as to update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the second fused error during the backpropagation process.

[0092] The second fused error can be the fusion of the first error, the second error, the third error, and the fourth error. The fourth error can characterize the error of the acoustic model in terms of voiceprint. Applying the component of the fourth error to the backpropagation of the acoustic model can improve the matching degree between the acoustic model and the input voiceprint feature.

[0093] In a specific implementation, the predicted acoustic model can be used as the input of the voiceprint model to obtain the predicted voiceprint model output by the voiceprint model.

[0094] Referring to Figure 2 , a schematic diagram of a training process of an acoustic model according to an embodiment of the present invention is shown. In the training process of the acoustic model, the input of the acoustic model may include: text, actual duration feature, actual prosody feature, voiceprint feature 1, and acoustic feature 1. In a specific implementation, the acoustic feature 1 corresponding to the training sample can be determined according to the traditional method for extracting acoustic features, and the acoustic feature 1 is input into the voiceprint model to obtain the voiceprint feature 1 output by the voiceprint model.

[0095] The acoustic model may specifically include: a duration prediction module, a prosody prediction module, and an acoustic prediction module. Among them, the output of the duration prediction model corresponds to a first error Loss1, the output of the prosody prediction module corresponds to a second error Loss2, and the output of the acoustic prediction module corresponds to a third error Loss3 and a fourth error Loss4. Fusing the first error, the second error, the third error, and the fourth error, the corresponding fusion methods may include but are not limited to: summation, weighted average, etc. Using the obtained second fusion error Loss in the backpropagation process of the entire acoustic model can make the parameters of the three modules meet the requirements in terms of duration prediction, prosody prediction, acoustic prediction, and voiceprint matching. Therefore, the matching degree between the acoustic model and the input voiceprint features can be improved.

[0096] Method Embodiment Three

[0097] This embodiment describes the training process of the vocoder.

[0098] In this embodiment, the vocoder can be used to convert acoustic features into playable speech waveforms. The mathematical model can be trained based on training samples to obtain a vocoder, and the vocoder can represent the mapping relationship between the input (such as acoustic features) and the output (speech).

[0099] The training process of the vocoder may include: using the trained acoustic model to determine the training acoustic features corresponding to the training samples and their voiceprint features; training the vocoder according to the training acoustic features; the vocoder is used to obtain audio according to acoustic features.

[0100] The acoustic model and the vocoder can use the same training samples, and based on these training samples, the acoustic model and the vocoder can be trained sequentially. Specifically, after the acoustic model is trained, the text and voiceprint features corresponding to the training samples are input into the trained acoustic model to obtain the training acoustic features output by the trained acoustic model. During the training process of the vocoder, these training acoustic features can be used as the input of the vocoder. It is also possible to determine the loss of the vocoder based on the output of the vocoder and the speech corresponding to the training samples, and perform backpropagation based on this loss to achieve the training convergence of the vocoder.

[0101] Referring to Figure 3 , a schematic structural diagram of a voice cloning model according to an embodiment of the present invention is shown. The voice cloning model may specifically include: a voiceprint model, an acoustic model, and a vocoder. Among them, the voiceprint model outputs voiceprint features according to the input speech; the acoustic model outputs acoustic features according to the input text and the voiceprint features output by the voiceprint model; the vocoder outputs the target audio according to the acoustic features output by the acoustic model.

[0102] In summary, the embodiment of the present invention uses voiceprint technology for voice cloning. Based on the training of the voiceprint model, a voiceprint space of a large number of users is constructed, which can save the adaptive training of the original audio.

[0103] Secondly, users only need to record the original audio of one sentence to achieve voice cloning, which can simplify the processing flow and improve the user experience.

[0104] Furthermore, since the adaptive training of the original audio can be saved, the training of the voice cloning model can be allowed offline. Compared with the traditional cloning technology that requires cloud training, it can reduce the usage cost and the performance requirements for the device.

[0105] Method Embodiment Four

[0106] Referring to Figure 4 , a flowchart of the steps of a voice cloning method according to an embodiment of the present invention is shown. The method may specifically include the following steps:

[0107] Step 401: Receive the text and the original audio of the cloning object;

[0108] Step 402: Determine the voiceprint features corresponding to the original audio;

[0109] Step 403: Input the text and the voiceprint features into the acoustic model to obtain the corresponding acoustic features; among them, the acoustic model can be obtained according to the voiceprint features corresponding to the training samples;

[0110] Step 404: Determine the corresponding target audio according to the acoustic features.

[0111] Figure 4The method embodiments shown can be executed by a client or a server. It can be understood that the embodiments of the present invention do not limit the specific execution entity of the method embodiments.

[0112] In step 401, the cloning object can be the speaker corresponding to the original audio. The original audio can correspond to one sentence or multiple sentences. The text can represent the text to be converted into the target audio. Both the text and the original audio of the cloning object can be determined by the user.

[0113] In step 402, the first acoustic features of the original audio can be extracted, and the first acoustic features of the original audio are input into the voiceprint model to obtain the voiceprint features output by the voiceprint model.

[0114] In step 403, an acoustic model is trained according to the voiceprint features corresponding to the training samples, so that the acoustic model can represent the mapping relationship between the input (such as text and voiceprint features) and the output (acoustic features).

[0115] In one implementation, the acoustic model can include: a duration prediction module, a prosody prediction module, and an acoustic prediction module; during the backpropagation process of training the acoustic model, the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module are updated according to the first fusion error corresponding to the duration prediction module, the prosody prediction module, and the acoustic prediction module.

[0116] In another implementation, the acoustic model can include: a duration prediction module, a prosody prediction module, and an acoustic prediction module; during the backpropagation process of training the acoustic model, the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module are updated according to the second fusion error corresponding to the duration prediction module, the prosody prediction module, the acoustic prediction module, and the voiceprint error; wherein, the voiceprint error represents the error between the predicted voiceprint features obtained based on the output of the acoustic prediction module and the voiceprint features corresponding to the training samples.

[0117] Refer to Figure 5 , which shows a schematic diagram of the usage process of an acoustic model according to an embodiment of the present invention. Among them, the input of the acoustic model includes: text and the voiceprint features corresponding to the original audio. The acoustic model can include: a duration prediction module, a prosody prediction module, and an acoustic prediction module. Among them, the duration prediction model can predict the duration features corresponding to the text and the voiceprint features; the prosody prediction module can predict the prosody features corresponding to the text and the voiceprint features; the acoustic prediction module can predict the acoustic features according to the text, the voiceprint features, the output of the duration prediction module, and the output of the prosody prediction module, and output the predicted acoustic features.

[0118] In step 404, the acoustic features can be input into a vocoder to obtain the target audio output by the vocoder. In practical applications, the target audio can be output to the user.

[0119] In summary, in the voice cloning method according to the embodiments of the present invention, an acoustic model is trained based on the voiceprint features corresponding to the training samples. Among them, the acoustic model can represent the mapping relationship between the input (such as text and voiceprint features) and the output (acoustic features), and can obtain an output that matches the input. Since the input of the acoustic model includes voiceprint features, therefore, the embodiments of the present invention can obtain acoustic features that match the voiceprint features based on the acoustic model.

[0120] In addition, in the voice cloning process of the embodiments of the present invention, the original audio of the cloning object is used to determine the voiceprint features. Since the determination of the voiceprint features has low requirements for the amount of audio data, therefore, the embodiments of the present invention can reduce the requirements for the original audio of the cloning object, that is, can reduce the amount of audio data of the cloning object. In practical applications, the cloning object can record one or more sentences to achieve voice cloning, so the user experience can be improved.

[0121] Moreover, the principle of the voice cloning process of the embodiments of the present invention is specifically as follows: Voice cloning is realized according to the voiceprint features corresponding to the original audio and the mapping relationship between the input (including voiceprint features) and the acoustic features represented by the acoustic model. Since the adaptive training corresponding to the original audio can be saved, the voice cloning process of the embodiments of the present invention can save the time of adaptive training and improve the processing efficiency; and the voice cloning process of the embodiments of the present invention can reduce the requirements for device performance, and can be applicable to both the server side and the client side, so the applicable range can be increased.

[0122] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of combinations of movement actions. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described order of movement actions, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the movement actions involved are not necessarily essential to the embodiments of the present invention.

[0123] Device embodiments

[0124] Referring to Figure 6 , a structural block diagram of an embodiment of a voice cloning device according to the present invention is shown. The above device may specifically include:

[0125] A receiving module 601, configured to receive text and the original audio of the cloning object;

[0126] A voiceprint determination module 602, configured to determine the voiceprint feature corresponding to the original audio;

[0127] An acoustic determination module 603, configured to input the text and the voiceprint feature into an acoustic model to obtain a corresponding acoustic feature; wherein, the acoustic model is obtained according to the voiceprint feature corresponding to the training sample;

[0128] An audio determination module 604, configured to determine a corresponding target audio according to the acoustic feature.

[0129] Optionally, the acoustic model may include: a duration prediction module, a prosody prediction module, and an acoustic prediction module;

[0130] During the backpropagation process of training the acoustic model, update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the first fusion error corresponding to the duration prediction module, the prosody prediction module, and the acoustic prediction module.

[0131] Optionally, the acoustic model may include: a duration prediction module, a prosody prediction module, and an acoustic prediction module;

[0132] During the backpropagation process of training the acoustic model, update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the second fusion error corresponding to the duration prediction module, the prosody prediction module, the acoustic prediction module, and the voiceprint error; wherein, the voiceprint error characterizes the error between the predicted voiceprint feature obtained based on the output of the acoustic prediction module and the voiceprint feature corresponding to the training sample.

[0133] Refer to Figure 7 , which shows a structural block diagram of an embodiment of a training device of the present invention. The above device may specifically include:

[0134] A voiceprint determination module 701, configured to determine the voiceprint feature corresponding to the training sample;

[0135] An acoustic training module 702, configured to train an acoustic model according to the voiceprint feature corresponding to the training sample.

[0136] Optionally, the acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module;

[0137] The training of the acoustic model includes:

[0138] According to the voiceprint feature corresponding to the training sample, determine the first error corresponding to the duration prediction module, determine the second error corresponding to the prosody prediction module, and determine the third error corresponding to the acoustic prediction module;

[0139] The first error, the second error, and the third error are fused to obtain a corresponding first fused error, so as to update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the first fused error during the backpropagation process.

[0140] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiments.

[0141] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0142] Regarding the devices in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0143] Figure 8 FIG. 13 is a block diagram of a device 1300 for voice cloning according to an exemplary embodiment. For example, the device 1300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0144] Referring to Figure 8 , the device 1300 may include one or more of the following components: a processing component 1302, a memory 1304, a power component 1306, a multimedia component 1308, an audio component 1310, an input / output (I / O) interface 1312, a sensor component 1314, and a communication component 1316.

[0145] The processing component 1302 generally controls the overall operation of the device 1300, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing element 1302 may include one or more processors 1320 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 1302 may include one or more modules to facilitate the interaction between the processing component 1302 and other components. For example, the processing component 1302 may include a multimedia module to facilitate the interaction between the multimedia component 1308 and the processing component 1302.

[0146] The memory 1304 is configured to store various types of data to support the operation of the device 1300. Examples of such data include instructions for any application or method operating on the device 1300, contact data, phone book data, messages, pictures, videos, etc. The memory 1304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0147] The power supply component 1306 provides power to various components of the device 1300. The power supply component 1306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 1300.

[0148] The multimedia component 1308 includes a screen that provides an output interface between the device 1300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1308 includes a front camera and / or a rear camera. When the device 1300 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0149] The audio component 1310 is configured to output and / or input audio signals. For example, the audio component 1310 includes a microphone (MIC) that is configured to receive external audio signals when the device 1300 is in an operating mode, such as a call mode, a recording mode, and a voice data processing mode. The received audio signals can be further stored in the memory 1304 or transmitted via the communication component 1316. In some embodiments, the audio component 1310 further includes a speaker for outputting audio signals.

[0150] The I / O interface 1312 provides an interface between the processing component 1302 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0151] The sensor assembly 1314 includes one or more sensors for providing a status assessment of various aspects of the device 1300. For example, the sensor assembly 1314 can detect the on / off state of the device 1300, the relative positioning of components, such as the display and keypad of the device 1300. The sensor assembly 1314 can also detect a change in the position of the device 1300 or a component of the device 1300, the presence or absence of user contact with the device 1300, the orientation or acceleration / deceleration of the device 1300, and a change in the temperature of the device 1300. The sensor assembly 1314 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1314 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1314 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0152] The communication component 1316 is configured to facilitate communication between the device 1300 and other devices in a wired or wireless manner. The device 1300 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1316 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency data processing (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0153] In an exemplary embodiment, the device 1300 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0154] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 1304 including instructions, is also provided. The above instructions can be executed by the processor 1320 of the device 1300 to complete the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0155] In addition, it should be noted here that the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer programs executed by the aforementioned first cross-device device and second cross-device device. The computer programs include program instructions. When the aforementioned processor executes the program instructions, it can execute the descriptions of the training method or voice cloning method in the corresponding embodiments mentioned above. Therefore, the descriptions will not be repeated here. In addition, the descriptions of the beneficial effects of using the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the descriptions of the method embodiments of the present application. Figure 1 and Figure 4 the descriptions of the training method or voice cloning method in the corresponding embodiments. Therefore, the descriptions will not be repeated here. In addition, the descriptions of the beneficial effects of using the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the descriptions of the method embodiments of the present application.

[0156] In addition, it should be noted that: the embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program may include computer instructions, and the computer instructions may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, so that the computer device executes the descriptions of the training method or voice cloning method in the corresponding embodiments mentioned above. Figure 1 and Figure 4 the descriptions of the training method or voice cloning method in the corresponding embodiments. Therefore, the descriptions will not be repeated here. In addition, the descriptions of the beneficial effects of using the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer program product or the computer program involved in the present application, please refer to the descriptions of the method embodiments of the present application.

[0157] Figure 9 FIG. 15 is a schematic structural diagram of a server in an embodiment of the present invention. The server 1900 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1922 (for example, one or more processors) and a memory 1932, and one or more storage media 1930 (for example, one or more mass storage devices) for storing application programs 1942 or data 1944. Among them, the memory 1932 and the storage media 1930 may be transient storage or persistent storage. The programs stored in the storage media 1930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 1922 may be configured to communicate with the storage media 1930 and execute a series of instruction operations in the storage media 1930 on the server 1900.

[0158] The server 1900 may further include one or more power supplies 1926, one or more wired or wireless network interfaces 1950, one or more input / output interfaces 1958, one or more keyboards 1956, and / or one or more operating systems 1941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0159] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and examples are only to be considered exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0160] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

[0161] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0162] The above has introduced in detail a training method, a voice cloning method, a training device, a voice cloning device, a device for training, a device for voice cloning, and a machine-readable medium provided by the embodiments of the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A voice cloning method, characterized in that, The method includes: Receiving the text and the original audio of the cloned object; Determining the voiceprint feature corresponding to the original audio; Inputting the text and the voiceprint feature into an acoustic model to obtain corresponding acoustic features; wherein, the acoustic model is obtained according to the voiceprint features corresponding to the training samples; the acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module; during the backpropagation process of training the acoustic model, the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module are updated according to the first fusion error corresponding to the duration prediction module, the prosody prediction module, and the acoustic prediction module; wherein, the first fusion error is obtained by fusing the first error corresponding to the duration prediction module, the second error corresponding to the prosody prediction module, and the third error corresponding to the acoustic prediction module; Determining the corresponding target audio according to the acoustic features.

2. The method according to claim 1, wherein The acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module; During the backpropagation process of training the acoustic model, the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module are updated according to the second fusion error corresponding to the duration prediction module, the prosody prediction module, the acoustic prediction module, and the voiceprint error; wherein, the voiceprint error represents the error between the predicted voiceprint feature obtained based on the output of the acoustic prediction module and the voiceprint feature corresponding to the training sample; the second fusion error is obtained by fusing the first error corresponding to the duration prediction module, the second error corresponding to the prosody prediction module, the third error corresponding to the acoustic prediction module, and the voiceprint error.

3. A training method, characterized in that, The method includes: Determining the voiceprint feature corresponding to the training sample; Training an acoustic model according to the voiceprint feature corresponding to the training sample; the acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module; The training of the acoustic model includes: According to the voiceprint feature corresponding to the training sample, determining the first error corresponding to the duration prediction module, determining the second error corresponding to the prosody prediction module, and determining the third error corresponding to the acoustic prediction module; Fusing the first error, the second error, and the third error to obtain the corresponding first fusion error, so as to update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the first fusion error during the backpropagation process.

4. The method according to claim 3, wherein The acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module; The training of the acoustic model includes: According to the voiceprint feature corresponding to the training sample, determining the first error corresponding to the duration prediction module, determining the second error corresponding to the prosody prediction module, and determining the third error corresponding to the acoustic prediction module; Determining the corresponding predicted voiceprint feature for the predicted acoustic feature output by the acoustic prediction module; Determining the fourth error according to the voiceprint feature and the predicted voiceprint feature; Fuse the first error, the second error, the third error, and the fourth error to obtain a corresponding second fused error, so as to update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the second fused error during the backpropagation process.

5. The method according to any one of claims 3 to 4, characterized in that, The method further includes: Using the trained acoustic model, determine the training acoustic features corresponding to the training samples and their voiceprint features; Train a vocoder according to the training acoustic features; the vocoder is used to obtain an audio according to the acoustic features.

6. A voice cloning device, characterized in that, The device includes: A receiving module, configured to receive text and the original audio of the cloning object; A voiceprint determination module, configured to determine the voiceprint features corresponding to the original audio; An acoustic determination module, configured to input the text and the voiceprint features into an acoustic model to obtain corresponding acoustic features; wherein, the acoustic model is obtained according to the voiceprint features corresponding to the training samples; the acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module; during the backpropagation process of training the acoustic model, update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the first fused error corresponding to the duration prediction module, the prosody prediction module, and the acoustic prediction module; wherein, the first fused error is obtained by fusing the first error corresponding to the duration prediction module, the second error corresponding to the prosody prediction module, and the third error corresponding to the acoustic prediction module; An audio determination module, configured to determine a corresponding target audio according to the acoustic features.

7. The device according to claim 6, characterized in that, The acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module; During the backpropagation process of training the acoustic model, update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the second fused error corresponding to the duration prediction module, the prosody prediction module, the acoustic prediction module, and the voiceprint error; wherein, the voiceprint error represents the error between the predicted voiceprint features obtained based on the output of the acoustic prediction module and the voiceprint features corresponding to the training samples.

8. A training device, characterized in that, The device includes: A voiceprint determination module, configured to determine the voiceprint features corresponding to the training samples; An acoustic training module, configured to train an acoustic model according to the voiceprint features corresponding to the training samples; the acoustic model includes: a duration prediction module, a prosody prediction module, and an acoustic prediction module; The training of the acoustic model includes: According to the voiceprint features corresponding to the training samples, determine the first error corresponding to the duration prediction module, determine the second error corresponding to the prosody prediction module, and determine the third error corresponding to the acoustic prediction module; Fuse the first error, the second error, and the third error to obtain a corresponding first fused error, so as to update the parameters of the duration prediction module, the prosody prediction module, and the acoustic prediction module according to the first fused error during the backpropagation process.

9. An apparatus for training a voice cloning model, characterized in that, It includes a memory, and one or more programs, where the one or more programs are stored in the memory, and when the one or more programs are executed by one or more processors, the steps of the method according to any one of claims 1 to 5 are implemented.

10. One or more machine-readable media, characterized in that, Instructions are stored thereon, and when executed by one or more processors, cause the device to perform the method according to one or more of claims 1 to 5.

11. A computer program product, characterized in that, The computer program product includes computer instructions, which are stored in a computer-readable storage medium and are adapted to be read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1 - 5.

Citation Information

Patent Citations

  • Voice cloning method and device based on single-speaker voice synthesis data set

    CN111048064A

  • Text-to-language conversion method and device

    CN111508469A

  • Voice cloning method, system and device based on neural network and storage medium

    CN112233646A

Cited By

  • High-fidelity AI voiceprint cloning method, system and its storage medium

    CN120783726B