Voice cloning method, device, equipment and storage medium
By extracting and training the multi-dimensional features of the acoustic model, the problem of insufficient cloning speech authenticity in existing speech cloning techniques is solved, and a more natural speech synthesis effect is achieved.
Patent Information
- Application Number
- CN202211045094.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-08-30
AI Technical Summary
The clone voice generated by existing voice cloning technology is poor in authenticity and cannot effectively simulate the characteristics of real people's voice.
By extracting the recording environment characteristics, timbre characteristics and pronunciation time characteristics of the reference speech, the acoustic model is trained to generate cloned speech that is more in line with the actual vocal characteristics of the cloned object, and a multi-dimensional sound feature extraction layer and adversarial training technology are used.
The authenticity of cloned speech is improved, making the synthesized speech more in line with the actual vocal characteristics of the cloned object, and enhancing the authenticity and nature of speech cloning.
Smart Images

Figure CN115394285B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing, and in particular to a voice cloning method, apparatus, device, and storage medium. Background Art
[0002] With the advancement of deep learning, speech synthesis technology has made tremendous progress. Realistic, natural speech synthesis technology has been applied to voice interaction systems such as mobile phone voice assistants, smart speakers, and in-car computers. At the same time, user demand for speech synthesis technology is increasing, and the technical requirements are also increasing. Users not only expect synthesized speech to be comparable to real people, but also to have a variety of voices, even the voices of family and friends.
[0003] Voice cloning technology was developed to meet the diverse needs of synthesized voices. It consists of three main components: a front-end system (used to convert text and symbols into phonemes); an acoustic model system (used to convert phonemes into acoustic features. Phonemes are the smallest units of speech divided according to the natural properties of speech); and a vocoder system (used to convert acoustic features into audio). The acoustic model system uses a reference speech segment of person A and the phonemes of the text to be cloned as input. The acoustic model system uses features extracted from the reference speech to synthesize the acoustic features of the cloned speech of person A reading the text. The acoustic features of the actual speech segment of person A reading the text are then used as training labels to train the acoustic model system.
[0004] The above method can only rigidly simulate real speech to generate cloned speech. The generated cloned speech does not fit the characteristics of the person's actual voice, and the authenticity of the cloned speech is poor. Summary of the Invention
[0005] The embodiments of the present application provide a voice cloning method, apparatus, device, and storage medium that can improve the authenticity of cloned voices. The technical solution is as follows.
[0006] According to one aspect of the present application, a voice cloning method is provided, the method comprising:
[0007] Get the phoneme information of the text to be cloned;
[0008] Extracting features from the reference speech to obtain speech features;
[0009] synthesizing a cloned voice of the phoneme information according to the voice features;
[0010] The speech features include recording environment features and timbre features, or the speech features include the recording environment features, the timbre features and rhythm duration features.
[0011] According to another aspect of the present application, a voice cloning device is provided, the device comprising:
[0012] Phoneme module, used to obtain the phoneme information of the text to be cloned;
[0013] A feature extraction module is used to extract features from the reference speech to obtain speech features;
[0014] A synthesis module, configured to synthesize the cloned speech of the phoneme information according to the speech features;
[0015] The speech features include recording environment features and timbre features, or the speech features include the recording environment features, the timbre features and rhythm duration features.
[0016] According to another aspect of the present application, a computer device is provided, comprising: a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the voice cloning method described above.
[0017] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the voice cloning method described above.
[0018] According to another aspect of the embodiments of the present disclosure, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice cloning method provided in the above-mentioned optional implementation.
[0019] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:
[0020] By training the acoustic model, it can extract the acoustic features of the reference speech in multiple practical dimensions such as timbre, recording environment, and rhythmic duration. The acoustic model generates a cloned speech based on the actual acoustic features of the reference speech. Compared with the method of bluntly imitating real speech, the sound output by the acoustic model in this method is more consistent with the actual vocal characteristics of the cloned object, thereby improving the authenticity of the cloned speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 is a block diagram of a computer device provided by an exemplary embodiment of the present application;
[0023] Figure 2 is a schematic diagram of a voice cloning model provided by another exemplary embodiment of the present application;
[0024] Figure 3 is a schematic diagram of a voice cloning model provided by another exemplary embodiment of the present application;
[0025] Figure 4 is a flow chart of a voice cloning method provided by another exemplary embodiment of the present application;
[0026] Figure 5 is a flow chart of a voice cloning method provided by another exemplary embodiment of the present application;
[0027] Figure 6 is a flow chart of a voice cloning method provided by another exemplary embodiment of the present application;
[0028] Figure 7 is a schematic diagram of a voice cloning model training method provided by another exemplary embodiment of the present application;
[0029] Figure 8 is a flow chart of a voice cloning method provided by another exemplary embodiment of the present application;
[0030] Figure 9 is a block diagram of a voice cloning device provided by another exemplary embodiment of the present application;
[0031] Figure 10 is a structural diagram of a server provided by another exemplary embodiment of the present application;
[0032] Figure 11 is a block diagram of a terminal provided by another exemplary embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0034] Figure 1A schematic diagram of a computer device 101 provided in an exemplary embodiment of the present application is shown. The computer device 101 may be a terminal or a server.
[0035] The terminal may include at least one of a digital camera, a smartphone, a laptop computer, a desktop computer, a tablet computer, a smart speaker, and an intelligent robot. Optionally, the terminal may also be a device with a camera, such as a facial recognition payment device, a surveillance device, or an access control device. In one optional implementation, the voice cloning method provided in this application may be applied to an application with voice cloning functionality, such as an audio processing application, a voice cloning application, a video processing application, an audio publishing application, a video publishing application, a social application, a shopping application, a live broadcast application, a forum application, an information application, a lifestyle application, an office application, and the like. Optionally, the terminal may have a client for the application installed.
[0036] Exemplarily, the terminal stores a voice cloning model 102. When the client needs to use the voice cloning function, the client can call the voice cloning model to complete the voice cloning. Exemplarily, the voice cloning process can be completed by the terminal or by the server.
[0037] The terminal and the server are connected to each other through a wired or wireless network.
[0038] The terminal includes a first memory and a first processor. The first memory stores a voice cloning model; the voice cloning model is called and executed by the first processor to implement the voice cloning method provided in the present application. The first memory may include, but is not limited to, the following: random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM).
[0039] The first processor may be composed of one or more integrated circuit chips. Alternatively, the first processor may be a general-purpose processor, such as a central processing unit (CPU) or a network processor (NP). Alternatively, the first processor may implement the voice cloning method provided in the present application by running a program or code.
[0040] The server includes a second memory and a second processor. The second memory stores a voice cloning model; the voice cloning model is invoked by the second processor to implement the voice cloning method provided herein. Optionally, the second memory may include, but is not limited to, RAM, ROM, PROM, EPROM, and EEPROM. Optionally, the second processor may be a general-purpose processor, such as a CPU or NP.
[0041] The computer device 101 stores a voice cloning model 102. When the computer device 101 needs to perform voice cloning, the voice cloning model 102 is called to perform voice cloning based on the reference voice and the text to be cloned to obtain a cloned voice.
[0042] Optional, such as Figure 2 As shown, the speech cloning model 102 includes a feature extraction layer 103 and an acoustic model 104. A computer device inputs a reference speech into the feature extraction layer 103 to extract reference speech features from the reference speech. The computer device then obtains the phoneme information of the text to be cloned and inputs the phoneme information and reference speech features into the acoustic model 104, which then outputs the cloned speech. This produces a cloned speech that reads the text to be cloned using the human voice of the reference speech.
[0043] In an optional implementation, Figure 3 As shown, the feature extraction layer 103 includes an environmental feature extraction layer, a timbre feature extraction layer, and a prosody duration feature extraction layer. The acoustic model 104 includes an encoder, a prosody duration estimation layer, and a decoder. The computer device inputs the reference speech into the environmental feature extraction layer to obtain recording environmental features, inputs the reference speech into the timbre feature extraction layer to obtain timbre features, and inputs the reference speech into the prosody duration feature extraction layer to obtain prosody duration features. The computer device obtains the phoneme information of the text to be cloned, inputs the phoneme information into the encoder to obtain phoneme coding information, inputs the phoneme coding information and prosody duration features into the prosody duration estimation layer to obtain the phoneme duration of each phoneme. The phoneme coding information, phoneme duration, recording environmental features, timbre features, and prosody duration features are input into the decoder to obtain the cloned speech.
[0044] In one application scenario, the voice cloning method provided in this application is used in a voice application to perform voice cloning on user-entered text to generate a cloned voice, which is then sent. For example, in a social application, a user first records an audio clip of a speech and then enters a text to be cloned. The terminal or server then feeds the recorded audio and the text to be cloned into a voice cloning model to generate the cloned voice. The user can then send the cloned voice to the chat partner.
[0045] In another application scenario, the voice cloning method provided in this application can be used in an audio or video editing application to perform voice cloning on user-provided text to produce a cloned voice. For example, a user submits a voice recording and a text segment. The terminal or server inputs the voice recording and text segment into a voice cloning model to produce a cloned voice. This allows the user to quickly obtain an audio segment of the target text being read, allowing them to use the cloned voice for audio or video editing.
[0046] Figure 4 FIG1 shows a flow chart of a voice cloning method provided by an exemplary embodiment of the present application. The method can be executed by a computer device, for example, Figure 1 The method includes the following steps.
[0047] Step 210: Obtain phoneme information of the text to be cloned.
[0048] The computer device converts the text to be cloned into phonemes to obtain phoneme information. The phoneme information includes at least one phoneme corresponding to the text to be cloned.
[0049] Phoneme is the smallest phonetic unit divided according to the natural properties of speech. It is analyzed based on the pronunciation actions in a syllable, and one action constitutes a phoneme.
[0050] The text to be cloned includes phrases, sentences, paragraphs, and chapters composed of characters and / or symbols. The language of the text to be cloned is not limited.
[0051] The computer device receives the text to be cloned input by the user, or the computer device obtains the text to be cloned through a specified path, for example, the computer device obtains the text to be cloned from a text library.
[0052] Step 220: Extract features from the reference speech to obtain speech features.
[0053] The reference voice is the voice of the person being cloned (the voice recorder). Alternatively, the reference voice is the voice of the person being cloned reading a text. Alternatively, the reference voice is an audio clip containing the voice of the person being cloned. For example, the reference voice is a recording of person A reading a text.
[0054] The computer device extracts speech features of the reference speech from multiple dimensions. These speech features may represent the vocalization characteristics, reading characteristics, or characteristics of the recording environment of the reference speech. For example, the computer device may extract speech features of the reference speech from the dimensions of recording environment, timbre, and rhythmic duration.
[0055] Optionally, the speech features include recording environment features and timbre features, or the speech features include recording environment features, timbre features and rhythm duration features.
[0056] Optional, such as Figure 2 As shown, the computer device calls the feature extraction layer to extract features from the reference speech to obtain recording environment features, timbre features, and rhythm duration features. Optionally, the computer device calls the environment feature extraction layer to extract features from the reference speech to obtain recording environment features; the computer device calls the timbre feature extraction layer to extract features from the reference speech to obtain timbre features; and the computer device calls the rhythm duration feature extraction layer to extract features from the reference speech to obtain rhythm duration features.
[0057] The recording environment feature is used to characterize the characteristics of the reference speech recording environment, the timbre feature is used to characterize the timbre characteristics of the cloned object, and the prosody duration feature is used to characterize the prosody characteristics of the reference speech. For example, the prosody duration feature is used to characterize the tone, duration, pitch, and other characteristics of the cloned object's speech, or the prosody duration feature is used to characterize the intonation characteristics of the cloned object's speech.
[0058] By extracting features from three dimensions: recording environment, timbre, and rhythm duration, the computer device can fully learn the acoustic features of the reference speech and synthesize the cloned speech based on the extracted features. For example, compared to the method that only extracts timbre and rhythm duration features, the embodiment of the present application further extracts recording environment features. If the recording environment of the reference speech is relatively noisy or has a reverberation effect, it can be extracted as a recording environment feature, and when synthesizing the cloned speech, the cloned speech will also have a similar recording environment effect. However, if the recording environment feature is not extracted, this part of the feature may be extracted into the timbre feature, making the timbre feature inaccurate, the synthesized cloned speech timbre is fake, and the recording environment effect is poor.
[0059] Step 230: synthesize the cloned speech of the phoneme information according to the speech features.
[0060] The computer device synthesizes a cloned voice of phoneme information based on recording environment features and timbre features extracted from the reference voice.
[0061] Alternatively, the computer device synthesizes a cloned voice of phoneme information based on recording environment features, timbre features, and prosodic duration features extracted from the reference voice.
[0062] Optionally, the computer device synthesizes acoustic features of phoneme information based on recording environment features, timbre features, and rhythmic duration features extracted from the reference speech, and converts the acoustic features into cloned speech (cloned audio) through a vocoder.
[0063] Optionally, the computer device inputs the recording environment characteristics, timbre characteristics, rhythm duration characteristics and phoneme information into the acoustic model to obtain acoustic characteristics, and inputs the acoustic characteristics into the vocoder to obtain cloned audio.
[0064] For example, the acoustic feature may be a mel spectrogram or a mel frequency cepstrum coefficient (MFCC) of the audio.
[0065] The cloned voice is an audio recording of the text to be cloned, using the voice of the target. The cloned voice is an audio recording of the text to be cloned, imitating the voice of the target (the human voice in the reference voice).
[0066] In summary, the method provided in this embodiment trains an acoustic model to extract the acoustic features of a reference speech in terms of timbre and recording environment, or in terms of timbre, recording environment, and rhythm duration. The acoustic model generates a cloned speech based on the acoustic features of the reference speech. Compared to methods that rigidly imitate real speech, the sound output by the acoustic model in this method is more consistent with the actual vocal characteristics of the cloned object, thereby improving the authenticity of the cloned speech.
[0067] Exemplarily, a voice cloning model is stored in a computer device, and the computer device executes the voice cloning method provided in the embodiment of the present application by calling the voice cloning model.
[0068] Optional, such as Figure 2 As shown, the voice cloning model includes a feature extraction layer and an acoustic model.
[0069] Optional, such as Figure 3 As shown, the feature extraction layer includes an environmental feature extraction layer, a timbre feature extraction layer and a rhythm duration feature extraction layer.
[0070] Optionally, the acoustic model includes an encoder and a decoder. Alternatively, Figure 3 As shown, the acoustic model includes an encoder, a decoder, and a prosodic duration estimation layer.
[0071] For example, a method based on Figure 3 The voice cloning method implemented by the voice cloning model shown.
[0072] Figure 5FIG1 shows a flow chart of a voice cloning method provided by an exemplary embodiment of the present application. The method can be executed by a computer device, for example, Figure 1 The method is executed by the terminal or server shown in the figure. The method includes the following steps.
[0073] Step 210: Obtain phoneme information of the text to be cloned.
[0074] For example, the computer device obtains input text to be cloned in response to a user's editing operation.
[0075] Alternatively, the computer device obtains the text audio and performs speech recognition on the text audio to obtain the text to be cloned, or performs phoneme extraction on the text audio to obtain phoneme information. For example, the computer device may record a user's speech through a microphone, perform speech recognition on the user's speech to obtain the text to be cloned, and convert the text to be cloned into phonemes to obtain phoneme information.
[0076] Step 221: Input the reference speech into the environment feature extraction layer to obtain the recording environment features.
[0077] Optionally, the environmental feature extraction layer includes x convolutional layers, where x is a positive integer less than the first threshold. For example, the environmental feature extraction layer includes three convolutional layers, each of which includes any number of convolution kernels. Of course, the structure of the environmental feature extraction layer is not limited to this implementation.
[0078] Exemplarily, the environmental feature extraction layer is a simple convolutional layer structure, which extracts shallow features in the reference speech through a simple convolutional structure, that is, the recording environment features that exist in the reference speech from beginning to end.
[0079] Step 222: Input the reference speech into the timbre feature extraction layer to obtain timbre features.
[0080] Optionally, the timbre feature extraction layer includes a convolutional layer and a recurrent neural network. For example, the timbre feature extraction layer includes three convolutional layers and one convolutional neural network. Of course, the structure of the timbre feature extraction layer is not limited to this implementation.
[0081] The timbre feature extraction layer extracts deep features from the reference speech through a complex network structure, and trains the timbre feature extraction layer to output the same features for the speech of the same cloned object during the training phase, so that the timbre feature extraction layer can accurately extract the timbre features of the cloned object.
[0082] Step 223: Input the reference speech into the prosody duration feature extraction layer to obtain prosody duration features.
[0083] Optionally, the rhythm duration feature extraction layer includes a convolutional layer and a recurrent neural network. For example, the rhythm duration feature extraction layer includes three convolutional layers and one convolutional neural network. Of course, the structure of the rhythm duration feature extraction layer is not limited to this implementation.
[0084] The prosodic duration feature extraction layer extracts deep features from the reference speech through a complex network structure, that is, the prosodic features of the phonemes in the reference speech when the cloned object speaks, to characterize the speaking characteristics of the cloned object.
[0085] Optionally, the method may include step 221 and step 222; or, the method may include step 221, step 222 and step 223.
[0086] Step 231: Input the phoneme information into the encoder to obtain phoneme coding information.
[0087] Optionally, the computer device inputs the phoneme information into an encoder, and the encoder performs phoneme encoding on the phoneme information to obtain phoneme encoding information.
[0088] The encoder can be implemented using a convolutional neural network (CNN), a transformer neural network (Transformer), or a convolution-enhanced transformer neural network (Conformer). For example, the encoder can be implemented using a four-layer transformer neural network (Transformer), but the encoder structure is not limited to this implementation.
[0089] Step 232: Input the phoneme coding information, recording environment characteristics, timbre characteristics and rhythm duration characteristics into a decoder to obtain cloned speech.
[0090] The computer device inputs the factor coding information and the voice features into a decoder to obtain the cloned voice.
[0091] Optionally, the computer device cascades or superimposes the phoneme coding information, recording environment characteristics and timbre characteristics and inputs them into a decoder for decoding processing to obtain the cloned voice.
[0092] Alternatively, the computer device cascades or superimposes the phoneme coding information, recording environment characteristics, timbre characteristics and rhythm duration characteristics and inputs them into a decoder for decoding processing to obtain the cloned voice.
[0093] The computer device cascades or superimposes the phoneme coding information, recording environment characteristics, timbre characteristics and rhythm duration characteristics, inputs them into a decoder to obtain acoustic characteristics, and inputs the acoustic characteristics into a vocoder to output cloned speech.
[0094] The decoder can be implemented using a convolutional neural network (CNN), a transformer neural network (Transformer), or a convolution-enhanced transformer neural network (Conformer). For example, the decoder can be implemented using a six-layer transformer neural network (Transformer), but the decoder structure is not limited to this implementation.
[0095] Optional, such as Figure 6 As shown, step 231 may be followed by step 233, and step 232 may be replaced by step 234.
[0096] Step 233: Input the phoneme encoding information and the prosody duration feature information into the prosody duration estimation layer to obtain the first phoneme duration of each phoneme in the phoneme information.
[0097] The first phoneme duration refers to the phoneme duration output according to the prosodic duration feature of the reference speech in step 210 .
[0098] The computer device cascades or superimposes the phoneme coding information and the rhythmic duration feature information and inputs them into the rhythmic duration estimation layer to perform rhythmic duration estimation and output the duration of each phoneme (phoneme duration).
[0099] The computer device outputs the phoneme duration of each phoneme in the phoneme encoding information based on the prosodic duration characteristics of the reference speech. For example, if the phoneme encoding information includes ten phonemes, the prosodic duration estimation layer outputs ten phoneme durations, with the ten phoneme durations corresponding one to each of the ten phonemes. The phoneme duration is the duration of each phoneme, for example, the duration of the first phoneme is 1 second, or the duration of the second phoneme is 3 frames.
[0100] Optionally, the rhythm duration estimation layer is composed of a convolutional layer. For example, the rhythm duration estimation layer is implemented using a three-layer convolutional neural network. Of course, the structure of the rhythm duration estimation layer is not limited to this implementation.
[0101] Step 234: input the phoneme coding information, the duration of the first phoneme, the recording environment characteristics, the timbre characteristics and the rhythm duration characteristics into a decoder to obtain the cloned speech.
[0102] The computer device cascades or superimposes the phoneme coding information, the duration of the first phoneme of each phoneme, the recording environment characteristics, the timbre characteristics and the rhythm duration characteristics, and inputs them into the decoder for decoding processing to obtain the cloned voice.
[0103] The computer device cascades or superimposes the phoneme coding information, the first phoneme duration of each phoneme, the recording environment characteristics, the timbre characteristics and the rhythm duration characteristics, inputs them into the decoder to obtain the acoustic characteristics, and inputs the acoustic characteristics into the vocoder to output the cloned voice.
[0104] The first phoneme duration includes the phoneme duration of each phoneme in the text to be cloned. For example, if the text to be cloned includes ten phonemes, the first phoneme duration includes the durations of 10 phonemes.
[0105] In summary, the method provided in this embodiment utilizes a multi-dimensional sound feature extraction layer, which efficiently extracts information about the recording environment, the voice of the recorder, and the duration of the rhythm. This information assists the acoustic model in accurately reproducing the acoustic features that contain this sound information. This improves the sound similarity of voice cloning, speeds up the fine-tuning training of the acoustic model, and enhances the robustness of the entire system. This improves the synthesis effect while reducing the requirements for the recording environment and the cost of model training.
[0106] The method provided in this embodiment provides a voice cloning acoustic model based on multi-dimensional sound feature extraction, which can accurately restore the voice recorder's timbre, reading rhythm, and the background noise and reverberation of the recording environment. Thanks to the multi-dimensional sound feature extraction module, this method can easily adjust the reading rhythm and environmental noise level of the synthesized audio.
[0107] Optionally, the discriminator is used to perform adversarial training on the voice cloning model. That is, the discriminator is used to adversarially train the feature extraction layer and the acoustic model.
[0108] Figure 7 FIG1 shows a schematic diagram of a voice cloning model training method provided by an exemplary embodiment of the present application. The method can be executed by a computer device, for example, Figure 1 The method comprises the following steps:
[0109] Step 1: Call the feature extraction layer 103 to extract features of the sample reference speech to obtain sample recording environment features, sample timbre features and sample rhythm duration features, and the recording of the second clone object of the sample reference speech.
[0110] Step 2: Obtain sample phoneme information of the sample text to be cloned.
[0111] Step three: call the acoustic model 104 to generate acoustic features by combining the sample phoneme information with the sample reference speech features to obtain the sample cloned speech.
[0112] Step 4: Input the sample cloned speech into the discriminator 105 to obtain a first discrimination result.
[0113] The discriminator is used to determine whether the input information is model-generated information (false label) or real information (true label). When the discriminator determines that the input speech is a cloned speech generated by the model, the discriminator outputs a false label (for example, 0). When the discriminator determines that the input speech is a real recorded speech, the discriminator outputs a true label (for example, 1).
[0114] Optionally, the acoustic model outputs acoustic features of the sample cloned speech, and the acoustic features of the sample cloned speech are input into the discriminator 105 to obtain a first discrimination result.
[0115] In adversarial training, the goal of training the discriminator is to make the first discrimination result output a false label. When training the voice cloning model, the goal is to make the first discrimination result output a true label.
[0116] Optionally, the discriminator is composed of convolutional layers. For example, the discriminator is implemented using a five-layer two-dimensional convolutional neural network. Of course, the structure of the discriminator is not limited to this implementation method.
[0117] Step 5: Input the sample real speech into the discriminator 105 to obtain a second discrimination result. The sample real speech is the speech of the sample text to be cloned interpreted by the voice of the second clone object.
[0118] The sample real voice is the voice recorded by the second clone object reading the sample text to be cloned.
[0119] Optionally, the acoustic features of the real speech sample are input into the discriminator to obtain a second discrimination result.
[0120] During adversarial training, the goal of training the discriminator is to make the first discriminant result output the true label.
[0121] Step 6: Calculate the first loss between the sample cloned speech and the sample real speech.
[0122] Call the loss function to calculate the loss between the sample cloned speech and the sample real speech.
[0123] Step 7: Calculate the second loss of the first discrimination result and the false label.
[0124] Call the loss function to calculate the loss between the first discriminant result and the false label.
[0125] Step 8: Calculate the third loss between the second discrimination result and the true label.
[0126] Call the loss function to calculate the loss between the second discriminant result and the true label.
[0127] Step nine: train the feature extraction layer, acoustic model, and discriminator based on the first loss, the second loss, and the third loss.
[0128] Exemplarily, the model parameters of the fixed feature extraction layer and the acoustic model remain unchanged, and the model parameters of the discriminator are trained using the back propagation algorithm based on the second loss and the third loss, so that the discriminator can accurately distinguish between the model cloned speech and the real recorded speech.
[0129] Exemplarily, the model parameters of the discriminator are fixed, the fourth loss between the first discrimination result and the true label is calculated, and the speech cloning model (including the feature extraction layer and the acoustic model) is trained using the back propagation algorithm based on the first loss, the third loss, and the fourth loss, so that the cloned speech output by the speech cloning model is close to the real speech, and the discriminator determines that the output cloned speech is the real speech (true label).
[0130] Exemplarily, the computer device repeatedly executes the above nine steps until the voice cloning model converges. Experimental verification shows that repeating the above steps 500 times can effectively converge the model. After the model converges, the trained voice cloning model can be used to execute the voice cloning method provided in the embodiments of this application.
[0131] Optionally, when the feature extraction layer includes a timbre feature extraction layer, the timbre feature similarity of the training timbre feature extraction layer to the reference voice output of the same cloned object is higher than a threshold.
[0132] The threshold value can be any value, for example, 90%. For example, the goal of training the timbre feature extraction layer is to make it output the same timbre features for the reference voice of the same clone object, but the multiple timbre features actually output may not be exactly the same.
[0133] While using the above-mentioned adversarial training method to train the voice cloning model, the following method is used to train the timbre feature extraction layer.
[0134] Calling the timbre feature extraction layer to extract features from the first sample reference speech to obtain the first sample timbre features, the first sample reference speech corresponds to the third clone object; calling the timbre feature extraction layer to extract features from the second sample reference speech to obtain the second sample timbre features, the second sample reference speech corresponds to the third clone object; calculating the fourth loss of the first sample timbre features and the second sample timbre features; and training the timbre feature extraction layer according to the fourth loss.
[0135] That is, different sample reference voices of the same cloned object are input into the timbre feature extraction layer to obtain at least two sample timbre features. Based on the loss of at least two sample timbre features, the timbre feature extraction layer is trained so that it outputs the same timbre features for the voice of the same cloned object.
[0136] In summary, the method provided in this embodiment uses adversarial training to train a speech cloning model, enabling the speech cloning model to learn the characteristics of real speech and output a cloned speech that the discriminator cannot distinguish between true and false. This effectively improves the problem of over-smoothing acoustic features and enhances the quality of synthesized sound.
[0137] Optionally, when the reading rhythm of the recorder of the reference voice (clone object) is not good, the rhythm duration features of other recorders can be used to generate the cloned voice of the recorder to improve the reading effect of the cloned voice.
[0138] Figure 8 FIG1 shows a flow chart of a voice cloning method provided by an exemplary embodiment of the present application. The method can be executed by a computer device, for example, Figure 1 The method is executed by the terminal or server shown in the figure. The method includes the following steps.
[0139] Step 210: Obtain phoneme information of the text to be cloned.
[0140] Step 220 , extracting features of the reference speech to obtain recording environment features, timbre features, and rhythm duration features.
[0141] Optionally, the reference voice is a voice recorded by the first clone object.
[0142] Step 235 : synthesizing the cloned speech of the phoneme information according to the recording environment characteristics, the timbre characteristics, and the rhythmic duration characteristics of the second clone object.
[0143] Optionally, feature extraction is performed on the reference speech of the second clone object to obtain the prosody duration feature of the second clone object. For example, the reference speech of the second clone object is input into a prosody duration feature extraction layer to obtain the prosody duration feature of the second clone object.
[0144] Then, the prosody duration characteristics of the second cloned object are used to replace the prosody duration characteristics of the first cloned object in step 220. That is, the timbre characteristics of the first cloned object, the recording environment characteristics of the first cloned object, and the prosody duration characteristics of the second cloned object are used to synthesize the cloned speech of the phoneme information. The cloned speech thus generated has the timbre of the first cloned object and the prosody of the second cloned object.
[0145] Optionally, the phoneme information is input into the encoder to obtain phoneme coding information; the phoneme coding information and the rhythmic duration features of the second cloned object are input into the rhythmic duration estimation layer to obtain the second phoneme duration of each phoneme in the phoneme information; the phoneme coding information, the second phoneme duration, the recording environment features, the timbre features, and the rhythmic duration features of the second cloned object are input into the decoder to obtain the cloned speech.
[0146] In summary, the method provided by this embodiment, since the feature extraction layer can extract the speech features of each acoustic feature dimension in the reference speech, when a certain recorder performs poorly in a certain dimension, the features of other recorders in that dimension can be used to replace the features of the recorder in that dimension to improve the quality of the cloned speech finally generated. For example, when the reading level of the recorder is stumbling, the rhythmic duration features of the recorder are used to replace the rhythmic duration features of the recorder, and then the cloned speech is generated, which can make the generated cloned speech fluent. For example, the timbre features of ordinary recording users can be combined with the rhythmic features (rhythmic duration features) of professional anchors' reading to synthesize a voice with a timbre similar to that of the user and a natural reading rhythm.
[0147] The following is an embodiment of the device of the present application. For details not described in detail in the embodiment of the device, reference can be made to the corresponding records in the above method embodiment, and no further details will be given herein.
[0148] Figure 9 The following is a schematic diagram of the structure of a voice cloning device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:
[0149] Phoneme module 401, used to obtain phoneme information of the text to be cloned;
[0150] A feature extraction module 402 is used to extract features from the reference speech to obtain speech features;
[0151] Synthesis module 403, configured to synthesize the cloned speech of the phoneme information according to the speech features;
[0152] The speech features include recording environment features and timbre features, or the speech features include the recording environment features, the timbre features and rhythm duration features.
[0153] In an optional embodiment, the feature extraction module 402 is configured to input the reference speech into an environment feature extraction layer to obtain recording environment features; input the reference speech into a timbre feature extraction layer to obtain timbre features;
[0154] Alternatively, the reference speech is input into the environmental feature extraction layer to obtain the recording environment feature; the reference speech is input into the timbre feature extraction layer to obtain the timbre feature; the reference speech is input into the rhythm duration feature extraction layer to obtain the rhythm duration feature.
[0155] In an optional embodiment, the environmental feature extraction layer includes x convolutional layers, where x is a positive integer less than a first threshold;
[0156] The timbre extraction layer includes a convolutional layer and a recurrent neural network;
[0157] The rhythm duration extraction layer includes a convolutional layer and a recurrent neural network.
[0158] In an optional embodiment, the synthesis module 403 is configured to input the phoneme information into an encoder to obtain phoneme coding information;
[0159] The synthesis module 403 is configured to input the phoneme coding information and the speech features into a decoder to obtain the cloned speech.
[0160] In an optional embodiment, the speech feature includes the prosodic duration feature;
[0161] The synthesis module 403 is configured to input the phoneme encoding information and the prosody duration feature information into a prosody duration estimation layer to obtain the first phoneme duration of each phoneme in the phoneme information;
[0162] The synthesis module 403 is configured to input the phoneme encoding information, the duration of the first phoneme, and the voice features into a decoder to obtain the cloned voice.
[0163] In an optional embodiment, the feature extraction module 402 is configured to call a feature extraction layer to extract features from the reference speech to obtain the speech features;
[0164] The synthesis module 403 is used to call the acoustic model to synthesize the cloned speech of the phoneme information according to the speech features;
[0165] The device further comprises:
[0166] The training module 404 is configured to use a discriminator to perform adversarial training on the feature extraction layer and the acoustic model.
[0167] In an optional embodiment, the feature extraction layer includes a timbre feature extraction layer;
[0168] The training module 404 is used to train the timbre feature extraction layer to output timbre features of the reference speech of the same cloned object with a similarity higher than a threshold.
[0169] In an optional embodiment, the reference voice is the voice of the first clone object;
[0170] The synthesis module 403 is configured to synthesize the cloned speech of the phoneme information according to the recording environment characteristics, the timbre characteristics, and the rhythmic duration characteristics of the second clone object.
[0171] In an optional embodiment, the synthesis module 403 is configured to input the phoneme information into an encoder to obtain phoneme coding information;
[0172] The synthesis module 403 is configured to input the phoneme encoding information and the prosodic duration feature of the second clone object into a prosodic duration estimation layer to obtain a second phoneme duration of each phoneme in the phoneme information;
[0173] The synthesis module 403 is configured to input the phoneme encoding information, the duration of the second phoneme, the recording environment characteristics, the timbre characteristics and the rhythm duration characteristics of the second clone object into a decoder to obtain the cloned speech.
[0174] Figure 10 Schematic diagram of the structure of a server provided by one embodiment of the present application. Specifically, server 800 includes a central processing unit (CPU) 801, a system memory 804 including random access memory (RAM) 802 and read-only memory (ROM) 803, and a system bus 805 connecting system memory 804 and CPU 801. Server 800 also includes a basic input / output system (I / O system) 806 that facilitates information transmission between various components within the computer, and a mass storage device 807 for storing an operating system 813, application programs 814, and other program modules 815.
[0175] The basic input / output system 806 includes a display 808 for displaying information and an input device 809, such as a mouse and keyboard, for user account input. Both the display 808 and the input device 809 are connected to the central processing unit 801 via an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include an input / output controller 810 for receiving and processing input from various other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides output to a display screen, printer, or other types of output devices.
[0176] The mass storage device 807 is connected to the central processing unit 801 via a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable media provide non-volatile storage for the server 800. In other words, the mass storage device 807 may include computer-readable media (not shown) such as a hard disk or a CD-ROM drive.
[0177] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other solid-state memory technology, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media are not limited to the above-mentioned types. The above-mentioned system memory 804 and mass storage device 807 can be collectively referred to as memory.
[0178] According to various embodiments of the present application, the server 800 may also be connected to a remote computer on a network such as the Internet for operation. That is, the server 800 may be connected to the network 812 via the network interface unit 811 connected to the system bus 805, or the network interface unit 811 may be used to connect to other types of networks or remote computer systems (not shown).
[0179] The present application also provides a terminal, which includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the voice cloning method provided by each of the above method embodiments. It should be noted that the terminal can be as follows Figure 11 The terminal provided.
[0180] Figure 11The following is a block diagram of a terminal 900 according to an exemplary embodiment of the present application. Terminal 900 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 900 may also be referred to as a user account device, portable terminal, laptop terminal, desktop terminal, or other similar names.
[0181] Typically, the terminal 900 includes a processor 901 and a memory 902 .
[0182] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may also include an AI (Artificial Intelligence) processor, which is used to handle computing operations related to machine learning.
[0183] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one instruction, which is used to be executed by the processor 901 to implement the voice cloning method or voice cloning method provided in the method embodiment of the present application.
[0184] In some embodiments, terminal 900 may optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, memory 902, and peripheral device interface 903 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 903 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, a positioning assembly 908, and a power supply 909.
[0185] The peripheral device interface 903 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0186] The RF circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 904 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Exemplarily, the RF circuit 904 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user account identity module card, and the like. The RF circuit 904 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 904 may also include circuitry related to Near Field Communication (NFC), although this application does not limit this.
[0187] Display screen 905 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. When display screen 905 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 905. These touch signals can be input as control signals to processor 901 for processing. Display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 905, located on the front panel of terminal 900. In other embodiments, there can be at least two display screens 905, located on different surfaces of terminal 900 or in a foldable design. In still other embodiments, display screen 905 can be a flexible display, located on a curved or foldable surface of terminal 900. Display screen 905 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 905 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0188] The camera assembly 906 is used to capture images or videos. Exemplarily, the camera assembly 906 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0189] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user account and the environment, and convert the sound waves into electrical signals that are input into the processor 901 for processing, or input into the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there can be multiple microphones, each located in different parts of the terminal 900. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 907 may also include a headphone jack.
[0190] Positioning component 908 is used to locate the current geographic location of terminal 900 to implement navigation or LBS (Location Based Service). Positioning component 908 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, or Russia's Galileo system.
[0191] Power supply 909 is used to power various components in terminal 900. Power supply 909 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 909 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0192] In some embodiments, the terminal 900 further includes one or more sensors 910 , including but not limited to: an acceleration sensor 911 , a gyroscope sensor 912 , a pressure sensor 913 , a fingerprint sensor 914 , an optical sensor 915 , and a proximity sensor 916 .
[0193] The accelerometer 911 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the terminal 900. For example, the accelerometer 911 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 901 can control the display screen 905 to display the user account interface in either a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 911. The accelerometer 911 can also be used to collect motion data for games or user accounts.
[0194] The gyroscope sensor 912 can detect the orientation and rotation angle of the terminal 900. It can also work with the accelerometer 911 to collect 3D motions of the user account on the terminal 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can implement the following functions: motion sensing (for example, changing the UI based on the user account's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0195] The pressure sensor 913 can be set on the side frame of the terminal 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is set on the side frame of the terminal 900, it can detect the user account's grip signal on the terminal 900, and the processor 901 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 913. When the pressure sensor 913 is set on the lower layer of the display screen 905, the processor 901 controls the operable controls on the UI interface based on the pressure operation of the user account on the display screen 905. Operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0196] The fingerprint sensor 914 is used to collect the fingerprint of the user account. The processor 901 identifies the identity of the user account based on the fingerprint collected by the fingerprint sensor 914, or the fingerprint sensor 914 identifies the identity of the user account based on the collected fingerprint. When the identity of the user account is identified as a trusted identity, the processor 901 authorizes the user account to perform relevant sensitive operations, such as unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 914 can be set on the front, back, or side of the terminal 900. When a physical button or manufacturer logo is set on the terminal 900, the fingerprint sensor 914 can be integrated with the physical button or manufacturer logo.
[0197] The optical sensor 915 is used to detect ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 based on the ambient light intensity detected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera assembly 906 based on the ambient light intensity detected by the optical sensor 915.
[0198] Proximity sensor 916, also known as a distance sensor, is typically located on the front panel of terminal 900. Proximity sensor 916 is used to detect the distance between the user account and the front of terminal 900. In one embodiment, when proximity sensor 916 detects that the distance between the user account and the front of terminal 900 is gradually decreasing, processor 901 controls display screen 905 to switch from the screen-on state to the screen-off state. When proximity sensor 916 detects that the distance between the user account and the front of terminal 900 is gradually increasing, processor 901 controls display screen 905 to switch from the screen-off state to the screen-on state.
[0199] Those skilled in the art will understand that Figure 11 The structure shown in the figure does not constitute a limitation on the terminal 900, and the terminal 900 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0200] The memory further includes one or more programs, which are stored in the memory and include instructions for performing the voice cloning method provided in the embodiment of the present application.
[0201] The present application also provides a computer device, which includes: a processor and a memory, wherein the storage medium stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the voice cloning method provided by the above-mentioned method embodiments.
[0202] The present application also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the voice cloning method provided by the above-mentioned method embodiments.
[0203] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice cloning method provided in the above-mentioned optional implementation.
[0204] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0205] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0206] The above are merely optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A voice cloning method, characterized in that: The method comprises: Get the phoneme information of the text to be cloned; Extracting features from the reference speech to obtain speech features; synthesizing a cloned voice of the phoneme information according to the voice features; The speech features include recording environment features and timbre features, or the speech features include the recording environment features, the timbre features and rhythm duration features; The extracting features of the reference speech to obtain speech features includes: Inputting the reference speech into the environment feature extraction layer to obtain the recording environment feature; inputting the reference speech into the timbre feature extraction layer to obtain the timbre feature; or, The reference speech is input into the environment feature extraction layer to obtain the recording environment feature; the reference speech is input into the timbre feature extraction layer to obtain the timbre feature; the reference speech is input into the rhythm duration feature extraction layer to obtain the rhythm duration feature.
2. The method according to claim 1, characterized in that The environmental feature extraction layer includes x convolutional layers, where x is a positive integer less than a first threshold; The timbre extraction layer includes a convolutional layer and a recurrent neural network; The rhythm duration extraction layer includes a convolutional layer and a recurrent neural network.
3. The method according to claim 1 or 2, characterized in that The step of synthesizing the cloned speech of the phoneme information according to the speech features comprises: Inputting the phoneme information into an encoder to obtain phoneme coding information; The phoneme coding information and the voice features are input into a decoder to obtain the cloned voice.
4. The method according to claim 3, characterized in that The speech feature includes the prosody duration feature, and the method further includes: Inputting the phoneme encoding information and the prosody duration feature information into a prosody duration estimation layer to obtain the first phoneme duration of each phoneme in the phoneme information; The step of inputting the phoneme coding information and the voice features into a decoder to obtain the cloned voice comprises: The phoneme encoding information, the duration of the first phoneme, and the voice feature are input into a decoder to obtain the cloned voice.
5. The method according to claim 1 or 2, characterized in that The extracting features of the reference speech to obtain speech features includes: Calling a feature extraction layer to perform feature extraction on the reference speech to obtain the speech feature; The step of synthesizing the cloned speech of the phoneme information according to the speech features comprises: Invoking an acoustic model to synthesize the cloned voice of the phoneme information according to the voice features; The method further comprises: The feature extraction layer and the acoustic model are trained adversarially using a discriminator.
6. The method according to claim 5, characterized in that The feature extraction layer includes a timbre feature extraction layer; the method further includes: The timbre feature extraction layer is trained to have a timbre feature similarity to a reference voice output of the same cloned object that is higher than a threshold.
7. The method according to claim 1 or 2, characterized in that The reference voice is the voice of the first clone object; The step of synthesizing the cloned speech of the phoneme information according to the speech features comprises: synthesizing the cloned speech of the phoneme information according to the recording environment characteristics, the timbre characteristics, and the rhythmic duration characteristics of the second clone object; The prosody duration feature of the second cloned object is obtained by inputting the reference speech of the second cloned object into the prosody duration feature extraction layer.
8. The method according to claim 7, characterized in that The step of synthesizing the cloned speech of the phoneme information according to the recording environment feature, the timbre feature, and the rhythmic duration feature of the second clone object includes: Inputting the phoneme information into an encoder to obtain phoneme coding information; Inputting the phoneme encoding information and the prosodic duration feature of the second clone object into a prosodic duration estimation layer to obtain a second phoneme duration of each phoneme in the phoneme information; The phoneme coding information, the duration of the second phoneme, the recording environment characteristics, the timbre characteristics and the rhythm duration characteristics of the second clone object are input into a decoder to obtain the cloned voice.
9. A voice cloning device, characterized in that: The device comprises: Phoneme module, used to obtain the phoneme information of the text to be cloned; A feature extraction module is used to extract features from the reference speech to obtain speech features; A synthesis module, configured to synthesize the cloned speech of the phoneme information according to the speech features; The speech features include recording environment features and timbre features, or the speech features include the recording environment features, the timbre features and rhythm duration features; The extracting features of the reference speech to obtain speech features includes: Inputting the reference speech into the environment feature extraction layer to obtain the recording environment feature; inputting the reference speech into the timbre feature extraction layer to obtain the timbre feature; or, The reference speech is input into the environment feature extraction layer to obtain the recording environment feature; the reference speech is input into the timbre feature extraction layer to obtain the timbre feature; the reference speech is input into the rhythm duration feature extraction layer to obtain the rhythm duration feature.
10. A computer device, characterized in that: The computer device includes: a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the voice cloning method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the voice cloning method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Speech synthesis method and device, equipment, storage medium and program product
CN114242032A
Method and apparatus for processing speech information using a phoneme environment
US5845047A