Speech clone model training method, training device and electronic equipment

By independently training the acoustic module and variable information adaptation module in the speech cloning model and combining them with the post-processing module to process the spectral information, the problem of large differences between the cloned speech and the target speech in the speech cloning model is solved, and the accuracy of the speech cloning model and the quality of the cloned speech are improved.

CN120690172APending Publication Date: 2025-09-23GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410326168.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the existing speech cloning model, during the target speaker's speech cloning process, there are differences between the cloned speech and the target speech due to factors such as poor recording environment, unclear pronunciation and inconsistent speaking style, resulting in poor cloning effect.

Method used

The encoding module and the variable information adaptation module are set up in parallel. By iteratively updating the parameters of the acoustic module and the variable information adaptation module, the acoustic module and the variable information adaptation module are trained independently. The spectral information is processed in combination with the post-processing module to improve the accuracy of the speech cloning model.

Benefits of technology

The deviation of the prediction data of the voice cloning model is reduced, and the accuracy of the prediction data of the voice cloning model and the quality of the cloned voice are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690172A_ABST
    Figure CN120690172A_ABST
Patent Text Reader

Abstract

The invention provides a voice clone model training method and device and electronic equipment, and relates to the technical field of voice synthesis. The method is applied to the electronic equipment, and comprises the following steps: acquiring target voice data comprising an original phoneme sequence and a target acoustic sequence; inputting the original phoneme sequence into a coding module for coding operation to obtain phoneme features; inputting the phoneme features into an acoustic module to obtain initial predicted acoustic features; inputting the phoneme features into a variable information adaptation module to obtain phoneme basic features; obtaining a target prediction acoustic sequence according to the initial prediction acoustic features and the phoneme basic features; and iteratively updating the parameters of the acoustic module and the parameters of the variable information adaptation module according to a difference value between the target prediction acoustic sequence and the target acoustic sequence to obtain a voice clone model. According to the invention, the deviation of the voice clone model prediction data can be reduced, so that the accuracy of the voice clone model prediction data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech synthesis technology, and more specifically, to a training method, a training device, and an electronic device for a speech cloning model in the field of speech synthesis technology. Background Art

[0002] In recent years, with the continuous development of artificial intelligence (AI) technology, voice cloning (VC) has been widely used in scenarios such as voice assistants, digital humans, intelligent customer service, and dubbing for film, television, and animation. Voice cloning, also known as sound cloning or personalized speech synthesis, can mimic the target speaker's timbre and speaking style based on a small number of speech samples, achieving a realistic effect.

[0003] In related technologies, voice cloning is mostly based on adaptive training in a multi-speaker Text To Speech (TTS) framework. However, when cloning the voice of a target speaker, the voice cloning model may be unstable due to factors such as poor recording environment, unclear pronunciation, discontinuous pronunciation, and inconsistent speaking style. As a result, there may be differences between the cloned voice and the target speaker's voice, resulting in poor cloning effect.

[0004] Therefore, how to improve the accuracy of voice cloning model prediction data is an urgent problem that needs to be solved. Summary of the Invention

[0005] The present application provides a training method, training device, and electronic device for a voice cloning model. The method can reduce the deviation of the voice cloning model prediction data, thereby improving the accuracy of the voice cloning model prediction data.

[0006] In a first aspect, a method for training a speech cloning model is provided. The method is applied to an electronic device. The speech cloning model includes an encoding module, an acoustic module, and a variable information adaptation module. The acoustic module is used to predict acoustic features, and the variable information adaptation module is used to extract basic phoneme features. The method includes:

[0007] The method includes obtaining target speech data, wherein the target speech data includes an original phoneme sequence and a target acoustic sequence; inputting the original phoneme sequence into an encoding module for encoding to obtain phoneme features; inputting the phoneme features into an acoustic module to obtain initial predicted acoustic features; inputting the phoneme features into a variable information adaptation module to obtain basic phoneme features; obtaining a target predicted acoustic sequence based on the initial predicted acoustic features and the basic phoneme features; and iteratively updating parameters of the acoustic module and parameters of the variable information adaptation module based on a difference between the target predicted acoustic sequence and the target acoustic sequence to obtain a speech cloning model.

[0008] In an embodiment of the present application, a phoneme sequence in target speech data is encoded by an encoding module to obtain phoneme features, which are respectively input into an acoustic module and a variable information adaptation module to obtain initial predicted acoustic features through the acoustic module and basic phoneme features through the variable information adaptation module; a target predicted acoustic sequence is then obtained based on the initial predicted acoustic features and the basic phoneme features, and the parameters of the acoustic module and the parameters of the variable information adaptation module are iteratively updated based on the difference between the target predicted acoustic sequence and the acoustic sequence in the target speech to obtain a speech cloning model. Because the phoneme features are input into the acoustic module and the variable information adaptation module separately, the output of the acoustic module and the output of the variable information adaptation module are independent of each other. That is, the acoustic module and the variable information adaptation module are arranged in parallel in the speech cloning model. Therefore, the output of the acoustic module and the output of the variable information adaptation module do not affect each other. Therefore, even if the output of the acoustic module has a large deviation, the output of the variable information adaptation module will not be affected by the output deviation of the acoustic module. The deviation of the target predicted acoustic sequence obtained by the output of the acoustic module and the output of the variable information adaptation module can be reduced. In other words, the deviation of the predicted data of the speech cloning model is reduced, thereby improving the accuracy of the predicted data of the speech cloning model.

[0009] Optionally, the phoneme-based features are used to characterize the duration and volume of the original phoneme sequence.

[0010] In conjunction with the first aspect, in certain implementations of the first aspect, iteratively updating parameters of the acoustic module and parameters of the variable information adaptation module based on a difference between a target predicted acoustic sequence and a target acoustic sequence to obtain a speech cloning model includes:

[0011] A decoding operation is performed on the initial predicted acoustic features to obtain an initial predicted acoustic sequence; based on the difference between the initial predicted acoustic sequence and the target acoustic sequence, the parameters of the acoustic module are iteratively updated to obtain a trained acoustic module; based on the difference between the target predicted acoustic sequence and the target acoustic sequence, the parameters of the variable information adaptation module are iteratively updated to obtain a trained variable information adaptation module; and based on the trained acoustic module and the trained variable information adaptation module, a speech cloning model is obtained.

[0012] In an embodiment of the present application, the parameters of the acoustic module are iteratively updated by using the difference value between the initial predicted acoustic sequence and the target acoustic sequence obtained after the initial predicted acoustic feature is decoded, so as to obtain the trained acoustic module; since the difference value is obtained after the initial predicted acoustic feature output by the acoustic module is decoded, the variable information adaptation module is not iteratively trained, that is, the acoustic module is iteratively trained only by using the difference value; and, the parameters of the variable information adaptation module are iteratively updated by using the difference value between the target predicted acoustic sequence and the target acoustic sequence, and the acoustic model is not iteratively trained; thereby realizing step-by-step training of the acoustic module and the variable information adaptation module; that is, the parameters of the acoustic model can be trained first, and then the parameters of the variable information adaptation module can be trained after the parameters of the acoustic model are fixed. Since the iterative training of the acoustic model and the variable information adaptation module are independent of each other, that is, the iterative training process of the acoustic model and the iterative training process of the variable information adaptation module do not affect each other; therefore, by training the acoustic model and the variable information adaptation module separately, the deviation of the speech cloning model's predicted data can be reduced, thereby improving the accuracy of the speech cloning model's predicted data.

[0013] In combination with the first aspect and the above implementations, in certain implementations of the first aspect, the speech cloning model further includes a decoding module and a post-processing module, wherein the post-processing module is configured to process spectral information of acoustic features output by the decoding module. Obtaining a target predicted acoustic sequence based on the initial predicted acoustic features and the basic phoneme features includes:

[0014] The initial predicted acoustic features and the basic phoneme features are input into the decoding module for decoding operation, and a first predicted acoustic sequence is output; the first predicted acoustic sequence is input into the post-processing module, and a second predicted acoustic sequence is output; the first predicted acoustic sequence and the second predicted acoustic sequence are weighted to obtain a target predicted acoustic sequence.

[0015] In an embodiment of the present application, the first predicted acoustic sequence output by the encoding module is processed by a post-processing network to obtain a second predicted acoustic sequence, and the first predicted acoustic sequence and the second predicted acoustic sequence are weighted to obtain a target predicted acoustic sequence; since the spectral information of the acoustic features output by the decoding module can be further processed by the post-processing module, problems such as uneven volume in the spectral information or partial missing of spectral information can be avoided; in addition, after weighted processing of the first predicted acoustic sequence and the second predicted acoustic sequence, the obtained target predicted acoustic sequence is more accurate; therefore, by processing the spectral information of the acoustic features output by the decoding module by the post-processing module, the accuracy of the target predicted acoustic sequence can be further improved.

[0016] In combination with the first aspect and the above implementations, in certain implementations of the first aspect, the voice cloning model further includes a user prosody module, which is configured to screen a prosody feature set to obtain target prosody features that meet preset conditions; and input the original phoneme sequence into the encoding module for encoding to obtain phoneme features, including:

[0017] The original phoneme sequence is input into the encoding module for encoding operation, and the first phoneme feature is output; the first phoneme feature and the target prosodic feature are weighted to obtain the phoneme feature.

[0018] In an embodiment of the present application, the phoneme feature (for example, the first phoneme feature) and the target prosodic feature output by the encoding module are weighted to obtain the phoneme feature; since the target prosodic feature is a prosodic feature that satisfies a preset condition by screening a set of prosodic features, that is, the target prosodic feature is a prosodic feature of better quality; therefore, after weighted processing of the target prosodic feature and the first phoneme feature, the obtained target predicted acoustic sequence can also have better prosody.

[0019] In combination with the first aspect and the above implementations, in some implementations of the first aspect, the user prosody module further includes user identification information corresponding to each prosody feature in the prosody feature set.

[0020] In an embodiment of the present application, the user prosody module may include user identification information corresponding to each prosody feature in the prosody feature set, that is, there is a one-to-one correspondence between the prosody feature and the user identification information, so that the corresponding prosody feature can be obtained through the user identification information, which can avoid errors in the obtained prosody features, thereby improving the accuracy of obtaining the prosody features.

[0021] In combination with the first aspect and the above implementations, in certain implementations of the first aspect, the user prosody module is configured to filter the prosody feature set to obtain target prosody features that meet preset conditions, including:

[0022] The duration and energy value of each rhythmic feature in the rhythmic feature set are calculated; and the rhythmic features whose duration is within a preset duration range and whose energy value is within a preset energy value range are taken as target rhythmic features.

[0023] In an embodiment of the present application, a rhythmic feature in a rhythmic feature set whose duration is within a preset duration range and whose energy value is within a preset energy value range is taken as a target rhythmic feature; since the duration and energy value in the target rhythmic feature are both within the preset range, that is, the target rhythmic feature has a better rhythm; therefore, on the basis that the target rhythmic feature has a better rhythm, the obtained target predicted acoustic sequence can have a better rhythm.

[0024] In combination with the first aspect and the above implementations, in certain implementations of the first aspect, the voice cloning model further includes a voiceprint recognition module; and obtaining target voice data includes:

[0025] Collect at least one voice data of a target user and at least one candidate voice data, wherein the user corresponding to the at least one candidate voice data is different from the target user; calculate the similarity between the acoustic features of each candidate voice data in the at least one candidate voice data and the acoustic features of each voice data in the at least one voice data through a voiceprint recognition module, and determine the target candidate voice data whose acoustic feature similarity is greater than or equal to a preset threshold; and use the target candidate voice data and the at least one voice data as the target voice data.

[0026] In an embodiment of the present application, the similarity between the acoustic features of each candidate voice data and the acoustic features of each voice data of the target user is calculated through the voiceprint recognition module, and the candidate voice data with a similarity greater than or equal to a preset threshold and the voice data of the target user are used as target candidate voice data; since the similarity between the acoustic features of the candidate voice data and the acoustic features of the voice data of the target user is greater than or equal to the preset threshold, it can be said that the acoustic features of the candidate voice data and the acoustic features of the voice data of the target user have a good overlap, when using the target candidate voice data as the target voice data, the data volume of the target user's voice data can be expanded, thereby further improving the accuracy of the target predicted acoustic sequence.

[0027] In a second aspect, a training device for a speech cloning model is provided. The device is configured in an electronic device. The speech cloning model includes an encoding module, an acoustic module, and a variable information adaptation module. The acoustic module is used to predict acoustic features, and the variable information adaptation module is used to extract basic phoneme features. The training device includes: an acquisition module for acquiring target speech data, wherein the target speech data includes an original phoneme sequence and a target acoustic sequence; an encoding module for inputting the original phoneme sequence into the encoding module for encoding to obtain phoneme features; a prediction module for inputting the phoneme features into the acoustic module to obtain initial predicted acoustic features; an extraction module for inputting the phoneme features into the variable information adaptation module to obtain basic phoneme features; a processing module for obtaining a target predicted acoustic sequence based on the initial predicted acoustic features and the basic phoneme features; and a training module for iteratively updating parameters of the acoustic module and parameters of the variable information adaptation module based on the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain the speech cloning model.

[0028] In combination with the second aspect, in certain implementations of the second aspect, the training module is specifically used to: perform a decoding operation on the initial predicted acoustic features to obtain an initial predicted acoustic sequence; iteratively update the parameters of the acoustic module based on the difference between the initial predicted acoustic sequence and the target acoustic sequence to obtain a trained acoustic module; iteratively update the parameters of the variable information adaptation module based on the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain a trained variable information adaptation module; and obtain a speech cloning model based on the trained acoustic module and the trained variable information adaptation module.

[0029] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the processing module is specifically used to: input the initial predicted acoustic features and the basic phoneme features into the decoding module for decoding operations, and output a first predicted acoustic sequence; input the first predicted acoustic sequence into the post-processing module, and output a second predicted acoustic sequence; perform weighted processing on the first predicted acoustic sequence and the second predicted acoustic sequence to obtain a target predicted acoustic sequence.

[0030] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the encoding module is specifically used to: input the original phoneme sequence into the encoding module for encoding operation, and output the first phoneme feature; perform weighted processing on the first phoneme feature and the target prosodic feature to obtain the phoneme feature.

[0031] In combination with the second aspect and the above implementations, in some implementations of the second aspect, the user prosody module further includes user identification information corresponding to each prosody feature in the prosody feature set.

[0032] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the training device also includes a calculation module for calculating the duration and energy value of each rhythmic feature in the rhythmic feature set; and the rhythmic features in each rhythmic feature whose duration is within a preset duration range and whose energy value is within a preset energy value range are taken as target rhythmic features.

[0033] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the acquisition module is specifically used to: collect at least one voice data and at least one candidate voice data of the target user, wherein the user corresponding to the at least one candidate voice data is different from the target user; calculate the similarity between the acoustic features of each candidate voice data in the at least one candidate voice data and the acoustic features of each voice data in the at least one voice data through the voiceprint recognition module, and determine the target candidate voice data whose similarity of acoustic features is greater than or equal to a preset threshold; and use the target candidate voice data and the at least one voice data as the target voice data.

[0034] In a third aspect, an electronic device is provided, comprising a memory and a processor. The memory is configured to store executable program code, and the processor is configured to retrieve and execute the executable program code from the memory, so that the electronic device executes the method of the first aspect or any possible implementation of the first aspect.

[0035] In a fourth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method in the first aspect or any possible implementation of the first aspect.

[0036] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of a scene of voice cloning in related technology.

[0038] Figure 2 This is a structural diagram of a voice cloning model in related technology.

[0039] Figure 3 This is a flow chart of a method for training a voice cloning model provided in an embodiment of the present application.

[0040] Figure 4 Schematic diagram of another voice cloning model provided in an embodiment of the present application.

[0041] Figure 5 This is a structural diagram of another voice cloning model provided in an embodiment of the present application.

[0042] Figure 6 2 is a schematic diagram of the structure of another voice cloning model provided in an embodiment of the present application.

[0043] Figure 7 4 is a flow chart of another method for training a voice cloning model provided in an embodiment of the present application.

[0044] Figure 8 Schematic diagram of the structure of the training device of the voice cloning model provided in the embodiment of the present application.

[0045] Figure 9 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] The following will clearly and thoroughly describe the technical solutions in this application in conjunction with the accompanying drawings. In the description of the embodiments of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more than two.

[0047] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0048] Figure 1 It is a schematic diagram of a scene of voice cloning in related technology.

[0049] For example, Figure 1 As shown, Figure 1 The system includes a terminal 110 and a server 120. The server 120 may include a speech cloning model 121. The terminal 110 may be used as a device for using the speech cloning model 121; and the server 120 may be used as an electronic device for training the speech cloning model.

[0050] The terminal 110 may be used to collect voice data of the user and send the collected voice data to the server 120 .

[0051] The server 120 may be configured to activate the voice cloning model 121 after receiving the voice data, and transmit the voice data to the voice cloning model 121 .

[0052] The voice cloning model 121 can be used to perform training based on received voice data, so that the trained voice cloning model can perform voice cloning to obtain cloned voice.

[0053] For example, the user's voice data can be collected through the terminal 110, and the terminal 110 can send the collected voice data to the server 120. After receiving the voice data, the server 120 can start the voice cloning model 121, so that the voice cloning model 121 can be trained according to the received voice data, so that the trained voice cloning model can perform voice cloning, so that voice cloning can be performed in the terminal 110 to obtain cloned voice.

[0054] It should be noted that the terminal 110 can be an intelligent device capable of collecting user voice data, including but not limited to: a personal computer, a tablet computer, a handheld device, an in-vehicle device, a wearable device, a computing device, or other processing device connected to a wireless modem. Terminal devices can be called different names in different networks, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, terminal device in 5G network or future evolution network, etc.; and the server 120 can be an electronic device capable of training the voice cloning model 121, such as a computer or mobile phone, etc., and the embodiments of the present application are not limited to this.

[0055] Figure 2 This is a structural diagram of a voice cloning model in related technology.

[0056] For example, Figure 2 As shown, the speech cloning model may include a phoneme embedding module 201, a first position encoding module 202, an encoding module 203, an acoustic module 204, a variable information adaptation module 205, a second position encoding module 206, a decoding module 207 and a linear module 208.

[0057] Optionally, when the user's voice data is less, Figure 2 The speech cloning model in the training is trained. As a result, when the trained speech cloning model uses non-training speech data for speech cloning, it may output speech with average timbre or the cloned speech may be unclear.

[0058] Exemplarily, all speech data included in the speech database can be used as candidate speech data, and all candidate speech data can be input into a pre-trained speaker feature encoder (Speaker Encoder), the vectors of the candidate speech data (which can be recorded as "candidate speech vectors") are extracted, and the extracted candidate speech vectors are stored in the speech database.

[0059] Exemplarily, a speaker feature encoder can be used to extract vectors (which can be recorded as "target voice vectors") from all voice data of a target speaker (which can be recorded as a "target user"); the similarity between each candidate voice vector stored in the voice database and the target voice vector is calculated. If the similarity between one or more candidate voice vectors and the target voice vector among the multiple candidate voice vectors is greater than a preset similarity threshold (for example, 95%), it can be said that the candidate voice and the target voice have a good degree of consistency. Then, one or more candidate voice vectors can be added to the voice data training set to serve as training data for training the voice cloning model, thereby increasing the training data of the voice cloning model.

[0060] For example, since speech data is composed of a phoneme sequence (for example, English phonetic symbols or pinyin finals, etc.) and an acoustic feature sequence (for example, tone, stress, pronunciation order and rhythm, etc.), and a predicted acoustic feature sequence can be obtained by inputting the phoneme sequence into the speech cloning model, the accuracy of the predicted acoustic feature sequence can be determined by comparing the predicted acoustic feature sequence with the original acoustic feature sequence in the speech data based on the difference between the predicted acoustic feature sequence and the original acoustic feature sequence.

[0061] refer to Figure 2 , the phoneme sequence in the speech data included in the speech data training set can be input into the phoneme embedding module 201, and the phoneme embedding module 201 performs embedding processing (Embedding) on ​​the phoneme sequence to obtain an embedding vector of the phoneme sequence, so that the embedding vector of the phoneme sequence can represent the tone and rhythm of each phoneme in the phoneme sequence; since the embedding vector of the phoneme sequence output by the phoneme embedding module 201 does not include the position feature of each phoneme in the phoneme sequence, the pronunciation order of each phoneme can be encoded by the first position encoding module 202 to obtain the position vector of the phoneme sequence; the embedding vector of the phoneme sequence and the position vector of the phoneme sequence are weighted and input into the encoding module 203 (also called "phoneme encoder") for encoding processing to obtain the phoneme feature vector of the phoneme sequence, so that the phoneme feature vector can represent the tone, rhythm and pronunciation order of each phoneme.

[0062] Furthermore, the phoneme feature vector output by the encoding module 203 can be input into the acoustic module 204 (also referred to as "acoustic conditional modeling"), and the acoustic module 204 extracts phoneme features for predicting the acoustic feature sequence, such as tone, stress, pronunciation order and rhythm, from the phoneme feature vector, and obtains a predicted initial acoustic feature sequence (which can be recorded as "initial predicted acoustic features") through these phoneme features; since the initial predicted acoustic features do not include the pitch, volume and duration of each phoneme, the initial predicted acoustic features are input into the variable information adaptation module 205, so that the variable information adaptation module 205 extracts the pitch, volume and duration of each phoneme, and adds the pitch, volume and duration of each phoneme to the initial predicted acoustic features to obtain a more complete acoustic feature (which can be recorded as "target predicted acoustic features").

[0063] Optionally, the variable information adaptation module 205 may include a duration extractor for predicting the duration of a phoneme, a pitch predictor for predicting the pitch of a phoneme, and an energy predictor for predicting the volume of a phoneme. The structure of the variable information adaptation module 205 may be set according to actual needs, and the embodiment of the present application is not limited to this.

[0064] Since each phoneme may correspond to multiple speech frames (also known as "speech segments"), the second position encoding module 206 upsamples the multiple speech frames and rearranges them to obtain a vocalization order. The obtained vocalization order is then encoded to obtain a position vector of the upsampled phoneme sequence. The target predicted acoustic features and the upsampled position vector of the phoneme sequence are weighted and input into the decoding module 207 for decoding to obtain the corresponding synthesized spectrum. Furthermore, the obtained synthesized spectrum is input into the linear module 208 through the decoding module 207, and the linear module 208 performs convolution processing on the input synthesized spectrum to obtain a convolved acoustic feature sequence (which can be recorded as the "target predicted acoustic sequence").

[0065] Optionally, the decoding module 207 may adopt a Mel-Spectrum (Mal) decoder, and the synthesized spectrum obtained by the Mal decoder may include a synthesized Mel-Spectrum.

[0066] For example, in order to improve the accuracy of speech cloning, the target predicted acoustic sequence can be compared with the original acoustic feature sequence in the speech data, and the network parameters of each module in the speech cloning model can be adjusted according to the difference between the target predicted acoustic sequence and the original acoustic feature sequence in the speech data, thereby improving the accuracy of the cloned speech obtained by the speech cloning model.

[0067] In related technologies, in order to avoid the influence of the previous level module on the next level module in the speech cloning model, a loss function can be set between the target predicted acoustic sequence and the original acoustic feature sequence in the speech data. The network parameters in each module of the speech cloning model are adjusted using the back propagation algorithm based on the loss function. This allows the speech cloning model after training to generate a target predicted acoustic sequence that is highly consistent with the original acoustic feature sequence in the speech data.

[0068] However, when cloning the voice of a target speaker, the data input into the trained voice cloning model for voice cloning may differ significantly from the training data due to factors such as a poor recording environment of the voice data, fuzzy pronunciation, discontinuous pronunciation, and inconsistent speaking style. This may result in a significant deviation between the voice cloned by the trained voice cloning model and the voice of the target speaker. For example, the cloned voice may have better sound quality but lower similarity to the voice of the target speaker; or the cloned voice may have higher similarity to the voice of the target speaker but lower sound quality; or the cloned voice may have better sound quality and higher similarity to the voice of the target speaker but lower prosody, resulting in poor voice cloning effect.

[0069] In order to solve the above-mentioned problem of poor effect of voice cloning, the present application proposes a training method, a training device and an electronic device for a voice cloning model.

[0070] The following combination Figures 3 to 7 The training method of the voice cloning model provided in the embodiment of the present application is described in detail.

[0071] Figure 3 300 is a flow chart of a method for training a voice cloning model provided in an embodiment of the present application. The method 300 may be executed by an electronic device, for example, Figure 1 Server 120 in.

[0072] For example, Figure 3 As shown, the method 300 includes the following implementation process:

[0073] S310, obtaining target voice data.

[0074] The target speech data includes an original phoneme sequence and a target acoustic sequence; the original phoneme sequence may include English phonetic symbols or pinyin finals of the target speaker's speech data, and the target acoustic sequence may include the tone, stress, articulation order and rhythm of the target speaker's speech data.

[0075] Exemplarily, target speech data including an original phoneme sequence and a target acoustic sequence may be acquired.

[0076] In one possible implementation, at least one voice data and at least one candidate voice data of a target user are collected; the similarity between the acoustic features of each candidate voice data in the at least one candidate voice data and the acoustic features of each voice data in the at least one voice data is calculated through a voiceprint recognition module, and the target candidate voice data whose acoustic feature similarity is greater than or equal to a preset threshold is determined; the target candidate voice data and the at least one voice data are used as the target voice data.

[0077] The user corresponding to at least one candidate voice data is different from the target user. It should be understood that the user corresponding to the candidate voice data is different from the target user. For example, if the target user is user A, the candidate voice data may be user B or user C.

[0078] Exemplarily, before training the voice cloning model, voice data of the user whose voice is to be cloned at the current moment (referred to as the "target user") may be collected, and at least one voice data set may be collected. Furthermore, all voice data collected before the current moment may be used as candidate voice data. A pre-trained speaker feature encoder (referred to as a "voiceprint recognition module") in the voice cloning model, such as an Extended Context-Aware Parallel Aggregation TDNN voiceprint recognition module, may be used to extract a speaker vector from each candidate voice data set, thereby obtaining at least one candidate voice vector and acquiring acoustic features from the at least one candidate voice vector. The output layer (referred to as "FC+BN") of the speaker feature encoder, which is composed of a fully connected layer (FC) and a batch normalization layer (BN), may be used as the output of the speaker vector.

[0079] Exemplarily, a speaker vector may be extracted from each speech data of the target user through a voiceprint recognition module to obtain at least one target speech vector, and acoustic features in the at least one target speech vector may be acquired.

[0080] Furthermore, the voiceprint recognition module can be used to respectively calculate the similarity between the acoustic features of each candidate voice data in at least one candidate voice data and the acoustic features of each voice data in at least one voice data of the target user; if there is one or more candidate voice data whose acoustic features have a similarity greater than or equal to a preset threshold with the acoustic features of any voice data of the target user, then the one or more candidate voice data and the at least one voice data of the target user can be used as target voice data, and the at least one target voice data can be added to the voice data training set to serve as training data for training the voice cloning model.

[0081] It should be noted that in order to ensure that the candidate voice data selected as the target voice data has a high similarity with the voice data of the target user, the preset threshold can be set to a higher value, for example, 97% or 98%, which is not limited in this embodiment of the present application.

[0082] Optionally, the distance between the acoustic features of each candidate speech data in at least one candidate speech data and the acoustic features of each speech data in at least one speech data of the target user can also be calculated by cosine similarity; if the cosine similarity between the acoustic features of one or more candidate speech data and the acoustic features of any speech data of the target user is greater than or equal to a preset cosine threshold, then the one or more candidate speech data can be used as target speech data.

[0083] It should be noted that in order to ensure that the similarity between the candidate voice data screened as the target voice data and the voice data of the target user is high, the preset cosine threshold can be set to a larger value, for example, 0.7 or 0.8, which is not limited in this embodiment of the present application.

[0084] In an embodiment of the present application, the similarity between the acoustic features of each candidate voice data and the acoustic features of each voice data of the target user is calculated through the voiceprint recognition module, and the candidate voice data with a similarity greater than or equal to a preset threshold and the voice data of the target user are used as target candidate voice data; since the similarity between the acoustic features of the candidate voice data and the acoustic features of the voice data of the target user is greater than or equal to the preset threshold, it can be said that the acoustic features of the candidate voice data and the acoustic features of the voice data of the target user have a good overlap, when using the target candidate voice data as the target voice data, the data volume of the target user's voice data can be expanded, thereby further improving the accuracy of the target predicted acoustic sequence. Furthermore, since the similarity between the acoustic features of each candidate voice data and the acoustic features of each voice data of the target user is calculated, the problem of large pronunciation differences between multiple voice data of the target user due to different emotions or different speaking styles can be avoided, which may result in the candidate voice data with smaller acoustic feature similarity to the voice data of the target user among multiple candidate voice data being used as the target candidate voice data, or the candidate voice data with larger acoustic feature similarity to the voice data of the target user among multiple candidate voice data being omitted, thereby causing the accuracy of the screened target candidate voice data to be poor; therefore, by calculating the similarity between the acoustic features of each candidate voice data and the acoustic features of each voice data of the target user, the accuracy of the screened target candidate voice data can be further improved.

[0085] S320: Input the original phoneme sequence into the encoding module for encoding to obtain phoneme features.

[0086] Among them, phoneme features can be used to represent pitch, stress, pronunciation order and rhythm, etc.

[0087] Exemplarily, the phoneme sequence included in the target speech data may be input into an encoding module for encoding, and the encoded phoneme sequence may be used as a phoneme feature.

[0088] Figure 4 Schematic diagram of another voice cloning model provided in an embodiment of the present application.

[0089] For example, Figure 4 As shown, the speech cloning model may include a phoneme embedding module 401, a first position encoding module 402, an encoding module 403, an acoustic module 404, a variable information adaptation module 405, a second position encoding module 406, and a decoding module 407. Figure 4 The voice cloning model shown includes Figure 2 The functions of the repeated modules in the voice cloning model shown are the same. Figure 4 The voice cloning model shown is similar to Figure 2 The voice cloning model shown can be obtained Figure 4 In the illustrated voice cloning model, the connection relationship between the acoustic module 404 and the variable information adaptation module 405 is improved, that is, the acoustic module 404 and the variable information adaptation module 405 are arranged in parallel.

[0090] Exemplarily, the phoneme sequence included in the target speech data (which can be recorded as the "original phoneme sequence") can be input into the phoneme embedding module 401, and the original phoneme sequence is embedded by the phoneme embedding module 401 to obtain an embedding vector of the original phoneme sequence, so that the embedding vector of the original phoneme sequence can represent the tone and rhythm of each phoneme in the original phoneme sequence; since the embedding vector of the original phoneme sequence output by the phoneme embedding module 401 does not include the positional features of each phoneme in the original phoneme sequence, the pronunciation order of each phoneme can be encoded by the first position encoding module 402 to obtain the position vector of the original phoneme sequence; the embedding vector of the original phoneme sequence and the position vector of the original phoneme sequence are weighted and input into the encoding module 403 for encoding processing to obtain the phoneme features of the original phoneme sequence (also referred to as "phoneme feature vectors"), so that the phoneme features of the original phoneme sequence can represent the tone, rhythm and pronunciation order of each phoneme.

[0091] S330: Input the phoneme features into the acoustic module to obtain initial predicted acoustic features.

[0092] Exemplarily, the phoneme features of the original phoneme sequence output by the encoding module 403 can be input into the acoustic module 404, and the acoustic module 404 extracts the phoneme features for predicting the acoustic feature sequence, such as pitch, stress, pronunciation order and rhythm, from the phoneme features of the original phoneme sequence; since the phoneme features such as pitch, stress, pronunciation order and rhythm cannot be obtained by labeling the phoneme features of the original phoneme sequence, it is necessary to generate the initial predicted acoustic features based on the phoneme features of the original phoneme sequence through the acoustic module 404.

[0093] S340: Input the phoneme features into the variable information adaptation module to obtain the phoneme basic features.

[0094] Exemplarily, the phoneme features of the original phoneme sequence output by the encoding module 403 can be input into the variable information adaptation module 405, and the variable information adaptation module 405 extracts the basic phoneme features included in the original phoneme sequence, such as pitch, volume, and duration, from the phoneme features of the original phoneme sequence; since the basic phoneme features such as pitch, volume, and duration can be directly obtained by labeling the phoneme features of the original phoneme sequence, the basic phoneme features extracted by the variable information adaptation module 405 are exactly the same as the basic phoneme features in the target speech data, and there will be no deviation.

[0095] Optionally, the basic phoneme features may also include bass frequency and treble frequency, etc., which is not limited in the embodiment of the present application.

[0096] If it is necessary to explain Figure 4 In the structure of the voice cloning model shown, S330 and S340 are executed. S330 and S340 can be executed simultaneously or at different times. However, the output of the acoustic module 404 and the output of the variable information adaptation module 405 are independent of each other and do not affect each other.

[0097] S350: Obtain a target predicted acoustic sequence based on the initial predicted acoustic features and the phoneme basic features.

[0098] Exemplarily, since each phoneme in the original phoneme sequence may correspond to multiple speech frames, the second position encoding module 406 upsamples the multiple speech frames and rearranges them to obtain a pronunciation order, and the obtained pronunciation order is encoded to obtain a position vector of the upsampled original phoneme sequence; the position vector of the upsampled original phoneme sequence and the initial predicted acoustic features output by the acoustic module 404 and the basic phoneme features output by the variable information adaptation module 405 are weighted and input into the decoding module 407 for decoding operation to obtain the corresponding synthetic spectrum, that is, the target predicted acoustic sequence, and the target predicted acoustic sequence is output by the decoding module 407.

[0099] In one possible implementation, the target predicted acoustic sequence is obtained based on the initial predicted acoustic features and the basic phoneme features, including: inputting the initial predicted acoustic features and the basic phoneme features into a decoding module for decoding operation, and outputting a first predicted acoustic sequence; inputting the first predicted acoustic sequence into a post-processing module, and outputting a second predicted acoustic sequence; and performing weighted processing on the first predicted acoustic sequence and the second predicted acoustic sequence to obtain a target predicted acoustic sequence.

[0100] Figure 5 This is a structural diagram of another voice cloning model provided in an embodiment of the present application.

[0101] For example, Figure 5 As shown, the speech cloning model may include a phoneme embedding module 501, a first position encoding module 502, an encoding module 503, an acoustic module 504, a variable information adaptation module 505, a second position encoding module 506, a decoding module 507 and a post-processing module 508. Figure 5 The voice cloning model shown includes Figure 4 The functions of the repeated modules in the voice cloning model shown are the same. Figure 5 The voice cloning model shown is similar to Figure 4 The voice cloning model shown can be obtained Figure 5 The illustrated speech cloning model includes a post-processing module 508 after the decoding module 507. The post-processing module 508 can be used to perform precision processing on the spectrum information of the acoustic features output by the decoding module 507 to improve the precision of the spectrum information of the acoustic features.

[0102] Optionally, the position vector after upsampling the original phoneme sequence, the initial predicted acoustic features output by the acoustic module 504, and the basic phoneme features output by the variable information adaptation module 505 are weighted and input into the decoding module 507 for decoding, and the corresponding synthesized spectrum is output, which can be recorded as a "first predicted acoustic sequence." Furthermore, the first predicted acoustic sequence output by the decoding module 507 after the decoding operation is input into the post-processing module 508, so that the post-processing module 508 performs precision processing on the input first predicted acoustic sequence to output a second predicted acoustic sequence. By having the post-processing module 508 perform precision processing on the input first predicted acoustic sequence, problems such as uneven volume or partial missing of spectrum information in the spectrum information corresponding to the first predicted acoustic sequence can be avoided.

[0103] Furthermore, the first predicted acoustic sequence output after the decoding operation of the decoding module 507 and the second predicted acoustic sequence output after the precision processing of the input first predicted acoustic sequence by the post-processing module 508 can be weighted to obtain a predicted target predicted acoustic sequence.

[0104] In an embodiment of the present application, the first predicted acoustic sequence output by the encoding module is processed by a post-processing network to obtain a second predicted acoustic sequence, and the first predicted acoustic sequence and the second predicted acoustic sequence are weighted to obtain a target predicted acoustic sequence; since the spectral information of the acoustic features output by the decoding module can be further processed by the post-processing module, problems such as uneven volume in the spectral information or partial missing of spectral information can be avoided; in addition, after weighted processing of the first predicted acoustic sequence and the second predicted acoustic sequence, the obtained target predicted acoustic sequence is more accurate; therefore, by processing the spectral information of the acoustic features output by the decoding module by the post-processing module, the accuracy of the target predicted acoustic sequence can be further improved.

[0105] S360 , iteratively updating the parameters of the acoustic module and the parameters of the variable information adaptation module according to the difference between the target predicted acoustic sequence and the target acoustic sequence, to obtain a speech cloning model.

[0106] For example, the difference between the target predicted acoustic sequence predicted by the speech cloning model and the target acoustic sequence in the target speech data, for example, the similarity difference, can be calculated. The parameters of the acoustic module and the parameters of the variable information adaptation module are iteratively updated based on the calculated difference value, thereby obtaining a trained speech cloning model.

[0107] In one possible implementation, the above-mentioned iterative updating of parameters of the acoustic module and the variable information adaptation module based on the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain the speech cloning model includes: decoding the initial predicted acoustic features to obtain the initial predicted acoustic sequence; iteratively updating the parameters of the acoustic module based on the difference between the initial predicted acoustic sequence and the target acoustic sequence to obtain a trained acoustic module; iteratively updating the parameters of the variable information adaptation module based on the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain a trained variable information adaptation module; and obtaining the speech cloning model based on the trained acoustic module and the trained variable information adaptation module.

[0108] Exemplarily, the initial predicted acoustic features output by the acoustic module 504 can be input separately to the decoding module 507 for decoding operation and then output the corresponding synthetic spectrum; then, the synthetic spectrum is input to the post-processing module 508, so that the post-processing module 508 performs precision processing on the input synthetic spectrum to output the initial predicted acoustic sequence. It should be understood that since the variable information adaptation module 505 does not participate in the calculation before the output of the acoustic module 504 during the training of the acoustic module 504, nor does it involve the iterative update of the acoustic module 504, the variable information adaptation module 505 does not participate in the training process of the acoustic module 504, and thus does not affect the parameters of the acoustic module 504; and the parameters of the phoneme embedding module 501, the first position encoding module 502, the encoding module 503 and the post-processing module 508 are fixed.

[0109] Furthermore, the parameters of the acoustic module 504 can be iteratively updated based on the difference between the initial predicted acoustic sequence and the target acoustic sequence in the target speech data to obtain a trained acoustic module, and the parameters in the trained acoustic module can be stored and fixed.

[0110] For example, the difference between the target predicted acoustic sequence predicted by the speech cloning model and the target acoustic sequence in the target speech data can be calculated, and the parameters of the variable information adaptation module 505 can be iteratively updated based on the difference to obtain a trained variable information adaptation module, and the parameters in the trained variable information adaptation module can be stored and fixed. It should be understood that since the decoding module 507 and the post-processing module 508 do not participate in the calculation before the output of the variable information adaptation module 505 during the training of the variable information adaptation module 505, nor are they involved in the iterative update of the variable information adaptation module 505, that is, the decoding module 507 and the post-processing module 508 do not participate in the training process of the variable information adaptation module 505, and thus will not affect the parameters of the variable information adaptation module 505. Moreover, since the training of the acoustic module 504 is completed, the parameters of the acoustic module 504 are fixed; and the parameters of the phoneme embedding module 501, the first position encoding module 502, and the encoding module 503 are fixed.

[0111] Furthermore, a trained speech cloning model can be obtained through the trained acoustic module and the trained variable information adaptation module.

[0112] In an embodiment of the present application, the parameters of the acoustic module are iteratively updated by using the difference value between the initial predicted acoustic sequence and the target acoustic sequence obtained after the initial predicted acoustic feature is decoded, so as to obtain the trained acoustic module; since the difference value is obtained after the initial predicted acoustic feature output by the acoustic module is decoded, the variable information adaptation module is not iteratively trained, that is, the acoustic module is iteratively trained only by using the difference value; and, the parameters of the variable information adaptation module are iteratively updated by using the difference value between the target predicted acoustic sequence and the target acoustic sequence, and the acoustic model is not iteratively trained; thereby realizing step-by-step training of the acoustic module and the variable information adaptation module; that is, the parameters of the acoustic model can be trained first, and then the parameters of the variable information adaptation module can be trained after the parameters of the acoustic model are fixed. Since the iterative training of the acoustic model and the variable information adaptation module are independent of each other, that is, the iterative training process of the acoustic model and the iterative training process of the variable information adaptation module do not affect each other; therefore, by training the acoustic model and the variable information adaptation module separately, the deviation of the speech cloning model's predicted data can be reduced, thereby improving the accuracy of the speech cloning model's predicted data.

[0113] Optionally, after obtaining the trained acoustic module and the trained variable information adaptation module, the decoding module 507 can be iteratively updated based on the difference between the target predicted acoustic sequence predicted by the speech cloning model and the target acoustic sequence in the target speech data to obtain a trained decoding module. The parameters of the trained decoding module are stored and fixed. Furthermore, the parameters of the phoneme embedding module 501, the first position encoding module 502, the encoding module 503, the acoustic module 504, the variable information adaptation module 505, the second position encoding module 506, the decoding module 507, and the post-processing module 508 remain fixed.

[0114] Furthermore, a trained voice cloning model can be obtained through the trained acoustic module, the trained variable information adaptation module and the trained decoding module.

[0115] In an embodiment of the present application, a phoneme sequence in target speech data is encoded by an encoding module to obtain phoneme features, which are respectively input into an acoustic module and a variable information adaptation module to obtain initial predicted acoustic features through the acoustic module and basic phoneme features through the variable information adaptation module; a target predicted acoustic sequence is then obtained based on the initial predicted acoustic features and the basic phoneme features, and the parameters of the acoustic module and the parameters of the variable information adaptation module are iteratively updated based on the difference between the target predicted acoustic sequence and the acoustic sequence in the target speech to obtain a speech cloning model. Because the phoneme features are input into the acoustic module and the variable information adaptation module separately, the output of the acoustic module and the output of the variable information adaptation module are independent of each other. That is, the acoustic module and the variable information adaptation module are arranged in parallel in the speech cloning model. Therefore, the output of the acoustic module and the output of the variable information adaptation module do not affect each other. Therefore, even if the output of the acoustic module has a large deviation, the output of the variable information adaptation module will not be affected by the output deviation of the acoustic module. The deviation of the target predicted acoustic sequence obtained by the output of the acoustic module and the output of the variable information adaptation module can be reduced. In other words, the deviation of the predicted data of the speech cloning model is reduced, thereby improving the accuracy of the predicted data of the speech cloning model.

[0116] Figure 6 2 is a schematic diagram of the structure of another voice cloning model provided in an embodiment of the present application.

[0117] For example, Figure 6 As shown, the voice cloning model may include a phoneme embedding module 601, a first position encoding module 602, an encoding module 603, an acoustic module 604, a variable information adaptation module 605, a second position encoding module 606, a decoding module 607, a post-processing module 608, and a speaker embedding layer 609 (also referred to as a "user prosody module"). Figure 6 The voice cloning model shown includes Figure 5 The functions of the repeated modules in the voice cloning model shown are the same. Figure 6 The voice cloning model shown is Figure 5 The voice cloning model shown can be obtained Figure 6 The voice cloning model shown in the figure adds a speaker embedding layer 609 after the encoding module 603. The speaker embedding layer 609 can be used to filter the prosodic feature set to obtain target prosodic features that meet preset conditions.

[0118] For example, if the target user's voice data has poor prosodic coherence, poor consistency, hesitant pronunciation, too fast or too slow speaking speed, or uneven volume, this may cause significant discrepancies between the target acoustic sequence predicted by the voice cloning model and the target acoustic sequence in the target voice data. To address these issues, the prosodic features of the target voice data can be replaced with better prosodic features from TTS.

[0119] In one possible implementation, the user prosody module is used to screen a set of prosody features to obtain target prosody features that meet preset conditions, calculate the duration and energy value of each prosody feature in the prosody feature set; and take the prosody features whose duration is within a preset duration range and whose energy value is within a preset energy value range as target prosody features.

[0120] For example, the parameters in the speaker embedding layer 609 may be trained to obtain parameters of prosodic duration and energy value with better prosodic features, and the prosodic duration and energy value may be used to replace the prosodic duration and energy in the target speech data.

[0121] Optionally, the duration and energy value can be calculated using formula (1) and formula (2):

[0122] d=(d0-μ0)×σ / σ0+μ Formula (1)

[0123] e=(e0-μ0)×σ / σ0+μ Formula (2)

[0124] Wherein, d0 included in formula (1) can represent the duration value predicted by the duration replacement network; μ0 and σ0 can respectively represent the duration mean and standard deviation of the target prosodic feature counted in TTS; μ and σ can respectively represent the duration mean and standard deviation of the prosodic feature counted in multiple speech data of the target user. e0 included in formula (2) can represent the energy value predicted by the energy replacement network; μ0 and σ0 can respectively represent the energy mean and standard deviation of the target prosodic feature counted in TTS; μ and σ can respectively represent the energy mean and standard deviation of the target prosodic feature counted in multiple speech data of the target user.

[0125] It should be noted that d0 and e0 can be obtained through the trained variable information adaptation module, and different target users may correspond to different d0 and e0, which is not limited in the embodiment of the present application.

[0126] Furthermore, the calculated duration and the preset duration range, as well as the energy value and the preset energy value range are compared respectively, and the rhythmic features whose duration is within the preset duration range, and whose energy value is within the preset energy value range, are obtained to replace the duration and energy value in multiple voice data of the target user, so that the target voice data has better rhythm, that is, the rhythmic features whose duration is within the preset duration range, and whose energy value is within the preset energy value range are used as the target rhythmic features.

[0127] It should be noted that the preset duration range and the preset energy value range can be set according to actual conditions, and the embodiments of the present application do not limit this.

[0128] Optionally, if the target user needs to preserve the prosody of his or her own voice in the cloned voice, the prosody features in the target user's own voice data may be directly used.

[0129] In an embodiment of the present application, a rhythmic feature in a rhythmic feature set whose duration is within a preset duration range and whose energy value is within a preset energy value range is taken as a target rhythmic feature; since the duration and energy value in the target rhythmic feature are both within the preset range, that is, the target rhythmic feature has a better rhythm; therefore, on the basis that the target rhythmic feature has a better rhythm, the obtained target predicted acoustic sequence can have a better rhythm.

[0130] Optionally, the user prosody module further includes user identification information corresponding to each prosody feature in the prosody feature set.

[0131] Among them, user identification information can be represented by identity information (ID), and a corresponding ID can be set for each prosodic feature corresponding to each target voice data in the voice database. For example, the ID corresponding to the target prosodic feature is set to "5"; and the ID corresponding to the prosodic features in the voice data of different users is set differently.

[0132] In an embodiment of the present application, the user prosody module may include user identification information corresponding to each prosody feature in the prosody feature set, that is, there is a one-to-one correspondence between the prosody feature and the user identification information, so that the corresponding prosody feature can be obtained through the user identification information, which can avoid errors in the obtained prosody features, thereby improving the accuracy of obtaining the prosody features.

[0133] In one possible implementation, the above-mentioned inputting the original phoneme sequence into the encoding module for encoding operation to obtain phoneme features includes: inputting the original phoneme sequence into the encoding module for encoding operation, outputting the first phoneme feature; weighting the first phoneme feature and the target prosodic feature to obtain the phoneme feature.

[0134] Exemplarily, the phoneme sequence included in the target speech data can be input into the phoneme embedding module 601, and the original phoneme sequence is embedded by the phoneme embedding module 601 to obtain the embedding vector of the original phoneme sequence, and the pronunciation order of each phoneme is encoded by the first position encoding module 602 to obtain the position vector of the original phoneme sequence; the embedding vector of the original phoneme sequence and the position vector of the original phoneme sequence are weighted and input into the encoding module 603 for encoding to obtain the first phoneme feature (i.e., the phoneme feature of the original phoneme sequence); and the first phoneme feature is weighted with the target prosody feature input through the speaker embedding layer 609 to obtain the weighted phoneme feature. Then, the phoneme feature is processed as follows: Figure 5 The steps shown are used to obtain the target predicted acoustic sequence.

[0135] For example, the trained voice cloning model can be used to clone the voice data of the target user to obtain a cloned voice that has a high similarity to the voice data of the target user.

[0136] In an embodiment of the present application, the phoneme feature (for example, the first phoneme feature) and the target prosodic feature output by the encoding module are weighted to obtain the phoneme feature; since the target prosodic feature is a prosodic feature that satisfies a preset condition by screening a set of prosodic features, that is, the target prosodic feature is a prosodic feature of better quality; therefore, after weighted processing of the target prosodic feature and the first phoneme feature, the obtained target predicted acoustic sequence can also have better prosody.

[0137] Optionally, to enable the voice cloning model to accurately generate the target user's cloned voice, a speaker feature encoder or a speaker feature encoding network in the TTS can be used to obtain the target user's speaker vector. However, if the TTS does not include a speaker feature encoding network, a pre-set ID can be used directly as the target user's ID. There is a one-to-one correspondence between the pre-set ID and the target user, meaning that different target users have different IDs.

[0138] It should be noted that the pre-set ID can be "0" or "9", and there is no duplication between multiple IDs, which is not limited in the embodiment of the present application.

[0139] Figure 7 4 is a flow chart of another method for training a voice cloning model provided in an embodiment of the present application.

[0140] For example, Figure 7 As shown, the method 700 includes the following implementation process:

[0141] S701: Collect at least one voice data and at least one candidate voice data of a target user.

[0142] Exemplarily, one or more speech data of the target user may be collected, and one or more candidate speech data may be collected from a speech database.

[0143] S702: Calculate the similarity between the acoustic features of each candidate voice data and the acoustic features of each voice data of the target user through the voiceprint recognition module, and determine the target candidate voice data whose acoustic feature similarity is greater than or equal to a preset threshold.

[0144] Exemplarily, the voiceprint recognition module can calculate the similarity between the acoustic features of each candidate voice data and the acoustic features of each voice data of the target user, and use the candidate voice data with a similarity greater than or equal to a preset threshold as the target candidate voice data.

[0145] S703: Use the target candidate voice data and at least one voice data as target voice data.

[0146] Exemplarily, the filtered target candidate voice data and one or more voice data of the target user may be used as the target voice data.

[0147] S704: Input the original phoneme sequence into the encoding module for encoding, and output a first phoneme feature.

[0148] refer to Figure 6 The phoneme sequence included in the target speech data can be input into the phoneme embedding module 601, and the original phoneme sequence is embedded by the phoneme embedding module 601 to obtain the embedding vector of the original phoneme sequence, and the pronunciation order of each phoneme is encoded by the first position encoding module 602 to obtain the position vector of the original phoneme sequence; the embedding vector of the original phoneme sequence and the position vector of the original phoneme sequence are weighted and input into the encoding module 603 for encoding processing to obtain the first phoneme feature.

[0149] S705 : Perform weighted processing on the first phoneme feature and the target prosodic feature to obtain a phoneme feature.

[0150] refer to Figure 6 The first phoneme feature and the target prosodic feature input through the speaker embedding layer 609 can be weighted to obtain the weighted phoneme feature.

[0151] S706: Input the phoneme features into the acoustic module to obtain initial predicted acoustic features.

[0152] refer to Figure 6 The phoneme features output by the encoding module 603 can be input to the acoustic module 604, so that the acoustic module 604 generates initial predicted acoustic features according to the phoneme features of the original phoneme sequence.

[0153] S707: Input the phoneme features into the variable information adaptation module to obtain the phoneme basic features.

[0154] refer to Figure 6 The phoneme features output by the encoding module 603 may be input to the variable information adaptation module 605 so as to extract the phoneme basic features from the phoneme features through the variable information adaptation module 605 .

[0155] It should be noted that S706 and S707 may be executed simultaneously or at different times; however, the execution processes of S706 and S707 are independent of each other and do not affect each other.

[0156] S708: Input the initial predicted acoustic features and the phoneme basic features into the encoding module for encoding, and output a first predicted acoustic sequence.

[0157] refer to Figure 6 , multiple speech frames can be upsampled and rearranged to obtain a pronunciation order through the second position encoding module 606, and the obtained pronunciation order is encoded to obtain the position vector of the upsampled original phoneme sequence; then, the position vector of the upsampled original phoneme sequence and the initial predicted acoustic features output by the acoustic module 604 and the basic phoneme features output by the variable information adaptation module 605 are weighted and input into the decoding module 607 for decoding operation to output the first predicted acoustic sequence.

[0158] S709: Input the first predicted acoustic sequence into a post-processing module, and output a second predicted acoustic sequence.

[0159] refer to Figure 6 , can be input to the post-processing module 608, so that the post-processing module 608 performs precision processing on the input first predicted acoustic sequence to output a second predicted acoustic sequence.

[0160] S710 , performing weighted processing on the first predicted acoustic sequence and the second predicted acoustic sequence to obtain a target predicted acoustic sequence.

[0161] refer to Figure 6 , weighted processing is performed on the first predicted acoustic sequence output by the decoding module 607 and the second predicted acoustic sequence output by the post-processing module 608 to obtain a predicted target predicted acoustic sequence.

[0162] Optionally, S711, S712 and S713 are used as Figure 3 Specifically, in S711, decoding is performed on the initial predicted acoustic features to obtain an initial predicted acoustic sequence.

[0163] refer to Figure 6The initial predicted acoustic features output by the acoustic module 604 can be separately input to the decoding module 607 for decoding operation, and the initial predicted acoustic sequence is output after the decoding operation of the initial predicted acoustic features is completed by the decoding module 607.

[0164] S712 , iteratively updating the parameters of the acoustic module according to the difference between the initial predicted acoustic sequence and the target acoustic sequence to obtain a trained acoustic module.

[0165] refer to Figure 6 , the parameters of the acoustic module 604 can be iteratively updated according to the difference between the initial predicted acoustic sequence and the target acoustic sequence in the target speech data to obtain a trained acoustic module.

[0166] S713 , iteratively updating the parameters of the variable information adaptation module according to the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain a trained variable information adaptation module.

[0167] refer to Figure 6 , the difference between the target predicted acoustic sequence predicted by the speech cloning model and the target acoustic sequence in the target speech data can be calculated, and the parameters of the variable information adaptation module 605 can be iteratively updated according to the difference value to obtain a trained variable information adaptation module.

[0168] S714: Obtain a speech cloning model based on the trained acoustic module and the trained variable information adaptation module.

[0169] refer to Figure 6 , can be obtained through the trained acoustic module and the trained variable information adaptation module. Figure 6 The voice cloning model shown has the function of cloning voice based on the voice data of the target user.

[0170] It should be understood that the above examples are intended to help those skilled in the art understand the embodiments of the present application, and are not intended to limit the embodiments of the present application to the specific numerical values ​​or specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or variations based on the above examples, and such modifications or variations also fall within the scope of the embodiments of the present application.

[0171] Combined with the above Figures 1 to 7 The training method of the voice cloning model provided by the embodiment of the present application is described in detail; Figure 8 and Figure 9 The device embodiments of the present application are described in detail. It should be understood that the devices in the embodiments of the present application can execute the various methods of the aforementioned embodiments of the present application, that is, the specific working processes of the following various products can refer to the corresponding processes in the aforementioned method embodiments.

[0172] Figure 8 This is a schematic diagram of the structure of a training device for a speech cloning model provided in an embodiment of the present application. The device is configured in an electronic device. The speech cloning model includes an encoding module, an acoustic module, and a variable information adaptation module. The acoustic module is used to predict acoustic features, and the variable information adaptation module is used to extract basic phoneme features.

[0173] For example, Figure 8 As shown, the training device 800 includes:

[0174] Acquisition module 810: used to acquire target speech data, wherein the target speech data includes an original phoneme sequence and a target acoustic sequence;

[0175] Coding module 820: used to input the original phoneme sequence into the coding module for encoding operation to obtain phoneme features;

[0176] Prediction module 830: used to input the phoneme features into the acoustic module to obtain initial predicted acoustic features;

[0177] Extraction module 840: used to input the phoneme features into the variable information adaptation module to obtain the phoneme basic features;

[0178] Processing module 850: used to obtain a target predicted acoustic sequence based on the initial predicted acoustic features and the phoneme basic features;

[0179] Training module 860: used to iteratively update the parameters of the acoustic module and the parameters of the variable information adaptation module according to the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain a speech cloning model.

[0180] In one possible implementation, the training module 860 is specifically used to: decode the initial predicted acoustic features to obtain an initial predicted acoustic sequence; iteratively update the parameters of the acoustic module based on the difference between the initial predicted acoustic sequence and the target acoustic sequence to obtain a trained acoustic module; iteratively update the parameters of the variable information adaptation module based on the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain a trained variable information adaptation module; and obtain a speech cloning model based on the trained acoustic module and the trained variable information adaptation module.

[0181] In one possible implementation, the processing module 850 is specifically used to: input the initial predicted acoustic features and the basic phoneme features into the decoding module for decoding operations, and output a first predicted acoustic sequence; input the first predicted acoustic sequence into the post-processing module, and output a second predicted acoustic sequence; and perform weighted processing on the first predicted acoustic sequence and the second predicted acoustic sequence to obtain a target predicted acoustic sequence.

[0182] In a possible implementation, the encoding module 820 is specifically configured to: input the original phoneme sequence into the encoding module for encoding, and output a first phoneme feature; and perform weighted processing on the first phoneme feature and the target prosodic feature to obtain a phoneme feature.

[0183] In a possible implementation, the user prosody module further includes user identification information corresponding to each prosody feature in the prosody feature set.

[0184] Optionally, the training device 800 further includes a calculation module for calculating the duration and energy value of each rhythmic feature in the rhythmic feature set; and taking the rhythmic features whose duration is within a preset duration range and whose energy value is within a preset energy value range as target rhythmic features.

[0185] In one possible implementation, the acquisition module 810 is specifically used to: collect at least one voice data of a target user and at least one candidate voice data, wherein the user corresponding to the at least one candidate voice data is different from the target user; calculate the similarity between the acoustic features of each candidate voice data in the at least one candidate voice data and the acoustic features of each voice data in the at least one voice data through the voiceprint recognition module, and determine the target candidate voice data whose acoustic feature similarity is greater than or equal to a preset threshold; and use the target candidate voice data and the at least one voice data as the target voice data.

[0186] It should be noted that the above-mentioned device 800 is embodied in the form of a functional module. The term "module" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.

[0187] For example, a "module" may be a software program, a hardware circuit, or a combination of the two that implements the above-described functions. The hardware circuit may include an application-specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., a shared processor, a dedicated processor, or a group processor, etc.) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functions.

[0188] Therefore, the modules of each example described in the embodiments of this application can be implemented with electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0189] Figure 9 It is a structural diagram of an electronic device provided in an embodiment of the present application.

[0190] For example, Figure 9 As shown, the electronic device 900 includes: a memory 910 and a processor 920, wherein the memory 910 stores an executable program code 9101, and the processor 920 is used to call and execute the executable program code 9101 to perform a training method for a voice cloning model.

[0191] This application can divide the electronic device into functional modules based on the above-mentioned method examples. For example, each functional module can be mapped to a specific functional module, or two or more functions can be integrated into a processing module. The above-mentioned integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative and is only a logical functional division. In actual implementation, other division methods may be used.

[0192] When the functional modules are divided according to their functions, the electronic device may include an acquisition module, an encoding module, a prediction module, an extraction module, a processing module, and a training module. It should be noted that all relevant content of the steps involved in the above method embodiments can be referred to in the functional descriptions of the corresponding functional modules and will not be repeated here.

[0193] The electronic device provided in this application is used to execute the above-mentioned training method of a voice cloning model, and thus can achieve the same effect as the above-mentioned implementation method.

[0194] In the case of an integrated unit, the electronic device may include a processing module and a storage module. The processing module may be used to control and manage the operation of the electronic device, and the storage module may be used to support the electronic device in executing mutual program codes and data.

[0195] The processing module may be a processor or controller that implements or executes the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processing (DSP) and a microprocessor, and the storage module may be a memory.

[0196] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the aforementioned methods. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD (Digital Video Disc), a CD-ROM (Compact Disc Read-Only Memory), a microdrive, a magneto-optical disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), an EPROM (Erasable Programmable Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a DRAM (Dynamic Random Access Memory), a VRAM (Video Random Access Memory), a flash memory device, a magnetic or optical card, a nanosystem (including a molecular memory IC), or any other type of medium or device suitable for storing instructions and / or data.

[0197] The present application also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the above-mentioned related steps to implement a training method of a voice cloning model in the above-mentioned embodiment.

[0198] In addition, the electronic device provided in the embodiments of the present application may specifically be a chip, component, or module, and the electronic device may include a connected processor and memory; wherein the memory is used to store instructions, and when the electronic device is running, the processor may call and execute the instructions to enable the chip to perform a training method for a voice cloning model in the above embodiment.

[0199] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0200] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0201] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0202] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for training a voice cloning model, characterized in that: The method is applied to an electronic device. The voice cloning model includes an encoding module, an acoustic module, and a variable information adaptation module. The acoustic module is used to predict acoustic features, and the variable information adaptation module is used to extract basic phoneme features. The training method includes: Acquiring target speech data, wherein the target speech data includes an original phoneme sequence and a target acoustic sequence; Inputting the original phoneme sequence into the encoding module for encoding to obtain phoneme features; Inputting the phoneme features into the acoustic module to obtain initial predicted acoustic features; Inputting the phoneme features into the variable information adaptation module to obtain the phoneme basic features; Obtaining a target predicted acoustic sequence according to the initial predicted acoustic features and the phoneme basic features; According to the difference value between the target predicted acoustic sequence and the target acoustic sequence, the parameters of the acoustic module and the parameters of the variable information adaptation module are iteratively updated to obtain the speech cloning model.

2. The method according to claim 1, characterized in that The iterative updating of the parameters of the acoustic module and the parameters of the variable information adaptation module according to the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain the speech cloning model includes: Performing a decoding operation on the initial predicted acoustic features to obtain an initial predicted acoustic sequence; Iteratively updating the parameters of the acoustic module according to the difference between the initial predicted acoustic sequence and the target acoustic sequence to obtain a trained acoustic module; Iteratively updating the parameters of the variable information adaptation module according to the difference between the target predicted acoustic sequence and the target acoustic sequence to obtain a trained variable information adaptation module; The speech cloning model is obtained according to the trained acoustic module and the trained variable information adaptation module.

3. The method according to claim 1, characterized in that The speech cloning model further includes a decoding module and a post-processing module. The post-processing module is used to process the spectrum information of the acoustic features output by the decoding module. The target predicted acoustic sequence is obtained based on the initial predicted acoustic features and the phoneme basic features, including: Inputting the initial predicted acoustic features and the phoneme basic features into the decoding module for decoding, and outputting a first predicted acoustic sequence; Inputting the first predicted acoustic sequence into the post-processing module and outputting a second predicted acoustic sequence; The first predicted acoustic sequence and the second predicted acoustic sequence are weighted to obtain the target predicted acoustic sequence.

4. The method according to claim 3, characterized in that The voice cloning model further includes a user prosody module, which is used to screen the prosody feature set to obtain target prosody features that meet preset conditions; The step of inputting the original phoneme sequence into the encoding module for encoding to obtain phoneme features includes: Inputting the original phoneme sequence into the encoding module to perform the encoding operation and output a first phoneme feature; The first phoneme feature and the target prosodic feature are weighted to obtain the phoneme feature.

5. The method according to claim 4, characterized in that The user prosody module further includes user identification information corresponding to each prosody feature in the prosody feature set.

6. The method according to claim 4, characterized in that The user prosody module is used to screen the prosody feature set to obtain target prosody features that meet preset conditions, including: Calculating the duration and energy value of each rhythmic feature in the rhythmic feature set; The rhythmic features whose duration is within a preset duration range and whose energy value is within a preset energy value range are used as the target rhythmic features.

7. The method according to any one of claims 1 to 6, characterized in that The voice cloning model also includes a voiceprint recognition module; the step of obtaining target voice data includes: Collecting at least one voice data of a target user and at least one candidate voice data, wherein the user corresponding to the at least one candidate voice data is different from the target user; Calculating, by the voiceprint recognition module, the similarity between the acoustic features of each candidate voice data in the at least one candidate voice data and the acoustic features of each voice data in the at least one voice data, and determining target candidate voice data whose acoustic feature similarity is greater than or equal to a preset threshold; The target candidate voice data and the at least one voice data are used as the target voice data.

8. A training device for a speech cloning model, characterized in that: The device is configured in an electronic device. The voice cloning model includes an encoding module, an acoustic module, and a variable information adaptation module. The acoustic module is used to predict acoustic features, and the variable information adaptation module is used to extract basic phoneme features. The training device includes: An acquisition module, configured to acquire target speech data, wherein the target speech data includes an original phoneme sequence and a target acoustic sequence; An encoding module, configured to input the original phoneme sequence into the encoding module for encoding to obtain phoneme features; A prediction module, configured to input the phoneme features into the acoustic module to obtain initial predicted acoustic features; An extraction module, configured to input the phoneme features into the variable information adaptation module to obtain the phoneme basic features; a processing module, configured to obtain a target predicted acoustic sequence according to the initial predicted acoustic features and the phoneme basic features; A training module is configured to iteratively update parameters of the acoustic module and parameters of the variable information adaptation module according to a difference value between the target predicted acoustic sequence and the target acoustic sequence, so as to obtain the speech cloning model.

9. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable program code; A processor, configured to call and run the executable program code from the memory, so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 is implemented.