Speech conversion model training methods, speech conversion methods, devices and media

By using a pre-defined masking strategy and an adversarial network to decouple the speech features in the speech conversion model, the problem of insufficient robustness of speech feature decoupling is solved, thereby improving training efficiency and model robustness.

CN115171666BActive Publication Date: 2026-04-07PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In the training process of existing speech conversion models, the decoupling of speech features is not robust enough, resulting in low training efficiency.

Method used

Speech sample features are extracted by an encoder, and decoupled using a pre-set masking strategy and a pre-set adversarial network. Adversarial loss and speech reconstruction loss are calculated to optimize the parameters of the speech conversion model.

Benefits of technology

This improves the training robustness and efficiency of the speech conversion model, ensuring the accuracy and consistency of speech features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171666B_ABST
    Figure CN115171666B_ABST
Patent Text Reader

Abstract

This application relates to the field of speech conversion technology, and provides a speech conversion model training method, speech conversion method, apparatus, and medium. The method includes: extracting speech sample features from preset speech samples using an encoder; then decoupling the speech samples based on a preset masking strategy to obtain sample feature representations; inputting the sample feature representations into a generator to reconstruct the Mel spectrogram of the speech samples based on the sample feature representations to obtain the Mel spectrogram of the target sample; calculating the speech reconstruction loss of the speech conversion model based on the Mel spectrogram of the target sample and the original sample Mel spectrograms corresponding to the preset speech samples; and optimizing the parameters in the speech conversion model based on adversarial loss and speech reconstruction loss to obtain a trained speech conversion model. By decoupling the speech sample features through a preset masking strategy and a preset adversarial network, the robustness of the speech conversion model is improved, thereby increasing training efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech conversion technology, and in particular to a speech conversion model training method, a speech conversion method, a speech conversion model training device, a speech conversion device, a computer device, and a storage medium. Background Technology

[0002] Speech conversion involves altering the source speaker's voice to sound like the target speaker's voice while preserving the linguistic information.

[0003] In the training process of existing speech conversion models, the deentanglement algorithms used, such as random resampling and temporary bottleneck layer size, are used to deentangle speech features. However, this method is difficult to ensure robust speech feature decoupling, which affects the entire training process and results in low training efficiency of speech conversion models. Summary of the Invention

[0004] This application provides a speech conversion model training method to address the problem of low training efficiency in existing speech conversion model training schemes.

[0005] A first aspect of this application provides a speech conversion model training method, the speech conversion model training method comprising:

[0006] The encoder extracts speech sample features from preset speech samples; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features.

[0007] The speech sample features are decoupled based on a preset masking strategy and a preset adversarial network to obtain a sample feature representation, and the adversarial loss in the decoupling process is calculated; the sample feature representation is used to characterize the enhanced speech sample features.

[0008] The sample feature representation is input into the generator to generate the Mel-spectrum map of the target sample;

[0009] The speech reconstruction loss is calculated based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample.

[0010] The parameters in the speech conversion model are optimized based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model.

[0011] A second aspect of this application provides a speech conversion method, including:

[0012] Extract speech information from the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information, and pitch information.

[0013] The speech information is input into a trained speech conversion model for speech conversion to obtain a target Mel spectrogram; wherein the trained speech conversion model is trained using the speech conversion model training method described above.

[0014] The target Mel spectrogram is converted into a waveform using a preset algorithm to obtain synthesized speech.

[0015] A third aspect of this application provides a speech conversion model training apparatus, the speech conversion model training apparatus comprising:

[0016] Extraction module: used to extract speech sample features from preset speech samples through an encoder; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features;

[0017] Decoupling module: used to decouple the speech sample features based on a preset masking strategy and a preset adversarial network, obtain the sample feature representation, and calculate the adversarial loss in the decoupling process; the sample feature representation is used to characterize the enhanced speech sample features;

[0018] Reconstruction module: used to input the sample feature representation into the generator to generate the Mel spectrum of the target sample;

[0019] Calculation module: used to calculate speech reconstruction loss based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample;

[0020] Training module: used to optimize the parameters in the speech conversion model based on the adversarial loss and the speech reconstruction loss, so as to obtain a trained speech conversion model.

[0021] A fourth aspect of this application provides a speech conversion device, including:

[0022] Extraction module: used to extract speech information from the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information and pitch information;

[0023] The first conversion module is used to input the speech information into a trained speech conversion model for speech conversion to obtain a target Mel spectrogram; wherein the trained speech conversion model is trained using the speech conversion model training method described above.

[0024] The second conversion module is used to convert the target Mel spectrogram into a waveform using a preset algorithm to obtain synthesized speech.

[0025] A fifth aspect of this application provides a computer device including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the above-described speech conversion model training method, or when the processor executes the computer-readable instructions, it implements the above-described speech conversion method.

[0026] A sixth aspect of this application provides one or more readable storage media storing computer-readable instructions, which, when executed by one or more processors, implement the speech conversion model training method described above, or, when executed by one or more processors, implement the speech conversion method described above.

[0027] This application provides a speech conversion model training method. An encoder extracts speech sample features from preset speech samples, including content features, timbre features, rhythm features, and pitch features. Then, the speech samples are decoupled based on a preset masking strategy, and speech enhancement is applied to the speech sample features to obtain a sample feature representation. Adversarial loss during the decoupling process is calculated to minimize distortion of the speech sample features, resulting in a more accurate representation. This aims to overcome the problem of feature mismatch when inputting the sample feature representation into the generator, which affects the robustness of the speech conversion model training. The decoupled sample feature representation is input into the generator, and the generator is trained to reconstruct the Mel spectrogram of the speech sample based on the sample feature representation, obtaining the target sample Mel spectrogram. The speech reconstruction loss of the speech conversion model is calculated based on the target sample Mel spectrogram and the original sample Mel spectrogram corresponding to the preset speech sample. The parameters in the speech conversion model are optimized based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model. By decoupling the speech sample features through a preset masking strategy and a preset adversarial network, distortion of the speech sample features is reduced, improving the robustness of the speech conversion model training and thus increasing training efficiency. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a schematic diagram of the application environment of the speech conversion model training method or the speech conversion method in the embodiments of this application;

[0030] Figure 2This is a schematic diagram illustrating the implementation process of the speech conversion model training method in the embodiments of this application;

[0031] Figure 3 This is an example diagram of a speech conversion model for the speech conversion model training method in the embodiments of this application;

[0032] Figure 4 This is an example diagram of the decoupled network for the speech conversion model training method in the embodiments of this application;

[0033] Figure 5 This is a schematic diagram illustrating the implementation process of the speech conversion method in the embodiments of this application;

[0034] Figure 6 This is a schematic diagram of the structure of the speech conversion model training device in the embodiments of this application;

[0035] Figure 7 This is a schematic diagram of the speech conversion device in the embodiments of this application;

[0036] Figure 8 This is a schematic diagram of a computer device in an embodiment of this application. Detailed Implementation

[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0038] Please see Figure 1 , Figure 1 This illustration shows an application environment diagram of the speech conversion model training method in an embodiment of this application, such as... Figure 1As shown, the user terminal can input and upload preset voice samples or voice information of the source speaker and target speaker, and the server can perform voice conversion model training and voice conversion. Alternatively, the user terminal, which includes a processor and computer storage media, can perform voice conversion model training and voice conversion. The user terminal includes, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be a standalone server or a server cluster composed of multiple servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. User terminals from different business systems can interact with the server simultaneously or with a specific server in the server cluster.

[0039] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0040] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0041] Please see Figure 2 , Figure 2 The diagram shown is a flowchart illustrating the implementation of the speech conversion model training method in this application embodiment, demonstrating the application of this method in... Figure 1 Taking the server in the example, the following steps are included:

[0042] The speech conversion model in this application embodiment includes an encoder, a decoupling network, and a generator. The decoupling network includes a preset masking strategy and a preset adversarial network.

[0043] S11: Extract speech sample features from preset speech samples using an encoder; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features.

[0044] In step S11, the encoder includes, but is not limited to, a content encoder, a timbre encoder, a rhythm encoder, and a pitch encoder.

[0045] In this embodiment, the content encoder is used to recognize the input speech and perform text conversion to obtain speech content information unrelated to the speaker. Automatic Speech Recognition (ASR) or similar technologies can be used as the content encoder. The timbre encoder is used to extract timbre features, and its input is a timbre vector. The rhythm encoder's input is speech containing a large amount of information; therefore, various kinds of disorganized information may be encoded into the rhythm features. Therefore, the input speech can be preprocessed before rhythm feature extraction to filter out rhythm-irrelevant information and improve the accuracy of rhythm feature extraction. The pitch encoder's input is the pitch contour information of the speech, or fundamental frequency information.

[0046] As an example, please refer to Figure 3 , Figure 3 The diagram shown is an example of a speech conversion model for the speech conversion model training method provided in this application embodiment. Figure 1 As shown, one-hot encoding is used to obtain the timbre vector of a preset speech sample. Then, the timbre variable is input into the timbre encoder, and deep learning is performed on the input timbre vector to obtain the sample timbre features. The timbre encoder can use word-embedding encoding to achieve deep learning of the input timbre vector. One-hot encoding generates one-hot labels for different preset speech samples by encoding 0-1 according to the number of speech samples in the training corpus. For example, if there are three preset speech samples 1, 2, and 3, the one-hot label of the first preset speech sample is

[100] , the one-hot label of the second preset speech sample is

[010] , and the one-hot label of the third preset speech sample is

[001] . The one-hot labels corresponding to the preset speech samples are used as the timbre vectors of the preset speech samples and input into the timbre encoder. In addition, before inputting the speech samples into the content encoder and pitch encoder, it is necessary to randomly sample the input speech samples and pitch contours in advance to improve the accuracy of the sample features during training.

[0047] S12: Decouple the speech sample features based on a preset masking strategy and a preset adversarial network to obtain a sample feature representation, and calculate the adversarial loss during the decoupling process; the sample feature representation is used to characterize the enhanced speech sample features.

[0048] In step S12, the decoupling network includes a preset masking strategy and a preset adversarial network. Decoupling refers to learning to separate speech sample features through adversarial training, thereby enhancing the speech sample features.

[0049] In this embodiment, the preset masking strategy randomly masks any feature in the speech sample features generated by the encoder using a random mask. The preset adversarial network is used to infer the masked feature based on other unmasked speech sample features, while simultaneously inversely stimulating the encoder to generate more accurate speech sample features containing fewer irrelevant features. By configuring the preset masking strategy and the preset adversarial network, speech sample features are separated through adversarial training, thereby enhancing the speech sample features. This avoids the difficulty in ensuring the robustness of decoupling, which is often achieved through random resampling and temporary bottleneck layer size adjustment, thus affecting the robustness of speech conversion model training.

[0050] As an embodiment of this application, the preset adversarial network includes a prediction layer and a gradient inversion layer; the process of decoupling the speech sample features based on the preset masking strategy and the preset adversarial network to obtain a sample feature representation and calculating the adversarial loss during the decoupling process includes: generating a random mask based on the preset masking strategy; the random mask is used to randomly mask one of the sample content features, sample timbre features, sample rhythm features, and sample pitch features, so that the prediction layer predicts the masked sample features based on the other three sample features besides the masked sample features; calculating the adversarial loss based on the random mask and the speech sample features; and decoupling the speech sample features based on the gradient inversion layer and the adversarial loss to obtain a sample feature representation.

[0051] In this embodiment, 0 represents masked sample features and 1 represents unmasked sample features. Following the aforementioned rules, the random mask includes (0, 1, 1, 1), (1, 0, 1, 1), (1, 1, 0, 1), and (1, 1, 1, 0). The random mask randomly masks one of the following sample features encoded by the encoder: sample content features, sample timbre features, sample rhythm features, and sample pitch features. The prediction layer in the adversarial network predicts the masked sample features based on the other three sample features. Then, the adversarial loss is calculated based on the random mask and the speech sample features. This adversarial loss is backpropagated to the encoder through a gradient inversion layer, encouraging the encoder to learn speech sample features containing as little mutual information as possible.

[0052] As an example, please refer to Figure 4 , Figure 4The diagram shows an example of the decoupled network in the speech conversion model training method provided in this application. The prediction layer of the adversarial network includes a fully connected layer, an activation function, layer normalization, and another fully connected layer. The gradient of the adversarial network is inverted by the gradient inversion layer before backpropagation to the encoder, encouraging the encoder to learn speech sample features containing as little mutual information as possible. By using a decoupled network with random mask prediction, speech sample features are separated, improving the robustness of multi-factor highly controllable style transfer during the training of the speech conversion model.

[0053] As an embodiment of this application, the calculation of the adversarial loss based on the random mask and the speech sample features includes:

[0054] The adversarial loss is calculated using the following formula:

[0055] L adv =||(1-M)·(Z-MAP(M·Z)||,

[0056] Where Z = (Z r Z c Z f Z u ), M∈(0,1,1,1),(1,0,1,1),(1,1,0,1),(1,1,1,0);

[0057] In the formula, L adv This refers to the adversarial loss; M refers to the random mask; Z refers to the adversarial loss. r This refers to the rhythmic characteristics of the samples, Z. c Z refers to the features of the sample content. f This refers to the pitch characteristics of the sample, Z. u This refers to the timbre characteristics of the sample; Z refers to Z. r Z c Z f Z u The concatenated vector; MAP refers to Mean Average Precision. It should be noted that MAP is an abbreviation for Mean Average Precision. It is used as a metric for measuring detection accuracy in object descriptions. The calculation formula is: MAP = Sum of the average precisions of all categories divided by the total number of categories.

[0058] S13: Input the sample feature representation into the generator to generate the Mel spectrum of the target sample.

[0059] In step S13, the sample feature representation exists in vector form, including but not limited to sample content representation, sample timbre representation, sample rhythm representation, and sample pitch representation.

[0060] In this embodiment, the decoupled sample content representation, sample timbre representation, sample rhythm representation, and sample pitch representation are extracted. These sample feature representations are then input into a generator for feature fusion to obtain a fused vector. Based on the characteristics of the Mel-frequency coefficients, this fused vector is decoded to obtain the Mel-frequency spectrogram of the target sample. It should be noted that the sample content representation, sample timbre representation, sample rhythm representation, and sample pitch representation can be representations of the same dimension or vectors of different dimensions. By fusing features from these representations, a higher-dimensional vector can be obtained. For example, if the sample content representation is a 128-dimensional vector, the sample timbre representation is a 64-dimensional vector, the sample rhythm representation is a 32-dimensional vector, and the sample pitch representation is a 32-dimensional vector, feature fusion yields a 512-dimensional fused vector.

[0061] S14: Calculate the speech reconstruction loss based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample.

[0062] In step S14, the original sample Mel spectrogram is obtained by passing a Mel spectrogram filter through the original speech sample features of the input preset speech sample.

[0063] In this embodiment, due to the changes in the features of the speech samples after encoder and adversarial training in the speech conversion model, the Mel spectrogram of the target sample synthesized by the speech conversion model also differs from the Mel spectrogram of the original sample. This difference is represented by speech reconstruction loss.

[0064] As an embodiment of this application, the step of calculating the speech reconstruction loss based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample includes:

[0065] The speech reconstruction loss is calculated using the following formula:

[0066]

[0067] In the formula, L recon This refers to the speech reconstruction loss; S refers to the Mel-spectrum of the original sample. This refers to the Mel spectrum of the target sample.

[0068] S15: Optimize the parameters in the speech conversion model based on the adversarial loss and the speech reconstruction loss to obtain the trained speech conversion model.

[0069] In step S15, the adversarial loss refers to the error between the sample feature representation and the speech sample features generated during the decoupling process of the speech sample features. The speech reconstruction loss refers to the error between the Mel spectrum of the target sample generated by the generator based on the sample feature representation and the Mel spectrum of the original sample of the corresponding input preset speech sample. The adversarial loss and the speech reconstruction loss constitute the model loss function of the speech conversion model.

[0070] In this embodiment, weights are assigned to the adversarial loss and the speech reconstruction loss respectively. The speech conversion model is trained based on the adversarial loss and the speech reconstruction loss. The parameters in the speech conversion model are optimized, and the weight values ​​of the two are adjusted so that the value of the model loss function can meet the model convergence condition, thus obtaining the trained speech conversion model.

[0071] As an embodiment of this application, the step of optimizing the parameters in the speech conversion model based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model includes:

[0072] The model loss is calculated using the following formula:

[0073] L=α*L adv +β*L recon ,

[0074] In the formula, L refers to the model loss, α refers to the weight of the adversarial loss, and β refers to the weight of the speech reconstruction loss. The values ​​of α and β are both in the range of [0, 1].

[0075] When the model loss reaches the preset convergence condition, the speech conversion model converges, and a trained speech conversion model is obtained. The preset convergence condition can be a specific value or a range of values. The magnitude or value of the convergence condition can be customized, with the aim of minimizing the model loss and improving the accuracy of the speech conversion model's data output.

[0076] This application provides a method for training a speech conversion model. The speech conversion model includes an encoder, a decoupling network, and a generator. The decoupling network includes a preset masking strategy and a preset adversarial network. The encoder extracts speech sample features from preset speech samples, including sample content features, sample timbre features, sample rhythm features, and sample pitch features. Then, based on the preset masking strategy, the speech samples are decoupled, and speech enhancement is performed on the speech sample features to obtain a sample feature representation. The adversarial loss during the decoupling process is calculated to minimize the distortion of the speech sample features, resulting in a more accurate sample feature representation. This aims to overcome the problem of feature mismatch after inputting the sample feature representation into the generator, which affects the robustness of the speech conversion model training. The decoupled sample feature representation is input into the generator, and the generator is trained to reconstruct the Mel spectrogram of the speech sample based on the sample feature representation to obtain the Mel spectrogram of the target sample. Based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample, the speech reconstruction loss of the speech conversion model is calculated. The parameters in the speech conversion model are optimized based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model. By decoupling speech sample features through a pre-defined masking strategy and a pre-defined adversarial network, the distortion of speech sample features is reduced, the robustness of speech conversion model training is improved, and thus the training efficiency is increased.

[0077] Please see Figure 5 , Figure 5 The diagram shown is a flowchart illustrating the implementation of a speech conversion method provided in an embodiment of this application, demonstrating the application of this method in... Figure 1 Taking the server in the example, the following steps are included:

[0078] S21: Extract the speech information of the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information and pitch information.

[0079] In step S21, the source speaker's voice information is also the voice information to be converted. When it is necessary to convert the source speaker's voice, the source speaker's voice is used as the voice to be converted.

[0080] In this embodiment, before converting the source speaker's speech, it is necessary to obtain the speech information of the source speaker and the target speaker. Specifically, complete or partial audio can be extracted from video files, audio files, etc., as the source speaker's speech or the target speaker's speech information. The speech information includes, but is not limited to, speech content information, timbre information, rhythm information, and pitch information.

[0081] S22: Input the speech information into the trained speech conversion model to perform speech conversion and obtain the target Mel spectrogram; wherein, the trained speech conversion model is trained using the above-mentioned speech conversion model training method.

[0082] In step S22, the target Mel spectrogram refers to the Mel spectrogram of the new speech obtained after speech conversion by the trained speech conversion model.

[0083] In this embodiment, inputting speech information into a trained speech conversion model to obtain a target Mel spectrogram includes: inputting the source speaker's speech information into the content encoder of the trained speech conversion model to extract content features unrelated to the speaker; inputting the target speaker's speech information into the timbre encoder, rhythm encoder, and pitch encoder of the trained speech conversion model to extract the target speaker's timbre, rhythm, and pitch features; and generating a target Mel spectrogram based on the above content features and the target speaker's timbre, rhythm, and pitch features using the trained speech conversion model. As one implementation, before inputting the target speaker's speech information into the timbre encoder, rhythm encoder, and pitch encoder of the trained speech conversion model, the target speaker's speech information can be preprocessed, for example, by extracting the pitch contour of the target speaker's speech information and randomly sampling the pitch contour before inputting it into the pitch encoder. This application does not limit the method of preprocessing.

[0084] S23: The target Mel spectrogram is converted into a waveform using a preset algorithm to obtain synthesized speech.

[0085] In step S23, the preset algorithm includes, but is not limited to, the Griffin_lim algorithm.

[0086] In this embodiment, the Griffin_lim algorithm works as follows: a phase spectrum is randomly initialized, and a new speech waveform is synthesized by inverse Fourier transform of the phase spectrum and the known target Mel spectrum. The synthesized speech is then subjected to short-time Fourier transform to obtain a new amplitude spectrum and a new phase spectrum. The known target Mel spectrum and the new phase spectrum are then synthesized by inverse Fourier transform of the synthesized speech. This process is repeated multiple times until the synthesized speech achieves a satisfactory effect.

[0087] This embodiment provides a speech conversion method that improves the conversion effect by adding the conversion of the target speaker's rhythm and pitch features during the speech conversion process, ensuring that the prosody of the source speaker and the target speaker remains consistent after the speech conversion. Furthermore, based on the decoupled speech representation network of adversarial learning in the trained speech conversion model, the content representation of the source speaker's speech and the timbre, rhythm, and pitch representation of the target speaker are extracted, improving the robustness of multi-factor, highly controllable style transfer during the speech conversion process.

[0088] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0089] In one embodiment, a speech conversion model training device 600 is provided, which corresponds one-to-one with the speech conversion model training method in the above embodiments. For example... Figure 6 As shown, the speech conversion model training device includes an extraction module 601, a decoupling module 602, a reconstruction module 603, a calculation module 604, and a training module 605. Detailed descriptions of each functional module are as follows:

[0090] Extraction module 601: used to extract speech sample features from preset speech samples through an encoder; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features;

[0091] Decoupling module 602: used to decouple the speech sample features based on a preset masking strategy and a preset adversarial network, obtain a sample feature representation, and calculate the adversarial loss in the decoupling process; the sample feature representation is used to characterize the enhanced speech sample features;

[0092] Reconstruction module 603: used to input the sample feature representation into the generator to generate the Mel spectrum of the target sample;

[0093] Calculation module 604: used to calculate speech reconstruction loss based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample;

[0094] Training module 605: used to optimize the parameters in the speech conversion model based on the adversarial loss and the speech reconstruction loss, so as to obtain a trained speech conversion model.

[0095] In one embodiment, a voice conversion device 700 is also provided, which corresponds one-to-one with the voice conversion methods in the above embodiments. For example... Figure 7As shown, the speech conversion device includes an extraction module 701, a first conversion module 702, and a second conversion module 703. Detailed descriptions of each functional module are as follows:

[0096] Extraction module 701: used to extract speech information of the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information and pitch information;

[0097] First conversion module 702: used to input the speech information into a trained speech conversion model for speech conversion to obtain a target Mel spectrogram; wherein, the trained speech conversion model is trained using the above-mentioned speech conversion model training method;

[0098] The second conversion module 703 is used to convert the target Mel spectrogram into a waveform using a preset algorithm to obtain synthesized speech.

[0099] Specific limitations regarding the speech conversion model training device can be found in the limitations regarding the speech conversion model training method described above, and specific limitations regarding the speech conversion device can be found in the limitations regarding the speech conversion method described above; they will not be repeated here. The aforementioned speech conversion model training device and each module within the speech conversion device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0100] In one embodiment, a computer device is provided, which may be a server. The computer device includes a processor, memory, a network interface, and a database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a readable storage medium and internal memory. The readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The database of the computer device stores data involved in a speech conversion model training method. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer-readable instructions are executed by the processor, a speech conversion model training method is implemented. The readable storage medium provided in this embodiment includes non-volatile readable storage media and volatile readable storage media.

[0101] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8As shown, the computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a readable storage medium and internal memory. The non-volatile storage medium stores the operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The network interface is used to communicate with an external server via a network connection. When the computer-readable instructions are executed by the processor, they implement a speech conversion model training method. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.

[0102] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor, when executing the computer-readable instructions, implements:

[0103] A method for training a speech conversion model includes:

[0104] The encoder extracts speech sample features from preset speech samples; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features.

[0105] The speech sample features are decoupled based on a preset masking strategy and a preset adversarial network to obtain a sample feature representation, and the adversarial loss in the decoupling process is calculated; the sample feature representation is used to characterize the enhanced speech sample features.

[0106] The sample feature representation is input into the generator to generate the Mel-spectrum map of the target sample;

[0107] The speech reconstruction loss is calculated based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample.

[0108] The parameters in the speech conversion model are optimized based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model.

[0109] And a speech conversion method, comprising:

[0110] Extract speech information from the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information, and pitch information;

[0111] The speech information is input into a trained speech conversion model for speech conversion to obtain a target Mel spectrogram; wherein the trained speech conversion model is trained using the speech conversion model training method described above.

[0112] The target Mel spectrogram is converted into a waveform using a preset algorithm to obtain synthesized speech.

[0113] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. The readable storage media stores computer-readable instructions, which are implemented when executed by one or more processors.

[0114] A method for training a speech conversion model includes:

[0115] The encoder extracts speech sample features from preset speech samples; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features.

[0116] The speech sample features are decoupled based on a preset masking strategy and a preset adversarial network to obtain a sample feature representation, and the adversarial loss in the decoupling process is calculated; the sample feature representation is used to characterize the enhanced speech sample features.

[0117] The sample feature representation is input into the generator to generate the Mel-spectrum map of the target sample;

[0118] The speech reconstruction loss is calculated based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample.

[0119] The parameters in the speech conversion model are optimized based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model.

[0120] And a speech conversion method, comprising:

[0121] Extract speech information from the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information, and pitch information;

[0122] The speech information is input into a trained speech conversion model for speech conversion to obtain a target Mel spectrogram; wherein the trained speech conversion model is trained using the speech conversion model training method described above.

[0123] The target Mel spectrogram is converted into a waveform using a preset algorithm to obtain synthesized speech.

[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0125] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0126] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for training a speech conversion model, characterized in that, The speech conversion model training method includes: The encoder extracts speech sample features from preset speech samples; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features. The speech sample features are decoupled based on a preset masking strategy and a preset adversarial network to obtain a sample feature representation, and the adversarial loss in the decoupling process is calculated; the sample feature representation is used to characterize the enhanced speech sample features. The sample feature representation is input into the generator to generate the Mel-spectrum map of the target sample; The speech reconstruction loss is calculated based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample. The parameters in the speech conversion model are optimized based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model. The preset adversarial network includes a prediction layer and a gradient inversion layer; the process of decoupling the speech sample features based on the preset masking strategy and the preset adversarial network to obtain the sample feature representation and calculating the adversarial loss during the decoupling process includes: A random mask is generated based on the preset masking strategy; the random mask is used to randomly mask one of the sample features, sample timbre features, sample rhythm features and sample pitch features, so that the prediction layer predicts the masked sample features based on the other three sample features besides the masked sample features. The adversarial loss is calculated based on the random mask and the speech sample features; Based on the gradient inverse layer and the adversarial loss, the speech sample features are decoupled to obtain the sample feature representation.

2. The speech conversion model training method according to claim 1, characterized in that, The calculation of the adversarial loss based on the random mask and the speech sample features includes: The adversarial loss is calculated using the following formula: , in, , ; In the formula, This refers to the aforementioned resistance loss; This refers to the random mask; This refers to the rhythmic characteristics of the samples. This refers to the characteristics of the sample content. This refers to the pitch characteristics of the sample. This refers to the timbre characteristics of the sample; It means , , , The concatenated vector; MAP refers to the mean average precision.

3. The speech conversion model training method according to claim 2, characterized in that, The step of calculating the speech reconstruction loss based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample includes: The speech reconstruction loss is calculated using the following formula: , In the formula, This refers to speech reconstruction loss; This refers to the Mel spectrum of the original sample; This refers to the Mel spectrum of the target sample.

4. The speech conversion model training method according to any one of claims 3, characterized in that, The process of optimizing the parameters in the speech conversion model based on the adversarial loss and the speech reconstruction loss to obtain a trained speech conversion model includes: The model loss is calculated using the following formula: , In the formula, This refers to model loss. This refers to the weight of the resistance loss. This refers to the weights of the speech reconstruction loss. , The values ​​of all are in the range of [0, 1]. When the model loss reaches the preset convergence condition, the speech conversion model converges, and the trained speech conversion model is obtained.

5. A speech conversion method, characterized in that, The speech conversion method includes: Extract speech information from the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information, and pitch information; The speech information is input into a trained speech conversion model for speech conversion to obtain a target Mel spectrogram; wherein the trained speech conversion model is trained using the speech conversion model training method as described in any one of claims 1-4; The target Mel spectrogram is converted into a waveform using a preset algorithm to obtain synthesized speech.

6. A speech conversion model training device, characterized in that, The speech conversion model training device includes: Extraction module: used to extract speech sample features from preset speech samples through an encoder; the speech sample features include sample content features, sample timbre features, sample rhythm features, and sample pitch features; Decoupling module: used to decouple the speech sample features based on a preset masking strategy and a preset adversarial network, obtain the sample feature representation, and calculate the adversarial loss in the decoupling process; the sample feature representation is used to characterize the enhanced speech sample features; Reconstruction module: used to input the sample feature representation into the generator to generate the Mel spectrum of the target sample; Calculation module: used to calculate speech reconstruction loss based on the Mel spectrogram of the target sample and the Mel spectrogram of the original sample corresponding to the preset speech sample; Training module: used to optimize the parameters in the speech conversion model based on the adversarial loss and the speech reconstruction loss, so as to obtain a trained speech conversion model; The preset adversarial network includes a prediction layer and a gradient inversion layer; the process of decoupling the speech sample features based on the preset masking strategy and the preset adversarial network to obtain the sample feature representation and calculating the adversarial loss during the decoupling process includes: A random mask is generated based on the preset masking strategy; the random mask is used to randomly mask one of the sample features, sample timbre features, sample rhythm features and sample pitch features, so that the prediction layer predicts the masked sample features based on the other three sample features besides the masked sample features. The adversarial loss is calculated based on the random mask and the speech sample features; Based on the gradient inverse layer and the adversarial loss, the speech sample features are decoupled to obtain the sample feature representation.

7. A voice conversion device, characterized in that, The speech conversion device includes: Extraction module: used to extract speech information from the source speaker and the target speaker; the speech information includes speech content information, timbre information, rhythm information and pitch information; First conversion module: used to input the speech information into a trained speech conversion model for speech conversion to obtain a target Mel spectrogram; wherein, the trained speech conversion model is trained using the speech conversion model training method as described in any one of claims 1-4; The second conversion module is used to convert the target Mel spectrogram into a waveform using a preset algorithm to obtain synthesized speech.

8. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and running on the processor, characterized in that, When the computer-readable instructions are executed by the processor, they implement the speech conversion model training method as described in any one of claims 1-4, or the computer-readable instructions are executed by the processor to implement the speech conversion method as described in claim 5.

9. One or more readable storage media, characterized in that, The readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the speech conversion model training method as described in any one of claims 1-4, or, when executed by a processor, implement the speech conversion method as described in claim 5.

Citation Information

Patent Citations

  • Air traffic control speech recognition method and device for small number of labeled samples

    CN111785257A

  • Voice conversion model training method, voice conversion method, device and medium

    CN116959465A