Speech conversion model training method, speech conversion method and device

By extracting the content characteristics and global features of the audio training data, and combining with the Mel spectrogram processing technology, the speech conversion model parameters are adjusted, and the problem of speech conversion models in the existing technology is difficult to take into account the tone similarity, noise robustness and expressiveness, and a better speech conversion effect is achieved.

CN120148485AActive Publication Date: 2025-06-13NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510045796.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-13
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing voice conversion models are difficult to balance tone similarity, noise robustness and expressiveness, resulting in poor voice conversion effects.

Method used

By obtaining multiple audio training data, extracting content features and global features, and combining the local mask processing and noise-added processing of the Mel spectrogram, input the speech conversion model to be trained, and adjusting the model parameters to obtain the trained speech conversion model.

Benefits of technology

The voice conversion model is realized in the balance between tone similarity, noise robustness and expressiveness, and the voice conversion effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148485A_ABST
    Figure CN120148485A_ABST
Patent Text Reader

Abstract

The invention discloses a voice conversion model training method and device, a voice conversion method and device, electronic equipment and a computer readable storage medium. The training method comprises the following steps: acquiring multiple pieces of audio training data, and extracting first feature training data and second feature training data; obtaining a Mel spectrogram corresponding to the audio training data, and obtaining a mask Mel picture and a noise Mel spectrogram corresponding to the Mel spectrogram; inputting the first feature training data, the second feature training data, the mask Mel spectrogram and the noise Mel spectrogram corresponding to the audio training data into a to-be-trained voice conversion model to obtain a predicted Mel spectrogram; and according to the Mel spectrogram and the predicted Mel spectrogram, performing model parameter adjustment on a to-be-trained voice conversion model to obtain a trained voice conversion model. According to the method, the technical problem that the voice conversion effect is poor due to the fact that tone similarity, noise robustness and expressive force cannot be taken into consideration in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to a method for training a voice conversion model, a voice conversion method, an apparatus thereof, an electronic device, and a computer-readable storage medium. Background Art

[0002] Voice conversion is a technology that changes the timbre of a speaker while keeping the content information unchanged, and is widely used in multiple fields. Usually, based on a voice conversion model, the source voice of a source speaker is converted into a target voice with the timbre of a target speaker.

[0003] Currently, voice conversion models are mainly trained using ASR (Automatic Speech Recognition) model features or SSL (Self Supervised Learning) model features. ASR model features remove most of the paralinguistic information, non-linguistic information, and noise information in the source voice, enabling the voice conversion model to achieve a high timbre similarity and noise robustness, but the output target voice often lacks expressiveness. SSL model features retain almost all the information in the source voice, which can effectively improve the expressiveness of the target voice. However, since there are often potential background noises in the SSL model features, the noise robustness of the voice conversion model is greatly reduced.

[0004] Therefore, there is an urgent need for a voice conversion model that can both ensure high timbre similarity and noise robustness and improve the expressiveness of the target voice, so as to solve the technical problem that the existing voice conversion models cannot balance timbre similarity, noise robustness, and expressiveness, resulting in poor voice conversion effects. Summary of the Invention

[0005] The present application provides a method for training a voice conversion model, a voice conversion method, an apparatus thereof, an electronic device, and a computer-readable storage medium, so as to solve the technical problem that the existing technology cannot balance timbre similarity, noise robustness, and expressiveness, resulting in poor voice conversion effects.

[0006] In a first aspect, an embodiment of the present application provides a method for training a voice conversion model, and the method further includes: obtaining a plurality of audio training data, and extracting first feature training data and second feature training data from the audio training data; wherein, the first feature training data is used to characterize the content features corresponding to the audio training data, and the second feature training data is used to characterize the global features corresponding to the audio training data; obtaining a Mel spectrogram corresponding to the audio training data, and performing local masking processing and noise addition processing on the Mel spectrogram to obtain a masked Mel spectrogram and a noisy Mel spectrogram corresponding to the Mel spectrogram; inputting the first feature training data, the second feature training data, the masked Mel spectrogram, and the noisy Mel spectrogram corresponding to the audio training data into a voice conversion model to be trained, and obtaining a predicted Mel spectrogram output by the voice conversion model to be trained; adjusting model parameters of the voice conversion model to be trained according to the Mel spectrogram and the predicted Mel spectrogram corresponding to the audio training data, and obtaining a trained voice conversion model.

[0007] In a second aspect, an embodiment of the present application provides a voice conversion method, and the method includes: in response to a voice conversion instruction, obtaining first audio data of a speaker to be converted and sample audio data of a specified speaker; extracting first feature data and second feature data from the first audio data, and extracting third feature data and fourth feature data from the sample audio data; wherein, the first feature data is used to characterize the content features corresponding to the first audio data, the second feature data is used to characterize the global features corresponding to the first audio data, the third feature data is used to characterize the content features corresponding to the sample audio data, and the fourth feature data is used to characterize the global features corresponding to the sample audio data; splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data; converting the first audio data of the speaker to be converted into second audio data of the specified speaker according to the fifth feature data, the sixth feature data, a first Mel spectrogram corresponding to the sample audio data, and a trained voice conversion model; wherein, the trained voice conversion model is obtained by training according to the method for training a voice conversion model.

[0008] In a third aspect, an embodiment of the present application provides a training device for a voice conversion model. The device further includes: a first data processing unit, a second data processing unit, a model training unit, and a model parameter adjustment unit. The first data processing unit is configured to obtain a plurality of audio training data, and extract first feature training data and second feature training data from the audio training data. Among them, the first feature training data is used to represent the content features corresponding to the audio training data, and the second feature training data is used to represent the global features corresponding to the audio training data. The second data processing unit is configured to obtain a Mel spectrogram corresponding to the audio training data, and perform local masking processing and noise addition processing on the Mel spectrogram to obtain a masked Mel spectrogram and a noisy Mel spectrogram corresponding to the Mel spectrogram. The model training unit is configured to input the first feature training data, the second feature training data, the masked Mel spectrogram, and the noisy Mel spectrogram corresponding to the audio training data into a voice conversion model to be trained, and obtain a predicted Mel spectrogram output by the voice conversion model to be trained. The model parameter adjustment unit is configured to adjust the model parameters of the voice conversion model to be trained according to the Mel spectrogram and the predicted Mel spectrogram corresponding to the audio training data, and obtain a trained voice conversion model.

[0009] In a fourth aspect, an embodiment of the present application provides a voice conversion device. The device further includes: a data acquisition unit, a feature extraction unit, a feature splicing unit, and a data conversion unit. The data acquisition unit is configured to obtain first audio data of a speaker to be converted and sample audio data of a specified speaker in response to a voice conversion instruction. The feature extraction unit is configured to extract first feature data and second feature data from the first audio data, and extract third feature data and fourth feature data from the sample audio data. Among them, the first feature data is used to represent the content features corresponding to the first audio data, the second feature data is used to represent the global features corresponding to the first audio data, the third feature data is used to represent the content features corresponding to the sample audio data, and the fourth feature data is used to represent the global features corresponding to the sample audio data. The feature splicing unit is configured to splice the first feature data and the third feature data to generate fifth feature data, and splice the second feature data and the fourth feature data to generate sixth feature data. The data conversion unit is configured to convert the first audio data of the speaker to be converted into second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model. Among them, the trained voice conversion model is obtained by training according to the training method of the voice conversion model.

[0010] Fifth aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to implement the above method.

[0011] Sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which one or more computer instructions are stored, and when the instructions are executed by a processor, the above method is executed.

[0012] Compared with the prior art, the training method of the voice conversion model provided by the present application includes: obtaining a plurality of audio training data, and extracting first feature training data and second feature training data from the audio training data; wherein, the first feature training data is used to represent the content features corresponding to the audio training data, and the second feature training data is used to represent the global features corresponding to the audio training data; obtaining the Mel spectrogram corresponding to the audio training data, and performing local masking processing and noise addition processing on the Mel spectrogram to obtain the masked Mel spectrogram and the noisy Mel spectrogram corresponding to the Mel spectrogram; inputting the first feature training data, the second feature training data, the masked Mel spectrogram, and the noisy Mel spectrogram corresponding to the audio training data into the voice conversion model to be trained, and obtaining the predicted Mel spectrogram output by the voice conversion model to be trained; according to the Mel spectrogram and the predicted Mel spectrogram corresponding to the audio training data, adjusting the model parameters of the voice conversion model to be trained to obtain the trained voice conversion model. First, this method uses the content features corresponding to the audio training data as a training data item, so that the voice conversion model to be trained can fully learn the timbre conversion logic during the training process, and improves the timbre similarity of the output result of the voice conversion model. Second, this method uses the global features corresponding to the audio training data as a training data item, so that the voice conversion model to be trained can learn the paralinguistic information and non-linguistic information in the audio training data during the training process, and improves the expressiveness of the output result of the voice conversion model. Third, this method uses the noisy Mel spectrogram corresponding to the audio training data as a training data item to predict the masked part of the masked Mel spectrogram, so that the voice conversion model to be trained can learn to reduce the interference of noise on the prediction result during the training process, and improves the noise robustness of the voice conversion model. Fourth, this method does not involve the pre-operation of time series alignment of the first feature training data, the second feature training data, and the Mel spectrogram, so that the voice conversion model to be trained needs to perform time series alignment modeling during the training process, and reduces the attention of the voice conversion model to non-important information such as noise. In summary, the voice conversion model trained based on the training method of the voice conversion model provided by the present application is a voice conversion model that can not only ensure high timbre similarity and noise robustness, but also improve the expressiveness of the target voice. It solves the technical problem in the prior art that the voice conversion effect is not good due to the inability to balance timbre similarity, noise robustness, and expressiveness. Brief Description of the Drawings

[0013] Figure 1 is an application system diagram of the training method of the voice conversion model provided by the embodiment of the present application;

[0014] Figure 2 is an application system diagram of the voice conversion method provided by the embodiment of the present application;

[0015] Figure 3 is a flowchart of the training method of the voice conversion model provided by the first embodiment of the present application;

[0016] Figure 4 is a schematic diagram of training a voice conversion model to be trained based on audio training data provided by the first embodiment of the present application;

[0017] Figure 5 is a flowchart of the voice conversion method provided by the second embodiment of the present application;

[0018] Figure 6 is a schematic structural diagram of the training device of the voice conversion model provided by the third embodiment of the present application;

[0019] Figure 7 is a schematic structural diagram of the voice conversion device provided by the fourth embodiment of the present application;

[0020] Figure 8 is a schematic structural diagram of the electronic device provided by the fifth embodiment of the present application. Detailed Description of the Embodiments

[0021] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0022] With the development of speech processing technology, voice conversion has emerged. Voice conversion is a technology that changes the speaker's voice while keeping the content information unchanged, and is widely used in many fields such as film and television drama dubbing, privacy protection, and personalized speech synthesis. Voice conversion is usually implemented based on a voice conversion model, and the voice conversion model can convert the source speech of the source speaker into the target speech with the voice of the target speaker.

[0023] At present, speech conversion models can be classified into speech conversion models based on parallel corpora and speech conversion models based on non-parallel corpora according to the type of training data. Since parallel corpora require two or more speakers to say the same content, the cost of collecting training data is high and it is difficult to actually operate. Therefore, speech conversion models based on non-parallel corpora are the mainstream in speech conversion applications. Speech conversion models based on non-parallel corpora are mainly trained based on the features of the ASR (Automatic Speech Recognition) model or the SSL (Self Supervised Learning) model.

[0024] To train a speech conversion model based on the features of the ASR model, specifically, the source speech is input into the ASR model to obtain the bottleneck features (i.e., ASR model features) during the process of the ASR model inferring the source speech, and then the speech conversion model is trained with these features. Since the ASR model features remove most of the paralinguistic information (such as the pronunciation style and pitch changes of the source speaker), non-linguistic information (such as laughter, coughing, and inhalation made by the source speaker), and noise information in the source speech. Therefore, training the language conversion model with these features can enable the language conversion model to achieve a high timbre similarity (i.e., the timbre in the target speech is highly similar to the timbre of the target speaker) and noise robustness (i.e., the target speech does not contain noise or only has slight noise, and the content information in the target speech is not interfered by the noise and is consistent with the content information in the source speech). However, the output target speech often lacks expressiveness (i.e., it cannot show the pronunciation style, pitch changes, etc. of the source speaker in the source speech, as well as the laughter, coughing, etc. made by the source speaker, and only shows the emotionless reading aloud of the target speaker).

[0025] To train a speech conversion model based on the features of the SSL model, specifically, the source speech is input into the SSL model to obtain the speech features (i.e., SSL model features) derived by the SSL model through self-supervised learning, and then the speech conversion model is trained with these features. Since the SSL model features almost retain all the information in the source speech (such as paralinguistic information and non-linguistic information). Therefore, training the language conversion model with these features can effectively improve the expressiveness of the target speech. However, since the SSL model features often also retain background noise, it greatly reduces the noise robustness of the speech conversion model.

[0026] Therefore, existing speech conversion models make a trade-off among timbre similarity, noise robustness, and expressiveness and cannot achieve all of them, resulting in the problem of poor speech conversion effect.

[0027] In view of this, the present application provides a method for training a voice conversion model. Based on this method, a voice conversion model can be obtained that can not only ensure high voice quality similarity and noise robustness, but also improve the expressiveness of the target voice. First, the method uses the content features corresponding to the audio training data as a training data item, enabling the voice conversion model to be trained to fully learn the voice quality conversion logic during the training process and improving the voice quality similarity of the output result of the voice conversion model. Second, the method uses the global features corresponding to the audio training data as a training data item, enabling the voice conversion model to be trained to learn the paralinguistic information and non-linguistic information in the audio training data during the training process and improving the expressiveness of the output result of the voice conversion model. Third, the method uses the noise Mel spectrogram corresponding to the audio training data as a training data item to predict the masked part of the masked Mel spectrogram, enabling the voice conversion model to be trained to learn to reduce the interference of noise on the prediction result during the training process and improving the noise robustness of the voice conversion model. Fourth, the method does not involve pre-operations for temporal alignment of the first feature training data, the second feature training data, and the Mel spectrogram, enabling the voice conversion model to be trained to perform temporal alignment modeling during the training process and reducing the attention of the voice conversion model to non-important information such as noise.

[0028] The following further elaborates in detail on the method for training a voice conversion model, the voice conversion method, and their devices, electronic devices, and computer-readable storage media according to the present application, in conjunction with specific embodiments and the accompanying drawings.

[0029] Figure 1 It is an application system diagram of the method for training a voice conversion model provided by an embodiment of the present application. As Figure 1 shown, the system includes a user terminal 101 and a server 102. The user terminal 101 can be any device such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant (Personal Digital Assistant, PDA), etc. The server 102 can be a cloud server communicatively connected to the user terminal 101. The method for training a voice conversion model provided by the present application and the voice conversion model to be trained are deployed on the server 102, and in response to receiving a model training instruction and a training sample set sent by the user terminal 101, the voice conversion model to be trained is trained based on this method.

[0030] Figure 2 It is an application system diagram of the voice conversion method provided by an embodiment of the present application. As Figure 2As shown, the system includes a first client 201, a first server 202, and a second server 203. The first client 201 can be any device such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant (PDA), etc. The first server 202 can be a service module inside the first client 201, a service device electrically connected to the first client 201, or a server communicatively connected to multiple first clients 201. The voice conversion method provided in this application is deployed on the first server 202, and a trained voice conversion model is deployed on the second server 203. In response to receiving a voice conversion instruction sent by the first client 201, the first server 202 invokes the trained voice conversion model from the second server 203 and converts the voice to be converted based on this method.

[0031] The first embodiment of this application provides a method for training a voice conversion model. This method is deployed in Figure 1 the server 102 shown and is used to provide training services for the voice conversion model.

[0032] The model can be understood as a bionic neural network that imitates the structure and function of the biological nervous system and can be applied to tasks such as recognition, classification, analysis, and conversion. The model consists of a large number of nodes (i.e., neurons). By continuously optimizing the parameters of each node during the training process, the model can learn the complex relationships in the training data and make accurate predictions for the data to be processed.

[0033] The voice conversion model is a neural network model that can be applied to voice conversion tasks. It can keep the content information of the source voice unchanged while changing the source speaker's tone to the specified speaker's tone. In this embodiment, the voice conversion model can also ensure that the target voice does not contain noise or only has slight noise, and the content information in the target voice is not interfered by the noise and is consistent with the content information in the source voice. Furthermore, the voice conversion model described in this embodiment can also ensure that the target voice includes the paralinguistic information and non-linguistic information in the source voice and has strong expressiveness.

[0034] The following details the method for training the voice conversion model described in this embodiment:

[0035] Figure 3 is a flowchart of the method for training the voice conversion model provided in this embodiment. The following combines Figure 3 to describe in detail the method for training the voice conversion model provided in this embodiment. The embodiments involved in the following description are used to explain the technical solutions of this application and are not used as limitations for actual use.

[0036] AsFigure 3 As shown in Figure 3 , the training method of the voice conversion model provided in this embodiment includes the following steps S310 to S340:

[0037] Step S310: Obtain a plurality of audio training data, and extract first feature training data and second feature training data from the audio training data; wherein, the first feature training data is used to represent the content features corresponding to the audio training data, and the second feature training data is used to represent the global features corresponding to the audio training data.

[0038] The content features can be understood as features converted from the language information in the audio training data. Exemplarily, a word vector that can represent its semantics converted from the language information "Hello World" in a piece of audio. In this embodiment, the content features extracted from the audio training data are defined as the first feature training data.

[0039] The global features can be understood as features converted from all information such as language information, non-language information, and paralinguistic information in the audio training data. Exemplarily, for a piece of audio, the global features include a word vector that can represent its semantics converted from the language information, an acoustic feature that can represent its pronunciation method and pitch change converted from the paralinguistic information, and an emotional feature that can represent its emotion converted from the non-language information. In this embodiment, the global features extracted from the audio training data are defined as the second feature training data.

[0040] In an optional implementation manner, extracting the first feature training data and the second feature training data from the audio training data may specifically include the following steps S311 to S312:

[0041] Step S311: Input the audio training data into a pre-trained automatic speech recognition (ASR) model, and obtain the content features generated by the ASR model during the inference process of the audio training data, and use the content features as the first feature training data.

[0042] The automatic speech recognition (ASR) model is a neural network model that can convert human speech into text. It recognizes and transcribes the words spoken by the speaker by processing and analyzing audio data. In this embodiment, the ASR model is a pre-trained model that has been trained and can accurately infer audio data to obtain text data. Optionally, the pre-trained ASR model can be a model trained by developers based on a large amount of data, or an existing open-source model, and specific limitations are not made.

[0043] In a specific implementation, the content feature is a bottleneck feature generated during the inference of audio data by a pre-trained ASR model, that is, a feature generated by the hidden layer of the ASR model. The hidden layer is located in the middle of the ASR model and is designed as a compressed representation for retaining the necessary information for speech recognition in the audio data, that is, language information, and converting the language information into features that can characterize its semantics, that is, content features.

[0044] Step S312: Input the audio training data into the pre-trained self-supervised learning model to obtain the global features generated after the self-supervised learning model infers the audio training data, and use the global features as the second feature training data.

[0045] The self-supervised learning (SSL) model is a learning method that does not require manually labeled data. By designing specific tasks (such as predicting and attempting to recover the masked part in the audio data or outputting a certain audio feature in the audio data), it self-supervisedly learns useful logic. In this embodiment, the SSL model is a pre-trained model that has been trained and can accurately infer the global features in the audio data. Optionally, the pre-trained SSL model can be a model trained by developers based on a large amount of data or an existing open-source model, and there is no specific limitation.

[0046] In a specific implementation, the global feature is an output feature generated during the inference of audio data by the pre-trained SSL model, that is, a feature generated by the output layer of the SSL model. The output layer is located at the outermost layer of the SSL model and is responsible for generating the final feature representation, that is, the global feature.

[0047] Step S320: Obtain the Mel spectrogram corresponding to the audio training data, and perform local masking processing and noise addition processing on the Mel spectrogram to obtain the masked Mel image and the noisy Mel spectrogram corresponding to the Mel spectrogram.

[0048] The Mel spectrogram is a representation method of audio signals, which shows the frequency components of the audio at different time points, but uses the Mel scale, which is a frequency scale that is more in line with human auditory perception. The Mel spectrogram is usually obtained by performing short-time Fourier transform (STFT) and Mel filter bank on the audio data. Extracting the Mel spectrogram corresponding to the audio data can adopt conventional techniques, and no specific description is given here.

[0049] In this embodiment, obtaining the Mel spectrogram corresponding to the audio training data can be understood as converting the audio training data into a Mel spectrogram for subsequent processing, including performing local masking processing and noise addition processing on the Mel spectrogram.

[0050] Local mask processing is a data augmentation technique that refers to selectively obscuring or masking certain regions in the mel spectrogram. This processing can simulate the missing or damaged parts that may exist in audio data to help the model learn to better handle incomplete information. In this embodiment, the mel spectrogram processed by local mask processing is defined as the masked mel spectrogram. Exemplarily, a continuous time window can be randomly selected on the time axis, and some continuous frequency channels can be randomly selected on the frequency axis, and then the values of these selected regions are set to zero or a specific value, thereby creating the masked mel spectrogram.

[0051] Noise addition processing is also a data augmentation technique that refers to adding artificially generated noise to the mel spectrogram. This processing can improve the noise robustness of the model, enabling it to better cope with background noise or other interference factors existing in the actual environment. The noise can be randomly generated or extracted from real environmental noise samples. The intensity and type of the added noise can be adjusted according to the specific application scenario. In this embodiment, the mel spectrogram processed by noise addition processing is defined as the noisy mel spectrogram.

[0052] Using the mel spectrograms processed by local mask processing and noise addition processing as part of the training data in the model training process can help the model learn more generalized capabilities, that is, it can still exhibit good conversion effects when facing audio data generated under complex conditions.

[0053] Step S330: Input the first feature training data, the second feature training data, the masked mel spectrogram, and the noisy mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtain the predicted mel spectrogram output by the speech conversion model to be trained.

[0054] After obtaining the first feature training data, the second feature training data, the masked mel spectrogram, and the noisy mel spectrogram corresponding to the audio training data, input these training data into the speech conversion model to be trained. The speech conversion model to be trained will then infer the masked part of the masked mel spectrogram based on the first feature training data, the second feature training data, the masked mel spectrogram, and the noisy mel spectrogram, and thus output the predicted mel spectrogram, that is, the combination of the masked mel spectrogram and the predicted masked part.

[0055] The method provided in this embodiment does not involve aligning the first feature training data, the second feature training data, and the mel spectrograms (including the masked mel spectrogram and the noise mel spectrogram) in time series. That is, after inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, the speech conversion model to be trained first needs to perform time series alignment modeling. That is, during the training process, it learns how to align the first feature training data, the second feature training data with the masked mel spectrogram and the noise mel spectrogram in time series. Thus, based on the time series alignment, based on the first feature training data and the second feature training data corresponding to the masked part, as well as the unmasked part of the masked mel spectrogram and the noise mel spectrogram, the masked part is predicted, and then the predicted mel spectrogram is output. As described above, since the speech conversion model to be trained needs to perform time series alignment modeling during the training process, it reduces its attention to the detailed parts other than speech in the audio training data (such as background noise), and improves the noise robustness of the speech conversion model.

[0056] Based on this, in an optional implementation manner, the speech conversion model to be trained can be a diffusion model, such as a flow matching model. Diffusion Models are a type of generative model with modeling and generation capabilities. In the method provided in this embodiment, using a diffusion model as the speech conversion model to be trained is to enable the model to automatically perform time series alignment of the input features and output features.

[0057] Step S340, according to the mel spectrogram and the predicted mel spectrogram corresponding to the audio training data, adjust the model parameters of the speech conversion model to be trained to obtain the trained speech conversion model.

[0058] After inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, the predicted mel spectrogram output by the speech conversion model to be trained can be obtained. This predicted mel spectrogram can be understood as the predicted value corresponding to the audio training data, while the mel spectrogram can be understood as the true value corresponding to the audio training data. Based on the difference between the predicted value and the true value, the model parameters of the speech conversion model to be trained are adjusted, and the speech conversion model trained by the audio training data can be obtained.

[0059] Specifically, adjusting the model parameters based on the predicted mel spectrogram and the mel spectrogram aims to minimize the difference between the predicted value and the true value, thereby improving the performance of the speech conversion model to be trained. Optionally, the model parameters can be adjusted based on the gradient descent method by minimizing the loss function between the predicted mel spectrogram and the mel spectrogram, gradually making the model parameters approach the optimal solution.

[0060] In an alternative implementation, the method provided in this embodiment may further include the following step S350:

[0061] Step S350: Determine whether the trained voice conversion model has reached a preset training goal. Specifically, in response to the trained voice conversion model reaching the training goal, use the trained voice conversion model as the voice conversion model that has completed training to perform voice conversion on the audio data to be converted; in response to the trained voice conversion model not reaching the training goal, use the trained voice conversion model as the voice conversion model to be trained and continue iterative training.

[0062] Optionally, during the process of training the voice conversion model to be trained based on the audio training data, it is detected at preset time intervals whether the trained voice conversion model has reached the training goal. If the training goal is reached, the trained voice conversion model is the voice conversion model that has completed training and can be put into use. If the training goal is not reached, the trained voice conversion model cannot be put into use and iterative training based on the audio training data is still required.

[0063] Optionally, the training goal is preset by the developer according to the application requirements of the voice conversion model. The training goal can be a loss function threshold preset for the loss function between the predicted value (i.e., the predicted mel spectrogram) and the true value (i.e., the mel spectrogram), or a similarity threshold set for the similarity between the predicted value (i.e., the predicted mel spectrogram) and the true value (i.e., the mel spectrogram), or an iteration number threshold set for the number of iterations of model training. There is no specific limitation here. Exemplarily, if the training goal is less than the loss function threshold of 0.05, then when the loss function between the predicted mel spectrogram output by the voice conversion model to be trained and the mel spectrogram corresponding to the audio training data is less than 0.05, it can be determined that the trained voice conversion model has reached the training goal, and training can be terminated and the voice conversion model that has completed training can be put into the task of converting the voice to be converted.

[0064] In an alternative implementation, before the step of inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the voice conversion model to be trained to obtain the predicted mel spectrogram output by the voice conversion model to be trained, the method provided in this embodiment may further include the following steps S11 to S13:

[0065] Step S11: Perform dimensionality reduction processing on the preset channel dimension of the first feature training data and the second feature training data corresponding to the audio training data respectively to obtain the first representation training data corresponding to the first feature training data and the second representation training data corresponding to the second feature training data.

[0066] Step S12: According to the temporal length of the Mel spectrogram corresponding to the audio training data, pad the first feature training data and the second feature training data to be of the same length as the Mel spectrogram.

[0067] Step S13: Concatenate the padded first feature training data, the padded second feature training data, the masked Mel spectrogram, and the noise Mel spectrogram along the channel dimension to obtain the concatenated feature training data corresponding to the audio training data.

[0068] Performing dimensionality reduction on the feature training data (including the first feature training data and the second feature training data) can be understood as a fully connected mapping, that is, mapping the high-dimensional feature training data to a low dimension. In this embodiment, the feature training data after dimensionality reduction is defined as the feature representation training data.

[0069] The preset channel dimension is an adjustable value. By setting different dimensionality reduction channel dimensions for different feature training data, the proportion of each feature training data in the overall feature data can be adjusted. Generally, the higher the proportion of a feature training data, the higher the possibility that it will be concerned by the speech conversion model to be trained, and the lower the proportion of a feature training data, the lower the attention of the speech conversion model to be trained. Therefore, developers can adjust the final training effect of the speech conversion model to be trained by adjusting the channel dimensions of different feature training data during dimensionality reduction. Exemplarily, if the first feature training data is dimensionally reduced to 256 dimensions and the second feature training data is dimensionally reduced to 64 dimensions, then the speech conversion model to be trained will pay more attention to the first feature training data during training, while reducing the attention to the second feature training data. The finally trained speech conversion model may have a high timbre similarity and noise robustness, but the expressiveness may be slightly weaker.

[0070] In this embodiment, the low-dimensional feature training data obtained by performing dimensionality reduction on the first feature training data is defined as the first feature representation training data, and the low-dimensional feature training data obtained by performing dimensionality reduction on the second feature training data is defined as the second feature representation training data. Optionally, the dimensionality reduction is a preprocessing operation based on an encoder. Specifically, the first feature training data is dimensionally reduced based on a content feature encoder to obtain the first feature representation training data, and the second feature training data is dimensionally reduced based on a global feature encoder to obtain the second feature representation training data. Optionally, the content feature encoder and the global feature encoder can be independent of the speech conversion model or combined with the speech conversion model as components that do not need to be trained in the model.

[0071] In this embodiment, after performing dimensionality reduction processing on the first feature training data and the second feature training data corresponding to the audio training data to obtain the first representation training data and the second representation training data, padding processing is further performed on the first representation training data and the second representation training data so that the first representation training data and the second representation training data are temporally equal in length to the mel spectrogram. The padding processing is different from the temporal alignment processing. The temporal alignment processing usually expands the data to be equal in length to the mel spectrogram through upsampling (such as linear interpolation, transposed convolution, etc.), that is, it changes the structure of the data itself to achieve temporal alignment with the mel spectrogram. The padding processing does not change the structure of the data itself. Only by padding special marker symbols or specific constants (such as padding "0") at the end of the data sequence, the feature data is supplemented to be equal in length to the mel spectrogram. That is, although the padded feature training data is equal in length to the mel spectrogram, it does not achieve temporal alignment.

[0072] In this embodiment, after obtaining the padded first representation training data and the padded second representation training data, the padded first representation training data and the padded second representation training data are also concatenated with the masked mel spectrogram and the noise mel spectrogram in the channel dimension to form a combined feature. In this embodiment, the combined feature after the concatenation processing is defined as the concatenated feature training data. Concatenating in the channel dimension can be understood as horizontal concatenation, that is, increasing the amount of data in the channel dimension without changing the time step or the spatial size. In this embodiment, the temporal length of the concatenated feature training data is still the same as the temporal length of the mel spectrogram.

[0073] Based on this, the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data are input into the speech conversion model to be trained, and the predicted mel spectrogram output by the speech conversion model to be trained is obtained. Specifically, it may include: inputting the concatenated feature training data corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained.

[0074] In an alternative implementation, to further enhance the expressiveness of the audio converted by the voice conversion model, the fundamental frequency data in the audio training data is separately used as an independent feature to participate in the model training. The fundamental frequency data is an inherent attribute of the audio data and is the lowest frequency component in the audio data, which determines the pitch height. Explicitly embodying the fundamental frequency data as an independent training data enables the voice conversion model to have a better conversion effect on special audio data such as singing, which has strong expressiveness and rich rhythm. Based on this, the first feature training data and the second feature training data are extracted from the audio training data, and specifically, it may further include: extracting the first feature training data, the second feature training data, and the fundamental frequency data from the audio training data. Inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the voice conversion model to be trained to obtain the predicted mel spectrogram output by the voice conversion model to be trained, and specifically, it may further include: inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the voice conversion model to be trained to obtain the predicted mel spectrogram output by the voice conversion model to be trained.

[0075] In an alternative implementation, before the step of inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the voice conversion model to be trained to obtain the predicted mel spectrogram output by the voice conversion model to be trained, the method provided in this embodiment further includes: performing multi-scale fundamental frequency modeling on the fundamental frequency data corresponding to the audio training data to obtain the multi-scale fundamental frequency data corresponding to the audio training data.

[0076] Multi-scale fundamental frequency modeling can be understood as downsampling the fundamental frequency data at different granularities to obtain the fundamental frequency data at different granularities. Exemplarily, N fundamental frequency values are sampled from 1 second of audio training data to achieve one fundamental frequency value corresponding to each content word in the audio data. The fundamental frequency data is downsampled to gradually coarsen the fundamental frequency data, so that one fundamental frequency value corresponds to each sentence in the audio data.

[0077] Based on this, the first feature training data, the second feature training data, the fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data are input into the speech conversion model to be trained, and the predicted mel spectrogram output by the speech conversion model to be trained is obtained. Specifically, it may include: inputting the first feature training data, the second feature training data, the multi-scale fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained. Training the speech conversion model based on the multi-scale fundamental frequency data can enhance the stability and accuracy of pronunciation in the target speech converted by the speech conversion model.

[0078] In an optional implementation manner, the multi-scale fundamental frequency data is concatenated with the padded first feature training data, the padded second feature training data, the masked mel spectrogram, and the noise mel spectrogram in the channel dimension to obtain the concatenated feature training data. Subsequently, the concatenated feature training data is input into the speech conversion model to be trained, and the predicted mel spectrogram output by the speech conversion model to be trained is obtained.

[0079] In an optional implementation manner, after the step of extracting the first feature training data and the second feature training data from the audio training data, the method provided in this embodiment may further include: removing a part of the second feature training data corresponding to the audio training data according to a preset ratio. Specifically, for each audio training data among multiple audio training data, after extracting the first feature training data and the second feature training data, a part of the second feature training data is discarded to reduce the dependence of the speech conversion model on the second feature training data. This is because the second feature training data is the global feature in the audio training data, which includes not only language information but also other paralinguistic information, non-linguistic information, background noise, etc. If the speech conversion model depends too much on the second feature training data during training, it will not be able to fully learn the method of voice conversion because it has to pay too much attention to the language information. The removal ratio of the second feature training data is an adjustable value, and developers can adjust it according to the output effect of the speech conversion model during training, and there is no specific limitation.

[0080] Based on this, for the audio training data from which the second feature training data has been removed, inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained. Specifically, it may further include: inputting the first feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained.

[0081] The following uses a specific implementation manner to give an exemplary description of how to train a speech conversion model to be trained based on audio training data.

[0082] Figure 4 It is a schematic diagram of training a speech conversion model to be trained based on audio training data provided in this embodiment. As Figure 4 shown, training the speech conversion model to be trained based on audio training data may specifically include the following steps S401 to step S412:

[0083] Step S401: Input the audio training data 410 into the pre-trained SSL model to obtain the second feature training data 411 corresponding to the audio training data 410.

[0084] Step S402: Input the audio training data 410 into the pre-trained ASR model to obtain the first feature training data 412 corresponding to the audio training data 410.

[0085] Step S403: Input the audio training data 410 into the fundamental frequency tracking algorithm to obtain the fundamental frequency data 413 corresponding to the audio training data 410.

[0086] Step S404: Input the second feature training data 411 into the global feature encoder for dimensionality reduction preprocessing to obtain the second representation training data 414 corresponding to the second feature training data 411.

[0087] Step S405: Input the first feature training data 412 into the content feature encoder for dimensionality reduction preprocessing to obtain the first representation training data 415 corresponding to the first feature training data 412.

[0088] Step S406: Input the fundamental frequency data 413 into the multi-scale fundamental frequency modeling encoder for multi-scale modeling to obtain the multi-scale fundamental frequency data 416 corresponding to the fundamental frequency data 413.

[0089] Step S407: Obtain the Mel spectrogram 417 corresponding to the audio training data 410.

[0090] Step S408: Perform local masking processing on the Mel spectrogram 417 to obtain the masked Mel spectrogram 418.

[0091] Step S409: Perform noise addition processing on the Mel spectrogram 417 to obtain the noisy Mel spectrogram 419.

[0092] Step S410: Pad the second representation training data 414 and the first representation training data 415 to be the same length as the mel spectrogram 417, and concatenate the padded second representation training data 414, the padded first representation training data 415, the multi-scale fundamental frequency data 416, the masked mel spectrogram 418, and the noise mel spectrogram 419 to obtain the concatenated feature training data 420 corresponding to the audio training data 410.

[0093] Step S411: Input the concatenated feature training data 420 into the speech conversion model to be trained, and obtain the predicted mel spectrogram 421 output by the speech conversion model to be trained.

[0094] Step S412: Based on the mel spectrogram 417 corresponding to the audio training data 410 and the predicted mel spectrogram 421, adjust the model parameters of the speech conversion model.

[0095] The above first embodiment provides an optional training method for a speech conversion model. First, this method uses the content features corresponding to the audio training data as a training data item, enabling the speech conversion model to be trained to fully learn the timbre conversion logic and improving the timbre similarity of the output result of the speech conversion model. Second, this method uses the global features corresponding to the audio training data as a training data item, enabling the speech conversion model to be trained to learn the paralinguistic information and non-linguistic information in the audio training data and improving the expressiveness of the output result of the speech conversion model. Third, this method uses the noise mel spectrogram corresponding to the audio training data as a training data item to predict the masked part of the masked mel spectrogram, enabling the speech conversion model to be trained to learn to reduce the interference of noise on the prediction result and improving the noise robustness of the speech conversion model. Fourth, this method does not involve pre-operations for temporal alignment of the first feature training data, the second feature training data, and the mel spectrogram, enabling the speech conversion model to be trained to perform temporal alignment modeling and reducing the attention of the speech conversion model to non-essential information such as noise. Fifth, in some implementation manners, this method also uses the fundamental frequency data / multi-scale fundamental frequency data corresponding to the audio training data as a training data item, enabling the speech conversion model to be trained to learn the pitch conversion logic and further improving the expressiveness of the output result of the speech conversion model. Especially for special audio data such as singing, it can ensure the stability and accuracy of singing pronunciation. In summary, the speech conversion model trained based on the training method of the speech conversion model provided in this application is a speech conversion model that can not only ensure high timbre similarity and noise robustness but also improve the expressiveness of the target speech.

[0096] It should be noted that the examples in the first embodiment are only for explaining the method described in this application, and do not serve as a limitation for actual use. The training method of the voice conversion model provided by this application includes but is not limited to the method described in the first embodiment.

[0097] A second embodiment of this application provides a voice conversion method, which is deployed in Figure 2 the first server 202 shown, and is used to provide a conversion service for the voice to be converted. Specifically, this method uses the voice conversion model trained in the first embodiment of this application to perform conversion processing on the source voice of the speaker to be converted, and generate the target voice of the specified speaker.

[0098] Figure 5 is a flowchart of the voice conversion method provided in this embodiment. The following combines Figure 5 to describe in detail the voice conversion method provided in this embodiment. The embodiments involved in the following description are used to explain the technical solutions of this application, and do not serve as a limitation for actual use.

[0099] As Figure 5 shown, the voice conversion method provided in this embodiment includes the following steps S510 to S540:

[0100] Step S510, in response to a voice conversion instruction, obtain the first audio data of the speaker to be converted and the sample audio data of the specified speaker.

[0101] The voice conversion instruction can be understood as the instruction information sent by the speaker to be converted or the user based on Figure 2 the first client 201 shown to the first server 202, which requires converting the source voice of the speaker to be converted into the target voice of the specified speaker. Exemplarily, the voice conversion instruction can be the instruction information generated and sent by the first client 201 in response to the speaker to be converted uploading the source voice and checking the specified speaker and sent to the second server 202. In this embodiment, the source voice of the speaker to be converted is defined as the first audio data, and the target voice of the specified speaker is defined as the second audio data.

[0102] The sample audio data of the specified speaker refers to the sample voice of the specified speaker prepared in advance, which is used to provide the timbre of the specified speaker. This sample audio data can be a very short segment of voice, such as a few seconds or a dozen seconds. After the speaker to be converted or the user specifies the speaker, the sample audio data of the specified speaker will be obtained while obtaining the first audio data, so as to convert the first audio data into the second audio data with the timbre of the specified speaker.

[0103] Step S520: Extract first feature data and second feature data from the first audio data, and extract third feature data and fourth feature data from the sample audio data. Herein, the first feature data is used to characterize the content features corresponding to the first audio data, the second feature data is used to characterize the global features corresponding to the first audio data, the third feature data is used to characterize the content features corresponding to the sample audio data, and the fourth feature data is used to characterize the global features corresponding to the sample audio data.

[0104] In an optional implementation manner, extracting the first feature data and the second feature data from the first audio data, and extracting the third feature data and the fourth feature data from the sample audio data may specifically include the following steps S521 to S524:

[0105] Step S521: Input the first audio data into a pre-trained speech recognition model, obtain the content features generated during the inference process of the speech recognition model for the first audio data, and use the content features as the first feature data.

[0106] Step S522: Input the first audio data into a pre-trained self-supervised learning model, obtain the global features generated after the self-supervised learning model infers the first audio data, and use the global features as the second feature data.

[0107] Step S523: Input the sample audio data into a pre-trained speech recognition model, obtain the content features generated during the inference process of the speech recognition model for the sample audio data, and use the content features as the third feature data.

[0108] Step S524: Input the sample audio data into a pre-trained self-supervised learning model, obtain the global features generated after the self-supervised learning model infers the sample audio data, and use the global features as the fourth feature data.

[0109] The specific method for extracting content features and global features from audio data based on the pre-trained speech recognition model and self-supervised learning model has been described in detail in the first embodiment of this application, and will not be elaborated herein.

[0110] Step S530: Concatenate the first feature data and the third feature data to generate fifth feature data, and concatenate the second feature data and the fourth feature data to generate sixth feature data.

[0111] The concatenation of the first feature data and the third feature data is performed in the time series dimension. Specifically, the third feature data is concatenated at the front end of the first feature data to form the fifth feature data. Similarly, the concatenation of the second feature data and the fourth feature data is also performed in the time series dimension. Specifically, the fourth feature data is concatenated at the front end of the second feature data to form the sixth feature data.

[0112] Step S540: According to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, convert the first audio data of the speaker to be converted into the second audio data of the specified speaker. The trained voice conversion model is obtained by training according to the method described in the first embodiment of the present application.

[0113] Specifically, use the fifth feature data, the sixth feature data, and the first Mel spectrogram corresponding to the sample audio data as output data and input them into the trained voice conversion model. The trained voice conversion model will continue to write the first Mel spectrogram and further perform audio conversion based on the Mel spectrogram written after the first Mel spectrogram, and then the second audio data with the timbre of the specified speaker can be obtained. Since the trained voice conversion model is trained based on multiple features (such as content features, global features, fundamental frequency), the second audio data not only has high timbre similarity and noise robustness, but also has high expressiveness.

[0114] In an optional implementation, before the step of concatenating the first feature data and the third feature data to generate the fifth feature data and concatenating the second feature data and the fourth feature data to generate the sixth feature data, the method provided in this embodiment may further include the following steps S21 to S24:

[0115] Step S21: Perform dimensionality reduction processing on the first feature data and the third feature data in a first preset channel dimension to obtain the first representation data corresponding to the first feature data and the third representation data corresponding to the third feature data.

[0116] Step S22: Perform dimensionality reduction processing on the second feature data and the fourth feature data in a second preset channel dimension to obtain the second representation data corresponding to the second feature data and the fourth representation data corresponding to the fourth feature data.

[0117] Step S23: According to the time sequence length of the first Mel spectrogram, pad the third representation data and the fourth representation data to be of the same length as the first Mel spectrogram.

[0118] Step S24: According to the time sequence length of the second Mel spectrogram, pad the first representation data and the second representation data to be of the same length as the second Mel spectrogram. The second Mel spectrogram is a full-mask Mel spectrogram preset according to the time sequence length of the first audio data.

[0119] The dimensionality reduction process has been described in detail in the first embodiment of the present application and will not be elaborated here. It should be noted that the first preset channel dimension and the second preset channel dimension need to be consistent with the channel dimensions preset for the first feature training data and the second feature training data when training the voice conversion model. Exemplarily, when training the voice conversion model, the channel dimension set for the first feature training data is 256, and the channel dimension set for the second feature training data is 128. Then, in this embodiment, the dimensionality reduction process for the first feature data and the third feature data needs to be performed in 256 dimensions, and the dimensionality reduction process for the second feature data and the fourth feature data needs to be performed in 128 dimensions. In this embodiment, the low-dimensional feature data obtained by performing the dimensionality reduction process on the first feature data is defined as the first representation data. Similarly, the low-dimensional feature data obtained by performing the dimensionality reduction process on the second feature data, the third feature data, and the fourth feature data are respectively defined as the second representation data, the third representation data, and the fourth representation data.

[0120] In this embodiment, after performing the dimensionality reduction process on the first feature data, the second feature data, the third feature data, and the fourth feature data, the representation data corresponding to these feature data will also be filled. Specifically, the first representation data and the second representation data are filled to be the same length as the second mel spectrogram, and the third representation data and the fourth representation data are filled to be the same length as the first mel spectrogram. The filling method has been described in detail in the first embodiment of the present application and will not be elaborated here. Exemplarily, the time series length corresponding to the first mel spectrogram is 10 frames. Then, "0" needs to be filled at the end of the sequences of the third representation data and the fourth representation data so that the time series lengths of the third representation data and the fourth representation data also reach 10 frames. The time series length corresponding to the second mel spectrogram is 1000 frames. Then, "0" needs to be filled at the end of the sequences of the first representation data and the second representation data so that the time series lengths of the first representation data and the third representation data also reach 1000 frames.

[0121] The second mel spectrogram is a fully masked mel spectrogram preset according to the time series length of the first audio data. Therefore, the second mel spectrogram can be understood as a blank mel spectrogram with a certain time series length, and this blank mel spectrogram is the content that the trained voice conversion model needs to predict.

[0122] Based on this, the first feature data and the third feature data are concatenated to generate fifth feature data, and the second feature data and the fourth feature data are concatenated to generate sixth feature data. Specifically, it may include: concatenating the filled third representation data before the filled first representation data to generate fifth feature data, and concatenating the filled fourth representation data before the filled second representation data to generate sixth feature data. Specifically, the filled third representation data is concatenated in the time series dimension before the filled first representation data, and the filled fourth representation data is concatenated in the time series dimension before the filled second representation data. The fifth feature data and the sixth feature data after concatenation are of equal length and equal to the sum of the time series lengths corresponding to the first Mel spectrogram and the second Mel spectrogram.

[0123] In an alternative implementation, before the step of converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, the method provided in this embodiment may further include the following steps S31 to S32:

[0124] Step S31: Concatenate the first Mel spectrogram before the second Mel spectrogram to generate a concatenated Mel spectrogram, where the second Mel spectrogram is a fully masked Mel spectrogram preset according to the time series length of the first audio data.

[0125] Step S32: Concatenate the fifth feature data, the sixth feature data, and the concatenated Mel spectrogram in the channel dimension to obtain concatenated feature data.

[0126] In this implementation, the first Mel spectrogram and the second Mel spectrogram are also concatenated. Specifically, the first Mel spectrogram is concatenated in the time series dimension before the second Mel spectrogram to generate a concatenated Mel spectrogram. It can be understood that predicting the second Mel spectrogram through the trained voice conversion model is actually a continuation of the first Mel spectrogram.

[0127] Based on the above steps to obtain the concatenated fifth feature data, sixth feature data, and concatenated Mel spectrogram, the fifth feature data, the sixth feature data, and the concatenated Mel spectrogram can be concatenated in the channel dimension to form a combined feature. In this embodiment, the combined feature after this concatenation process is defined as concatenated feature data.

[0128] Based on this, in an alternative implementation, according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, converting the first audio data of the speaker to be converted into the second audio data of the specified speaker may specifically include the following steps S41 to S42:

[0129] Step S41: Input the splicing feature data into the trained voice conversion model to obtain the predicted Mel spectrogram corresponding to the second Mel spectrogram output by the trained voice conversion model.

[0130] Step S42: Perform audio conversion on the predicted Mel spectrogram to generate second audio data.

[0131] After the above splicing feature data is input into the trained voice conversion model, the model will predict the fully masked second Mel spectrogram and output the prediction result of the second Mel spectrogram. In this embodiment, the prediction result of the second Mel spectrogram is defined as the predicted Mel spectrogram. After obtaining the predicted Mel spectrogram, the corresponding audio waveform, that is, the second audio data, can be synthesized through a pre-trained vocoder. The function of the vocoder is to convert the Mel spectrogram into the corresponding audio data. The pre-trained vocoder can be pre-trained by developers or an open-source model, and there is no specific limitation. Based on the conversion of the Mel spectrogram into audio data by the vocoder being relatively mature, no detailed description will be given here.

[0132] In an optional implementation manner, when the first audio data is audio data with strong expressiveness and rich rhythm such as singing, in order to enhance the conversion effect of the trained voice conversion model, fundamental frequency data can also be extracted from the first audio data, and the fundamental frequency data is used as an input data item to ensure the stability and accuracy of the pronunciation of the second audio data. Based on this, first feature data and second feature data are extracted from the first audio data, and third feature data and fourth feature data are extracted from the sample audio data. Specifically, it may include: extracting first feature data, second feature data, and first fundamental frequency data from the first audio data, and extracting third feature data, fourth feature data, and second fundamental frequency data from the sample audio data. Subsequently, while splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data, the first fundamental frequency data and the second fundamental frequency data are spliced to generate third fundamental frequency data. Specifically, the second fundamental frequency data is spliced before the first fundamental frequency data in the time series dimension to generate third fundamental frequency data. Further, according to the fifth feature data, sixth feature data, third fundamental frequency data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, the first audio data of the speaker to be converted can be converted into the second audio data of the specified speaker.

[0133] In an alternative implementation, to further improve the pronunciation stability and accuracy of the second audio data, multi-scale fundamental frequency modeling can be performed on the fundamental frequency data. The multi-scale fundamental frequency modeling has been described in the first embodiment of this application and will not be elaborated here. Specifically, before the steps of concatenating the first feature data and the third feature data to generate the fifth feature data, concatenating the second feature data and the fourth feature data to generate the sixth feature data, and concatenating the first fundamental frequency data and the second fundamental frequency data to generate the third fundamental frequency data, the method provided in this embodiment may further include: performing multi-scale fundamental frequency modeling on the first fundamental frequency data and the second fundamental frequency data to obtain the first multi-scale fundamental frequency data corresponding to the first fundamental frequency data and the second multi-scale fundamental frequency data corresponding to the second fundamental frequency data. Subsequently, while concatenating the first feature data and the third feature data to generate the fifth feature data and concatenating the second feature data and the fourth feature data to generate the sixth feature data, concatenate the first multi-scale fundamental frequency data and the second multi-scale fundamental frequency data to generate the third multi-scale fundamental frequency data. Further, according to the fifth feature data, the sixth feature data, the third multi-scale fundamental frequency data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, convert the first audio data of the speaker to be converted into the second audio data of the specified speaker.

[0134] The following uses a specific implementation to exemplarily illustrate how to convert the first audio data of the speaker to be converted into the second audio data of the specified speaker based on the trained voice conversion model.

[0135] Converting the first audio data of the speaker to be converted into the second audio data of the specified speaker based on the trained voice conversion model may specifically include the following steps S601 to S613:

[0136] Step S601, in response to receiving a voice conversion instruction, obtain the first audio data of the speaker to be converted and the sample audio data of the specified speaker.

[0137] Step S602, input the first audio data into the pre-trained SAR model to obtain the first feature data corresponding to the first audio data; input the first audio data into the pre-trained SSL model to obtain the second feature data corresponding to the first audio data; input the first audio data into the fundamental frequency tracking algorithm to obtain the first fundamental frequency data corresponding to the first audio data.

[0138] Step S603, input the sample audio data into the pre-trained SAR model to obtain the third feature data corresponding to the sample audio data; input the sample audio data into the pre-trained SSL model to obtain the fourth feature data corresponding to the sample audio data; input the sample audio data into the fundamental frequency tracking algorithm to obtain the second fundamental frequency data corresponding to the sample audio data.

[0139] Step S604: Input the first feature data and the third feature data into the content feature encoder for dimensionality reduction preprocessing to obtain the first representation data corresponding to the first feature data and the third representation data corresponding to the third feature data.

[0140] Step S605: Input the second feature data and the fourth feature data into the global feature encoder for dimensionality reduction preprocessing to obtain the second representation data corresponding to the second feature data and the fourth representation data corresponding to the fourth feature data.

[0141] Step S606: Input the first fundamental frequency data and the second fundamental frequency data into the multi-scale fundamental frequency modeling encoder for multi-scale modeling to obtain the first multi-scale fundamental frequency data corresponding to the first fundamental frequency data and the second multi-scale fundamental frequency data corresponding to the second fundamental frequency data.

[0142] Step S607: Obtain the first Mel spectrogram corresponding to the sample audio data.

[0143] Step S608: Construct a fully masked second Mel spectrogram according to the temporal length of the first audio data.

[0144] Step S609: Pad the first representation data and the second representation data to be of the same length as the second Mel spectrogram, and pad the third representation data and the fourth representation data to be of the same length as the first Mel spectrogram.

[0145] Step S610: Concatenate the padded third representation data before the padded first representation data in the temporal dimension to generate the fifth feature data; concatenate the padded fourth representation data before the padded second representation data in the temporal dimension to generate the sixth feature data; concatenate the second multi-scale fundamental frequency data before the first multi-scale fundamental frequency data in the temporal dimension to generate the third multi-scale fundamental frequency data; concatenate the first Mel spectrogram before the second Mel spectrogram in the temporal dimension to generate the concatenated Mel spectrogram.

[0146] Step S611: Perform concatenation processing on the fifth feature data, the sixth feature data, the third multi-scale fundamental frequency data, and the concatenated Mel spectrogram in the channel dimension to generate the concatenated feature data.

[0147] Step S612: Input the concatenated feature data into the trained voice conversion model to obtain the predicted Mel spectrogram output by the model.

[0148] Step S613: Input the predicted Mel spectrogram into the pre-trained vocoder to obtain the second audio data converted by the vocoder.

[0149] The above second embodiment provides an alternative voice conversion method. This method selects a combination of content features (i.e., ASR features extracted by an ASR model) and global features (i.e., SSL features extracted by an SSL model) to provide sufficient content information and acoustic information. On this basis, multi-scale fundamental frequency modeling is added to improve the pronunciation stability of the conversion result, which is particularly obvious in singing conversion. It should be noted that the voice conversion method provided in this embodiment is different from the current conversion scheme based on a voice conversion model in that: the method provided in this embodiment cancels the strong prior information that the content information and the acoustic information are completely aligned in time sequence, and the voice conversion model needs to complete the alignment modeling among the content information, the acoustic information, and the prediction result by itself. This not only facilitates the voice conversion model to transfer the pronunciation characteristics and styles of a specified speaker, but also can effectively reduce the attention of the voice conversion model to background noise, thereby achieving a highly expressive voice conversion effect with strong robustness.

[0150] It should be noted that the examples in the second embodiment are only for explaining the method described in this application and are not used as a limitation for actual use. The voice conversion method provided in this application includes but is not limited to the method described in the second embodiment.

[0151] The third embodiment of this application provides a training device for a voice conversion model. Figure 6 It is a schematic structural diagram of the training device for the voice conversion model provided in this embodiment.

[0152] As Figure 6 shown, the training device for the voice conversion model provided in this embodiment includes: a first data processing unit 601, a second data processing unit 602, a model training unit 603, and a model parameter adjustment unit 604.

[0153] The first data processing unit 601 is configured to obtain a plurality of audio training data, and extract first feature training data and second feature training data from the audio training data; wherein, the first feature training data is used to represent the content features corresponding to the audio training data, and the second feature training data is used to represent the global features corresponding to the audio training data.

[0154] Optionally, the extracting the first feature training data and the second feature training data from the audio training data includes:

[0155] Inputting the audio training data into a pre-trained speech recognition model, obtaining the content features generated by the speech recognition model during the inference process of the audio training data, and using the content features as the first feature training data;

[0156] Input the audio training data into a pre-trained self-supervised learning model to obtain the global features generated by the self-supervised learning model after reasoning on the audio training data, and use the global features as the second feature training data.

[0157] The second data processing unit 602 is configured to obtain the Mel spectrogram corresponding to the audio training data, and perform local masking processing and noise addition processing on the Mel spectrogram to obtain the masked Mel spectrogram and the noisy Mel spectrogram corresponding to the Mel spectrogram.

[0158] The model training unit 603 is configured to input the first feature training data, the second feature training data, the masked Mel spectrogram, and the noisy Mel spectrogram corresponding to the audio training data into a speech conversion model to be trained, and obtain the predicted Mel spectrogram output by the speech conversion model to be trained.

[0159] Optionally, before the step of inputting the first feature training data, the second feature training data, the masked Mel spectrogram, and the noisy Mel spectrogram corresponding to the audio training data into a speech conversion model to be trained and obtaining the predicted Mel spectrogram output by the speech conversion model to be trained, it is further configured to:

[0160] Perform dimensionality reduction processing on the first feature training data and the second feature training data corresponding to the audio training data in a preset channel dimension to obtain the first representation training data corresponding to the first feature training data and the second representation training data corresponding to the second feature training data;

[0161] According to the time sequence length of the Mel spectrogram corresponding to the audio training data, fill the first representation training data and the second representation training data to be of the same length as the Mel spectrogram;

[0162] Perform channel dimension concatenation on the filled first representation training data, the filled second representation training data, the masked Mel spectrogram, and the noisy Mel spectrogram to obtain the concatenated feature training data corresponding to the audio training data;

[0163] The step of inputting the first feature training data, the second feature training data, the masked Mel spectrogram, and the noisy Mel spectrogram corresponding to the audio training data into a speech conversion model to be trained and obtaining the predicted Mel spectrogram output by the speech conversion model to be trained includes:

[0164] Input the concatenated feature training data corresponding to the audio training data into the speech conversion model to be trained, and obtain the predicted Mel spectrogram output by the speech conversion model to be trained.

[0165] Optionally, extracting the first feature training data and the second feature training data from the audio training data includes:

[0166] Extracting the first feature training data, the second feature training data, and the fundamental frequency data from the audio training data;

[0167] Inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained includes:

[0168] Inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained.

[0169] Optionally, before the step of inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained and obtaining the predicted mel spectrogram output by the speech conversion model to be trained, it is also used for:

[0170] Performing multi-scale fundamental frequency modeling on the fundamental frequency data corresponding to the audio training data to obtain the multi-scale fundamental frequency data corresponding to the audio training data;

[0171] Inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained includes:

[0172] Inputting the first feature training data, the second feature training data, the multi-scale fundamental frequency data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained.

[0173] Optionally, after the step of extracting the first feature training data and the second feature training data from the audio training data, it is also used for:

[0174] Removing a part of the second feature training data corresponding to the audio training data according to a preset ratio;

[0175] Inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained, includes:

[0176] Inputting the first feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining the predicted mel spectrogram output by the speech conversion model to be trained.

[0177] The model parameter adjustment unit 604 is configured to adjust the model parameters of the speech conversion model to be trained according to the mel spectrogram and the predicted mel spectrogram corresponding to the audio training data, and obtain the trained speech conversion model.

[0178] Optionally, the apparatus further includes a judgment unit; the judgment unit is configured to:

[0179] Judge whether the trained speech conversion model reaches a preset training target;

[0180] In response to the trained speech conversion model reaching the training target, using the trained speech conversion model as the speech conversion model that has completed training to perform speech conversion on the audio data to be converted;

[0181] In response to the trained speech conversion model not reaching the training target, using the trained speech conversion model as the speech conversion model to be trained and continuing iterative training.

[0182] The fourth embodiment of the present application provides a speech conversion apparatus. Figure 7 It is a schematic structural diagram of the speech conversion apparatus provided in this embodiment.

[0183] As Figure 7 shown, the speech conversion apparatus provided in this embodiment includes: a data acquisition unit 701, a feature extraction unit 702, a feature splicing unit 703, and a data conversion unit 704.

[0184] The data acquisition unit 701 is configured to acquire the first audio data of the speaker to be converted and the sample audio data of the specified speaker in response to a speech conversion instruction.

[0185] The feature extraction unit 702 is configured to extract first feature data and second feature data from the first audio data, and extract third feature data and fourth feature data from the sample audio data; wherein, the first feature data is used to characterize the content features corresponding to the first audio data, the second feature data is used to characterize the global features corresponding to the first audio data, the third feature data is used to characterize the content features corresponding to the sample audio data, and the fourth feature data is used to characterize the global features corresponding to the sample audio data.

[0186] Optionally, the extracting first feature data and second feature data from the first audio data, and extracting third feature data and fourth feature data from the sample audio data includes:

[0187] Input the first audio data into a pre-trained speech recognition model, obtain the content features generated during the inference process of the speech recognition model for the first audio data, and use the content features as the first feature data;

[0188] Input the first audio data into a pre-trained self-supervised learning model, obtain the global features generated after the self-supervised learning model infers the first audio data, and use the global features as the second feature data;

[0189] Input the sample audio data into the speech recognition model, obtain the content features generated during the inference process of the speech recognition model for the sample audio data, and use the content features as the third feature data;

[0190] Input the sample audio data into the self-supervised learning model, obtain the global features generated after the self-supervised learning model infers the sample audio data, and use the global features as the fourth feature data.

[0191] The feature splicing unit 703 is configured to splice the first feature data and the third feature data to generate fifth feature data, and splice the second feature data and the fourth feature data to generate sixth feature data.

[0192] Optionally, before the step of splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data, it is further configured to:

[0193] Perform dimensionality reduction processing on the first feature data and the third feature data in a first preset channel dimension to obtain a first representation data corresponding to the first feature data and a third representation data corresponding to the third feature data;

[0194] Perform dimensionality reduction processing on the second feature data and the fourth feature data in a second preset channel dimension to obtain second representation data corresponding to the second feature data and fourth representation data corresponding to the fourth feature data;

[0195] According to the time series length of the first Mel spectrogram, pad the third representation data and the fourth representation data to be of the same length as the first Mel spectrogram;

[0196] According to the time series length of the second Mel spectrogram, pad the first representation data and the second representation data to be of the same length as the second Mel spectrogram, where the second Mel spectrogram is a full mask Mel spectrogram preset according to the time series length of the first audio data;

[0197] The step of splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data includes:

[0198] Splice the padded third representation data before the padded first representation data to generate the fifth feature data, and splice the padded fourth representation data before the padded second representation data to generate the sixth feature data.

[0199] The data conversion unit 704 is configured to convert the first audio data of the to-be-converted speaker into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model; wherein, the trained voice conversion model is obtained by training according to the training method of the voice conversion model.

[0200] Optionally, the step of extracting the first feature data and the second feature data from the first audio data, and extracting the third feature data and the fourth feature data from the sample audio data includes:

[0201] Extract the first feature data, the second feature data, and first fundamental frequency data from the first audio data, and extract the third feature data, the fourth feature data, and second fundamental frequency data from the sample audio data;

[0202] The step of splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data includes:

[0203] Concatenate the first feature data and the third feature data to generate the fifth feature data, concatenate the second feature data and the fourth feature data to generate the sixth feature data, and concatenate the first fundamental frequency data and the second fundamental frequency data to generate the third fundamental frequency data;

[0204] Converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, includes:

[0205] Converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the third fundamental frequency data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model.

[0206] Optionally, before the steps of concatenating the first feature data and the third feature data to generate the fifth feature data, concatenating the second feature data and the fourth feature data to generate the sixth feature data, and concatenating the first fundamental frequency data and the second fundamental frequency data to generate the third fundamental frequency data, it is also used for:

[0207] Perform multi-scale fundamental frequency modeling on the first fundamental frequency data and the second fundamental frequency data to obtain the first multi-scale fundamental frequency data corresponding to the first fundamental frequency data and the second multi-scale fundamental frequency data corresponding to the second fundamental frequency data;

[0208] The steps of concatenating the first feature data and the third feature data to generate the fifth feature data, concatenating the second feature data and the fourth feature data to generate the sixth feature data, and concatenating the first fundamental frequency data and the second fundamental frequency data to generate the third fundamental frequency data, include:

[0209] Concatenate the first feature data and the third feature data to generate the fifth feature data, concatenate the second feature data and the fourth feature data to generate the sixth feature data, and concatenate the first multi-scale fundamental frequency data and the second multi-scale fundamental frequency data to generate the third multi-scale fundamental frequency data;

[0210] Converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, includes:

[0211] According to the fifth feature data, the sixth feature data, the third multi-scale fundamental frequency data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, convert the first audio data of the speaker to be converted into the second audio data of the specified speaker.

[0212] Optionally, before the step of converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, it is further configured to:

[0213] Generate a spliced Mel spectrogram before splicing the first Mel spectrogram to a second Mel spectrogram, where the second Mel spectrogram is a full-mask Mel spectrogram preset according to the time sequence length of the first audio data;

[0214] Splice the fifth feature data, the sixth feature data, and the spliced Mel spectrogram in the channel dimension to obtain spliced feature data.

[0215] Optionally, the step of converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model includes:

[0216] Input the spliced feature data into the trained voice conversion model to obtain a predicted Mel spectrogram corresponding to the second Mel spectrogram output by the trained voice conversion model;

[0217] Perform audio conversion on the predicted Mel spectrogram to generate the second audio data.

[0218] The fifth embodiment of the present application provides an electronic device, Figure 8 which is a schematic structural diagram of the electronic device provided in this embodiment.

[0219] As Figure 8 shown, the electronic device provided in this embodiment includes: a memory 801 and a processor 802;

[0220] The memory 801 is used to store computer instructions for executing the training method and / or the voice conversion method of the voice conversion model;

[0221] The processor 802 is used to execute the computer instructions stored in the memory 801 to perform the following operations:

[0222] Obtain multiple audio training data, and extract first feature training data and second feature training data from the audio training data; wherein, the first feature training data is used to characterize the content features corresponding to the audio training data, and the second feature training data is used to characterize the global features corresponding to the audio training data;

[0223] Obtain the Mel spectrogram corresponding to the audio training data, and perform local masking processing and noise addition processing on the Mel spectrogram to obtain the masked Mel picture and the noisy Mel spectrogram corresponding to the Mel spectrogram;

[0224] Input the first feature training data, the second feature training data, the masked Mel spectrogram, and the noisy Mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtain the predicted Mel spectrogram output by the speech conversion model to be trained;

[0225] According to the Mel spectrogram and the predicted Mel spectrogram corresponding to the audio training data, adjust the model parameters of the speech conversion model to be trained to obtain the trained speech conversion model.

[0226] Optionally, also execute:

[0227] Judge whether the trained speech conversion model reaches a preset training target;

[0228] In response to the trained speech conversion model reaching the training target, use the trained speech conversion model as the speech conversion model that has completed training to perform speech conversion on the audio data to be converted;

[0229] In response to the trained speech conversion model not reaching the training target, use the trained speech conversion model as the speech conversion model to be trained and continue iterative training.

[0230] Optionally, the extracting the first feature training data and the second feature training data from the audio training data includes:

[0231] Input the audio training data into a pre-trained speech recognition model, obtain the content features generated during the inference process of the speech recognition model for the audio training data, and use the content features as the first feature training data;

[0232] Input the audio training data into a pre-trained self-supervised learning model, obtain the global features generated after the self-supervised learning model infers the audio training data, and use the global features as the second feature training data.

[0233] Optionally, before the step of inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained to obtain the predicted mel spectrogram output by the speech conversion model to be trained, the following is also performed:

[0234] Perform dimensionality reduction processing on the first feature training data and the second feature training data corresponding to the audio training data in a preset channel dimension to obtain the first representation training data corresponding to the first feature training data and the second representation training data corresponding to the second feature training data;

[0235] According to the time series length of the mel spectrogram corresponding to the audio training data, pad the first representation training data and the second representation training data to be the same length as the mel spectrogram;

[0236] Concatenate the padded first representation training data, the padded second representation training data, the masked mel spectrogram, and the noise mel spectrogram in the channel dimension to obtain the concatenated feature training data corresponding to the audio training data;

[0237] The step of inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained to obtain the predicted mel spectrogram output by the speech conversion model to be trained includes:

[0238] Input the concatenated feature training data corresponding to the audio training data into the speech conversion model to be trained to obtain the predicted mel spectrogram output by the speech conversion model to be trained.

[0239] Optionally, the extracting the first feature training data and the second feature training data from the audio training data includes:

[0240] Extract the first feature training data, the second feature training data, and the fundamental frequency data from the audio training data;

[0241] The step of inputting the first feature training data, the second feature training data, the masked mel spectrogram, and the noise mel spectrogram corresponding to the audio training data into the speech conversion model to be trained to obtain the predicted mel spectrogram output by the speech conversion model to be trained includes:

[0242] Input the first feature training data, the second feature training data, the fundamental frequency data, the masked Mel spectrogram, and the noise Mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtain the predicted Mel spectrogram output by the speech conversion model to be trained.

[0243] Optionally, before the step of inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked Mel spectrogram, and the noise Mel spectrogram corresponding to the audio training data into the speech conversion model to be trained and obtaining the predicted Mel spectrogram output by the speech conversion model to be trained, the following is also performed:

[0244] Perform multi-scale fundamental frequency modeling on the fundamental frequency data corresponding to the audio training data to obtain the multi-scale fundamental frequency data corresponding to the audio training data;

[0245] The step of inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked Mel spectrogram, and the noise Mel spectrogram corresponding to the audio training data into the speech conversion model to be trained and obtaining the predicted Mel spectrogram output by the speech conversion model to be trained includes:

[0246] Input the first feature training data, the second feature training data, the multi-scale fundamental frequency data, the masked Mel spectrogram, and the noise Mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtain the predicted Mel spectrogram output by the speech conversion model to be trained.

[0247] Optionally, after the step of extracting the first feature training data and the second feature training data from the audio training data, the following is also performed:

[0248] Remove a part of the second feature training data corresponding to the audio training data according to a preset ratio;

[0249] The step of inputting the first feature training data, the second feature training data, the masked Mel spectrogram, and the noise Mel spectrogram corresponding to the audio training data into the speech conversion model to be trained and obtaining the predicted Mel spectrogram output by the speech conversion model to be trained includes:

[0250] Input the first feature training data, the masked Mel spectrogram, and the noise Mel spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtain the predicted Mel spectrogram output by the speech conversion model to be trained.

[0251] Or, perform the following operations:

[0252] In response to a voice conversion instruction, obtain first audio data of the speaker to be converted and sample audio data of the specified speaker;

[0253] Extract first feature data and second feature data from the first audio data, and extract third feature data and fourth feature data from the sample audio data; wherein, the first feature data is used to characterize the content feature corresponding to the first audio data, the second feature data is used to characterize the global feature corresponding to the first audio data, the third feature data is used to characterize the content feature corresponding to the sample audio data, and the fourth feature data is used to characterize the global feature corresponding to the sample audio data;

[0254] Perform splicing processing on the first feature data and the third feature data to generate fifth feature data, and perform splicing processing on the second feature data and the fourth feature data to generate sixth feature data;

[0255] According to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, convert the first audio data of the speaker to be converted into the second audio data of the specified speaker; wherein, the trained voice conversion model is obtained by training according to the training method of the voice conversion model.

[0256] Optionally, the extracting first feature data and second feature data from the first audio data, and extracting third feature data and fourth feature data from the sample audio data includes:

[0257] Input the first audio data into a pre-trained speech recognition model, obtain the content feature generated by the speech recognition model during the inference process of the first audio data, and use the content feature as the first feature data;

[0258] Input the first audio data into a pre-trained self-supervised learning model, obtain the global feature generated by the self-supervised learning model after the inference of the first audio data, and use the global feature as the second feature data;

[0259] Input the sample audio data into the speech recognition model, obtain the content feature generated by the speech recognition model during the inference process of the sample audio data, and use the content feature as the third feature data;

[0260] Input the sample audio data into the self-supervised learning model, obtain the global feature generated by the self-supervised learning model after the inference of the sample audio data, and use the global feature as the fourth feature data.

[0261] Optionally, before the step of splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data, the following is also performed:

[0262] Perform dimensionality reduction processing on the first feature data and the third feature data in a first preset channel dimension to obtain first representation data corresponding to the first feature data and third representation data corresponding to the third feature data;

[0263] Perform dimensionality reduction processing on the second feature data and the fourth feature data in a second preset channel dimension to obtain second representation data corresponding to the second feature data and fourth representation data corresponding to the fourth feature data;

[0264] According to the time series length of the first Mel spectrogram, pad the third representation data and the fourth representation data to be of the same length as the first Mel spectrogram;

[0265] According to the time series length of the second Mel spectrogram, pad the first representation data and the second representation data to be of the same length as the second Mel spectrogram, where the second Mel spectrogram is a full-mask Mel spectrogram preset according to the time series length of the first audio data;

[0266] The step of splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data includes:

[0267] Splice the padded third representation data before the padded first representation data to generate the fifth feature data, and splice the padded fourth representation data before the padded second representation data to generate the sixth feature data.

[0268] Optionally, the step of extracting first feature data and second feature data from the first audio data, and extracting third feature data and fourth feature data from the sample audio data includes:

[0269] Extract the first feature data, the second feature data, and first fundamental frequency data from the first audio data, and extract the third feature data, the fourth feature data, and second fundamental frequency data from the sample audio data;

[0270] The step of splicing the first feature data and the third feature data to generate fifth feature data, and splicing the second feature data and the fourth feature data to generate sixth feature data includes:

[0271] Concatenate the first feature data and the third feature data to generate the fifth feature data, concatenate the second feature data and the fourth feature data to generate the sixth feature data, and concatenate the first fundamental frequency data and the second fundamental frequency data to generate the third fundamental frequency data;

[0272] Converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, includes:

[0273] Converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the third fundamental frequency data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model.

[0274] Optionally, before the steps of concatenating the first feature data and the third feature data to generate the fifth feature data, concatenating the second feature data and the fourth feature data to generate the sixth feature data, and concatenating the first fundamental frequency data and the second fundamental frequency data to generate the third fundamental frequency data, the following is also performed:

[0275] Perform multi-scale fundamental frequency modeling on the first fundamental frequency data and the second fundamental frequency data to obtain the first multi-scale fundamental frequency data corresponding to the first fundamental frequency data and the second multi-scale fundamental frequency data corresponding to the second fundamental frequency data;

[0276] The steps of concatenating the first feature data and the third feature data to generate the fifth feature data, concatenating the second feature data and the fourth feature data to generate the sixth feature data, and concatenating the first fundamental frequency data and the second fundamental frequency data to generate the third fundamental frequency data, include:

[0277] Concatenate the first feature data and the third feature data to generate the fifth feature data, concatenate the second feature data and the fourth feature data to generate the sixth feature data, and concatenate the first multi-scale fundamental frequency data and the second multi-scale fundamental frequency data to generate the third multi-scale fundamental frequency data;

[0278] Converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, includes:

[0279] According to the fifth feature data, the sixth feature data, the third multi-scale fundamental frequency data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, convert the first audio data of the speaker to be converted into the second audio data of the specified speaker.

[0280] Optionally, before the step of converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model, the following is also performed:

[0281] Generate a spliced Mel spectrogram before splicing the first Mel spectrogram to the second Mel spectrogram, where the second Mel spectrogram is a full-mask Mel spectrogram preset according to the time series length of the first audio data;

[0282] Splice the fifth feature data, the sixth feature data, and the spliced Mel spectrogram in the channel dimension to obtain spliced feature data.

[0283] Optionally, the step of converting the first audio data of the speaker to be converted into the second audio data of the specified speaker according to the fifth feature data, the sixth feature data, the first Mel spectrogram corresponding to the sample audio data, and the trained voice conversion model includes:

[0284] Input the spliced feature data into the trained voice conversion model to obtain a predicted Mel spectrogram corresponding to the second Mel spectrogram output by the trained voice conversion model;

[0285] Perform audio conversion on the predicted Mel spectrogram to generate the second audio data.

[0286] The sixth embodiment of the present application provides a computer-readable storage medium, which includes computer instructions that are used to implement the methods described in the embodiments of the present application when executed by a processor.

[0287] It should be noted that the relational terms such as "first" and "second" in this article are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, words such as "including", "having", "containing", and "comprising" have the same meaning, and at the end of any one or more items after any of the above words, it is open-ended. None of the above nouns means that the one or more items have been listed exhaustively, or are limited to these listed one or more items.

[0288] As used herein, unless otherwise expressly stated, the term "or" includes all possible combinations, except where infeasible. For example, if it is stated that a database may include A or B, then unless otherwise specifically provided or infeasible, it may include database A, or B, or both A and B. As a second example, if it is stated that a certain database may include A, B, or C, then unless otherwise specifically provided or infeasible, the database may include database A, or B, or C, or both A and B, or both A and C, or both B and C, or all of A, B, and C.

[0289] It should be noted that the above embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When executed by a processor, this software can perform the methods disclosed above. The computing units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above multiple modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.

[0290] In the above detailed description, the embodiments have been described with reference to many specific details, which may vary depending on the implementation. Certain adaptations and modifications can be made to the embodiments. For those skilled in the art, some other embodiments can be clearly obtained from the specific implementation manners disclosed in this application. This specification and examples are for illustrative purposes only, and the true scope and essence of this application are defined by the claims. The step order shown in the drawings is also for explanatory purposes only and does not mean to be limited to any specific steps or order. Therefore, those skilled in the art will realize that when implementing the same method, these steps can be executed in a different order.

[0291] In the drawings and detailed description of this application, exemplary embodiments are disclosed. However, many variations and modifications can be made to these embodiments. Accordingly, although specific terms are used, these terms are only general and descriptive, and not for the purpose of limitation.

Claims

1. A method for training a speech conversion model, characterized in that: The method further comprises: Acquire multiple audio training data, and extract first feature training data and second feature training data from the audio training data; wherein the first feature training data is used to characterize content features corresponding to the audio training data, and the second feature training data is used to characterize global features corresponding to the audio training data; Obtain a mel-spectrogram corresponding to the audio training data, and perform local masking and noise processing on the mel-spectrogram to obtain a masked mel-spectrogram and a noise mel-spectrogram corresponding to the mel-spectrogram; Inputting the first feature training data, the second feature training data, the masked mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining a predicted mel-spectrogram output by the speech conversion model to be trained; According to the mel-spectrogram corresponding to the audio training data and the predicted mel-spectrogram, model parameters of the speech conversion model to be trained are adjusted to obtain a trained speech conversion model.

2. The method according to claim 1, characterized in that: The method further comprises: Determining whether the trained speech conversion model has achieved a preset training goal; In response to the trained speech conversion model reaching the training target, using the trained speech conversion model as a trained speech conversion model for performing speech conversion on the audio data to be converted; In response to the trained speech conversion model not reaching the training target, the trained speech conversion model is used as the speech conversion model to be trained, and iterative training is continued.

3. The method according to claim 1, characterized in that The extracting the first feature training data and the second feature training data from the audio training data comprises: Inputting the audio training data into a pre-trained speech recognition model, obtaining content features generated by the speech recognition model in a process of inferring the audio training data, and using the content features as the first feature training data; The audio training data is input into a pre-trained self-supervised learning model to obtain global features generated by the self-supervised learning model after inferring the audio training data, and the global features are used as the second feature training data.

4. The method according to claim 1, characterized in that Before the step of inputting the first feature training data, the second feature training data, the masked mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained to obtain the predicted mel-spectrogram output by the speech conversion model to be trained, the method further includes: Performing dimensionality reduction processing of preset channel dimensions on the first feature training data and the second feature training data corresponding to the audio training data, respectively, to obtain first representation training data corresponding to the first feature training data, and second representation training data corresponding to the second feature training data; According to the time series length of the mel-spectrogram corresponding to the audio training data, padding the first representation training data and the second representation training data to be equal in length to the mel-spectrogram; splicing the padded first representation training data, the padded second representation training data, the masked Mel-spectrogram, and the noise Mel-spectrogram in the channel dimension to obtain spliced ​​feature training data corresponding to the audio training data; The step of inputting the first feature training data, the second feature training data, the masked mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining a predicted mel-spectrogram output by the speech conversion model to be trained, comprises: The concatenated feature training data corresponding to the audio training data is input into the speech conversion model to be trained to obtain the predicted mel-spectrogram output by the speech conversion model to be trained.

5. The method according to claim 1, characterized in that: The extracting the first feature training data and the second feature training data from the audio training data comprises: Extracting the first feature training data, the second feature training data, and fundamental frequency data from the audio training data; The step of inputting the first feature training data, the second feature training data, the masked mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining a predicted mel-spectrogram output by the speech conversion model to be trained, comprises: The first feature training data, the second feature training data, the fundamental frequency data, the mask Mel-spectrogram, and the noise Mel-spectrogram corresponding to the audio training data are input into the speech conversion model to be trained to obtain the predicted Mel-spectrogram output by the speech conversion model to be trained.

6. The method according to claim 5, characterized in that Before the step of inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained to obtain the predicted mel-spectrogram output by the speech conversion model to be trained, the method further includes: Performing multi-scale fundamental frequency modeling on the fundamental frequency data corresponding to the audio training data to obtain the multi-scale fundamental frequency data corresponding to the audio training data; The step of inputting the first feature training data, the second feature training data, the fundamental frequency data, the masked Mel-spectrogram, and the noise Mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained to obtain the predicted Mel-spectrogram output by the speech conversion model to be trained comprises: The first feature training data, the second feature training data, the multi-scale fundamental frequency data, the masked Mel-spectrogram, and the noise Mel-spectrogram corresponding to the audio training data are input into the speech conversion model to be trained to obtain the predicted Mel-spectrogram output by the speech conversion model to be trained.

7. The method according to claim 1, characterized in that After the step of extracting the first feature training data and the second feature training data from the audio training data, the method further includes: According to a preset ratio, removing the second feature training data corresponding to a portion of the audio training data; The step of inputting the first feature training data, the second feature training data, the masked mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtaining a predicted mel-spectrogram output by the speech conversion model to be trained, comprises: The first feature training data, the mask mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data are input into the speech conversion model to be trained to obtain the predicted mel-spectrogram output by the speech conversion model to be trained.

8. A voice conversion method, characterized in that: The method comprises: In response to the voice conversion instruction, obtaining first audio data of the speaker to be converted and sample audio data of the designated speaker; Extracting first feature data and second feature data from the first audio data, and extracting third feature data and fourth feature data from the sample audio data; wherein the first feature data is used to characterize content features corresponding to the first audio data, the second feature data is used to characterize global features corresponding to the first audio data, the third feature data is used to characterize content features corresponding to the sample audio data, and the fourth feature data is used to characterize global features corresponding to the sample audio data; The first feature data and the third feature data are concatenated to generate fifth feature data, and the second feature data and the fourth feature data are concatenated to generate sixth feature data; According to the fifth feature data, the sixth feature data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model, the first audio data of the speaker to be converted is converted into the second audio data of the designated speaker; wherein the trained speech conversion model is obtained by training according to the method described in any one of claims 1 to 7.

9. The method according to claim 8, characterized in that The extracting the first feature data and the second feature data from the first audio data, and extracting the third feature data and the fourth feature data from the sample audio data, comprises: Inputting the first audio data into a pre-trained speech recognition model, obtaining content features generated by the speech recognition model in a process of inferring the first audio data, and using the content features as the first feature data; Inputting the first audio data into a pre-trained self-supervised learning model, obtaining a global feature generated by the self-supervised learning model after inferring the first audio data, and using the global feature as the second feature data; Input the sample audio data into the speech recognition model, obtain content features generated by the speech recognition model in the process of inferring the sample audio data, and use the content features as the third feature data; The sample audio data is input into the self-supervised learning model to obtain the global features generated by the self-supervised learning model after inferring the sample audio data, and the global features are used as the fourth feature data.

10. The method according to claim 8, characterized in that Before the step of concatenating the first feature data with the third feature data to generate the fifth feature data, and concatenating the second feature data with the fourth feature data to generate the sixth feature data, the method further includes: Performing dimensionality reduction processing of a first preset channel dimension on the first feature data and the third feature data to obtain first characterization data corresponding to the first feature data and third characterization data corresponding to the third feature data; Performing dimensionality reduction processing of a second preset channel dimension on the second feature data and the fourth feature data to obtain second characterization data corresponding to the second feature data and fourth characterization data corresponding to the fourth feature data; According to the time series length of the first mel-spectrogram, padding the third characterization data and the fourth characterization data to be equal in length to the first mel-spectrogram; According to the time sequence length of the second mel-spectrogram, the first representation data and the second representation data are filled to be equal to the length of the second mel-spectrogram, where the second mel-spectrogram is a fully masked mel-spectrogram preset according to the time sequence length of the first audio data; The step of concatenating the first feature data with the third feature data to generate the fifth feature data, and concatenating the second feature data with the fourth feature data to generate the sixth feature data includes: The padded third characterization data is concatenated before the padded first characterization data to generate the fifth characterization data, and the padded fourth characterization data is concatenated before the padded second characterization data to generate the sixth characterization data.

11. The method according to claim 8, characterized in that The extracting the first feature data and the second feature data from the first audio data, and extracting the third feature data and the fourth feature data from the sample audio data, comprises: Extracting the first feature data, the second feature data, and the first fundamental frequency data from the first audio data, and extracting the third feature data, the fourth feature data, and the second fundamental frequency data from the sample audio data; The step of concatenating the first feature data with the third feature data to generate the fifth feature data, and concatenating the second feature data with the fourth feature data to generate the sixth feature data includes: The first feature data and the third feature data are concatenated to generate the fifth feature data, the second feature data and the fourth feature data are concatenated to generate the sixth feature data, and the first fundamental frequency data and the second fundamental frequency data are concatenated to generate the third fundamental frequency data; The method of converting the first audio data of the speaker to be converted into the second audio data of the designated speaker according to the fifth feature data, the sixth feature data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model comprises: According to the fifth feature data, the sixth feature data, the third fundamental frequency data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model, the first audio data of the speaker to be converted is converted into the second audio data of the designated speaker.

12. The method according to claim 11, characterized in that Before the steps of concatenating the first feature data with the third feature data to generate the fifth feature data, concatenating the second feature data with the fourth feature data to generate the sixth feature data, and concatenating the first fundamental frequency data with the second fundamental frequency data to generate the third fundamental frequency data, the method further includes: Performing multi-scale fundamental frequency modeling on the first fundamental frequency data and the second fundamental frequency data to obtain first multi-scale fundamental frequency data corresponding to the first fundamental frequency data and second multi-scale fundamental frequency data corresponding to the second fundamental frequency data; The step of concatenating the first feature data with the third feature data to generate the fifth feature data, concatenating the second feature data with the fourth feature data to generate the sixth feature data, and concatenating the first baseband data with the second baseband data to generate the third baseband data includes: The first feature data and the third feature data are concatenated to generate the fifth feature data, the second feature data and the fourth feature data are concatenated to generate the sixth feature data, and the first multi-scale fundamental frequency data and the second multi-scale fundamental frequency data are concatenated to generate the third multi-scale fundamental frequency data; The method of converting the first audio data of the speaker to be converted into the second audio data of the designated speaker according to the fifth feature data, the sixth feature data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model comprises: According to the fifth feature data, the sixth feature data, the third multi-scale fundamental frequency data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model, the first audio data of the speaker to be converted is converted into the second audio data of the designated speaker.

13. The method according to claim 8, characterized in that Before the step of converting the first audio data of the speaker to be converted into the second audio data of the designated speaker according to the fifth feature data, the sixth feature data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model, the method further includes: Splicing the first mel-spectrogram to a second mel-spectrogram to generate a spliced ​​mel-spectrogram, where the second mel-spectrogram is a fully masked mel-spectrogram preset according to the time sequence length of the first audio data; The fifth feature data, the sixth feature data, and the spliced ​​Mel-spectrogram are spliced ​​in the channel dimension to obtain spliced ​​feature data.

14. The method according to claim 13, characterized in that The method of converting the first audio data of the speaker to be converted into the second audio data of the designated speaker according to the fifth feature data, the sixth feature data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model comprises: Inputting the concatenated feature data into the trained speech conversion model to obtain a predicted mel-spectrogram corresponding to the second mel-spectrogram output by the trained speech conversion model; Perform audio conversion on the predicted Mel-spectrogram to generate the second audio data.

15. A training device for a speech conversion model, characterized in that: The device also includes: a first data processing unit, a second data processing unit, a model training unit, and a model parameter adjustment unit; The first data processing unit is used to obtain multiple audio training data and extract first feature training data and second feature training data from the audio training data; wherein the first feature training data is used to characterize content features corresponding to the audio training data, and the second feature training data is used to characterize global features corresponding to the audio training data; The second data processing unit is used to obtain a mel-spectrogram corresponding to the audio training data, and perform local masking and noise processing on the mel-spectrogram to obtain a masked mel-image and a noise mel-spectrogram corresponding to the mel-spectrogram; The model training unit is used to input the first feature training data, the second feature training data, the mask mel-spectrogram, and the noise mel-spectrogram corresponding to the audio training data into the speech conversion model to be trained, and obtain the predicted mel-spectrogram output by the speech conversion model to be trained; The model parameter adjustment unit is used to adjust the model parameters of the speech conversion model to be trained according to the mel-spectrogram and the predicted mel-spectrogram corresponding to the audio training data to obtain the trained speech conversion model.

16. A speech conversion device, characterized in that: The device also includes: a data acquisition unit, a feature extraction unit, a feature splicing unit, and a data conversion unit; The data acquisition unit is used to acquire the first audio data of the speaker to be converted and the sample audio data of the designated speaker in response to the voice conversion instruction; The feature extraction unit is used to extract first feature data and second feature data from the first audio data, and to extract third feature data and fourth feature data from the sample audio data; wherein the first feature data is used to characterize content features corresponding to the first audio data, the second feature data is used to characterize global features corresponding to the first audio data, the third feature data is used to characterize content features corresponding to the sample audio data, and the fourth feature data is used to characterize global features corresponding to the sample audio data; The feature concatenation unit is used to concatenate the first feature data with the third feature data to generate fifth feature data, and to concatenate the second feature data with the fourth feature data to generate sixth feature data; The data conversion unit is used to convert the first audio data of the speaker to be converted into the second audio data of the designated speaker according to the fifth feature data, the sixth feature data, the first mel-spectrogram corresponding to the sample audio data, and the trained speech conversion model; wherein the trained speech conversion model is obtained by training according to the method described in any one of claims 1 to 7.

17. An electronic device, characterized in that: include: Memory, processor; The memory is used to store one or more computer instructions; The processor is used to execute the one or more computer instructions to implement the method according to any one of claims 1-14.

18. A computer-readable storage medium having one or more computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the method according to any one of claims 1 to 14 is performed.

Citation Information

Patent Citations

  • Voice conversion and related model training method, electronic equipment and storage device

    CN112786018A

  • Sound scene classification method based on multi-modal feature fusion

    CN116543795A

  • Speech enhancement method based on artificial intelligence and related equipment

    CN116631426A

  • Speech emotion recognition method based on bimodal and attention mechanism

    CN116705073A

  • Clockwork Hierarchical Variational Encoder

    US20190348020A1