A singing voice conversion method and a singing voice conversion system

By separating and slicing the original audio data, combining the combined loss functions of Mel feature loss, pitch loss and KL divergence loss, feature extraction is achieved using the BERT model and Adapter module, efficient and stable singing voice conversion is achieved, and the naturalness and realism of the generated audio is improved, and the shortcomings in the existing technology are solved.

CN119181370BActive Publication Date: 2025-07-04JIANGNAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411689547.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-07-04
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

The existing singing voice conversion technology has insufficient vocal separation and preprocessing, computational complexity and training stability, limitations of emotional expression and tone conversion, model complexity and data dependence, resulting in low and unstable audio quality generated.

Method used

By separating and slicing the original audio data, a singing voice conversion model is constructed, a combined loss function of Mel feature loss, pitch loss and KL divergence loss is used, feature extraction is combined with the BERT model and Adapter module, and audio synthesis is used using the Flow stream model to optimize the model training process.

Benefits of technology

It improves the quality of vocal separation, reduces data redundancy, improves model training efficiency and naturalness and realism of audio generation, reduces training time, and enhances the retention ability of emotional and style characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181370B_ABST
    Figure CN119181370B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of audio processing, and in particular to a singing voice conversion method and a singing voice conversion system. The method includes: performing vocal separation on the acquired original audio data to obtain clean vocal data; performing slicing processing on the clean vocal data to remove silent sounds and obtain vocal slice data; using the vocal slice data as a training data set to construct a singing voice conversion model, aiming to minimize the value of the loss function, and training the singing voice conversion model through the training data set to obtain a trained singing voice conversion model; inputting the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice. The present invention combines fine audio preprocessing, innovative model architectures and feature extraction methods, and flexible loss function design to achieve efficient and high-quality singing voice conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular to a singing voice conversion method and a singing voice conversion system. Background Art

[0002] The singing voice conversion technology is a technology that can convert the timbre of a source singer into that of a target singer. The core lies in achieving precise timbre conversion while keeping the singing content unchanged. This technology has shown great application potential and value in many fields such as film and television music production, artistic creation, the entertainment industry, and human-computer interaction. Through singing voice conversion, production personnel can easily replace the voice of a certain singer with that of another singer to meet specific creative needs or achieve personalized customization.

[0003] In existing singing voice conversion technologies, efficient vocoders such as WaveNet and Parallel WaveGAN are widely used to generate high-quality audio. These vocoders model audio data through deep learning models and can generate audio with naturalness and realism. However, although these technologies have made significant progress in sound quality, they still face some challenges and limitations:

[0004] First, insufficient vocal separation and preprocessing: The original audio data usually contains various components such as accompaniment, reverberation, inhalation sounds, and environmental noise, which will interfere with the effective extraction of the human voice. Existing technologies often have deficiencies in vocal separation and preprocessing, resulting in low-quality extracted vocal data, which in turn affects the subsequent singing voice conversion effect.

[0005] Second, computational complexity and training stability: Models such as WaveNet adopt an autoregressive structure, resulting in a step-by-step process for generating audio, with low computational efficiency. Although Parallel WaveGAN has improved the generation speed, it still requires a complex training and inference process. In addition, the model is prone to falling into local optimal solutions during training, resulting in unstable audio quality generated.

[0006] Third, limitations in emotional expression and timbre conversion: Existing singing voice conversion technologies can often only achieve timbre conversion and cannot fully retain and convey the emotional and style characteristics of the source singer. This leads to deficiencies in the emotional expression of the generated audio.

[0007] Fourth, model complexity and data dependence: Existing singing voice conversion models have complex structures and high dependence on data, requiring a large amount of high-quality data for training. This increases the cost and difficulty in practical applications.

[0008] Therefore, in order to overcome the deficiencies in the prior art and improve the sound quality, stability, and emotional expression ability of voice conversion, the industry needs to continuously explore new voice conversion methods and algorithms. By optimizing the model structure, improving the training strategy, and introducing techniques such as emotional feature extraction, a more efficient, accurate, and natural voice conversion effect can be achieved. This will inject new vitality and impetus into the development of fields such as film and television music production, artistic creation, the entertainment industry, and human-computer interaction. Summary of the Invention

[0009] To solve the above technical problems, the present invention provides a voice conversion method, including the following steps:

[0010] S1: Perform voice separation on the acquired original audio data to obtain clean voice data;

[0011] S2: Perform slicing processing on the clean voice data to remove silent sounds and obtain voice slice data;

[0012] S3: Use the voice slice data as a training data set to construct a voice conversion model. With the goal of minimizing the value of the loss function, train the voice conversion model through the training data set to obtain a trained voice conversion model; wherein, the loss function includes Mel feature loss, pitch loss, and KL divergence loss;

[0013] S4: Input the audio data to be converted into the trained voice conversion model to obtain the final target voice; wherein, in step S4, the steps to obtain the final target voice are as follows:

[0014] S41: Respectively perform feature extraction on the audio data to be converted to obtain embedded semantic features, pitch features, and Mel spectrogram features;

[0015] S42: Process the embedded semantic features and the pitch features to extract the audio text of the speaker; process the Mel spectrogram features to obtain the timbre features of the speaker;

[0016] S43: Perform audio synthesis on the timbre features and the audio text to obtain the final target voice.

[0017] In an embodiment of the present invention, in S1, the steps to obtain clean voice data are as follows:

[0018] S11: Separate the accompaniment of the original audio data to obtain initial voice data;

[0019] S12: Based on the initial voice data, remove reverberation, inhalation sounds, and environmental noise to obtain the clean voice data.

[0020] In one embodiment of the present invention, in S42, the steps of extracting the audio text of the speaker are as follows:

[0021] S421: Construct an audio content extraction model, perform iterative training on the audio content extraction model to obtain a trained audio content extraction model, including:

[0022] The audio content extraction model includes a BERT model and an Adapter module. Freeze the parameters of the multi-head attention layer and the feed-forward neural network layer in the BERT model, and only iteratively update the parameters of the Adapter module;

[0023] S422: Use the trained audio content extraction model to process the embedded semantic features and the pitch features, and perform feature extraction through the BERT model and the Adapter module to obtain the audio text of the speaker.

[0024] In one embodiment of the present invention, in S422, obtaining the audio text of the speaker :

[0025] ,

[0026] wherein, represents the input of the Adapter module; represents the two-dimensional weight matrix for dimensionality reduction of the input ; represents the bias matrix for dimensionality reduction of the input ; represents the weight matrix for converting the dimensionality-reduced features into the original dimension; is the bias matrix for converting the dimensionality-reduced features into the original dimension.

[0027] In one embodiment of the present invention, the Adapter module is composed of two layers of linear networks, a ReLU activation function, and a residual network.

[0028] In one embodiment of the present invention, in S42, the steps of obtaining the timbre feature of the speaker are as follows:

[0029] S423: Construct a timbre feature extraction model, and the timbre feature extraction model includes stacked causal convolutional layers, residual connection layers, and gated activation modules;

[0030] S424: At each time node, obtain the temporal dependence features in the Mel spectrogram features through the causal convolutional layer and the gated activation module, calculate the probability distribution value of the next sampling point based on the current Mel spectrogram features and the temporal dependence features of the previous time nodes, and sample to obtain the next sampling point;

[0031] S425: Repeatedly execute step S424 to continuously generate new sampling points until a speech feature sequence is obtained;

[0032] S426: Extract the timbre feature of the speaker based on the speech feature sequence.

[0033] In an embodiment of the present invention, the expression of the loss function is:

[0034] ,

[0035] where L is the total loss function value, is the attenuation coefficient, represents the pitch loss, represents the KL divergence loss, represents the Mel feature loss.

[0036] In an embodiment of the present invention, the Mel feature loss is calculated as follows:

[0037] ,

[0038] where represents the actual Mel spectrogram feature of the input; represents the Mel spectrogram feature of the audio generated by the model;

[0039] The pitch loss is calculated as follows:

[0040] ,

[0041] where represents the embedded semantic feature after frequency conversion; represents the predicted embedded semantic feature obtained after f0 passing through the decoder; n represents the number of samples;

[0042] The KL divergence loss is calculated as follows:

[0043] ,

[0044] where and are the embedded representations of the i-th sample after being encoded by enc_p, where represents the output of the first half of the encoding result, represents the output of the second half of the encoding result; and are the embedded representation and mask representation of the i-th sample after being encoded by enc_q, , represents the embedding representation generated by the HuBERT model for the i-th audio sample; represents the embedding representation obtained through the flow model, .

[0045] Based on the same inventive concept, the present invention also provides a singing voice conversion system, which includes the following modules:

[0046] A clean vocal data acquisition module, which is used to perform vocal separation on the acquired original audio data to obtain clean vocal data;

[0047] A vocal slice data acquisition module, which is used to perform slicing processing on the clean vocal data to remove silent sounds and obtain vocal slice data;

[0048] A singing voice conversion model construction module, which is used to use the vocal slice data as a training data set to construct a singing voice conversion model, aiming to minimize the value of the loss function, and train the singing voice conversion model through the training data set to obtain a trained singing voice conversion model; wherein, the loss function includes Mel feature loss, pitch loss, and KL divergence loss;

[0049] An audio synthesis module, which is used to input the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice; wherein, the method for obtaining the final target singing voice is as follows:

[0050] Respectively perform feature extraction on the audio data to be converted to obtain embedding semantic features, pitch features, and Mel spectrogram features; process the embedding semantic features and the pitch features to extract the audio text of the speaker; process the Mel spectrogram features to obtain the timbre features of the speaker; perform audio synthesis on the timbre features and the audio text to obtain the final target singing voice.

[0051] The above technical solutions of the present invention have the following advantages compared with the prior art:

[0052] 1. Efficient vocal separation and preprocessing: The present invention finds the optimal combination method through the combined operation of multiple UVR5 models. Thus, high-quality clean vocal data is separated from the original audio data, and then conventional slicing processing is performed to remove silent sounds, etc., further effectively reducing data redundancy, improving the model training efficiency and the real-time performance of singing voice conversion, and providing a high-quality data basis for subsequent processing.

[0053] 2. Optimize the loss function for model training: Measure the difference between the audio generated by the model and the real audio from multiple perspectives, including Mel loss, pitch loss, and KL divergence loss, so as to ensure that the audio generated by the model approaches the real audio continuously as the training progresses. Meanwhile, considering the different degrees of dependence on each loss at different stages of training, an adjustment attenuation coefficient is introduced to linearly attenuate the weight of the pitch loss function in the later stage of training. By balancing the influence of different loss terms during the training process, the training process of the final model can converge smoothly.

[0054] 3. Efficient fine-tuning: Different from traditional fine-tuning of all model parameters, in the present invention, the parameters of the multi-head attention and feed-forward layers in BERT are frozen, and an Adapter module is added for training. The experimental verification results reveal the excellent performance of the present invention in enhancing audio processing efficiency, specifically reflected in the significant improvement in the model training speed. When the model loss reaches the preset threshold, compared with the traditional fine-tuning method, the average time consumed by the present invention is reduced by 15%, fully demonstrating its efficiency and practicality. Description of the Drawings

[0055] To make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention and in combination with the drawings, wherein,

[0056] Figure 1 is a flowchart of a singing voice conversion method provided in Embodiment 1 of the present invention;

[0057] Figure 2 is the training loss function curve of the singing voice conversion model in Embodiment 1 of the present invention, where (a) represents the KL divergence loss, (b) represents the mean square error (mse) loss, and (c) represents the L1 loss;

[0058] Figure 3 is the result comparison of the singing voice conversion model proposed by the present invention and other 5 existing singing voice conversion models (RVC-Flow, SoVITS-Flow, SoVITS-Diff, CoMoSVC, Amphion) in terms of performance based on 5 evaluation index values;

[0059] Figure 4 is a schematic structural diagram of a singing voice conversion system provided in Embodiment 2 of the present invention;

[0060] Explanation of the reference numerals in the drawings: 100, clean vocal data acquisition module; 200, vocal slice data acquisition module; 300, singing voice conversion model construction module; 400, audio synthesis module. Detailed Embodiments

[0061] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.

[0062] Embodiment 1

[0063] Referring to Figure 1 As shown, the present invention provides a singing voice conversion method, and the method includes the following steps:

[0064] S1: Perform vocal separation on the acquired original audio data through the UVR5 model to obtain clean vocal data;

[0065] S2: Perform slicing processing on the clean vocal data through the slicer tool to remove redundant silent sounds and obtain vocal slice data;

[0066] S3: Use the vocal slice data as a training data set to construct a singing voice conversion model, aiming to minimize the value of the loss function, and train the singing voice conversion model through the training data set to obtain a trained singing voice conversion model;

[0067] S4: Input the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice.

[0068] Furthermore, in step S1, three UVR5 models are used to perform vocal separation on the acquired original audio data to obtain clean vocal data, and the specific steps of the method are as follows:

[0069] S11: Separate the accompaniment of the original audio data through the MDX-Net Kim Vocal 2 model to obtain initial vocal data;

[0070] S12: Based on the initial vocal data, use the VR Architecture 9_HP2-UVR model to remove reverberation, and use the VR Architecture UVR-De-Echo-Aggeressive model to remove the inhalation sound and ambient noise in the singing voice to obtain the clean vocal data.

[0071] Furthermore, in step S4, input the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice, and the method steps are as follows:

[0072] S41: Respectively use the ContentVec model, the rmvpe model, and the mel spectrogram feature extraction algorithm to perform feature extraction on the audio data to be converted to obtain embedding semantic features, pitch ( f0), features and mel-spectrum features;

[0073] S42: The feature extraction main structure of the Enc_p encoder is a pre-trained BERT model. First, freeze the main parameters of the BERT model and add an Adapter module to help the model adapt to timbre learning. Process the embedded semantic features and the pitch features through the Enc_p encoder to extract the audio text of the speaker; process the mel-spectrum features through the Enc_q encoder to obtain the timbre features of the speaker;

[0074] S43: Process the timbre features through the Flow flow model, and synthesize the processed timbre features and the audio text through the Nsf_vocoder to obtain the final target singing voice.

[0075] Specifically, in step S42, the steps of extracting the audio text of the speaker are as follows:

[0076] S421: Construct an audio content extraction model, and perform iterative training on the audio content extraction model to obtain a trained audio content extraction model, including:

[0077] The audio content extraction model includes a BERT model and an Adapter module. Freeze the parameters of the multi-head attention layer and the feed-forward neural network layer in the BERT model, and only iteratively update the parameters of the Adapter module; among them, the Adapter module is composed of two layers of linear networks, a ReLU activation function, and a residual network;

[0078] S422: Use the trained audio content extraction model to process the embedded semantic features and the pitch features, and perform feature extraction through the BERT model and the Adapter module to obtain the audio text of the speaker :

[0079] ,

[0080] Among them, represents the input of the Adapter module; represents the two-dimensional weight matrix for reducing the dimension of the input ; represents the bias matrix for reducing the dimension of the input ; represents the weight matrix for converting the reduced-dimensional features into the original dimension; is the bias matrix for converting the reduced-dimensional features into the original dimension.

[0081] In this embodiment, in S42, the embedded semantic features and the pitch features are processed by the Enc_p encoder to obtain the timbre features of the speaker, and the specific steps are as follows:

[0082] S423: constructing a timbre feature extraction model based on the WaveNet network structure, wherein the timbre feature extraction model includes a stacked atrous causal convolution layer, a residual connection layer, and a gated activation module;

[0083] S424: At each time node, the mel-spectrogram feature is used as a conditional input, and the temporal dependency feature in the mel-spectrogram feature is obtained through the causal convolution layer and the gated activation module. Based on the current mel-spectrogram feature and the temporal dependency feature of the previous time node, the probability distribution value of the next sampling point is calculated, and the next sampling point is sampled;

[0084] S425: looping through step S424 to continuously generate new sampling points until a complete speech feature sequence is obtained;

[0085] S426: Extracting the timbre features of the speaker based on the speech feature sequence.

[0086] In summary, by extracting embedded semantic features, pitch features, and mel-spectrogram features and processing them separately, the audio text and timbre features of the speaker are extracted. This processing method not only achieves the conversion of timbre, but also retains the emotional and style characteristics of the source singer as much as possible, thereby improving the naturalness and realism of the generated audio.

[0087] In this embodiment, the Flow model shows wide adaptability to data distributions of various complexity, and can achieve fine control over the generation process in data generation tasks. The core of the model adopts the Residual Coupling Block module design based on the concept of Residual Network (ResNet). Its working mechanism is to divide the input data into two subsets, apply a transformation operation to one of the subsets, and then merge the transformed subset with the other untransformed subset to construct an efficient reversible mapping mechanism. This design gives the Flow model a powerful ability to process high-dimensional and complex data structures (such as audio signals).

[0088] When training the Flow model, we mainly focus on the Mel feature loss of the audio feature extraction in the enc_q module, the pitch loss of the audio content extraction in the enc_p module, and the KL divergence loss of the generator. Therefore, the expression of the loss function is:

[0089] ,

[0090] Among them, L is the value of the total loss function. represents the pitch loss. represents the KL divergence loss. represents the Mel feature loss. is the attenuation coefficient, which shows a linear decreasing trend from 1 to 0 as the number of training rounds increases. However, as the number of training rounds increases, the value of will gradually decrease, resulting in a decrease in the weight of the pitch extraction module in the total loss function. Correspondingly, the losses of the Flow model and the Nsf_vocoder in generating real audio will account for a larger proportion, thus guiding the model to pay more attention to the performance improvement of these two modules.

[0091] Among them, the Mel feature loss is calculated based on the L1_Loss function :

[0092] ,

[0093] Among them, represents the actual Mel spectrogram features of the input; represents the Mel spectrogram features of the audio generated by the model; the Mel spectrogram features help the model reduce the error of the extracted audio features during the training process, thereby improving the prediction performance of the model.

[0094] The pitch loss is constructed using the mean squared error (mse) , and its calculation method is:

[0095] ,

[0096] Among them, represents the embedded semantic features after frequency conversion; represents after f0 the predicted embedded semantic features obtained by the decoder;

[0097] The KL divergence is used to measure the difference degree of the generation results of the enc_p module, the enc_q module, and the Flow model. Therefore, the calculation method of the KL divergence loss is:

[0098] ,

[0099] Among them, and are the embedded representations of the i-th sample after being encoded by enc_p, where represents the output of the first half of the encoding result, represents the output of the second half of the encoding result; and is the embedded representation and masked representation of the i-th sample after being encoded by enc_q, , represents the embedded representation generated by the HuBERT model for the i-th audio sample; represents the embedded representation obtained through the flow model, .

[0100] Collect the audio data of 4 singers and the audio data from the publicly available high-quality Mandarin singing corpus Opencpop as the audio dataset for training. The specific content of each audio dataset is shown in Table 1 below.

[0101] Table 1

[0102]

[0103] According to the above technical solution, first, separately perform voice separation on the above-obtained different types of audio datasets through the UVR5 model to obtain clean voice data; second, perform slicing processing on the clean voice data through the slicer tool to remove redundant silent sounds and obtain voice slice data; then, use the voice slice data as the training dataset to construct a singing voice conversion model, aiming to minimize the value of the loss function, and train the singing voice conversion model through the training dataset to obtain the trained singing voice conversion model; finally, input the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice.

[0104] When using the singing voice conversion method of the present invention to perform model training on different types of audio datasets, the average training duration is reduced by 15%, which means the significant advantage of the present invention in improving audio processing efficiency. Specifically, for the audio dataset of singer 1, the model training time is reduced by 16.3% using the singing voice conversion method of the present invention; for the audio dataset of singer 2, the model training time is reduced by 14.8% using the singing voice conversion method of the present invention; for the audio dataset of singer 3, the model training time is reduced by 14.2% using the singing voice conversion method of the present invention; for the audio dataset of singer 4, the model training time is reduced by 15.4% using the singing voice conversion method of the present invention; for the audio dataset of Opencpop, the model training time is reduced by 16.7% using the singing voice conversion method of the present invention.

[0105] As Figure 2 shown, three loss curves are adopted - KL divergence loss, mse loss (i.e., the " " mentioned above), L1 loss (i.e., the " ”), to comprehensively evaluate the performance of the singing voice conversion model proposed in the present invention. By observing Figure 2 the experimental results in (a) to (c) below, the following conclusions can be drawn:

[0106] After the training process progresses to a certain stage, the loss value gradually stabilizes, which indicates that the model has successfully reached the convergence state after balancing the influences of different loss terms, that is, a (local) optimal solution has been found. The realization of this state benefits from the fine control of the training process by the optimized loss function in the present invention.

[0107] Among them, the mse loss and the L1 loss are used to quantify the gap between the reconstructed audio pitch features and Mel features and the real features, so as to guide the model decoder to generate higher-quality sounds. As the number of training epochs increases, both of these loss values show a trend of gradually decreasing and stabilizing. On the one hand, this indicates that the generation performance of the model is continuously improving, and on the other hand, it also provides a basis for stopping training at an appropriate training inflection point in order to achieve the expected effect by optimizing the training duration.

[0108] The KL divergence loss, as a measurement tool for measuring the difference between two probability distributions, is used to compare the difference between the latent distributions generated by the encoder enc_q, enc_p and the flow module and the standard normal distribution (mean is 0, variance is 1). By introducing the KL divergence loss, the distribution structure of the latent space can be effectively constrained, thereby improving the generalization ability of the model. During the actual training process, the KL divergence loss shows a trend of slightly fluctuating and decreasing, which fully proves that the generalization ability of the model in generating audio is continuously enhanced.

[0109] However, when selecting the appropriate number of iterations to generate the final model, the actual synthesis quality of the audio also needs to be comprehensively considered. Therefore, the present invention combines the mse loss and the L1 loss for comprehensive evaluation to ensure that the obtained model not only has excellent performance but also can meet the requirements of actual application scenarios.

[0110] To comprehensively evaluate the effect of singing voice conversion, STOI (Short-Time Objective Intelligibility), PESQ (Perceptual Evaluation of Speech Quality), FPC (F0 Pearson Correlation, used to evaluate pitch accuracy), CER (Character Error Rate, reflecting audio quality and clarity), and SIM (Similarity Index) are used as evaluation indicators to quantitatively analyze the converted singing voice from different dimensions.

[0111] exist Figure 3 In this paper, five existing singing voice conversion models, namely RVC-Flow, SoVITS-Flow, SoVITS-Diff, CoMoSVC and Amphion, are selected and compared with the singing voice conversion model proposed in the present invention. Among them, the STOI index is used to measure the clarity of the audio. The closer its value is to 1, the closer the clarity is to the original real audio; the PESQ index is used to evaluate the voice quality. The higher its value is, the better the quality of the generated audio is; the FPC index evaluates whether the audio is out of tune by calculating the F0 Pearson correlation coefficient. The larger its value is, the more similar it is to the real audio pitch; the CER index obtains the text information of the audio through the whisper algorithm and calculates the character error rate. The smaller its value is, the higher the quality and intelligibility of the generated audio is; the SIM index is expressed by the embedding of the generated audio through expnet2, and the cosine distance is calculated to evaluate the similarity between the generated audio and the real audio. The larger its value is, the higher the similarity is.

[0112] from Figure 3 The evaluation results show that the singing voice conversion model proposed in the present invention has shown excellent performance in all evaluation indicators, and most indicators are better than other comparison models. This result fully proves the advancement and practicality of the present invention in the field of singing voice conversion, and provides new ideas and methods for the development of singing voice conversion technology.

[0113] Embodiment 2

[0114] Based on the same inventive concept as the method described in Example 1, the present invention also provides a singing voice conversion system. Figure 4 As shown, the system includes the following modules:

[0115] The clean vocal data acquisition module 100 is used to separate the vocals from the acquired original audio data to obtain clean vocal data;

[0116] The vocal slice data acquisition module 200 is used to slice the clean vocal data to remove the silent sound and obtain the vocal slice data;

[0117] The singing voice conversion model construction module 300 is used to use the vocal slice data as a training data set to construct a singing voice conversion model, and to train the singing voice conversion model through the training data set with the goal of minimizing the value of the loss function to obtain a trained singing voice conversion model; wherein the loss function includes Mel feature loss, pitch loss and KL divergence loss;

[0118] The audio synthesis module 400 is used to input the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice. The method for obtaining the final target singing voice is as follows:

[0119] Extract the feature vectors, pitch features, and Mel spectrogram features from the audio data to be converted respectively; process the feature vectors and pitch features to extract the audio text of the speaker; process the Mel spectrogram features to obtain the timbre features of the speaker; synthesize the timbre features and the audio text to obtain the final target singing voice.

[0120] A singing voice conversion system proposed in this embodiment is used to implement the foregoing singing voice conversion method. Therefore, the specific implementation manners in the singing voice conversion system can be seen in the embodiment part of the foregoing singing voice conversion method. For example, the clean vocal data acquisition module 100, the vocal slice data acquisition module 200, the singing voice conversion model construction module 300, and the audio synthesis module 400 are respectively used to implement steps S1, S2, S3, and S4 in the singing voice conversion method in Embodiment 1. Therefore, the specific implementation manners can refer to the descriptions of the corresponding individual embodiments. To avoid redundancy, they will not be elaborated here.

[0121] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0122] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more processes and / or blocks Figure 1 specified in the function.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more processes and / or blocks Figure 1 specified in the function of the block or blocks.

[0125] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A singing voice conversion method, characterized in that, It includes the following steps: S1: Perform voice separation on the obtained original audio data to obtain clean voice data; S2: Perform slicing processing on the clean voice data, remove silent sounds, and obtain voice slice data; S3: Use the voice slice data as a training data set to construct a singing voice conversion model. With the goal of minimizing the value of the loss function, train the singing voice conversion model through the training data set to obtain a trained singing voice conversion model; among them, the loss function includes Mel feature loss, pitch loss, and KL divergence loss; S4: Input the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice, including: S41: Respectively perform feature extraction on the audio data to be converted to obtain embedded semantic features, pitch features, and Mel spectrogram features; S42: Process the embedded semantic features and the pitch features to extract the audio text of the speaker; process the Mel spectrogram features to obtain the timbre features of the speaker; S43: Synthesize the timbre features and the audio text to obtain the final target singing voice; Among them, in S42, the steps to extract the audio text of the speaker are as follows: S421: Construct an audio content extraction model, perform iterative training on the audio content extraction model to obtain a trained audio content extraction model, including: The audio content extraction model includes a BERT model and an Adapter module. Freeze the parameters of the multi-head attention layer and the feed-forward neural network layer in the BERT model, and only iteratively update the parameters of the Adapter module; S422: Use the trained audio content extraction model to process the embedded semantic features and the pitch features, and perform feature extraction through the BERT model and the Adapter module to obtain the audio text of the speaker; The steps to obtain the timbre features of the speaker are as follows: S423: Construct a timbre feature extraction model, and the timbre feature extraction model includes stacked causal convolutional layers, residual connection layers, and gated activation modules; S424: At each time node, obtain the temporal dependence features in the Mel spectrogram features through the causal convolutional layer and the gated activation module. Based on the current Mel spectrogram features and the temporal dependence features of the previous time nodes, calculate the probability distribution value of the next sampling point, and sample to obtain the next sampling point; S425: Loop and execute step S424 to continuously generate new sampling points until a speech feature sequence is obtained; S426: Based on the speech feature sequence, extract the timbre features of the speaker.

2. The singing voice conversion method according to claim 1, characterized in that In S1, the steps to obtain clean voice data are as follows: S11: Separate the accompaniment of the original audio data to obtain initial voice data; S12: Based on the initial voice data, remove reverberation, inhalation sounds, and environmental noises to obtain the clean voice data.

3. The singing voice conversion method according to claim 1, wherein In S422, the audio text x of the speaker is obtained o : x o = x i + ReLU(W d x i + b d )W u + b u Among them, x i represents the input of the Adapter module; W d represents the two-dimensional weight matrix for dimensionality reduction of the input x i ; b d represents the bias matrix for dimensionality reduction of the input x i ; W u represents the weight matrix for converting the dimensionality-reduced features back to the original dimension; b u is the bias matrix for converting the dimensionality-reduced features back to the original dimension.

4. The singing voice conversion method according to claim 3, wherein: The Adapter module is composed of two layers of linear networks, a ReLU activation function, and a residual network.

5. The singing voice conversion method according to claim 1, wherein The expression of the loss function is: L = αf0_Loss + kl_Loss + mel_Loss Among them, L is the total loss function value, α is the attenuation coefficient, f0_Loss represents the pitch loss, kl_Loss represents the KL divergence loss, and mel_Loss represents the Mel feature loss.

6. The singing voice conversion method according to claim 5, characterized in that: The calculation method of the Mel feature loss Mel_Loss is as follows: Among them, represents the actual mel spectrogram features of the input; represents the mel spectrogram features of the audio generated by the model; The calculation method of the pitch loss f0_Loss is as follows: Among them, lf0 i represents the embedded semantic features after frequency conversion; pred_lf0 i represents the predicted embedded semantic features obtained by the f0 decoder for lf0 i ; n represents the number of samples; The calculation method of the KL divergence loss kl_Loss is as follows: Among them, m_p i and logs_p i are the embedded representations of the i-th sample after being encoded by enc_p, where m_p i represents the output of the first half of the encoding result, and logs_p i represents the output of the second half of the encoding result; logs_q i and z_mask i are the embedded representation and the mask representation of the i-th sample after being encoded by enc_q, logs_q i = enc_q(spec_i), where spec_i represents the embedded representation of the i-th audio sample generated by the HuBERT model; z_p i represents the embedded representation obtained through the flow model, z_p i = flow(z_i, z_mask i ).

7. A singing voice conversion system, characterized in that, To implement the steps of the singing voice conversion method according to any one of claims 1 to 6, the system includes the following modules: A clean vocal data acquisition module, which is used to perform vocal separation on the acquired original audio data to obtain clean vocal data; A vocal slice data acquisition module, which is used to perform slicing processing on the clean vocal data to remove silent sounds and obtain vocal slice data; A singing voice conversion model construction module, which is used to use the vocal slice data as a training data set to construct a singing voice conversion model, aiming to minimize the value of the loss function, and train the singing voice conversion model through the training data set to obtain a trained singing voice conversion model; among them, the loss function includes Mel feature loss, pitch loss and KL divergence loss; An audio synthesis module, which is used to input the audio data to be converted into the trained singing voice conversion model to obtain the final target singing voice; among them, the method for obtaining the final target singing voice is as follows: Respectively perform feature extraction on the audio data to be converted to obtain embedded semantic features, pitch features and Mel spectrogram features; process the embedded semantic features and the pitch features to extract the audio text of the speaker; process the Mel spectrogram features to obtain the timbre features of the speaker; perform audio synthesis on the timbre features and the audio text to obtain the final target singing voice.

Citation Information

Patent Citations

  • Neural machine translation method fusing Bert pre-training language knowledge

    CN117313750A

  • Song timbre conversion method, computer equipment and storage medium

    CN117672241A