Training method of singing sound conversion system, method for generating audio and related device

By introducing a tone encoder and a tone perception attention mechanism module in the singing conversion system, using multiple reference audio to encode the target tone, the problem of poor tone conversion effect in the prior art is solved, and the similarity of the tone after synthesis is significantly improved.

CN119993117AActive Publication Date: 2025-05-13TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510235997.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-13
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

When the existing zero-sample singing voice conversion system synthesizes audio, the similarity between the tone and the target tone is low, and the tone characteristics of the song audio to be converted cannot be effectively retained.

Method used

The timbre encoder and timbre perception attention mechanism module are added to the singing voice conversion system. By inputting multiple reference audios to encode the target tone, the target tone adapted to the song to be converted is learned and synthesized.

Benefits of technology

The similarity between the timbre of the synthesis and the timbre of the timbre to be converted is improved, ensuring the conversion effect of the timbre.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993117A_ABST
    Figure CN119993117A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a training method of a singing sound conversion system, a method for generating audio based on the singing sound conversion system and a related device, which are used for improving the similarity between the timbre of a synthesized singing sound and the timbre of a to-be-converted singing sound. The method provided by the embodiment of the invention comprises the following steps: acquiring a plurality of reference audios of a first target timbre, and inputting the plurality of reference audios into a timbre encoder to obtain a timbre encoding vector; inputting the phoneme posterior probability and the fundamental frequency of the to-be-converted singing sound into a text encoder to obtain a prior distribution parameter of the to-be-converted singing sound content; sampling according to the prior distribution parameter to obtain a text sampling value vector of the singing content to be converted; inputting the text sampling value vector and the timbre coding vector into a timbre perception attention mechanism module to determine a new timbre coding vector; and taking the new tone coding vector as new input added in the singing sound conversion system, calculating the reconstruction loss of the singing sound conversion system, and training the singing sound conversion system according to the reconstruction loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio data processing, and in particular to a conversion sample method of a singing voice conversion system, a method for generating audio based on the singing voice conversion system, and related devices. Background Art

[0002] Singing voice conversion refers to the technology of inputting a singing voice audio to be converted and an audio with a target timbre, and converting the timbre of the singing voice audio to be converted to the target timbre while retaining the singing content in the singing voice audio to be converted. This singing voice conversion system can help people easily synthesize any song with a specified timbre for use in the tuning function of karaoke software.

[0003] The prior art generally uses a zero-shot singing voice conversion system (Zero-shot SVC) for singing voice conversion. This zero-shot singing voice conversion system does not use target timbre data to train the singing voice conversion system, but directly inputs the target timbre data in the inference stage to convert the timbre of the singing voice to be converted.

[0004] However, since the singing voice conversion model is not fine-tuned before being used in the zero-sample singing voice conversion system, the timbre of the synthesized audio obtained by using the zero-sample singing voice conversion system has a low similarity with the target timbre of the singer to be converted. Summary of the invention

[0005] The embodiments of the present invention provide a training method for a singing voice conversion system, a method for generating audio based on the singing voice conversion system, and related devices. A timbre encoder and a timbre perception attention mechanism module are added to the singing voice system, and multiple reference audios of a first target timbre are input into the timbre encoder. The multiple reference audios include multiple audios of the singer of the singing voice to be converted in different vocal ranges. Because the resonance methods used for singing in different vocal ranges are different, the timbre of the singer of the singing voice to be converted in the multiple reference audios is also different. Therefore, the singing voice conversion system after converting the sample can learn the target timbre suitable for the singing voice to be converted from the multiple reference audios according to the singing content representation vector of the singing voice to be converted, and then synthesize the timbre of the singing voice to be converted according to the target timbre suitable for the singing voice to be converted, thereby improving the similarity between the timbre of the synthesized singing voice and the timbre of the singing voice to be converted.

[0006] In a first aspect, an embodiment of the present application provides a training method for a singing voice conversion system, wherein the singing voice conversion system comprises at least a text encoder, a timbre encoder, and a timbre perception attention mechanism module, and the method comprises:

[0007] Acquire multiple reference audios of the first target timbre, wherein the multiple reference audios at least include multiple audios of the singer to be converted in different voice ranges;

[0008] Inputting the plurality of reference audios into the timbre encoder to perform timbre encoding on the plurality of reference audios, and obtaining timbre encoding vectors for the plurality of reference audios output by the timbre encoder;

[0009] Inputting the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the song content to be converted output by the text encoder;

[0010] According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the song content to be converted;

[0011] Input the text sampling value vector of the singing content to be converted and the timbre coding vectors of the multiple reference audios into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of the multiple reference audios and the text sampling value vector of the singing content to be converted;

[0012] Using the new timbre coding vector as a new input added to the singing voice conversion system, and calculating the reconstruction loss of the singing voice conversion system;

[0013] The singing voice conversion system is trained according to reconstruction loss and back propagation algorithm.

[0014] A second aspect of an embodiment of the present application provides a method for generating audio based on a singing voice conversion system, wherein the singing voice conversion system comprises at least a text encoder, a timbre encoder, a timbre perception attention mechanism module and a decoder, and the method comprises:

[0015] Acquire multiple reference audios of a second target timbre, wherein the multiple reference audios of the second target timbre at least include multiple audios of different singers in different vocal ranges;

[0016] Inputting the plurality of reference audios into the timbre encoder to perform timbre encoding on the plurality of reference audios, and obtaining timbre encoding vectors for the plurality of reference audios output by the timbre encoder;

[0017] Inputting the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the song content to be converted output by the text encoder;

[0018] According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the song content to be converted;

[0019] Input the text sampling value vector of the singing content to be converted and the timbre coding vectors of the multiple reference audios into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of the multiple reference audios and the text sampling value vector of the singing content to be converted;

[0020] The text sampling value vector of the singing content to be converted and the new timbre coding vector are used as inputs of the decoder to obtain synthesized singing audio.

[0021] A third aspect of an embodiment of the present application provides a computer device, comprising a processor, which, when executing a computer program stored in a memory, is used to implement the training method for the singing voice conversion system provided in the first aspect of the embodiment of the present application, or the method for generating audio based on the singing voice conversion system provided in the second aspect of the embodiment of the present application.

[0022] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the training method of the singing voice conversion system provided in the first aspect of the embodiments of the present application, or the method for generating audio based on the singing voice conversion system provided in the second aspect of the embodiments of the present application.

[0023] The fifth aspect of the embodiments of the present application provides a computer program product on which a computer program is stored, characterized in that when the computer program is executed by a processor, it is used to implement the training method of the singing voice conversion system provided in the first aspect of the embodiments of the present application, or the method for generating audio based on the singing voice conversion system provided in the second aspect of the embodiments of the present application.

[0024] It can be seen from the above technical solutions that the embodiments of the present invention have the following advantages:

[0025] The embodiment of the present application provides a training method for a singing voice conversion system, wherein the singing voice conversion system includes at least a text encoder, a timbre encoder and a timbre perception attention mechanism module, and the method includes: obtaining multiple reference audios of a first target timbre, wherein the multiple reference audios include at least multiple audios of a singer of a singing voice to be converted in different voice ranges; inputting the multiple reference audios into the timbre encoder to perform timbre encoding on the multiple reference audios, and obtaining timbre encoding vectors of the multiple reference audios output by the timbre encoder; inputting the PPG phoneme posterior probability of the singing voice to be converted and the Pitch fundamental frequency of the singing voice to be converted into the text encoder, and obtaining the timbre encoding vectors of the singing voice to be converted output by the text encoder; Prior distribution parameters of the content; according to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted to obtain the text sampling value vector of the song content to be converted; the text sampling value vector of the song content to be converted and the timbre coding vectors of multiple reference audios are input into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of multiple reference audios and the text sampling value vector of the song content to be converted; the new timbre coding vector is used as a new input added to the singing voice conversion system, the reconstruction loss of the singing voice conversion system is calculated, and the singing voice conversion system is trained according to the reconstruction loss and the back propagation algorithm.

[0026] Because the embodiment of the present application adds a timbre encoder and a timbre perception attention mechanism module to the singing voice conversion system, and inputs multiple reference audios of the first target timbre into the timbre encoder, and the multiple reference audios include at least multiple audios of the singer to be converted in different pitch ranges, and because the resonance methods used for singing in the high pitch range and the low pitch range are different, the timbre of the singer to be converted in the multiple reference audios is also different. Therefore, the singing voice conversion system after the conversion sample can learn the target timbre suitable for the singing voice to be converted from the multiple reference audios according to the singing content representation vector of the singing voice to be converted, and then synthesize the timbre of the singing voice to be converted according to the target timbre suitable for the singing voice to be converted, thereby improving the similarity between the timbre of the synthesized singing voice and the timbre of the singing voice to be converted. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 A schematic diagram of the system architecture of a training singing voice conversion system provided in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of an embodiment of a training method for a singing voice conversion system in an embodiment of the present application;

[0029] Figure 3 This is a schematic diagram of the architecture of the singing voice conversion system in the embodiment of the present application;

[0030] Figure 4This is a schematic diagram of the architecture of a singing voice conversion system based on the VITS variational autoencoder in an embodiment of the present application;

[0031] Figure 5 A schematic diagram of the training process of a singing voice conversion system based on a VITS variational autoencoder in an embodiment of the present application;

[0032] Figure 6 Another schematic diagram of the architecture of the singing voice conversion system based on the VITS variational autoencoder in the embodiment of the present application;

[0033] Figure 7 Schematic diagram of another training process of a singing voice conversion system based on a VITS variational autoencoder in an embodiment of the present application;

[0034] Figure 8 Schematic diagram of the architecture of the Fastspeeh fast speech synthesis model in the embodiment of the present application;

[0035] Fig. 9 A schematic diagram of the training process of the Fastspeeh fast speech synthesis model in an embodiment of the present application;

[0036] Fig.10 Another schematic diagram of the Fastspeeh fast speech synthesis model in the embodiment of the present application;

[0037] Fig.11 This is a schematic diagram of another training process of the Fastspeeh fast speech synthesis model in an embodiment of the present application;

[0038] Fig.12 This is another model architecture diagram of the Diffusion diffusion generation model in the embodiment of the present application;

[0039] Fig.13 This is another schematic diagram of a training process based on a Diffusion generation model in an embodiment of the present application;

[0040] Fig.14 A schematic diagram of an embodiment of a method for generating audio based on a singing voice conversion system in an embodiment of the present application;

[0041] Fig.15 This is a schematic diagram of the architecture of the singing voice conversion system in the embodiment of the present application;

[0042] Fig.16 This is a schematic diagram of the architecture of a singing voice conversion system based on the VITS variational autoencoder in an embodiment of the present application;

[0043] Fig.17 A schematic diagram of an embodiment of a method for generating audio by a singing voice conversion system based on a VITS variational autoencoder in an embodiment of the present application;

[0044] Fig.18 This is a schematic diagram of an embodiment of a method for generating audio based on the Fastspeeh fast speech synthesis model or based on the Diffusion diffusion generation model in an embodiment of the present application. DETAILED DESCRIPTION

[0045] The embodiments of the present invention provide a training method for a singing voice conversion system, a method for generating audio based on the singing voice conversion system, and related devices, which are used to improve the similarity between the timbre of the synthesized singing voice and the timbre of the singing voice to be converted.

[0046] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0047] The terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0048] The embodiment of the present application provides a first aspect of a training method for a singing voice conversion system, which adds a timbre encoder and a timbre perception attention mechanism module to the original singing voice conversion system. The general principle of the training method is: multiple reference audios of a first target timbre are input into the timbre encoder to perform timbre encoding on the multiple reference audios, and timbre encoding vectors for the multiple reference audios output by the timbre encoder are obtained, wherein the multiple reference audios include at least multiple audios of the singer of the singing voice to be converted in different pitch ranges, and then the PPG phoneme posterior probability of the singing voice to be converted and the Pitch fundamental frequency of the singing voice to be converted are input into the text encoder in the singing voice conversion system to obtain the timbre encoding vector output by the text encoder. The prior distribution parameters of the singing content to be converted are sampled from the PPG phoneme posterior probability of the singing content to be converted and the Pitch fundamental frequency of the singing content to be converted according to the prior distribution parameters to obtain the text sampling value vector of the singing content to be converted, and finally the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios are input into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of multiple reference audios and the text sampling value vector of the singing content to be converted, and the new timbre coding vector is used as a new input added to the singing conversion system, the reconstruction loss of the singing conversion system is calculated, and then the singing conversion system is trained according to the reconstruction loss and the back propagation algorithm. The training method for the singing voice conversion system provided in the embodiment of the present application adds a timbre encoder and a timbre perception attention mechanism module to the singing voice conversion system, so that the timbre encoder and the timbre perception attention mechanism module can determine a new timbre coding vector based on the timbre coding vectors of multiple reference audios and the text sampling value vector of the singing content to be converted, and use the new timbre coding vector as a new input of the singing voice conversion system to train the singing voice conversion system, so that the trained singing voice conversion system can learn the target timbre suitable for the singing content to be converted according to the text sampling value vector of the singing content to be converted and the multiple reference audios, and then synthesize the singer's singing content according to the target timbre suitable for the singing content to be converted, thereby improving the similarity between the timbre of the synthesized singing and the timbre of the singer's singing.

[0049] In order to better implement the training method of the singing voice conversion system, the present application embodiment provides a system for training the singing voice conversion system. Figure 1 , Figure 1A schematic diagram of the system architecture of a training singing voice conversion system provided for an embodiment of the present application. The system of the training singing voice conversion system may include at least one terminal device 101 and a server 102; different types of applications may be installed on the terminal device 101, for example, a karaoke program, an instant messaging application, a live broadcast application, a conference communication application, etc. may be installed on the terminal device 101; the terminal device 101 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart car, etc. The server 102 may be used to store application data and audio data generated by different types of applications of the terminal device 101. The server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms, etc.

[0050] Wherein, the training method of the singing voice conversion system is executed by the terminal device 101 or the server 102. When the training method of the singing voice conversion system is executed by the terminal device 101, the audio data generated by the terminal device 101 in different types of applications can be included in the server. Then, when the terminal device 101 needs to train the singing voice conversion system, the terminal device 101 can obtain multiple reference audios of the samples to be converted from the server 102, wherein the multiple reference audios at least include multiple audios of the singer of the singing voice to be converted in different voice zones. Then, after the terminal device 101 obtains the multiple reference audios of the first target timbre from the server 102, the multiple reference audios can be input into the timbre encoder to obtain the timbre coding vectors of the multiple reference audios output by the timbre encoder, and then Then, the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted are input into the text encoder in the singing voice conversion system to obtain the prior distribution parameters of the singing voice content to be converted output by the text encoder, and according to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted to obtain the text sampling value vector of the singing voice content to be converted, and finally, the text sampling value vector of the singing voice content to be converted and the timbre coding vectors of multiple reference audios are input into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of multiple reference audios and the text sampling value vector of the singing voice content to be converted, and the new timbre coding vector is used as a new input added by the singing voice conversion system to train the singing voice conversion system.

[0051] For ease of understanding, the following describes the training method of the singing voice conversion system in the embodiment of the present application, wherein the training method can be applied to a terminal device or a server, and can also be applied to a small program installed in the terminal device or the server. The following describes the training process of the singing voice conversion system in detail, please refer to Figure 2 , Figure 2 This is a schematic diagram of an embodiment of the training method of the singing voice conversion system in the embodiment of the present application:

[0052] Specifically, the singing voice conversion system in the embodiment of the present application adds a timbre encoder and a timbre perception attention mechanism module to the original singing voice conversion system. For ease of understanding, Figure 3 A schematic diagram of the architecture of the singing voice conversion system is given.

[0053] 201. Acquire multiple reference audios of a first target timbre, wherein the multiple reference audios at least include multiple audios of a singer of a to-be-converted singing voice in different voice ranges;

[0054] Different from the prior art that does not fine-tune the original singing voice conversion model, resulting in a problem in which the timbre of the synthesized audio obtained by the original singing voice conversion system is less similar to the target timbre of the singer to be converted, the embodiment of the present application adds a timbre encoder and a timbre perception attention mechanism module to the original singing voice conversion system, and uses multiple reference audios of the first target timbre as conversion sample inputs of the timbre encoder and the timbre perception attention mechanism module.

[0055] Specifically, in the embodiment of the present application, the multiple reference audios of the first target timbre are multiple audios in different vocal ranges of the singer of the song to be converted. For example, when the song to be converted is the song sung by user A, the multiple reference audios of the first target timbre are multiple audios in different vocal ranges sung by user A. Of course, if the song to be converted is the song sung by user B, the multiple reference audios of the first target timbre are multiple audios in different vocal ranges sung by user B. No specific restriction is made on the singer of the song to be converted here.

[0056] Furthermore, the multiple reference audios here are songs sung or spoken content by the singer of the song to be converted, because the same user uses different resonance methods when singing in different vocal ranges (such as high pitch, middle pitch and bass range), and the timbre is also different. In order to select the best timbre suitable for the song content to be converted, it is generally required that the timbre of the multiple reference audios be as rich as possible. Therefore, the multiple reference audios in this application at least include multiple audios of the singer of the song to be converted in different vocal ranges. For example, when the sample to be converted is user A, the multiple reference audios here include multiple audios of user A in different vocal ranges, such as user A's audio in the high pitch range, audio in the middle pitch range and audio in the bass range.

[0057] It should be noted here that the treble audio and bass audio here can be the treble part and bass part of the same song sung by the singer of the song to be converted, or the treble part and bass part of different songs sung by the singer of the song to be converted, and the treble audio and bass audio in the embodiment of the present application can be the same song as the song to be converted, or can be a different song from the song to be converted, and no specific limitation is made here.

[0058] 202. Input the plurality of reference audios into the timbre encoder to perform timbre encoding on the plurality of reference audios, and obtain timbre encoding vectors for the plurality of reference audios output by the timbre encoder;

[0059] After obtaining the multiple reference audios, the multiple reference audios are input into a timbre encoder to perform timbre encoding on the multiple reference audios, and timbre encoding vectors of the multiple reference audios output by the timbre encoder are obtained.

[0060] Specifically, the timbre encoder in the embodiment of the present application can be a stacked multi-head attention mechanism to simultaneously focus on different parts of the input, and capture different levels of abstract information through multiple independent attention heads, and each head has its own weight matrix to capture different feature combinations of the input respectively. The parallel processing capability of this multi-head attention mechanism improves the global perception ability of the model and can better understand and process contextual relationships.

[0061] 203. Inputting the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the song content to be converted output by the text encoder;

[0062] In the embodiment of the present application, the singing-voice-conversion model (SVC) refers to a model that can retain the singing content of a singing audio sample to be converted and an audio of a target timbre in the output audio and convert the timbre of the singing into the target timbre. In the embodiment of the present application, the singing-voice-conversion model can be a FastSpeech model, a VITS model or a diffusion model, etc., and there is no specific limitation on the singing-voice-conversion model here.

[0063] Furthermore, the singing voice conversion model includes at least a text encoder. Therefore, the embodiment of the present application inputs the phonetic posterior grams (PPG) of the singing voice to be converted and the fundamental frequency (Pitch) of the singing voice to be converted into the text encoder to obtain the prior distribution parameters of the singing voice content to be converted output by the text encoder. The prior distribution parameters here are functions used to predict the timbre and fundamental frequency distribution. When the prior distribution is a normal distribution, the prior distribution parameters can be the mean and variance.

[0064] Furthermore, the text encoder in the embodiment of the present application can be a multi-head attention mechanism module with the same architecture as the timbre encoder.

[0065] 204. According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the content of the song to be converted;

[0066] After obtaining the prior distribution parameters, the prior distribution parameters can be used to sample from the phoneme posterior probability (PPG) of the song to be converted and the fundamental frequency (Pitch) of the song to be converted to obtain a text sampling value vector of the song content to be converted.

[0067] The singing voice to be converted in the embodiment of the present application may be multiple segments of the same song sung by the singer of the singing voice to be converted, or multiple segments of different songs sung by the singer of the singing voice to be converted. Further, after obtaining the singing voice to be converted, the embodiment of the present application also needs to extract the phoneme posterior probability (PPG) of the singing voice to be converted and the fundamental frequency (Pitch) of the singing voice to be converted. When proposing the phoneme posterior probability (PPG), the singing voice to be converted may be input into the Whisper model, the Hubert model or the Librosa model to obtain the phoneme posterior probability (PPG) of the singing voice to be converted, and when extracting the fundamental frequency (Pitch) of the singing voice to be converted, the singing voice to be converted may be input into the singing melody extraction model based on the recurrent neural network of Rmvte to extract the Pitch fundamental frequency of the singing voice to be converted.

[0068] It should be noted here that when using the above-mentioned model to extract the phoneme posterior probability (PPG) of the song to be converted and the fundamental frequency (Pitch) of the song to be converted, the extracted phoneme posterior probability (PPG) and the fundamental frequency (Pitch) of the song to be converted are both in units of audio frames. Therefore, in the embodiment of the present application, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted according to the prior distribution parameters, and the obtained text sampling value vector is also a frame-level singing content encoding.

[0069] 205. Inputting the text sampling value vector of the singing content to be converted and the timbre coding vectors of the multiple reference audios into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of the multiple reference audios and the text sampling value vector of the singing content to be converted;

[0070] After obtaining the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios, the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios are input into the timbre perception attention mechanism module, wherein the timbre perception attention mechanism module in the embodiment of the present application is the attention mechanism in the transformer, wherein the formula of the attention mechanism is:

[0071]

[0072] Among them, the content encoding vector is Q, is the dimension of the timbre coding vector K, and the timbre coding vectors are K and V respectively.

[0073] In combination with the above formula, the process of determining a new timbre coding vector based on the timbre coding vectors of multiple reference audios and the text sampling value vector of the singing content to be converted is described below:

[0074] 1. Based on the description of step 204, it can be known that the text sampling value vector of the singing content to be converted is used to characterize the content coding of the singing, that is, the text sampling value vector of the singing content to be converted is Q in the above formula, and the timbre coding vectors of the multiple reference audios are respectively used as K and V in the above formula;

[0075] 2. In the above formula, the attention score is calculated first. The attention score here is mainly used to calculate the similarity between vector Q and vector K. The attention score is as follows:

[0076] in, The dimension of the timbre encoding vector is used to scale the dot product to prevent the softmax input value from being too large, resulting in a small gradient.

[0077] 3. After obtaining the attention score, calculate the weight of the attention score. Assuming that the weight of the attention score is α, calculate α according to the following formula:

[0078] α=softmax(score(Q,K))

[0079] 4. Weighted summation: After obtaining the attention weight, the timbre coding vector is weighted summed according to the attention weight to obtain a new timbre coding vector. The weighted summation formula is as follows:

[0080] Attention(Q,K,V)=αV

[0081] Through the above process, a new timbre coding vector can be determined based on the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios.

[0082] That is, after the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios are input into the timbre perception attention mechanism module, the timbre perception attention mechanism module can determine a new timbre coding vector based on the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios.

[0083] 206. The new timbre coding vector is used as a new input added to the singing voice conversion system, the reconstruction loss of the singing voice conversion system is calculated, and the singing voice conversion system is trained according to the reconstruction loss and the back propagation algorithm.

[0084] After obtaining the new timbre coding vector output by the timbre perception attention mechanism module, the new timbre coding vector is used as the new input added to the singing voice conversion system, the reconstruction loss of the singing voice conversion system is calculated, and the singing voice conversion system is trained based on the reconstruction loss and the back propagation algorithm.

[0085] Specifically, when calculating the reconstruction loss of the singing voice conversion system, the embodiment of the present application calculates the loss between the Mel-spectrogram of the singing voice to be converted and the Mel-spectrogram of the synthesized audio, and the horizontal axis of the Mel-spectrogram is time (used to reflect the singing time of the singing voice to be converted), and the vertical axis is frequency (used to reflect the pitch of the singing voice to be converted). In order to reduce the loss between the Mel-spectrogram of the singing voice to be converted and the Mel-spectrogram of the synthesized audio during the training of the singing voice conversion system, the embodiment of the present application uses the same singing voice to be converted and multiple reference audios when training the singing voice conversion system. Human audio, that is, if the song to be converted is sung by user A, then the multiple reference audios are also multiple vocal range fragments sung by user A, and it is easy to understand that what the singing voice conversion system learns after training is the ability to generate the timbre of synthetic audio using the timbre of multiple reference audios. For example, in the model inference stage, when the song to be converted is the audio sung by user A, and the multiple reference audios are the audio sung by the target singer (such as Liu Xhua), the generated synthetic audio is the content of the song to be converted sung by user A using the timbre of the target singer (Liu Xhua).

[0086] Specifically, because the singing voice conversion system adds a timbre encoder and a timbre perception attention mechanism module to the original singing voice conversion, and the original singing voice conversion system can be a FastSpeech model, a VITS model or a diffusion model, etc., when the singing voice conversion system is converted using the new timbre coding vector, the training method used is different depending on each model. The following describes the training process of the singing voice conversion system when the singing voice conversion system includes different original singing voice conversion systems, which will not be repeated here.

[0087] In the embodiment of the present application, because a timbre encoder and a timbre perception attention mechanism module are added to the singing system, and multiple reference audios of the first target timbre are input into the timbre encoder, and the multiple reference audios include multiple audios of the singer of the singing voice to be converted in different pitch ranges, because the resonance methods used for singing in the high pitch range and the low pitch range are different, the timbre of the singer of the singing voice to be converted in the multiple reference audios is also different, so that the trained singing voice conversion system can learn the target timbre suitable for the singing voice to be converted from the multiple reference audios according to the singing content representation vector of the singing voice to be converted, and then synthesize the timbre of the singing voice to be converted according to the target timbre suitable for the singing voice to be converted, thereby improving the similarity between the timbre of the synthesized singing voice and the timbre of the singing voice to be converted.

[0088] The following describes the training process of the singing voice conversion system when the singing voice conversion system includes different singing voice conversion systems:

[0089] 1. If the singing voice conversion system includes the VITS singing voice conversion system based on variational autoencoder;

[0090] Specifically, when the singing voice conversion system includes the VITS singing voice conversion system based on variational autoencoder, for ease of understanding, Figure 4 A schematic diagram of the architecture of a singing voice conversion system including VITS is provided as an optional embodiment. Figure 4 The VITS also includes: a posteriori encoder, FLOW flow module and decoder. Figure 2 Based on the embodiment, the following continues to describe the VITS training process, please refer to Figure 5 :

[0091] 501. Inputting the linear spectrum of the song to be converted into a posterior encoder to obtain the posterior distribution parameters output by the posterior encoder module;

[0092] Specifically, in the embodiment of the present application, in addition to inputting the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted into the text encoder, and inputting multiple reference audios of the sample to be converted into the timbre encoder, the linear spectrum of the song to be converted can also be input into the posterior encoder to obtain the posterior distribution parameters output by the posterior encoder module, wherein the role of the posterior encoder in the embodiment of the present application is to make the posterior distribution parameters output by the posterior encoder module as close as possible to the prior distribution parameters output by the text encoder, so that the loss value between the first output of the song to be converted after passing through the posterior encoder and the second output after passing through the text encoder is as small as possible.

[0093] Among them, the posterior encoder in the embodiment of the present application may include Wavnet, which is a deep convolutional neural network architecture. Unlike the traditional recurrent neural network (RNN), WavNet uses a causal convolution layer (CausalConvolution), which allows the model to only consider the previous input samples during forward propagation of the network, avoiding the problems caused by information loops, which is very effective when processing variable-length audio signals.

[0094] The core idea of ​​WavNet is to use temporal convolutional blocks (TCB) and residual connections to process audio sequences. Each temporal block contains multiple small convolution kernels that slide frame by frame on the time axis while maintaining a global perception of the input. In addition, due to its sparse connection design, WavNet can capture long-term dependencies and learn more complex local features.

[0095] As an optional embodiment, after obtaining the song to be converted, the embodiment of the present application can input the song to be converted into the Librosa library to obtain the linear spectrum of the song to be converted, and further input the linear spectrum of the song to be converted into the posterior encoder.

[0096] 502. Sampling from the linear spectrum according to the posterior distribution parameters to obtain a sampling value vector of the linear spectrum;

[0097] After obtaining the posterior distribution parameters output by the posterior encoder in step 501, the posterior distribution parameters are used to sample from the linear spectrum of the song to be converted, thereby obtaining a sampling value vector of the linear spectrum.

[0098] As an optional embodiment, the posterior distribution parameter in the embodiment of the present application may also be a distribution parameter including a mean value and a variance.

[0099] 503. Input the sampling value vector of the linear spectrum to the FLOW flow module to map the sampling value vector of the linear spectrum from the complex distribution of the audio to the simple distribution of the text, so as to obtain the sampling value vector of the mapped linear spectrum;

[0100] After obtaining the sampling values ​​of the linear spectrum, the embodiment of the present application further inputs the sampling value vector of the linear spectrum into the FLOW flow module to map the sampling value vector of the linear spectrum from the complex distribution of audio to the simple distribution of text to obtain the sampling value vector of the mapped linear spectrum.

[0101] Specifically, the FLOW flow module, as a mapping function, maps a numerical distribution from one distribution to another. Because it is itself a mapping function, the FLOW flow module is reversible. When the FLOW flow module converts a first numerical distribution into a second numerical distribution, the inverse FLOW flow module converts the second numerical distribution into the first numerical distribution.

[0102] Because the linear spectrum in the embodiment of the present application is a kind of audio data, and the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted are a kind of text data, the audio data has a higher complexity than the text data. In order to have the same data distribution when calculating the loss value, the function of the FLOW flow module in the embodiment of the present application is to map the sampling value vector of the linear spectrum from the complex distribution of the audio to the simple distribution of the text, so as to obtain the sampling value vector of the mapped linear spectrum.

[0103] 504. According to the prior distribution of the text sampling value vector of the song content to be converted and the distribution of the sampling value vector of the mapped linear spectrum, calculate the KL divergence between the posterior distribution and the prior distribution to obtain the KL loss;

[0104] After step 503 obtains the sampling value vector of the mapped linear spectrum, the KL divergence between the prior distribution and the posterior distribution is further calculated based on the prior distribution of the text sampling value vector of the singing content to be converted and the posterior distribution of the sampling value vector of the mapped linear spectrum to obtain the KL loss of the model.

[0105] Specifically, KL divergence (Kullback-Leibler Divergence), also known as information gain or relative entropy, is a statistic that measures the difference between two probability distributions. If distribution P is completely equal to distribution Q, KL divergence is 0; when the difference between distribution P and distribution Q is large, the KL divergence is larger.

[0106] 505. Concatenate or add the sampling value vector of the linear spectrum and the new phoneme vector to obtain a fused singing vector;

[0107] After the above steps obtain the sampling value vector of the linear spectrum and the new phoneme vector, the sampling value vector of the linear spectrum and the new phoneme vector are further fused to obtain a fused singing vector.

[0108] Specifically, when the sampling value vector of the linear spectrum and the new phoneme vector are fused, the fusion operation may be concatenation or addition, and the specific means of the fusion operation are not limited here.

[0109] 506. Input the fused singing voice vector into a decoder to obtain synthesized singing voice audio output by the decoder, wherein the decoder includes an upsampling module;

[0110] After obtaining the fused singing voice vector, the fused singing voice vector is input into the decoder to obtain the synthesized singing voice audio output by the decoder.

[0111] As an optional embodiment, the decoder in the embodiment of the present application includes an upsampling module to decode the fused singing vector to obtain the synthesized singing audio.

[0112] 507. Calculate the reconstruction loss according to the Mel-spectrogram of the song to be converted and the Mel-spectrogram of the synthesized song audio;

[0113] After the decoder outputs the synthesized singing audio, in order to facilitate the calculation of reconstruction loss, the embodiment of the present application further extracts the Mel-spectrogram features of the synthesized singing audio and the Mel-spectrogram features of the singing to be converted.

[0114] As an optional embodiment, when extracting the Mel spectrum features of singing, the following operations may be performed: first, the time domain signal of the singing is converted to the frequency domain by Fourier transform to obtain the frequency domain features of the singing, and then the frequency domain signal of the singing is processed using a filter group with a Mel frequency scale to obtain the Mel spectrum.

[0115] Further, after obtaining the mel-spectrogram features of the synthesized singing audio and the mel-spectrogram features of the singing to be converted, a preset loss function is used to calculate the loss between the mel-spectrogram features of the synthesized singing audio and the mel-spectrogram features of the singing to be converted. The loss function here can be a mean square error loss function, a cross entropy loss function or a logarithmic loss function, etc. There is no specific restriction on the loss function here.

[0116] 508. The VITS singing voice conversion system based on variational autoencoder is trained according to reconstruction loss, KL loss and back propagation algorithm.

[0117] After the reconstruction loss is obtained in step 507, the reconstruction loss, KL loss and back propagation algorithm are further used to train the VITS singing voice conversion system based on the variational autoencoder to obtain a trained singing voice conversion system.

[0118] In an embodiment of the present application, when the singing voice conversion system includes VITS, the training process of the VITS system is described in detail, and the FLOW module is used in the VITS system, thereby improving the accuracy of calculating the KL loss. Furthermore, the embodiment of the present application uses reconstruction loss and KL loss to train the VITS system, which also improves the timbre similarity between the synthesized singing audio output by the VITS system and the input singing audio.

[0119] based on Figure 5 In order to further improve the accuracy of the synthesized singing audio output by the VITS system, the embodiment of the present application can also add a discriminator after the decoder. For ease of understanding, Figure 6 Another architectural diagram of the VITS system is given. Figure 2 Based on the examples, Figure 7 Another embodiment of the VITS system training process in the embodiment of the present application:

[0120] 701. Inputting the linear spectrum of the song to be converted into a posterior encoder to obtain the posterior distribution parameters output by the posterior encoder module;

[0121] 702. Sampling from the linear spectrum according to the posterior distribution parameters to obtain a sampling value vector of the linear spectrum;

[0122] 703. Input the sampling value vector of the linear spectrum to the FLOW flow module to map the sampling value vector of the linear spectrum from the complex distribution of the audio to the simple distribution of the text, so as to obtain the sampling value vector of the mapped linear spectrum;

[0123] 704. According to the prior distribution of the text sample value vector of the song content to be converted and the distribution of the sample value vector of the mapped linear spectrum, calculate the KL divergence between the posterior distribution and the prior distribution to obtain the KL loss.

[0124] 705. Concatenate or add the sampling value vector of the linear spectrum and the new phoneme vector to obtain a fused singing vector;

[0125] 706. Input the fused singing voice vector to a decoder to obtain synthesized singing voice audio output by an up-sampling module, wherein the decoder includes an up-sampling module;

[0126] 707. Calculate the reconstruction loss according to the Mel-spectrogram features of the song to be converted and the Mel-spectrogram features of the synthesized song audio;

[0127] It should be noted that steps 701 to 707 in the embodiment of the present application are similar to Figure 5Steps 501 to 507 in the embodiment are similar and will not be described again here.

[0128] 708. Input the synthesized singing audio to the discriminator to obtain a predicted identification result output by the discriminator;

[0129] The improved VITS system with the addition of the discriminator is equivalent to a GAN adversarial neural network. In order to improve the accuracy of the synthesized singing audio output by the decoder, the embodiment of the present application further adds a discriminator to the back end of the decoder, and inputs the synthesized singing audio output by the decoder into the discriminator, so that the discriminator can identify the synthesized singing audio output by the decoder to obtain the predicted identification result output by the discriminator.

[0130] Specifically, when no discriminator is added, that is, Figure 5 The improved VITS system in is equivalent to the generator, while Figure 7 In Figure 5 A discriminator is added after the generator, and the synthesized singing audio generated by the generator is discriminated to obtain the predicted identification result output by the discriminator.

[0131] Specifically, the prediction and identification result in the embodiment of the present application can be a binary classification result, such as whether the synthesized singing audio is the audio of a certain singer, if so, output 1, if not, output 0.

[0132] 709. Calculate adversarial loss according to the real identification result of the synthesized singing audio and the predicted identification result;

[0133] After obtaining the prediction results output by the discriminator, the predicted identification results and the true identification results are further used to calculate the adversarial loss, and the adversarial loss is used to train the improved VITS system.

[0134] 710. According to the reconstruction loss, the KL loss and the adversarial loss, the improved VITS singing voice conversion system based on variational autoencoder is trained.

[0135] Because the improved VITS system in the embodiment of the present application introduces reconstruction loss, KL loss and adversarial loss respectively, after obtaining the above losses, the improved VITS system can be reversely updated using the above losses and the back propagation algorithm to continuously update the model parameters in the improved VITS system until the improved VITS system converges.

[0136] In the embodiment of the present application, a discriminator is further introduced into the improved VITS system, and the adversarial loss of the discriminator is used to update the model of the improved VITS system, thereby further improving the accuracy of the synthesized singing audio output by the improved VITS system.

[0137] 2. If the singing voice conversion system includes the Fastspeeh fast speech synthesis model

[0138] Specifically, if the singing voice conversion system includes a Fastspeeh fast speech synthesis model, the Fastspeeh fast speech synthesis model also includes a text encoder and a decoder, wherein both the text encoder and the decoder adopt a stacked multi-head attention mechanism module. For the convenience of description, it is assumed that the text encoder includes a first multi-head attention mechanism module, and the decoder includes a second multi-head attention mechanism module. Figure 8 The Fastspeeh fast speech synthesis model architecture diagram is given. Figure 2 Based on the example, the training process of the Fastspeeh fast speech synthesis model is described in detail below. Fig. 9 :

[0139] 901. Concatenate or add the text sampling value vector of the singing content to be converted and the new timbre coding vector to obtain a fused singing vector, and input the fused singing vector to a decoder to obtain a Mel spectrum of the synthesized singing audio output by the decoder;

[0140] exist Figure 2 On the basis of the embodiment, after obtaining the new timbre coding vector output by the timbre perception attention mechanism module and the text sampling value vector of the singing content to be converted, the embodiment of the present application splices or adds the text sampling value vector of the singing content to be converted and the new timbre coding vector to obtain a fused singing vector, and inputs the fused singing vector into the decoder, that is, inputs it into the second multi-head attention mechanism module to obtain the Mel spectrum of the synthesized singing audio output by the second multi-head attention mechanism module.

[0141] It should be noted that due to the different architectures of the models, when the singing voice conversion system includes the VITS system, the output of the decoder is the synthesized singing audio, and when the singing voice conversion system includes Fastspeeh, the output of the decoder is the Mel-spectrogram of the synthesized singing audio.

[0142] 902. Calculate the reconstruction loss according to the Mel-spectrogram of the synthesized singing audio, the Mel-spectrogram of the singing to be converted, and a preset loss function;

[0143] After obtaining the Mel-spectrogram of the synthesized singing audio output by the decoder, the reconstruction loss is further calculated based on the Mel-spectrogram of the synthesized singing audio, the Mel-spectrogram of the singing to be converted and a preset loss function.

[0144] The process of extracting the mel spectrum features and the preset loss function are similar to those described in the above embodiment and will not be described again here.

[0145] 903. According to the reconstruction loss and back propagation algorithm, the Fastspeeh fast speech synthesis model is trained.

[0146] When the singing voice conversion system includes the Fastspeeh model, the loss only includes the reconstruction loss between the Mel-spectrogram features of the synthesized singing voice audio and the Mel-spectrogram features of the singing voice to be converted, so that the Fastspeeh system can be trained quickly based on the reconstruction loss.

[0147] Because in the embodiment of the present application, when the singing voice conversion model includes Fastspeeh, the loss function only includes reconstruction loss, and the Fastspeeh model structure is simpler than the VITS model structure, thereby reducing the configuration requirements for terminal processing and improving the training process of the improved Fastspeeh model.

[0148] based on Fig. 9 In order to further improve the accuracy of the synthesized singing audio Mel spectrum output by the Fastspeeh model, the embodiment described above can also be used in Fig. 9 The discriminator is added to the back end of the decoder for easy understanding. Fig.10 The architecture diagram of the Fastspeeh fast speech synthesis model is given. Fig.11 Another embodiment of the training process of the Fastspeeh fast speech synthesis model is provided:

[0149] 1101. Concatenate or add the text sampling value vector of the singing content to be converted and the new timbre coding vector to obtain a fused singing vector, and input the fused singing vector to a decoder to obtain a Mel spectrum of the synthesized singing audio output by the decoder;

[0150] 1102. Calculate the reconstruction loss according to the Mel-spectrogram of the synthesized singing audio, the Mel-spectrogram of the singing to be converted, and a preset loss function;

[0151] Among them, steps 1101 to 1102 are similar to the description of steps 901 to 902, and are not repeated here.

[0152] 1103. Inputting the synthesized Mel spectrum of the singing audio into the discriminator to obtain a predicted identification result output by the discriminator;

[0153] 1104. Calculate the adversarial loss according to the real identification result of the Mel spectrum of the synthesized singing audio and the predicted identification result;

[0154] The description of step 1103 to step 1104 is similar to the description of step 708 to step 709 and will not be repeated here.

[0155] 1105. The Fastspeeh fast speech synthesis model is trained based on reconstruction loss, adversarial loss and back propagation algorithm.

[0156] Because the reconstruction loss is obtained in step 1102 and the adversarial loss is obtained in step 1104, the model parameters of the Fastspeeh fast speech synthesis model are trained using the reconstruction loss and the adversarial loss, as well as the back propagation algorithm, until the model of the Fastspeeh fast speech synthesis model converges.

[0157] The embodiments of this application are Fig. 9 On the basis of the embodiment, a discriminator and adversarial loss are further introduced to train the Fastspeeh fast speech synthesis model, thereby improving the accuracy of the mel-spectrogram features of the synthesized singing audio output by the Fastspeeh fast speech synthesis model.

[0158] 3. If the singing voice conversion system includes a diffusion generation model

[0159] Specifically, if the singing voice conversion system includes a Diffusion diffusion generation model, the Diffusion diffusion generation model also includes a text encoder and a decoder, wherein the text encoder includes a third multi-head attention mechanism module, and the decoder includes a U-net module or a DiT module, wherein the model architecture diagram of the Diffusion diffusion generation model is the same as Figure 8 The architecture diagram of the Fastspeeh fast speech synthesis model in is similar, except that the modules in the decoder are slightly different. The following describes the training process of the Diffusion diffusion generation model. Fig.12 :

[0160] 1201. Concatenate or add the text sampling value vector of the singing content to be converted and the new timbre coding vector to obtain a fused singing vector, and input the fused singing vector to a decoder to obtain a Mel spectrum of the synthesized singing audio output by the decoder;

[0161] Because the singing voice conversion system in the embodiment of the present application includes a Diffusion diffusion generation model, which is different from the Fastspeeh fast speech synthesis model, the decoder in the embodiment of the present application adopts a U-net module or a DiT module. Because the U-net module adopts a residual training network, when the decoder adopts the U-net model, the convergence speed of the model can be accelerated and the training time of the Diffusion diffusion generation model can be shortened.

[0162] 1202. Calculate reconstruction loss according to the Mel-spectrogram features of the synthesized singing audio, the Mel-spectrogram features of the singing audio to be converted, and a preset loss function;

[0163] 1203. Train the Diffusion generation model according to the reconstruction loss and back propagation algorithm.

[0164] The process from step 1202 to step 1203 is similar to the description from step 902 to step 903, and will not be repeated here.

[0165] In the embodiment of the present application, when the singing voice conversion system includes a Diffusion diffusion generation model, the training process of the Diffusion diffusion generation model is described in detail, and when the decoder in the Diffusion diffusion generation model is a U-net module, the convergence speed of the model can be accelerated and the training time of the Diffusion diffusion generation model can be shortened.

[0166] based on Fig.12 In order to improve the accuracy of the singing synthesis audio output by the Diffusion diffusion generation model, the embodiment of the present application can further add a discriminator after the decoder in the Diffusion diffusion generation model, wherein another schematic diagram of the Diffusion diffusion generation model is similar to Fig.10 The architecture diagram of the Fastspeeh fast speech synthesis model is similar to that of the Fastspeeh model. Fig.13 Another embodiment of the training process of the Diffusion diffusion generation model is given:

[0167] 1301. Concatenate or add the text sampling value vector of the singing content to be converted and the new timbre coding vector to obtain a fused singing vector, and input the fused singing vector to a decoder to obtain a Mel spectrum of the synthesized singing audio output by the decoder;

[0168] 1302. Calculate reconstruction loss according to the Mel-spectrogram features of the synthesized singing audio, the Mel-spectrogram features of the singing audio to be converted, and a preset loss function;

[0169] The description of step 1301 to step 1302 is similar to the description of step 901 to step 902, and will not be repeated here.

[0170] 1303. Inputting the synthesized Mel spectrum of the singing audio into a discriminator to obtain a predicted identification result output by the discriminator;

[0171] 1304. Calculate the adversarial loss according to the real identification result of the Mel spectrum of the synthesized singing audio and the predicted identification result;

[0172] 1305. Train the Diffusion generation model according to the reconstruction loss, the adversarial loss and the back propagation algorithm.

[0173] The description of steps 1303 to 1305 is similar to the description of steps 1103 to 1105 and will not be repeated here.

[0174] The embodiments of this application are Fig.12 On the basis of the embodiment, a discriminator and adversarial loss are further introduced to train the Diffusion diffusion generation model, thereby improving the accuracy of the synthesized singing audio mel-spectrogram features output by the Diffusion diffusion generation model.

[0175] The above describes in detail the training method of the singing voice conversion system in the embodiment of the present application. Next, the method for generating audio based on the singing voice conversion system in the embodiment of the present application is described. Fig.14 :

[0176] Specifically, the singing voice conversion system in the embodiment of the present application includes at least a text encoder, a timbre encoder, a timbre perception attention mechanism module and a decoder. For ease of understanding, Fig.15 A schematic diagram of the architecture of singing voice conversion is given, and a method for generating audio based on the singing voice conversion system includes:

[0177] 1401. Acquire multiple reference audios of a second target timbre, wherein the multiple reference audios at least include multiple audios of different singers in different vocal ranges;

[0178] The multiple reference audios of the second target timbre in the embodiment of the present application at least include multiple audios of different singers in different vocal ranges. Specifically, when using the singing voice conversion system to generate the audio of the target timbre, the singing voice of the sample to be converted can be converted into the timbre of different singers, such as converting the singing voice of the sample to be converted into the timbre of Liu Xhua, or into the timbre of Zhou Y, that is, converting the singing voice to be converted into the timbre of other singers. Therefore, the multiple reference audios in the embodiment of the present application at least include multiple audios of different singers in different vocal ranges. In addition, because each singer uses different resonance methods when singing in different vocal ranges, that is, the resonance methods used when singing in the high pitch and low pitch are different, and in order to convert the singing voice to be converted into the most suitable target timbre, multiple audios of different singers in different vocal ranges (such as the high pitch, middle pitch and low pitch) are required, so that during the conversion process, the best target timbre suitable for the singing voice to be converted can be found.

[0179] Furthermore, multiple audios of different singers in different pitch ranges may be audios of different singers singing the same song, or may be audios of different singers singing different songs. No specific restrictions are imposed on the specific contents of the multiple audios of different singers in different pitch ranges.

[0180] 1402. Input the plurality of reference audios into a timbre encoder to perform timbre encoding on the plurality of reference audios, and obtain timbre encoding vectors for the plurality of reference audios output by the timbre encoder;

[0181] After obtaining multiple reference audios of different singers in different vocal ranges, the multiple reference audios are input into a timbre encoder to perform timbre encoding on the multiple reference audios of the singer's singing voice, and obtain timbre encoding vectors of the multiple reference audios output by the timbre encoder.

[0182] Specifically, the timbre encoder in the embodiment of the present application can be a stacked multi-head attention mechanism, because the multi-head attention mechanism can pay attention to different parts of the input at the same time, and capture different levels of abstract information through multiple independent attention heads, and each head has its own weight matrix to capture different feature combinations of the input respectively. The parallel processing capability of this multi-head attention mechanism improves the global perception ability of the model and can better understand and process contextual relationships.

[0183] 1403. Input the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the content of the song to be converted output by the text encoder;

[0184] Specifically, in the embodiment of the present application, the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted are both used to characterize the singing content of the singer. In the embodiment of the present application, after obtaining the singing of the sample to be converted, the song to be converted can be input into the Whisper model, the Hubert model or the Librosa model to obtain the phoneme posterior probability (PPG) of the song to be converted, and the song to be converted is input into the Rmvte singing melody extraction model based on the recurrent neural network to extract the Pitch fundamental frequency of the song to be converted.

[0185] The text encoder in the embodiment of the present application is a multi-head attention mechanism module with the same architecture as the timbre encoder. When the posterior probability of phonemes (PPG) of the singer's singing and the fundamental frequency (Pitch) of the singer's singing are input into the text encoder, the prior distribution parameters of the predicted phoneme distribution and fundamental frequency distribution in the singer's singing output by the text encoder can be obtained. When the prior distribution is a normal distribution, the prior distribution parameters can be the mean and variance.

[0186] Furthermore, the samples to be converted in the embodiment of the present application are users who need to perform singing conversion. For example, if the singing of user A needs to be converted into the target timbre, the sample to be converted is user A. If the singing of user B needs to be converted into the target timbre, the user to be converted is user B. Therefore, the samples to be converted in the embodiment of the present application are different users who need to perform timbre conversion.

[0187] 1404. According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the content of the song to be converted;

[0188] After obtaining the prior distribution parameters, the prior distribution parameters are used to sample from the phoneme posterior probability (PPG) of the song to be converted and the fundamental frequency (Pitch) of the song to be converted to obtain a text sampling value vector of the song content to be converted.

[0189] It should be noted here that, because the Whisper model, Hubert model or Librosa model is used in the embodiments of the present application, the extracted phoneme posterior probability (PPG) of the song to be converted is at the frame level, and Rmvte is a singing melody extraction model based on a recurrent neural network, and the extracted fundamental frequency (Pitch) of the song to be converted is also at the frame level. Therefore, the text sampling value vector of the singing content to be converted obtained in the embodiments of the present application is also a frame-level singing content encoding.

[0190] 1405. Input the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of the multiple reference audios and the text sampling value vector of the singing content to be converted;

[0191] Because a timbre-perceived attention mechanism module is added in the embodiment of the present application, after the text sampling value vector of the singer's singing content and the timbre coding vectors of multiple reference audios are input into the timbre-perceived attention mechanism module, a new timbre coding vector matching the text sampling value vector of the singer's singing content output by the timbre-perceived attention mechanism module can be obtained.

[0192] Because the timbre perception attention mechanism module in the embodiment of the present application is the attention mechanism in the transformer, where the formula of the attention mechanism is:

[0193]

[0194] Among them, the text sampling value vector of the singing content to be converted is Q, and the timbre code is K and V. is the dimension of the timbre encoding vector K.

[0195] After the above operations, a new timbre coding vector matching the singing content to be converted can be obtained.

[0196] 1406. Use the text sampling value vector of the singing content to be converted and the new timbre coding vector as inputs of the decoder to obtain synthesized singing audio.

[0197] After obtaining the text sampling value vector and the new timbre coding vector of the singing content to be converted, the text sampling value vector and the new timbre coding vector of the singing content to be converted are input into the decoder, and the synthesized singing audio output by the decoder can be obtained.

[0198] Because the new timbre coding vector in the embodiment of the present application is a new timbre coding vector that matches the text sampling value vector of the singing content determined based on the timbre coding vectors of multiple reference audios and the text sampling value vector of the singing content to be converted, the timbre vector of the synthesized singing audio output by the decoder has a higher timbre similarity to the singer himself.

[0199] The singing voice conversion system includes different singing voice conversion systems, and the singing voice conversion system is described in detail below:

[0200] 1. If the singing voice conversion system includes VITS singing voice conversion system based on variational autoencoder

[0201] Specifically, when the singer conversion system includes the VITS singing voice conversion system based on the variational autoencoder, VITS also includes an inverse FLOW flow module. Fig.16 The following is a schematic diagram of the architecture of the singing voice conversion system including VITS. The following is a schematic diagram of the process of VITS outputting synthesized singing audio. Fig.17 :

[0202] 1701. Acquire multiple reference audios of a second target timbre, wherein the multiple reference audios at least include multiple audios of different singers in different vocal ranges;

[0203] 1702. Input the plurality of reference audios into a timbre encoder to perform timbre encoding on the plurality of reference audios, and obtain timbre encoding vectors for the plurality of reference audios output by the timbre encoder;

[0204] 1703. Input the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the song content to be converted output by the text encoder;

[0205] 1704. According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the song content to be converted;

[0206] 1705. Input the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of the multiple reference audios and the text sampling value vector of the singing content to be converted;

[0207] It should be noted that the description of steps 1701 to 1705 is similar to that of steps 1401 to 1405 and will not be repeated here.

[0208] 1706. Input the text sampling value vector of the singing content to be converted into the inverse FLOW flow module to map the simple distribution of the text sampling value vector of the singing content to be converted to the complex distribution of the audio to obtain the mapped text sampling value vector of the singing content to be converted;

[0209] Because when the singing voice conversion system includes VITS, the FLOW flow module is introduced during the training process, so in the deduction stage, VITS also includes an inverse FLOW flow module to map the simple distribution of the text sampling value vector of the singing content to be converted to the complex distribution of the audio, so as to obtain the mapped text sampling value vector of the singing content to be converted.

[0210] It should be noted that the inverse FLOW flow module in the embodiment of the present application is the inverse operation of the FLOW flow module in the training process. When the FLOW flow module is the first function, the inverse FLOW flow module is the inverse function of the first function, and the function of the FLOW flow module is to convert the complex distribution of the audio into a simple distribution of the text sampling value vector, and the function of the inverse FLOW flow module is to convert the simple distribution of the text sampling value vector into a complex distribution of the audio, so as to obtain the text sampling value vector of the singer's singing content after mapping.

[0211] 1707. Concatenate or add the mapped text sampling value vector of the singing content to be converted and the new timbre coding vector to obtain a fused singing vector;

[0212] After obtaining the mapped text sampling value vector of the singing content to be converted and the new timbre coding vector, the mapped text sampling value vector of the singing content to be converted and the new timbre coding vector are further fused to obtain a fused singing vector, wherein the fusion operation here can be vector splicing or addition.

[0213] 1708. Input the fused singing vector into the decoder to obtain the synthesized singing audio.

[0214] After obtaining the fused singing voice vector, the fused singing voice vector is input into the decoder to obtain the synthesized singing voice audio.

[0215] Specifically, the decoder in the embodiment of the present application is an upsampling module.

[0216] In the embodiment of the present application, when the singing voice conversion system includes VITS based on a variational autoencoder, the process of generating the synthesized audio is described in detail, and in the embodiment of the present application, VITS is based on a variational autoencoder and can screen out new timbre coding vectors that match the singing content from multiple reference audios of different singers in different pitch ranges, so the synthesized singing audio has a high timbre similarity with the singer's singing.

[0217] 2. If the singing voice conversion system includes Fastspeeh fast speech synthesis model or Diffusion diffusion generation model

[0218] Specifically, if the singing voice conversion system includes a Fastspeeh fast speech synthesis model or a Diffusion diffusion generation model, the architecture diagram of the Fastspeeh fast speech synthesis model or the Diffusion diffusion generation model is as follows: Fig.15 As shown, the process of Fastspeeh fast speech synthesis model or Diffusion diffusion generation model generating synthetic audio is as follows Fig.18 As shown:

[0219] 1801. Acquire multiple reference audios of a second target timbre, wherein the multiple reference audios at least include multiple audios of different singers in different vocal ranges;

[0220] 1802. Input the plurality of reference audios into a timbre encoder to perform timbre encoding on the plurality of reference audios, and obtain timbre encoding vectors for the plurality of reference audios output by the timbre encoder;

[0221] 1803. Input the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the song content to be converted output by the text encoder;

[0222] 1804. According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the song content to be converted;

[0223] 1805. Input the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of the multiple reference audios and the text sampling value vector of the singing content to be converted;

[0224] It should be noted that the description of steps 1801 to 1805 is similar to that of steps 1401 to 1405 and will not be repeated here.

[0225] And when the singing is converted into the Fastspeeh fast speech synthesis model or the Diffusion diffusion generation model, the text encoders therein are all multi-head attention mechanism modules.

[0226] 1806. Concatenate or add the text sampling value vector of the singing content to be converted and the new timbre coding vector to obtain a fused singing vector;

[0227] When the singing voice conversion system includes the Fastspeeh fast speech synthesis model or the Diffusion diffusion generation model, if the text sampling value vector of the singing content to be converted and the new timbre coding vector are obtained, the text sampling value vector of the singer's singing content and the new timbre coding vector are fused to obtain a fused singing vector. The fusion operation here can be vector addition or splicing.

[0228] 1807. Input the fused singing vector into the decoder to obtain synthesized singing audio.

[0229] After obtaining the fused singing voice vector, the fused singing voice vector is input into the decoder to obtain the synthesized singing voice audio.

[0230] It should be noted here that if the singing voice conversion system is a Fastspeeh fast speech synthesis model, the decoder is also a multi-head attention mechanism module, and if the singing voice conversion system is a Diffusion diffusion generation model, the decoder is a U-net module or a DiT module.

[0231] In the embodiment of the present application, when the singing voice conversion system includes the Fastspeeh fast speech synthesis model or the Diffusion diffusion generation model, the process of synthesizing audio is described in detail, and the Fastspeeh fast speech synthesis model or the Diffusion diffusion generation model in the embodiment of the present application can generate a new timbre coding vector that matches the singing content based on multiple reference audios of different singers in different pitch ranges, so the synthesized singing audio has a high timbre similarity with the singer's singing.

[0232] Furthermore, the Fastspeeh fast speech synthesis model or Diffusion diffusion generation model in the embodiments of the present application has a simpler model structure than the VITS singing voice conversion system based on variational autoencoder, and reduces the configuration requirements for the terminal processor.

[0233] The embodiment of the present application further provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the method steps in the above method embodiment.

[0234] The present application also provides a computer device, which includes:

[0235] Processor and memory;

[0236] The memory is used to store computer programs, and when the processor is used to execute the computer programs stored in the memory, it is used to implement the method steps in the above method embodiment.

[0237] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that a processor and a memory are merely examples of computer devices and do not constitute a limitation of the computer device. The computer device may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, and the like.

[0238] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device.

[0239] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the computer device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0240] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the method steps in the above method embodiment.

[0241] It is understandable that if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned corresponding embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0242] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A training method for a singing voice conversion system, characterized in that: The singing voice conversion system comprises at least a text encoder, a timbre encoder and a timbre perception attention mechanism module, and the method comprises: Acquire multiple reference audios of a first target timbre, wherein the multiple reference audios of the first target timbre at least include multiple audios of a singer in different voice ranges of the singing voice to be converted; Inputting the plurality of reference audios of the first target timbre into the timbre encoder to perform timbre encoding on the plurality of reference audios of the first target timbre, and obtaining timbre encoding vectors of the plurality of reference audios of the first target timbre output by the timbre encoder; Inputting the PPG phoneme posterior probability of the song to be converted and the pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the song content to be converted output by the text encoder; According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the song content to be converted; Input the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios of the first target timbre into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of multiple reference audios of the first target timbre and the text sampling value vector of the singing content to be converted; Using the new timbre coding vector as a new input added to the singing voice conversion system, and calculating the reconstruction loss of the singing voice conversion system; The singing voice conversion system is trained according to the reconstruction loss and the back propagation algorithm.

2. The method according to claim 1, characterized in that If the singing voice conversion system includes a VITS singing voice conversion system based on a variational autoencoder, the VITS singing voice conversion system based on a variational autoencoder further includes: a posterior encoder and a FLOW flow module, wherein the text encoder includes a multi-head attention mechanism module; The method further comprises: Inputting the linear spectrum of the song to be converted into the posterior encoder to obtain the posterior distribution parameters output by the posterior encoder module; Sampling from the linear spectrum according to the posterior distribution parameters to obtain a sampling value vector of the linear spectrum; Inputting the sampling value vector of the linear spectrum into the FLOW flow module to map the sampling value vector of the linear spectrum from the complex distribution of the audio to the simple distribution of the text, so as to obtain the sampling value vector of the mapped linear spectrum; According to the prior distribution of the text sampling value vector of the singing content to be converted and the posterior distribution of the sampling value vector of the mapped linear spectrum, the KL divergence between the posterior distribution and the prior distribution is calculated to obtain the KL loss.

3. The method according to claim 2, characterized in that The VITS singing voice conversion system based on variational autoencoder further includes: a decoder, wherein the decoder includes an upsampling module; The method of using the new timbre coding vector as a new input added to the singing voice conversion system and calculating the reconstruction loss of the singing voice conversion system comprises: Concatenate or add the sampling value vector of the linear spectrum and the new timbre coding vector to obtain a fused singing vector; Inputting the fused singing voice vector into the up-sampling module to obtain the synthesized singing voice audio output by the up-sampling module; Calculating the reconstruction loss according to the Mel-spectrogram of the song to be converted and the Mel-spectrogram of the synthesized song audio; The method of training the singing voice conversion system according to the reconstruction loss and the back propagation algorithm comprises: The VITS singing voice conversion system based on variational autoencoder is trained according to the reconstruction loss, the KL loss and the back propagation algorithm.

4. The method according to claim 3, characterized in that: The VITS singing voice conversion system based on variational autoencoder also includes: a discriminator; The method further comprises: Inputting the synthesized singing audio into the discriminator to obtain a predicted identification result output by the discriminator; Calculating adversarial loss according to the real identification result of the synthesized singing audio and the predicted identification result; The VITS singing voice conversion system based on the variational autoencoder is trained according to the reconstruction loss, the KL loss and the back propagation algorithm, including: The VITS singing voice conversion system based on variational autoencoder is trained according to the reconstruction loss, the KL loss and the adversarial loss, and the back propagation algorithm.

5. The method according to claim 2, characterized in that: Before inputting the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted into the text encoder, the method further includes: Inputting the song to be converted into a first preset model to extract the PPG phoneme posterior probability of the song to be converted, wherein the first preset model includes a Whisper model, a Hubert model, or a Librosa model; Input the song to be converted into a song melody extraction model based on a recurrent neural network of Rmvte to extract the pitch fundamental frequency of the song to be converted; Before the linear spectrum of the song to be converted is input into the a posteriori encoder, the method further comprises: The song to be converted is input into the Librosa library to obtain a linear spectrum of the song to be converted.

6. The method according to claim 1, characterized in that If the singing voice conversion system includes a Fastspeeh fast speech synthesis model, the Fastspeeh fast speech synthesis model also includes a text encoder and a decoder, wherein the text encoder includes a first multi-head attention mechanism module, and the decoder includes a second multi-head attention mechanism module; The new timbre coding vector is used as a new input added to the singing voice conversion system, and the reconstruction loss of the singing voice conversion system is calculated, including: The text sampling value vector of the singing content to be converted and the new timbre coding vector are concatenated or added to obtain a fused singing vector; Inputting the fused singing voice vector into the decoder to obtain the Mel spectrum of the synthesized singing voice audio output by the decoder; The reconstruction loss is calculated according to the Mel-spectrogram of the synthesized singing audio, the Mel-spectrogram of the singing audio to be converted and a preset loss function.

7. The method according to claim 6, characterized in that The Fastspeeh fast speech synthesis model also includes: a discriminator; The method further comprises: Inputting the synthesized Mel spectrum of the singing audio into the discriminator to obtain a predicted identification result output by the discriminator; Calculating the adversarial loss according to the real identification result of the Mel spectrum of the synthesized singing audio and the predicted identification result; According to the reconstruction loss and the back propagation algorithm, the Fastspeeh fast speech synthesis model is trained, including: The Fastspeeh fast speech synthesis model is trained according to the reconstruction loss, the adversarial loss and the back propagation algorithm.

8. The method according to claim 1, characterized in that If the singing voice conversion system includes a Diffusion diffusion generation model, the Diffusion diffusion generation model also includes a text encoder and a decoder, wherein the text encoder includes a third multi-head attention mechanism module, and the decoder includes a U-net module or a DiT module; The new timbre coding vector is used as a new input added to the singing voice conversion system, and the reconstruction loss of the singing voice conversion system is calculated, including: The text sampling value vector of the singing content to be converted and the new timbre coding vector are concatenated or added to obtain a fused singing vector; Inputting the fused singing voice vector into the decoder to obtain the Mel spectrum of the synthesized singing voice audio output by the decoder; The reconstruction loss is calculated according to the Mel-spectrogram of the synthesized singing audio, the Mel-spectrogram of the singing audio to be converted and a preset loss function.

9. The method according to claim 8, characterized in that The Diffusion diffusion generation model also includes: a discriminator; The method further comprises: Inputting the synthesized Mel spectrum of the singing audio into the discriminator to obtain a predicted identification result output by the discriminator; Calculating the adversarial loss according to the real identification result of the Mel spectrum of the synthesized singing audio and the predicted identification result; The Diffusion generation model is trained according to the reconstruction loss and the back propagation algorithm, including: The Diffusion diffusion generation model is trained according to the reconstruction loss, the adversarial loss and the back propagation algorithm.

10. A method for generating audio based on a singing voice conversion system, characterized in that: The singing voice conversion system comprises at least a text encoder, a timbre encoder, a timbre perception attention mechanism module and a decoder, and the method comprises: Acquire multiple reference audios of a second target timbre, wherein the multiple reference audios of the second target timbre at least include multiple audios of different singers in different vocal ranges; Inputting the plurality of reference audios of the second target timbre into the timbre encoder to perform timbre encoding on the plurality of reference audios of the second target timbre, and obtaining timbre encoding vectors of the plurality of reference audios of the second target timbre output by the timbre encoder; Inputting the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted into the text encoder to obtain the prior distribution parameters of the song content to be converted output by the text encoder; According to the prior distribution parameters, sampling is performed from the PPG phoneme posterior probability of the song to be converted and the Pitch fundamental frequency of the song to be converted to obtain a text sampling value vector of the song content to be converted; Input the text sampling value vector of the singing content to be converted and the timbre coding vectors of multiple reference audios of the second target timbre into the timbre perception attention mechanism module to determine a new timbre coding vector based on the timbre coding vectors of multiple reference audios of the second target timbre and the text sampling value vector of the singing content to be converted; The text sampling value vector of the singing content to be converted and the new timbre coding vector are used as inputs of the decoder to obtain synthesized singing audio.

11. The method according to claim 10, characterized in that If the singing voice conversion system includes a VITS singing voice conversion system based on a variational autoencoder, then the VITS singing voice conversion system based on a variational autoencoder further includes an inverse FLOW flow module; The method further comprises: Input the text sampling value vector of the singing content to be converted into the inverse FLOW flow module to map the simple distribution of the text sampling value vector of the singing content to be converted to the complex distribution of the audio to obtain the mapped text sampling value vector of the singing content to be converted; The method of using the new timbre coding vector as the input of the decoder to obtain synthesized singing audio comprises: The mapped text sampling value vector of the singing content to be converted and the new timbre coding vector are concatenated or added to obtain a fused singing vector; The fused singing vector is input into the decoder to obtain the synthesized singing audio.

12. The method according to claim 10, characterized in that If the singing voice conversion system includes a Fastspeeh fast speech synthesis model or a Diffusion diffusion generation model; The text sampling value vector of the song content to be converted and the new timbre coding vector are used as inputs of the decoder to obtain synthesized song audio, including: The text sampling value vector of the singing content to be converted and the new timbre coding vector are concatenated or added to obtain a fused singing vector; The fused singing vector is input into the decoder to obtain the synthesized singing audio.

13. The method according to claim 11, characterized in that If the singing voice conversion system includes a VITS singing voice conversion system based on a variational autoencoder, the text encoder includes a first multi-head attention mechanism module, and the decoder includes an upsampling module; If the singing voice conversion system includes a Fastspeeh fast speech synthesis model, the text encoder includes a second multi-head attention mechanism module, and the decoder includes a third multi-head attention mechanism module; If the singing voice conversion system includes a Diffusion generation model, the text encoder includes a fourth multi-head attention mechanism module, and the decoder includes a U-net module or a DiT module.

14. A computer device comprising a processor, characterized in that: When executing the computer program stored in the memory, the processor is used to implement the training method of the singing voice conversion system as described in any one of claims 1 to 9, or the method of generating audio based on the improved singing voice conversion system as described in any one of claims 10 to 13.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the training method of the singing voice conversion system as described in any one of claims 1 to 9, or the method of generating audio based on the singing voice conversion system as described in any one of claims 10 to 13.

16. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the training method of the singing voice conversion system as described in any one of claims 1 to 9, or the method of generating audio based on the singing voice conversion system as described in any one of claims 10 to 13.

Citation Information

Patent Citations

  • Tone conversion method and device, electronic equipment and readable storage medium

    CN113611309A

  • Voice conversion model training method and device, electronic equipment and medium

    CN113689866A

  • Speech synthesis method, related device, electronic equipment and storage medium

    CN113793591A

  • Rap audio generation method, device and equipment and readable storage medium

    CN116013248A

  • Speech synthesis model training method, speech synthesis method, device and equipment

    CN116994553A

Cited By

  • Audio processing method and device

    CN120510853A