An End-to-End Multi-Scale Style Transfer Method and System for Singing Voice Conversion
Through the end-to-end multi-scale style transfer singing conversion method, using technical means such as feature extraction modules and residual style encoders, the problem of difficulty in modeling fine-grained singing styles in the existing technology is solved, high-quality style transfer is achieved, and sound quality and naturalness are improved.
Patent Information
- Application Number
- CN202410944150.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-07-15
AI Technical Summary
The existing singing voice conversion technology is difficult to effectively model and transform singing styles, especially fine-grained singing and expression techniques, which leads to poor style transfer results and cannot meet users' personalized needs.
A singing voice conversion method for end-to-end multi-scale style transfer is proposed. Through technical means such as feature extraction module, residual style encoder, uncertain style example normalization module, and other technical means, the singing voice characteristics are automatically extracted, global and local styles are modeled, style leakage is reduced, and multi-scale style transfer is realized.
It improves the style similarity of singing voice migration, reduces the style leakage of target voice, improves the sound quality and nature, and meets users' needs for personalized singing styles.
Smart Images

Figure CN118969013B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of singing voice conversion, and particularly relates to an end-to-end multi-scale style transfer singing voice conversion method and system. Background Art
[0002] Singing voice conversion (SVC) aims to convert the identity of a source singer to that of a target singer without changing the content. With the rapid development of SVC, the demand for personalized singing styles is increasing, posing challenges to current SVC models. Different from traditional SVC tasks, singing voice conversion with style transfer aims not only to transfer the identity of the target singer but also to transfer the singing style. Personal singing styles mainly include singing methods (such as pop, ethnic), techniques (such as vibrato, glissando), and pronunciation. However, traditional SVC methods lack effective mechanisms for modeling and converting these styles.
[0003] To model the style of singing voices, most studies start from a decoupled perspective. Some methods use singer identity information to adjust formant-related and pitch-related decoders in an attempt to separate timbre and singing style, but this method can only be trained on parallel data, which is costly for singing. Or some methods attempt to model the overall style through a style encoder, but this method can only model singer timbre. The singing style is usually fine-grained, and the details of singing methods cannot be modeled by just one style encoder. Moreover, these methods only focus on limited aspects of singing style and cannot capture the expressive techniques of singing style. Currently, due to a lot of redundancy in the style modeling information, style leakage occurs, so the effects cannot meet user satisfaction in terms of both objective evaluation and subjective scoring. Summary of the Invention
[0004] The purpose of the present invention is to overcome the defects of the prior art and propose an end-to-end multi-scale style transfer singing voice conversion method and system.
[0005] To achieve the above purpose, the present invention proposes an end-to-end multi-scale style transfer singing voice conversion method, including:
[0006] Collect the target singing voice to be converted and perform preprocessing to remove the accompaniment sound;
[0007] Input the preprocessed target singing voice and the reference singing voice of the intended style into a pre-established and trained singing voice conversion model, and output a synthesized singing voice with the style of the reference singing voice to achieve style transfer;
[0008] The singing voice conversion model is used to extract the content vector and MIDI from the preprocessed target singing voice, extract the global and local style vectors, pitch, and CQT spectrum from the reference singing voice, and obtain the singing voice waveform through end-to-end processing.
[0009] Preferably, the singing voice conversion model includes: a feature extraction module, a residual style encoder, a content encoder, an uncertainty style instance normalization module, a pitch predictor, a prior encoder, and a neural vocoder; wherein,
[0010] The feature extraction module is used to extract a content vector and MIDI from the preprocessed target singing voice, and extract the pitch F0 and the first-order difference ΔF0 of F0, the global style vector s including the timbre vector Gt and the singing style vector Gs, and the CQT spectrum from the reference singing voice;
[0011] The residual style encoder is used to extract the local style vector ls including ornaments at the MIDI scale according to the CQT spectrum and ΔF0 of the reference singing voice;
[0012] The content encoder is used to encode the content vector from 768 dimensions to 192 dimensions, encode MIDI to 128 dimensions, and after splicing, input it into a 4-layer FFT network to obtain the encoded content vector Ec, which is respectively input into the uncertainty style instance normalization module and the residual style encoder;
[0013] The uncertainty style instance normalization module is used to perform instance normalization on the encoded content vector Ec, calculate the channel covariance matrix of the scale and bias vectors for the global style vector s, and obtain the content representation after style perturbation;
[0014] The pitch predictor is used to fuse the content representation after style perturbation and the global style vector s through adaptive instance normalization, and then output the predicted pitch through a 4-layer FFT, and constrain the pitch within the set range of MIDI pitch to avoid out-of-tune;
[0015] The prior encoder uses a 4-layer FFT to predict the mean and variance of the prior distribution according to the predicted pitch and the content vector after style perturbation by the uncertainty style instance normalization module, calculate the multivariate Gaussian distribution, and sample from it to obtain the sampled feature of the 192-dimensional prior distribution;
[0016] The neural vocoder uses a structure based on nsf HiFiGAN to output the predicted singing voice waveform according to the sampled feature of the prior distribution.
[0017] Preferably, the feature extraction module includes: a pitch extractor, a Content Vec unit, a global style encoder, and a MIDI extractor; wherein,
[0018] The pitch extractor is used to extract the pitch F0 of the target singing voice and the reference singing voice, and calculate the first-order difference ΔF0 of the reference singing voice F0;
[0019] The Content Vec unit is used to extract the content vector from the preprocessed target singing voice;
[0020] The global style encoder is used to extract the timbre vector Gt and the singing style vector Gs of the reference singing voice;
[0021] The MIDI extractor is used to measure the beats of the target singing voice, quantize F0 according to the beats, remove short notes and silent parts, interpolate to the same frame rate as the Mel spectrogram, and calculate the CQT spectrogram representing the pitch characteristics.
[0022] Preferably, the global style encoder is based on the MERT model, adding two branches with the same structure, and each branch consists of a linear layer and a pooling layer. Its processing process includes:
[0023] Downsample the reference singing voice with a sampling rate of 24K by 320 times to obtain low-frame-rate features;
[0024] Input into the two branches simultaneously to obtain the timbre vector Gt and the singing style vector Gs respectively.
[0025] Preferably, the residual style encoder includes: a front-end convolutional network, a residual quantization module, and a cross-attention fusion network; its processing process includes:
[0026] Input the CQT spectrogram, ΔF0, and note onset identifier of the reference singing voice into the front-end convolutional network and the residual quantization module in sequence, and output the two-layer local style vector ls, where the note onset symbol is obtained according to the change position of MIDI;
[0027] The cross-attention fusion network fuses the local style vector ls and the encoded content vector Ec output by the content encoder, and outputs the content vector Ec' with style fusion having the same length as Ec.
[0028] Preferably, the processing process of the uncertainty style instance normalization module includes:
[0029] Calculate the instance normalization of the content vector Ec;
[0030] Add the timbre vector Gt and the singing style vector Gs as the global style vector s, predict the scale vector γ(s) and the bias vector β(s) respectively through the linear layer, and then calculate the channel covariance matrix ∈ of the scale vector according to the following formula γ (s) and the channel covariance matrix ∈ of the bias vector β (s):
[0031]
[0032] Among them, T represents transpose, B represents the training batch size, and E(*) represents expectation;
[0033] Balance the original style and the perturbed style of the predicted scale vector γ(s) and the bias vector β(s) through the parameter λ according to the following formula:
[0034] γ perb (s) = γ(s) + λ * ∈ γ (s)
[0035] β perb (s) = β(s) + λ * ∈ β (s)
[0036] where γ perb (s), β perb (s) represent the scale vector and the bias vector after style perturbation respectively, λ ~ Beta(α, α), α ∈ (0, ∞),
[0037] During the training process, activate the uncertainty style instance normalization module with a set probability, and the content representation USIN(Ec, s) after style perturbation is:
[0038]
[0039] where μ(Ec) and σ(Ec) are the expectation and standard deviation of the content vector Ec along the time dimension respectively.
[0040] Preferably, the method further includes the training step of the singing voice conversion model; including:
[0041] Add a posterior encoder and a discriminator before and after the neural vocoder of the singing voice conversion model respectively; where, the posterior encoder is an 8-layer convolutional block based on the stacking of convolution, ReLU and LayerNorm; the discriminator includes a multi-period discriminator, a multi-scale discriminator and a multi-scale multi-subband complex STFT discriminator;
[0042] Obtain the audio data sets of songs with different styles, perform source separation to remove harmonics and voice activity detection, and establish a training set;
[0043] Input the training set into the singing voice conversion model with the added posterior encoder and discriminator in sequence;
[0044] Adjust the parameters of the model by calculating the KL Loss between the posterior distribution output by the posterior encoder and the prior distribution output by the prior encoder, and calculate the feature loss function for the predicted singing voice waveform output by the neural vocoder and the real singing voice waveform by the discriminator to adjust the parameters of the model until the training requirements are met, and obtain the trained singing voice conversion model.
[0045] On the other hand, the present invention provides an end-to-end multi-scale style transfer singing voice conversion system, including:
[0046] An acquisition and preprocessing module, which is used to acquire the target singing voice to be converted and perform preprocessing to remove the accompaniment sound;
[0047] A style transfer module, which is used to input the preprocessed target singing voice and the reference singing voice of the style to be adopted into a pre-established and trained singing voice conversion model, and output a synthesized singing voice with the style of the reference singing voice to achieve style transfer;
[0048] The singing voice conversion model is used to extract the content vector and MIDI from the preprocessed target singing voice, extract the global and local style vectors, pitch and CQT spectrum from the reference singing voice, and obtain the singing voice waveform through end-to-end processing.
[0049] Compared with the prior art, the advantages of the present invention are as follows:
[0050] 1. The feature extraction module of the present invention automatically extracts the features of the singing voice. Compared with traditional singing voice conversion, the pre-trained global style encoder models the timbre and singing method of the reference singing voice, improving the singing voice style similarity of singing voice migration; the MIDI extractor extracts the MIDI information of the target singing voice, removing the style features of the target singing voice in terms of pitch and avoiding the occurrence of out-of-tune phenomena;
[0051] 2. The residual style encoder of the present invention models the fine-grained singing voice style, solving the problem that the fine-grained singing voice style cannot be modeled in classical singing voice conversion;
[0052] 3. The uncertainty style instance normalization module of the present invention effectively reduces the style leakage of the target singing voice in singing voice migration by disturbing the style information of the content vector of the target singer;
[0053] 4. The method of the present invention performs multi-scale style modeling on singing, solves the problem of poor style similarity in classical singing voice conversion, and further improves the sound quality and the naturalness of the singing voice. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a training framework diagram of an end-to-end multi-scale style transfer singing voice conversion system;
[0055] Figure 2 is an inference framework diagram of an end-to-end multi-scale style transfer singing voice conversion system;
[0056] Figure 3 Structural diagram of the residual style encoder;
[0057] Figure 4 Structural diagram of the feature extraction module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Residual Vector Quantization (RVQ) was first used for information compression and is currently widely used in the fields of Text-to-Speech (TTS) and Voice Conversion (VC) for decoupling and information extraction, etc.
[0059] The present invention attempts to construct a high-quality and expressive SVC system to solve the problem that traditional singing voice conversion cannot model fine-grained singing styles, and further improve the sound quality and naturalness. The method includes the following steps:
[0060] Collect the target singing voice to be converted and perform preprocessing to remove the accompaniment sound;
[0061] Input the preprocessed target singing voice and the reference singing voice of the intended style into a pre-established and trained singing voice conversion model, and output a synthesized singing voice with the style of the reference singing voice to achieve style transfer;
[0062] The singing voice conversion model is used to extract the content vector and MIDI from the preprocessed target singing voice, extract the global and local style vectors, pitch, and CQT spectrum from the reference singing voice, and obtain the singing voice waveform through end-to-end processing.
[0063] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0064] Embodiment 1
[0065] The present invention proposes a singing voice conversion method for end-to-end multi-scale style transfer, which can not only effectively model complex singing styles in singing voices, but also has higher sound quality and naturalness compared with the current singing voice conversion systems.
[0066] The singing voice conversion model for end-to-end multi-scale style transfer proposed by us specifically includes: a feature extraction module, a residual style encoder, an uncertainty style instance normalization module, a neural content encoder, a pitch predictor, a prior encoder, and a neural vocoder. The posterior encoder and discriminator module are added during model training. The specific block diagram is as Figure 1 shown. For the trained model used in actual use, the block diagram is as Figure 2 shown.
[0067] As Figure 3 shown is the residual style encoder, Figure 4 and this is the feature extraction module.
[0068] The method includes:
[0069] Step 1) Obtain a multi-style song audio dataset from the Internet, perform source separation, remove harmony, and conduct voice activity detection (VAD). Then, manually spot-check and remove the uncleaned data. Since the existing singing data has a single style and it is difficult to obtain multi-style singing data, the training set of the present invention consists of multi-style singing data. First, obtain multi-style song data from the Internet, unify the audio sampling rate to 44.1K, use the open-source UVR5 tool for source separation to retain the singing voice; use the open-source tool silero-vad for VAD on the singing voice; use the open-source pyannote tool to perform speaker logging first and discard the audio of multiple singers; manually spot-check some data and delete the data with poor quality, finally forming a multi-style singing dataset.
[0070] Step 2) Train the global style encoder MERT based on the training set, and extract the global style for each piece of speech in the training set. Divide the global style into timbre and singing method (e.g., pop singing method, bel canto). Use the pre-trained Content Vec to extract the content vector, the pre-trained RMVPE to extract the pitch and the first-order difference information of the pitch, and obtain the Musical Instrument Digital Interface (MIDI) notes according to the pitch quantization. Finally, calculate the CQT spectrum; the global style encoder is based on the MERT model, adding two modules composed of a linear layer and a pooling layer, and fine-tuning for the singer and singing style classification tasks on the multi-style singing dataset. The input of the global style encoder is the singing voice of 24K. The MERT model downsamples the audio by 320 times to obtain low-frame-rate features, and then inputs them into the modules composed of a linear layer and a pooling layer respectively to obtain the timbre coding vector and the singing style feature, and uses the cross-entropy loss for optimization. The trained global style encoder extracts the timbre vector Gt and the singing style vector Gs for each piece of singing voice in the training set respectively, and the sum of the two is used as the final global style information s.
[0071] The MIDI extractor first measures the tempo of the singing voice audio. According to the tempo, quantize the pitch according to the 1 / 2 note, and then quantize it according to the 1 / 4 and 1 / 8 notes in turn. Remove the shorter notes and the silent parts, and finally interpolate to the same frame rate as the Mel spectrum. In addition, calculate the first-order difference pitch information of the reference audio as the dynamic feature, which represents the vibrato and ornamentation-related features when the singer is singing, and calculate the CQT spectrum of the audio, where the CQT dimension is 108, which shows the pitch-related features compared with the Mel spectrum.
[0072] Step 3) Based on the global style vector, content vector, pitch, first-order difference of pitch, CQT, and MIDI obtained in Step 2), input to the end-to-end singing voice conversion model to output the singing voice waveform;
[0073] The content encoder is based on the Feed-Forward Transformer (FFT) structure. The input content vector is encoded from 768 dimensions to 192 dimensions, and the MIDI information of the source singing voice is encoded to 128 dimensions. After concatenating the two, they are sent into a 4-layer FFT network, and finally the encoded content vector Ec is output.
[0074] The Uncertainty Style Instance Normalization (USIN) module. The inputs of this module are the encoded content vector Ec and the global style vector s. First, the instance normalization of the content vector is calculated. The global style vector is used to predict the scale vector γ(s) and the bias vector β(s) respectively through a linear layer, and then the channel covariance matrices of the scale and bias vectors are calculated. The formula is:
[0075]
[0076] where T represents transpose, B represents the training batch size, and E(*) represents expectation;
[0077] It represents the style covariance matrix between samples in a batch. Based on this, the original bias and scale vectors are balanced between the original style and the perturbed style through a parameter. The formula is:
[0078] γ perb (s) = γ(s) + λ * ∈ γ (s)
[0079] β perb (s) = β(s) + λ * ∈ β (s)
[0080] where γ perb (s), β perd (s) represent the scale vector and the bias vector after style perturbation respectively. The parameter λ ~ Beta(α, α), where α ∈ (0, ∞), and α is set to 0.2. In actual operation, in order to balance the sound quality and the style transfer effect, USIN is activated with a set probability, and the probability is set to 0.2. The final content representation USIN(Ec, s) after style perturbation is:
[0081]
[0082] where μ(Ec) and σ(Ec) are the expectation and standard deviation of the content vector Ec along the time dimension respectively.
[0083] The input of the residual style encoder is the CQT spectrum of the reference audio, ΔF0, and the note onset identifier, where the note onset is obtained based on the change positions of MIDI. Each feature is first separately passed through a one-dimensional convolution and then concatenated, and then all features are fused through a convolution. Next, the features are fed into a convolution module, which includes a WaveNet network and a 5-layer convolution block composed of depthwise separable convolutions and Snake activation functions. Then, average pooling is performed according to the MIDI boundaries to obtain the style information at the MIDI scale. Finally, it is input into the Residual Vector Quantize (RVQ) to obtain the two-layer style information. These two layers of information are fused with the content information through the cross-attention mechanism of the FFT decoder. The local style vector ls and the encoded content vector Ec output by the content encoder are fused to output the content vector Ec' with style fusion of the same length as Ec.
[0084] The input of the pitch predictor is the output of the residual style encoder. The global style information is fused through adaptive instance normalization. The model structure is based on a 4-layer FFT, which outputs the predicted pitch. Then, the pitch is constrained within ±100 cents of the MIDI pitch to avoid out-of-tune.
[0085] The input of the prior encoder is the feature obtained by adding the predicted pitch information and the content information, which predicts the mean and variance of the prior distribution. The model structure is based on a 4-layer FFT. The dimensions are all 192. A multivariate Gaussian distribution is calculated according to the mean and variance, and a 192-dimensional hidden layer feature is sampled from the distribution.
[0086] The posterior encoder model structure is based on an 8-layer convolution block stacked with convolutions, ReLU, and LayerNorm, and is only used during the training process. The input is a 128-dimensional Mel spectrum. First, it is mapped to 192 dimensions through a convolution layer, then detailed modeling is performed through the stacked convolution block, and finally, a convolution layer is used to predict the mean and variance of the posterior distribution of the 192-dimensional feature. During the training process, it is optimized by calculating the KL (Kullback-Leibler) divergence Loss between the prior distribution.
[0087] The neural vocoder adopts a structure based on nsf HiFiGAN, which inputs the sampled features of the prior distribution and outputs the predicted singing waveform.
[0088] The discriminator adopts a multi-period discriminator, a multi-scale discriminator, and a multi-scale multi-subband complex STFT discriminator. The end-to-end model is trained based on GAN, so the discriminator is only used during the model training process. The multi-period and multi-scale discriminators are the same as those in HiFiGAN. For the multi-scale multi-subband complex STFT discriminator, the multi-scale includes three sets of parameters: FFT length, frame shift length, and window length of 1024, 120, 600; 2048, 240, 1200; 512, 50, 240. First, calculate the STFT complex spectrum of the waveform, and then divide it into multiple subbands according to the frequency ratios of (0.0, 0.1), (0.1, 0.25), (0.25, 0.5), (0.5, 0.75), (0.75, 1.0). These subbands are respectively input into two-dimensional convolutions with a kernel size of (3, 9) for 4 times and a kernel size of (3, 3) for 1 time to calculate the feature map, and then the final output is flattened into a one-dimensional feature to be jointly used as the feature of the discriminator. During training, the predicted waveform and the real waveform are respectively optimized by calculating the feature loss through the discriminator.
[0089] Embodiment 2
[0090] Embodiment 2 of the present invention provides an end-to-end multi-scale style transfer singing voice conversion system, which is implemented based on the method of Embodiment 1. The system includes:
[0091] An acquisition and preprocessing module, which is used to acquire the target singing voice to be converted and perform preprocessing to remove the accompaniment sound;
[0092] A style transfer module, which is used to input the preprocessed target singing voice and the reference singing voice of the intended style into a pre-established and trained singing voice conversion model, and output a synthesized singing voice with the style of the reference singing voice to achieve style transfer;
[0093] The singing voice conversion model is used to extract the content vector and MIDI from the preprocessed target singing voice, extract the global and local style vectors, pitch, and CQT spectrum from the reference singing voice, and obtain the singing voice waveform through end-to-end processing.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An end-to-end multi-scale style transfer singing voice conversion method, comprising: Collect the target singing voice to be converted and pre-process it to remove the accompaniment sound; The pre-processed target singing voice and the reference singing voice of the intended style are input into a pre-established and trained singing voice conversion model, and a synthesized singing voice with the style of the reference singing voice is output to achieve style transfer; The singing voice conversion model is used to extract content vectors and MIDI from the pre-processed target singing voice, extract global and local style vectors, pitch and CQT spectrum from the reference singing voice, and obtain the singing voice waveform through end-to-end processing; The singing voice conversion model includes: a feature extraction module, a residual style encoder, a content encoder, an uncertainty style instance normalization module, a pitch predictor, a priori encoder and a neural vocoder; wherein, The feature extraction module is used to extract the content vector and MIDI of the pre-processed target singing voice, and extract the pitch F0 and the first-order difference ΔF0 of F0, the global style vector s including the timbre vector Gt and the singing style vector Gs, and the CQT spectrum of the reference singing voice; The residual style encoder is used to extract a local style vector ls including modified sounds in MIDI scale according to the CQT spectrum and ΔF0 of the reference singing voice; The content encoder is used to encode the content vector from 768 dimensions to 192 dimensions, encode the MIDI to 128 dimensions, concatenate and input into a 4-layer FFT network, and obtain the encoded content vector Ec, which is respectively input into an uncertainty style instance normalization module and a residual style encoder; The uncertainty style instance normalization module is used to normalize the encoded content vector Ec instance, calculate the channel covariance matrix of the scale and bias vector of the global style vector s, and obtain the content representation after style perturbation; The pitch predictor is used to fuse the content representation after style disturbance and the global style vector s through adaptive instance normalization, and then output the predicted pitch through 4-layer FFT to constrain the pitch within the set range of MIDI pitch to avoid out-of-tune; The prior encoder uses a 4-layer FFT to calculate a multivariate Gaussian distribution based on the predicted pitch and the mean and variance of the content vector predicted prior distribution after the style perturbation of the uncertainty style instance normalization module, and obtains the sampling features of the 192-dimensional prior distribution by sampling; The neural vocoder adopts a structure based on nsf HiFiGAN, which is used to output a predicted singing waveform according to the sampling features of the prior distribution.
2. The singing voice conversion method of end-to-end multi-scale style transfer according to claim 1, characterized in that: The feature extraction module includes: a pitch extractor, a Content Vec unit, a global style encoder and a MIDI extractor; wherein, The pitch extractor is used to extract the pitch F0 of the target singing voice and the reference singing voice, and calculate the first-order difference ΔF0 of the reference singing voice F0; The Content Vec unit is used to extract a content vector from the pre-processed target singing voice; The global style encoder is used to extract the timbre vector Gt and the singing style vector Gs of the reference singing voice; The MIDI extractor is used to measure the beat of the target song, quantize F0 according to the beat, remove shorter notes and silent parts, interpolate to the same frame rate as the Mel spectrum, and calculate the CQT spectrum that expresses the pitch characteristics.
3. The singing voice conversion method of end-to-end multi-scale style transfer according to claim 2, characterized in that: The global style encoder is based on the MERT model and adds two branches with the same structure. Each branch consists of a linear layer and a pooling layer. The processing process includes: The reference singing voice with a sampling rate of 24K was downsampled 320 times to obtain the low frame rate features; Input the two branches at the same time to obtain the timbre vector Gt and the singing style vector Gs respectively.
4. The singing voice conversion method of end-to-end multi-scale style transfer according to claim 1, characterized in that: The residual style encoder includes: a front-end convolutional network, a residual quantization module and a cross-attention fusion network; its processing process includes: The CQT spectrum, ΔF0 and note start identifier of the reference singing voice are input into the front-end convolutional network and residual quantization module in sequence, and a 2-layer local style vector ls is output, where the note start identifier is obtained according to the change position of MIDI; The local style vector ls and the encoded content vector Ec output by the content encoder are fused by the cross-attention fusion network, and a content vector Ec′ with style fusion of the same length as Ec is output.
5. The singing voice conversion method of end-to-end multi-scale style transfer according to claim 1, characterized in that: The processing process of the uncertainty style instance normalization module includes: Calculate the instance normalization of the content vector Ec; The timbre vector Gt and the singing style vector Gs are added as the global style vector s, and the scale vector γ(s) and the bias vector β(s) are predicted through the linear layer respectively. Then, the channel covariance matrix ∈ of the scale vector is calculated according to the following formula: γ (s) and the channel covariance matrix ∈ of the bias vector β (s): Where T represents transpose, B represents the training batch size, and E(*) represents expectation; The original style and the perturbed style are balanced by the parameter λ for the predicted scale vector γ(s) and the bias vector β(s) according to the following formula: c perb (s)=γ(s)+λ*∈ γ (s) b perb (s)=β(s)+λ*∈ β (s) Among them, γ perb (s), β perb (s) represent the scale vector and bias vector after style perturbation, λ~Beta(α,α), α∈(0,∞), During the training process, the uncertainty style instance normalization module is activated with a set probability, and the content representation USIN(Ec, s) after style perturbation is: Among them, μ(Ec) and σ(Ec) are the expectation and standard deviation of the content vector Ec along the time dimension respectively.
6. The singing voice conversion method of end-to-end multi-scale style transfer according to claim 1, characterized in that: The method also includes a step of training a singing voice conversion model; comprising: A posterior encoder and a discriminator are added before and after the neural vocoder of the singing voice conversion model respectively; wherein the posterior encoder is an 8-layer convolution block based on convolution and ReLU and LayerNorm stacking; the discriminator includes a multi-cycle discriminator, a multi-scale discriminator and a multi-scale multi-subband complex STFT discriminator; Obtain audio data sets of songs of different styles, perform sound source separation, remove harmony, and detect voice activity to establish a training set; The training set is sequentially input into the singing voice conversion model with added posterior encoder and discriminator; The model parameters are adjusted by calculating the KL Loss between the posterior distribution output by the posterior encoder and the prior distribution output by the prior encoder. The discriminator calculates the feature loss function between the predicted singing waveform output by the neural vocoder and the real singing waveform, and the model parameters are adjusted until the training requirements are met to obtain a trained singing conversion model.
7. A system for singing voice conversion based on the end-to-end multi-scale style transfer method of claim 1, characterized in that it includes: The acquisition preprocessing module is used to acquire the target singing voice to be converted and perform preprocessing to remove the accompaniment sound; A style transfer module is used to input the pre-processed target singing voice and the reference singing voice of the intended style into a pre-established and trained singing voice conversion model, and output a synthesized singing voice with the style of the reference singing voice to achieve style transfer; The singing voice conversion model is used to extract content vectors and MIDI from the preprocessed target singing voice, extract global and local style vectors, pitch and CQT spectrum from the reference singing voice, and obtain the singing voice waveform through end-to-end processing.