The invention discloses a singing voice
conversion method based on VITS. The method comprises the following steps of: 1, respectively extracting a
bottleneck feature, a
fundamental frequency, a singing style feature and a linear
spectrogram from the singing of a source singer by utilizing a Whisper superficial layer
encoder, an automatic singing transcription model AST and an AutoVC-based
pitch-energy
encoder, and taking the
bottleneck feature, the
fundamental frequency, the singing style feature and the linear
spectrogram as the input of a VITS-based singing voice conversion model; reconstructing the singing voice by using a VITS-based singing voice conversion model in combination with the voiceprint characteristics of the singer to obtain converted singing voice of the target singer; and 2, performing
pitch adjustment according to the
pitch difference between the source singer and the target singer through a pitch shifter, and generating a target singing sound with the characteristics of the target singer by taking the
bottleneck characteristics, the F0 and the singing style characteristics as the input of a voice conversion model. The method has the characteristics of multi-level refined acoustic feature decoupling, explicit singing style migration, high-fidelity waveform generation and dynamic adaptive pitch conversion.