The application discloses a song generation method and device, equipment and storage medium, belong to
artificial intelligence technical field. The application can generate a song with specific human voice according to user demand. In the application, the song generation request is first obtained, and the song attribute information and the reference sound sample are included in the song generation request. Then, the song is generated according to the song attribute information. Since the user demand is considered when generating the song, the generated song is more consistent with the user expectation, and the song quality is improved. In addition, the scheme also includes a sound
imitation process, that is, the
encoder of the sound
imitation model is called to map the reference sound sample to a
latent vector, and the decoder of the sound
imitation model is called to generate an imitation audio consistent with the sound characteristics of the human voice in the reference sound sample according to the
latent vector. Next, the initial audio and the imitation audio are synthesized to generate a song with specific human voice, which enriches the song generation method.