The invention relates to the technical field of voice
semantics, can be applied to business
system platforms of financial science and technology,
medical treatment and health and the like, and discloses a controllable zero sample voice
conversion method, device, equipment and medium, the method comprises the following steps: carrying out self-supervised voice learning on unlabeled
voice data to obtain self-supervised voice representation; the method comprises the following steps of: extracting a content
feature vector and a
rhythm style vector represented by self-supervised speech, converting the content
feature vector and the
rhythm style vector into a discrete content token and a discrete
rhythm token, performing
mask generation on the discrete rhythm token to obtain a target rhythm token, obtaining reference speech of a target user, extracting user style embedding in the reference speech, and obtaining a target user. And performing
stream matching on the discrete content token, the target rhythm token and the user style embedding to generate a target Mel
spectrogram, and performing voice waveform reconstruction and optimization on the target Mel
spectrogram to obtain a zero sample voice conversion result. According to the invention, under the condition of no annotated
voice data, personalized, high-fidelity and style-consistent zero-sample voice conversion is realized.