The application discloses an
active defense method, device and equipment for
diffusion speech conversion and a medium, and relates to the technical field of speech security. The
active defense method comprises the following steps: obtaining source speech and reference speech, introducing a constrained protection disturbance into the reference speech, constructing protected speech, and ensuring that the protection disturbance satisfies a disturbance imperceptibility constraint. The source speech and the protected speech are input into a speech conversion model with the reference speech as a speaker condition, content
information analysis, speaker condition constraint acoustic generation and waveform synthesis are completed by the speech conversion model, and a protection generated speech is output. A joint
loss function is constructed based on speaker embedding representations of the protection generated speech and the reference speech, gradients of a current disturbance position and a predicted midpoint position are calculated according to the joint
loss function and are fused to obtain an update direction, the protection disturbance is iteratively optimized, and projection or clipping is performed to satisfy the disturbance imperceptibility constraint. Finally, the protected speech is output.