The invention relates to the technical field of
computer vision and
artificial intelligence security, in particular to a multi-scale 3D field speaking face generation defense method,
system, equipment and medium, which actively
resist the risk of privacy abuse caused by a speaking face generation technology and solve the problems of visual quality loss, insufficient robustness and the like in the existing defense technology. Based on a space-
frequency domain mixed attention mechanism, the method specifically comprises the following steps: firstly, decomposing an input video into a
frame sequence, performing significance weighting on
noise in a
frequency domain, and inversely transforming the
noise back to a
spatial domain for screening to obtain basic defense
noise; performing multi-round iterative optimization on the basic noise through a projection
gradient descent loop; and finally, the consistency of noise on a time axis is ensured through a bidirectional
time sequence optimization strategy, and the robustness of the noise is enhanced by adopting anti-purification optimization, so that lip motion rendering of the generation model can be effectively interfered, dynamic
abnormality of a forged video is generated, and the
identifiability of deep forged contents is improved on the premise of keeping the quality of the original video.