The invention relates to an
artificial intelligence technology, can be applied to business
system platforms of
medical health, financial science and technology and the like, and discloses a speaking face generation method, device, equipment and medium, and the method comprises the steps: obtaining a reference face image and an original video
frame sequence, carrying out the coding, and generating a potential representation, performing lip
mask processing on each frame of image in the original video
frame sequence to obtain a local mouth image, and encoding the local mouth image to generate lip features; obtaining random sampling
noise, and generating joint potential features based on the potential characterization, the lip features and the random sampling
noise; on the basis of the reference face image features, combining the potential features and the input audio features, generating lip alignment features; and after the audio and lip alignment features are decoded, lip synchronization frames are generated and assembled, and a speaking face video is generated. According to the method, high-resolution visual quality can be kept, meanwhile, deep alignment of audio content, facial expressions and
time sequence dynamic states is achieved, and the speaking face generation quality and generalization ability are improved.