The invention provides a method, a device and equipment for matching and repairing a
human mouth shape and voice content in a video, and a storage medium, and belongs to the field of
image processing, and the method comprises the steps: obtaining original video data and to-be-supplemented text information, carrying out the voice synthesis, generating voice feature parameters, and generating a
mouth shape animation sequence synchronized with the voice content. The semantic segmentation model is utilized to perform pixel-level identification on the mouth area, the physical constraint model is combined to control the tooth area, and if the tooth area exceeds the lip contour, dynamic regression adjustment is performed to ensure the naturalness and rationality of the generated
mouth shape. Furthermore, through dense cross-layer connection, a feature
pyramid structure, tooth region specialized jump connection and a tooth key point
heat map gating mechanism, the tooth region generation precision is improved; a space-channel double-attention
algorithm is introduced into a decoder, and feature expression of a mouth key area is enhanced; and a tooth region focus
loss function is adopted to optimize a generation result, so that the model robustness is improved. And simulating a
mass spring system and an antagonistic physical constraint term through a differential physical engine to realize physical rationality control of lip key points. And finally, the generated
mouth shape animation is fused with the original video, and the repaired video is output, so that the consistency and reality of the voice and the mouth shape in the video are remarkably improved.