This application provides a method, apparatus, device, and storage medium for matching and repairing lip movements and speech content in a video, belonging to the field of
image processing. It acquires original video data and supplementary text information, performs
speech synthesis, generates speech feature parameters, and then generates a lip movement
animation sequence synchronized with the speech content. A semantic segmentation model is used for pixel-level recognition of the
mouth region, and a physical constraint model is combined to control the tooth region. If the tooth region exceeds the lip contour, dynamic regression adjustment is performed to ensure the naturalness and plausibility of the generated lip movement. Furthermore, the accuracy of tooth region generation is improved through dense cross-layer connections, a feature
pyramid structure, specialized skip connections for the tooth region, and a heatmap gating mechanism for key tooth points. A spatial-channel dual attention
algorithm is introduced into the decoder to enhance the feature representation of key mouth regions. A tooth region focus
loss function is used to optimize the generation results and improve model robustness. A differential
physics engine simulates a point
mass spring system and adversarial physical constraints to achieve physical plausibility control of key lip points. Finally, the generated lip movement
animation is fused with the original video to output the repaired video, significantly improving the consistency and realism of the speech and lip movement in the video.