A method, apparatus, electronic device, and storage medium for matching and repairing lip movements and speech content in a video.

By combining speech synthesis with lip-sync animation generation and introducing semantic segmentation and physical constraint models, the problem of mismatch between lip movements and speech content was solved, achieving high-quality video restoration and improving the realism and naturalness of the video.

CN121353126BActive Publication Date: 2026-05-26CLOUD ATTACK NETWORK TECH HEBEI CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CLOUD ATTACK NETWORK TECH HEBEI CO LTD
Filing Date
2025-09-01
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing video processing technologies suffer from poor viewing experience and reduced video realism due to mismatches between lip movements and audio content, particularly in terms of teeth protrusion and the naturalness of lip closure.

Method used

By combining speech synthesis with lip-sync animation generation, a semantic segmentation model and a physical constraint model are introduced to control the tooth region to not exceed the lip contour. Dense cross-layer connections, feature pyramid structure and spatial-channel dual attention algorithm are adopted, combined with tooth region focus loss function and differential physics engine to achieve synchronization and natural restoration of lip shape and speech content.

Benefits of technology

It significantly improves the realism of video restoration and the consistency between speech and vision, enhances the semantic accuracy and visual clarity of mouth shape generation, ensures that the tooth area does not exceed the lip contour, and improves the quality of the restored video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353126B_ABST
    Figure CN121353126B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, device, and storage medium for matching and repairing lip movements and speech content in a video, belonging to the field of image processing. It acquires original video data and supplementary text information, performs speech synthesis, generates speech feature parameters, and then generates a lip movement animation sequence synchronized with the speech content. A semantic segmentation model is used for pixel-level recognition of the mouth region, and a physical constraint model is combined to control the tooth region. If the tooth region exceeds the lip contour, dynamic regression adjustment is performed to ensure the naturalness and plausibility of the generated lip movement. Furthermore, the accuracy of tooth region generation is improved through dense cross-layer connections, a feature pyramid structure, specialized skip connections for the tooth region, and a heatmap gating mechanism for key tooth points. A spatial-channel dual attention algorithm is introduced into the decoder to enhance the feature representation of key mouth regions. A tooth region focus loss function is used to optimize the generation results and improve model robustness. A differential physics engine simulates a point mass spring system and adversarial physical constraints to achieve physical plausibility control of key lip points. Finally, the generated lip movement animation is fused with the original video to output the repaired video, significantly improving the consistency and realism of the speech and lip movement in the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of video processing and voice-driven animation generation, specifically to a method, storage medium, apparatus, and device for matching and repairing the lip movements of characters in a video with the voice content. Background Technology

[0002] In the field of video content editing, existing technologies address the mismatch between lip movements and spoken content by typically employing voice-driven models to generate lip animations corresponding to the spoken content, which are then integrated into the original video for restoration. Common voice-driven models can generate relatively natural lip movements based on speech features, and combined with facial landmark detection technology, they achieve alignment between the lip movements and the original video's face. However, during the generation process, due to insufficient understanding of lip closure, abnormal phenomena often occur where teeth "penetrate" beyond the outer edge of the lips, severely impacting the video's realism and visual quality. Furthermore, existing technologies lack modeling of the physical deformation patterns of the lips, resulting in unnatural lip movements during closure. Therefore, current video lip restoration systems still have significant shortcomings in controlling tooth protrusion, ensuring natural lip closure, and achieving overall integration. A comprehensive optimization scheme combining semantic segmentation, physical constraint modeling, and deep learning is urgently needed to improve restoration quality. Summary of the Invention

[0003] Based on this, in order to solve the technical problem of poor viewing experience and reduced video authenticity caused by the mismatch between the lip movements of people and the audio content in existing video processing technologies, this invention proposes a method, storage medium, device and equipment for matching and repairing the lip movements of people and the audio content in videos.

[0004] This invention protects a method for matching and repairing the lip movements of a person in a video with the corresponding audio content. The method involves acquiring the original video data and the text information to be supplemented; performing speech synthesis on the text information to generate a corresponding audio signal and extracting audio feature parameters; generating a lip movement animation sequence that matches the audio content based on the audio feature parameters; controlling the tooth region in the generated lip movement dialogue sequence based on a semantic segmentation model and a physical constraint model; if the tooth region is determined to exceed the lip contour, controlling the tooth region to return to the lip contour range; and fusing the generated lip movement animation with the original video frames to output the repaired video data.

[0005] Furthermore, controlling the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model further includes: training the semantic segmentation model using training data, which includes a face image and its corresponding mouth annotation map, with annotations accurate to the teeth, upper and lower lips, and the internal oral cavity region. The annotation map is used to guide the model to recognize mouth details; the trained semantic segmentation model is set in the encoder and decoder, where the encoder is used to extract high-level semantic features of the mouth region, and performs pixel-level semantic segmentation of the generated mouth region in the encoder to distinguish the lips, teeth, and the internal oral cavity region; spatial restoration is performed in the decoder to output a semantic segmentation map with the same resolution as the input image.

[0006] Furthermore, dense cross-layer connections are introduced between the encoder and decoder to achieve feature splicing and multi-layer feature fusion; a feature pyramid structure is added to the cross-layer connection path to perform weighted fusion of low-level and high-level features; specialized skip connections for the tooth region are added to the last layer of the decoder, and the tooth key point heatmap is used as the gating signal.

[0007] Furthermore, a spatial-channel dual attention algorithm is introduced into the decoder to enhance the attention of the semantic segmentation results of the mouth region. Channel features are extracted through global average pooling and then transformed through multiple convolutions to generate channel attention weights, thereby enhancing the channel response of key mouth regions. The maximum and average values ​​of the input feature map in the channel dimension are concatenated as spatial attention input and then transformed through convolution to generate spatial attention weights, which are used to highlight the spatial position of the mouth movement region. The final output is the element-wise product of the channel attention weights and the spatial attention weights, which is then multiplied by the original feature map to achieve dual attention weighting enhancement in the mouth shape generation process.

[0008] Furthermore, in the process of generating mouth animation sequences, the tooth region focus loss function is used to optimize the generation results of the tooth region. The tooth region focus loss function is constructed based on cross-entropy loss and a preset weight is assigned to the tooth category. When calculating the focus loss, the cross-entropy loss is calculated based on the prediction results and the target label, and the probability term is obtained through exponential operation. The weighted cross-entropy loss is weighted based on the probability term and the focus loss adjustment factor to reduce the weight of easily classified samples.

[0009] Furthermore, the process of controlling the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model further includes: implementing a mass-spring system to physically simulate the key points of the lips through a differential physics engine. The differential physics engine is built using the PyTorch framework and includes learnable spring stiffness and damping parameters. During the forward propagation process, the tooth region mask and the coordinates of the key points of the lips output by the semantic segmentation model are input, and the position of the key points of the lips is adjusted through the physical simulation process, outputting a set of lip points under physical constraints. An adversarial physical constraint term is introduced during the physical simulation process to apply a reverse force to the tooth region that exceeds the lip contour, controlling the regression of the tooth region into the lip contour.

[0010] Furthermore, a discriminant network is used to determine whether the lip movements conform to physical laws. The discriminator network adopts a three-dimensional convolutional neural network structure, which includes an input layer, multiple convolutional layers, and a fully connected output layer. The convolutional layers contain 3D convolution operations, activation functions, and normalization processing steps, and extract temporal spatial features. The high-dimensional features output by the convolutional layers are mapped to physical rationality scores through the fully connected layers. The scores are used to evaluate whether the lip movements in the generated mouth shape sequence conform to physical laws.

[0011] This invention protects an apparatus for matching and repairing the lip movements of a person in a video with the audio content, comprising: a video and text acquisition module for acquiring original video data and text information to be supplemented; a lip movement animation generation module for performing speech synthesis processing on the text information to generate corresponding speech signals and extracting speech feature parameters; generating a lip movement animation sequence that matches the audio content based on the speech feature parameters; a tooth region control module for controlling the tooth region in the generated lip movement animation sequence based on a semantic segmentation model and a physical constraint model, and controlling the tooth region to return to the lip contour range if it is determined that the tooth region exceeds the lip contour; and an image fusion module for fusing the generated lip movement animation with the original video frames to output the repaired video data.

[0012] This invention protects an electronic device, comprising: a processor, a memory, and a communication bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory interact with each other via the communication bus. The machine-readable instructions are executed by the processor to perform the steps of the method for matching and repairing the lip movements and speech content of a person in a video as described above.

[0013] This invention protects a computer-readable storage medium storing a computer program, which, when run by a processor, performs steps of a method for matching and repairing the lip movements of a person in a video with the audio content.

[0014] This invention protects a method for matching and repairing the lip movements of a person in a video with the corresponding audio content. By combining speech synthesis and lip-sync animation generation technologies, it achieves semantic content repair of the lip movements in the original video, ensuring synchronization and consistency with the newly generated audio content. The method acquires the original video data and the text information to be supplemented, generates a speech signal through speech synthesis, extracts speech feature parameters, and then generates a lip-sync animation sequence that matches the audio content. Furthermore, it combines a semantic segmentation model and a physical constraint model to control the tooth region, ensuring it does not exceed the lip contour. Finally, the generated lip-sync animation is fused with the original video frames to output the repaired video, significantly improving the realism of the video repair and the consistency between audio and visual information. Further, it introduces dense cross-layer connections and a feature pyramid structure between the encoder and decoder. Effective fusion of multi-scale features was achieved. Furthermore, the precise control over tooth region generation was enhanced through specialized skip connections and heatmap gating mechanisms for key tooth points, improving the accuracy of restoration details. Further, a spatial-channel dual attention algorithm was introduced into the decoder. By applying dual attention weighting to the semantic segmentation results in both channel and spatial dimensions, the feature representation of key mouth regions was effectively enhanced, improving the semantic accuracy and visual clarity of mouth shape generation. In addition, by introducing a differential physics engine built on PyTorch to simulate a point mass spring system, combined with learnable physical parameters and adversarial physical constraints, physical simulation of lip key points and dynamic regression control of tooth regions were achieved, further improving the restoration effect in terms of physical plausibility and visual naturalness. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0016] Figure 1 The prior art in this application contains images of exposed teeth during video fusion;

[0017] Figure 2 This application provides a flowchart of an apparatus for matching and repairing the lip movements of a person in a video with the audio content.

[0018] Figure 3 This application provides a detailed flowchart of a video lip-syncing and audio content matching and repair process.

[0019] Figure 4 This application provides a flowchart of a video repair process based on a semantic segmentation model and a physical constraint model.

[0020] Figure 5This application provides a flowchart of a three-level semantic segmentation optimization process.

[0021] Figure 6 This application provides a schematic diagram of the structure of a device for matching and repairing the lip movements of a person in a video with the audio content.

[0022] Figure 7 This application provides a schematic diagram of the structure of an electronic device. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0024] Currently, in video scenarios, the person in the original video doesn't actually open their mouth to say a certain word, but the text requires that word. This necessitates using AI technology to "fill in" the mouth movement, making it appear as if they are saying the word. However, if not handled properly, this could result in strange visuals where teeth are visible outside the mouth, so this issue still needs to be addressed.

[0025] Research has found that current video restoration technologies, when addressing the mismatch between lip movements and audio content, typically rely on manual adjustments or simple image synthesis techniques. These methods struggle to achieve automated restoration, and the resulting lip movements often lack synchronization and naturalness with the audio content. Furthermore, during the lip animation generation process, the teeth area tends to extend beyond the lip contour, leading to visual unrealistic appearances and impacting the overall video quality. Figure 1 This image screenshot demonstrates a situation where teeth are visible outside the lips during video fusion. In existing technologies, teeth often appear outside the lips during video fusion. Therefore, there is an urgent need for a method that can automatically correct the mismatch between the mouth shape and the audio content in a video and effectively control the position of the teeth.

[0026] Based on this, please refer to Figure 2 and 3This application provides a method for matching and repairing the lip movements of a person in a video with the audio content. The method includes: acquiring original video data and text information to be supplemented; performing speech synthesis processing on the text information to generate a corresponding speech signal and extracting speech feature parameters; generating a lip movement animation sequence that matches the speech content based on the speech feature parameters; controlling the tooth region in the generated lip movement dialogue sequence based on a semantic segmentation model and a physical constraint model, and if it is determined that the tooth region exceeds the lip contour, controlling the tooth region to return to the lip contour range; and performing image fusion between the generated lip movement animation and the original video frames to output the repaired video data.

[0027] This application provides a method for matching and restoring the lip movements of a person in a video to the audio content. Through speech synthesis and speech feature parameter extraction, it can automatically generate a lip movement animation sequence that highly matches the audio content, improving the synchronization and naturalness of the lip movements and speech. Simultaneously, by introducing a semantic segmentation model and a physical constraint model to dynamically control the tooth area, it ensures that the tooth area does not exceed the lip contour, enhancing the realism and visual effect of the generated lip movements. Finally, image fusion technology seamlessly embeds the generated lip movement animation into the original video frames, achieving a high-quality video restoration effect, demonstrating good practicality and application prospects.

[0028] In step S101, the original video data and the text information to be supplemented are obtained;

[0029] Specifically, this step is completed collaboratively by the video acquisition module and the text input interface. Raw video data is read from local storage or external devices via the video acquisition module. This raw video data includes video frame sequences and their corresponding raw audio tracks. The text information to be supplemented is received through the text input interface, which supports multiple methods such as manual input by the user, file import, or API interface transmission, ensuring the accuracy and completeness of the text information. When acquiring raw video data, the system parses the video format and extracts metadata information, including frame rate, resolution, encoding format, and audio sampling rate, providing a basis for subsequent parameter configuration. During the acquisition of the text information to be supplemented, the system preprocesses the text content, including removing redundant spaces, standardizing punctuation, and verifying semantic integrity, to ensure that the text information can accurately drive the subsequent speech synthesis process. In one embodiment, the system also provides a text proofreading interface, allowing users to edit and confirm the input text, thereby improving the matching degree between the final generated lip movements and the speech content.

[0030] In step S102, the text information is processed by speech synthesis to generate a corresponding speech signal, and speech feature parameters are extracted. Based on the speech feature parameters, a mouth-shape animation sequence matching the speech content is generated. Specifically, this step converts the text information to be supplemented into a continuous speech signal by calling a text-to-speech engine. The TTS engine uses a deep learning-based speech synthesis model such as Tacotron2 or FastSpeech2, which can generate natural, fluent, and richly intonation-rich speech audio. The generated speech signal is then input to the speech feature extraction module to extract speech feature parameters closely related to mouth shape changes. These speech feature parameters include, but are not limited to, phoneme sequences, fundamental frequency, energy, MFCC (Melbourne frequency cepstral coefficients), and time alignment information of speech frames. In one embodiment, the speech feature extraction module uses the Wav2Vec2 model to encode the speech signal and outputs a high-dimensional speech feature vector, which can effectively characterize the temporal structure and pronunciation features of the speech content. Based on the extracted speech feature parameters, the system further drives a lip-sync generation model to generate a lip-sync animation sequence that matches the speech content. This lip-sync generation model employs a speech-driven deep learning model such as Wav2Lip or First Order Motion Model (FOMM), capable of generating a sequence of lip-sync animation frames with high temporal synchronization and natural facial movements based on the speech feature parameters. The generated lip-sync animation sequence is output as an image sequence or video clip for subsequent fusion processing with the original video frames.

[0031] In step S103, the tooth region in the generated mouth-shaped dialogue sequence is controlled based on a semantic segmentation model and a physical constraint model. If the tooth region is determined to exceed the lip contour, it is controlled to return to the lip contour range. Specifically, this step uses a combination of a semantic segmentation model and a physical constraint model to precisely control the tooth region in the generated mouth-shaped animation sequence, preventing the tooth region from exceeding the lip contour during generation and causing visual anomalies. The semantic segmentation model performs pixel-level semantic segmentation of the generated mouth region, identifying regions such as the lips, teeth, and the inside of the mouth, providing accurate region boundary information for subsequent physical constraints. The physical constraint model dynamically simulates the closed state of the lips based on the physical deformation laws of the lips. When the tooth region is detected to exceed the lip contour, the position of key points on the lips is adjusted through a physical constraint mechanism to force the tooth region back to the lip contour range.

[0032] The control of the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model further includes:

[0033] In step S1031, the semantic segmentation model is trained using training data, which includes face images and their corresponding mouth annotation maps. The annotations are accurate down to the teeth, upper and lower lips, and the interior of the oral cavity. These annotation maps guide the model to recognize mouth details. Specifically, the semantic segmentation model employs an improved UNet architecture. The training dataset includes publicly available datasets and a self-constructed, finely annotated mouth region dataset, with annotation accuracy down to the pixel level, ensuring the model can accurately identify the boundary between the lips and teeth. During training, data augmentation processing is applied to the input images, including random cropping, rotation, and lighting changes, to improve the model's generalization ability. The loss function uses a weighted combination of Dice Loss and cross-entropy loss to improve the model's segmentation accuracy of edge regions, especially its ability to recognize the boundary between the lips and teeth.

[0034] In step S1032, the trained semantic segmentation model is incorporated into the encoder and decoder. The encoder extracts high-level semantic features of the mouth region, performing pixel-level semantic segmentation on the generated mouth region to distinguish between the lips, teeth, and the internal oral cavity. Specifically, the encoder uses ResNet-50 as the backbone network to extract multi-scale features from the input image and extracts high-level semantic information through multi-layer convolution and pooling operations. The decoder uses upsampling operations to restore spatial resolution and outputs a semantic segmentation map with the same resolution as the input image, ensuring the localization accuracy of the tooth region.

[0035] In step S10321, dense cross-layer connections are introduced between the encoder and decoder to achieve feature stitching and multi-layer feature fusion. Specifically, this cross-layer connection structure uses Dense Block modules, which enhance the model's ability to perceive local details through layer-by-layer feature stitching, thereby improving the recognition accuracy of the lip edge and tooth region.

[0036] Specifically, a dense cross-layer connection design is adopted, with each Dense Block containing 4 convolutional layers and a dense connection mode of growth_rate=32. The output of the previous layer is concatenated with the input of all subsequent layers through channel dimension, forming a composite feature representation that includes low-level visual features, mid-level geometric features, and high-level semantic features. This design significantly improves the model's ability to recognize the lip-tooth junction region, improving the segmentation F1-score for this region by 12.6%.

[0037] In step S10322, a feature pyramid structure is added to the cross-layer connection path to perform weighted fusion of low-level and high-level features. Low-level features retain the edge and texture information of the image, while high-level features contain semantic information. Through the weighted fusion strategy, the model can simultaneously take into account details and overall structure during semantic segmentation, thereby improving the recognition accuracy of the tooth region.

[0038] Specifically, the following feature pyramid structure is adopted: dynamic weights are generated through learnable 1×1 convolutions and a sigmoid activation function to perform weighted fusion of low-level edge features and high-level semantic features. The low-level features (from conv2_x) preserve edge details with a precision of 2px, while the high-level features (from conv5_x) provide semantic context information. This multi-scale feature fusion strategy enables the model to simultaneously capture the fine contours of teeth and their overall spatial relationships.

[0039] In step S10323, specialized skip connections for the tooth region are added to the last layer of the decoder, and a heatmap of tooth key points is used as a gating signal. This gating signal is generated by a pre-trained key point detection model, with the preferred model being MediaPipe, which guides the model to perform more refined segmentation in the tooth region, thereby improving the localization accuracy and boundary clarity of the tooth region.

[0040] The specific gating signal is controlled as follows: First, MediaPipe is used to generate 78 facial landmarks, and Gaussian kernel diffusion is used to generate a heatmap of the tooth region. At the decoder end, this heatmap is used as the gating signal to perform a Hadamard product operation with the segmentation feature map, where the heatmap guidance intensity coefficient λ is set to 1.5. This design significantly improves the Dice coefficient of the tooth region from 0.83 to 0.91. Simultaneously, we also use a conditional random field as a post-processing module, further improving the visual coherence of the tooth boundaries through color similarity and spatial proximity constraints.

[0041] Therefore, based on the features extracted by the encoder, multi-scale information is fused across layers, and finally, a decoder guided by thermal tooth images performs refined segmentation. Dense cross-layer connections and a feature pyramid structure are introduced between the encoder and decoder to achieve effective fusion of multi-scale features. Simultaneously, through tooth region-specific skip connections and a thermal image gating mechanism for key tooth points, precise control over tooth region generation is enhanced, improving the accuracy of restoration details.

[0042] The process of controlling the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model further includes:

[0043] In step S1033, a differential physics engine is used to physically simulate key points of the lips using a point-spring system. This differential physics engine, built using the PyTorch framework, includes learnable spring stiffness and damping parameters. Specifically, the physical constraint model is constructed based on the physical deformation laws of the lips, employing a point-spring system to model key points of the lips. The spring stiffness and damping coefficients are learned parameters, continuously optimized during training to enable the model to more realistically simulate the closing and opening of the lips.

[0044] Among them, the spring stiffness parameter controls the elasticity of lip deformation, while the damping coefficient affects the energy dissipation characteristics during movement. During training, these two parameters are automatically adjusted based on the tooth position information provided by the semantic segmentation module. When a potential risk of tooth penetration is detected, the system dynamically increases the stiffness coefficient of the local area, so that the lips can more "firmly" wrap around the teeth.

[0045] In step S1034, during the forward propagation process, the tooth region mask and lip keypoint coordinates output by the semantic segmentation model are input. The positions of the lip keypoints are adjusted through a physical simulation process, and a set of lip points under physical constraints is output. Specifically, the physical simulation process uses the semantic segmentation results as input constraints and combines them with the lip keypoint coordinates to simulate the deformation trajectory of the lips in a closed state. When a tooth region is detected to be located outside the lips, the lip deformation path is corrected through a reverse force or constraint function to ensure that the teeth are always within the lip contour.

[0046] This step achieves a deep fusion of semantic information and physical simulation. Specifically, the system converts the tooth mask into a distance field, where the distance is represented by `Distance_Field` in the code. When the distance between the lip particles and the tooth boundary is less than a threshold, a repulsive force calculation based on Hooke's law is triggered. Simultaneously, to maintain the natural deformation of the lips as a whole, the system employs a layered constraint strategy: inner-layer particles are subject to strong physical constraints to prevent masking, while outer-layer particles retain more space for free deformation. This fine control enables sub-pixel level protection accuracy for the tooth region at 4K resolution, for example, an error of <0.5px.

[0047] During the physical simulation, the system first converts the binarized tooth mask into a signed distance field (DF). By calculating the shortest Euclidean distance from each lip keypoint to the tooth mask boundary, a continuous distance field representation is generated: positive values ​​indicate that the particle is located outside the tooth, and negative values ​​indicate that it penetrates the tooth. This distance field is efficiently implemented using PyTorch's vectorization operations, supporting automatic differentiation to be embedded in the training process. When the distance between the particle and the tooth boundary is less than a preset threshold (e.g., 0.5 pixels), the system triggers the calculation of a repulsive force based on the distance field. The magnitude of the force is linearly related to the penetration depth, following Hooke's law, and its direction is along the gradient direction of the distance field, forcing the particle to return to the safe region.

[0048] In step S1035, a discriminant network is used to determine whether the lip movements conform to physical laws. The discriminant network adopts a three-dimensional convolutional neural network structure, which includes an input layer, multiple convolutional layers, and a fully connected output layer. This discriminant network is used to evaluate the lip movements during the physical simulation process to ensure that the generated mouth animation is physically reasonable and to avoid abnormal phenomena such as teeth wearing through the mold.

[0049] This discriminative network receives multiple consecutive frames of lip region images as input and extracts spatiotemporal features through 3D convolutional layers. Each convolutional layer includes 3D convolutional kernel operations, a LeakyReLU activation function, and instance normalization. This combination effectively captures the temporal dynamics of lip movements while maintaining training stability. Deeper layers employ dilated 3D convolutions to expand the receptive field, paying particular attention to the interaction between the lips and teeth during pronunciation. Finally, fully connected layers map the extracted spatiotemporal features to a physical plausibility score between 0 and 1. This score reflects not only the naturalness of the overall movement but also detects anomalies such as tooth distortion. In practical applications, the discriminator performs sliding detection in 5-frame windows. When the score falls below a threshold (e.g., 0.6), the system automatically triggers the physical constraint module's reinforcement correction mechanism. During training, the discriminator uses special sample augmentation strategies, including simulating motion blur during rapid pronunciation and changes in tooth visibility under extreme lighting conditions, giving it robust judgment capabilities for various edge cases. Compared to traditional rule-based detection methods, this learning-based discriminator can capture more subtle physical anomalies, reducing the false negative rate of tooth profiling by 63%. Furthermore, the features of the last two layers of the network are visualized for debugging purposes, allowing engineers to visually see which areas' motion patterns are judged to be illogical.

[0050] A discriminant network is used to determine whether lip movements conform to physical laws. This discriminant network employs a three-dimensional convolutional neural network structure, comprising an input layer, multiple convolutional layers, and a fully connected output layer. This discriminant network evaluates adversarial constraints introduced during the physical simulation process, ensuring that while improving control of the tooth area, it does not affect the naturalness and coherence of the overall mouth animation.

[0051] In step S1036, an adversarial physics constraint is introduced during the physical simulation process to apply a reverse force to the tooth region that extends beyond the lip contour, controlling the regression of the tooth region back into the lip contour. Specifically, this adversarial physics constraint, as part of the loss function, guides the model to correct the tooth region during training, ensuring that the generated mouth animation is visually natural and conforms to physical laws.

[0052] This constraint term, as part of the loss function, consists of three parts:

[0053] 1. Penalty for clipping (scheme code uses Penetration Loss) Lp Based on the distance field of the tooth mask, the depth of the penetration area is calculated. Smooth L1 Loss is used to penalize the tooth region that exceeds the lip contour to ensure that the mass point returns to a safe position.

[0054] 2. Motion Naturalness Loss (Procedure code uses Motion Naturalness Loss) Lm The mean square error (MSE) is calculated from the physical plausibility score output by the discriminant network, forcing the simulation results to conform to the dynamic characteristics of real lip movements (such as inertia and elastic recovery).

[0055] 3. Deformation Smoothness Constraint (The solution code uses Deformation Smoothness Loss) Ls ): By calculating the second-order difference of the displacements of adjacent key points, unreasonable local abrupt changes are suppressed, and the overall continuity of deformation is maintained.

[0056] The final loss function is a weighted combination:

[0057] Ltotal = λpLp + λmLm + λsLs

[0058] Among them, the weight λp is dynamically increased (up to 5 times) when clipping is detected, while λm and λs are adaptively adjusted according to the pronunciation speed, and the smoothness constraint is reduced to allow for more drastic deformation when pronunciation is fast.

[0059] The working mechanism of adversarial constraints: During forward simulation, abnormal displacement in the tooth region triggers high gradient backpropagation by the APC (Adversarial Process Control), which adjusts the spring stiffness and damping parameters in reverse through a differentiable physics engine, locally enhancing the physical constraint force. Simultaneously, the adversarial gradients provided by the discriminant network further correct the motion trajectory, making it approximate the distribution of real pronunciation data while avoiding tooth penetration. Experiments show that this scheme reduces the tooth penetration rate to below 0.3% while maintaining 98% naturalness.

[0060] In step S104, the generated lip-shape animation is fused with the original video frames to output the repaired video data. Specifically, this step uses an image fusion module to seamlessly fuse the generated lip-shape animation sequence with the original video frames at the pixel level, ensuring that the generated area and the original video are consistent in terms of lighting, skin tone, and boundary transitions, thereby achieving a natural and visually imperceptible repair effect. During the image fusion process, the system first extracts key point information of the facial region in the original video frames based on a facial key point detection model, including the contours of the eyes, nose, and mouth, to achieve precise alignment between the generated lip-shape animation and the original facial structure. Subsequently, the system uses image fusion technology to fuse the generated lip-shape region with the original facial region, so that the fused lip-shape blends naturally with the surrounding area in terms of color, texture, and lighting.

[0061] The system employs Laplacian pyramid decomposition technology, performing feature fusion at three different scales (original resolution, 1 / 2 downsampling, and 1 / 4 downsampling) to ensure a smooth transition from macroscopic contours to microscopic textures. Each scale layer uses an adaptive weighted blending strategy, with the high-frequency layer (detail layer) biased towards generating mouth shapes, and the low-frequency layer (base layer) biased towards the original video. The fusion process integrates a physically based rendering lighting model, dynamically adjusting the lighting properties of the mouth shape generation region by analyzing the ambient occlusion map and specular reflection parameters of the original video frames. Specifically, it generates matching subsurface scattering effects in real time for different mouth shapes (such as the internal shadows produced by an "O" shaped mouth). A CUDA-accelerated parallel fusion pipeline optimizes the traditional Poisson fusion algorithm into block-based processing, combining NVIDIA's Tensor Cores for blending precision calculations, keeping the single-frame fusion time at 4K resolution within 16ms (60fps real-time requirement).

[0062] In one implementation, the system further introduces a super-resolution reconstruction model (such as Real-ESRGAN) to enhance the details of the fused region, improving the clarity and realism of the generated region to match the quality of the original video. Finally, the fused video frame sequence is encapsulated by a video encoder according to the encoding format and frame rate of the original video, outputting the repaired video data, thus completing the entire lip-sync and speech content matching and repair process.

[0063] In step S1032, the trained semantic segmentation model is set in the encoder and decoder. The encoder is used to extract high-level semantic features of the mouth region. The encoder performs pixel-level semantic segmentation of the generated mouth region, distinguishing between the lips, teeth, and the internal oral cavity region. Specifically, the semantic segmentation model introduces a tooth region focus loss function during training to optimize the generated tooth region, thereby improving the accuracy and robustness of the model in tooth region recognition and localization.

[0064] The tooth region focus loss function is constructed based on cross-entropy loss and assigns preset weights to tooth categories. Specifically, this focus loss function enhances the model's attention to the tooth region by introducing category weights and focus adjustment factors on top of cross-entropy loss, maintaining high recognition accuracy even when tooth edges are blurred or occluded. The weight of the tooth category is set to a multiple of other categories (such as upper and lower lips, and the interior of the mouth), for example, 10 times, to compensate for the training bias caused by the small proportion of the tooth region in the overall mouth region.

[0065] When calculating the focus loss, the cross-entropy loss is calculated based on the prediction results and the target label, and the probability term is obtained through exponential operation. Specifically, the probability value of each pixel output by the model belonging to a certain category is used to calculate the cross-entropy loss term. Then, the focus adjustment factor is generated through exponential operation of the probability term. This adjustment factor is dynamically adjusted according to the difficulty of sample classification, so that the model pays more attention to tooth region samples that are difficult to classify during training.

[0066] The weighted cross-entropy loss is weighted based on the probability term and the focus loss adjustment factor, reducing the weight of easily classified samples. Specifically, the focus loss function strengthens the model's learning of the tooth region through a triple mechanism: first, the standard cross-entropy loss is calculated as the basic error measure; then, the prediction confidence pt is obtained through exponential operation, and a dynamic weighting system is constructed—a fixed weight of 10 times is applied to the tooth category, with a base weight of 1 times and an additional 9 times. Simultaneously, (1− pt ) gamma The focus adjustment factor, where gamma defaults to 2, enables the model to automatically focus on low-confidence, hard-to-classify samples (such as tooth edges / occluded areas). Finally, the class weights, focus factor, and base loss are multiplied and averaged to create a gradient amplification effect. This design combines hard weight allocation with soft dynamic adjustment to address training bias caused by the small pixel proportion in the tooth region while maintaining normal learning in other areas, particularly optimizing segmentation accuracy in scenes with blurred edges. The specific code for the loss function is as follows:

[0067] class TeethFocalLoss(nn.Module):

[0068] def __init__(self, gamma=2):

[0069] super().__init__()

[0070] self.gamma = gamma

[0071] def forward(self, pred, target):

[0072] ce_loss = F.cross_entropy(pred, target, reduction='none')

[0073] pt = torch.exp(-ce_loss)

[0074] # Increase the weight of tooth categories (e.g., class=2) by 10 times.

[0075] focal_weight = (target == 2).float() * 9 + 1

[0076] loss = (focal_weight * (1-pt)**self.gamma * ce_loss).mean()

[0077] return loss

[0078] By introducing this loss function, the model can effectively suppress the dominant role of easily classified samples in gradient updates during training, thereby improving the boundary clarity and localization accuracy of the tooth region in the semantic segmentation map, and providing a more reliable tooth region mask input for subsequent physical constraint models.

[0079] By introducing a tooth region focus loss function during the training of the semantic segmentation model, the system can more accurately identify the boundaries of the tooth region when generating mouth animation sequences. This allows for more effective judgment of whether the teeth exceed the lip contour in the physical constraint model, and corresponding regression control, further improving the realism and naturalness of the overall restoration effect.

[0080] Please see Figure 6 , Figure 6 This is a schematic diagram of a device for matching and repairing the lip movements of a person in a video with the audio content, provided as an embodiment of this application. Figure 6 As shown, the repair device 200 includes:

[0081] The video and text acquisition module 210 is used to acquire raw video data and text information to be supplemented;

[0082] The lip-sync animation generation module 220 is used to perform speech synthesis processing on the text information, generate a corresponding speech signal, and extract speech feature parameters; and generate a lip-sync animation sequence that matches the speech content based on the speech feature parameters.

[0083] The tooth region control module 230 is used to control the tooth region in the generated mouth animation sequence based on the semantic segmentation model and the physical constraint model. If it is determined that the tooth region exceeds the lip contour, the tooth region is controlled to return to the lip contour range.

[0084] The image fusion module 240 is used to fuse the generated lip-sync animation with the original video frames and output the repaired video data.

[0085] Furthermore, when the tooth region control module 230 controls the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model, it specifically includes:

[0086] The semantic segmentation model is trained using training data, which includes face images and their corresponding mouth annotation maps. The annotations are accurate to the teeth, upper and lower lips, and the internal area of ​​the oral cavity. The annotation maps are used to guide the model to recognize mouth details.

[0087] The trained semantic segmentation model is set in the encoder and decoder. The encoder is used to extract high-level semantic features of the mouth region. The generated mouth region is semantically segmented at the pixel level in the encoder to distinguish the lips, teeth and the internal oral cavity region.

[0088] Spatial restoration is performed in the decoder, and the output is a semantic segmentation map with the same resolution as the input image.

[0089] Furthermore, the tooth region control module 230, when used to control the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model, also includes:

[0090] Dense cross-layer connections are introduced between the encoder and decoder to achieve feature splicing and multi-layer feature fusion;

[0091] A feature pyramid structure is added to the cross-layer connection path to perform weighted fusion of low-level and high-level features;

[0092] Specialized skip connections for the tooth region are added to the last layer of the decoder, and a heatmap of the tooth key points is used as the gating signal.

[0093] Furthermore, the tooth region control module 230, when used to control the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model, also includes:

[0094] A spatial-channel dual attention algorithm is introduced into the decoder to enhance the attention of the semantic segmentation results of the mouth region.

[0095] Channel features are extracted by global average pooling, and channel attention weights are generated after multi-layer convolution transformation to enhance the channel response of key mouth regions.

[0096] By concatenating the maximum and average values ​​of the input feature map in the channel dimension, the spatial attention input is used as the spatial attention input. The spatial attention weights are generated through convolution operations to highlight the spatial location of the mouth movement region.

[0097] The final output is the element-wise product of the channel attention weights and the spatial attention weights, which is then multiplied by the original feature map to achieve dual attention weighting enhancement in the mouth shape generation process.

[0098] Furthermore, the tooth region control module 230, when used to control the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model, also includes:

[0099] In the process of generating mouth animation sequences, the tooth region focus loss function is used to optimize the generation results of the tooth region;

[0100] The tooth region focus loss function is constructed based on cross-entropy loss and assigns preset weights to tooth categories;

[0101] When calculating the focus loss, the cross-entropy loss is calculated based on the prediction results and the target label, and the probability term is obtained through exponential operation.

[0102] The weighted cross-entropy loss is weighted based on the probability term and the focus loss adjustment factor to reduce the weight of easily classified samples.

[0103] Furthermore, the tooth region control module 230, when used to control the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model, also includes:

[0104] A differential physics engine is used to simulate the physical characteristics of key points on the lips using a point mass spring system. The differential physics engine is built using the PyTorch framework and includes learnable spring stiffness and damping parameters.

[0105] During the forward propagation process, the tooth region mask and lip key point coordinates output by the semantic segmentation model are input, and the positions of the lip key points are adjusted through a physical simulation process to output a set of lip points under physical constraints.

[0106] An adversarial physics constraint is introduced during the physics simulation to apply a reverse force to the tooth region that extends beyond the lip contour, thereby controlling the regression of the tooth region back into the lip contour.

[0107] Furthermore, the tooth region control module 230, when used to control the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model, also includes:

[0108] The discriminator network determines whether the lip movements conform to physical laws. The discriminator network adopts a three-dimensional convolutional neural network structure, which includes an input layer, multiple convolutional layers, and a fully connected output layer.

[0109] The convolutional layer includes 3D convolution operations, activation functions, and normalization steps, and the convolutional layer extracts temporal spatial features.

[0110] The fully connected layer maps the high-dimensional features output by the convolutional layer to a physical plausibility score, and the score is used to evaluate whether the lip movements in the generated mouth shape sequence conform to physical laws.

[0111] The apparatus for matching and repairing the lip movements of a person in a video with the audio content provided in this application embodiment acquires the original video data and the text information to be supplemented; performs speech synthesis processing on the text information to generate a corresponding speech signal and extracts speech feature parameters; generates a lip movement animation sequence that matches the speech content based on the speech feature parameters; controls the tooth region in the generated lip movement animation sequence based on a semantic segmentation model and a physical constraint model; if it is determined that the tooth region exceeds the lip contour, it controls the tooth region to return to the lip contour range; and performs image fusion with the generated lip movement animation and the original video frames to output the repaired video data. In this way, it can automatically repair the problem of mismatch between the lip movements of a person in a video and the audio content, improve the realism and naturalness of the video content, and enhance the user experience.

[0112] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 300 includes a processor 310, a memory 320, and a communication bus 330.

[0113] The memory 320 stores machine-readable instructions that can be executed by the processor 310. When the electronic device 300 is running, the processor 310 and the memory 320 interact with each other via the communication bus 330. When the machine-readable instructions are executed by the processor 310, the steps of the method for matching and repairing the lip movements of people in the video with the voice content in the above method embodiment can be performed. For specific implementation methods, please refer to the method embodiment, which will not be repeated here.

[0114] This application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it can execute the steps of the method for matching and repairing the lip movements of a person in a video with the audio content as described in the above method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0115] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0116] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0117] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0119] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0120] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for matching and repairing the lip movements of a person in a video with the audio content, characterized in that, Obtain the raw video data and the text information to be supplemented; The text information is processed by speech synthesis to generate a corresponding speech signal, and speech feature parameters are extracted; a mouth shape animation sequence matching the speech content is generated based on the speech feature parameters. Based on semantic segmentation and physical constraint models, the tooth region in the generated mouth animation sequence is controlled. If it is determined that the tooth region exceeds the lip contour, the tooth region is controlled to return to the lip contour range. The generated lip-sync animation is fused with the original video frames to output the repaired video data. The semantic segmentation model is trained using training data, which includes face images and their corresponding mouth annotation maps. The annotations are accurate to the teeth, upper and lower lips, and the internal area of ​​the oral cavity. The annotation maps guide the model to recognize mouth details. The trained semantic segmentation model is set in the encoder and decoder. The encoder is used to extract high-level semantic features of the mouth region. The generated mouth region is semantically segmented at the pixel level in the encoder to distinguish the lips, teeth and the internal oral cavity region. Spatial restoration is performed in the decoder, and the output is a semantic segmentation map with the same resolution as the input image. Dense cross-layer connections are introduced between the encoder and decoder to achieve feature splicing and multi-layer feature fusion; A feature pyramid structure is added to the cross-layer connection path to perform weighted fusion of low-level and high-level features; Specialized skip connections for the tooth region are added to the last layer of the decoder, and a heatmap of the tooth key points is used as the gating signal.

2. The method for matching and repairing lip movements and audio content in a video according to claim 1, characterized in that, A spatial-channel dual attention algorithm is introduced into the decoder to enhance the attention of the semantic segmentation results of the mouth region. Channel features are extracted by global average pooling, and channel attention weights are generated after multi-layer convolution transformation to enhance the channel response of key mouth regions. By concatenating the maximum and average values ​​of the input feature map in the channel dimension, the spatial attention input is used as the spatial attention input. The spatial attention weights are generated through convolution operations to highlight the spatial position of the mouth movement region. The final output is the element-wise product of the channel attention weights and the spatial attention weights, which is then multiplied by the original feature map to achieve dual attention weighting enhancement in the mouth shape generation process.

3. The method for matching and repairing lip movements and audio content in a video according to claim 2, characterized in that, In the process of generating mouth animation sequences, the tooth region focus loss function is used to optimize the generation results of the tooth region; The tooth region focus loss function is constructed based on cross-entropy loss and assigns preset weights to tooth categories; The cross-entropy loss is calculated based on the prediction results and the target label, and the probability term is obtained through exponential operation. The weighted cross-entropy loss is weighted based on the probability term and the focus loss adjustment factor to reduce the weight of easily classified samples.

4. The method for matching and repairing lip movements and audio content in a video according to claim 1, characterized in that, The process of controlling the tooth region in the generated mouth shape based on the semantic segmentation model and the physical constraint model further includes: A differential physics engine is used to simulate the physical characteristics of key points on the lips using a point mass spring system. The differential physics engine is built using the PyTorch framework and includes learnable spring stiffness and damping parameters. During the forward propagation process, the tooth region mask and lip key point coordinates output by the semantic segmentation model are input, and the positions of the lip key points are adjusted through a physical simulation process to output a set of lip points under physical constraints. An adversarial physics constraint is introduced during the physics simulation to apply a reverse force to the tooth region that extends beyond the lip contour, thereby controlling the regression of the tooth region back into the lip contour.

5. The method for matching and repairing lip movements and audio content in a video according to claim 4, characterized in that, The discriminant network determines whether the lip movements conform to physical laws. The discriminant network adopts a three-dimensional convolutional neural network structure, which includes an input layer, multiple convolutional layers, and a fully connected output layer. The convolutional layer includes 3D convolution operations, activation functions, and normalization steps, and the convolutional layer extracts temporal spatial features. The high-dimensional features output by the convolutional layer are mapped to a physical plausibility score through the fully connected output layer. The score is used to evaluate whether the lip movements in the generated mouth shape sequence conform to physical laws.

6. A device for matching and repairing the lip movements of a person in a video with the audio content, the device using the method described in any one of claims 1-5, characterized in that, include: The video and text acquisition module is used to acquire raw video data and supplementary text information; The lip-sync animation generation module is used to perform speech synthesis processing on the text information, generate corresponding speech signals, and extract speech feature parameters; and generate a lip-sync animation sequence that matches the speech content based on the speech feature parameters. The tooth region control module is used to control the tooth region in the generated mouth animation sequence based on the semantic segmentation model and the physical constraint model. If it is determined that the tooth region exceeds the lip contour, the tooth region is controlled to return to the lip contour range. The image fusion module is used to fuse the generated lip-sync animation with the original video frames and output the repaired video data.

7. An electronic device, characterized in that, include: The device includes a processor, a memory, and a communication bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory interact with each other via the communication bus. The machine-readable instructions are executed by the processor to perform the steps of the method for matching and repairing lip movements and speech content in a video as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method for matching and repairing the lip movements of a person in a video with the audio content as described in any one of claims 1 to 5.