Video processing method and device, electronic equipment and storage medium

By extracting standardized images and temporal deformation fields, the problems of low efficiency and discontinuity in facial animation video generation are solved, achieving efficient and smooth facial animation video generation, which is suitable for video processing of complex facial movements and expressions.

CN121999404APending Publication Date: 2026-05-08BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2024-11-08
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies are computationally inefficient and inconsistent when generating animated facial videos. In particular, when dealing with large facial movements, they are prone to producing videos that are disjointed and awkward, affecting the smoothness of the video and the user experience.

Method used

By extracting canonical images and temporal deformation fields, including facial motion information, from reference face videos, and stylizing the canonical images, facial animation videos are generated. This avoids processing each video frame individually, improving generation efficiency and maintaining the overall smoothness of the video.

Benefits of technology

It significantly improves the efficiency of generating animated facial videos, avoids inconsistencies and awkwardness, and enhances the overall smoothness and visual quality of the videos. It is especially suitable for handling large-scale head movements and complex expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999404A_ABST
    Figure CN121999404A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video processing method and device, electronic equipment and a storage medium. The method comprises the following steps: extracting a standard image from a reference face video; extracting a time deformation field of each video frame relative to the standard image from the reference face video, wherein the time deformation field comprises face motion information; performing specified style processing on the standard image to obtain a stylized image; and generating a face animation video according to the stylized image and the time deformation field. According to the embodiment of the invention, the generation efficiency of the face animation video is improved, the incoherence and violation of the face animation video are avoided, and the overall fluency of the video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a video processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] While current technologies have made some progress in the field of facial animation video generation, they still face numerous challenges and limitations in practical applications. Firstly, long generation times and low computational efficiency are significant problems, as the entire video needs to be processed. Although some methods can generate high-quality video content, these often involve substantial computational resource consumption, resulting in slow generation speeds that fail to meet the demands of real-time applications. This bottleneck in computational efficiency severely hinders the widespread application of the technology, making it imperative to improve the efficiency and real-time performance of the models a pressing issue.

[0003] Furthermore, existing technologies often exhibit video discontinuity and awkwardness when handling large facial movements. Due to limitations in capturing complex motion details, especially when dealing with dramatic head movements or multi-angle changes, the generated videos often suffer from misalignment and image distortion. This unnatural inter-frame motion not only affects the overall smoothness of the video but also degrades the user's viewing experience. Summary of the Invention

[0004] This application provides a video processing method, apparatus, electronic device, and storage medium, which helps to improve the generation efficiency of animated facial videos and enhance the overall smoothness of the videos.

[0005] To address the aforementioned problems, in a first aspect, embodiments of this application provide a video processing method, including:

[0006] Extract canonical images from reference face videos;

[0007] Extract the temporal deformation field of each video frame relative to the canonical image from the reference face video, the temporal deformation field including facial motion information;

[0008] The standardized image is processed in a specified style to obtain a stylized image;

[0009] An animated video of a face is generated based on the stylized image and the temporal deformation field.

[0010] Secondly, embodiments of this application provide a video processing apparatus, including:

[0011] The standardized image extraction module is used to extract standardized images from reference face videos;

[0012] A temporal deformation field extraction module is used to extract the temporal deformation field of each video frame relative to the canonical image from the reference face video, wherein the temporal deformation field includes facial motion information;

[0013] The stylization processing module is used to process the standard image in a specified style to obtain a stylized image.

[0014] The video generation module is used to generate a facial animation video based on the stylized image and the time deformation field.

[0015] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video processing method described in embodiments of this application.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the video processing method disclosed in embodiments of this application.

[0017] The video processing method, apparatus, electronic device, and storage medium provided in this application extract a standardized image from a reference face video, extract the temporal deformation field of each video frame relative to the standardized image from the reference face video (the temporal deformation field includes facial motion information), perform stylization processing on the standardized image to obtain a stylized image, and generate a face animation video based on the stylized image and the temporal deformation field. Since the standardized image is stylized after extraction, the temporal deformation field can be applied to the entire video without needing to stylize each video frame separately, thus improving the generation efficiency of face animation videos and avoiding discontinuity and awkwardness in face animation videos, thereby improving the overall smoothness of the video. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a video processing method provided in an embodiment of this application;

[0020] Figure 2 This is a flowchart of a video processing method provided in an embodiment of this application;

[0021] Figure 3This is a flowchart of a video processing method provided in an embodiment of this application;

[0022] Figure 4 This is an overall technical roadmap of the video processing method provided in the embodiments of this application;

[0023] Figure 5 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Figure 1 This is a flowchart of a video processing method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes steps 110 to 140.

[0027] Step 110: Extract a standardized image from the reference face video.

[0028] In one exemplary embodiment, the reference face video is the original video to be processed, specifically the original face video requiring style transfer. A canonical image is a standardized image representation that aggregates static content (such as texture, shape, and color) throughout the video to provide a unified base image for each frame, enabling time-consistent processing and deformation propagation.

[0029] In an optional embodiment, the various video frames of the reference face video are compared, and static content that does not change over time (such as background, basic shape of the subject, etc.) is extracted from each video frame. This static content may include texture, shape, and color information from the reference face video. All extracted static content is then rendered to obtain a canonical image. For example, a canonical image from the reference face video can be extracted by referring to a CoDeF canonical content field. Through this canonical content field, each frame of the reference face video can be considered as being "rendered" from this static standard field.

[0030] Step 120: Extract the temporal deformation field of each video frame relative to the canonical image from the reference face video, the temporal deformation field including facial motion information.

[0031] In one exemplary embodiment, the facial motion information may include facial expression motion information and head posture motion information, etc.

[0032] In one exemplary embodiment, the temporal deformation field records the transformation process from a canonical image to each video frame in the reference face video. For example, implicit keypoints in LivePortrait can be used to capture facial motion information (such as facial expressions and head poses) in each video frame of the reference face video. LivePortrait is an open-source portrait animation generation framework focused on efficiently and controllably transferring the expressions and poses driving videos to static or dynamic portraits, creating expressive videos. LivePortrait is implemented through an implicit keypoint framework, utilizing large-scale, high-quality training data and hybrid training strategies to improve the model's generalization ability and motion control accuracy.

[0033] In one exemplary embodiment, implicit keypoints are a dynamic representation extracted through deep learning to capture key motion features in facial animation. These keypoints do not require explicit annotation; instead, the model adaptively learns and represents changes in facial expression and pose. Unlike explicit keypoints, implicit keypoints can effectively represent motion in a more compact form. This approach ensures a balance between computational efficiency and realism when generating animations.

[0034] In one exemplary embodiment, a temporal deformation field is used to record the deformation of each video frame in a reference face video relative to a canonical image, including information such as the object's motion, pose changes, and non-rigid body deformation. This temporal deformation field can be encoded using a three-dimensional hash table to store the deformation of each video frame relative to the canonical image and to describe these changes in a smooth manner.

[0035] Step 130: Process the standardized image according to a specified style to obtain a stylized image.

[0036] In one exemplary embodiment, a stylized image can be obtained by stylizing a standard image with a specified style using a preset stylization algorithm (e.g., ControlNet). This stylized image is the stylized standard image. The specified style can be obtained together with a reference face video. For example, the specified style could be a specified uniform, anime style, etc.

[0037] Step 140: Generate a facial animation video based on the stylized image and the temporal deformation field.

[0038] In an exemplary embodiment, a stylized image is propagated throughout the video based on a temporal deformation field to ultimately generate a stylized face animation video. Specifically, the temporal deformation field is applied to the stylized image according to the temporal order to deform the stylized image, resulting in each stylized video frame. The stylized video frames are then assembled into a face animation video according to the temporal order.

[0039] The video processing method provided in this application extracts a standardized image from a reference face video, extracts the temporal deformation field of each video frame relative to the standardized image from the reference face video (the temporal deformation field includes facial motion information), performs stylization processing on the standardized image to obtain a stylized image, and generates a face animation video based on the stylized image and the temporal deformation field. Since the standardized image is stylized after extraction, the temporal deformation field can be applied to the entire video without needing to stylize each video frame separately, thus improving the generation efficiency of the face animation video and avoiding discontinuity and awkwardness in the face animation video, thereby improving the overall smoothness of the video.

[0040] Based on the above technical solution, the step of extracting the temporal deformation field of each video frame relative to the canonical image from the reference face video includes: extracting image features from each video frame of the reference face video; identifying implicit keypoints representing facial motion from the image features; and determining the temporal deformation field based on the implicit keypoints of each video frame and the canonical image.

[0041] In one exemplary embodiment, image features can be extracted from video frames of a reference face video using a convolutional neural network. These image features can be a three-dimensional appearance feature volume, representing information such as texture, color, and contour in the image. After feature extraction, implicit keypoints representing motion are identified from the image features of each video frame of the reference face video. These implicit keypoints summarize the global motion of the face (such as expression changes, head rotation, etc.).

[0042] In an exemplary embodiment, image features extracted from video frames may include canonical keypoints, head pose features, facial expression features, etc., and then implicit keypoints are determined based on the canonical keypoints, head resource features, and facial expression features. Implicit keypoints are learned and calculated using a model of the temporal deformation field. The mathematical formula for calculating implicit keypoints is as follows:

[0043] X = S·(X) CR+δ)+t

[0044] Where X represents the implicit keypoint of the video frame; X C R represents the canonical keypoints of the video frame, which are learned by the model from the video frame; δ represents the head pose rotation matrix of the video frame; t represents the facial expression change parameter, used to capture subtle changes in the face; and S represents the translation parameter, which represents the translation of the implicit keypoints.

[0045] After extracting the implicit keypoints of each video frame, the temporal deformation field can be determined based on the implicit keypoints and the canonical image of each video frame.

[0046] By extracting image features from each frame of a reference face video, implicit keypoints are identified from these features. Then, a temporal deformation field is determined based on the implicit keypoints and the canonical image. This extracted temporal deformation field can maintain good generation results even when dealing with large-scale head movements, ensuring that facial features are not lost during style transfer. At the same time, it maintains high consistency between video frames, avoids facial feature drift, improves the consistency between frames in the video and the preservation of identity features, and significantly enhances the overall visual quality and naturalness of the video.

[0047] Figure 2 This is a flowchart of a video processing method provided in an embodiment of this application. Based on the above embodiments, this embodiment determines the time deformation field based on implicit keypoints and canonical images in the video frames. Figure 2 As shown, the method includes steps 210 to 270.

[0048] Step 210: Extract a standardized image from the reference face video.

[0049] Step 220: Extract image features from each video frame of the reference face video.

[0050] Step 230: Identify implicit keypoints representing facial movement from the image features.

[0051] Step 240: Determine the motion vector of the implicit keypoint in each video frame relative to the implicit keypoint in the canonical image to obtain sparse motion cue information.

[0052] In an exemplary embodiment, after identifying implicit keypoints representing facial motion from image features, the motion vectors of the implicit keypoints in each video frame relative to the corresponding implicit keypoints in the normalized image are calculated. The motion vectors corresponding to each video frame are arranged in chronological order to form motion vectors on the time axis. These motion vectors of the implicit keypoints serve as sparse motion cue information.

[0053] Step 250: Generate the time deformation field based on the sparse motion cue information.

[0054] In one exemplary embodiment, the temporal deformation field can be a two-dimensional vector field that characterizes the displacement of each pixel in the video frame relative to the corresponding pixel in the canonical image.

[0055] In one exemplary embodiment, based on sparse motion cue information, the motion information therein is progressively expanded so that the sparse motion cue information can cover the entire image area, thereby obtaining the temporal deformation field of the entire video.

[0056] Step 260: Process the standardized image according to a specified style to obtain a stylized image.

[0057] Step 270: Generate a facial animation video based on the stylized image and the temporal deformation field.

[0058] The video processing method provided in this embodiment uses implicit keypoints representing facial motion in each video frame as sparse motion cue information when generating the temporal deformation field, and generates the temporal deformation field based on the sparse motion cue information. The facial animation video generated based on the temporal deformation field and the stylized image can avoid facial feature drift, maintain good generation effect when processing large-scale head movements, ensure that facial features are not lost during style transfer, and maintain high consistency between video frames, significantly improving the overall visual quality and naturalness of the video.

[0059] Figure 3 This is a flowchart of a video processing method provided in an embodiment of this application. Based on the above embodiments, this embodiment can generate controllable facial animation videos based on manually controlled trajectories. Figure 3 As shown, the method includes steps 310 to 370.

[0060] Step 310: Obtain the manual control trajectory of a specified facial organ in the reference face video.

[0061] In one exemplary embodiment, when acquiring a reference face video and a specified style, manual control trajectories for specified facial features in the reference face video can be obtained. These specified facial features may include, for example, the nose, mouth, and eyes. The manual control trajectory is a control trajectory drawn by the user on the screen.

[0062] Step 320: Extract a standardized image from the reference face video.

[0063] Step 330: Extract image features from each video frame of the reference face video.

[0064] Step 340: Identify implicit keypoints representing facial motion from the image features.

[0065] Step 350: Determine the temporal deformation field based on the implicit keypoints of each video frame, the manual control trajectory, and the canonical image.

[0066] In an exemplary embodiment, the motion vector of the manually controlled trajectory on the time axis can be determined, and the motion vector of the implicit keypoints in each video frame relative to the canonical image on the time axis can be determined. Combining these two motion vectors, a temporal deformation field is obtained.

[0067] In an optional embodiment of this application, determining the temporal deformation field based on the implicit keypoints of each video frame, the manual control trajectory, and the canonical image includes: discretizing the manual control trajectory to obtain a first motion vector of each discrete point on the time axis; determining a second motion vector of the implicit keypoints in each video frame relative to the implicit keypoints in the canonical image; superimposing the first motion vector onto the implicit keypoints representing the specified facial organs in the second motion vector to obtain sparse motion cue information; and generating the temporal deformation field based on the sparse motion cue information.

[0068] In an exemplary embodiment, the manual control trajectory is interpolated into multiple discrete points, and the first motion vector of these discrete points on the time axis is calculated. The motion vector of the implicit keypoint in each video frame relative to the corresponding implicit keypoint in the normalized image is calculated, and the motion vectors corresponding to the implicit keypoints in each video frame are arranged in chronological order to form a second motion vector on the time axis. The first motion vector is superimposed on the implicit keypoints representing the specified facial organs in the second motion vector, and the superimposed motion vector is used as sparse motion cue information. Based on the sparse motion cue information, the motion information is gradually expanded so that the sparse motion cue information can cover the entire image area to obtain the temporal deformation field of the entire video.

[0069] By superimposing the first motion vector, which is a discretized manual control trajectory, onto the second motion vector determined based on implicit key points, the resulting temporal deformation field can realize the generation of controllable image animation based on the manual control trajectory.

[0070] In one exemplary embodiment, controllable image animation refers to a technique for dynamically manipulating static images through user input, aiming to generate a continuous sequence of frames with specific motion or facial expression changes. Its core lies in combining deep learning models to derive coherent animation effects from a small number of keyframes or input conditions. Through precise control of shape, posture, or expression, highly realistic dynamic visual effects are achieved while maintaining the structural and detail consistency of the input image.

[0071] Step 360: Process the standardized image according to a specified style to obtain a stylized image.

[0072] Step 370: Generate a facial animation video based on the stylized image and the temporal deformation field.

[0073] The video processing method provided in this embodiment obtains the manual control trajectory of specified facial organs in a reference face video, and then determines the temporal deformation field based on the implicit keypoints, manual control trajectory, and normalized image of each video frame when determining the temporal deformation field. This ensures that the determined temporal deformation field includes manual control information, thereby realizing the generation of controllable facial animation videos. Based on the implicit keypoints and manual control trajectory, precise control can be exercised over facial details (such as micro-expressions and local movements), which helps to generate high-quality facial animations. It is particularly suitable for applications involving expression recognition and emotional expression.

[0074] Based on the above technical solution, the temporal deformation field of each video frame is a two-dimensional vector field, which represents the displacement of each pixel in the video frame relative to the corresponding pixel in the standard image.

[0075] The step of generating the time deformation field based on the sparse motion cue information includes: generating the time deformation field through a sparse-to-dense motion generation network based on the sparse motion cue information.

[0076] In one exemplary embodiment, a sparse-to-dense motion generation (S2D) network is a model that generates a dense motion field from a small number of keypoints. This model derives global motion information using the sparse keypoint locations and generates a high-resolution dense motion representation through optical flow or deformation fields. This method significantly reduces computational complexity while preserving motion details and is commonly used in tasks such as video generation and image animation, effectively balancing generation accuracy and computational efficiency. Optical flow is a technique used to estimate the velocity and direction of pixel motion in an image sequence. It aims to derive motion information in a scene by analyzing brightness changes between adjacent frames. Its core assumption is that the brightness of objects remains constant over a short period, and it describes the dynamic changes of the scene by solving for the motion vector field of pixels.

[0077] In one exemplary embodiment, the role of the sparse-to-dense motion generation network is to transform sparse motion information (such as manual control trajectories or facial keypoint sequences) into a dense motion stream to drive motion generation across the entire image. The sparse-to-dense motion generation network processes this sparse motion cue information through multiple convolutional layers, progressively expanding the motion information so that the sparse motion cue information covers the entire image region. The final output dense motion field is a two-dimensional vector field representing the displacement of each pixel in the video frame relative to the canonical image. This displacement information is used to guide feature warping of the video frame, ensuring that the motion of each pixel is smooth and consistent.

[0078] In an optional embodiment, the training objective of the sparse-to-dense motion generation network during training includes minimizing the square of the difference between a first difference and the motion field change values ​​at two adjacent time points, where the first difference is the difference between the dense motion fields at two adjacent time points. This training objective aims to make the difference between the position vectors of the dense motion fields at two adjacent time points approximately equal to the motion field change values ​​at those two adjacent time points. The output data of the sparse-to-dense motion generation network, i.e., the temporal deformation field, includes the dense motion fields at each time point and the motion field change values ​​at two adjacent time points.

[0079] The sparse-guided temporal consistency loss function can be used to control the sparse-to-dense motion generation network to achieve the above training objectives during training. The sparse-guided temporal consistency loss function is expressed as follows:

[0080] L SGTC =λ1∑||F pred (x,y,t)-F pred (x t+1 ,y t+1 ,t+1)-ΔF pred (t→t+1)|| 2 ·M hint

[0081] Among them, F pred (x,y,t) represents the position vector at position (x,y) of the dense motion field generated by the sparse-to-dense motion generation network at time t. F pred (x t+1 ,y t+1 (t+1) represents the position vector of the dense motion field at time t+1 predicted by the sparse-to-dense motion generation network, ΔF pred (t→t+1) represents the change in the motion field from time t to time t+1, M hinThe mask represents the sparse cue information, and λ1 is a hyperparameter. The mask of the sparse cue information characterizes the positional information provided by the sparse cue in the motion field (temporal deformation field), and λ1 is used to control the impact of the sparse-guided temporal consistency loss function on the overall loss function.

[0082] The training of the sparse-guided temporal consistency loss is used to guide the generation of the temporal deformation field in the dense motion generation network. The generated temporal deformation field is a dense motion field, which effectively provides optimization for the temporal deformation field. Moreover, it ensures that the difference between the motion field position vectors of two adjacent time steps in the dense motion field is close to the motion field variation value of the two adjacent time steps, thus ensuring the smoothness and consistency of the deformation field on the time axis.

[0083] Face animation videos generated based on optimized temporal deformation fields are a type of spatiotemporally consistent video. Spatiotemporally consistent video refers to dynamic videos that maintain visual and structural continuity between frames in the temporal dimension. The purpose of spatiotemporally consistent video is to avoid visual artifacts such as flickering, jumps, or unstable textures between frames during video generation. By combining temporal information and deep learning models, spatiotemporally consistent algorithms ensure smooth transitions and coherence in video when there are changes in motion, lighting, color, and texture.

[0084] Figure 4 This is an overall technical roadmap of the video processing method provided in the embodiments of this application, such as... Figure 4 As shown, feature extraction is performed on the reference face video to obtain the implicit keypoints of each video frame. Based on the implicit keypoints of each video frame, a sparse-to-dense motion generation (S2D) network is used to generate a temporal deformation field of each video frame relative to the normalized image. This temporal deformation field can be a controllable motion field obtained by combining manually controlled trajectories and implicit keypoints. The normalized image is processed with a specified style to obtain a stylized image. The temporal deformation field is applied to the stylized image to obtain the face animation video.

[0085] This application not only optimizes generation time and computational efficiency, significantly shortening the generation time, but also maintains the integrity of facial features and high consistency between video frames when handling large-scale motion and style transfer. Furthermore, this application excels in controllable facial animation generation, accurately responding to complex motion control requirements, enabling fine-tuning of faces, significantly improving the naturalness and visual quality of generated videos, and meeting the needs of various application scenarios.

[0086] Figure 5 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application, as shown below. Figure 5 As shown, the device includes:

[0087] The standardized image extraction module 410 is used to extract standardized images from the reference face video;

[0088] The temporal deformation field extraction module 420 is used to extract the temporal deformation field of each video frame relative to the canonical image from the reference face video, wherein the temporal deformation field includes facial motion information.

[0089] The stylization processing module 430 is used to process the standard image in a specified style to obtain a stylized image.

[0090] The video generation module 440 is used to generate a facial animation video based on the stylized image and the time deformation field.

[0091] Optionally, the time deformation field extraction module includes:

[0092] The feature extraction unit is used to extract image features from each video frame of the reference face video;

[0093] An implicit key point recognition unit is used to identify implicit key points representing facial motion from the image features.

[0094] The temporal deformation field determination unit is used to determine the temporal deformation field based on the implicit keypoints of each video frame and the canonical image.

[0095] The time-deformation field determination unit includes:

[0096] The sparse cue determination subunit is used to determine the motion vector of the implicit keypoint in each video frame relative to the implicit keypoint in the canonical image, thereby obtaining sparse motion cue information;

[0097] The time-deformation field generation subunit is used to generate the time-deformation field based on the sparse motion cue information.

[0098] Optionally, the device further includes:

[0099] The control trajectory acquisition module is used to acquire the manual control trajectory of a specified facial organ in the reference face video;

[0100] The time-deformation field determination unit is specifically used for:

[0101] The temporal deformation field is determined based on the implicit keypoints of each video frame, the manual control trajectory, and the canonical image.

[0102] Optionally, the time-deformation field determination unit includes:

[0103] The first motion vector determination subunit is used to discretize the manual control trajectory to obtain the first motion vector of each discrete point on the time axis.

[0104] The second motion vector determination subunit is used to determine the second motion vector of the implicit keypoint in each video frame relative to the implicit keypoint in the canonical image.

[0105] A sparse cue determination subunit is used to superimpose the first motion vector onto the implicit key points representing the specified facial organs in the second motion vector to obtain sparse motion cue information;

[0106] The time-deformation field generation subunit is used to generate the time-deformation field based on the sparse motion cue information.

[0107] Optionally, the temporal deformation field of each video frame is a two-dimensional vector field, representing the displacement of each pixel in the video frame relative to the corresponding pixel in the canonical image;

[0108] The time-deformation field generation subunit is specifically used for:

[0109] The time deformation field is generated by a sparse-to-dense motion generation network based on the sparse motion cue information.

[0110] Optionally, the loss function of the sparse-to-dense motion generation network is expressed as follows:

[0111] L sGTc =λ1∑||F pred (x,y,t)-F pred (t→t+1),y+F pred (t→t+1),t+1|| 2 ·M hint

[0112] Among them, F pred (x,y,t) represents the position vector at position (x,y) of the dense motion field generated by the sparse-to-dense motion generation network at time t. F pred (t→t+1) represents the predicted motion field value from time t to time t+1, M hint The mask representing the sparse cue information, where λ1 is a hyperparameter.

[0113] The video processing apparatus provided in this application embodiment is used to implement the various steps of the video processing method described in this application embodiment. The specific implementation methods of each module of the apparatus are described in the corresponding steps, and will not be repeated here.

[0114] The video processing apparatus provided in this application extracts a standardized image from a reference face video, extracts a temporal deformation field of each video frame relative to the standardized image from the reference face video (the temporal deformation field includes facial motion information), performs stylization processing on the standardized image to obtain a stylized image, and generates a face animation video based on the stylized image and the temporal deformation field. Since the standardized image is stylized after extraction, the temporal deformation field can then be applied to the entire video, eliminating the need to stylize each video frame separately. This improves the generation efficiency of face animation videos and avoids discontinuity and awkwardness in face animation videos, thus improving the overall smoothness of the video.

[0115] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the electronic device 500 may include one or more processors 510 and one or more memories 520 connected to the processors 510. The electronic device 500 may also include an input interface 530 and an output interface 540 for communicating with another device or system. Program code executed by the processor 510 may be stored in the memory 520.

[0116] The processor 510 in the electronic device 500 calls the program code stored in the memory 520 to execute the video processing method in the above embodiment.

[0117] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the video processing method described in this application.

[0118] This application also provides a computer program product that, when executed by a processor, implements the steps of the video processing method described in this application.

[0119] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are fundamentally similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0120] The foregoing has provided a detailed description of a video processing method, apparatus, electronic device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

Claims

1. A video processing method, characterized in that, include: Extract canonical images from reference face videos; Extract the temporal deformation field of each video frame relative to the canonical image from the reference face video, the temporal deformation field including facial motion information; The standardized image is processed in a specified style to obtain a stylized image; An animated video of a face is generated based on the stylized image and the temporal deformation field.

2. The method according to claim 1, characterized in that, Extracting the temporal deformation field of each video frame relative to the canonical image from the reference face video includes: Extract image features from each video frame of the reference face video; Identify implicit keypoints representing facial movement from the image features; The temporal deformation field is determined based on the implicit keypoints of each video frame and the canonical image.

3. The method according to claim 2, characterized in that, Determining the temporal deformation field based on the implicit keypoints of each video frame and the canonical image includes: Determine the motion vector of the implicit keypoint in each video frame relative to the implicit keypoint in the canonical image to obtain sparse motion cue information; The time deformation field is generated based on the sparse motion cue information.

4. The method according to claim 2, characterized in that, Before determining the temporal deformation field based on the implicit keypoints of each video frame and the canonical image, the method further includes: Obtain the manual control trajectory of a specified facial organ in the reference face video; Determining the temporal deformation field based on the implicit keypoints of each video frame and the canonical image includes: The temporal deformation field is determined based on the implicit keypoints of each video frame, the manual control trajectory, and the canonical image.

5. The method according to claim 4, characterized in that, Determining the temporal deformation field based on the implicit keypoints of each video frame, the manual control trajectory, and the canonical image includes: The manual control trajectory is discretized to obtain the first motion vector of each discrete point on the time axis; Determine the second motion vector of the implicit keypoints in each video frame relative to the implicit keypoints in the canonical image; The first motion vector is superimposed onto the implicit key points representing the specified facial organs in the second motion vector to obtain sparse motion cue information; The time deformation field is generated based on the sparse motion cue information.

6. The method according to claim 3 or 5, characterized in that, The temporal deformation field of each video frame is a two-dimensional vector field, representing the displacement of each pixel in the video frame relative to the corresponding pixel in the canonical image; The step of generating the time-deformation field based on the sparse motion cue information includes: The time deformation field is generated by a sparse-to-dense motion generation network based on the sparse motion cue information.

7. The method according to claim 6, characterized in that, The training objective of the sparse-to-dense motion generation network during the training process includes: minimizing the square of the difference between the first difference and the motion field change value between two adjacent time points, where the first difference is the difference between the dense motion fields between two adjacent time points.

8. A video processing device, characterized in that, include: The standardized image extraction module is used to extract standardized images from reference face videos; A temporal deformation field extraction module is used to extract the temporal deformation field of each video frame relative to the canonical image from the reference face video, wherein the temporal deformation field includes facial motion information; The stylization processing module is used to process the standard image in a specified style to obtain a stylized image. The video generation module is used to generate a facial animation video based on the stylized image and the time deformation field.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the video processing method according to any one of claims 1 to 7.