Method, apparatus, storage medium and computer device for driving human expressions
By identifying and aligning the key points of faces in character videos, and using the target expression-driven model of the three-dimensional convolutional neural network for expression-driven, the problem of high limitations in character expression-driven methods in the existing technology is solved, and a high-quality and natural video expression-driven effect is achieved.
Patent Information
- Application Number
- CN202411784360.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-12-06
AI Technical Summary
In the prior art, the character expression driving method has high limitations and cannot meet the video expression driving needs of film and television level.
A character expression driving method is adopted. By obtaining the video frame image of the character video, the target face detection model is used to identify the key points of the face and align the face. Then, the target expression driving model composed of a three-dimensional convolutional neural network is used to drive the standard face image with expression.
It realizes high-quality and natural video expression drivers, which can accurately track facial expressions at various angles in the video frame image and edit and drive them to ensure smooth and natural expression changes.
Smart Images

Figure CN119478161B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, device, storage medium, and computer device for driving human expressions. Background Art
[0002] With the wide application of AIGC (Artificial Intelligence Generated Content) technology in the video field, especially in video scenarios with high production costs such as movies, TV dramas, micro short dramas, and animations, the demand for driving and editing human expressions in videos is increasing day by day. This demand aims to achieve the editing of human expressions in a more flexible and rapid manner, thereby effectively reducing shooting costs, shortening the production cycle, and improving the efficiency of the team in video content creation.
[0003] Existing methods for driving human expressions highly rely on the training of a large amount of data, have high limitations on the application scope, and the application effects in high-quality scenarios are not ideal; while the method for driving human expressions of 3D models performs well in 3D modeling scenarios, but due to its technical characteristics, it is not suitable for directly applying to the driving of human videos. Therefore, the current methods for driving human expressions have high limitations and cannot meet the requirements for driving video expressions at the film and television level. Summary of the Invention
[0004] The purpose of this application aims to at least solve one of the above technical defects, especially the technical defect that the limitations of existing methods for driving human expressions are relatively high and cannot meet the requirements for driving video expressions at the film and television level.
[0005] This application provides a method for driving human expressions, and the method includes:
[0006] Obtain a human video containing a human face, as well as the expression driving target of the human video, and parse to obtain consecutive video frame images in the human video;
[0007] Use a target face detection model to identify the face key points in the video frame images, and perform face alignment on the video frame images based on the face key points to obtain standard face images;
[0008] Determine a target expression driving model; the target expression driving model is composed of a three-dimensional convolutional neural network and is used to map two-dimensional images into a three-dimensional latent space for feature encoding and expression driving;
[0009] Input the standard face image and the expression driving target into the target expression driving model, and use the target expression driving model to drive the expression of the standard face image based on the expression driving target to output an expression-driven video.
[0010] Optionally, the step of identifying the face key points in the video frame image by using the target face detection model includes:
[0011] Determine a target face detection model, where the target face detection model includes a face recognition network and a key point detection network;
[0012] Use the face recognition network to perform face recognition on the video frame image to obtain a face position box in the video frame image;
[0013] Based on the face position box in the key point detection network, perform key point detection on the original face image to obtain face key points.
[0014] Optionally, the training process of the target face detection model includes:
[0015] Obtain a sample face image labeled with a true face position box, where the true face position box is labeled with true face key points;
[0016] Input the sample face image into a preset initial face detection model to obtain a predicted face position box output by the initial face detection model for the sample face image, as well as the predicted confidence of the predicted face position box and predicted face key points;
[0017] Take the predicted face position box and the predicted face key points approaching the true face position box and the true face key points respectively as the goals, and train the initial face detection model;
[0018] When the initial face detection model meets the preset training conditions, perform lightweight processing on the trained initial face detection model to obtain a target face detection model.
[0019] Optionally, the step of taking the predicted face position box and the predicted face key points approaching the true face position box and the true face key points respectively as the goals and training the initial face detection model includes:
[0020] Determine the position loss value of the predicted face position box based on the true face position box, and calculate the confidence loss value of the predicted confidence according to the position loss value;
[0021] Determine the key point loss value of the predicted face key points based on the true face key points;
[0022] Update the parameters in the initial face detection model according to the position loss value, the confidence loss value, and the key point loss value.
[0023] Optionally, the face alignment of the video frame image based on the face key points to obtain a standard face image includes:
[0024] Locate the face region in the video frame image according to the face key points, and use the affine transformation technology to align the located face region to a preset standard position to form a standard face image.
[0025] Optionally, the target expression driving model includes a latent space mapping module, a key point generation module, a spatial local deformation module, and a feature decoding module;
[0026] The expression driving of the standard face image based on the expression driving target by using the target expression driving model includes:
[0027] Use the latent space mapping module to extract the low-order high-frequency features of the standard face image, and perform latent space encoding and feature decoupling on the low-order high-frequency features to generate a three-dimensional image corresponding to the standard face image in the latent space;
[0028] Determine multiple deformation regions in the three-dimensional image and the deformation features corresponding to each deformation region through the key point generation module;
[0029] Use the spatial local deformation module to determine the deformation region and deformation direction corresponding to the expression driving target, and perform feature editing on the deformation features of the deformation region based on the deformation direction to generate an expression-driven image;
[0030] Input the expression-driven image into the feature decoding module, so that the feature decoding module uses a generator composed of a SPADE network to perform upsampling decoding on the expression-driven image to map the expression-driven image to a two-dimensional space and generate an expression-driven video.
[0031] Optionally, each deformation region is composed of multiple deformation key points;
[0032] The determination of multiple deformation regions in the three-dimensional image and the deformation features corresponding to each deformation region through the key point generation module includes:
[0033] Determine multiple deformation regions in the three-dimensional image and the three-dimensional coordinates of each deformation key point in each deformation region through the key point generation module;
[0034] For each deformation region, use the Gaussian distribution technology in the key point generation module to calculate the influence of the three-dimensional coordinates of each deformation key point in the deformation region to obtain the feature influence value of each deformation key point in the three-dimensional image, and perform weighted summation on the respective feature influence values to form the deformation feature of the deformation region.
[0035] The present application also provides a device for driving human expressions, including:
[0036] A data acquisition module, configured to acquire a human video containing a human face, as well as an expression driving target of the human video, and parse to obtain a video frame image of the human video;
[0037] A face recognition module, configured to use a target face detection model to recognize face key points in the video frame image, and perform face alignment on the video frame image based on the face key points to obtain a standard face image;
[0038] A model determination module, configured to determine a target expression driving model; the target expression driving model is composed of a three-dimensional convolutional neural network, and is used to map a two-dimensional image into a three-dimensional latent space for feature encoding and expression driving;
[0039] An expression driving module, configured to input the standard face image and the expression driving target into the target expression driving model, and use the target expression driving model to perform expression driving on the standard face image based on the expression driving target, so as to output an expression-driven video.
[0040] The present application also provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for driving human expressions according to any one of the above embodiments.
[0041] The present application also provides a computer device, including: one or more processors, and a memory;
[0042] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the method for driving human expressions according to any one of the above embodiments are executed.
[0043] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0044] The method, device, storage medium, and computer equipment for driving human expressions provided in this application can, when driving the expressions of a person in a human video, first determine the expression driving target of the human video and parse to obtain consecutive video frame images in the human video, so that expression driving can be realized on the video frame images to ensure the naturalness of the expression driving effect; then, a target face detection model can be used to identify the face key points in the video frame images, and the video frame images can be face-aligned based on the face key points to obtain standard face images, so as to perform expression driving on the standard face images based on the expression driving target through the target expression driving model and output them as an expression-driven video, thereby improving the quality of the expression driving effect. Since the target expression driving model is composed of a three-dimensional convolutional neural network and is used to map two-dimensional images into a three-dimensional latent space for feature encoding and expression driving, the target expression driving model can accurately track faces at various angles in the video frame images and perform expression editing and driving, realizing natural expression connection between frames of the video frame images without affecting the quality of the human video. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1 It is a schematic flowchart of a method for driving human expressions provided in an embodiment of the present application;
[0047] Figure 2 It is a schematic flowchart of a face key point detection process provided in an embodiment of the present application;
[0048] Figure 3 It is a schematic flowchart of a process for updating the parameters of an initial face detection model provided in an embodiment of the present application;
[0049] Figure 4 It is a schematic flowchart of a process for applying a target expression driving model provided in an embodiment of the present application;
[0050] Figure 5 It is a schematic structural diagram of a device for driving human expressions provided in an embodiment of the present application;
[0051] Figure 6 It is a schematic internal structure diagram of a computer equipment provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0053] Existing methods for driving human facial expressions highly rely on the training of a large amount of data, have high limitations on the application scope, and the application effects in high-quality scenarios are not ideal. Although the method for driving human facial expressions with 3D models performs well in 3D modeling scenarios, due to its technical characteristics, it is not suitable for directly driving human videos. Therefore, the current methods for driving human facial expressions have high limitations and cannot meet the requirements for driving video facial expressions at the film and television level.
[0054] Based on this, the present application proposes the following technical solutions. For details, please refer to the following:
[0055] In one embodiment, as Figure 1 shown, Figure 1 is a schematic flowchart of a method for driving human facial expressions provided by an embodiment of the present application. The present application provides a method for driving human facial expressions. For details, please refer to the following:
[0056] S110: Obtain a human video containing a human face and an expression driving target of the human video, and parse to obtain consecutive video frame images in the human video.
[0057] In this step, when a user needs to drive the facial expressions of a film and television character, the human video where the film and television character is located can be uploaded to a computer device, and the expression driving target of the human video can be set in the computer device. Then, the computer device can parse the human video to obtain consecutive video frame images, and further, facial expression driving can be implemented on the video frame images to ensure the naturalness of the facial expression driving effect.
[0058] It can be understood that the expression driving target refers to the facial expression changes or action effects that the user hopes to achieve during the process of driving human facial expressions, such as specific expressions like smiling, being surprised, or frowning. Here, the expression driving target can not only include the action changes of facial features but also the overall expression dynamic effects, such as the transition from a neutral expression to a smile, so as to achieve a high-quality facial expression driving effect for film and television characters.
[0059] In addition, the video frame images of the present application refer to the sequence of frame-by-frame images after parsing a person video. To elaborate, the essence of a person video is composed of a series of rapidly switched static images (frames). When the frame rate is fast enough, a visual effect of continuous movement will be generated. Therefore, when performing expression driving on a person video in the present application, the facial features of each frame image can be parsed, and then the expression driving target can be applied to these frames to ensure that the facial expressions of the person in the person video can change as expected, thereby avoiding abrupt jumps or incoherent expression effects, and thus maintaining the smoothness and naturalness of the person's expression changes.
[0060] Specifically, the computer device can use a video processing tool or library to parse and obtain the continuous video frame images in the person video, and then apply the expression driving target set by the user to the video frame images. By adjusting the expression features of the person's face, such as the dynamic changes of the eyes, eyebrows, and lips, the computer can accurately drive the expression changes of the person in each frame, ensuring that the transitions of these expressions between frames are smooth and natural. Therefore, the present application can achieve specific expression effects while maintaining the original natural state of the person video.
[0061] S120: Use the target face detection model to identify the face key points in the video frame image, and perform face alignment on the video frame image based on the face key points to obtain a standard face image.
[0062] In this step, after obtaining the video frame images of the person video through step S110, the computer device can obtain the pre-set target face detection model, and then use the target face detection model to identify the face key points in the video frame image, so that face alignment can be performed on the video frame image based on the face key points to obtain a standard face image. Through the standard face image, the present application can further improve the naturalness of expression driving.
[0063] Among them, the target face detection model refers to a model that performs face detection on the input video frame image and identifies the face key points; the face key points here refer to several important points distributed in specific parts of the face, such as the eyes, nose, mouth, etc. These points are used to describe the geometric structure and morphological features of the face. Through the face key points, the computer device can accurately identify the pose, shape, and expression of the face, thereby improving the natural effect of expression driving.
[0064] Specifically, after identifying the facial key points in the video frame image, the computer device can align the face in the image based on these facial key points, thereby standardizing the face in the video frame image, reducing the image differences caused by shooting conditions or lighting conditions, and maintaining a unified scale, ensuring that subtle changes in facial expressions can be better captured during the subsequent expression driving process, making the changes in facial expressions more accurate and coherent.
[0065] S130: Determine the target expression driving model; the target expression driving model is composed of a three-dimensional convolutional neural network and is used to map a two-dimensional image into a three-dimensional latent space for feature encoding and expression driving.
[0066] In this step, after obtaining the standard face image through step S120, the computer device can also determine the target expression driving model to drive the facial expression in the standard face image through this target expression driving model.
[0067] It can be understood that the target expression driving model of this application is composed of a three-dimensional convolutional neural network and is used to map a two-dimensional image into a three-dimensional latent space for feature encoding and expression driving. Among them, compared with the traditional two-dimensional convolutional network, the three-dimensional convolutional neural network can process more complex spatial and temporal features. Specifically, the three-dimensional convolutional neural network can not only extract the features of facial expressions in a single-frame two-dimensional image, but also combine the multi-frame information in the video to capture the changes in expressions in the time dimension. Therefore, the target expression driving model can map a two-dimensional image into a three-dimensional latent space, thereby performing a more abundant feature encoding on the facial expression. This feature encoding can not only describe the static expressions of the face, but also reflect the dynamic changes of expressions, and then achieve fine expression driving.
[0068] It should be noted that due to the powerful three-dimensional space perception ability of the target expression driving model, it can accurately track the changes in facial expressions of the face at different angles and poses in the video frame image, enabling the model to not only process frontal faces, but also perform high-precision expression capture and editing on faces in various complex poses such as side, top-down, and up-looking. Therefore, in this application, through the learning and reasoning of the target expression driving model, no matter at what angle the person appears in the person video, it can track the subtle changes in their facial expressions and drive according to the preset expression driving target.
[0069] In addition, the target expression-driven model is not limited to simple expression recognition, but also has the ability to perform in-depth editing and generation of facial expressions. Through feature encoding in the three-dimensional latent space, the model can achieve rich expression operations, such as exaggerating a certain expression, smoothing the expression transition, or applying the change of a specific expression to the sequence of entire video frame images, and this process will not affect the overall quality of the video. Therefore, when the model performs expression driving, it can ensure that the expression changes between frames are naturally connected, avoiding abrupt or discontinuous expression transitions, making the facial expressions of the characters in the character video more real and natural.
[0070] S140: Input the standard face image and the expression driving target into the target expression-driven model, and use the target expression-driven model to perform expression driving on the standard face image based on the expression driving target, so as to output an expression-driven video.
[0071] In this step, after determining the target expression-driven model through step S130, the computer device can input the standard face image and the expression driving target into the target expression-driven model, and use the target expression-driven model to perform expression driving on the standard face image based on the expression driving target, so as to output a high-quality and natural expression-driven video.
[0072] Specifically, after receiving the standard face image and the expression driving target, the target expression-driven model can perform frame-by-frame processing and editing on the input expression face image according to the expression driving target, and finally output and form an expression-driven video. Since the model has been trained with a large amount of face data and expression changes and has highly intelligent feature extraction and transformation capabilities, it can accurately capture and manipulate the subtle changes in facial expressions. Therefore, when the model performs frame-by-frame editing on the facial expressions of the standard face image in the time dimension, it can ensure that the naturalness of the human face is not damaged, and the expression changes are smooth and conform to the dynamic change law of real facial expressions.
[0073] In the above embodiments, when driving the expression of a person in a person video, the expression driving target of the person video can be determined first, and continuous video frame images in the person video can be parsed, so that expression driving can be realized on the video frame images, ensuring the naturalness of the expression driving effect; then, the face key points in the video frame images can be recognized by using the target face detection model, and the video frame images can be aligned based on the face key points to obtain standard face images, so that the target expression driving model can drive the standard face images based on the expression driving target and output them as expression-driven videos, thereby improving the quality of the expression driving effect. Since the target expression driving model is composed of a three-dimensional convolutional neural network and is used to map two-dimensional images into a three-dimensional latent space for feature encoding and expression driving, the target expression driving model can accurately track the faces at various angles in the video frame images and perform expression editing and driving, realizing the natural connection of expressions between frames of the video frame images without affecting the quality of the person video.
[0074] In one embodiment, as Figure 2 shown, Figure 2 is a schematic flowchart of a face key point detection process provided by an embodiment of the present application; Figure 2 In it, in step S120, recognizing the face key points in the video frame images by using the target face detection model may include:
[0075] S121: Determine the target face detection model, and the target face detection model includes a face recognition network and a key point detection network.
[0076] S122: Perform face recognition on the video frame images by using the face recognition network to obtain the face position boxes in the video frame images.
[0077] S123: Perform key point detection on the original face images based on the face position boxes in the key point detection network to obtain the face key points.
[0078] In this embodiment, when the computer device performs face key point recognition on the video frame images, it can first determine the target face detection model, and then use this model to recognize the face key points in the video frame images.
[0079] Specifically, since the target face detection model includes a face recognition network and a key point detection network, after the computer device inputs the video frame image into the target face detection model, the target face detection model can first deeply analyze each frame of the video frame image through the face recognition network to identify whether there is a face in the image. When the face recognition network detects a face, it generates a face position box, that is, frames the specific area where the face is located in that frame of the image, ensuring that the computer device can accurately locate the face part in the video frame for subsequent processing. Then, based on the already identified face position box, the key point detection network focuses on analyzing specific parts of the face, such as eyes, nose, mouth, eyebrows, etc., and identifies the geometric positions of the face key points to accurately depict the geometric structure and expression features of the human face.
[0080] Therefore, in the process of identifying face key points, this application can, through the close cooperation between the face recognition network and the key point detection network, first determine the overall position of the face and then perform key point detection on this basis; in this way, this application can significantly improve the detection efficiency and accuracy, avoid unnecessary calculations on the entire image, and further improve the overall efficiency of human expression driving.
[0081] In one embodiment, the training process of the target face detection model in step S120 may include:
[0082] S124: Obtain a sample face image labeled with a true face position box, and the true face position box is labeled with true face key points.
[0083] S125: Input the sample face image into a preset initial face detection model to obtain the predicted face position box output by the initial face detection model for the sample face image, the predicted confidence of the predicted face position box, and the predicted face key points.
[0084] S126: Take the predicted face position box and the predicted face key points approaching the true face position box and the true face key points respectively as the goal, and train the initial face detection model.
[0085] S127: When the initial face detection model meets the preset training conditions, lightweight the trained initial face detection model to obtain the target face detection model.
[0086] In this embodiment, when training the target face detection model, the computer device can first obtain sample face images annotated with real face position frames and real face key points, such as consecutive images in a person video in different scenarios. This image can be a preprocessed image, such as face alignment processing. Then, the processed sample face images are input into a preset initial face detection model for iterative training, and the trained model is lightweighted to obtain the target face detection model. Therefore, the size of the target face detection model can be within 3M, and the inference speed can be within 10ms.
[0087] Specifically, during the training process, the computer device obtains the preprocessed sample face images, which are annotated with corresponding sample labels, that is, real face position frames and real face key points within the real face position frames. Therefore, the training samples with sample labels are input into the preset initial face detection model for forward propagation to train the model, and a preset loss function is used to optimize the model parameters during the backpropagation process of the model. When the model meets certain training conditions or parameter convergence conditions, such as the number of iterations reaches the set value, it is regarded as the training is completed. At this time, the trained model can be used as the final target face detection model. In addition, this application can also store the trained target face detection model so that when performing face detection later, the pre-stored target face detection model can be directly called to perform face detection on video frame images.
[0088] Furthermore, the target face detection model in this application can select Mobilenetv3 (a lightweight deep learning network) as the basic model for improvement and training. When improving the network structure of Mobilenetv3, the parameters of channel, block, and linear can be adjusted to be smaller, thereby reducing the computational amount and operation time of the model during the entire training and inference processes, and improving the usage range of the model by reducing the requirement for computing power; while the output parameters of the network can be adjusted to 5, which are the X coordinate of the upper left corner of the face detection frame, the Y coordinate of the upper left corner of the face detection frame, the X coordinate of the lower right corner of the face detection frame, the Y coordinate of the lower right corner of the face detection frame, and the confidence of the face detection frame, so as to use the training process of the confidence of the face detection frame for supervision.
[0089] In one embodiment, as Figure 3 shown, Figure 3 is a schematic flowchart of the parameter update process of an initial face detection model provided by an embodiment of this application; Figure 3 In it, in step S126, the initial face detection model is trained with the goal that the predicted face position frame and the predicted face key points approach the real face position frame and the real face key points respectively, including:
[0090] S1261: Determine the position loss value of the predicted face position box based on the real face position box, and calculate the confidence loss value of the predicted confidence based on the position loss value.
[0091] S1262: Determine the key point loss value of the predicted face key points based on the real face key points.
[0092] S1263: Update the parameters in the initial face detection model according to the position loss value, the confidence loss value, and the key point loss value.
[0093] In this embodiment, in each iteration process, the computer device can determine the position loss value of the predicted face position box based on the real face position box, then calculate the confidence loss value of the predicted confidence based on the position loss value. At the same time, the computer device can also determine the key point loss value of the predicted face key points based on the real face key points. Finally, the computer device can update the parameters in the initial face detection model according to the position loss value, the confidence loss value, and the key point loss value, so as to complete the training of the model in the current round.
[0094] Specifically, when updating the parameters of the model, the computer device can use a position loss function to calculate the position loss value between the predicted face position box and the real face position box. This position loss value reflects the position box accuracy in the current iteration. Therefore, the parameters of the model can be further adjusted through this position loss value. After determining the position loss function, the computer device can also use a confidence loss function to calculate the position loss value and the predicted confidence, and obtain the confidence loss value of the predicted confidence. This confidence loss value can be used to measure the accuracy of the model when judging whether a certain area actually contains a face. It reflects the confidence degree of the model in its own prediction results. Therefore, by adjusting the parameters of the model through the confidence loss value, the reliability of the model prediction results can be improved.
[0095] At the same time, the computer device can also use a key point loss function to calculate the key point loss value between the predicted face key points and the real face key points. This key point loss value can be used to measure the accuracy of the model when predicting face details, especially the specific positions of facial features (such as eyes, nose, mouth, etc.). Therefore, the computer device can also further adjust the parameters of the model through this key point value.
[0096] Furthermore, after calculating the position loss value, the confidence loss value, and the key point loss value, the computer device can perform comprehensive calculations based on these loss values, and then update the parameters in the model through the backpropagation algorithm. Through backpropagation, the computer device can aim to reduce the loss value and gradually adjust each weight and parameter of the model, so that in the next iteration, the model can make more accurate predictions.
[0097] In one embodiment, the step of performing face alignment on the video frame image based on face key points in step S120 to obtain a standard face image may include:
[0098] S128: Locate the face region in the video frame image according to the face key points, and use the affine transformation technique to align the located face region to a preset standard position to form a standard face image.
[0099] In this embodiment, after the computer device identifies the face key points in the video frame image, it can first locate the face region in the video frame image according to the face key points, and then use the affine transformation technique to align the located face region to a preset standard position to form a standard face image, so that expression driving can be performed on the standard face image, thereby improving the face expression driving effect.
[0100] Among them, the affine transformation technique refers to a geometric transformation method that can perform linear transformation on an image while keeping the straight lines and proportional relationships in the image unchanged. Specifically, the affine transformation can perform operations such as translation, scaling, rotation, shearing, and flipping on the image. And during these transformation processes, the originally parallel straight lines in the image still remain parallel, and attributes such as angles and distances will change accordingly, but the linearity and relative proportional relationships remain unchanged, making the faces in all images have a consistent structure.
[0101] It can be understood that the face key points usually include the coordinates of facial features such as eyes, nose, and corners of the mouth. Locating the face region through the face key points can provide an accurate reference for subsequent affine transformation. Then, the computer device can use the affine transformation technique to align the located face region with the preset standard position to ensure the consistency between cross-frame images in the video frame image, thereby improving the accuracy and naturalness of subsequent expression driving.
[0102] In one embodiment, as Figure 4 shown, Figure 4 is a schematic flowchart of the application process of a target expression driving model provided by an embodiment of the present application; Figure 4 In, the target expression driving model in step S140 includes a latent space mapping module, a key point generation module, a spatial local deformation module, and a feature decoding module; among them, the step of using the target expression driving model to perform expression driving on the standard face image based on the expression driving target may include:
[0103] S141: Use the latent space mapping module to extract the low-order high-frequency features of the standard face image, and perform latent space encoding and feature decoupling on the low-order high-frequency features to generate a three-dimensional image corresponding to the standard face image in the latent space.
[0104] S142: Determine multiple deformation regions in the three-dimensional image and the deformation features corresponding to each deformation region through the key point generation module.
[0105] S143: Use the spatial local deformation module to determine the deformation region and deformation direction corresponding to the expression driving target, and perform feature editing on the deformation features of the deformation region based on the deformation direction to generate an expression driving image.
[0106] S144: Input the expression driving image into the feature decoding module, so that the feature decoding module uses the generator composed of the SPADE network to perform upsampling decoding on the expression driving image to map the expression driving image to the two-dimensional space and generate an expression driving video.
[0107] In this embodiment, the target expression driving model is composed of multiple modules, namely, the latent space mapping module, the key point generation module, the spatial local deformation module, and the feature decoding module. During the process of the target expression driving model driving the standard face image, each module performs its corresponding function, so that the final target expression driving model can output an expression driving video.
[0108] Among them, the latent space mapping module is mainly responsible for extracting the low-order high-frequency features of the standard face image, performing latent space encoding and feature decoupling on the low-order high-frequency features to generate a three-dimensional image corresponding to the standard face image in the latent space; the key point generation module is mainly responsible for determining multiple deformation regions in the three-dimensional image and the deformation features corresponding to each deformation region; the spatial local deformation module is mainly responsible for determining the deformation region and deformation direction corresponding to the expression driving target, and performing feature editing on the deformation features of the deformation region based on the deformation direction to generate an expression driving image; the feature decoding module is mainly responsible for using the generator composed of the SPADE network to perform upsampling decoding on the expression driving image to map the expression driving image to the two-dimensional space and generate an expression driving video.
[0109] It can be understood that in the feature decoding module, the generator for upsampling decoding is constructed by the SPADE network (Spatially-Adaptive Denormalization), which is mainly used for image generation tasks, especially tasks of generating images guided by semantic segmentation maps or other spatial structures. Here, SPADE is a spatially adaptive normalization method that can dynamically adjust the features of the generated image according to the input semantic mask map. Since traditional normalization methods normalize the feature maps of each pixel equally, resulting in loss of spatial information, this application can introduce the SPADE network, which introduces different adjustment parameters (scale and offset) according to the semantic category of each pixel during normalization, so as to ensure that spatial information is retained during normalization, making the generated image more realistic and natural while maintaining detail and structural consistency.
[0110] Specifically, since an RGB image can be regarded as a low-order high-frequency feature with strong coupling, the low-order high-frequency features of a standard face image can be extracted through the latent space mapping module, and then the low-order high-frequency features are encoded in the latent space and feature decoupled to deform the spatial size of the face region of the standard face image in the latent space. After mapping the standard face image to the three-dimensional latent space, the key point generation module can determine multiple deformation regions in the latent space, such as regions of eyes, nose, mouth, etc., and then can generate a set of key points corresponding to each deformation region, and further can calculate the deformation features of each deformation region based on each set of key points; therefore, the key point generation module can control the deformation features of each deformation region by controlling each set of key points.
[0111] In addition, to support global and local facial expression editing drives, the spatial local deformation module can change the deformation features of its corresponding deformation region by controlling any one or more sets of key points, and thus can achieve facial expression editing of the corresponding deformation region without affecting the pixels of other regions. For example, the spatial local deformation module can control the opening and closing of a person's mouth through a set of key points, and this set of key points cannot control the deformation of regions other than the mouth, so as to achieve constraints on the range of expression drive and ensure the high quality of the output image. When the spatial local deformation module generates an expression-driven image, the feature decoding module can decode the three-dimensional latent space features of the image into an RGB image by means of upsampling decoding and output to form an expression-driven video, which is the video of the person with user-defined expressions.
[0112] In one embodiment, each deformation region in step S143 is composed of multiple deformation key points; wherein, the step of determining multiple deformation regions in the three-dimensional image and the deformation features corresponding to each deformation region through the key point generation module may include:
[0113] S1431: Determine multiple deformation regions in the three-dimensional image and the three-dimensional coordinates of each deformation key point in each deformation region through the key point generation module.
[0114] S1432: For each deformation region, use the Gaussian distribution technology in the key point generation module to calculate the influence on the three-dimensional coordinates of each deformation key point in this deformation region, obtain the feature influence value of each deformation key point in the three-dimensional image, and perform weighted summation on each feature influence value to form the deformation feature of this deformation region.
[0115] In this embodiment, after mapping the standard face image to the three-dimensional latent space, the key point generation module can determine multiple deformation regions in the three-dimensional image and the three-dimensional coordinates of each deformation key point in each deformation region in the latent space. Then, for each deformation region, the computer device can use the Gaussian distribution technology to calculate the influence on the three-dimensional coordinates of each deformation key point in this deformation region, obtain the feature influence value of each deformation key point in the three-dimensional image, and perform weighted summation on each feature influence value to form the deformation feature of this deformation region.
[0116] It can be understood that the Gaussian distribution is also called the normal distribution (Gaussian Distribution or Normal Distribution), which is used to describe the distribution of data near the average value and has the characteristics of a symmetric bell-shaped curve. In this application, the key point generation module can use the Gaussian distribution to quantify the influence of the generated key points, ensuring that the deformation features calculated through each key point are based on the comprehensive influence of multiple points, rather than relying solely on a single key point, so as to ensure the smoothness and naturalness of the facial expression driving effect.
[0117] For example, the key point generation module can generate a set of key points with coordinates (x, y, z) in the latent space to control the position deformation of the deformation region features in the latent space. The key points corresponding to each deformation region can be 20. Each key point has a direct influence on the image features around it in the three dimensions of x, y, and z. At this time, the key point generation module can use a Gaussian distribution to represent the influence range and influence degree of each point, and finally perform weighted summation on the influences of these points to obtain the final deformation feature.
[0118] Next, the human expression driving device provided by the embodiment of the present application will be described. The human expression driving device described below can be correspondingly referred to the human expression driving method described above.
[0119] In one embodiment, as Figure 5 shown, Figure 5Schematic structural diagram of a human expression driving device provided by an embodiment of the present application; the present application also provides a human expression driving device, including a data acquisition module 210, a face recognition module 220, a model determination module 230, and an expression driving module 240, specifically including the following:
[0120] The data acquisition module 210 is configured to acquire a human video containing a face, as well as an expression driving target of the human video, and parse to obtain video frame images of the human video.
[0121] The face recognition module 220 is configured to use a target face detection model to identify face key points in the video frame image, and perform face alignment on the video frame image based on the face key points to obtain a standard face image.
[0122] The model determination module 230 is configured to determine a target expression driving model; the target expression driving model is composed of a three-dimensional convolutional neural network, and is used to map a two-dimensional image into a three-dimensional latent space for feature encoding and expression driving.
[0123] The expression driving module 240 is configured to input the standard face image and the expression driving target into the target expression driving model, and use the target expression driving model to perform expression driving on the standard face image based on the expression driving target, so as to output an expression driving video.
[0124] In the above embodiment, when performing expression driving on a person in a human video, the expression driving target of the human video can be determined first, and continuous video frame images in the human video can be parsed, so that expression driving can be realized on the video frame images, ensuring the naturalness of the expression driving effect; then the target face detection model can be used to identify face key points in the video frame image, and perform face alignment on the video frame image based on the face key points to obtain a standard face image, so as to perform expression driving on the standard face image based on the expression driving target through the target expression driving model, and output it as an expression driving video, thereby improving the quality of the expression driving effect. Since the target expression driving model is composed of a three-dimensional convolutional neural network and is used to map a two-dimensional image into a three-dimensional latent space for feature encoding and expression driving, the target expression driving model can accurately track faces at various angles in the video frame image and perform expression editing and driving, realizing natural expression connection between frames of the video frame image without affecting the quality of the human video.
[0125] In one embodiment, the face recognition module 220 may include:
[0126] A model determination sub-model, configured to determine a target face detection model, and the target face detection model includes a face recognition network and a key point detection network.
[0127] A face recognition sub-model, which is used to perform face recognition on video frame images by using a face recognition network to obtain a face position box in the video frame images.
[0128] A key point detection sub-model, which is used to perform key point detection on the original face image based on the face position box in a key point detection network to obtain face key points.
[0129] In one embodiment, the face recognition module 220 may further include:
[0130] A sample acquisition sub-model, which is used to acquire a sample face image marked with a real face position box, and the real face position box is marked with real face key points.
[0131] A model prediction sub-model, which is used to input the sample face image into a preset initial face detection model to obtain a predicted face position box output by the initial face detection model for the sample face image, as well as the predicted confidence of the predicted face position box and predicted face key points.
[0132] A model training sub-model, which is used to train the initial face detection model with the goal that the predicted face position box and predicted face key points approach the real face position box and real face key points respectively.
[0133] A model lightweighting sub-model, which is used to lightweight the trained initial face detection model to obtain a target face detection model when the initial face detection model meets the preset training conditions.
[0134] In one embodiment, the model training sub-model includes:
[0135] A first loss value calculation unit, which is used to determine the position loss value of the predicted face position box based on the real face position box and calculate the confidence loss value of the predicted confidence according to the position loss value.
[0136] A second loss value calculation unit, which is used to determine the key point loss value of the predicted face key points based on the real face key points.
[0137] A parameter update unit, which is used to update the parameters in the initial face detection model according to the position loss value, confidence loss value and key point loss value.
[0138] In one embodiment, the face recognition module 220 may further include:
[0139] A position transformation sub-model, which is used to locate the face area in the video frame image according to the face key points and use the affine transformation technology to align the located face area to a preset standard position to form a standard face image.
[0140] In one embodiment, the target expression driving model in the expression driving module 240 includes a latent space mapping module, a key point generation module, a spatial local deformation module, and a feature decoding module; the expression driving module 240 may include:
[0141] A latent space mapping sub-model, configured to extract low-order high-frequency features of a standard face image by using the latent space mapping module, perform latent space encoding and feature decoupling on the low-order high-frequency features, so as to generate a three-dimensional image corresponding to the standard face image in the latent space.
[0142] A deformation feature determination sub-model, configured to determine multiple deformation regions in the three-dimensional image and the deformation features corresponding to each deformation region through the key point generation module.
[0143] A feature editing sub-model, configured to determine the deformation region and deformation direction corresponding to the expression driving target by using the spatial local deformation module, and perform feature editing on the deformation features of the deformation region based on the deformation direction to generate an expression driving image.
[0144] A feature decoding sub-model, configured to input the expression driving image into the feature decoding module, so that the feature decoding module uses a generator composed of a SPADE network to perform upsampling decoding on the expression driving image, so as to map the expression driving image to a two-dimensional space and generate an expression driving video.
[0145] In one embodiment, the feature editing sub-model may include:
[0146] A three-dimensional coordinate determination unit, configured to determine multiple deformation regions in the three-dimensional image and the three-dimensional coordinates of each deformation key point in each deformation region through the key point generation module.
[0147] A deformation feature calculation unit, configured to perform influence calculation on the three-dimensional coordinates of each deformation key point in the deformation region by using a Gaussian distribution technique in the key point generation module for each deformation region, obtain the feature influence value of each deformation key point in the three-dimensional image, and perform weighted summation on the respective feature influence values to form the deformation feature of the deformation region.
[0148] In one embodiment, the present application further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the human expression driving method according to any one of the above embodiments.
[0149] In one embodiment, the present application further provides a computer device, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the human expression driving method according to any one of the above embodiments.
[0150] Schematically, as Figure 6 shown, Figure 6 is a schematic internal structure diagram of a computer device provided by an embodiment of the present application. The computer device 300 can be provided as a server. Referring to Figure 6 , the computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by a memory 301 for storing instructions executable by the processing component 302, such as application programs. The application programs stored in the memory 301 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the human facial expression driving method of any of the above embodiments.
[0151] The computer device 300 may further include a power component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 can operate based on an operating system stored in the memory 301, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM, or the like.
[0152] Those skilled in the art can understand that Figure 6 the structure shown in
[0153] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0154] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0155] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for driving a character's expression, characterized in that: The method comprises: Acquire a person video containing a human face and an expression driving target of the person video, and parse to obtain continuous video frame images in the person video; Using a target face detection model to identify key face points in the video frame image, and performing face alignment on the video frame image based on the key face points to obtain a standard face image; Determine a target expression driving model; the target expression driving model is composed of a three-dimensional convolutional neural network, which is used to map a two-dimensional image into a three-dimensional latent space for feature encoding and expression driving; Inputting the standard face image and the expression driving target into the target expression driving model, and using the target expression driving model to perform expression driving on the standard face image based on the expression driving target, so as to output an expression driving video; The target expression driving model includes a latent space mapping module, a key point generation module, a spatial local deformation module and a feature decoding module; The step of using the target expression driving model to perform expression driving on the standard face image based on the expression driving target includes: The latent space mapping module is used to extract low-order and high-frequency features of the standard face image, and the low-order and high-frequency features are latently encoded and feature decoupled to generate a three-dimensional image corresponding to the standard face image in the latent space; Determining a plurality of deformation regions in the three-dimensional image and a deformation feature corresponding to each deformation region by means of the key point generation module; Determine the deformation area and deformation direction corresponding to the expression-driven target by using the spatial local deformation module, and perform feature editing on the deformation features of the deformation area based on the deformation direction to generate an expression-driven image; The expression-driven image is input into the feature decoding module, so that the feature decoding module uses a generator composed of a SPADE network to upsample and decode the expression-driven image, so as to map the expression-driven image to a two-dimensional space and generate an expression-driven video.
2. The character expression driving method according to claim 1, characterized in that: The step of using the target face detection model to identify key face points in the video frame image includes: Determine a target face detection model, wherein the target face detection model includes a face recognition network and a key point detection network; Performing face recognition on the video frame image using the face recognition network to obtain a face position frame in the video frame image; In the key point detection network, key point detection is performed on the original face image based on the face position frame to obtain face key points.
3. The character expression driving method according to claim 1 or 2, characterized in that: The training process of the target face detection model includes: Acquire a sample face image marked with a real face position frame, wherein the real face position frame is marked with real face key points; Inputting the sample face image into a preset initial face detection model, obtaining a predicted face position frame output by the initial face detection model for the sample face image, and a predicted confidence level and predicted face key points of the predicted face position frame; Training the initial face detection model with the goal of making the predicted face position frame and the predicted face key points approach the real face position frame and the real face key points respectively; When the initial face detection model meets the preset training conditions, the trained initial face detection model is lightweighted to obtain a target face detection model.
4. The character expression driving method according to claim 3, characterized in that: The training of the initial face detection model with the goal of making the predicted face position frame and the predicted face key points approach the real face position frame and the real face key points respectively comprises: Determine a position loss value of the predicted face position frame based on the real face position frame, and calculate a confidence loss value of the prediction confidence according to the position loss value; Determine a key point loss value of the predicted face key point based on the real face key point; Update the parameters in the initial face detection model according to the position loss value, the confidence loss value and the key point loss value.
5. The character expression driving method according to claim 1, characterized in that: The step of performing face alignment on the video frame image based on the face key points to obtain a standard face image includes: The face region in the video frame image is located according to the face key points, and the located face region is aligned to a preset standard position using affine transformation technology to form a standard face image.
6. The character expression driving method according to claim 1, characterized in that: Each deformation region is composed of multiple deformation key points; The determining of the plurality of deformation regions in the three-dimensional image and the deformation features corresponding to each deformation region by the key point generation module includes: Determine, by means of the key point generation module, a plurality of deformation regions in the three-dimensional image and the three-dimensional coordinates of each deformation key point in each deformation region; For each deformation area, Gaussian distribution technology is used in the key point generation module to calculate the influence of the three-dimensional coordinates of each deformation key point in the deformation area, obtain the characteristic influence value of each deformation key point in the three-dimensional image, and perform weighted summation on each characteristic influence value to form the deformation feature of the deformation area.
7. A character expression driving device, characterized in that: include: A data acquisition module, used to acquire a person video containing a human face and an expression driving target of the person video, and parse to obtain continuous video frame images in the person video; A face recognition module, used to identify key points of faces in the video frame image using a target face detection model, and perform face alignment on the video frame image based on the key points of faces to obtain a standard face image; A model determination module is used to determine a target expression driving model; the target expression driving model is composed of a three-dimensional convolutional neural network, which is used to map a two-dimensional image into a three-dimensional latent space for feature encoding and expression driving; An expression driving module, used for inputting the standard face image and the expression driving target into the target expression driving model, and using the target expression driving model to perform expression driving on the standard face image based on the expression driving target, so as to output an expression driving video; The target expression driving model in the expression driving module includes a latent space mapping module, a key point generation module, a spatial local deformation module and a feature decoding module; The expression driving module also includes: The latent space mapping module is used to extract low-order and high-frequency features of the standard face image, and the low-order and high-frequency features are latently encoded and feature decoupled to generate a three-dimensional image corresponding to the standard face image in the latent space; Determining a plurality of deformation regions in the three-dimensional image and a deformation feature corresponding to each deformation region by means of the key point generation module; Determine the deformation area and deformation direction corresponding to the expression-driven target by using the spatial local deformation module, and perform feature editing on the deformation features of the deformation area based on the deformation direction to generate an expression-driven image; The expression-driven image is input into the feature decoding module, so that the feature decoding module uses a generator composed of a SPADE network to upsample and decode the expression-driven image, so as to map the expression-driven image to a two-dimensional space and generate an expression-driven video.
8. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the character expression driving method as described in any one of claims 1 to 6.
9. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the character expression driving method according to any one of claims 1 to 6 are performed.
Citation Information
Patent Citations
Real-time video generation method and device based on AIGC, equipment and storage medium
CN119031212A