Digital life generation method, device and storage medium
The method addresses the challenge of creating scene-matched digital humans by preprocessing and neural network-enhanced editing, ensuring high interaction and personalization in digital human technologies.
Patent Information
- Application Number
- CN202210541984.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-18
AI Technical Summary
In the prior art, models cannot fully match the time and actions of the interactive scene when recording image materials, resulting in the need of authenticity and personalization of digital human image customization.
By acquiring the video and adjusting the resolution and frame rate, combining the neural network model for super-resolution and video interpolation processing, editing the character image, expression and actions based on the interactive scene information, and generating digital human videos that match the interactive scene.
It realizes the precise matching of digital human image, expression and actions with interactive scenes, improves the authenticity and personalized customization efficiency of digital human videos, and reduces the differences in model recording time and action jump problems.
Smart Images

Figure CN114863533B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a digital human generation method, an apparatus, and a storage medium. Background Art
[0002] Driven by new technologies such as artificial intelligence and virtual reality, the performance of digital humans has been improved. Digital humans represented by virtual anchors and virtual employees have successfully entered the public eye and shone brightly in many fields such as film and television, games, media, culture and tourism, and finance in diverse forms.
[0003] The customization of digital human images aims for authenticity and personalization. Under the requirement of photo-realistic ultra-high definition, every detail of the digital human image will be concerned by users. This poses high requirements for models when recording image materials. However, after all, models are not robots and cannot achieve perfect matching with the interaction scenarios used for the image in terms of time and action positioning. Summary of the Invention
[0004] Embodiments of the present disclosure perform editing processing on the characters in a video according to the character customization information corresponding to the interaction scenario, and generate a digital human video that matches the interaction scenario through character editing.
[0005] Some embodiments of the present disclosure provide a digital human generation method, including:
[0006] Obtaining a first video;
[0007] Performing editing processing on the characters in each frame image of the first video according to the character customization information corresponding to the interaction scenario;
[0008] Outputting a second video according to the processed frame images in the first video.
[0009] In some embodiments, the first video is obtained by preprocessing the original video, and the preprocessing includes one or more of resolution adjustment, inter-frame smoothing processing, and frame rate adjustment.
[0010] In some embodiments, the resolution adjustment includes:
[0011] If the resolution of the original video is higher than the required preset resolution, downsampling the original video according to the preset resolution to obtain a first video with the preset resolution;
[0012] If the resolution of the original video is lower than the required preset resolution, processing the original video using a super-resolution model to obtain a first video with the preset resolution, where the super-resolution model is used to increase the resolution of the input video to the preset resolution.
[0013] In some embodiments, the super-resolution model is obtained by training a neural network. During the training process, the first video frame from a high-definition video is downsampled to a second video frame according to a preset resolution. The second video frame is used as the input of the neural network, and the first video frame is used as the supervision information of the output of the neural network to train the neural network to obtain the super-resolution model.
[0014] In some embodiments, the frame rate adjustment includes:
[0015] If the frame rate of the original video is higher than the required preset frame rate, the original video is decimated according to the ratio information of the frame rate of the original video to the preset frame rate to obtain a first video with the preset frame rate;
[0016] If the frame rate of the original video is lower than the required preset frame rate, the original video is interpolated to a first frame rate using a video interpolation model. The first frame rate is the least common multiple of the frame rate before interpolation of the original video and the preset frame rate. The interpolated original video is decimated according to the ratio information of the first frame rate to the preset frame rate to obtain a first video with the preset frame rate. The video interpolation model is used to generate transitional frames between any two frames of images.
[0017] In some embodiments, the video interpolation model is obtained by training a neural network. During the training process, three consecutive frames in the training video frame sequence are used as a triple. The first frame and the third frame in the triple are used as the input of the neural network, and the second frame in the triple is used as the supervision information of the output of the neural network to train the neural network to obtain the video interpolation model.
[0018] In some embodiments, the input of the neural network includes: the visual feature information and depth information of the first frame and the third frame, and the optical flow information and deformation information between the first frame and the third frame.
[0019] In some embodiments, the editing process of the characters in each frame image of the first video according to the character customization information corresponding to the interaction scenario includes one or more of the following:
[0020] Editing the character images in each frame image of the first video according to the character image customization information corresponding to the interaction scenario;
[0021] Editing the character expressions in each frame image of the first video according to the character expression customization information corresponding to the interaction scenario;
[0022] Editing the character actions in each frame image of the first video according to the character action customization information corresponding to the interaction scenario.
[0023] In some embodiments, the editing process of the character images in each frame of the first video according to the character image customization information corresponding to the interaction scenario includes: determining the character image adjustment parameters according to the character image adjustments made by the user in some video frames of the first video, and editing the character images in the remaining video frames of the first video according to the character image adjustment parameters.
[0024] In some embodiments, the editing process of the character images in the remaining video frames of the first video according to the character image adjustment parameters includes:
[0025] According to the target part of the character image adjustment in the character image adjustment parameters, locate the target part of the character in the remaining video frames of the first video through key point detection;
[0026] According to the amplitude information or position information of the character image adjustment in the character image adjustment parameters, adjust the amplitude or position of the located target part through graphics transformation.
[0027] In some embodiments, the character expression customization information includes preset classification information corresponding to the target expression. The editing process of the character expressions in each frame of the first video according to the character expression customization information corresponding to the interaction scenario includes:
[0028] Obtain the feature information of each frame image in the first video, the feature information of the face key points, and the classification information of the original expression;
[0029] Fuse the feature information of each frame image, the feature information of the face key points, and the classification information of the original expression with the preset classification information corresponding to the target expression to obtain the feature information of the fused image corresponding to each frame image;
[0030] Generate the fused image corresponding to each frame image according to the feature information of the fused image corresponding to each frame image, and all the fused images form a second video with the face expression being the target expression.
[0031] In some embodiments, the obtaining the feature information of each frame image in the first video, the feature information of the face key points, and the classification information of the original expression includes:
[0032] Input each frame image in the first video into a face feature extraction model to obtain the feature information of the output each frame image;
[0033] Input the feature information of each frame image into a face key point detection model to obtain the coordinate information of the face key points of each frame image, and use the principal component analysis method to reduce the dimension of the coordinate information of all face key points to obtain information in a preset dimension as the feature information of the face key points;
[0034] Input the feature information of each frame of the image into the expression classification model to obtain the classification information of the original expression of each frame of the image.
[0035] In some embodiments, the fusion of the feature information of each frame of the image, the feature information of the facial key points, the classification information of the original expression, and the preset classification information corresponding to the target expression includes:
[0036] Sum and average the classification information of the original expression of each frame of the image and the preset classification information corresponding to the target expression to obtain the classification information of the fused expression corresponding to each frame of the image;
[0037] Concatenate the feature information of the facial key points of each frame of the image multiplied by the first weight obtained through training, the feature information of each frame of the image multiplied by the second weight obtained through training, and the classification information of the fused expression corresponding to each frame of the image.
[0038] In some embodiments, generating the fused image corresponding to each frame of the image according to the feature information of the fused image corresponding to each frame of the image includes:
[0039] Input the feature information of the fused image corresponding to each frame of the image into the decoder, and output the fused image corresponding to each frame of the image generated;
[0040] Wherein, the facial feature extraction model includes a convolutional layer, and the decoder includes a transposed convolutional layer.
[0041] In some embodiments, input the first video with the original expression of the facial expression and the preset classification information corresponding to the target expression into the expression generation model, and output the second video with the target expression of the facial expression; the training method of the expression generation model includes:
[0042] Obtain the training pairs composed of each frame of the image of the first training video and each frame of the image of the second training video;
[0043] Input each frame of the image of the first training video into the first generator, obtain the feature information of each frame of the image of the first training video, the feature information of the facial key points, and the classification information of the original expression, fuse the feature information of each frame of the image of the first training video, the feature information of the facial key points, the classification information of the original expression, and the preset classification information corresponding to the target expression, obtain the feature information of each frame of the fused image corresponding to the first training video, and obtain each frame of the fused image corresponding to the first training video output by the first generator according to the feature information of each frame of the fused image corresponding to the first training video;
[0044] Input each frame image of the second training video into the second generator to obtain the feature information of each frame image of the second training video, the feature information of facial key points, and the classification information of the target expression. Fuse the feature information of each frame image of the second training video, the feature information of facial key points, the classification information of the target expression, and the preset classification information corresponding to the original expression to obtain the feature information of each frame of fused image corresponding to the second training video. According to the feature information of each frame of fused image corresponding to the second training video, obtain each frame of fused image corresponding to the second training video output by the second generator;
[0045] Determine the adversarial loss and the cycle consistency loss according to each frame of fused image corresponding to the first training video and each frame of fused image corresponding to the second training video;
[0046] Train the first generator and the second generator according to the adversarial loss and the cycle consistency loss. After the first generator is trained, it is used as an expression generation model.
[0047] In some embodiments, it further includes: determining the pixel-to-pixel loss according to the pixel difference between every two adjacent frames of fused images corresponding to the first training video and the pixel difference between every two adjacent frames of fused images corresponding to the second training video;
[0048] Wherein, the training of the first generator and the second generator according to the adversarial loss and the cycle consistency loss includes:
[0049] Train the first generator and the second generator according to the adversarial loss, the cycle consistency loss, and the pixel-to-pixel loss.
[0050] In some embodiments, the determining the adversarial loss according to each frame of fused image corresponding to the first training video and each frame of fused image corresponding to the second training video includes: inputting each frame of fused image corresponding to the first training video into the first discriminator to obtain the first discrimination result of each frame of fused image corresponding to the first training video;
[0051] Input each frame of fused image corresponding to the second training video into the second discriminator to obtain the second discrimination result of each frame of fused image corresponding to the second training video;
[0052] Determine the first adversarial loss according to the first discrimination result of each frame of fused image corresponding to the first training video, and determine the second adversarial loss according to the second discrimination result of each frame of fused image corresponding to the second training video.
[0053] In some embodiments, inputting each frame fusion image corresponding to the first training video into a first discriminator to obtain a first discrimination result for each frame fusion image corresponding to the first training video includes:
[0054] Inputting each frame fusion image corresponding to the first training video into a first face feature extraction model in the first discriminator to obtain feature information of each frame fusion image corresponding to the first training video output;
[0055] Inputting the feature information of each frame fusion image corresponding to the first training video into a first expression classification model in the first discriminator to obtain classification information of the expressions of each frame fusion image corresponding to the first training video as a first discrimination result;
[0056] Inputting each frame fusion image corresponding to the second training video into a second discriminator to obtain a second discrimination result for each frame fusion image corresponding to the second training video includes:
[0057] Inputting each frame fusion image corresponding to the second training video into a second face feature extraction model in the second discriminator to obtain feature information of each frame fusion image corresponding to the second training video output;
[0058] Inputting the feature information of each frame fusion image corresponding to the second training video into a second expression classification model in the second discriminator to obtain classification information of the expressions of each frame fusion image corresponding to the second training video as a second discrimination result.
[0059] In some embodiments, the cycle-consistency loss is determined by the following method:
[0060] Inputting each frame fusion image corresponding to the first training video into the second generator to generate reconstructed images of each frame of the first training video, and inputting each frame fusion image corresponding to the second training video into the first generator to generate reconstructed images of each frame of the second training video;
[0061] Determine the cycle-consistency loss according to the difference between the reconstructed images of each frame of the first training video and the images of each frame of the first training video, and the difference between the reconstructed images of each frame of the second training video and the images of each frame of the second training video.
[0062] In some embodiments, the pixel-to-pixel loss is determined by the following method:
[0063] For each position in every two adjacent frame fusion images corresponding to the first training video, determine the distance between the representation vectors of the two pixels at this position in the two adjacent frame fusion images, and sum up the distances corresponding to all positions to obtain a first loss;
[0064] For each position in every two adjacent frame fusion images corresponding to the second training video, determine the distance between the representation vectors of the two pixels at the position in the two adjacent frame fusion images, and sum up the distances corresponding to all positions to obtain the second loss;
[0065] Sum up the first loss and the second loss to obtain the pixel-to-pixel loss.
[0066] In some embodiments, the obtaining of the feature information of each frame image, the feature information of facial key points, and the classification information of the original expression of the first training video includes: inputting each frame image in the first training video into the third facial feature extraction model in the first generator to obtain the feature information of the output each frame image; inputting the feature information of each frame image into the first facial key point detection model in the first generator to obtain the coordinate information of the facial key points of each frame image; using the principal component analysis method to reduce the dimension of the coordinate information of all facial key points to obtain the first information of a preset dimension as the feature information of the facial key points of each frame image in the first training video; inputting the feature information of each frame image in the first training video into the third expression classification model in the first generator to obtain the classification information of the original expression of each frame image in the first training video;
[0067] The obtaining of the feature information of each frame image, the feature information of facial key points, and the classification information of the target expression of the second training video includes: inputting each frame image in the second training video into the fourth facial feature extraction model in the second generator to obtain the feature information of the output each frame image; inputting the feature information of each frame image into the second facial key point detection model in the second generator to obtain the coordinate information of the facial key points of each frame image; using the principal component analysis method to reduce the dimension of the coordinate information of all facial key points to obtain the second information of a preset dimension as the feature information of the facial key points of each frame image in the second training video; inputting the feature information of each frame image in the second training video into the fourth expression classification model in the second generator to obtain the classification information of the target expression of each frame image in the second training video.
[0068] In some embodiments, the fusion of the feature information of each frame image of the first training video, the feature information of the facial key points, the classification information of the original expression, and the preset classification information corresponding to the target expression includes: adding and averaging the classification information of the original expression of each frame image of the first training video and the preset classification information corresponding to the target expression to obtain the classification information of the fused expression corresponding to each frame image of the first training video; splicing the feature information of the facial key points of each frame image of the first training video multiplied by the first weight to be trained, the feature information of each frame image of the first training video multiplied by the second weight to be trained, and the classification information of the fused expression corresponding to each frame image of the first training video;
[0069] The fusion of the feature information of each frame image of the second training video, the feature information of the facial key points, the classification information of the target expression, and the preset classification information corresponding to the original expression includes: adding and averaging the classification information of the target expression of each frame image of the second training video and the preset classification information corresponding to the original expression to obtain the classification information of the fused expression corresponding to each frame image of the second training video; splicing the feature information of the facial key points of each frame image of the second training video multiplied by the third weight to be trained, the feature information of each frame image of the second training video multiplied by the fourth weight to be trained, and the classification information of the fused expression corresponding to each frame image of the second training video.
[0070] In some embodiments, the training of the first generator and the second generator according to the adversarial loss, the cycle-consistent loss, and the pixel-to-pixel loss includes: performing a weighted sum of the adversarial loss, the cycle-consistent loss, and the pixel-to-pixel loss to obtain a total loss; training the first generator and the second generator according to the total loss.
[0071] In some embodiments, the editing process of the human actions in each frame image of the first video according to the human action customization information corresponding to the interaction scenario includes:
[0072] Adjusting the first human key points of the human in the original first key frame of the first video during the first action to obtain the second human key points of the human during the second action, which are used as the human action customization information;
[0073] Extracting the feature information of the neighborhood of each second human key point from the original first key frame;
[0074] Inputting each second human key point and its neighborhood feature information into an image generation model to output the target first key frame of the human during the second action.
[0075] In some embodiments, the method for obtaining the image generation model includes: using the training video frames and the human key points of the people in the training video frames as a pair of training data, using the feature information of the human key points and their neighborhoods in the training video frames in the training data as the input of the image generation network, using the training video frames in the training data as the supervision information for the output of the image generation network, and training the image generation network to obtain the image generation model.
[0076] In some embodiments, the first human key points include the human contour feature points when the person makes a first action, and the second human key points include the human contour feature points when the person makes a second action.
[0077] Some embodiments of the present disclosure propose a digital human generation device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the digital human generation methods described in the various embodiments based on instructions stored in the memory.
[0078] Some embodiments of the present disclosure propose a digital human generation device, including:
[0079] An acquisition unit configured to acquire a first video;
[0080] A customization unit configured to edit the people in each frame image of the first video according to the character customization information corresponding to the interaction scenario;
[0081] An output unit configured to output a second video according to each frame image in the processed first video.
[0082] Some embodiments of the present disclosure propose a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the digital human generation methods described in the various embodiments are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The following will briefly introduce the drawings required for use in the embodiments or related art descriptions. The present disclosure can be more clearly understood according to the following detailed description with reference to the drawings.
[0084] Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0085] Figure 1A A flowchart showing the digital human generation method according to some embodiments of the present disclosure.
[0086] Figure 1B A flowchart showing the digital human generation method according to other embodiments of the present disclosure.
[0087] Figure 2 Schematic diagram showing video preprocessing according to some embodiments of the present disclosure.
[0088] Figure 3A Flow schematic diagram showing the method for generating expressions according to some embodiments of the present disclosure.
[0089] Figure 3B Schematic diagram showing the method for generating expressions according to other embodiments of the present disclosure.
[0090] Figure 3C Flow schematic diagram showing the training method of the expression generation model according to some embodiments of the present disclosure.
[0091] Figure 3D Schematic diagram showing the training method of the expression generation model according to some embodiments of the present disclosure.
[0092] Figure 4A Schematic diagram showing the human body contour feature points of a person in a first action according to some embodiments of the present disclosure.
[0093] Figure 4B Schematic diagram showing the human body contour feature points of a person in a second action according to some embodiments of the present disclosure.
[0094] Figure 4C Schematic diagram showing multiple key points and multiple key connection lines on a person according to some embodiments of the present disclosure.
[0095] Figure 5 Schematic diagram showing the structure of the digital human generation device according to some embodiments of the present disclosure.
[0096] Figure 6 Schematic diagram showing the structure of the digital human generation device according to other embodiments of the present disclosure. Detailed implementation manners
[0097] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure.
[0098] Unless otherwise specified, the descriptions such as "first" and "second" in the present disclosure are used to distinguish different objects, and do not represent meanings such as size or time sequence.
[0099] Figure 1A Flow schematic diagram showing the digital human generation method according to some embodiments of the present disclosure.
[0100] As Figure 1A shown, the digital human generation method of this embodiment includes the following steps.
[0101] In step S110, a first video is acquired.
[0102] The first video can be, for example, the original recorded video or the one obtained by preprocessing the original video. The preprocessing includes one or more of resolution adjustment, inter-frame smoothing processing, and frame rate adjustment.
[0103] In step S120, according to the character customization information corresponding to the interaction scenario, the characters in each frame image of the first video are edited.
[0104] The editing of the characters in each frame image of the first video according to the character customization information corresponding to the interaction scenario includes one or more of the following: according to the character image customization information corresponding to the interaction scenario, the character images in each frame image of the first video are edited to generate a digital human image matching the interaction scenario; according to the character expression customization information corresponding to the interaction scenario, the character expressions in each frame image of the first video are edited to generate a digital human expression matching the interaction scenario; according to the character action customization information corresponding to the interaction scenario, the character actions in each frame image of the first video are edited to generate a digital human action matching the interaction scenario.
[0105] In step S130, according to each frame image in the processed first video, a second video is output.
[0106] That is, the frame images in the processed first video are combined to form the second video, and the second video is a digital human video matching the interaction scenario.
[0107] In the above embodiment, the characters in the video are edited according to the character customization information corresponding to the interaction scenario, and a digital human video matching the interaction scenario is generated through character editing. For example, a digital human image, digital human expression, digital human action, etc. matching the interaction scenario are generated.
[0108] Figure 1B The flowchart showing some other embodiments of the digital human generation method of the present disclosure.
[0109] As Figure 1B shown, the digital human generation method of this embodiment includes the following steps.
[0110] In step S210, custom logic control is performed.
[0111] The custom logic control is used to control whether to execute the custom logics such as video preprocessing, image customization, expression customization, and action customization, and the execution order, etc.
[0112] The content edited in each part such as video preprocessing, image customization, expression customization, and motion customization is independent, and there is no strong dependency between them. Therefore, the execution order of each part can be swapped, and the basic effect of generating a digital human video matching the interaction scenario can be achieved. However, there is still a certain mutual influence between the parts. According to the execution order of S220 - S250 in this embodiment, the mutual influence between the parts can be minimized, and the final presentation effect of the character image is better.
[0113] In step S220, perform video preprocessing.
[0114] Video preprocessing is to preprocess the recorded original video to obtain a first video, and the preprocessing includes one or more of resolution adjustment, inter-frame smoothing processing, and frame rate adjustment.
[0115] In some embodiments, as Figure 2 shown, the preprocessing is sequentially performed in the order of resolution adjustment, inter-frame smoothing processing, and frame rate adjustment. The effect of video preprocessing is better, which can retain the visual information of the original video to the greatest extent, ensure that the preprocessed video does not have quality problems such as blurring and distortion, and minimize the impact of frame rate adjustment and resolution adjustment on the subsequent digital human customization process.
[0116] The resolution adjustment includes: if the resolution of the original video is higher than the required preset resolution, downsample the original video according to the preset resolution to obtain a first video with the preset resolution; if the resolution of the original video is lower than the required preset resolution, use a super-resolution model to process the original video to obtain a first video with the preset resolution, and the super-resolution model is used to increase the resolution of the input video to the preset resolution; if the resolution of the original video is equal to the required preset resolution, the resolution adjustment step can be skipped.
[0117] Through resolution adjustment, the first video after preprocessing can be made consistent in terms of resolution, and the impact of the differential resolution of the original video on the digital human customization effect can be reduced.
[0118] The super-resolution model is, for example, obtained by training a neural network. During the training process, the first video frame from a high-definition video is downsampled according to a preset resolution to obtain a second video frame. The second video frame is used as the input of the neural network, and the first video frame is used as the supervision information for the output of the neural network. The neural network is trained to obtain the super-resolution model. Among them, the difference information between the video frame output by the neural network and the first video frame is used as the loss function, and the parameters of the neural network are iteratively updated according to the loss determined by the loss function until the loss meets certain conditions and the training is completed. At this time, the video frame output by the neural network is very close to the first video frame, and the trained neural network is used as the super-resolution model. Among them, the neural network is a large class of models, such as including but not limited to convolutional neural networks, recurrent networks based on optical flow methods, generative adversarial networks, etc.
[0119] For example, the key frames of a high-definition video (1080p) are downsampled to obtain a second video frame with a lower resolution (such as 360p / 480p / 720p, etc.). According to the above training method, a super-resolution model is obtained. Using this super-resolution model, a first video with a resolution of 480p / 720p / 1080p, etc. can be obtained from the original video with any resolution. Among them, 360p / 480p / 720p / 1080p is a video display format, and P represents progressive scanning. For example, the picture resolution of 1080p is 1920 multiplied by 1080.
[0120] After the resolution is adjusted, there may be a certain gap in the texture information between two frames in the frame sequence generated or downsampled by the super-resolution model. Therefore, frame interpolation smoothing is adopted here to ensure that there are no jagged edges or moiré patterns in the texture, human edges, etc. during video playback, avoiding visual impacts.
[0121] Frame interpolation smoothing can, for example, adopt an average smoothing method. For example, the image information of three consecutive frames is averaged, and this average value is used as the image information of the middle frame among these three consecutive frames.
[0122] The frame rate adjustment includes: if the frame rate of the original video is higher than the required preset frame rate, the original video is decimated according to the ratio information between the frame rate of the original video and the preset frame rate to obtain a first video with the preset frame rate; if the frame rate of the original video is lower than the required preset frame rate, the original video is interpolated to the first frame rate using a video interpolation model, where the first frame rate is the least common multiple of the frame rate before interpolation of the original video and the preset frame rate. The interpolated original video is decimated according to the ratio information between the first frame rate and the preset frame rate to obtain a first video with the preset frame rate. The video interpolation model is used to generate transitional frames between any two frames of images; if the frame rate of the original video is equal to the required preset frame rate, the frame rate adjustment step can be skipped.
[0123] Through frame rate adjustment, the first video after preprocessing can be made consistent in terms of frame rate, reducing the impact of the differential frame rate of the original video on the digital human customization effect. Moreover, the frame interpolation operation can effectively solve the problem of jumps between two actions. For example, when the digital human finishes action A and then performs action B, playing the video without frame interpolation processing will make users feel that the character's actions are jumping and not very realistic. In this embodiment, frame interpolation will insert several transition frames between the key frames of the two actions, making the users feel that the character's actions transition naturally and realistically when playing the video after frame interpolation processing.
[0124] The video frame interpolation model is, for example, obtained by training a neural network. During the training process, three consecutive frames in the training video frame sequence are used as a triple. The first frame and the third frame in the triple are used as the input of the neural network, and the second frame in the triple is used as the supervised information of the output of the neural network to train the neural network to obtain the video frame interpolation model. Among them, the difference information between the video frame output by the neural network based on the first frame and the third frame in the input triple and the second frame in the triple is used as the loss function, and the parameters of the neural network are iteratively updated according to the loss determined by the loss function until the loss meets certain conditions and the training is completed. At this time, the video frame output by the neural network is very close to the second frame in the triple. Using the trained neural network as the video frame interpolation model can generate transition frames between any two frames of images. Among them, the neural network is a large class of models, such as including but not limited to convolutional neural networks, recurrent networks based on optical flow methods, generative adversarial networks, etc.
[0125] Among them, the input of the neural network includes, for example: the visual feature information and depth information of the first frame and the third frame, as well as the optical flow information and deformation information between the first frame and the third frame. Through the fusion of these four parts of information, the transition frames to be inserted between the two frames inferred can make the video transition more smoothly.
[0126] In step S230, image customization.
[0127] According to the character image customization information corresponding to the interaction scenario, edit the character images in each frame of the first video to meet the user's needs for digital human beauty and body shaping. Among them, image customization includes, for example, beauty and body shaping operations such as skin smoothing, face slimming, eye enlarging, adjustment of the positions of facial features, adjustment of body proportions, such as slimming and leg stretching.
[0128] In some embodiments, according to the character image adjustment made by the user in some video frames of the first video, the character image adjustment parameters are determined, and the character images in the remaining video frames of the first video are edited according to the character image adjustment parameters. Among them, the "partial video frames" may be, for example, one or several key frames in the first video. First, the image customization of all video digital humans can be completed through a small amount of editing work, improving the digital human customization efficiency and customization cost.
[0129] The editing the character images in the remaining video frames of the first video according to the character image adjustment parameters includes: according to the target part of the character image adjustment in the character image adjustment parameters, the target part of the character in the remaining video frames of the first video is located through key point detection. The target part is, for example, facial features or the human body, etc.; according to the amplitude information or position information of the character image adjustment in the character image adjustment parameters, the amplitude or position of the located target part is adjusted through graphics transformation.
[0130] For example, if the user enlarges the eyes of the character in some key frames, first, the face is detected through face detection technology, then, the eyes of the character in the remaining video frames are located through key point detection technology, and then, according to the amplitude information of the user enlarging the eyes, for example, the amplitude of the adjustment of the distance between the upper and lower eyelids, the amplitude of the eyes of the character in the remaining video frames is adjusted through graphics transformation, achieving the beauty effect of big eyes for the characters in all frames of the video.
[0131] In step S240, expression customization.
[0132] Expression customization refers to an expression generation method for editing the expressions of the characters in each frame image of the first video according to the character expression customization information corresponding to the interaction scenario, such as the preset classification information corresponding to the target expression, realizing the control of the digital human's facial expressions in the interaction scenario, and can transfer one expression state of the digital human to another target expression state, while ensuring that only the facial expressions of the digital human change, and the speaking mouth shape, head movement, etc. are not affected. Thus, when the digital human expresses the corresponding language content, the expression can change accordingly with the language content.
[0133] Figure 3A It is a flowchart of some embodiments of the expression generation method of the present disclosure. As Figure 3A shown, the method of this embodiment includes: steps S310 to S330.
[0134] In step S310, the feature information of each frame image, the feature information of the face key points, and the classification information of the original expression in the first video are obtained.
[0135] The facial expression in the first video is the original expression. That is, the facial expression in each frame image of the first video is mainly the original expression, and the original expression is, for example, a calm expression.
[0136] In some embodiments, each frame image in the first video is input into a facial feature extraction model to obtain the feature information of each output frame image; the feature information of each frame image is input into a facial key point detection model to obtain the coordinate information of the facial key points of each frame image; the principal component analysis (PCA) is used to reduce the dimension of the coordinate information of all facial key points to obtain the information of a preset dimension as the feature information of the facial key points; the feature information of each frame image is input into an expression classification model to obtain the classification information of the original expression of each frame image.
[0137] The overall expression generation model includes an encoder and a decoder. The encoder can include a facial feature extraction model, a facial key point detection model, and an expression classification model. The facial feature extraction model is connected to the facial key point detection model and the expression classification model. The facial feature extraction model can adopt an existing model. For example, deep learning models with feature extraction functions such as VGG-19, ResNet, and Transformer can be used. The part before VGG-19 block 5 can be used as the facial feature extraction model. The facial key point detection model and the expression classification model can also adopt existing models, such as MLP (multi-layer perceptron), specifically a 3-layer MLP. After the expression generation model is trained, it is used to generate expressions, and the training process will be described in detail later.
[0138] The feature information of each frame image in the first video is, for example, the feature map output by the facial feature extraction model. The key points include, for example, 68 key points such as the chin, the center of the eyebrows, and the corners of the mouth. Each key point is represented by the horizontal and vertical coordinates of its location. After obtaining the coordinate information of each key point through the facial key point detection model, in order to reduce redundant information and improve efficiency, the PCA is used to reduce the dimension of the coordinate information of all facial key points to obtain the information of a preset dimension (for example, 6 dimensions, which can achieve the best effect) as the feature information of the facial key points. The expression classification model can output the classification of several expressions such as neutral, happy, and sad, and can be represented by a one-hot encoded vector. The classification information of the original expression can be the one-hot encoding of the classification of the original expression of each frame image in the first video obtained through the expression classification model.
[0139] In step S320, the feature information of each frame image, the feature information of the facial key points, and the classification information of the original expression are fused with the preset classification information corresponding to the target expression to obtain the feature information of the fused image corresponding to each frame image.
[0140] In some embodiments, the classification information of the original expression of each frame of image is added to and averaged with the preset classification information corresponding to the target expression to obtain the classification information of the fused expression corresponding to each frame of image; the feature information of the facial key points of each frame of image multiplied by the first weight obtained through training, the feature information of each frame of image multiplied by the second weight obtained through training, and the classification information of the fused expression corresponding to each frame of image are concatenated.
[0141] The target expression is different from the original expression. For example, it is a smiling expression, and the preset classification information corresponding to the target expression is, for example, the preset one-hot encoding of the target expression. The preset classification information does not need to be obtained through the model and can be directly encoded using the preset encoding rule (one-hot). For example, a calm expression is encoded as 1000, and a smiling expression is encoded as 0100. The aforementioned classification information of the original expression is obtained through an expression classification model, and this classification information can be different from the preset classification information corresponding to the original expression. For example, the original expression is a calm expression, the preset one-hot encoding is 1000, but the one-hot encoding obtained by the expression classification model can be 0.8 0.2 0 0.
[0142] The encoder may further include a feature fusion model. The feature information of each frame of image, the feature information of the facial key points, the classification information of the original expression, and the preset classification information corresponding to the target expression are input into the feature fusion model for fusion. The parameters to be trained in the feature fusion model include the first weight and the second weight. For each frame of image, the first weight obtained through training is multiplied by the feature information of the facial key points of this image to obtain a first feature vector, and the second weight obtained through training is multiplied by the feature information of this image to obtain a second feature vector. The first feature vector, the second feature vector, and the classification information of the fused expression corresponding to this image are concatenated to obtain the feature information of the fused image corresponding to this image. The first weight and the second weight can unify the value ranges of the three types of information.
[0143] In step S330, according to the feature information of the fused image corresponding to each frame of image, the fused image corresponding to each frame of image is generated, and all the fused images are combined to form a second video with the target expression as the facial expression.
[0144] In some embodiments, the feature information of the fused image corresponding to each frame of image is input into the decoder, and the fused image corresponding to each frame of image generated is output. The facial feature extraction model includes a convolutional layer, and the decoder includes a transposed convolutional layer, and an image can be generated based on the features. The decoder is, for example, block 5 of VGG-19, and the last convolutional layer is replaced with a transposed convolutional layer. The fused image is an image with the target expression as the facial expression, and each frame of fused image forms a second video.
[0145] The following combines Figure 3B to describe some application examples of the present disclosure.
[0146] As shown in Figure 3B , for a frame image in the first video, after feature extraction, a feature map is obtained. Based on the feature map, face key point detection and expression classification are respectively performed. The feature information of each key point obtained from the face key point detection is subjected to PCA, and the dimension is reduced to the information of a preset dimension as the key point feature. The classification information of the original expression is subjected to one-hot encoding and fused with the preset classification information corresponding to the target expression to obtain an expression classification vector (the classification information of the fused expression). Furthermore, the feature map of the face, the expression classification vector, and the key point feature are fused to obtain the feature information of the fused image. The feature information of the fused image is subjected to feature decoding to obtain the face image of the target expression.
[0147] The solution of the above embodiment extracts the feature information of each frame image, the feature information of the face key points, and the classification information of the original expression in the first video, fuses the extracted information with the preset classification information corresponding to the target expression, obtains the feature information of the fused image corresponding to each frame image, and then generates the fused image corresponding to each frame image according to the feature information of the fused image corresponding to each frame image. All the fused images can form a second video with the face expression being the target expression. In the above embodiment, by extracting the feature information of the face key points and using it for feature fusion, the expression in the fused image is made more real and smooth. Through the fusion of the preset classification information corresponding to the target expression, the generation of the target expression is directly realized, and it is compatible with the facial movements and lip shapes of the characters in the original image, without affecting the lip shapes, head movements, etc. of the characters and without affecting the clarity of the original image, making the generated video stable, clear, and smooth.
[0148] Figure 3C It is a flowchart of some embodiments of the training method of the expression generation model of the present disclosure. The expression generation model can output a second video with the face expression being the target expression according to the first video with the face expression being the original expression and the preset classification information corresponding to the target expression input.
[0149] As shown in Figure 3C , the method of this embodiment includes steps S410 to S450.
[0150] In step S410, a training pair composed of each frame image of the first training video and each frame image of the second training video is obtained.
[0151] The first training video is a video with the face expression being the original expression, and the second training video is a video with the face expression being the target expression. Each frame image of the first training video and each frame image of the second training video do not need to correspond one by one. The classification information of the original expression and the classification information of the target expression are labeled.
[0152] Using videos of a large number of people speaking with different expressions as training data, cross-domain transfer learning (Domain Transfer Learning) is carried out through deep learning to learn a first generator for converting from one expression state to another, and then the expression generation result is integrated with the entire digital human.
[0153] In step S420, each frame image of the first training video is input into the first generator to obtain the feature information of each frame image of the first training video, the feature information of the face key points, and the classification information of the original expression. The feature information of each frame image of the first training video, the feature information of the face key points, the classification information of the original expression, and the preset classification information corresponding to the target expression are integrated to obtain the feature information of each frame of the fused image corresponding to the first training video. According to the feature information of each frame of the fused image corresponding to the first training video, each frame of the fused image corresponding to the first training video output by the first generator is obtained.
[0154] After the first generator is trained, it is used as an expression generation model. In some embodiments, each frame image in the first training video is input into the third face feature extraction model in the first generator to obtain the feature information of the output frame images; the feature information of each frame image is input into the first face key point detection model in the first generator to obtain the coordinate information of the face key points of each frame image; the principal component analysis method is used to reduce the dimension of the coordinate information of all face key points to obtain the first information of the preset dimension as the feature information of the face key points of each frame image of the first training video; the feature information of each frame image in the first training video is input into the third expression classification model in the first generator to obtain the classification information of the original expression of each frame image in the first training video.
[0155] Perform principal component analysis (PCA) on the coordinate information of the face key points, and reduce the key point coordinate information to 6 dimensions (6 dimensions is the best effect obtained through a large number of experiments). PCA does not involve training parameters (the feature extraction of PCA and the corresponding relationship of the front and back feature dimensions do not change with training. When backpropagating the gradient, only the feature correspondence obtained by the initial PCA is used to pass the gradient to the previous parameters).
[0156] In some embodiments, the classification information of the original expression of each frame image of the first training video is added and averaged with the preset classification information corresponding to the target expression to obtain the classification information of the fused expression corresponding to each frame image of the first training video; the feature information of the face key points of each frame image of the first training video multiplied by the first weight to be trained, the feature information of each frame image of the first training video multiplied by the second weight to be trained, and the classification information of the fused expression corresponding to each frame image of the first training video are spliced to obtain the feature information of each frame of the fused image corresponding to the first training video.
[0157] The first generator includes a first feature fusion model, and the first weight and the second weight are parameters to be trained in the first feature fusion model. The processes of the above-mentioned feature extraction and feature fusion can refer to the foregoing embodiments.
[0158] The first generator includes a first encoder and a first decoder. The first encoder includes: a third face feature extraction model, a first face key point detection model, a third expression classification model, and a first feature fusion model. The feature information of each frame of the fused image corresponding to the first training video is input into the first decoder to obtain each frame of the fused image corresponding to the generated first training video.
[0159] In step S430, each frame image of the second training video is input into the second generator to obtain the feature information of each frame image of the second training video, the feature information of the face key points, and the classification information of the target expression. The feature information of each frame image of the second training video, the feature information of the face key points, the classification information of the target expression, and the preset classification information corresponding to the original expression are fused to obtain the feature information of each frame of the fused image corresponding to the second training video. According to the feature information of each frame of the fused image corresponding to the second training video, each frame of the fused image corresponding to the second training video output by the second generator is obtained.
[0160] The second generator is the same as or similar to the first generator in structure. The training objective of the second generator is to generate a video with the same expression as the first training video based on the second training video.
[0161] In some embodiments, each frame image of the second training video is input into the fourth face feature extraction model in the second generator to obtain the feature information of the output frame images; the feature information of each frame image is input into the second face key point detection model in the second generator to obtain the coordinate information of the face key points of each frame image; the principal component analysis method is used to reduce the dimension of the coordinate information of all face key points to obtain the second information of the preset dimension as the feature information of the face key points of each frame image of the second training video. The feature information of each frame image of the second training video is input into the fourth expression classification model in the second generator to obtain the classification information of the target expression of each frame image in the second training video.
[0162] The dimension of the feature information of the face key points of each frame image of the second training video is the same as that of the feature information of the face key points of each frame image of the first training video. For example, 6 dimensions.
[0163] In some embodiments, the classification information of the target expression of each frame image of the second training video is added to and averaged with the preset classification information corresponding to the original expression to obtain the classification information of the fused expression corresponding to each frame image of the second training video; the feature information of the face key points of each frame image of the second training video after being multiplied by the third weight to be trained, the feature information of each frame image of the second training video after being multiplied by the fourth weight to be trained, and the classification information of the fused expression corresponding to each frame image of the second training video are spliced to obtain the feature information of each frame of the fused image corresponding to the second training video.
[0164] The preset classification information corresponding to the original expression does not need to be obtained through the model and can be directly encoded using the preset encoding rule. The second generator includes a second feature fusion model, and the third weight and the fourth weight are the parameters to be trained in the second feature fusion model. The above processes of feature extraction and feature fusion can refer to the foregoing embodiments and will not be elaborated here.
[0165] The second generator includes a second encoder and a second decoder. The second encoder includes: a fourth face feature extraction model, a second face key point detection model, a fourth expression classification model, and a second feature fusion model. The feature information of each frame of the fused image corresponding to the second training video is input into the second decoder to obtain each frame of the fused image corresponding to the generated second training video.
[0166] In step S440, the adversarial loss and the cycle consistency loss are determined according to each frame of the fused image corresponding to the first training video and each frame of the fused image corresponding to the second training video.
[0167] End-to-end training based on generative adversarial learning and cross-domain transfer learning can improve the accuracy of the model and the training efficiency.
[0168] In some embodiments, the adversarial loss is determined by the following method: each frame of the fused image corresponding to the first training video is input into the first discriminator to obtain the first discrimination result of each frame of the fused image corresponding to the first training video; each frame of the fused image corresponding to the second training video is input into the second discriminator to obtain the second discrimination result of each frame of the fused image corresponding to the second training video; the first adversarial loss is determined according to the first discrimination result of each frame of the fused image corresponding to the first training video, and the second adversarial loss is determined according to the second discrimination result of each frame of the fused image corresponding to the second training video.
[0169] Further, in some embodiments, the fused images of each frame corresponding to the first training video are input into the first face feature extraction model in the first discriminator to obtain the feature information of the fused images of each frame corresponding to the first training video; the feature information of the fused images of each frame corresponding to the first training video is input into the first expression classification model in the first discriminator to obtain the classification information of the expressions of the fused images of each frame corresponding to the first training video, which is used as the first discrimination result; the fused images of each frame corresponding to the second training video are input into the second face feature extraction model in the second discriminator to obtain the feature information of the fused images of each frame corresponding to the second training video; the feature information of the fused images of each frame corresponding to the second training video is input into the second expression classification model in the second discriminator to obtain the classification information of the expressions of the fused images of each frame corresponding to the second training video, which is used as the second discrimination result.
[0170] During the training process, the overall model includes two sets of generators and discriminators. The structures of the first discriminator and the second discriminator are the same or similar, and both include a face feature extraction model and an expression classification model. The structures of the first face feature extraction model, the second face feature extraction model are the same or similar to those of the third face feature extraction model and the fourth face feature extraction model, and the structures of the first expression classification model, the second expression classification model are the same or similar to those of the third expression classification model and the fourth expression classification model.
[0171] For example, the data of the first video is represented by X = {x i}, and the data of the second video is represented by Y = {y i}. The first generator G is used to implement X→Y, and during training, G(x) is made to be as close as possible to Y. The first discriminator D Y is used to discriminate the authenticity of the fused images of each frame corresponding to the first training video. The first adversarial loss can be expressed by the following formula:
[0172]
[0173] The second generator F is used to implement Y→X, and during training, F(Y) is made to be as close as possible to X. The second discriminator D X is used to discriminate the authenticity of the fused images of each frame corresponding to the second training video. The second adversarial loss can be expressed by the following formula:
[0174]
[0175] In some embodiments, the cycle consistency losses are determined as follows: The fused images of each frame corresponding to the first training video are input into the second generator to generate the reconstructed images of each frame of the first training video, and the fused images of each frame corresponding to the second training video are input into the first generator to generate the reconstructed images of each frame of the second training video; The cycle consistency losses are determined based on the differences between the reconstructed images of each frame of the first training video and the images of each frame of the first training video, and the differences between the reconstructed images of each frame of the second training video and the images of each frame of the second training video.
[0176] To further improve the accuracy of the model, the images generated by the first generator are input into the second generator to obtain the reconstructed images of each frame of the first training video. It is expected that the reconstructed images of each frame of the first training video generated by the second generator are as consistent as possible with the images of each frame of the first training video, that is, F(G(x))≈x. The images generated by the second generator are input into the first generator to obtain the reconstructed images of each frame of the second training video. It is expected that the reconstructed images of each frame of the second training video generated by the first generator are as consistent as possible with the images of each frame of the second training video, that is, G(F(y))≈y.
[0177] The difference between the reconstructed images of each frame of the first training video and the images of each frame of the first training video can be determined as follows: For each reconstructed image of each frame of the first training video and the image of the first training video corresponding to the reconstructed image, the distance (such as the Euclidean distance) between the representation vectors of the pixels at each same position of the reconstructed image and the corresponding image is determined, and all the distances are summed up.
[0178] The difference between the reconstructed images of each frame of the second training video and the images of each frame of the second training video can be determined as follows: For each reconstructed image of each frame of the second training video and the image of the second training video corresponding to the reconstructed image, the distance (such as the Euclidean distance) between the representation vectors of the pixels at each same position of the reconstructed image and the corresponding image is determined, and all the distances are summed up.
[0179] In step S450, the first generator and the second generator are trained according to the adversarial loss and the cycle consistency loss.
[0180] The first adversarial loss, the second adversarial loss, and the cycle consistency loss can be weighted and summed to obtain the total loss, and the first generator and the second generator are trained according to the total loss. For example, the total loss can be determined by the following formula:
[0181] L = L GAN (G, D Y , X, Y) + L GAN (F, D X , Y, X) + λL cyc(G,F) (3)
[0182] Among them, L cyc (G,F) represents the cycle consistency loss, and λ is a weight that can be obtained through training.
[0183] To further improve the accuracy of the model and ensure the stable continuity of the output video results, the loss brought by the pixel difference between two consecutive frames of the video is added during the training process. In some embodiments, according to the pixel differences between every two adjacent frame fusion images corresponding to the first training video and the pixel differences between every two adjacent frame fusion images corresponding to the second training video, the pixel-to-pixel loss is determined, and the first generator and the second generator are trained according to the adversarial loss, the cycle consistency loss, and the pixel-to-pixel loss.
[0184] Furthermore, in some embodiments, for each position in every two adjacent frame fusion images corresponding to the first training video, the distance between the representation vectors of the two pixels at this position in the two adjacent frame fusion images is determined, and the distances corresponding to all positions are summed to obtain the first loss; for each position in every two adjacent frame fusion images corresponding to the second training video, the distance between the representation vectors of the two pixels at this position in the two adjacent frame fusion images is determined, and the distances corresponding to all positions are summed to obtain the second loss; the first loss and the second loss are summed to obtain the pixel-to-pixel loss. The pixel-to-pixel loss can make the change between two consecutive frames of the generated video not too large.
[0185] In some embodiments, the adversarial loss, the cycle consistency loss, and the pixel-to-pixel loss are weighted and summed to obtain the total loss; the first generator and the second generator are trained according to the total loss. For example, the total loss can be determined by the following formula:
[0186] L = L GAN (G,D Y ,X,Y) + L GAN (F,D X ,Y,X) + λ1L cyc (G,F) + λ2L P2P (G(x i ),G(x i+1 )) + λ3L P2P (F(y j ),F(y j+1 )) (4)
[0187] Among them, λ1, λ2, and λ3 are weights that can be obtained through training, and L P2P (G(x i ),G(x i+1 )) represents the first loss, and L P2P (F(y j), F(y j+1 ) represents the second loss.
[0188] As Figure 3D shown, before performing end-to-end training, pre-training can be performed on the models of each part. For example, first, a large number of open-source face recognition data are selected to pre-train the face recognition model, and the part before its output feature map is selected as the face feature extraction model (the method of this part is not unique. Taking vgg-19 as an example, the part before block5 is selected, and a feature map of 8×8×512 dimensions can be output). After that, the face feature extraction model and its parameters are fixed, and then it is divided into two branches at the back. The two branches are the face key point detection model and the expression classification model. The face key point detection dataset and the expression classification data are used to fine-tune each branch respectively, and only the parameters in the model structures of these two parts are trained. The face key point detection model is not unique, as long as it is a model based on a convolutional network model that can obtain accurate key points, it can be connected to this solution; the expression classification model is a single-label classification task based on a convolutional network model. After pre-training, the end-to-end training process can be performed based on the foregoing embodiments. This can improve the training efficiency.
[0189] The method of the above embodiment uses adversarial loss, cycle-consistent loss, and pixel loss between two adjacent frames of the video to train the overall model, which can improve the accuracy of the model, and the end-to-end training process can improve the efficiency and save computing resources.
[0190] The solution of the present disclosure is applicable to the editing of human face expressions in videos. The present disclosure adopts a unique deep learning model, integrates technologies such as expression recognition and key point detection, learns the rules of the movement of human face key points under different expressions through data training, and finally controls the facial expression state output by the model by inputting the classification information of the target expression. And the expression exists as a style state, and when the person speaks or makes actions such as tilting the head or blinking, good effects can be superimposed, so that the finally output human face action video is natural and not obtrusive. The output result can have the same resolution and detail level as the input image, and still maintain stable, clear, and flawless output results at 1080p or even 2k resolution.
[0191] In step S250, action customization.
[0192] Action customization refers to editing and processing the human actions in each frame image of the first video according to the human action customization information corresponding to the interaction scenario, so as to realize the editing and control of the digital human actions in the interaction scenario.
[0193] In some embodiments, customizing information according to the corresponding human actions in the interaction scenario and editing the human actions in each frame image of the first video includes: adjusting the first human body key points of the human in the original first key frame of the first video during the first action to obtain the second human body key points of the human during the second action, which are used as the customized information of the human action; a feature extraction model, such as a convolutional kernel model, can be used to extract the feature information of the neighborhood of each second human body key point from the original first key frame; inputting each second human body key point and its neighborhood feature information into an image generation model to output the target first key frame of the human during the second action.
[0194] The first human body key points include the human body contour feature points of the human during the first action, such as Figure 4A the 14 pairs of white dots shown, and the second human body key points include the human body contour feature points of the human during the second action, such as Figure 4B the 14 pairs of white dots shown.
[0195] Using the human body contour feature points for human action editing generates more accurate human actions compared to using the human body skeleton feature points for human action editing, and is less likely to exhibit phenomena such as deformation and distortion, improving the quality of the generated images.
[0196] Before adjusting the human body contour feature points of the human during the first action, first extract the human body contour feature points of the human during the first action. Extracting the human body contour feature points of the human during the first action, for example, includes: using a semantic segmentation network model to extract the contour line of the human; using an object detection network model to extract multiple key points on the human, such as Figure 4C the black dots shown; according to the structural information of the human, connecting the multiple key points to determine multiple key connecting lines, such as Figure 4C the white straight lines shown; according to the intersection points of the perpendicular lines of the multiple key connecting lines and the contour line, determining multiple pairs of human body contour feature points of the human during the first action.
[0197] The method for obtaining the image generation model includes: using the training video frames and the human key points of the people in the training video frames as a pair of training data, using the human key points and the feature information of their neighborhoods in the training video frames in the training data as the input of the image generation network, using the training video frames in the training data as the supervision information of the output of the image generation network, and training the image generation network to obtain the image generation model. Among them, the difference information between the video frames output by the image generation network based on the input data and the training video frames is used as the loss function, and the parameters of the image generation network are iteratively updated according to the loss determined by the loss function until the loss meets certain conditions and the training is completed. At this time, the video frames output by the image generation network are very close to the training video frames, and the trained image generation network is used as the image generation model. Among them, the image generation network is a large class of models, such as including but not limited to convolutional neural networks, recurrent networks based on optical flow methods, generative adversarial networks, etc. If the image generation network is a generative adversarial network, the total loss function also includes the discriminant loss function of the image discriminant network.
[0198] In step S260, render and output.
[0199] Model the character image using the material results processed in each of steps S220 to 250. Different rendering techniques can be selected according to the application scenario, and combined with artificial intelligence techniques such as intelligent dialogue, speech recognition, speech synthesis, and motion interaction to form a complete digital human video (i.e., the second video) that can interact with the scene and output.
[0200] In the above embodiments, the characters in the video are edited according to the character customization information corresponding to the interaction scenario, and a digital human video matching the interaction scenario is generated through character editing. For example, a digital human image, digital human expression, digital human action, etc. matching the interaction scenario are generated. According to the method of the embodiments of the present disclosure, by recording a set of character image videos, multiple sets of videos with different character image styles in different scenarios can be quickly produced. Moreover, there is no need for professional engineers to intervene, and users can adjust the image, expression, action, etc. of the characters according to the scene needs.
[0201] Figure 5 The structural schematic diagram of a digital human generation device showing some embodiments of the present disclosure is as follows Figure 5 As shown, the digital human generation device 500 of this embodiment includes units 510 to 530.
[0202] The acquisition unit 510 is configured to acquire the first video, and for details, refer to step S220.
[0203] The customization unit 520 is configured to edit the characters in each frame image of the first video according to the character customization information corresponding to the interaction scenario, and for details, refer to steps S230 to 250.
[0204] The customization unit 520 includes, for example, an image customization unit 521, an expression customization unit 522, an action customization unit 523, etc. The image customization unit 521 is configured to edit the character image in each frame of the first video according to the character image customization information corresponding to the interaction scenario. For details, refer to step S230. The expression customization unit 522 is configured to edit the character expression in each frame of the first video according to the character expression customization information corresponding to the interaction scenario. For details, refer to step S240. The action customization unit 523 is configured to edit the character action in each frame of the first video according to the character action customization information corresponding to the interaction scenario. For details, refer to step S250.
[0205] The output unit 530 is configured to output a second video according to each frame image in the processed first video. For details, refer to step S260.
[0206] Figure 6 The structural schematic diagram of the digital human generation device showing some other embodiments of the present disclosure is as follows. Figure 6 As shown, the digital human generation device 600 of this embodiment includes: a memory 610 and a processor 620 coupled to the memory 610. The processor 620 is configured to execute the digital human generation method in any of the foregoing embodiments based on the instructions stored in the memory 610.
[0207] Among them, the memory 610 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.
[0208] Among them, the processor 620 can be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, or discrete hardware components such as transistors.
[0209] The device 600 may further include an input / output interface 630, a network interface 640, a storage interface 650, etc. These interfaces 630, 640, 650 and the memory 610 and the processor 620 may be connected, for example, through a bus 660. Among them, the input / output interface 630 provides connection interfaces for input / output devices such as displays, mice, keyboards, touchscreens, etc. The network interface 640 provides connection interfaces for various networking devices. The storage interface 650 provides connection interfaces for external storage devices such as SD cards and USB flash drives. The bus 660 may use any bus structure among various bus structures. For example, the bus structure includes, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.
[0210] Some embodiments of the present disclosure propose a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the digital human generation method in any of the foregoing embodiments are implemented.
[0211] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more non-transitory computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer program code.
[0212] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.
[0213] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram.
[0214] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus so that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram.
[0215] The foregoing are only preferred embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for generating digital lives, characterized in that, Including: Obtain a first video; According to the personalized information of the corresponding person in the interaction scenario, perform editing processing on the people in each frame image of the first video, including: according to the personalized information of the corresponding person's expression in the interaction scenario, perform editing processing on the expressions of the people in each frame image of the first video, and the personalized information of the person's expression includes the preset classification information corresponding to the target expression; According to each frame image in the processed first video, output a second video, wherein, input the first video with the original expression on the face and the preset classification information corresponding to the target expression into an expression generation model, and output a second video with the target expression on the face. The training method of the expression generation model includes: Obtain a training pair composed of each frame image of a first training video and each frame image of a second training video; Input each frame image of the first training video into a first generator, obtain the feature information of each frame image of the first training video, the feature information of the facial key points, and the classification information of the original expression, fuse the feature information of each frame image of the first training video, the feature information of the facial key points, the classification information of the original expression, and the preset classification information corresponding to the target expression, obtain the feature information of each frame of the fused image corresponding to the first training video, and according to the feature information of each frame of the fused image corresponding to the first training video, obtain each frame of the fused image corresponding to the first training video output by the first generator; Input each frame image of the second training video into a second generator, obtain the feature information of each frame image of the second training video, the feature information of the facial key points, and the classification information of the target expression, fuse the feature information of each frame image of the second training video, the feature information of the facial key points, the classification information of the target expression, and the preset classification information corresponding to the original expression, obtain the feature information of each frame of the fused image corresponding to the second training video, and according to the feature information of each frame of the fused image corresponding to the second training video, obtain each frame of the fused image corresponding to the second training video output by the second generator; Determine the adversarial loss and the cycle consistency loss according to each frame of the fused image corresponding to the first training video and each frame of the fused image corresponding to the second training video; According to the adversarial loss and the cycle consistency loss, train the first generator and the second generator, and use the first generator as an expression generation model after training is completed.
2. The method according to claim 1, characterized in that, The first video is obtained by preprocessing the original video, and the preprocessing includes one or more of resolution adjustment, inter-frame smoothing processing, and frame rate adjustment.
3. The method according to claim 2, wherein The resolution adjustment includes: If the resolution of the original video is higher than the required preset resolution, downsample the original video according to the preset resolution to obtain a first video with the preset resolution; If the resolution of the original video is lower than the required preset resolution, process the original video using a super-resolution model to obtain a first video with the preset resolution, and the super-resolution model is used to increase the resolution of the input video to the preset resolution.
4. The method according to claim 3, wherein The super-resolution model is obtained by training a neural network. During the training process, the first video frame from a high-definition video is downsampled according to a preset resolution to obtain a second video frame. The second video frame is used as the input of the neural network, and the first video frame is used as the supervision information of the output of the neural network to train the neural network to obtain the super-resolution model.
5. The method according to claim 2, wherein The frame rate adjustment includes: If the frame rate of the original video is higher than the required preset frame rate, the original video is decimated according to the ratio information of the frame rate of the original video to the preset frame rate to obtain a first video with the preset frame rate; If the frame rate of the original video is lower than the required preset frame rate, the original video is interpolated to a first frame rate using a video interpolation model. The first frame rate is the least common multiple of the frame rate before interpolation of the original video and the preset frame rate. The interpolated original video is decimated according to the ratio information of the first frame rate to the preset frame rate to obtain a first video with the preset frame rate. The video interpolation model is used to generate transitional frames between any two frames of images.
6. The method according to claim 5, characterized in that, The video interpolation model is obtained by training a neural network. During the training process, three consecutive frames in the training video frame sequence are used as a triple. The first frame and the third frame in the triple are used as the input of the neural network, and the second frame in the triple is used as the supervision information of the output of the neural network to train the neural network to obtain the video interpolation model.
7. The method according to claim 6, characterized in that, The input of the neural network includes: visual feature information and depth information of the first frame and the third frame, as well as optical flow information and deformation information between the first frame and the third frame.
8. The method according to claim 1, wherein The editing process of the characters in each frame image of the first video according to the character customization information corresponding to the interaction scenario further includes one or more of the following: Editing the character image in each frame image of the first video according to the character image customization information corresponding to the interaction scenario; Editing the character actions in each frame image of the first video according to the character action customization information corresponding to the interaction scenario.
9. The method according to claim 8, wherein The editing process of the character image in each frame image of the first video according to the character image customization information corresponding to the interaction scenario includes: Determining character image adjustment parameters according to the character image adjustments made by the user in some video frames of the first video, and editing the character images in the remaining video frames of the first video according to the character image adjustment parameters.
10. The method according to claim 9, characterized in that, The editing process of the character images in the remaining video frames of the first video according to the character image adjustment parameters includes: Locating the target part of the character in the remaining video frames of the first video through key point detection according to the target part of the character image adjustment in the character image adjustment parameters; Adjusting the amplitude or position of the located target part through graphics transformation according to the amplitude information or position information of the character image adjustment in the character image adjustment parameters.
11. According to the method described in claim 1, wherein The character expression customization information includes preset classification information corresponding to the target expression, The editing process of the character expressions in each frame image of the first video according to the character expression customization information corresponding to the interaction scenario includes: Obtain the feature information of each frame image, the feature information of facial key points, and the classification information of the original expression in the first video; Fuse the feature information of each frame image, the feature information of facial key points, and the classification information of the original expression with the preset classification information corresponding to the target expression to obtain the feature information of the fused image corresponding to each frame image; Generate the fused image corresponding to each frame image according to the feature information of the fused image corresponding to each frame image, and all the fused images form a second video with the facial expression being the target expression.
12. The method according to claim 11, wherein The obtaining the feature information of each frame image, the feature information of facial key points, and the classification information of the original expression in the first video includes: Input each frame image in the first video into a facial feature extraction model to obtain the feature information of the output each frame image; Input the feature information of each frame image into a facial key point detection model to obtain the coordinate information of the facial key points of each frame image, and use the principal component analysis method to reduce the dimension of the coordinate information of all facial key points to obtain information of a preset dimension as the feature information of the facial key points; Input the feature information of each frame image into an expression classification model to obtain the classification information of the original expression of each frame image.
13. The method according to claim 11, wherein The fusing the feature information of each frame image, the feature information of facial key points, and the classification information of the original expression with the preset classification information corresponding to the target expression includes: Sum and average the classification information of the original expression of each frame image and the preset classification information corresponding to the target expression to obtain the classification information of the fused expression corresponding to each frame image; Stitch together the feature information of the facial key points of each frame image multiplied by the first weight obtained through training, the feature information of each frame image multiplied by the second weight obtained through training, and the classification information of the fused expression corresponding to each frame image.
14. The method according to claim 12, wherein The generating the fused image corresponding to each frame image according to the feature information of the fused image corresponding to each frame image includes: Input the feature information of the fused image corresponding to each frame image into a decoder to output the generated fused image corresponding to each frame image; Among them, the facial feature extraction model includes a convolutional layer, and the decoder includes a transposed convolutional layer.
15. The method according to claim 1, characterized in that It further includes: Determine the pixel-to-pixel loss according to the pixel difference between every two adjacent fused images corresponding to the first training video and the pixel difference between every two adjacent fused images corresponding to the second training video; Among them, the training the first generator and the second generator according to the adversarial loss and the cycle consistency loss includes: Train the first generator and the second generator according to the adversarial loss, the cycle consistency loss, and the pixel-to-pixel loss.
16. The method according to claim 1 or 15, characterized in that The determining the adversarial loss according to each frame of fused image corresponding to the first training video and each frame of fused image corresponding to the second training video includes: Input each frame of fused image corresponding to the first training video into a first discriminator to obtain the first discrimination result of each frame of fused image corresponding to the first training video; Input each frame fusion image corresponding to the second training video into the second discriminator to obtain the second discrimination result of each frame fusion image corresponding to the second training video; Determine the first adversarial loss according to the first discrimination result of each frame fusion image corresponding to the first training video, and determine the second adversarial loss according to the second discrimination result of each frame fusion image corresponding to the second training video.
17. The method according to claim 16, wherein Input each frame fusion image corresponding to the first training video into the first discriminator to obtain the first discrimination result of each frame fusion image corresponding to the first training video, including: Input each frame fusion image corresponding to the first training video into the first face feature extraction model in the first discriminator to obtain the feature information of each frame fusion image corresponding to the first training video output; Input the feature information of each frame fusion image corresponding to the first training video into the first expression classification model in the first discriminator to obtain the classification information of the expressions of each frame fusion image corresponding to the first training video as the first discrimination result; Input each frame fusion image corresponding to the second training video into the second discriminator to obtain the second discrimination result of each frame fusion image corresponding to the second training video, including: Input each frame fusion image corresponding to the second training video into the second face feature extraction model in the second discriminator to obtain the feature information of each frame fusion image corresponding to the second training video output; Input the feature information of each frame fusion image corresponding to the second training video into the second expression classification model in the second discriminator to obtain the classification information of the expressions of each frame fusion image corresponding to the second training video as the second discrimination result.
18. The method according to claim 1 or 15, characterized in that, The cycle-consistency loss is determined by the following method: Input each frame fusion image corresponding to the first training video into the second generator to generate the reconstructed images of each frame of the first training video, and input each frame fusion image corresponding to the second training video into the first generator to generate the reconstructed images of each frame of the second training video; Determine the cycle-consistency loss according to the difference between the reconstructed images of each frame of the first training video and the images of each frame of the first training video, and the difference between the reconstructed images of each frame of the second training video and the images of each frame of the second training video.
19. The method according to claim 15, wherein The pixel-to-pixel loss is determined by the following method: For each position in every two adjacent frame fusion images corresponding to the first training video, determine the distance between the representation vectors of the two pixels at this position in the two adjacent frame fusion images, and sum up the distances corresponding to all positions to obtain the first loss; For each position in every two adjacent frame fusion images corresponding to the second training video, determine the distance between the representation vectors of the two pixels at this position in the two adjacent frame fusion images, and sum up the distances corresponding to all positions to obtain the second loss; Sum up the first loss and the second loss to obtain the pixel-to-pixel loss.
20. The method according to claim 1, characterized in that Obtaining the feature information of each frame image of the first training video, the feature information of the facial key points, and the classification information of the original expression includes: Input the frame images in the first training video into the third face feature extraction model in the first generator to obtain the feature information of the output frame images; input the feature information of the frame images into the first face key point detection model in the first generator to obtain the coordinate information of the face key points of the frame images; use the principal component analysis method to reduce the dimension of the coordinate information of all face key points to obtain the first information of a preset dimension as the feature information of the face key points of the frame images in the first training video; input the feature information of the frame images in the first training video into the third expression classification model in the first generator to obtain the classification information of the original expressions of the frame images in the first training video; The obtaining of the feature information of the frame images, the feature information of the face key points, and the classification information of the target expressions of the second training video includes: Input the frame images in the second training video into the fourth face feature extraction model in the second generator to obtain the feature information of the output frame images; input the feature information of the frame images into the second face key point detection model in the second generator to obtain the coordinate information of the face key points of the frame images; use the principal component analysis method to reduce the dimension of the coordinate information of all face key points to obtain the second information of a preset dimension as the feature information of the face key points of the frame images in the second training video; input the feature information of the frame images in the second training video into the fourth expression classification model in the second generator to obtain the classification information of the target expressions of the frame images in the second training video.
21. The method according to claim 1, wherein The fusion of the feature information of the frame images, the feature information of the face key points, the classification information of the original expressions, and the preset classification information corresponding to the target expressions of the first training video includes: Add and average the classification information of the original expressions of the frame images in the first training video and the preset classification information corresponding to the target expressions to obtain the classification information of the fused expressions corresponding to the frame images in the first training video; splice the feature information of the face key points of the frame images in the first training video multiplied by the first weight to be trained, the feature information of the frame images in the first training video multiplied by the second weight to be trained, and the classification information of the fused expressions corresponding to the frame images in the first training video; The fusion of the feature information of the frame images, the feature information of the face key points, the classification information of the target expressions, and the preset classification information corresponding to the original expressions of the second training video includes: Add and average the classification information of the target expressions of the frame images in the second training video and the preset classification information corresponding to the original expressions to obtain the classification information of the fused expressions corresponding to the frame images in the second training video; splice the feature information of the face key points of the frame images in the second training video multiplied by the third weight to be trained, the feature information of the frame images in the second training video multiplied by the fourth weight to be trained, and the classification information of the fused expressions corresponding to the frame images in the second training video.
22. The method according to claim 15, wherein Training the first generator and the second generator according to the adversarial loss, the cycle consistency loss, and the pixel-to-pixel loss includes: Performing a weighted sum of the adversarial loss, the cycle consistency loss, and the pixel-to-pixel loss to obtain a total loss; Training the first generator and the second generator according to the total loss.
23. The method according to claim 8, wherein Editing the human actions in each frame image of the first video according to the human action customization information corresponding to the interaction scenario includes: Adjusting the first human body key points of the human in the original first key frame of the first video during the first action to obtain the second human body key points of the human during the second action, as the human action customization information; Extracting the feature information of the neighborhoods of the respective second human body key points from the original first key frame; Inputting the respective second human body key points and their neighborhood feature information into an image generation model to output the target first key frame of the human during the second action.
24. The method according to claim 23, wherein The method for obtaining the image generation model includes: Using the training video frames and the human body key points of the humans in the training video frames as a pair of training data, using the human body key points in the training data and the feature information of their neighborhoods in the training video frames as the input of an image generation network, and using the training video frames in the training data as the supervision information of the output of the image generation network, and training the image generation network to obtain the image generation model.
25. The method according to claim 23, wherein The first human body key points include the human body contour feature points of the human during the first action, and the second human body key points include the human body contour feature points of the human during the second action.
26. A digital human generation device, comprising: A memory; And a processor coupled to the memory, the processor being configured to execute the digital human generation method according to any one of claims 1-25 based on instructions stored in the memory.
27. A digital life generation device, characterized in that, Including: An acquisition unit configured to acquire a first video; A customization unit configured to edit the humans in each frame image of the first video according to the human customization information corresponding to the interaction scenario, including: editing the human expressions in each frame image of the first video according to the human expression customization information corresponding to the interaction scenario, where the human expression customization information includes the preset classification information corresponding to the target expression; An output unit configured to output a second video according to each frame image of the processed first video, wherein inputting the first video with the original expression of the human face and the preset classification information corresponding to the target expression into an expression generation model, and outputting the second video with the target expression of the human face, and the training method of the expression generation model includes: Obtaining a training pair composed of each frame image of a first training video and each frame image of a second training video; Input the frame images of the first training video into the first generator to obtain the feature information of the frame images of the first training video, the feature information of the facial key points, and the classification information of the original expression. Integrate the feature information of the frame images of the first training video, the feature information of the facial key points, the classification information of the original expression, and the preset classification information corresponding to the target expression to obtain the feature information of the frame fused images corresponding to the first training video. According to the feature information of the frame fused images corresponding to the first training video, obtain the frame fused images corresponding to the first training video output by the first generator; Input the frame images of the second training video into the second generator to obtain the feature information of the frame images of the second training video, the feature information of the facial key points, and the classification information of the target expression. Integrate the feature information of the frame images of the second training video, the feature information of the facial key points, the classification information of the target expression, and the preset classification information corresponding to the original expression to obtain the feature information of the frame fused images corresponding to the second training video. According to the feature information of the frame fused images corresponding to the second training video, obtain the frame fused images corresponding to the second training video output by the second generator; Determine the adversarial loss and the cycle consistency loss according to the frame fused images corresponding to the first training video and the frame fused images corresponding to the second training video; Train the first generator and the second generator according to the adversarial loss and the cycle consistency loss. After the first generator is trained, it is used as an expression generation model.
28. A non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the digital human generation method according to any one of claims 1-25 are implemented.
Citation Information
Patent Citations
Human waist bodybuilding processing method and device in video and electronic equipment
CN111311519A
Method and system for realizing video and audio driven face animation by combining modal particle characteristics
CN112614212A
Virtual character synthesis method and device, equipment and storage medium
CN112967212A