Face dynamic post-processing method, system, computer device and storage medium thereof

Through the combination of face recognition network and convolutional network, the problem of poor visual perception of multiple faces and visual in static face dynamic processing is solved, and a natural dynamic effect is achieved.

CN114241577BActive Publication Date: 2025-08-19SHENZHEN WONDERSHARE SOFTWARE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111612190.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-08-19
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

In the prior art, the dynamic processing of static faces cannot process multiple faces at the same time, and the dynamic images generated are not natural enough and have poor visual perception.

Method used

The pre-trained face recognition network is used to detect face key points and deflection angle information, and the face frame information is cut and rotated, and convolution and channel stitching are combined with the target convolution network to generate natural dynamic images and perform edge processing.

Benefits of technology

The dynamic processing of multiple faces is realized, and the generated dynamic images are more natural and the visual effects are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241577B_ABST
    Figure CN114241577B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, computer device and storage medium for post-processing of face animation. The method comprises: performing face detection on each frame of an image to obtain facial key point information and pitch and deflection information of the face; obtaining facial frame information and cutting it accordingly to obtain an initial face image; calculating the face rotation angle and rotating it to obtain a target face image, splicing it with the generated key point map, and then inputting it into a target convolution network for convolution, performing channel splicing based on the pitch and deflection information of the face to obtain a generated image; reversely rotating the generated image and fitting it to the corresponding frame image and performing edge processing to obtain a target motion video. The present invention uses face detection to obtain all faces and performs convolution processing to obtain corresponding generated images, and then continues edge processing after fitting the generated images to obtain a target motion video with better visual appearance. The overall processing process is more efficient and can simultaneously meet the requirements of multiple face animation processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a face dynamic post-processing method, system, computer equipment and storage medium thereof. Background Art

[0002] Currently, some domestic short video editing software and short video dating software (such as Lico Short Video and Jianying) have the function of making static images dynamic. For faces, the main method of these algorithms is to use face swapping, that is, to extract the face of each frame in the action video, replace or fuse it into the static image separately, and then synthesize the static image into a video, thereby realizing the dynamic function of the static image. Other apps mainly splice 8 different dynamic face videos with static images in different spatial orientations to make the overall video look dynamic. Although existing APP short video can realize the AI static image dynamic function, there is still a big gap compared to the actual needs of users: 1. Users need the face in the static image to reproduce the face in the action video. The method of extracting the face from the action video and pasting it on the static image can retain some facial features in the original static image by using face fusion, but the main body of the face is obviously changed from the visual perspective, which is far from the actual needs of users. 2. Splicing different facial motion videos together spatially with the face to generate a video only makes the overall video appear animated, but the static images in between are not animated, which falls far short of actual user needs. 3. The current algorithm primarily animates single static faces and cannot support animating multiple faces within a static image. 4. Because the animated faces are generated using a GAN, directly pasting them back into the original image results in a poor visual experience. Summary of the Invention

[0003] The embodiments of the present invention provide a face animation post-processing method, system, computer device and storage medium thereof, aiming to solve the problems in the prior art that the animation process of static faces cannot animate multiple faces at the same time, and the animated images are not natural enough and have poor visual perception.

[0004] In a first aspect, an embodiment of the present invention provides a method for post-processing a dynamic face, comprising:

[0005] Use the pre-trained face recognition network to detect faces in each frame of the initial motion video, and obtain the facial key point information and the face pitch angle information, rotation angle information, and deflection angle information in each frame;

[0006] Acquire facial frame information according to the facial key point information, and cut the corresponding frame image according to the facial frame information to obtain an initial facial image;

[0007] Calculating the face rotation angle in the initial face image according to the face key point information and rotating the initial face image to obtain a target face image;

[0008] Generate a key point map corresponding to each frame of the initial motion video, and splice the key point map with the target face image and input it into the target convolutional network for convolution. Based on the pitch angle information, rotation angle information, and deflection angle information of the face, the convolution results in the target convolutional network are channel-spliced to obtain a generated image.

[0009] The generated image is reversely rotated according to the face rotation angle and attached to the corresponding frame image, and each attached frame image is edge processed to obtain the final target motion video.

[0010] In a second aspect, an embodiment of the present invention provides a face dynamic post-processing system, comprising:

[0011] A face detection unit is used to perform face detection on each frame of the initial motion video using a pre-trained face recognition network, and obtain facial key point information and facial depression angle information, rotation angle information, and deflection angle information in each frame;

[0012] An initial face image acquisition unit is used to acquire face frame information according to the face key point information, and cut the corresponding frame image according to the face frame information to obtain an initial face image;

[0013] a target face image acquisition unit, configured to calculate a face rotation angle in the initial face image based on the face key point information and rotate the initial face image to obtain a target face image;

[0014] A generated image acquisition unit is used to generate a key point map corresponding to each frame of the initial motion video, and to splice the key point map with the target face image and input the resulting image into a target convolutional network for convolution. The convolution results in the target convolutional network are then channel-spliced based on the pitch angle information, rotation angle information, and deflection angle information of the face to obtain a generated image.

[0015] The target motion video acquisition unit is used to reversely rotate the generated image according to the face rotation angle and fit it to the corresponding frame image, and perform edge processing on each frame image after fitting to obtain the final target motion video.

[0016] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the face dynamic post-processing method described in the first aspect is implemented.

[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the face dynamic post-processing method described in the first aspect above.

[0018] The embodiment of the present invention provides a face dynamic post-processing method, system, computer device and storage medium thereof, the method comprising: using a pre-trained face recognition network to perform face detection on each frame of an initial motion video, obtaining face key point information and face depression angle information, rotation angle information and deflection angle information in each frame; obtaining face frame information based on the face key point information, cutting the corresponding frame image based on the face frame information to obtain an initial face image; calculating the face rotation angle in the initial face image based on the face key point information and performing face rotation on the face frame image; The initial facial image is rotated to obtain a target facial image; a key point map corresponding to each frame of the initial motion video is generated, and the key point map is spliced with the target facial image and then input into a target convolutional network for convolution. The convolution results in the target convolutional network are channel-spliced based on the pitch angle information, rotation angle information, and deflection angle information of the face to obtain a generated image; the generated image is reversely rotated according to the face rotation angle and attached to the corresponding frame image, and each attached frame image is edge processed to obtain a final target motion video. In this embodiment of the present invention, face detection is first performed to obtain facial key point information and pitch, rotation, and deflection angle information, the face is cut using the facial key point information and convolution processing is performed, and the final generated image is obtained by combining the pitch, rotation, and deflection angle information. Finally, the generated image is attached back to the corresponding frame image and edge processed to obtain the target motion video. The overall processing process is more efficient and can simultaneously meet the dynamic processing requirements of multiple faces in a static image. The dynamic image is more natural, and the visual appearance of the video is improved due to the edge processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 A schematic flow chart of a face dynamic post-processing method provided by an embodiment of the present invention;

[0021] Figure 2 A schematic diagram of a face recognition network framework for a face dynamic post-processing method provided by an embodiment of the present invention;

[0022] Figure 3 A schematic diagram of the target convolutional network framework of the face dynamic post-processing method provided by an embodiment of the present invention;

[0023] Figure 4 A schematic diagram of a network framework for obtaining true value probabilities in the face dynamic post-processing method provided by an embodiment of the present invention;

[0024] Figure 5 A schematic block diagram of a face animation post-processing system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0027] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0028] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0029] See also Figure 1 , Figure 1 A flowchart of a face dynamic post-processing method provided by an embodiment of the present invention includes steps S101 to S105.

[0030] S101, using a pre-trained face recognition network to perform face detection on each frame of the initial motion video, and obtain facial key point information and facial depression angle information, rotation angle information, and deflection angle information in each frame;

[0031] In this step, each frame of the initial motion video is input into the pre-trained face recognition network for face detection, and the facial key points and corresponding facial key point information in each frame, as well as the face's pitch angle information, rotation angle information, and deflection angle information are extracted.

[0032] In one embodiment, step S101 includes:

[0033] Inputting each frame of the initial motion video into two consecutive first convolutional layers for convolution processing in sequence to obtain a first convolution result;

[0034] The first convolution result is input into three consecutive second convolution layers for convolution processing to obtain a second convolution result, and the second convolution result is input into a fully connected layer for convolution to obtain facial key point information and facial depression angle information, rotation angle information and deflection angle information.

[0035] In this embodiment, if Figure 2 As shown, each frame of the initial motion video (that is, each frame of the initial motion video as Figure 2 The input image in ( ) is input into two consecutive first convolutional layers of size 3*3*64 for convolution processing, and then input into three consecutive second convolutional layers of size 3*3*256 for convolution processing, and finally input into the fully connected layer for convolution, thereby obtaining 68 facial key points and corresponding facial key point information, as well as the face's depression angle information, rotation angle information and deflection angle information. The overall structure of the face recognition network consists of 5 convolutional layers and 1 fully connected layer, including two consecutive first convolutional layers and three consecutive second convolutional layers. Each frame of the picture input into the face recognition network is convolved through the first convolutional layer and the second convolutional layer in turn, and then input into the fully connected layer for convolution, thereby outputting the probability of the presence of a face in the frame. If the probability of the presence of a face is greater than a preset threshold, it is determined that the frame has a face and outputs the 68 facial key points corresponding to the face and the face's depression angle information, rotation angle information and deflection angle information. The face recognition network is trained based on the loss function, specifically including face classification loss L2, key point loss L3, and face pitch angle loss L4, wherein, is the final output face, p i is the real graph; L3=∑∑|y ij -y ij T |,y ij T is the real key point label coordinate, y ijis the key point coordinate output by the face recognition network, where i∈[1,68], j=1,2 (j=1 is the horizontal coordinate, j=2 is the vertical coordinate); L4=∑∑|y i -y i T |, where y i T is the true angle of the face, y i is the output angle of the face recognition network, where i=1, 2, and 3 represent the face's depression angle, rotation angle, and deflection angle, respectively. Then the loss function L of the face recognition network is f =a*L2+b*L3+c*L4, where a=1, b=1, c=1, and gradient descent and back propagation are used to train the face recognition network.

[0036] S102, acquiring face frame information according to the face key point information, and cutting the corresponding frame image according to the face frame information to obtain an initial face image;

[0037] In this step, the face frame information is determined based on the face key point information detected by the face recognition network, and the face position is calculated based on the face frame information, and then cut to obtain an initial face image.

[0038] In one embodiment, step S102 includes:

[0039] Acquire coordinate information of each facial key point according to the facial key point information, and select a minimum facial key point with the smallest coordinate value and a maximum facial key point with the largest coordinate value from the coordinate information of the facial key points;

[0040] The width and height of the face frame are calculated according to the minimum face key point and the maximum face key point, and the corresponding frame image is cut based on the coordinate information of the minimum face key point and the width and height of the face frame to obtain an initial face image.

[0041] In this embodiment, based on the acquired facial key point information, the minimum facial key point corresponding to the minimum coordinate information of the facial key point and the maximum facial key point corresponding to the maximum coordinate information are screened out to confirm the facial frame information, and then the width and height of the facial frame are calculated based on the minimum facial key point and the maximum facial key point. The corresponding frame picture is cut based on the width and height of the facial frame to obtain the initial facial image.

[0042] Specifically, confirm the minimum facial key point (X min , Y min ), and the largest facial landmark (X max , Y max ), where X min、Y min Respectively expressed as the minimum X coordinate and minimum Y coordinate of all key points, X max 、Y max Represents the maximum X coordinate and maximum Y coordinate of all key points respectively. Then the height of the face frame H = Y max -Y min , width W = X max -X min .

[0043] The sample face image is cropped based on the width and height of the face frame. The width and height of the face frame are first compensated, and then the width and height of the initial face image are calculated based on the compensation results. Specifically, the width of the initial face image is calculated according to the following formula: w = W + offsetW, where w is the width of the initial face image, W is the width of the face frame, and offsetW is the width compensation result. The height of the initial face image is calculated according to the following formula: h = H + offsetH, where h is the height of the initial face image, H is the height of the face frame, and offsetH is the height compensation result. In this embodiment, offsetW is 3 / 8W, and offsetH is 4 / 9H.

[0044] S103, calculating the face rotation angle in the initial face image according to the face key point information and rotating the initial face image to obtain a target face image;

[0045] In this step, the coordinates of the left eye key point and the right eye key point in the facial key points are obtained, and the face rotation angle is calculated based on the coordinates of the left eye key point and the right eye key point, and the target face image is rotated. Specifically, the coordinates of the left eye key point are set to (x1, y1) and the coordinates of the right eye are set to (x2, y2). Then, the rotation angle of the face can be calculated as angle = arctan ((y1-y2) / (x1-x2)). Then, the bilinear interpolation method is used to rotate the initial face image according to the face rotation angle angle to obtain the target face image.

[0046] S104: generating a key point map corresponding to each frame of the initial motion video, splicing the key point map with the target face image, and inputting the resulting image into a target convolutional network for convolution. Channel splicing is performed on the convolution results in the target convolutional network based on the pitch angle information, rotation angle information, and deflection angle information of the face to obtain a generated image.

[0047] In this step, a key point map corresponding to each frame of the initial motion video is generated and spliced with the corresponding target face image. The spliced image is then input into the target convolutional network for convolution. The convolution result is then channel-spliced with the face's pitch angle information, rotation angle information, and deflection angle information to obtain a generated image.

[0048] In one embodiment, the step S104 includes:

[0049] Generate an all-zero-value image of the same size as each frame of the initial motion video, and set the pixel values corresponding to the key points in the all-zero-value image to a specified size to obtain a key point map;

[0050] splicing the key point map with the target face image, and inputting the convolution operation into a plurality of consecutive third convolution layers to obtain a third convolution result, and inputting the third convolution result into a plurality of consecutive fourth convolution layers to obtain a fourth convolution result;

[0051] The fourth convolution result is channel-spliced with the depression angle information, rotation angle information and deflection angle information of the face, and the spliced fourth convolution result is input into multiple consecutive upsampling layers for upsampling to obtain a generated image.

[0052] In this embodiment, if Figure 3 As shown, first generate an all-zero value image Image_keypoint with the same size as each frame image, then set the pixel value corresponding to the key point in Image_keypoint to 255 to obtain a key point map, and then splice the key point map with the corresponding target face image, and the spliced image (the spliced image is Figure 3 The input image shown in FIG1 is input into two consecutive third convolutional layers of size 3*3*64 and stride 1 for convolution processing, and then continues to be input into three consecutive fourth convolutional layers of size 3*3*256 and stride 2 for convolution processing to obtain the fourth convolution result. At this time, the depression angle information, rotation angle information and deflection angle information of the face are channel-joined with the fourth convolution result, and then the fourth convolution result after channel splicing is input into three consecutive upsampling layers of size 3*3*256 and stride 2 for upsampling to obtain the generated image (the generated image is Figure 3 Output image shown).

[0053] In one embodiment, after performing channel stitching on the convolution results in the target convolution network based on the depression angle information, rotation angle information, and deflection angle information of the face to obtain the generated image, the following steps are included:

[0054] The generated image is input into multiple convolution layers with convolution kernels of 3*3*64 for convolution, and the convolution results are input into multiple convolution layers with convolution kernels of 3*3*256 for convolution, and the convolution results are input into the average pooling layer for pooling to obtain the true value probability of the generated image.

[0055] In this embodiment, if Figure 4 As shown, the generated image (ie, the generated image as Figure 4 The input image shown in FIG1 is input into two consecutive convolutional layers of size 3*3*64 for convolution, and then input into three consecutive convolutional layers of size 3*3*256 for convolution, and finally input into the average pooling layer for pooling to obtain the true value probability of the generated image, thereby realizing supervised training of the target convolutional network.

[0056] In one embodiment, the channel stitching of the convolution results in the target convolution network based on the depression angle information, rotation angle information, and deflection angle information of the face to obtain the generated image includes:

[0057] Input the generated image and the corresponding frame picture of the initial motion video into the trained VGG19 network for convolution processing, obtain the output results of each layer of the VGG19 network, and calculate the L1 loss function based on the output results of each layer of the VGG19 network;

[0058] The target convolutional network is trained by gradient descent back propagation using the L1 loss function to obtain an optimized target convolutional network.

[0059] In this embodiment, the generated image and the corresponding frame of the initial motion video are input into the VGG19 network for convolution processing, the output value of each layer in the VGG19 network is obtained, and the L1 difference is used as the loss function. The target convolution network is trained by gradient descent back propagation using the L1 loss function. Let VGGn(f) and VGGn(g) represent the output of the nth layer r and g of VGG19 respectively, then: L vgg_loss =∑|VGG n (r)-VGG n (g)|

[0060] S105 , reversely rotating the generated image according to the face rotation angle and fitting it to the corresponding frame image, and performing edge processing on each frame image after fitting to obtain the final target motion video.

[0061] In this step, the generated image corresponding to each frame is rotated inversely according to the face rotation angle and then attached to the corresponding frame. The edge processing is then performed on the attached part of each frame to obtain the final target motion video.

[0062] In one embodiment, performing edge processing on each frame of the assembled image to obtain the final target motion video includes:

[0063] Obtain the edge blocks of each frame after the alignment, and perform linear superposition processing from the outside to the inside and from the inside to the outside on each edge block;

[0064] Edge Gaussian filtering is performed on the edge blocks after the linear superposition processing to obtain the final target motion video; wherein the Gaussian kernel size of the Gaussian filtering is 3*3 and the standard deviation is 0.

[0065] In this embodiment, the fitting edges of each frame of the image are processed to obtain the upper edge block, lower edge block, left edge block, and right edge block of each frame of the image. Each edge block is first subjected to linear superposition processing from the outside to the inside, and then to linear superposition processing from the inside to the outside, and finally to edge Gaussian filtering processing to obtain the final target motion video.

[0066] When performing linear superposition processing from outside to inside on each edge block, taking the edge block above as an example, it is assumed that the face edge block has 10 rows and N columns. Rows 0 to 4 are outside, and rows 5 to 9 are inside.

[0067] Then the pixel values of rows 5, 6, 7, 8, and 9 are:

[0068] Row5=0.8*Row4+0.2*Row5

[0069] Row6=0.6*Row4+0.4*Row6

[0070] Row7=0.4*Row4+0.6*Row7

[0071] Row8=0.2*Row4+0.8*Row8

[0072] Row9=0.0*Row4+1.0*Row9

[0073] When performing linear superposition processing on each edge block from the inside out, taking the edge block above as an example, assuming that the face edge block has 10 rows and N columns, the pixel values of 3, 2, 1, and 0 are:

[0074] Row3=0.75*Row4+0.25*Row3

[0075] Row2=0.50*Row4+0.50*Row2

[0076] Row1=0.25*Row4+0.75*Row1

[0077] Row0=0.0*Row4+1.0*Row0

[0078] After completing the inner linear superposition processing from outside to inside and the linear superposition processing from inside to outside, edge Gaussian filtering is performed on the upper edge block, lower edge block, left edge block, and right edge block, where the Gaussian kernel size of the Gaussian filter is 3*3 and the standard deviation is 0.

[0079] See also Figure 5 , Figure 5 A schematic block diagram of a face animation post-processing system provided in an embodiment of the present invention, the face animation post-processing system 200 includes:

[0080] The face detection unit 201 is used to perform face detection on each frame of the initial motion video using a pre-trained face recognition network, and obtain facial key point information and facial depression angle information, rotation angle information, and deflection angle information in each frame;

[0081] An initial face image acquisition unit 202 is configured to acquire face frame information based on the face key point information, and cut the corresponding frame image based on the face frame information to obtain an initial face image;

[0082] a target facial image acquisition unit 203, configured to calculate a facial rotation angle in the initial facial image based on the facial key point information and rotate the initial facial image to obtain a target facial image;

[0083] A generated image acquisition unit 204 is configured to generate a key point map corresponding to each frame of the initial motion video, and to concatenate the key point map with the target face image, inputting the convolutional network into the target convolutional network for convolution. Furthermore, the convolutional results in the target convolutional network are channel-concatenated based on the pitch, rotation, and deflection information of the face to obtain a generated image.

[0084] The target motion video acquisition unit 205 is configured to reversely rotate the generated image according to the face rotation angle and fit it to the corresponding frame image, and perform edge processing on each frame image after fitting to obtain the final target motion video.

[0085] In one embodiment, the face detection unit 201 includes:

[0086] a first convolution result obtaining unit, configured to sequentially input each frame of the initial motion video into two consecutive first convolution layers for convolution processing to obtain a first convolution result;

[0087] A key point information acquisition unit is used to input the first convolution result into three consecutive second convolution layers for convolution processing to obtain a second convolution result, and input the second convolution result into a fully connected layer for convolution to obtain facial key point information and the face's depression angle information, rotation angle information, and deflection angle information.

[0088] In one embodiment, the initial face image acquisition unit 202 includes:

[0089] a facial key point maximum value obtaining unit, configured to obtain coordinate information of each facial key point according to the facial key point information, and to select a minimum facial key point with the smallest coordinate value and a maximum facial key point with the largest coordinate value from the coordinate information of the facial key points;

[0090] A facial image cutting unit is used to calculate the width and height of the facial frame according to the minimum facial key point and the maximum facial key point, and to cut the corresponding frame image based on the coordinate information of the minimum facial key point and the width and height of the facial frame to obtain an initial facial image.

[0091] In one embodiment, the image acquisition unit 204 includes:

[0092] a key point map acquisition unit, configured to generate an all-zero-value image of the same size as each frame of the initial motion video, and set the pixel values corresponding to the key points in the all-zero-value image to a specified size to obtain a key point map;

[0093] a key point map splicing unit, configured to splice the key point map with the target face image, input the spliced key point map into a plurality of consecutive third convolutional layers for convolution to obtain a third convolution result, and input the third convolution result into a plurality of consecutive fourth convolutional layers for convolution to obtain a fourth convolution result;

[0094] An upsampling processing unit is used to perform channel splicing on the fourth convolution result with the depression angle information, rotation angle information and deflection angle information of the face, and input the spliced fourth convolution result into multiple consecutive upsampling layers for upsampling to obtain a generated image.

[0095] In one embodiment, the image generation unit 204 includes:

[0096] A true value probability acquisition unit is used to input the generated image into multiple convolution layers with convolution kernels of 3*3*64 for convolution, and input the convolution results into multiple convolution layers with convolution kernels of 3*3*256 for convolution, and input the convolution results into the average pooling layer for pooling to obtain the true value probability of the generated image.

[0097] In one embodiment, the image acquisition unit 204 includes:

[0098] An L1 loss function calculation unit is used to input the generated image and the corresponding frame picture of the initial motion video into a trained VGG19 network for convolution processing, obtain the output result of each layer of the VGG19 network, and calculate the L1 loss function according to the output result of each layer of the VGG19 network;

[0099] The target convolutional network optimization unit is used to perform gradient descent back propagation training on the target convolutional network using the L1 loss function to obtain an optimized target convolutional network.

[0100] In one embodiment, the target motion video acquisition unit 205 includes:

[0101] A linear superposition processing unit is used to obtain edge blocks of each frame of the assembled image, and perform linear superposition processing from outside to inside and from inside to outside on each edge block;

[0102] The edge Gaussian filter processing unit is used to perform edge Gaussian filtering on the edge blocks after linear superposition processing to obtain the final target motion video; wherein the Gaussian kernel size of the Gaussian filter is 3*3 and the standard deviation is 0.

[0103] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned face dynamic post-processing method is implemented.

[0104] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for post-processing face animation as described above is implemented.

[0105] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

[0106] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

Claims

1. A face dynamic post-processing method, characterized in that: include: Use the pre-trained face recognition network to detect faces in each frame of the initial motion video, and obtain the facial key point information and the face pitch angle information, rotation angle information, and deflection angle information in each frame; Acquire facial frame information according to the facial key point information, and cut the corresponding frame image according to the facial frame information to obtain an initial facial image; Calculating the face rotation angle in the initial face image according to the face key point information and rotating the initial face image to obtain a target face image; Generate a key point map corresponding to each frame of the initial motion video, and splice the key point map with the target face image and input it into the target convolutional network for convolution. Based on the pitch angle information, rotation angle information, and deflection angle information of the face, the convolution results in the target convolutional network are channel-spliced to obtain a generated image. The generated image is reversely rotated according to the face rotation angle and attached to the corresponding frame image, and each attached frame image is edge processed to obtain the final target motion video; The method generates a key point map corresponding to each frame of the initial motion video, splices the key point map with the target face image, and inputs the spliced map into a target convolution network for convolution, and performs channel splicing on the convolution results in the target convolution network based on the pitch angle information, rotation angle information, and deflection angle information of the face to obtain a generated image, including: generating an all-zero-value image with the same size as each frame of the initial motion video, and setting the pixel values corresponding to the key points in the all-zero-value image to a specified size to obtain a key point map; splicing the key point map with the target face image, and inputting the convolution into multiple consecutive third convolution layers for convolution to obtain a third convolution result, and inputting the third convolution result into multiple consecutive fourth convolution layers for convolution to obtain a fourth convolution result; channel splicing the fourth convolution result with the pitch angle information, rotation angle information, and deflection angle information of the face, and inputting the spliced fourth convolution result into multiple consecutive upsampling layers for upsampling to obtain a generated image; The edge processing of each of the assembled frames to obtain the final target motion video includes: obtaining edge blocks of each of the assembled frames, and performing linear superposition processing from the outside to the inside and from the inside to the outside on each edge block; performing edge Gaussian filtering on the edge blocks after the linear superposition processing to obtain the final target motion video; wherein the Gaussian kernel size of the Gaussian filtering is 3*3 and the standard deviation is 0; The convolution results in the target convolution network are channel-stitched based on the depression angle information, rotation angle information, and deflection angle information of the face to obtain a generated image. The method further includes: inputting the generated image into multiple convolution layers with convolution kernels of 3*3*64 for convolution, and inputting the convolution results into multiple convolution layers with convolution kernels of 3*3*256 for convolution, and inputting the convolution results into the average pooling layer for pooling to obtain the true value probability of the generated image.

2. The face dynamic post-processing method according to claim 1, characterized in that: The method of using a pre-trained face recognition network to perform face detection on each frame of the initial motion video to obtain facial key point information and facial depression angle information, rotation angle information, and deflection angle information in each frame includes: Inputting each frame of the initial motion video into two consecutive first convolutional layers for convolution processing in sequence to obtain a first convolution result; The first convolution result is input into three consecutive second convolution layers for convolution processing to obtain a second convolution result, and the second convolution result is input into a fully connected layer for convolution to obtain facial key point information and facial depression angle information, rotation angle information and deflection angle information.

3. The face dynamic post-processing method according to claim 1, characterized in that: The step of obtaining face frame information according to the face key point information and cutting the corresponding frame image according to the face frame information to obtain an initial face image includes: Acquire coordinate information of each facial key point according to the facial key point information, and select a minimum facial key point with the smallest coordinate value and a maximum facial key point with the largest coordinate value from the coordinate information of the facial key points; The width and height of the face frame are calculated according to the minimum face key point and the maximum face key point, and the corresponding frame image is cut based on the coordinate information of the minimum face key point and the width and height of the face frame to obtain an initial face image.

4. The face dynamic post-processing method according to claim 1, characterized in that: The convolution results in the target convolution network are channel-stitched based on the depression angle information, rotation angle information, and deflection angle information of the face to obtain a generated image, including: Input the generated image and the corresponding frame picture of the initial motion video into the trained VGG19 network for convolution processing, obtain the output results of each layer of the VGG19 network, and calculate the L1 loss function based on the output results of each layer of the VGG19 network; The target convolutional network is trained by gradient descent back propagation using the L1 loss function to obtain an optimized target convolutional network.

5. A face dynamic post-processing system, characterized in that: include: A face detection unit is used to perform face detection on each frame of the initial motion video using a pre-trained face recognition network, and obtain facial key point information and facial depression angle information, rotation angle information, and deflection angle information in each frame; An initial face image acquisition unit is used to acquire face frame information according to the face key point information, and cut the corresponding frame image according to the face frame information to obtain an initial face image; a target face image acquisition unit, configured to calculate a face rotation angle in the initial face image based on the face key point information and rotate the initial face image to obtain a target face image; A generated image acquisition unit is used to generate a key point map corresponding to each frame of the initial motion video, and to splice the key point map with the target face image and input the resulting image into a target convolutional network for convolution. The convolution results in the target convolutional network are then channel-spliced based on the pitch angle information, rotation angle information, and deflection angle information of the face to obtain a generated image. a target motion video acquisition unit, configured to reversely rotate the generated image according to the face rotation angle and fit it to the corresponding frame image, and perform edge processing on each frame image after fitting to obtain the final target motion video; The generated image acquisition unit includes: a key point map acquisition unit, configured to generate an all-zero-value image of the same size as each frame of the initial motion video, and set the pixel values corresponding to the key points in the all-zero-value image to a specified size to obtain a key point map; a key point map splicing unit, configured to splice the key point map with the target face image, input the spliced key point map into a plurality of consecutive third convolutional layers for convolution to obtain a third convolution result, and input the third convolution result into a plurality of consecutive fourth convolutional layers for convolution to obtain a fourth convolution result; an upsampling processing unit, configured to perform channel splicing on the fourth convolution result and the depression angle information, rotation angle information, and deflection angle information of the face, and input the spliced fourth convolution result into multiple consecutive upsampling layers for upsampling to obtain a generated image; The target motion video acquisition unit includes: A linear superposition processing unit is used to obtain edge blocks of each frame of the assembled image, and perform linear superposition processing from outside to inside and from inside to outside on each edge block; An edge Gaussian filter processing unit is used to perform edge Gaussian filter processing on the edge blocks after the linear superposition processing to obtain the final target motion video; wherein the Gaussian kernel size of the Gaussian filter is 3*3 and the standard deviation is 0; The generated image acquisition unit includes: A true value probability acquisition unit is used to input the generated image into multiple convolution layers with convolution kernels of 3*3*64 for convolution, and input the convolution results into multiple convolution layers with convolution kernels of 3*3*256 for convolution, and input the convolution results into the average pooling layer for pooling to obtain the true value probability of the generated image.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the face animation post-processing method according to any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to perform the face animation post-processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Head attitude angle detection method and device, electronic equipment and storage medium

    CN112668480A

  • Static face dynamic method and system, computer equipment and readable storage medium

    CN113688753A

  • Multi-angle side face frontage method based on feature mapping

    CN113705358A