A human posture transfer method and system based on motion drive
By using high-precision key point detection model and perceived loss function in human posture migration, the problem of unstable key point detection and insufficient character clarity in high-resolution videos is solved, and a more stable and clear pose migration effect is achieved.
Patent Information
- Application Number
- CN202111304351.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-05
AI Technical Summary
The prior art has problems of unstable detection of key points and insufficient clarity of characters in high-resolution videos in human posture migration.
A high-precision human key point detection model is adopted, and a perceived loss function based on the target human body is introduced in model training, combined with the discriminator loss function, improve the stability and clarity of the model's posture migration of human body.
It improves the stability of human posture migration and the clarity of characters in high-resolution videos, enhancing the migration effect.
Smart Images

Figure CN114049652B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image synthesis processing technology, and in particular to a method and system for human body posture migration based on action drive. Background Art
[0002] Human posture transfer is to generate a new human posture video given a target person image and a human posture video, so that the target person in the generated video is the person in the given image, and the human posture in the generated video is consistent with the human posture in the given video. This process is called posture transfer. With the development of self-media, the application of posture transfer is becoming more and more extensive. The existing posture transfer generally obtains the key point information of the image first, obtains the mapping relationship between the transferred object and the source object based on the key point information, and finally outputs the transferred result through the model, such as the patent application with publication number CN111598977A, named a method and system for expression transfer and animation.
[0003] The prior art still has the following deficiencies:
[0004] 1. The key points of the human body are obtained through training, and a large amount of data training is required to obtain stable results. At the same time, when the human body parts in the video are blocked, the detected key positions are not stable, which will affect the relative mapping relationship between the key points, resulting in incoherent character movements. Especially when the size of the generated image is large, the incoherence of the character's movements is more obvious.
[0005] 2. The existing technology generally uses reconstruction loss as the loss function, that is, the overall image contrast is used as the loss function; when the character's action is transferred, the overall clarity of the character will not be additionally enhanced during the training process, so that when the output video resolution is high, the character is relatively blurred as a whole, and the blurring of the limbs and limb edges performing the action is more obvious. Summary of the invention
[0006] To solve the above problems, the present invention provides a method and system for human posture transfer based on action-driven, proposes a high-precision human key point detection model, explicitly separates the human key point detection module in the model training and inference process, and increases the stability of human posture transfer; at the same time, a perceptual loss function based on the target human body is added during the training process, so that the model focuses on the clarity of the human body after transfer and reconstruction, increases the clarity of the characters in the image at high resolution, and improves the effect of transfer.
[0007] The present invention provides a method for human posture migration based on action drive, and the specific technical scheme is as follows:
[0008] S1: Acquire human action video data, extract image frames of the video data to obtain a number of continuous pictures, and screen the extracted pictures to obtain a target picture;
[0009] S2: Detect key points of the human body on the target image to obtain the coordinates of the key points;
[0010] S3: randomly extracting two images from the target image as a source image and a driving image respectively, and calculating a transformation relationship between the driving image and the source image according to the obtained key point coordinates;
[0011] S4: input the transformation relationship into the motion estimation model, and output the corresponding optical flow map and redrawing map;
[0012] S5: inputting the optical flow map and the redrawing map into an action generation model to obtain a posture transfer generation image;
[0013] S6: Based on the driving image and the posture transfer generation image, a loss function is calculated. The specific process is as follows:
[0014] S601: Calculate the discriminator loss function L through the discriminator network model D D ;
[0015] S602: Calculate the perceptron loss function L through the human recognition model I ;
[0016] S603: Combine the discriminator loss function with the perceptron loss function, and output a final overall loss function L.
[0017] Furthermore, the human body motion video data is single person motion video data, and the screening is to delete incomplete human body video data.
[0018] Furthermore, step S2 can also re-screen the image through human key point detection, and delete the data where human key points cannot be detected or key points of multiple people are detected.
[0019] Furthermore, in step S5, the action generation model adopts a generative adversarial network, and the specific process of obtaining the image generated by posture transfer is as follows:
[0020] Input the source image into the action generation model to obtain the hidden feature map of the source image and concatenate it with the optical flow map;
[0021] The obtained splicing result is multiplied with the redrawing image, and the result output after the multiplication is input into the decoder of the model to output the posture migration generated image.
[0022] Furthermore, in step S601, the discriminator network model adopts the VGG16 model, and the loss function adopts the cross entropy loss function, and the formula is as follows:
[0023] L D =-ylog(D(x))-(1-y)log(1-D(x))
[0024] Among them, x is the input image and y is the image label.
[0025] Furthermore, in step S602, the perceptron loss function L I The calculation process is as follows:
[0026] Extracting model hidden features of the human body in the posture transfer generated image and the driving image through a human body recognition model;
[0027] The feature difference between the posture transfer generated image and the corresponding extracted driving image is calculated as the perceptron loss function, and the formula is as follows:
[0028] L J =||J(D g )-J(Q)||
[0029] Where D g Generates the image for posture transfer, and Q is the driving image.
[0030] Furthermore, the feature difference is the distance between the hidden feature vector of the last layer of the model when the posture transfer generated image and the driving image are input into the model.
[0031] Furthermore, in step S603, the perceptron loss function and the discriminator loss function are combined as follows:
[0032] L=w1L J +w2L D
[0033] Where w1, w2 are the weight coefficients of each loss function.
[0034] The present invention also provides a human posture migration system, the system comprising a data module, an action estimation module, an action generation module, and a loss function module;
[0035] The data module is used to collect human action videos and randomly extract image frames to obtain source images, drive image data and corresponding human key point coordinates;
[0036] The action estimation module is connected to the data module, and is used to receive the source image, the driving image and the coordinate data of the key points of the human body, and output the optical flow field and the redrawing map;
[0037] The action generation module is connected to the action estimation module and the data module, and is used to receive the optical flow map, the redrawing map and the source image, splice the hidden layer feature map of the source image with the optical flow map, and multiply the splicing result with the redrawing map, and finally output the posture migration generated image;
[0038] The loss function module is connected to the action generation module and the data module, and is used to receive the driving image and the generated image, calculate the perceptron loss function and the discriminator loss function, and combine and output the overall loss function.
[0039] The beneficial effects of the present invention are as follows:
[0040] 1. The character action video is used to drive the source character image to obtain the character posture transfer video. Based on the source image and the driving image, a high-precision key point human key point detection model is used to obtain the coordinates of the human key points, thereby increasing the stability of the human posture transfer and reducing the amount of data for the model to learn the key point information.
[0041] 2. During the model training process, the perception loss function of the target human body is combined with the discriminator loss function to form the final loss function, so that the model pays attention to the human body information and improves the clarity of the human body after migration and reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the structure of the model of the present invention;
[0043] Figure 2 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0044] The following description clearly and completely describes the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] The technical content of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] Example 1
[0047] The embodiment of the present invention provides a method for human body posture migration based on action drive, such as Figure 2 As shown, the method comprises the following steps:
[0048] S1: Acquire human action video data, extract image frames of the video data to obtain a number of continuous pictures, and screen the extracted pictures to obtain a target picture;
[0049] The human body motion video data is single-person motion video data, and is not limited to specific types of motion. The video preferably has a resolution of 1080P or above to ensure that the human body occupies most of the video and is complete. The screening is to delete incomplete video data of the human body.
[0050] S2: Detect key points of the human body on the target image to obtain the coordinates of the key points;
[0051] Perform human body detection on the target object through a human body key point detection model to obtain key point information. In this embodiment, a model with higher accuracy, such as Densepose, Mediapipe, etc., is used for key point detection;
[0052] In order to speed up subsequent training, key point detection can be performed on the data in advance, and the key point detection results can be synchronously input during training. Key point detection can also be used to perform initial screening and preprocessing on the data, deleting data where human key points cannot be detected or where multiple key points are detected, and then cropping is performed based on the detected human key points to delete data where the human body area accounts for too small a proportion. If the detection effect of the human key point model deviates on individual data, the key point position coordinates can be manually corrected to ensure the accuracy of the data.
[0053] S3: randomly extracting two images from the target image as a source image and a driving image respectively, and calculating a transformation relationship between the driving image and the source image according to the obtained key point coordinates;
[0054] The key point coordinates k of the source image o , the key point coordinates in the driving image are k t , then k o to k t The affine transformation of is:
[0055]
[0056] Among them, A, is the affine transformation parameter, A is the linear mapping matrix, is the translation parameter. Solving the above matrix gives A, Each pair of key points has a corresponding set of affine transformation parameters.
[0057] S4: input the transformation relationship into the motion estimation model, and output the corresponding optical flow map and redrawing map;
[0058] In this embodiment, the motion estimation model is composed of basic neural network structures such as convolutional layers, fully connected layers, activation layers, pooling layers, and normalization layers. The specific structure can adopt a UNet network structure or other Encoder-Decoder model structures.
[0059] The input of the motion estimation model is the driving image and the affine transformation parameters between the key points. After calculations by the network's convolutional layer, fully connected layer, activation layer, pooling layer, normalization layer, etc., the final output is the optical flow map L and the redrawing map M; the optical flow map represents the transformation relationship between the driving image and each pixel of the source image, and the redrawing map represents the area that needs to be redrawn.
[0060] S5: inputting the optical flow map and the redrawing map into an action generation model to obtain a posture transfer generation image;
[0061] The action generation model adopts a generative adversarial network and is composed of a super-resolution model, which includes an encoder E and a decoder G; wherein the input of the encoder E is the driving image, and the output is the hidden layer feature f of the driving image. E , concatenate the optical flow map L with the output features of the encoder E, and then multiply it with the remap map M to get the input of the decoder G, and finally generate the migrated image D through the decoder g , the formula is as follows:
[0062] D g =G(M⊙(E(S),L))
[0063] S6: Based on the driving image and the posture transfer generation image, a loss function is calculated. The specific process is as follows:
[0064] S601: Calculate the discriminator loss function L through the discriminator network model D D ;
[0065] The discriminator network model is a binary classification neural network model used to judge the authenticity of an input image. In this embodiment, the VGG16 model is selected as the network model of the discriminator, and the loss function adopts the cross entropy loss function.
[0066] The specific formula is as follows:
[0067] L D =-ylog(D(x))-(1-y)log(1-D(x))
[0068] Where x is the input image and y is the image label; if the image is the original image, y is 1; if x is the generated image of the action generation model, y is 0.
[0069] S602: Calculate the perceptron loss function L through the human recognition model I ;
[0070] The hidden features of the posture transfer generated image and the hidden features of the human body model in the driving image are extracted through a human body recognition model; the human body recognition model can adopt any human body detection model, such as a CMU human body detection model, in which the main network adopts a ResNet50 network;
[0071] The feature difference between the posture transfer generated image and the corresponding extracted driving image is calculated as the perceptron loss function; the specific formula is as follows:
[0072] L J =||J(D g )-J(Q)||
[0073] Among them, D g Generates the image for posture transfer, and Q is the driving image.
[0074] The feature difference is the distance between the hidden feature vector of the last layer of the model when the posture transfer generated image and the driving image are input into the model.
[0075] S603: Combine the discriminator loss function with the perceptron loss function to output a final overall loss function L. The combination of the perceptron loss function and the discriminator loss function is as follows:
[0076] L=w1L J +w2L D
[0077] Among them, w1 and w2 are the weight coefficients of each loss function, which are manually set according to the specific situation.
[0078] Back propagation is performed based on the obtained loss function, and the SGD method is used to optimize the model parameter weights. The training reaches the set rounds or the loss is reduced to a given threshold to complete the model training.
[0079] Example 2
[0080] Embodiment 2 of the present invention provides a human posture migration system based on action drive, such as Figure 1 As shown, the system includes a data module, an action estimation module, an action generation module, and a loss function module;
[0081] The data module is used to collect human action videos and randomly extract image frames to obtain source images, drive image data and corresponding human key point coordinates;
[0082] The action estimation module is connected to the data module, and is used to receive the source image, the driving image and the coordinate data of the key points of the human body, and output the optical flow field and the redrawing map;
[0083] The action generation module is connected to the action estimation module and the data module, and is used to receive the optical flow map, the redrawing map and the source image, splice the hidden layer feature map of the source image with the optical flow map, and multiply the splicing result with the redrawing map, and finally output the posture migration generated image;
[0084] The loss function module is connected to the action generation module and the data module, and is used to receive the driving image and the generated image, calculate the perceptron loss function and the discriminator loss function, and combine and output the overall loss function.
[0085] The present invention is not limited to the above-mentioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.
Claims
1. A human posture transfer method based on action drive, characterized in that: The methods include the following: S1: Acquire human action video data, extract image frames of the video data to obtain a number of continuous pictures, and screen the extracted pictures to obtain a target picture; S2: Detect key points of the human body on the target image to obtain the coordinates of the key points; S3: randomly extracting two images from the target image as a source image and a driving image respectively, and calculating a transformation relationship between the driving image and the source image according to the obtained key point coordinates; S4: input the transformation relationship into the motion estimation model, and output the corresponding optical flow map and redrawing map; S5: inputting the optical flow map and the redrawing map into an action generation model to obtain a posture transfer generation image; The action generation model is a generative adversarial network, which is composed of a super-resolution model. The super-resolution model includes an encoder E and a decoder G. The input of the encoder E is the driving image, and the output is the hidden layer feature f of the driving image. E , concatenate the optical flow map L with the output features of the encoder E, and then multiply it with the remap map M to get the input of the decoder G, and finally generate the migrated image D through the decoder g , the formula is as follows: D g =G(M⊙(E(S),L)); The specific process of obtaining the generated image by posture transfer is as follows: Input the source image into the action generation model to obtain the hidden feature map of the source image and concatenate it with the optical flow map; Multiplying the obtained stitching result with the redrawing image, and inputting the multiplied output result into the decoder of the model, and outputting the posture transfer generated image; S6: Based on the driving image and the posture transfer generation image, a loss function is calculated. The specific process is as follows: S601: Calculate the discriminator loss function L through the discriminator network model D D ; S602: Calculate the perceptron loss function L through the human recognition model I ; S603: Combine the discriminator loss function with the perceptron loss function, and output a final overall loss function L.
2. The human body posture migration method according to claim 1, characterized in that: The human body motion video data is single person motion video data, and the screening is to delete incomplete human body video data.
3. The human body posture migration method according to claim 1, characterized in that: Step S2 can also re-screen the image through human key point detection, and delete the data where human key points cannot be detected or key points of multiple people are detected.
4. The human body posture migration method according to claim 1, characterized in that: In step S601, the discriminator network model adopts the VGG16 model, and the loss function adopts the cross entropy loss function, and the formula is as follows: L D =-ylog(D(x))-(1-y)log(1-D(x) Among them, x is the input image and y is the image label.
5. The human body posture migration method according to claim 4, characterized in that: In step S602, the perceptron loss function L I The calculation process is as follows: Extracting model hidden features of the human body in the posture transfer generated image and the driving image through a human body recognition model; The feature difference between the posture transfer generated image and the corresponding extracted driving image is calculated as the perceptron loss function, and the formula is as follows: L J =||J(D g )-J(Q)|| Where D g Generates the image for posture transfer, and Q is the driving image.
6. The human body posture transfer method according to claim 5, characterized in that: The feature difference is the distance between the hidden feature vector of the last layer of the model when the posture transfer generated image and the driving image are input into the model.
7. The human body posture transfer method according to claim 5, characterized in that: In step S603, the perceptron loss function and the discriminator loss function are combined as follows: L=w1L J +w2L D Where w1, w2 are the weight coefficients of each loss function.
8. A human body posture migration system, characterized in that: The system includes a data module, an action estimation module, an action generation module, and a loss function module; The data module is used to collect human action videos and randomly extract image frames to obtain source images, drive image data and corresponding human key point coordinates; The action estimation module is connected to the data module, and is used to receive the source image, the driving image and the coordinate data of the key points of the human body, calculate the transformation relationship between the driving image and the source image according to the obtained key point coordinates, input the transformation relationship into the action estimation model, and output the optical flow map and the redrawing map; The action generation module is connected to the action estimation module and the data module, and is used to receive the optical flow map, the redrawing map and the source image, splice the hidden layer feature map of the source image with the optical flow map, and multiply the splicing result with the redrawing map, and finally output the posture migration generated image; The loss function module is connected to the action generation module and the data module, and is used to receive the driving image and the generated image, calculate the perceptron loss function and the discriminator loss function, and combine and output the overall loss function.
Citation Information
Patent Citations
Expression migration and animation method and system
CN111598977A
Image driving model training, image generation method, device, equipment and medium
CN111797753A
Face image automatic generation method based on face contour
CN111931908A
Image facial expression migration method and device, electronic equipment and readable storage medium
CN112800869A