A human body image generation method based on pose guidance, style, and shape feature constraints

The proposed method addresses the limitations of existing human image generation by using separate encoders for style, pose, and shape features, enabling flexible and precise control over semantic regions, resulting in more realistic and accurate image synthesis.

CN113160035BActive Publication Date: 2025-07-08ZHEJIANG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110413125.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-16
Publication Date
2025-07-08
Estimated Expiration
2041-04-16

AI Technical Summary

Technical Problem

Existing human image generation technologies lack the ability to independently extract and control specific semantic region features for style and shape, limiting their flexibility and control over image synthesis.

Method used

A method that employs a generator network with separate encoders for style, pose, and shape features, using semantic segmentation to independently manipulate and combine these features, along with discriminators for adversarial training and loss functions to optimize the image generation process.

Benefits of technology

Enables flexible and precise control over semantic region features in human image generation, allowing for more nuanced image synthesis by independently manipulating style and shape, enhancing the realism and accuracy of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113160035B_ABST
    Figure CN113160035B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating human body images based on pose guidance, style, and shape feature constraints, including: (1) Collecting and obtaining the source human body image I s and the target human body image I t , and calculating their pose images P s , P t and the human body semantic segmentation images S s , S t ; (2) Constructing the generator G and the discriminator D I , D P ; (3) Inputting I s and S s into the style encoder, P t into the pose encoder, and S t into the shape encoder; Inputting the obtained style features, pose features, and shape features into the decoder to obtain the virtual target human body image I f ; (4) Taking (I s , I t ) and (I s , I f ) as the inputs of the discriminator D I , taking (P t , I t ) and (P t , I f ) as the inputs of the discriminator D P , calculating the adversarial losses respectively, and calculating the image reconstruction loss, perceptual loss, and semantic loss based on I f and I t to optimize G; (5) Performing iterative training to obtain the generator G for the application of generating human body images. Using the present invention, style features can be extracted according to semantic regions to control the pose and shape of human body images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of human body image generation, and in particular relates to a human body image generation method based on pose guidance, style and shape feature constraints. Background Art

[0002] Human body image generation is an important branch in the field of computer vision and can be widely applied to data augmentation for pedestrian re-identification, movie character production, virtual fitting, augmented reality and other fields. Pose-guided human body image generation refers to generating a target human body image with the style features of the source image in the target pose given a target pose and a source image (or a set of source images).

[0003] For example, the Chinese patent document with the publication number CN112116673A discloses a virtual human body image generation method based on structural similarity under pose guidance; the Chinese patent document with the publication number CN109191366A discloses a multi-view human body image synthesis method and device based on human body pose.

[0004] There are currently two problems in human body image generation: (1) In style feature extraction, a global style feature is often extracted with the entire source image as the input, and the features of specific semantic regions cannot be extracted separately. (2) The control method is single, and only the pose of the source image can be changed, while the style and shape of specific semantic regions cannot be controlled.

[0005] Therefore, there is an urgent need for a human body image generation method that can provide diverse image synthesis control methods. Summary of the Invention

[0006] The present invention provides a human body image generation method based on pose guidance, style and shape feature constraints, which can extract style features according to semantic regions and control the pose and shape of human body images.

[0007] A human body image generation method based on pose guidance, style and shape feature constraints includes the following steps:

[0008] (1) Collect and obtain the source human body image I s and the target human body image I t , and respectively obtain their pose images P s , P t and the human body semantic segmentation images S s , S t ;

[0009] (2) Construct a generator G and a discriminator D I , D P , wherein the generator G includes a style encoder Encoderstyle , Pose Encoder pose , Shape Encoder shape and Decoder; Discriminator D I for discriminating the generated virtual target image I f and the source human image I s for texture similarity; Discriminator D P for discriminating the generated virtual target image I f and the target pose P t for consistency;

[0010] (3) Input the source human image I s and the source human semantic segmentation image S s into the Style Encoder style , input the target pose image P t into the Pose Encoder pose , input the target human semantic segmentation image S t into the Shape Encoder shape ;

[0011] Input the sequentially extracted style features, pose features, and shape features into the Decoder to obtain the virtual target human image I f ;

[0012] (4) Take (I s , I t ) and (I s , I f ) as the inputs of the Discriminator D I respectively, take (P t , I t ) and (P t , I f ) as the inputs of the Discriminator D P respectively, calculate the adversarial loss L adv respectively, and calculate the image reconstruction loss L f based on I t and I reconstruction , perceptual loss L perceptual and semantic loss L CX to optimize G;

[0013] (5) Loop through steps (3) and (4). After reaching the preset number of iterations, obtain the trained Generator G and use it for generating virtual target images in the real scenario.

[0014] In step (1), the number of key points of the pose image N = 18, and the number of categories of the human semantic segmentation image C = 8.

[0015] The specific steps of step (2) are as follows:

[0016] (2-1) Construct a style encoder Encoder style

[0017] Encoder style It includes a pre-trained VGG network with 5 3×3 convolutional layers. The sizes of the feature maps extracted by the first 4 convolutional layers correspond to the sizes of the {1_1, 2_1, 3_1, 4_1} feature maps in VGG respectively; the features extracted by the convolutional layers are combined in sequence with the features extracted by the VGG network and then input into the next convolutional layer; for the last convolutional layer, the features are mapped from 1024 dimensions to 64 dimensions;

[0018] During use, first use semantic segmentation to segment the image into 8 independent images Then input the 8 independent semantic images into Encoder style respectively to output the corresponding style features, and finally concatenate them in sequence to obtain the final 512-dimensional style features. (2-2) Construct a pose encoder Encoder pose and a shape encoder Encoder shape

[0019] Encoder pose and Encoder shape have the same network structure, both including 4 3×3 convolutional layers, and the activation layer is the ReLU layer, extracting 512-dimensional pose features and shape features;

[0020] (2-3) Construct a decoder Decoder

[0021] Using the pose features as the input, calculate the normalization parameters using the style features and shape features; first pass through 4 ResBlocks while keeping the number of channels unchanged; then pass through 3 groups of upsampling layers and ResBlock layers; except that the activation layer of the last layer is tanh, the other activation layers are all ReLU layers.

[0022] (2-4) Construct a discriminator D I 、D P

[0023] Use PatchGAN as the discriminator, including 4 3×3 convolutional layers and 3 residual blocks, and the Dropout of the discriminator is set to 0.5.

[0024] In step (4), the definition of the adversarial loss function is:

[0025]

[0026] In the formula, E represents expectation.

[0027] In step (4), the image reconstruction loss L reconstruction is the L1 loss between the virtual target image and the real target image, which is defined as:

[0028] L reconstruction = ||G(I s , S s , P t , S t ) - I t ||1.

[0029] The image perception loss is defined as:

[0030]

[0031]

[0032] where represents the Gram matrix, represents the l-th layer feature map of I t extracted by the pre-trained VGG19 network, and l = relu{3_2, 4_2};

[0033] The semantic loss L CX is defined as:

[0034]

[0035] In the formula, represents the l-th layer feature map of I f extracted by the pre-trained VGG19 network. In step (5), during the training process, the learning rate is initially 0.0001 and linearly decays to 0 in 1000 iterations.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. In the human body image generation method based on pose guidance, style and shape feature constraints provided by the present invention, the style encoder based on the semantic segmentation image can independently extract the features of each semantic region and combine them into style features in a preset order, so that the features between different semantic regions are independent. In the case of a set of source images, style feature recombination can be achieved, which is more flexible in practical applications.

[0038] 2. In the human body image generation method based on pose guidance, style and shape feature constraints provided by the present invention, the decoder uses the shape features of the target semantic segmentation image for normalization and can output an image that conforms to the target semantic segmentation. Compared with the prior art of human body image generation methods based on pose guidance, the present invention provides a control method for modifying the generated image by modifying the semantic segmentation image. Brief Description of the Drawings

[0039] Figure 1 It is a schematic flowchart of the method of the present invention;

[0040] Figure 2 It is a schematic diagram of the human body image posture of the present invention;

[0041] Figure 3 It is a schematic diagram of the semantic segmentation of the human body image of the present invention;

[0042] Figure 4 It is a schematic diagram of the style encoder of this method. Detailed Description of the Invention

[0043] The present invention will be further described in detail below with reference to the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not limit it in any way.

[0044] As Figure 1 shown, a method for generating a human body image based on pose guidance, style, and shape feature constraints includes the following steps:

[0045] Step 1, collect and obtain the source human body image I s and the target human body image I t ; respectively obtain their pose images P s , P t and the human body semantic segmentation image S s , S t .

[0046] Specifically, as Figure 2 shown, the number of key points of the pose image N = 18; as Figure 3 shown, the number of categories of the human body semantic segmentation image C = 8.

[0047] Step 2, construct the generator G and the discriminator D I , D P , where the generator G includes a style encoder Encoder styte , a pose encoder Encoder pose , a shape encoder Encoder shape and a decoder Decoder.

[0048] The specific steps are as follows:

[0049] Step 2.1, construct Encoder style , as Figure 4 shown.

[0050] Encoder stvleThe pre-trained VGG network contains 5 convolutional layers of 3×3. The sizes of the feature maps extracted by the first 4 convolutional layers correspond to the sizes of the feature maps {1_1, 2_1, 3_1, 4_1} in VGG respectively. Combine the features extracted by the convolutional layers and the features extracted by the VGG network in sequence, and then input them into the next convolutional layer. The last convolutional layer maps the features from 1024 dimensions to 64 dimensions.

[0051] During use, first use semantic segmentation to segment the image into 8 independent images. Then input the 8 independent semantic images into the Encoder respectively. stvle Output the corresponding style features, and finally concatenate them in sequence to obtain the final style features of 512 dimensions.

[0052] Step 2.2, construct the Encoder pose and the Encoder shape

[0053] The Encoder pose and the Encoder shape have the same network structure, both including 4 convolutional layers of 3×3, and the activation layer is the ReLU layer, extracting 512-dimensional pose features and shape features.

[0054] Step 2.3, construct the Decoder

[0055] Use the pose features as the input, and calculate the normalization parameters using the style features and shape features.

[0056] First pass through 4 ResBlocks, keeping the channels unchanged; then pass through 3 groups of upsampling layers and ResBlock layers. Except for the last activation layer which is tanh, the rest of the activation layers are all ReLU layers.

[0057] Step 3.4, construct the discriminator D I 、D P

[0058] Use PatchGAN as the discriminator, including 4 convolutional layers of 3×3 and 3 residual blocks, and the Dropout of the discriminator is set to 0.5.

[0059] Step 3, input the source human body image I s and the source human body semantic segmentation image S s obtained in Step 1 into the style encoder Encoder style , input the target pose image P t into the pose encoder Encoder pose , input the target human body semantic segmentation image S t into the shape encoder Encoder shape; Input the sequentially extracted style features, pose features, and shape features into the decoder Decoder to obtain the virtual target human body image I f .

[0060] Step 4, take (I s , I t ) and (I s , I f ) as the inputs of the discriminator D I respectively, take (P t , I t ) and (P t , I f ) as the inputs of the discriminator D P respectively, calculate the adversarial loss L adv respectively, and calculate the image reconstruction loss L f , perceptual loss L t , and semantic loss L reconstruction based on I perceptual and I CX to optimize G

[0061] Specifically, the adversarial loss function is defined as:

[0062]

[0063] where E represents expectation

[0064] The image reconstruction loss is the L1 loss between the virtual target image and the real target image, defined as:

[0065] L reconstruction = ||G(I s , S s , P t , S t ) - I t ||1

[0066] The image perceptual loss is defined as:

[0067]

[0068]

[0069] where represents the Gram matrix, represents the l-th layer feature map of I t extracted by the pre-trained VGG19 network, and l = relu{3_2, 4_2}

[0070] The semantic loss L CX is defined as:

[0071]

[0072] Step 5: Loop through Step 3 and Step 4. After reaching the preset number of iterations, obtain the trained generator G for generating virtual target images in the real scenario.

[0073] Specifically, during the training process, the learning rate is initially 0.0001 and linearly decays to 0 in 1000 iterations.

[0074] The above-described embodiments have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, and equivalent replacements made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating human body images based on pose guidance, style, and shape feature constraints, characterized in that It includes the following steps: (1)Collect and obtain the source human body image I s and the target human body image I t , and respectively obtain their pose images P s , P t and the human body semantic segmentation image S s , S t ; The number of key points of the pose image N = 18, and the number of categories of the human body semantic segmentation image C = 8; (2) Construct the generator G and the discriminator D I and D P , where the generator G includes a style encoder Encoder style , a pose encoder Encoder pose , a shape encoder Encoder shape and a decoder Decoder; the discriminator D I is used to discriminate the texture similarity between the generated virtual target image I f and the source human image O s ; the discriminator D P is used to discriminate the consistency between the generated virtual target image I f and the target pose P t ; the specific steps are as follows: (2-1) Construct the style encoder Encoder style Encoder style It contains 5 convolutional layers of 3×3 and a pre-trained VGG network. The sizes of the feature maps extracted by the first 4 convolutional layers correspond to the feature map sizes of the {1_1, 2_1, 3_1, 4_1} layers in VGG respectively; the features extracted by the convolutional layers are combined in sequence with the features extracted by the VGG network and then input into the next convolutional layer; in the last convolutional layer, the features are mapped from 1024 dimensions to 64 dimensions; During use, first use semantic segmentation to segment the image into 8 independent images Then input the 8 independent semantic images into the Encoder style to output the corresponding style features, and finally concatenate them in sequence to obtain the final 512-dimensional style features; (2-2) Construct the pose encoder Encoder pose and the shape encoder Encoder shape Encoder pose has the same network structure as Encoder shape and both include four 3×3 convolutional layers. The activation layer is the ReLU layer, which extracts 512-dimensional pose features and shape features; (2-3) Construct a decoder Decoder Using the pose feature as the input, calculate the normalization parameters using the style feature and the shape feature; first pass through 4 ResBlocks while keeping the number of channels unchanged; then pass through 3 groups of upsampling layers and ResBlock layers; except for the last activation layer which is tanh, the rest of the activation layers are ReLU layers; (2-4) Construct discriminator D I , D P Use PatchGAN as the discriminator, which includes 4 3×3 convolutional layers and 3 residual blocks, and the Dropout of the discriminator is set to 0.5; (3) Input the source human body image I obtained in step (1) s and the source human body semantic segmentation image S s into the style encoder Encoder style , input the target pose image P t into the pose encoder Encoder pose , and input the target human body semantic segmentation image S t into the shape encoder Encoder shape ; Input the successively extracted style features, pose features, and shape features into the decoder Decoder to obtain the virtual target human body image I f ; (4) Take (I s , I t ) and (I s , I f ) as the inputs of discriminator D I respectively, take (P t , I t ) and (P t , I f ) as the inputs of discriminator D P respectively, calculate the adversarial loss L adv respectively, and calculate the image reconstruction loss L f , perceptual loss L t , and semantic loss L reconstruction based on I perceptual and I CX to optimize G; the definition of the adversarial loss function is as follows: In the formula, E represents the expectation; Image reconstruction loss L reconstruction is the L1 loss between the virtual target image and the real target image, defined as: L reconstruction = ||G(I s ,S s ,P t ,S t ) - I t ||1 The perceptual loss of the image is defined as: Among them, represents the Gram matrix, represents the l-th layer feature map of I t extracted by the pre-trained VGG19 network, where l = relu{3_2, 4_2}; Semantic loss L CX is defined as: wherein, represents the l-th layer feature map of I f extracted by the pre-trained VGG19 network; (5) Repeat steps (3) and (4). After reaching the preset number of iterations, obtain the trained generator G and use it to generate virtual target images in the real scenario.

2. The method for generating a human body image based on pose guidance, style, and shape feature constraints according to claim 1, wherein In step (5), during the training process, the initial learning rate is 0.0001 and linearly decays to 0 in 1000 iterations.

Citation Information

Patent Citations

  • Method and device for synthesizing multi-view human body image based on human body posture

    CN109191366A

  • Attitude-guided virtual human body image generation method and system based on structural similarity and electronic equipment

    CN112116673A

  • Image replacement method and device, computer equipment and storage medium

    CN111402118A