A human body image sequence generation method based on pose-shape-content inference

By combining pose manifold networks and transfer attention networks, the robustness and quality issues of human image sequence generation in existing technologies are solved, achieving high-quality pose transformation and image generation, which can be applied to clothing dynamic display and virtual try-on.

CN115311142BActive Publication Date: 2025-12-12ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210942446.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-12-12
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

Existing technologies struggle to generate high-quality continuous human image sequences between two end images, especially lacking robustness in pose transformation under arbitrary views. Furthermore, existing methods lack reasonable feature classification, leading to network overload and low-quality predictions.

Method used

A pose manifold network is used to control pose feature interpolation, and a transfer attention network is used to infer indirect features from pose to image. Through multi-scale feature-level optical flow estimation and feature code injection mechanism, a high-quality human image sequence is generated.

Benefits of technology

Robust pose transformation under any view was achieved, generating high-quality continuous human image sequences and enhancing consumers' deep perception of clothing wearing status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a human body image sequence generation method based on pose-shape-content reasoning. According to the top-down feature inference sequence, the pose interpolation path is limited in the pose manifold, and the pose interpolation is realized by using the linear control parameters of the pose manifold network. The interpolated pose features are converted into end shape features by using the transfer attention network, and the attention mechanism is adopted to emphasize the high-span spatial relationship. By comparing the target shape and the end shape, the content conversion network estimates the multi-scale feature level optical flow, and respectively deforms the content features of the clothing cluster and the human body cluster, and meanwhile, the feature code injection mechanism is adopted to assist the content inference of the passive features. The application is suitable for continuous human body image generation, and the method is helpful for solving the poor clothing effect display effect of the static model image in the clothing online retail, is helpful for realizing the dynamic display of the virtual fitting and the clothing selling, and is helpful for the pose-driven animation film production.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of human posture image conversion, and particularly relates to a human image sequence generation method based on posture-shape-content reasoning. BACKGROUND

[0002] In online clothing retail, merchants usually use multiple model pictures on the website to show the wearing effect of clothes. But consumers can only perceive the wearing state of clothes in a specific posture or view in static display. Dynamic display can provide consumers with a full range of clothing wearing state perception, which needs to take two model images as end images to infer the continuous human image sequence therebetween. This task requires the generation model to perform top-down feature reasoning and be robust to reasoning of any posture. In addition, the human image sequence synthesis method can also be applied to virtual fitting, film making and human re-identification.

[0003] Existing human posture conversion research mainly focuses on the reasoning of any posture between standard views, and is not robust to the conversion of any posture in any view. Since the remapping from posture to image is sparse, in order to ensure the quality of the generated image, a reasonable feature reasoning hierarchy must be established. At the same time, in order to obtain a continuous image sequence, the posture cannot be used as an a priori feature, and a parameter-driven posture interpolation method is urgently needed. Existing human posture conversion research usually inputs the overall content of the image into the network, resulting in network overload and low-quality prediction results. It lacks reasonable feature classification to help the network emphasize feature priorities in content conversion. In the spatial mapping of content conversion, it is necessary to embed an effective reasoning mechanism in the conversion of active and passive features to improve the generation quality of images.

[0004] In summary, the existing technology cannot achieve the inference task of continuous human image sequence between two end images. SUMMARY

[0005] The present application aims at the deficiencies of the prior art, and provides a human image sequence generation method based on posture-shape-content reasoning.

[0006] The purpose of the present application is achieved by the following technical solution: a human image sequence generation method based on posture-shape-content reasoning, comprising the following steps:

[0007] Step (1): input the starting interpolation end posture probability heat map and the ending interpolation end posture probability heat map of the clothing image into the posture manifold network, control the focus movement of the posture feature through the interpolation parameter t, and obtain the interpolated posture probability heat map;

[0008] Step (2): input the interpolated posture probability heat map and the end instance-level segmentation map into the migration attention network to obtain the target instance-level segmentation map;

[0009] Step (3): Estimate the multi-scale feature-level optical flow by comparing the target instance-level segmentation map and the end instance-level segmentation map obtained in step (2);

[0010] Step (4): Extract feature codes from the end instance-level segmentation map;

[0011] Step (5): Input the human body image into the image synthesis module and finally output the human body image under the interpolation parameter t.

[0012] Further, step (1) includes the following sub-steps:

[0013] (1.1) Input the initial interpolation end pose probability heatmap p0 of the clothing image into the pose encoder E. p In the posture encoder E p End-point generation of starting pose feature map The terminal interpolation end pose probability heatmap of the clothing image p1 input pose encoder E p In the posture encoder E p End-generated termination pose feature map

[0014] (1.2) The initial posture feature map Termination posture feature map Parallel input of the first bottleneck layer combination, which includes 3×3 bottleneck layers; initial pose feature map Output after 3 bottleneck layers Termination posture feature map Output after 3 bottleneck layers

[0015] (1.3) Then the initial pose feature map Termination posture feature map The input is fed into the pose manifold module for pose interpolation to obtain the interpolated pose feature map. Where t is the interpolation parameter, and t is the sum. Coefficient matrices of the same size, with ⊙ representing the element-wise multiplication sign; interpolation pose feature map. Output after 1 bottleneck layer

[0016] Initial pose feature map Termination posture feature map The input is fed into the pose manifold module for pose interpolation to obtain the interpolated pose feature map. and Add the inputs to the output of the next bottleneck layer.

[0017] Initial pose feature map Termination posture feature map The input is fed into the pose manifold module for pose interpolation to obtain the interpolated pose feature map. and Add the inputs to the output of the next bottleneck layer.

[0018] (1.4) Finally, Input gesture decoder D p In the process, the interpolated pose feature map is remapped into an interpolated pose probability heatmap p. t .

[0019] Furthermore, step (2) includes the following sub-steps:

[0020] (2.1) The interpolated pose probability heatmap p t Input posture encoder E p In the posture encoder E p End-generated interpolation pose feature map End instance-level segmentation map S0 input shape encoder E s In and in shape encoder E s End-generated shape feature map

[0021] (2.2) Interpolation pose feature map Input the second bottleneck layer combination, which includes 2×3 bottleneck layers, and interpolate the pose feature map. Output after 3 bottleneck layers in, This is the first pose feature map. This is the second pose feature map. This is a third pose feature map;

[0022] (2.3) Then the interpolation pose feature map Shape feature map The input is fed into the style transfer module for normalization and matching to obtain the first target shape feature map. σ() is the variance function, μ() is the mean function; first target shape feature map Output after 1 bottleneck layer

[0023] Will and first pose feature map The input is fed into the style transfer module for normalization matching to obtain the second target shape feature map. Second target shape feature map Output after 1 bottleneck layer

[0024] Will and the second pose feature map input to the style transfer module to obtain a third target shape feature map the third target shape feature map output the target shape feature map through one bottleneck layer

[0025] (2.4) interpolate the pose feature map and the target shape feature map input to the attention module to obtain wherein is a matrix multiplication symbol, softmax() is a normalization function, is a learnable parameter;

[0026] input the shape feature map and the target shape feature map to the attention module to obtain

[0027] then add β1 and β2 to obtain the shape feature map β after attention emphasis: β = β1 + β2;

[0028] (2.5) finally, input the shape feature map β after attention emphasis to the shape decoder D s to map the shape feature map β after attention emphasis to the target instance-level segmentation map S t .

[0029] Further, the step (3) comprises the following sub-steps:

[0030] (3.1) input the target instance-level segmentation map S t to the encoder to obtain the feature map

[0031] input the end instance-level segmentation map S0 to the encoder to obtain the feature map

[0032] (3.2) combine the feature map the feature map , pass the combined feature map through a Conv2d convolution layer to reduce the channel number to 16, and then generate a multi-scale feature-level optical flow f1 through a convolution layer combination {Conv2d + Tanh} and a reuse mask m1 through a convolution layer combination {Conv2d + Sigmoid};

[0033] (3.3) combine the feature map the feature map The merged features are combined through a Conv2d convolution layer to reduce the channel number of the merged features to 16, and then a convolution layer combination {Conv2d+Tanh} is used to generate a multi-scale feature-level optical flow f2.

[0034] (3.4) The feature map The feature map The merged features are combined through a Conv2d convolution layer to reduce the channel number of the merged features to 16, and then a convolution layer combination {Conv2d+Tanh} is used to generate a multi-scale feature-level optical flow f3.

[0035] Further, the step (4) includes the following sub-steps:

[0036] (4.1) Extracting background end shape, hair end shape, face end shape, upper body skin end shape, leg end shape, upper garment end shape and lower garment end shape from the end instance-level segmentation map S0; the background end shape, hair end shape, face end shape, upper body skin end shape, leg end shape, upper garment end shape and lower garment end shape are respectively input into the encoder E b , which are input into the background instance-level feature map The hair instance-level feature map The face instance-level feature map The upper body skin instance-level feature map The leg instance-level feature map The upper garment instance-level feature map The lower garment instance-level feature map

[0037] (4.2) The background instance-level feature map The hair instance-level feature map The face instance-level feature map The upper body skin instance-level feature map The leg instance-level feature map Parallel pooling is performed respectively, and corresponding feature codes are generated; five feature codes are synthesized to form a feature code vector c b The channel number of the feature code vector c b is 5;

[0038] (4.3) The upper garment instance-level feature map The lower garment instance-level feature map Parallel pooling is performed respectively, and corresponding feature codes are generated; two feature codes are synthesized to form a feature code vector c c The channel number of the feature code vector c c is 2;

[0039] (4.4) The feature code vector c b and the feature code vector c cThe features are merged to form a feature vector c, which has 7 channels; then, the formula is used... Obtain features The Υ[·] layer consists of two convolutional layers, increasing the number of channels from 7 to 64. The feature code... The number of channels is 64.

[0040] Furthermore, step (5) includes the following sub-steps:

[0041] (5.1) Segment the human body image I0 into human body clusters I b Clothing Clustering I c ;

[0042] Then, human clustering I b Input to encoder E b The encoder E b It includes 4 downsampling residual blocks, and outputs after passing through one downsampling residual block. Output after a downsampling residual block Then The input f1 is fed into the deformation module (WM) to obtain γ′. e 1 : Among them, warp() is the mesh deformation function;

[0043] γ′ e 1 Output after a downsampling residual block Will f2 is input to the deformation module to obtain γ′ e 2 :

[0044] γ′ e 2 Output after a downsampling residual block Will f3 is input to the deformation module to obtain γ′ e 3 :

[0045] (5.2) Introducing the Sobel gradient I for clothing clustering g To enhance texture quality and image structure during content conversion, Ig is calculated using the following formula:

[0046]

[0047] Where g() is the grayscale conversion function; k x Let k be the Sobel convolution factor along the x-axis. yis the Sobel convolution factor in y-axis direction; I c is the clothing cluster;

[0048] (5.3) the Sobel gradient of the clothing cluster I c is input into the encoder E c , the encoder E c includes 4 down-sampling residual blocks, and the output of one down-sampling residual block is the output of one down-sampling residual block is then and f1 are input into the deformation module to obtain δ' e 1 :

[0049] δ' e 1 the output of one down-sampling residual block is is input into the deformation module to obtain δ' and f2 are input into the deformation module to obtain δ' e 2 :

[0050] δ' e 2 the output of one down-sampling residual block is is input into the deformation module to obtain δ' and f3 are input into the deformation module to obtain δ' e 3 :

[0051] (5.4) the Sobel gradient of the clothing cluster I g is input into the encoder E g , the encoder E g includes 4 down-sampling residual blocks, and the output of one down-sampling residual block is the output of one down-sampling residual block is then and f1 are input into the deformation module to obtain η' e 1 :

[0052] η' e 1 the output of one down-sampling residual block is is input into the deformation module to obtain η' and f2 are input into the deformation module to obtain η' e 2 :

[0053] η' e 2 the output of one down-sampling residual block is The and f3 are input into the deformation module to obtain η' e 3 :

[0054] (5.5) δ' e 3 and η' e 3 are input into the gradient fusion module to obtain δ' d 0 : δ' d 0 = δ' e 3 + η' e 3 ;

[0055] δ' d 0 is output through a down-sampling residual block Then, δ' e 2 , η' e 2 and are input into the gradient fusion module to obtain δ' d 1 : wherein, is a merging operation;

[0056] δ' d 1 is output through a down-sampling residual block Then, δ' e 1 , η' e 1 and are input into the gradient fusion module to obtain δ' d 2 :

[0057] δ' d 2 is output through a down-sampling residual block

[0058] (5.6) γ' e 3 is merged with the feature code , then passes through a bottleneck layer, and is merged with the feature code ; then, is input into four up-sampling residual blocks, is merged with after passing through the first up-sampling residual block, and is input into the second up-sampling residual block; is merged with The feature merging is performed and input to a third up-sampling residual block; after passing through the third up-sampling residual block, the human body image I The feature merging is performed and input to a fourth up-sampling residual block; after passing through the fourth up-sampling residual block, the human body image I t .

[0059] The present application has the beneficial effects that: the present application takes the continuous human body image sequence generation method for the dynamic display of clothing as the main method, the method can help to solve the reasonable content reasoning problem in the parameterized interpolation two-dimensional pose and pose conversion, and realize the pose-driven human body image sequence generation. The method can be used for the dynamic display of online clothing or virtual fitting, so as to improve the depth perception of consumers on the wearing state of clothing. BRIEF DESCRIPTION OF DRAWINGS

[0060] Figure 1 is a flowchart of the present application;

[0061] Figure 2 is a structure diagram of the pose manifold network;

[0062] Figure 3 is a structure diagram of the migration attention network;

[0063] Figure 4 is a structure diagram of the content conversion network;

[0064] Figure 5 is a structure diagram of the feature code injection module;

[0065] Figure 6 is a shape-guided content conversion result;

[0066] Figure 7 is a pose-driven human body image sequence generation result. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical scheme and advantages of the present application more clear and clear, the present application is further described in detail in combination with the drawings and examples, and it should be understood that the specific examples described herein are only used to explain the present application, rather than all examples. Based on the examples in the present application, all other examples obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0068] Example 1

[0069] As Figure 1As shown, this invention proposes a method for generating human image sequences based on pose-shape-content reasoning. Following a top-down feature inference order, the pose interpolation path is constrained within a pose manifold, and a pose manifold network is used to control the interpolation parameters to achieve pose interpolation. Style normalization in a transfer attention network is used to transfer the interpolated pose features to the end shape, and an attention mechanism is used to emphasize high-span spatial relationships. By comparing the differences between the target shape and the end shape, a content transfer network is used to estimate multi-scale feature-level optical flow to achieve deformation of human body clustering and clothing clustering, and a feature code injection mechanism is used to assist in content reasoning based on passive features.

[0070] The present invention specifically includes the following steps:

[0071] Step (1): Input the starting interpolation end pose probability heatmap and the ending interpolation end pose probability heatmap of the clothing image into the pose manifold network, and control the focus movement of the pose features through the interpolation parameters to obtain the interpolated pose probability heatmap.

[0072] The pose interpolation position between two poses is determined using a pose manifold network. The 2D pose is constrained to a low-dimensional space, and a unique linear interpolation path is determined. The focus movement of the pose feature is controlled by the interpolation parameters to achieve end-to-end 2D pose interpolation.

[0073] Step (1) specifically includes the following sub-steps:

[0074] (1.1) Input the initial interpolation end pose probability heatmap p0 of the clothing image into the pose encoder E. p In the posture encoder E p End-point generation of starting pose feature map The terminal interpolation end pose probability heatmap of the clothing image p1 input pose encoder E p In the posture encoder E p Termination pose feature map generated at the end The posture encoder E p It includes four downsampling convolutional layers; compared to the pose probability heatmap (p0, p1), the pose feature map... It has smaller spatial dimensions and lower degrees of freedom, while still retaining the original location and structural features;

[0075] (1.2) The initial posture feature map Termination posture feature map Parallel input of the first bottleneck layer combination, which includes 3×3 bottleneck layers; initial pose feature map Output after 3 bottleneck layers Termination posture feature map Output after 3 bottleneck layers

[0076] (1.3) Subsequently, the starting pose feature map the ending pose feature map is input into the pose manifold module for pose interpolation to obtain an interpolated pose feature map where t is an interpolation parameter, t is and a coefficient matrix of the same size, and is an element multiplication symbol; the interpolated pose feature map is output after passing through a bottleneck layer

[0077] the starting pose feature map the ending pose feature map is input into the pose manifold module for pose interpolation to obtain an interpolated pose feature map is added to is input into the next bottleneck layer to output

[0078] the starting pose feature map the ending pose feature map is input into the pose manifold module for pose interpolation to obtain an interpolated pose feature map is added to is input into the next bottleneck layer to output

[0079] (1.4) Finally, the input pose decoder D p re-maps the interpolated pose feature map to an interpolated pose probability heat map p t ; the pose decoder D p has four up-sampling convolutional layers.

[0080] Step (2): input the interpolated pose probability heat map and the end instance-level segmentation map into the transfer attention network to obtain a target instance-level segmentation map.

[0081] The transfer attention network is used to convert the interpolated pose feature into an end shape to generate a target shape. Since the correspondence between the pose and the image is sparse, if only the pose feature is used to infer the image, the multi-modal problem will lead to unreliable inference results. Therefore, the transfer attention network is used to infer the indirect feature of the pose to the image: the instance-level semantic segmentation, so as to improve the final generation quality of the image.

[0082] The step (2) specifically includes the following sub-steps:

[0083] (2.1) input the interpolated pose probability heat map p t into the pose encoder E p and generate an interpolated pose feature map at the end of the pose encoder E p ​ End instance level split graph S0 input shape encoder E s In the shape encoder E s End generate shape feature map

[0084] (2.2) interpolate pose feature map Input into the second bottleneck layer combination, the second bottleneck layer combination includes 2x3 bottleneck layers, interpolate pose feature map Through 3 bottleneck layers output Wherein, The first pose feature map, The second pose feature map, The third pose feature map;

[0085] (2.3) then interpolate pose feature map Shape feature map Input to the style transfer module normalization matching to get the first target shape feature map σ() is the variance function, μ() is the mean function; the first target shape feature map Through 1 bottleneck layer output

[0086] The first target shape feature map And the first pose feature map Input to the style transfer module normalization matching to get the second target shape feature map The second target shape feature map Through 1 bottleneck layer output

[0087] The second target shape feature map And the second pose feature map Input to the style transfer module normalization matching to get the third target shape feature map The third target shape feature map Through 1 bottleneck layer output target shape feature map

[0088] The target shape feature map Has both pose feature and shape feature;

[0089] (2.4) interpolate pose feature map And target shape feature map Input to the attention module to get Wherein The matrix multiplication symbol, softmax() is the normalization function, Is a learnable parameter;

[0090] The shape feature map and target shape feature map input into the attention module to obtain

[0091] Then β1 and β2 are added to obtain the shape feature map β after attention emphasis: β = β1 + β2.

[0092] The target shape feature map is emphasized by using the attention module and the interpolated pose feature map The relationship between the shape feature map is used to promote the reasonable distribution of joint structure.

[0093] (2.5) Finally, the shape feature map β after attention emphasis is input into the shape decoder D s to map the shape feature map β after attention emphasis into the target instance-level segmentation map S t .

[0094] Step (3): Estimate the multi-scale feature-level optical flow by comparing the target instance-level segmentation map obtained in step (2) and the end instance-level segmentation map;

[0095] (3.1) Input the target instance-level segmentation map S t into the encoder to obtain the feature map

[0096] Input the end instance-level segmentation map S0 into the encoder to obtain the feature map

[0097] (3.2) Merge the feature map and the feature map , pass the merged feature map through a Conv2d convolution layer to reduce the channel number to 16, and then generate the multi-scale feature-level optical flow f1 through a convolution layer combination {Conv2d + Tanh} and generate the reuse mask m1 through a convolution layer combination {Conv2d + Sigmoid};

[0098] (3.3) Merge the feature map and the feature map , pass the merged feature map through a Conv2d convolution layer to reduce the channel number to 16, and then generate the multi-scale feature-level optical flow f2 through a convolution layer combination {Conv2d + Tanh};

[0099] (3.4) Merge the feature map and the feature map The merged feature maps are convoluted by a Conv2d layer to reduce the channel number to 16, and then a multi-scale feature-level optical flow f3 is generated by a convolution layer combination {Conv2d+Tanh}.

[0100] Step (4): Extracting feature codes from the end-instance-level segmentation map;

[0101] (4.1) Extracting background end shape, hair end shape, face end shape, upper body skin end shape, leg end shape, upper garment end shape, and lower garment end shape from the end-instance-level segmentation map S0; the background end shape, hair end shape, face end shape, upper body skin end shape, leg end shape, upper garment end shape, and lower garment end shape are respectively input into the encoder E b , and after passing through two down-sampling residual blocks, the background instance-level feature map hair instance-level feature map face instance-level feature map upper body skin instance-level feature map leg instance-level feature map upper garment instance-level feature map lower garment instance-level feature map

[0102] (4.2) The background instance-level feature map hair instance-level feature map face instance-level feature map upper body skin instance-level feature map leg instance-level feature map are respectively parallel-pooled to generate corresponding feature codes; and five feature codes are synthesized to form a feature code vector c b , wherein the channel number of the feature code vector c b is 5.

[0103] (4.3) The upper garment instance-level feature map lower garment instance-level feature map are respectively parallel-pooled to generate corresponding feature codes; and two feature codes are synthesized to form a feature code vector c c , wherein the channel number of the feature code vector c c is 2.

[0104] (4.4) Merging the feature code vector c b and the feature code vector c c to form a feature code vector c, wherein the channel number of the feature code vector c is 7; and then obtaining a feature by the formula wherein Υ[·] is composed of two convolution layers to increase the channel number from 7 to 64, and the channel number of the feature code is 64.

[0105] Step (5): Input the human image I0 into the image synthesis module, and finally output the human image I under the interpolation parameter t. t .

[0106] Visual features in human images consist of two main clusters: the human body cluster and the clothing cluster. The visual features of the human body cluster include hair, face, upper body skin, and legs, emphasizing structural representation. The visual features of the clothing cluster include upper and lower garments, emphasizing texture representation. Because the features of the human body and clothing clusters have different emphases in content transformation, if a single model is used to learn and transform all features, the model will suffer from multitasking overload, failing to effectively represent either human structure or clothing texture in inference.

[0107] (5.1) Segment the human body image I0 into human body clusters I b Clothing Clustering I c ;

[0108] Then, human body clustering I b Input to encoder E b The encoder E b It includes 4 downsampling residual blocks, and the output is after passing through one downsampling residual block. Output after a downsampling residual block Then The input f1 is fed into the deformation module (WM) to obtain γ′. e 1 : Among them, warp() is the mesh deformation function;

[0109] γ′ e 1 Output after a downsampling residual block Will f2 is input to the deformation module to obtain γ′ e 2 :

[0110] γ′ e 2 Output after a downsampling residual block Will f3 is input to the deformation module to obtain γ′ e 3 :

[0111] (5.2) Introducing the Sobel gradient I for clothing clustering g To enhance texture quality and image structure during content conversion, Ig is calculated using the following formula:

[0112]

[0113] where g() is a gray scale conversion function; k x is a Sobel convolution factor in the x-axis direction, k y is a Sobel convolution factor in the y-axis direction; I c is a clothing cluster;

[0114] (5.3) the clothing cluster I c is input into an encoder E c , the encoder E c includes 4 down-sampling residual blocks, and the output of one down-sampling residual block is the output of one down-sampling residual block is Subsequently, I and f1 are input into a deformation module to obtain δ' e 1 :

[0115] δ' e 1 the output of one down-sampling residual block is I and f2 are input into a deformation module to obtain δ' e 2 :

[0116] δ' e the output of one down-sampling residual block is I and f3 are input into a deformation module to obtain δ' e 3 :

[0117] (5.4) the Sobel gradient of the clothing cluster I g is input into an encoder E g , the encoder E g includes 4 down-sampling residual blocks, and the output of one down-sampling residual block is the output of one down-sampling residual block is Subsequently, I and f1 are input into a deformation module to obtain η' e 1 :

[0118] η' e 1 the output of one down-sampling residual block is I and f2 are input into a deformation module to obtain η' e2 :

[0119] η′ e 2 Output after a downsampling residual block Will f3 is input to the deformation module to obtain η′ e 3 :

[0120] (5.5) δ′ e 3 and η′ e 3 The input is fed into the gradient fusion module to obtain δ′. d 0 :δ′ d 0 =δ′ e 3 +η′ e 3 ;

[0121] δ′ d 0 Output after a downsampling residual block Then δ′ e 2 η′ e 2 and The input is fed into the gradient fusion module to obtain δ′. d 1 : in, For merging operations;

[0122] δ′ d 1 Output after a downsampling residual block Then δ′ e 1 η′ e 1 and The input is fed into the gradient fusion module to obtain δ′. d 2 :

[0123] δ′ d 2 Output after a downsampling residual block

[0124] (5.6) γ′ e 3 With signature The process involves merging, then passing through a bottleneck layer, and finally combining with the feature code. merge; then input to four up-sampling residual blocks, after the first up-sampling residual block, and merge, input to the second up-sampling residual block; after the second up-sampling residual block, and merge, input to the third up-sampling residual block; after the third up-sampling residual block, and merge, input to the fourth up-sampling residual block; after the fourth up-sampling residual block, output the human body image I t .

[0125] Embodiment 2

[0126] The embodiment of the present application uses the frame image of the model fitting video to realize the self-supervised training and testing of the model. The training set and the testing set respectively include 473 and 99 videos, all the frame images are extracted from the videos, and are cropped and unified to 256x256.

[0127] The embodiment of the present application uses the Adam optimizer to optimize the network training, sets the learning rate to be dynamically changed, divides all iterations into four sub-periods, the decay rate between the sub-periods is 50%, and sets the initial learning rate to be 4x10 -4 .

[0128] The embodiment of the present application tests the robustness of the content conversion network in image reasoning guided by shape. As shown in Figure 6 , for the same end image, 10 target shapes with slight differences are listed. The content conversion network can effectively perceive the slight differences of the target shape, realize fine feature deformation and passive feature reasoning within the instance, and generate a high-quality human body image sequence.

[0129] The embodiment of the present application tests the integration effect of the three networks and the robustness of the human body image sequence reasoning. As shown in Figure 7 , given two end poses, the pose manifold network can generate a continuous pose sequence by controlling the interpolation parameter t. Based on the pose sequence and the end image, the migration attention network and the content conversion network can infer the corresponding shape sequence and image sequence. Finally, the same pose sequence drives the corresponding image sequence to be generated without human body. Therefore, the present application will be beneficial to online retailing, virtual fitting and animation production of clothes and the like.

[0130] The above only describes the preferred embodiments of the present application and should not be used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A human body image sequence generation method based on pose-shape-content inference, characterized in that, The method comprises the following steps: Step (1): inputting a starting interpolation end pose probability heat map and a terminal interpolation end pose probability heat map of a garment image into a pose manifold network, controlling focus movement of pose features through an interpolation parameter t, and obtaining an interpolated pose probability heat map; Step (2): inputting the interpolated pose probability heat map and an end instance-level segmentation map into a transfer attention network to obtain a target instance-level segmentation map; The step (2) comprises the following sub-steps: (2.1) input the interpolated pose probability heat map p t input pose encoder E p into the pose encoder E p and generate an interpolated pose feature map at the end of the pose encoder E s into the shape encoder E s and generate a shape feature map at the end of the shape encoder E (2.2) interpolate the pose feature maps input a second bottleneck layer combination, the second bottleneck layer combination comprising a bottleneck layer, interpolate the pose feature maps output by the 3 bottleneck layers : wherein, is a first pose feature map, is a second pose feature map, is a third pose feature map; (2.3) Then the interpolation pose feature map Shape feature diagram The input is fed into the style transfer module for normalization and matching to obtain the first target shape feature map. : , It is the variance function. The mean function; the first target shape feature map Output after 1 bottleneck layer ; Will and the first pose feature map Input to the style transfer module is normalized to match the second target shape feature map : ; the second target shape feature map After 1 bottleneck layer output ; will be described below with reference to the accompanying drawings. In the following description, well-known functions or constructions are not described in detail because they would obscure the understanding of the present application. As used herein, the term "and / or" includes combinations thereof, for example, A and / or B includes all of the following: A; B; A and B. and the second pose feature map input to the style transfer module is normalized to obtain a third target shape feature map : ; the third target shape feature map passes through a bottleneck layer to output a target shape feature map ; (2.4) inputting the interpolated pose feature map and the target shape feature map into an attention module to obtain wherein is a matrix multiplication symbol, softmax() is a normalization function, are learnable parameters; inputting the shape feature map and the target shape feature map into an attention module to obtain ; Subsequently, the shape feature map after attention emphasis is obtained by adding the shape feature map after attention emphasis and the shape feature map after attention emphasis and : ;​ (2.5) Finally, the shape feature maps with attention emphasized Input shape decoder D s In the middle, the shape feature maps with attention emphasized are remapped into the target instance-level segmentation map S t ; Step (3): estimating a multi-scale feature-level optical flow by comparing the target instance-level segmentation map obtained in the step (2) and the end instance-level segmentation map; The step (3) comprises the following sub-steps: (3.1) split the target instance-level segmentation map S into a foreground mask F and a background mask B t input encoder, to obtain a feature map through 4 down-sampling residual blocks 、 、 、 : ; The end instance level segmentation graph S0 is input into an encoder, and a feature map is obtained through 4 down-sampling residual blocks 、 、 、 : ; (3.2) merge the feature maps , the feature maps , the channel number of the merged feature maps is reduced to 16 through a Conv2d convolution layer, and a multi-scale feature level optical flow f1 is generated through a convolution layer combination {Conv2d+Tanh}, and a reuse mask m1 is generated through a convolution layer combination {Conv2d+Sigmoid}; (3.3) merge the feature maps , the feature maps , and pass the merged feature maps through a Conv2d convolution layer to reduce the number of channels to 16, and then generate a multi-scale feature-level optical flow f2 through a convolution layer combination {Conv2d+Tanh}. (3.4) merge the feature maps , the feature maps , and generate multi-scale feature-level optical flow f3 by a Conv2d convolution layer to reduce the channel number of the merged feature maps to 16, and then by a convolution layer combination {Conv2d+Tanh}; Step (4): extracting a feature code from the end instance-level segmentation map; Step (5): inputting a human body image into an image synthesis module to finally output a human body image under the condition of the interpolation parameter t; The step (5) comprises the following sub-steps: (5.1) segmenting the human body image I0 into a human body cluster I b and a garment cluster I c ; Then, human body clustering I b Input to encoder E b The encoder E b It includes 4 downsampling residual blocks, and the output is after passing through one downsampling residual block. , Output after a downsampling residual block ; then The f1 input is obtained by the deformation module. : Where warp() is the mesh deformation function; through a down-sampling residual block , f2 and f2 are input to a warping module to obtain : ; through a down-sampling residual block , and and f3 are input to a warping module to obtain : ; (5.2) Sobel gradient I of garment cluster is introduced g to enhance the texture quality and image structure in content conversion, which is calculated by the following formula: ; where g() is a gray scale conversion function; k x is a Sobel convolution factor in the x-axis direction, k y is a Sobel convolution factor in the y-axis direction; I c is a clothing cluster; (5.3) Clustering the garments I c Input to the encoder E c The encoder E c Comprises 4 down-sampling residual blocks, after one down-sampling residual block outputs , After one down-sampling residual block outputs ; then input and f1 to the deformation module to obtain : ; through a down-sampling residual block , the and f2 are input to a deformation module to obtain : ; through a down-sampling residual block , the and f3 are input to a deformation module to obtain : ; (5.4) Sobel gradient I for clothing clustering g Input to encoder E g Encoder E g It includes 4 downsampling residual blocks, and the output is after passing through one downsampling residual block. , Output after a downsampling residual block ; then The f1 input is obtained by the deformation module. : ; through a down-sampling residual block , f2 and f2 are input to a warping module to obtain : ; through a down-sampling residual block , f2, and f3 are input to a deformation module to obtain : : ; (5.5) to and input to the gradient fusion module, obtaining : ; through a down-sampling residual block ; then the , and are input into a gradient fusion module to obtain : wherein, is a merging operation; through a down-sampling residual block ; then input to a gradient fusion module to obtain , and : ;​ through a down-sampling residual block output ; (5.6) will With signature The process involves merging, then passing through a bottleneck layer, and finally combining with the feature code. Merge; then input to four upsampled residual blocks, after passing through the first upsampled residual block and... Feature merging is performed, and the input is fed into the second upsampled residual block; after passing through the second upsampled residual block, it is combined with... Feature merging is performed, and the input is fed into the third upsampled residual block; after passing through the third upsampled residual block, it is combined with... Feature merging is performed, and the result is fed into the fourth upsampled residual block; after the fourth upsampled residual block, the output is the human image under the interpolation parameter t. .

2. The human body image sequence generation method based on pose-shape-content reasoning according to claim 1, characterized in that, The step (1) comprises the following sub-steps: (1.1) Input the initial interpolation end pose probability heatmap p0 of the clothing image into the pose encoder E. p In the posture encoder E p End-point generation of starting pose feature map The terminal interpolation end pose probability heatmap of the clothing image p1 input pose encoder E p In the posture encoder E p End-generated termination pose feature map ; (1.2) a start-pose feature map , an end-pose feature map a first bottleneck layer combination including one bottleneck layer; a start-pose feature map output through 3 bottleneck layers : , an end-pose feature map output through 3 bottleneck layers : ; (1.3) the start pose feature map and the end pose feature map are input into the pose manifold module to obtain an interpolated pose feature map : where t is an interpolation parameter, t is a coefficient matrix of the same size as , and is an element multiplication symbol; the interpolated pose feature map is output after 1 bottleneck layer ; a starting pose feature map , a terminating pose feature map , input to a pose manifold module for pose interpolation to obtain an interpolated pose feature map : ; and addition input to the next bottleneck layer output ; starting pose feature map , ending pose feature map input to the pose manifold module for pose interpolation to obtain an interpolated pose feature map : ; and addition input to the next bottleneck layer output ; (1.4) Finally, the interpolated pose feature maps are re-mapped to interpolated pose probability heatmaps p Input pose decoder D p In the middle, the interpolated pose feature maps are re-mapped to interpolated pose probability heatmaps p t .

3. The human body image sequence generation method based on pose-shape-content reasoning according to claim 2, characterized in that, The step (4) comprises the following sub-steps: (4.1) extracting a background end shape, a hair end shape, a face end shape, an upper body skin end shape, a leg end shape, an upper garment end shape and a lower garment end shape from the end instance-level segmentation map S0; background end shape, hair end shape, face end shape, upper body skin end shape, leg end shape, upper garment end shape, and lower garment end shape are respectively input into an encoder E b , all of which are input into background instance-level feature maps , hair instance-level feature maps , face instance-level feature maps , upper body skin instance-level feature maps , leg instance-level feature maps , upper garment instance-level feature maps , and lower garment instance-level feature maps after two down-sampling residual blocks (4.2) Background instance-level feature map , head instance-level feature map , face instance-level feature map , upper body skin instance-level feature map , leg instance-level feature map Parallel pooling is respectively performed, and corresponding feature codes are generated; five feature codes are synthesized to form a feature code vector c b The channel number of the feature code vector c b is 5. (4.3) Upper garment example level feature map , lower garment example level feature map Parallel pooling is performed respectively, and corresponding feature codes are generated; two feature codes are synthesized to form a feature code vector c c The channel number of the feature code vector c c is 2. (4.4) The feature code vector c b and the feature code vector c c are merged to form a feature code vector c with 7 channels; then the feature code is obtained by the formula where consists of two convolutional layers, which increase the number of channels from 7 to 64, and the feature code has 64 channels.