Two-stage improved portrait attitude migration generation method

Through the two-stage improved portrait pose migration generation method, combined with the GAN network and diffusion model, the ControlNet network is used for optimization and improvement, which solves the problems of the existing pose migration algorithm with low resolution, blurred details and unstable generated content, and achieves high-precision pose migration effect.

CN120219146APending Publication Date: 2025-06-27GUANGZHOU ZIWEIYUN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510278606.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing pose migration algorithms have problems such as low resolution, blurred details, poor generalization performance, and unstable content generated.

Method used

The two-stage improved portrait pose migration generation method is adopted. First, the GAN network is used to generate the preliminary pose migration effect, and then the second stage optimization and improvement is carried out through the diffusion model, combining with the ControlNet network to improve the generation effect.

Benefits of technology

The resolution and detail clarity of pose migration are improved, the instability and inconsistency of the content generated by the diffusion model are overcome, and the high-precision pose migration effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219146A_ABST
    Figure CN120219146A_ABST
Patent Text Reader

Abstract

The invention relates to a two-stage improved portrait attitude migration generation method. Comprising the following steps: a first stage process: generating a final high-precision attitude migration effect picture (Ifinal) through a Controlnet network control diffusion model; in the second stage process, a final high-precision attitude migration effect picture (Ifinal) is output through a decoder; and the third stage process is a training diffusion process, a sampling method adopts DDIM sampling, and training is carried out according to a Deep Fashon data set until the network converges. According to the method disclosed by the invention, the second-stage effect optimization and improvement (Refinement Stage) is carried out by aiming at any model of GAN (Generic Area Network) network attitude migration. And according to a first-stage result output by the GAN network, using a diffusion model to carry out combination to improve the effect. According to the method, the problems of low resolution and fuzzy details of a picture generated by the GAN network are solved, and the defect of instability and inconsistency of content generated by a diffusion model is overcome. And the advantages of the two are complemented, so that a high-precision attitude migration effect is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of human pose transfer, and specifically relates to a two-stage improved human pose transfer generation method. Background Art

[0002] At present, the pose transfer algorithm technology is generally divided into two directions: the generation method of combining pictures with human pose features based on the adversarial generative network (GAN) model and the human pose transfer generation method based on the diffusion model. Among them, the adversarial generative network (GAN) model inputs according to the input human RGB image combined with the pose features of the key bone points. By training the discriminator / generator architecture network, an effect picture after pose transfer is generated according to the input bone pose features. On the other hand, the mainstream generation scheme architecture of the human pose transfer model based on the diffusion model is the image-to-image architecture. It encodes the input human RGB image into a noisy latent variable Zt, and through the reverse diffusion process from Zt to Z0 by combining the pose control input features. The images generated by the pose transfer method based on the diffusion model are more realistic and clear in terms of image details than the GAN network model. However, there are still some defects, which are listed as follows:

[0003] When the diffusion network model generates a person, it is easy to generate new human features (such as background, clothing) that are different from the original picture.

[0004] The effect is often better on the training set, but the accuracy is insufficient in the real clothing picture test set.

[0005] The facial information of the person is deformed and cannot maintain the consistency of the expression and facial features with the original picture.

[0006] Compared with the pose transfer model based on the diffusion model network, the human pose transfer model based on the GAN network architecture has the following defects:

[0007] Low resolution: The output resolution of the GAN network model is not high, and the output image is blurred and unclear.

[0008] Poor generalization performance, good effect on the training set, but the transferred images output in the actual scenario are blurred and have local artifacts.

[0009] Difficult to converge. Adversarial training is a game process between the generator and the discriminator, with a long training time and slow network convergence. Summary of the Invention

[0010] To solve the above technical problems, the present invention provides a two-stage improved method for generating human pose transfer, which optimizes the effect of the second stage (Refinement Stage) for any model of GAN network pose transfer. According to the first-stage results output by the GAN network, a diffusion model is used for combination to improve the effect. It makes up for the problems of low resolution and blurred details of the pictures generated by the GAN network, and at the same time overcomes the defects of instability and inconsistency of the content generated by the diffusion model. By leveraging the advantages of both and compensating for their weaknesses, a high-precision pose transfer effect is achieved.

[0011] To achieve the above object, the technical solution adopted by the present invention is:

[0012] A two-stage improved method for generating human pose transfer, including

[0013] The first-stage process: First, given an RGB human image I and a skeletal key point feature map P, based on these two features, the existing GAN network pose generation algorithm is used to obtain a preliminary pose transfer generation effect Iraw; the original image is passed into the feature extractor module to extract the human Mesh feature Imesh and the face feature Iface. Combining the previous skeletal feature map P, these three feature maps are concatenated to obtain the diffusion model control input C; where the dimension of C is (3*3*H*W); H is the height of the feature map and W is the width; the control input feature map C is passed into the VGG19 pre-trained encoder to obtain the control input C', and then the Controlnet network is used to control the diffusion model to generate the final high-precision pose transfer effect picture (Ifinal);

[0014] The second-stage process: The diffusion model for pose transfer is also called the optimization and improvement stage (RefinementStage); in the improvement stage, the imperfect pose transfer picture Iraw generated by the GAN network is targeted, and the diffusion model is used to generate Iraw again; the diffusion model control network uses the Controlnet structure. The input feature of ControlNet is C'. Iraw is passed into the image encoder (Encoder) to obtain Zt, whose dimension is (1*4*64*64). According to the diffusion control process, Z0 is obtained, and finally the final high-precision pose transfer effect picture (Ifinal) is output through the decoder;

[0015] The third-stage process: To train the diffusion process, the DDIM sampling method is used, and training is performed based on the Deep Fashion dataset until the network converges.

[0016] Furthermore, it also includes a VGG19 encoder, where the last fully connected layer is removed, and only the 7*7*512 feature map is output and combined with the subsequent convolutional layer as the input of ControlNet; the input dimension of the VGG encoder is 3*3*224*224; batch = 3 is the number of channels after reshaping and alignment of the human body mesh, human body bone key points, and face features output by the feature extraction network.

[0017] Furthermore, it also includes an encoder of the diffusion model, and the decoder adopts an autoencoder AE architecture; the autoencoder is a neural network based on unsupervised learning, aiming to reconstruct the input samples after dimensional compression by continuously adjusting parameters; the mapping from the input layer to the middle layer is called encoding, and the mapping from the middle layer to the output layer is called decoding. The autoencoder usually first obtains a compressed vector through encoding and then reconstructs it through decoding;

[0018] The encoder part of the autoencoder is y = enc(x); the decoder part is x' = dec(y)

[0019] The following loss function is used to train the autoencoder:

[0020] where N is the number of features.

[0021] Furthermore, the feature extraction network mainly consists of two parts: face detection and human body network SMPL model estimation;

[0022] Existing algorithms are used for face detection and human body SMPL model estimation. RetinaFace is used for face detection, and related methods such as MeshFormer and GraphCMR are used for SMPL model estimation.

[0023] Furthermore, the ControlNet network is a denoising depth model network in the diffusion model. The U-Net architecture introduces an additional control unit, and the original denoising depth model U-Net network ε θ (z t , t, τ) is adjusted to ε θ (z t , t, τ, c); where c is the control feature input; the control feature input passes through a "zero convolution" unit, is added to the input latent variable x = z t and is operated through the internal network units (convolutional layer, Transformer, etc. blocks) of the U-Net, and finally passes through the "zero convolution" unit to perform a merging operation with the output of the original U-Net network.

[0024] Furthermore, the ControlNet unit copies the neural structure of the internal network unit of U-Net as a trainable copy, while the original network retains the original weights without training, and finally superimposes and outputs to the result y; it should be noted that the "zero convolution" unit is actually a convolutional layer with both weights and biases being 0. When the network is initialized, W = 0, B = 0. For the input control feature map I ∈ R h×w×c , the forward propagation process of the "zero convolution" unit can be given in the following form:

[0025] Z(I; {W, B}) = I * W + B.

[0026] Furthermore, * is the convolution operation. To theoretically analyze the gradient update property of the "zero convolution" layer, for the sake of convenience, take the 1x1 convolution kernel as an example (n×n, n>1 and so on). For any position p in the convolutional feature space and channel position i, formula (5-14) can be rewritten in the following form:

[0027]

[0028] Furthermore, when the network is initialized, W = 0, B = 0. Therefore, when the input I p,i is non-zero, the gradients of each parameter are calculated backward as shown in the following formula:

[0029]

[0030] Furthermore, when the "zero convolution" unit performs gradient backpropagation, it does not have a gradient update effect of 0 on W and B; therefore, the weights of the network can be updated normally by backpropagating the gradient; for the loss function L and the learning rate lr, the update of the weights is given by the following formula:

[0031]

[0032] where is the Hadamard product (element-wise multiplication); for the input I p,i when it is non-zero, the "zero convolution" can normally perform backpropagation to update the network weights and enable the network to learn normally after the first gradient update; the significance of introducing the zero convolution unit in ControlNet is that for the input I p,i when it is equal to zero, the output result of the ControlNet side lobe network is 0 and will not affect the output of the original network. In addition, it can be considered that the ControlNet network is a finetune network, which introduces additional control features to partially guide the generation result of the diffusion model, but does not change the output of the original network to a large extent.

[0033] Compared with the prior art, the advantages of the present invention are as follows: The second-stage effect optimization improvement (Refinement Stage) is carried out on any model for pose transfer of the GAN network. The diffusion model is used for combination and improvement of the effect according to the first-stage result output by the GAN network. It makes up for the problems of low resolution and blurred details of the pictures generated by the GAN network, and at the same time overcomes the defects of instability and inconsistency of the content generated by the diffusion model. By making use of the advantages of both and compensating for their disadvantages, a high-precision pose transfer effect is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a system flowchart of the pose transfer algorithm of the present invention;

[0036] Figure 2 It is a structural diagram of the VGG19 encoder of the present invention;

[0037] Figure 3 It is an encoder diagram of the diffusion model of the present invention;

[0038] Figure 4 It is a schematic structural diagram of the ControlNet network control unit of the present invention;

[0039] Figure 5 It is a structural diagram of the original ControlNet network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0041] The following will describe the specific embodiments of the present invention in conjunction with the drawings:

[0042] As Figure 1 shown, a two-stage improved portrait pose transfer generation method includes

[0043] The first-stage process: First, given an RGB portrait I and a skeletal key-point feature map P, the existing GAN network pose generation algorithm is used based on these two features to obtain a preliminary pose transfer generation effect Iraw. The original image is passed into the feature extractor module to extract the human body Mesh feature Imesh and the face feature Iface. Combining with the previous skeletal feature map P, these three feature maps are concatenated to obtain the diffusion model control input C. The dimension of C is (3 * 3 * H * W), where H is the height of the feature map and W is the width. The control input feature map C is passed into the pre-trained VGG19 encoder to obtain the control input C', and then the Controlnet network is used to control the diffusion model to generate the final high-precision pose transfer effect diagram (Ifinal).

[0044] The second-stage process: The diffusion model for pose transfer is also called the RefinementStage. In the refinement stage, the imperfect pose transfer image Iraw generated by the GAN network is targeted, and the diffusion model is used to generate Iraw again. The diffusion model control network uses the Controlnet structure. The input feature of ControlNet is C'. Iraw is passed into the image encoder (Encoder) to obtain Zt, whose dimension is (1 * 4 * 64 * 64). Z0 is obtained according to the diffusion control process, and finally the final high-precision pose transfer effect diagram (Ifinal) is output through the decoder.

[0045] The third-stage process: For training the diffusion process, the DDIM sampling method is used, and training is performed based on the Deep Fashion dataset until the network converges.

[0046] Furthermore, it also includes a VGG19 encoder, as Figure 2 shown, where the last fully connected layer is removed, and only the 7 * 7 * 512 feature map is output and combined with the subsequent convolutional layer as the controlnet input. The input dimension of the VGG encoder is 3 * 3 * 224 * 224, and batch = 3 is the number of channels after reshaping and alignment of the human body mesh, human body skeletal key points, and face features output by the feature extraction network.

[0047] Furthermore, it also includes an encoder of the diffusion model, as Figure 3 shown. The decoder adopts the autoencoder AE architecture. The autoencoder is a neural network based on unsupervised learning, aiming to reconstruct the input samples after dimensional compression by continuously adjusting parameters. The mapping from the input layer to the middle layer is called encoding, and the mapping from the middle layer to the output layer is called decoding. The autoencoder usually first obtains a compressed vector through encoding and then reconstructs it through decoding.

[0048] The encoder part of the autoencoder is \(y = enc(x)\); the decoder part is \(x' = dec(y)\).

[0049] The autoencoder is trained using the following loss function:

[0050] where \(N\) is the number of features.

[0051] Furthermore, the feature extraction network mainly consists of two parts: face detection and human body network SMPL model estimation;

[0052] Face detection and human body SMPL model estimation use existing algorithms. RetinaFace is used for face detection, and related methods such as MeshFormer and GraphCMR are used for SMPL model estimation.

[0053] Furthermore, as Figure 4 、 5 shown, in the ControlNet network, in the denoising depth model network of the diffusion model, the U-Net architecture introduces an additional control unit, and the original denoising depth model U-Net network \(\epsilon\) θ (z t ,t,\(\tau\)) is adjusted to \(\epsilon\) θ (z t ,t,\(\tau\),c); where \(c\) is the control feature input; the control feature input passes through the "zero convolution" unit, is added to the input latent variable \(x = z\) t , and is operated through the internal network units (convolutional layer, Transformer, etc. blocks) of the U-Net, and finally passes through the "zero convolution" unit to merge with the output of the original U-Net network.

[0054] Furthermore, the ControlNet unit copies the neural structure of the internal network units of the U-Net as the trainable part, while the original network retains the original weights without training, and finally superimposes and outputs to the result \(y\); it should be noted that the "zero convolution" unit is actually a convolutional layer with both weights and biases equal to 0. When the network is initialized, \(W = 0\), \(B = 0\). For the input control feature map \(I\in R\) h×w×c , the forward propagation process of the "zero convolution" unit can be given in the following form:

[0055] \(Z(I;\{W,B\}) = I*W + B\).

[0056] Further, it is a convolution operation. For the purpose of theoretically analyzing the gradient update property of the "zero convolution" layer, for convenience, a 1x1 convolution kernel is taken as an example (n×n, n>1 and so on). For any position p in the convolution feature space and channel position i, formula (5-14) can be rewritten in the following form:

[0057]

[0058] Further, at the initial stage of the network, W = 0 and B = 0. Therefore, when the input I p,i is non-zero, the gradients of each parameter are calculated backward as shown in the following formula:

[0059]

[0060] Further, when the "zero convolution" unit performs gradient backpropagation, it does not have an update effect with a gradient of 0 on W and B; therefore, the weights of the network can be updated by normal gradient backpropagation; for the loss function L and learning rate lr, the update of the weights is given by the following formula:

[0061]

[0062] where is the Hadamard product (element-wise multiplication); for the input I p,i when it is non-zero, after the first gradient update, the "zero convolution" can normally perform backpropagation to update the network weights and enable the network to learn normally; the significance of introducing the zero convolution unit in ControlNet is that for the input I p,i when it is equal to zero, the output result of the side lobe network of ControlNet is 0 and does not affect the output of the original network. In addition, it can be considered that the ControlNet network is a fine-tuning network, which introduces additional control features to play a partial guiding role in the generation result of the diffusion model, but does not change the output of the original network to a large extent.

[0063] For the diffusion network model, first, the model portrait picture and the clothing segmentation picture are passed through the encoder structure. The encoder (Encoder) module consists of 3 convolutional layers and 3 deconvolutional layers in cascade. Here, we adopt a simple Enc-Dec structure. Given the image data of CxHxW, the final output is a feature map of size C1xH / 4xW / 4. We fuse the features of the above-mentioned encoded clothes and models through the feature fusion module. Perform Concat (a simple channel merge of C1+C2). And expand the merged features into a C1+C2-dimensional feature (flatten operation, merging w and h). At this time, the sampling point XT follows Normal distribution. According to the principle of the diffusion model, reverse sampling and the diffusion process are carried out using the following formula

[0064]

[0065] where beta is the diffusion coefficient at time t. Theta is the weight distribution of the diffusion model network.

[0066] Loss function

[0067] Given any time t, for the reverse network distribution We define the following loss function:

[0068]

[0069] where the first term is the latent variable generated by the diffusion process based on x0, and the second term is the feature learned by the reverse network. Take the square of the L2 norm; C is an irrelevant constant term.

[0070] Training principle

[0071] The training of the diffusion model is shown in the appendix and is briefly explained here. First, sample any time t from 0 to T. First, through the diffusion process, sample from x_0 to get

[0072]

[0073] Here, alpha is the coefficient.

[0074] Based on the q obtained from the diffusion process, then train the network P of the reverse procedure θ . By minimizing the mean squared error (MSE) or L2 norm between the latent variable at the current t - 1 moment and the output of the reverse network as the loss.

[0075] The beneficial effects of the present invention are as follows: By performing a second-stage effect optimization improvement (Refinement Stage) on any model for pose transfer of the GAN network. According to the first-stage result output by the GAN network, the diffusion model is used for combined improvement of the effect. It makes up for the problems of low resolution and blurred details of the pictures generated by the GAN network, and at the same time overcomes the defects of instability and inconsistency of the content generated by the diffusion model. By taking the advantages of both and making up for the disadvantages, a high-precision pose transfer effect is achieved.

[0076] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A two-stage improved portrait pose transfer generation method, characterized by: include The first stage process: First, given the RGB portrait I and the skeleton key point feature map P, these two features are passed into the existing GAN network posture generation algorithm to obtain the preliminary posture transfer generation effect Iraw; the original image is passed into the feature extractor module to extract the human body Mesh feature Imesh and the face feature Iface, combined with the previous skeleton feature map P, these three feature maps are concatenated to obtain the diffusion model control input C; the dimension of C is (3*3*H*W); H is the height of the feature map and W is the width; the control input feature map C is passed into the VGG19 pre-trained encoder to obtain the control input C' and then the Controlnet network is used to control the diffusion model to generate the final high-precision posture transfer effect map (Ifinal); The second stage process: used for the attitude transfer diffusion model is also called the refinement stage; In the improvement stage, the imperfect posture transfer image Iraw generated by the GAN network is regenerated using the diffusion model; the diffusion model control network uses the Controlnet structure. The input feature of ControlNet is C', and Iraw is passed to the image encoder (Encoder) to obtain Zt, whose dimension is (1*4*64*64). According to the diffusion control process, Z0 is obtained, and finally the final high-precision posture transfer effect image (Ifinal) is output through the decoder; The third stage process: To train the diffusion process, the sampling method uses DDIM sampling and trains according to the Deep Fashion dataset until the network converges.

2. The two-stage improved portrait posture transfer generation method according to claim 1, characterized in that: It also includes the VGG19 encoder, in which the last fully connected layer is removed, and only the 7*7*512 feature map is output and combined with the subsequent convolutional layer as the controlnet input; the input dimension of the VGG encoder is 3*3*224*224; batch=3 is the number of channels of the human mesh, human skeleton key points, and facial features reshaped and aligned output by the feature extraction network 3.

3. The two-stage improved portrait posture transfer generation method according to claim 2, characterized in that: It also includes the encoder of the diffusion model, and the decoder adopts the autoencoder AE architecture; the autoencoder is a neural network based on unsupervised learning, the purpose of which is to reconstruct the dimensionally compressed input samples by continuously adjusting parameters; the mapping between the input layer and the middle layer is called encoding, and the mapping between the middle layer and the output layer is called decoding. The autoencoder usually obtains the compressed vector by encoding first, and then reconstructs it by decoding; The encoder part of the autoencoder is y=enc(x); the decoder part is x'=dec(y) The autoencoder is trained using the following loss function: Where N is the number of features.

4. The two-stage improved portrait posture transfer generation method according to claim 3 is characterized in that: Feature extraction The network mainly consists of two parts: face detection and human network SMPL model estimation; Face detection and human SMPL model estimation use existing algorithms. Face detection uses RetinaFace, and SMPL model estimation is implemented using MeshFormer, GraphCMR and other related methods.

5. The two-stage improved portrait posture transfer generation method according to claim 4, characterized in that: The denoising deep model network of the ControlNet network in the diffusion model, the U-Net architecture introduces an additional control unit to convert the original denoising deep model U-Net network ε θ (z t ,t,τ) is adjusted to ε θ (z t ,t,τ,c); where c is the control feature input; the control feature input passes through the "zero convolution" unit and is combined with the input latent variable x=z t The outputs are added together and calculated through the internal network units of U-Net (convolutional layer, Transformer and other blocks), and finally passed into the "zero convolution" unit to be merged with the output of the original U-Net network.

6. The two-stage improved portrait posture transfer generation method according to claim 5, characterized in that: The ControlNet unit copies the neural structure of the U-Net internal network unit as a trainable part (trainable copy), while the original network retains the original weights without training, and finally superimposes the output to the result y; it is worth noting that the "zero convolution" unit is actually a convolution layer with both weight and bias 0. When the network is initialized, W = 0, B = 0. For the input control feature map I∈R h×w×c , the forward propagation process of the "zero convolution" unit can be given by the following form: Z(I;{W,B})=I*W+B.

7. The two-stage improved portrait posture transfer generation method according to claim 6, characterized in that: is the convolution operation. To theoretically analyze the back propagation gradient update properties of the "zero convolution" layer, for convenience, we take the 1x1 convolution kernel as an example (n×n, n>1 and so on). For any position p in the convolution feature space and channel position i, formula (5-14) can be rewritten as follows:

8. The two-stage improved portrait posture transfer generation method according to claim 7, characterized in that: When the network is initially connected, W = 0, B = 0, so when the input I p,i When it is non-zero, the gradient of each parameter is calculated in reverse as shown in the following formula:

9. The two-stage improved portrait posture transfer generation method according to claim 8, characterized in that: The "zero convolution" unit does not produce a zero gradient update effect on W and B during gradient backpropagation; therefore, the network weights can backpropagate the updated gradient normally; for the loss function L and the learning rate lr, the weight update is given by the following formula: Where. is the Hadamard product (element multiplication); for input I p,i When it is non-zero, "zero convolution" can perform backpropagation normally after the first gradient update to update the network weights and enable the network to learn normally; The significance of introducing zero convolution unit in ControlNet is that for input I p,i When it is equal to zero, the output of the ControlNet sidelobe network is 0, and it does not affect the output of the original network. In addition, it can be seen that the ControlNet network is a fine-tuned network, which introduces additional control features to partially guide the generation results of the diffusion model, but does not change the output of the original network to a great extent.

Citation Information

Patent Citations

  • Image style migration method based on diffusion loop generative adversarial network

    CN116433466A

  • Image generation method based on diffusion model and generative adversarial network

    CN116563399A

  • Virtual fitting generation method and system based on human skeleton postures

    CN117635883A

  • Latent prior embedded network for restoration and enhancement of images

    WO2024049441A1

  • KR20240136705A