Conditional diffusion model-based non-cross attention attitude migration method

By adopting a conditional diffusion model without cross attention mechanism and self-attention mechanism in posture migration, the problem of distortion of the details of the character image in the existing methods is solved, and the quality and consistency of the generated images are improved through the training method of independent vacancies condition information.

CN120070160APending Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510143844.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing pose transfer methods are prone to distortion of face and texture details when generating character images, and the probability of setting two conditions empty during training is too high, which affects the difference between the output of a single condition in the model learning model.

Method used

Using a cross-no-attention pose transfer method based on conditional diffusion model, the input image is uniformly encoded through a pre-trained VAE encoder, the complex relationships within the fusion feature are captured using the self-attention mechanism, and the condition information is independently nullified during training to promote the influence of network learning of a single condition.

Benefits of technology

The image generation quality and consistency of pose migration are improved, the problem of insufficient noise training is avoided, and the fidelity and consistency of generated images are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070160A_ABST
    Figure CN120070160A_ABST
Patent Text Reader

Abstract

The invention discloses a non-cross attention attitude migration method based on a conditional diffusion model, and the method comprises the steps: carrying out the unified coding of an input through a pre-trained VAE encoder, taking three parallel U-Nets as a basis, extracting the multi-level features of different inputs through different U-Nets, carrying out the direct addition at a trunk U-Net, and carrying out the feature fusion through a self-attention mechanism. During training, two condition images are randomly set to be empty with independent probabilities. During sampling, on the basis of prediction noise corresponding to a single condition, differences caused by the two conditions relative to the single condition to output are mixed. And the model learning difficulty is reduced through unified coding. And the self-attention mechanism calculates the relationship between different information in the fusion features, so that the high-correlation features are utilized more effectively. The independent null strategy promotes the network to learn differences introduced by a single condition. The mixed noise guides the basic noise to be corrected in a direction with stronger condition control, thereby avoiding insufficient noise training when the two conditions are both empty, and improving the fidelity and consistency of the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image processing, and particularly to conditional image generation and pose transfer technologies. Specifically, it is a cross-attention-free pose transfer method based on a conditional diffusion model. Background Art

[0002] Diffusion Models (DMs) are a widely used generative model. Conditional diffusion models are models that control the generated content by inputting additional conditional information based on general diffusion models. Conditional diffusion models are applied to conditional image generation tasks and have advantages such as high-quality generated images and stable training.

[0003] Pose Transfer is a conditional image generation task whose goal is to generate an image of a reference person in a target pose based on a given reference person image and a target pose image. Currently, conditional diffusion models perform well in pose transfer tasks. Pose transfer based on conditional diffusion models refers to using conditional diffusion models to achieve pose transfer tasks. This model has two conditions, namely the reference person image and the target pose image. Due to the diversity of the reference person image and the target pose image, existing pose transfer methods have problems such as distortion of the facial and texture details of the generated person image.

[0004] PIDM is the first pose transfer method based on diffusion models. The target pose image is input into the backbone network by splicing. The feature information of the reference person image is incorporated into the backbone network through the cross-attention mechanism, and the Disentangled Classifier-Free Guidance method is used to improve the consistency between the generated image and the reference person image. UPGPT is a Latent Diffusion Models (LDMs), which encodes the input image with a Variational Auto-Encoder (VAE) to obtain a latent vector. The model takes this latent vector as an input and predicts the output latent vector corresponding to the target image. In this model, the pose information is represented by SMPL parameters and the reinforced person mask (RPM). The reference person image is divided into nine sub-images according to different parts of the person and encoded by CLIP. At the same time, the text of the target image is also encoded by CLIP. The input of the network is the concatenation of the target latent vector and RPM. On this basis, the SMPL parameters, the encoded reference person image after division, and the encoded text of the target image are used as conditions and incorporated into the network through a Multimodal Fusing Block similar to the cross-attention mechanism to guide the generation of the target image. PCDMs adopt a three-stage training method from coarse to fine. The network parameters, the processing of conditional information, and the fusion method of conditional information used in each stage are different to meet the requirements of different stages: In the first stage, a Prior Conditional Diffusion Model is trained to predict the CLIP encoding of the target image; In the second stage, the output of the prior conditional diffusion model is decoded by the CLIP decoder and used as the network input to train an Inpainting Conditional Diffusion Model to predict a more accurate target image; In the third stage, the output of the inpainting conditional diffusion model is used to train a Refining Conditional Diffusion Model to predict texture details.

[0005] In the above methods, the target pose features are incorporated into the backbone network through direct addition or cross-attention mechanism, while the reference person features are all incorporated into the backbone network through the cross-attention mechanism. Since the backbone network contains the generated features and the target pose features, the cross-attention only calculates the relationship between the reference person features and the generated features or the target pose features. The model does not calculate the relationship between the generated features and the target pose, and cannot well associate the generated features with specific pose positions. In addition, these methods both set two conditions to be empty with a certain probability during training, which is not conducive to the model learning the differences brought by a single condition to the network output.

[0006] The present invention discloses a cross-attention-free pose transfer method based on a conditional diffusion model. This method efficiently fuses multi-level features of different inputs in a direct addition manner, and only uses the self-attention mechanism to capture the complex relationships inside the fused features, improving the consistency of pose transfer and the quality of image generation. Before inputting the reference person image, the target pose image, and the target image into the network, a pre-trained VAE encoder is first used for encoding to obtain a unified low-dimensional representation, so as to improve the training efficiency and reduce the learning difficulty of the model. During training, two condition information are independently set to be empty with a certain probability. This operation aims to promote the network to learn the differences introduced by a single condition information. During sampling, based on the predicted noise corresponding to a single condition, the differences caused by the two conditions relative to a single condition to the output are mixed, guiding the base noise to be corrected in the direction of stronger condition control, avoiding insufficient noise training when both conditions are set to be empty, thereby improving the fidelity and consistency of the generated images. Summary of the Invention

[0007] In view of this, an embodiment of the present invention provides a cross-attention-free pose transfer method based on a conditional diffusion model to achieve pose transfer of a reference person image.

[0008] Pose transfer is a conditional image generation task, which focuses on generating an image of a reference person in a target pose based on a given reference person image and a target pose image. Among them, the reference person image provides the appearance features of the person, while the target pose image provides the target pose information. The two are used as conditional inputs together to generate the target output image. In the pose transfer task based on the conditional diffusion model, U-Net is commonly used as the backbone network. This network takes the noisy target person image as the input and outputs the predicted noise. The reference person image and the target pose image are used as conditional information, and relevant features are obtained through the feature extraction stage and integrated into the backbone network in the feature fusion stage, so as to control the predicted noise of the output. In order to improve the network training efficiency and reduce the learning difficulty, VAE is usually used to encode the input image. In addition, for conditional information, features at different abstraction levels need to be extracted, and the output of the backbone network is accurately controlled at different levels through an efficient feature fusion method without cross-attention. During training, two conditional images are randomly set to be empty with independent probabilities. During sampling, based on the predicted noise corresponding to a single condition, the differences caused by the two conditions relative to the single condition to the output are mixed.

[0009] To achieve the above object, the embodiments of the present invention provide the following solutions:

[0010] A cross-attention-free pose transfer method based on a conditional diffusion model includes two stages: training and sampling.

[0011] For the training stage, it is characterized by including the following 6 steps:

[0012] Step 1. Set the parameters of the diffusion model and the network structure

[0013] This method is based on the diffusion model. The total number of time steps T is set to 1000, and a linear noise addition schedule is adopted, with β 1 = 0.00085, β T = 0.012. The backbone U-Net uses the Stable Diffusion model structure, but replaces the spatial cross-attention module in it with a self-attention module. Each stage of the side U-Net consists of a convolutional layer and a SiLU activation function. For both the backbone U-Net and the side U-Net, channel_mult = [1, 2, 4] is set, that is, the downsampling and upsampling parts are repeated 3 times. The basic number of channels is set to 192, and the number of channels corresponding to the 3 stages of downsampling is set to 1 times, 2 times, and 4 times the basic number of channels according to channel_mult respectively. For each stage, num_res_blocks = 2 is set.

[0014] Step 2. Scaling and normalizing the image

[0015] Scale the reference person image, target pose image, and target person image to 256×256. Then, normalize the values of the images to the range of -1 to 1.

[0016] Step 3. Randomly blank out the conditional images independently

[0017] Before each round of training, independently set the two conditional images to all-zero tensors with a probability of η%. This operation is beneficial for the network to learn the influence of different conditions on the output.

[0018] Step 4. Uniformly encode the input using the pre-trained VAE encoder

[0019] Use the pre-trained VAE encoder to encode the reference person image, target pose image, and target person image. The encoded target person image is denoted as z. Then, multiply these encodings by a scaling factor so that the encoding values are in the range close to -1 to 1.

[0020] Step 5. Mix standard Gaussian noise into the encoded target person image

[0021] Randomly sample a Gaussian noise ∈ from the standard Gaussian distribution t , and then obtain the noise-added encoded target person image according to the following formula:

[0022]

[0023] where t is a value randomly selected from a uniform distribution in the range of 1 to 1000, ∈ t obeys the standard Gaussian distribution.

[0024] Step 6. Forward propagation, backward propagation, and gradient descent of the network

[0025] First, freeze the parameters of the VAE encoder. The reference person image and target pose image are processed by the side U-Net1 and side U-Net2 respectively, while the noise-added encoded target person image is processed by the backbone U-Net. At each stage of the U-Net execution, the features extracted by the side U-Net are directly added to the same-level features of the backbone U-Net until the end of the last downsampling stage. After the post-processing stage of the backbone U-Net, the predicted noise is obtained to complete the forward propagation process. Finally, use the mean squared error between the predicted noise and the real noise as the loss, and perform backward propagation and gradient descent. Among them, the AdamW algorithm is used for gradient descent.

[0026] For the sampling stage, it is characterized by the following 9 steps:

[0027] Step 1. Read the checkpoint

[0028] First, read the checkpoint of the model and load the relevant parameters.

[0029] Step 2. Scaling and normalization of the images

[0030] Scale the sizes of the two conditional images to 256×256. Then, normalize their values to the range from -1 to 1.

[0031] Step 3. Set the initial value of sampling as standard Gaussian noise

[0032] Sample a Gaussian noise z from the standard Gaussian distribution T , which serves as the initial input to the backbone U-Net.

[0033] Step 4. Obtain different conditional inputs according to whether the conditional images are set to be empty

[0034] According to whether the two conditional images input to the network are set to be empty, four conditional inputs (s, p), (0, p), (s, 0), and (0, 0) are obtained.

[0035] Step 5. Encode the conditional images using the pre-trained VAE encoder

[0036] Use the VAE encoder to encode the reference person image and the target pose image. Then, multiply the two encodings by the same scaling factor as in the training stage.

[0037] Step 6. Obtain different predicted noises through forward propagation

[0038] At the current time step, perform forward propagation. Since four conditional inputs are obtained in Step 3, four different predicted noises are obtained: ∈ sp , ∈ s , ∈ p and ∈ un .

[0039] Step 7. Mix the noises

[0040] Utilize the specific mixing of ∈ s , ∈ p and ∈ sp to calculate the final predicted noise:

[0041]

[0042] where λ 1 and λ 2 are adjustment coefficients greater than or equal to 1. By adjusting λ 1 and λ 2 , the influence strengths of the reference person image and the target pose image on the generated image can be changed.

[0043] Step 8. Execute the DDIM sampling algorithm

[0044] According to the DDIM sampling algorithm, the following sampling formula is executed:

[0045]

[0046] where ∈ follows a standard Gaussian distribution, and

[0047]

[0048] In addition, τ is an increasing arithmetic sequence with a length of 50, and its last value is equal to T, that is, 1000. z T is the Gaussian noise obtained in step 3. Each time an iteration is executed, step 5 will be returned until the iteration ends to obtain the iteration result.

[0049] Step 9. Generate an image

[0050] The iteration result is decoded by the VAE decoder and then divided by the scaling factor to obtain the generated image. After the image is generated, depending on the usage, the image can be scaled to the size before preprocessing.

[0051] Preferably, the batch size during the training of the algorithm is set to 8.

[0052] Preferably, the algorithm is iterated 1,388,100 times during training.

[0053] Preferably, the initial learning rate during the training of the algorithm is 0.0001 and is set to 0.00001 after 925,400 iterations.

[0054] Preferably, η during the training of the algorithm is set to 15.

[0055] Preferably, λ during the sampling of the algorithm 1 = λ 2 = 4.

[0056] The present invention discloses a cross-attention-free pose transfer method based on a conditional diffusion model, aiming to efficiently achieve the pose transfer of a target person. This method uses a unified VAE to encode the reference person image, the target pose image, and the target image, and mixes standard Gaussian noise in the encoding of the target image. Subsequently, the encoded noisy target image is input into the backbone U-Net, while the encoded reference person image and the encoded target pose image are respectively input into two side U-Nets. The side U-Nets extract multi-level features and integrate them into the backbone U-Net in a direct addition manner to achieve feature fusion. Each stage of the side U-Net consists only of convolutional layers and activation functions, while each stage of the backbone U-Net contains residual blocks and self-attention layers, where the self-attention layers are used to capture the complex relationships within the fused features. In the training stage, by independently setting the encoded reference person image and the encoded target pose image to zero tensors, the network is prompted to learn the influence of a single condition on the output. In the sampling stage, depending on whether the two condition information is set to empty, four different predicted noises will be obtained. Since the proportion of training samples corresponding to both condition information being set to empty is small, to avoid image distortion and lack of consistency caused by insufficient learning, the mixture of the other three predicted noises is used as the prediction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0058] Figure 1 The structural diagram corresponding to the cross-attention-free pose transfer network based on the conditional diffusion model provided by the example of the present invention during training;

[0059] Figure 2 The structural diagram corresponding to the cross-attention-free pose transfer network based on the conditional diffusion model provided by the example of the present invention during sampling;

[0060] Figure 3 The training flowchart of the cross-attention-free pose transfer network based on the conditional diffusion model provided by the example of the present invention;

[0061] Figure 4 The network structure diagram of the backbone U-Net in the cross-attention-free pose transfer network based on the conditional diffusion model provided by the example of the present invention;

[0062] Figure 5The network structure diagram of the side U-Net in the cross-attention-free pose transfer network based on the conditional diffusion model provided by the examples of the present invention;

[0063] Figure 6 The network structure diagram of the ResBlock in the backbone U-Net provided by the embodiments of the present invention;

[0064] Figure 7 The network structure diagram of the Self-Attention in the backbone U-Net provided by the embodiments of the present invention;

[0065] Figure 8 The sampling flow chart of the cross-attention-free pose transfer network based on the conditional diffusion model provided by the examples of the present invention; Detailed implementation manners

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0067] Pose Transfer is a conditional image generation task, the goal of which is to generate an image of the reference person in the target pose based on the given reference person image and the target pose image. The key to this process is to jointly control the output of the network through the appearance of the reference person and the target pose features. The present invention uniformly encodes the input into a low-dimensional space and extracts multi-level features of different inputs. By only using the self-attention mechanism, the present invention realizes simple and effective feature fusion and pose transfer. As Figure 1 shown, the pose transfer network consists of a pre-trained VAE, a backbone U-Net, and two side U-Nets. The backbone U-Net is responsible for outputting the predicted noise, and the side U-Nets are used for multi-level feature extraction. The core idea of this method is to uniformly encode the input into a low-dimensional space, and different U-Nets extract multi-level features of different inputs. The multi-level features of the side U-Nets can be integrated into the backbone U-Net by direct addition. The backbone U-Net contains a self-attention mechanism for capturing the complex relationships inside the fused features. During sampling, as Figure 2 shown, this method obtains four different predicted noises by independently nullifying the conditional information input. The noises other than the predicted noises corresponding to the cases where both conditional informations are nullified are mixed, and a single-step denoising result is iteratively generated through the DDIM algorithm.

[0068] An embodiment of the present invention discloses a non-cross attention pose transfer method based on a conditional diffusion model, which is divided into two stages: training and sampling, where the sampling stage is the image generation stage.

[0069] The training stage is as Figure 3 shown. First, training data is randomly read, and the input images are scaled and normalized. After the conditional images are randomly emptied, they are encoded by the VAE encoder to uniformly map all images into a low-dimensional space. During this process, noise is added to the encoding of the target person's image. Then, the model performs forward propagation and feature fusion. After that, the model outputs the predicted noise, calculates the loss, and performs backpropagation and gradient descent until the model converges or reaches the maximum number of iterations. After training is completed, the checkpoint is saved.

[0070] See Figure 3 , the training stage of the above method includes the following 6 steps:

[0071] Step 1. Set the parameters of the diffusion model and the network structure

[0072] This method is based on a diffusion model. The total number of time steps T is set to 1000, and a linear noise addition schedule is adopted, with β 1 = 0.00085 and β T = 0.012. As shown in Figure 1, the backbone U-Net uses the Stable Diffusion model structure, but only has about one-third of the parameters, and replaces the spatial cross-attention module with a self-attention module. The backbone U-Net is divided into four stages: upsampling, middle, downsampling, and post-processing. Among them, both the upsampling and downsampling stages are composed of a residual block ResBlock, a self-attention module Self-Attention, a downsampling convolutional layer Conv(D), and an upsampling convolutional layer Conv(U). The post-processing stage includes group normalization, the SiLU activation function, and a convolutional layer. The ResBlock structure is as Figure 6 shown. It not only has a skip connection from the input to the output but also introduces time embedding information, where time_emb is obtained by encoding time t with sine and cosine and processed by a multi-layer perceptron. The structure of Self-Attention is as Figure 7 shown, and it is calculated using the multi-head attention method. The two U-Nets on the side are as Figure 5As shown, each stage consists of a convolutional layer and a SiLU activation function. For the backbone U-Net and the side U-Net, channel_mult = [1, 2, 4] is set. Since the length of channel_mult is 3, the downsampling and upsampling parts of all U-Nets are repeated 3 times. The base number of channels is set to 192, and the number of channels corresponding to the 3 stages of downsampling is set to 1 times, 2 times, and 4 times the base number of channels according to channel_mult respectively. For each stage, num_res_blocks = 2 is set.

[0073] Step 2. Image Scaling and Normalization

[0074] A training triplet consists of two conditional images and one target image. Specifically, the two conditional images refer to the reference person image and the target pose image, denoted as s and p respectively; the target image refers to the target person image that the network needs to generate. First, the image size is scaled to 256×256 by bicubic interpolation as needed. Then, the values of all input images are normalized to the range from -1 to 1.

[0075] Step 3. Randomly Blank the Conditional Images Independently

[0076] Before each round of training, the two conditional images are independently set to a tensor of all 0s with a probability of 15%. This operation is beneficial for the network to learn the influence of different conditions on the output.

[0077] Step 4. Uniformly Encode the Inputs with a Pre-trained VAE Encoder

[0078] As Figure 1 shown, the same pre-trained VAE encoder is used to encode the reference person image, the target pose image, and the target person image into encodings that are only 1 / 8 the size of the original. Among them, the encoding of the target person image is denoted as z. Using the same VAE encoder, the three inputs are encoded into the same low-dimensional Gaussian distribution. Then, these encodings are multiplied by a scaling factor so that the encoding values are in the range close to -1 to 1, aiming for numerical stability. Compared with using different encoders respectively, this scheme reduces the complexity of the inputs that the U-Net needs to process, making it easier for the network to learn in the subsequent stages.

[0079] Step 5. Mix Standard Gaussian Noise into the Encoding of the Target Person Image

[0080] As Figure 1 shown, for the encoding z of the target person image, a Gaussian noise needs to be mixed according to the following formula to obtain the noisy encoding z of the target person image t :

[0081]

[0082] where t is a value randomly selected from a uniform distribution with a range of 1 to 1000, ∈ t obeys a standard Gaussian distribution.

[0083] Step 6. Forward propagation, backpropagation, and gradient descent of the network

[0084] As Figure 1 shown, when training the pose transfer network, first freeze the parameters of the VAE encoder and only update the rest. The batch size is set to 8, and the AdamW optimization method is used. The learning rate is initially set to 0.0001 and then set to 0.00001 after 925400 iterations. The two conditional image encodings are respectively input into the side U-Net1 and side U-Net2; the noisy target image encoding z t is input into the backbone U-Net. Assume that the original feature maps generated by the backbone U-Net at a certain stage of upsampling, middle, or downsampling are F 1 , F 2 , …, F l , and the feature maps generated by side U-Net1 at this stage are P 1 , P 2 , …, P l , and the feature maps generated by side U-Net2 at this stage are Q 1 , Q 2 , …, Q l , then after directly adding the feature maps of the side U-Net to the corresponding feature maps of the backbone network, the fused feature M 1 , M 2 , …, M l is obtained, where M i = F i + P i + Q i , i = 1, 2, …, l. Since the fused feature contains the reference person image feature, pose feature, and the feature generated by the backbone U-Net, the Self-Attention module at each stage can capture the complex relationships between these information, thus improving the learning effect of the network. This process is carried out in the order of upsampling, middle, and downsampling stages, that is, addition is performed once for each stage of feature map calculation until all stages are completed, and the forward propagation process is completed. After execution, the predicted noise ∈ θ [z t , t, s, p] is obtained. Where θ is the parameter of the network, and s and p are determined by step 3 whether they are 0. Calculate the loss according to the following loss function:

[0085]

[0086] Among them, ∈t is the true mixed noise corresponding to the noisy target person image. After the loss calculation is completed, backpropagation and gradient descent are performed to update the network parameters. The above processes of forward propagation, backpropagation, and gradient descent need to be executed iteratively until the loss converges or the maximum number of iterations is reached. Finally, save the checkpoint of the model.

[0087] The sampling stage is as Figure 8 shown. Its core is to complete the pose transfer task of the target person by continuously denoising the input image using the predicted noise under the guidance of the input image. First, read the checkpoint saved in the training stage. Then, scale and normalize the input image. The conditional images can be selected to be empty or retained and encoded through the VAE. The input of the backbone network starts from a standard Gaussian noise and generates four different predicted noises after being processed by the model. Subsequently, iterative denoising is performed by combining the mixed noise strategy and the DDIM algorithm until the final image is generated.

[0088] See Figure 8 , the above sampling stage includes the following 9 steps:

[0089] Step 1. Read the checkpoint

[0090] First, read the checkpoint of the model and load the parameters of the diffusion model and the network.

[0091] Step 2. Scaling and normalization of the image

[0092] Scale the sizes of the two conditional images to 256×256 through bicubic interpolation. Then, normalize their values to the range from -1 to 1.

[0093] Step 3. Set the initial value of sampling as a standard Gaussian noise

[0094] Sample a Gaussian noise z T from the standard Gaussian distribution as the initial input of the backbone U-Net.

[0095] Step 4. Obtain different conditional inputs according to whether the conditional images are set to be empty

[0096] According to whether the two conditional images input to the network are set to be empty, four conditional inputs (s, p), (0, p), (s, 0), and (0, 0) are obtained.

[0097] Step 5. Encode the conditional images with the pre-trained VAE encoder

[0098] Use the same VAE encoder as in the training stage to encode the reference person image and the target pose image. Multiply the two encodings by the same scaling factor as in the training stage to keep the numerical ranges of the conditional image encodings consistent.

[0099] Step 6. Obtain different predicted noises through forward propagation

[0100] At the current time step, perform the forward propagation process. Since four conditional inputs are obtained in Step 3, four different predicted noises are obtained: ∈ sp , ∈ s , ∈ p and ∈ un , corresponding to ∈ θ [z t , t, s, p] obtained when both encodings are non-empty, ∈ θ [z t , t, s, 0] obtained when the reference person image encoding is non-empty, ∈ θ [z t , t, 0, p] obtained when the target pose image encoding is non-empty, and ∈ θ [z t , t, 0, 0] obtained under the condition that both encodings are empty.

[0101] Step 7. Mix the noises

[0102] Since the number of samples for training ∈u n only accounts for 15% × 15% = 2.25% of the total number of training samples, the proportion is too small. Therefore, to avoid the decline and instability of the fidelity and consistency of the generated images caused by insufficient training samples, ∈u n will not be used as the basic term of the classifier-free guidance method. The present invention uses ∈ s , ∈ p and ∈ sp to calculate the final predicted noise through a specific mixture:

[0103]

[0104] This method is called the mixed classifier-free guidance. Among them, λ 1 and λ 2 are adjustment coefficients greater than or equal to 1. This method avoids using ∈ un , and when both λ 1 and λ 2 are taken as 1, the obtained predicted noise is exactly ∈ sp , which means that the guidance gradually increases from 0 based on ∈ sp . By adjusting λ 1 and λ 2 , the influence strengths of the reference person image and the target pose image on the generated image can be changed specifically. Here, λ 1 = λ 2 = 4.

[0105] Step 8. Execute the DDIM sampling algorithm

[0106] According to the DDIM sampling algorithm, the following sampling formula is executed:

[0107]

[0108] where, ∈ follows a standard Gaussian distribution, and

[0109]

[0110] In addition, τ is an increasing arithmetic sequence with a length of 50, and its last value is equal to T, that is, 1000. z T is the Gaussian noise obtained in step 3. Each time an iteration is executed, step 5 will be returned until the iteration ends to obtain the iteration result.

[0111] Step 9. Generate an image

[0112] The iteration result is decoded by the VAE decoder and then divided by the scaling factor to obtain the generated image. After the image is generated, depending on the usage, the image can be scaled to the size before preprocessing.

[0113] The present invention uses the Deepfashion: In-shop Clothes Retrieval dataset to train and test the network. It is a subset of the DeepFashion dataset, containing a large number of images of the same person wearing the same clothing in different poses. According to the scheme provided by Zhu et al., 101,966 pairs of training data and 8,570 pairs of validation data are made, and the target pose images are generated by the DWPose model.

[0114] The present invention utilizes FID, LPIPS, SSIM, and PSNR as quantitative evaluation metrics. FID measures the similarity of the distributions of two image sets by calculating the means and variances of the feature vectors generated by the last pooling layer of the Inception v3 network for the generated images and the training images. The smaller the value, the closer the distribution of the generated images is to the distribution of the training images. LPIPS calculates the weighted average of the mean squared errors of the feature maps of two images at each layer of the VGG network. The smaller the value, the more similar the two images are under human judgment. SSIM calculates the similarity between the reconstructed image and the ground truth image from three aspects: the mean, standard deviation, and correlation coefficient of the entire image. PSNR is the ratio of the peak power of the signal to the mean squared error between the signal and the noise. The higher the SSIM and PSNR metrics, the better the image reconstruction quality. The PCDMs method uses a three-stage training, and each of the last two stages uses a complete Stable Diffusion model, with a huge number of parameters. The present invention discloses a quantitative comparison of the proposed method with CASD, UPGPT, and PIDM on the Deepfashion: In-shop Clothes Retrieval validation set. This comparison is carried out at an image size of 176×256. Therefore, the generated images are first scaled to the corresponding size using bicubic interpolation before the comparison. As shown in Table 1, the proposed method of the present invention reaches 0.7161 and 17.912 in the SSIM and PSNR metrics respectively, which is better than other models, indicating that the proposed method of the present invention has better performance in terms of structural similarity and pixel-level similarity. At the same time, in terms of the FID metric, the proposed method of the present invention obtains a score of 7.781, slightly higher than 6.440 of PIDM, but better than other models.

[0115] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined in the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in the present invention, but rather will conform to the widest scope consistent with the principles and novel features disclosed in the present invention.

[0116] Table 1 Experimental results of the method disclosed in the present invention

[0117]

Claims

1. A non-cross-attention posture transfer method based on a conditional diffusion model, characterized in that: It includes two stages: training and sampling. The training stage includes the following 6 steps: Step 1. Set the parameters of the diffusion model and network structure; This method is based on a diffusion model, with a total time step T set to 1000 and a linear noise scheme, β1 = 0.00085, β T =0.012; the backbone U-Net uses the Stable Diffusion model structure, and the spatial cross attention module is replaced by the self-attention module; each stage of the side U-Net consists of a convolutional layer and a SiLU activation function; for both the backbone U-Net and the side U-Net, channel_mult=[1,2,4] is set, that is, the downsampling and upsampling parts are repeated 3 times; the number of basic channels is set to 192, and the number of channels corresponding to the three stages of downsampling is set to 1 times, 2 times, and 4 times the number of basic channels according to channel_mult; for each stage, num_res_blocks=2; Step 2. Scaling and standardization of images; The reference person image, the target pose image, and the target person image are scaled to 256×256; then, the image values ​​are normalized to the range of -1 to 1; Step 3. Randomly blank the conditional images independently; Before each round of training, the two conditional images are independently set to all-0 tensors with a probability of η%; Step 4. Use the pre-trained VAE encoder to uniformly encode the input; Use the pre-trained VAE encoder to encode the reference person image, the target pose image and the target person image, and the target person image encoding is recorded as z; these encodings are multiplied by a scaling factor so that the encoding values ​​are close to the range of -1 to 1; Step 5. Encode the target person image with mixed standard Gaussian noise; Randomly sample a Gaussian noise ∈ from a standard Gaussian distribution t , and then the noisy target person image encoding is obtained according to the following formula: where t is a value randomly selected from a uniform distribution ranging from 1 to 1000. ∈ t Obey the standard Gaussian distribution; Step 6. Forward propagation, back propagation and gradient descent of the network; First, the parameters of the VAE encoder are frozen; the reference person image and the target pose image are processed by the side U-Net1 and side U-Net2 respectively, while the encoding of the noisy target person image is processed by the backbone U-Net; at each stage of the U-Net execution, the features extracted by the side U-Net are directly added to the features of the same level of the backbone U-Net until the end of the last downsampling stage; after that, the predicted noise is obtained through the post-processing stage of the backbone U-Net, completing the forward propagation process; finally, the average mean square error between the predicted noise and the real noise is used as the loss, and back propagation and gradient descent are performed; among them, the gradient descent uses the AdamW algorithm; The sampling phase consists of the following 9 steps: Step 1. Read the checkpoint; First, read the model's checkpoint and load the relevant parameters; Step 2. Scaling and standardization of images; The sizes of the two conditional images are scaled to 256×256; after that, their values ​​are normalized to the range of -1 to 1; Step 3. Set the initial value of the sampling to standard Gaussian noise; Sample a Gaussian noise z from a standard Gaussian distribution T , as the initial input of the backbone U-Net; Step 4. Get different conditional inputs by setting the conditional image to blank or not; Depending on whether the two conditional images of the input network are set to empty, four conditional inputs are obtained: (s, p), (0, p), (s, 0), and (0, 0); Step 5. Encode the conditional image using the pre-trained VAE encoder. Use the VAE encoder to encode the reference person image and the target pose image, and multiply the two codes by the same scaling factor as in the training phase; Step 6. Obtain different prediction noises through forward propagation; At the current time step, forward propagation is performed; by obtaining four conditional inputs, four different prediction noises are obtained: ∈ sp ,∈ s ,∈ p and ∈ un ; Step 7. Mixing Noise Take advantage of s ,∈ p and ∈ sp Calculate the final prediction noise for a specific mixture of: Among them, λ1 and λ2 are adjustment coefficients greater than or equal to 1; by adjusting λ1 and λ2, the influence of the reference person image and the target posture image on the generated image is changed; Step 8. Execute DDIM sampling algorithm; According to the DDIM sampling algorithm, the following sampling formula is executed: Among them, ∈ obeys the standard Gaussian distribution, and In addition, τ is an increasing arithmetic sequence of length 50, and its last value is equal to T, that is, 1000; z T is the Gaussian noise obtained in step 3; each time an iteration is performed, it will return to step 5 until the iteration is completed and the iteration result is obtained; Step 9. Generate image; The iterative result is decoded by the VAE decoder and then divided by the scaling factor to obtain the generated image. After the image is generated, the image is scaled to the size before preprocessing depending on the usage.

2. The method for non-cross-attention posture transfer based on conditional diffusion model as claimed in claim 1, characterized in that: The batch size during training is set to 8.

3. The method for non-cross-attention posture transfer based on conditional diffusion model as claimed in claim 1, characterized in that: A total of 1,388,100 iterations were performed during training.

4. The method for non-cross-attention posture transfer based on conditional diffusion model as claimed in claim 1, characterized in that: The initial learning rate during training was 0.0001 and was set to 0.00001 after 925400 iterations.

5. The method for non-cross-attention posture transfer based on conditional diffusion model as claimed in claim 1, characterized in that: The η during training is set to 15.

6. The method for non-cross-attention posture transfer based on conditional diffusion model as claimed in claim 1, characterized in that: During sampling, λ1=λ2=4.

Citation Information

Cited By

  • Conditional diffusion model-based digital human posture action generation method

    CN121304873A