A Human Pose Transfer Method Based on Image Redrawing

By independently processing the human body and background of the image in human body posture migration, and redrawing the human body image on the background image using the diffusion model and reference attention mechanism, the unreal details caused by the difference in the appearance and background of the human body in the existing methods are solved, and a highly consistent and authentic human body posture migration image generation is achieved.

CN119672794BActive Publication Date: 2025-06-20NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411234512.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-06-20
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

The existing human posture migration methods do not fully consider the obvious differences in the migration process between the human appearance and the background in the source image, resulting in the possibility of introducing unreal details in the generated image background.

Method used

The human body posture migration method based on image redrawing is adopted to independently process the human body and background parts of the image, and the human body image with specified poses and appearance is redrawn on the background image using the diffusion model. The human body appearance and background features are integrated through the reference attention mechanism to generate a highly consistent human body posture migration image.

Benefits of technology

It realizes the consistency of the background in the migration of human poses, generates high-quality images with real details, improves the authenticity and coherence of visual effects, and can replace the background area in the image while migrating.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672794B_ABST
    Figure CN119672794B_ABST
Patent Text Reader

Abstract

The present invention discloses a human pose transfer method based on image redrawing, belonging to the field of artificial intelligence image generation. The background and human appearance in the source image are separated through foreground and background separation technology; the pose encoder is used to obtain the target human pose features to guide the pose transfer of the source human body; rich human body detail features are obtained through the human appearance extraction network; the reference attention mechanism is used to guide the image redrawing network based on the diffusion model to effectively fuse the human appearance and background features; and the final generated image is output. The present invention independently processes the human body and background parts of the image and utilizes the powerful image generation ability of the diffusion model to optimize the naturalness and detail performance of human pose transfer. This method can maintain the consistency of the background during human pose transfer, while generating high-quality images with real details, improving the authenticity and coherence of the visual effect; in addition, it can also replace the background area in the image while realizing human pose transfer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence image generation, and particularly relates to a human pose transfer method based on image redrawing. Background Art

[0002] Human pose transfer is a challenging task in the field of computer vision. This task mainly refers to transferring the human body in the source human image from the source pose to an arbitrary target pose. Related technologies have great application value in artistic creation and the task of generating training data for downstream tasks, so it has become one of the most concerned research directions in recent years.

[0003] During the human pose transfer process, the source human image needs to undergo various changes including complex spatial changes and the generation of invisible regions. For the complex spatial changes during human pose transfer, early methods attempted to simulate this spatial change through geometric changes and flow changes to achieve human pose transfer, but these methods often perform poorly when faced with the transfer between complex poses. Subsequently, attention mechanisms were proposed to simulate this change. Thanks to the global vision of the attention mechanism, these methods can well achieve the transfer between complex poses. For the generation of invisible parts during human pose transfer, previous methods mainly relied on GANs. Recently, diffusion models have demonstrated remarkable image generation capabilities in the fields of text-to-image, image-to-image conversion, and image editing. Therefore, diffusion models began to be used to solve the problem of generating invisible regions in human pose transfer. For example, two human pose transfer methods, DreamPose and MagicDance, with the help of the powerful generation ability of the diffusion model, have shown excellent performance in the human transfer task. However, these methods do not take into account that there are obvious differences in the changes between the human body and the background part of the source human image during the human pose transfer process. Specifically, the background of the source human image usually only involves the generation of invisible regions during the entire human pose transfer process, but the human body part of the source human image involves complex spatial changes in addition to the generation of invisible regions, that is, the originally visible human body parts need to undergo complex geometric changes in terms of geometric position and shape. To improve the human appearance and pose effect of the generated image, these methods usually use neural networks to perform complex spatial transformations on the entire image. Although these transformations help with the pose conversion of human body parts, they may also have unexpected effects on the background area, thus introducing unrealistic details in the generated image background. A small number of methods, such as Disco, separate the human body and the background in the source human image and adopt different mechanisms for processing to avoid the impact of complex geometric transformations of human body parts on the background area, but due to the reason of its model design, there are still deficiencies in the consistent generation of human appearance and background. Summary of the Invention

[0004] In view of the fact that existing human pose transfer methods do not fully consider the obvious differences in changes between the human appearance and the background in the source image during the transfer process, the present invention proposes a method for human pose transfer that regards human pose transfer as an image redrawing task, and defines this process as an image redrawing-based human pose transfer method for redrawing a human image with a specified pose and appearance on a background image.

[0005] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0006] An image redrawing-based human pose transfer method independently processes the human body and background parts of the image, utilizes the image generation ability of the diffusion model, redraws a human image with a specified pose and appearance on the background image, generates a human pose transfer image with high consistency, and realizes the optimization of the naturalness and detail performance of human pose transfer. Specifically, it includes the following steps:

[0007] S1: Separate the background and human appearance in the source image through foreground-background separation technology;

[0008] S2: Use a pose encoder to extract the target human pose features to accurately guide the source human body for pose transfer;

[0009] S3: Obtain rich human body detail features through a human appearance extraction network;

[0010] S4: Use a reference attention mechanism to guide the image redrawing network based on the diffusion model, and effectively fuse human appearance and background features;

[0011] S5: Output the final generated image.

[0012] Preferably, in step S1, the specific content of separating the background and human appearance in the source image through foreground-background separation technology is:

[0013] First, use the DINO model and the SAM model to obtain a mask M of the corresponding human body region from the source human body image, and then use the mask M of the human body region to separate the source human body appearance and the background, obtaining the source human body appearance image I App and the corresponding background image I Back .

[0014] Preferably, in step S2, the specific content of using a pose encoder to extract the target human pose features to accurately guide the source human body for pose transfer is:

[0015] First, input the target human pose image I pose into a lightweight pose encoder composed of 5 convolutional layers and the corresponding activation function SiLU to obtain the corresponding target human pose feature map F pose , and finally the target human pose feature map Fpose They are respectively input into the human body appearance extraction network and the image redrawing network based on the diffusion model to accurately guide the migration of the source human body posture. Through the accurate posture feature extraction in step S2, a solid foundation is laid for the subsequent recovery of human body appearance details and posture migration.

[0016] Preferably, in step S3, the specific content of obtaining rich human body detail features through the human body appearance extraction network is as follows:

[0017] S3.1: Introduce a unet based on ResnetBlock and TransformerBlock into the human body appearance extraction network. This unet specifically includes an input layer composed of convolutional layers and a feature extraction module composed of several ResnetBlock and TransformerBlock.

[0018] S3.2: Through the VAE encoder part of Stable diffusion, map the human body appearance image I App from the pixel space to the low-dimensional latent space as the human body appearance latent space variable ε(I App ).

[0019] S3.3: Use the random noise Z t and the human body appearance latent space variable ε(I App ) and the mask M of the scaled human body area after channel concatenation as the input of the human body appearance extraction network. The specific method is as follows:

[0020]

[0021] Among them, γ represents the input of the human body appearance extraction network, Z t represents the random noise, M represents the mask of the human body area, I App represents the source human body appearance image, ε(I App ) represents the human body appearance latent space variable after I App is mapped to the latent space, represents that the final input dimension is 9×h×w, where h = H / 8, w = W / 8, and H and W are the height and width of the original input image respectively.

[0022] S3.4: Add the human target body posture feature map F pose obtained in step S2 to the output of the input layer of the human body appearance extraction network by matrix element addition and use it as the input of the feature extraction module, so as to achieve the preliminary guidance of the human body posture and realize the small-range transformation of the human body posture. The specific method is as follows:

[0023]

[0024] Among them, μ represents the input of the feature extraction module, and F pose represents the human body pose feature map of the human target obtained in S2, and F App represents the output of the input layer of the human appearance extraction network. It means that the final input dimension is 320×h×w, where h = H / 8, w = W / 8, and H and W are the height and width of the original input image respectively.

[0025] S3.5: Use the key K of the self-attention layer in the TransformerBlock of the feature extraction module App and the value V App as the extraction of fine-grained human appearance features and input them into the image redrawing network based on the diffusion model.

[0026] Preferably, in S4, the reference attention mechanism is used to guide the image redrawing network based on the diffusion model. The specific content of effectively fusing the human appearance and background features is as follows:

[0027] S4.1: Introduce a unet based on ResnetBlock and TransformerBlock as the image redrawing network based on the diffusion model. This unet specifically includes an input layer composed of convolutional layers and a feature fusion module composed of several ResnetBlocks and TransformerBlocks, where the self-attention layer in the TransformerBlock is replaced by a reference attention layer.

[0028] S4.2: Through the VAE encoder part of Stable diffusion, map the background image I Back from the pixel space to the low-dimensional latent space as the background latent space variable ε(I Back ).

[0029] S4.3: Use the random noise Z t and the human appearance latent space variable ε(I Back ) and the mask M of the scaled human body area after channel concatenation as the input of the image redrawing network. The specific method is as follows:

[0030]

[0031] Among them, τ represents the input of the image redrawing network, Z t represents the random noise, M represents the mask of the human body area, I Back represents the source human appearance image, and ε(I Back ) represents the background latent space variable after I Back is mapped to the latent space. It is indicated that the final input dimension is 9×h×w, where h = H / 8, w = W / 8, and H and W are the height and width of the original input image respectively.

[0032] S4.4: Add the human body pose feature map F obtained in step S2 pose to the output of the input layer of the image redrawing network by matrix element addition and use it as the input of the subsequent feature fusion module, which provides precise guidance for the fusion of subsequent human appearance features and background features, ensures that the human appearance is introduced into the specified area of the background image, and realizes the precise transfer of human poses. The specific calculation method is as follows:

[0033]

[0034] where, represents the input of the feature extraction module, F pose represents the human body pose feature map of the human target obtained in S2, F Back represents the output of the input layer of the image redrawing network, It is indicated that the final input dimension is 320×h×w, where h = H / 8, w = W / 8, and H and W are the height and width of the original input image respectively.

[0035] S4.5: With the help of the reference attention mechanism, guide the fine-grained human appearance features extracted by the human appearance extraction network into the specified area of the background image to realize the fusion of human appearance features and background features, so as to realize the redrawing of the background image and then generate the final human pose transfer image. The calculation method of the reference attention is as follows:

[0036]

[0037] where Referrnce_Attn represents the calculation result of the reference attention, Softmax() represents the sofmax function, Q Back , K Back value V Back respectively represent the query, key and value of the background features in the image redrawing network, K App and V App represent the key and value of the human appearance features in the human appearance extraction network, [;] represents the feature splicing operation, · represents the vector multiplication, and d represents the number of channels of the Q Back channel dimension.

[0038] S4.6: After completing the feature fusion of the human appearance and the background with the help of the reference attention mechanism, the image redrawing network based on the diffusion model completes the redrawing of the human image by continuously iterating to denoise and generates the final virtual try-on image.

[0039] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0040] (1) The human pose transfer method based on image redrawing proposed by the present invention can generate highly consistent human pose transfer images by redrawing a human body image with a specified pose and appearance on a background image.

[0041] (2) The present invention uses a human appearance extraction network to extract multi-scale human appearance features with rich details from the source human appearance image, which prompts the image redrawing network to generate a target human appearance that is more consistent with the source human appearance, making the generated image more realistic and natural.

[0042] (3) The present invention uses a carefully designed reference attention mechanism to fully fuse the human appearance features extracted from the appearance extraction network and the background features in the image redrawing network. This reference attention mechanism can accurately introduce the human appearance features into specific regions of the background image under the guidance of the target pose, enabling the model to generate pose transfer images that better meet the requirements of the target pose.

[0043] (4) Thanks to the decoupling of the human appearance and the background, the present invention can replace the background area in the image while realizing human pose transfer. Brief Description of the Drawings

[0044] Figure 1 is the overall flowchart of the present invention;

[0045] Figure 2 is the detailed network architecture diagram of the present invention;

[0046] Figure 3 is the human pose transfer effect diagram of the present invention under the TikTok dataset;

[0047] Figure 4 is the effect diagram of simultaneously realizing human pose transfer and background replacement of the present invention under the TikTok dataset. Detailed Embodiments

[0048] The following further clarifies the present invention in conjunction with specific embodiments. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0049] As Figure 1-2 shown, the human pose transfer method based on image redrawing provided in this embodiment includes the following steps:

[0050] S1: Separate the background and human appearance in the source image through foreground and background separation technology;

[0051] First, use the DINO (DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection) model and the SAM (Segment Anything Model) model to obtain the mask M of the corresponding human body region from the source human body image, and then use the mask M of the human body region to separate the source human body appearance and the background, obtaining the source human body appearance image I App and the corresponding background image I Back .

[0052] S2: Use the pose encoder to extract the target human body pose features to precisely guide the pose transfer of the source human body;

[0053] First, input the target human body pose image I pose into a lightweight pose encoder composed of 5 convolutional layers and the corresponding activation function SiLU to obtain the corresponding target human body pose feature map F pose , and finally input the target human body pose feature map F pose into the human body appearance extraction network and the image redrawing network based on the diffusion model respectively to precisely guide the pose transfer of the source human body.

[0054] S3: Obtain rich human body detail features through the human body appearance extraction network;

[0055] S3.1: Introduce a unet based on ResnetBlock and TransformerBlock as the human body appearance extraction network. This unet specifically includes an input layer composed of convolutional layers and a feature extraction module composed of several ResnetBlocks and TransformerBlocks.

[0056] S3.2: First, through the VAE encoder part of Stable diffusion, map the human body appearance image I App from the pixel space to the low-dimensional latent space as the human body appearance latent space variable ε(I App ).

[0057] S3.3: Concatenate the random noise Z t and the human body appearance latent space variable ε(I App ) and the scaled mask M of the human body region in the channel dimension as the input of the human body appearance extraction network. The specific method is as follows:

[0058]

[0059] where γ represents the input of the human body appearance extraction network, Z tdenotes random noise, M denotes the mask of the human body region, and I App denotes the source human body appearance image, and ε(I app ) denotes the human body appearance latent space variable after I App is mapped to the latent space, indicating that the final input dimension is 9×h×w, where h = H / 8 and w = W / 8, and H and W are the height and width of the original input image respectively.

[0060] S3.4: Add the human body target human pose feature map F obtained in step S2 pose to the output of the input layer of the human body appearance extraction network by matrix element addition and use it as the input of the feature extraction module, so as to achieve the preliminary guidance of the human body pose and realize a small range transformation of the human body pose. The specific method is as follows:

[0061]

[0062] where μ represents the input of the feature extraction module, and F pose represents the human body target human pose feature map obtained in step S2, and F App represents the output of the input layer of the human body appearance extraction network, indicating that the final input dimension is 320×h×w, where h = H / 8 and w = W / 8, and H and W are the height and width of the original input image respectively.

[0063] S3.5: Use the key K of the self-attention layer in the TransformerBlock of the feature extraction module App and the value V App as the extraction of fine-grained human body appearance features and input them into the image redrawing network based on the diffusion model.

[0064] S4: Use the reference attention mechanism to guide the image redrawing network based on the diffusion model to effectively fuse human body appearance and background features;

[0065] S4.1: Introduce a unet based on ResnetBlock and TransformerBlock as the image redrawing network based on the diffusion model. This unet specifically includes an input layer composed of convolutional layers and a feature fusion module composed of several ResnetBlocks and TransformerBlocks, where the self-attention layer in the TransformerBlock is replaced by a reference attention layer.

[0066] S4.2: First, through the VAE encoder part of Stable diffusion, map the background image I Back from the pixel space to the low-dimensional latent space as the background latent space variable ε(IBack )。

[0067] S4.3: Concatenate the random noise Z t and the human appearance latent space variable ε(I Back ) and the mask M of the scaled human region after channel concatenation as the input of the image redrawing network, and the specific method is as follows:

[0068]

[0069] where τ represents the input of the image redrawing network, Z t represents the random noise, M represents the mask of the human region, I Back represents the source human appearance image, and ε(I Back ) represents the background latent space variable after I Back is mapped to the latent space. It is shown that the final input dimension is 9×h×w, where h = H / 8, w = W / 8, and H and W are the height and width of the original input image respectively.

[0070] S4.4: Add the human target human pose feature map F obtained in S2 pose to the output of the input layer of the image redrawing network by matrix element addition and use it as the input of the subsequent feature fusion module. It provides precise guidance for the fusion of subsequent human appearance features and background features, ensuring that the human appearance is introduced into the specified area of the background image, realizing the precise transfer of human poses, and its specific calculation method is as follows:

[0071]

[0072] where, represents the input of the feature extraction module, F pose represents the human target human pose feature map obtained in step S2, and F Back represents the output of the input layer of the image redrawing network. It is shown that the final input dimension is 320×h×w, where h = H / 8, w = W / 8, and H and W are the height and width of the original input image respectively.

[0073] S4.5: Use the reference attention mechanism to guide the fine-grained human appearance features extracted by the human appearance extraction network into the specified area of the background image, realize the fusion of human appearance features and background features, so as to redraw the background image, and then generate the final human pose transfer image. The calculation method of the reference attention mechanism is as follows:

[0074]

[0075] Among them, Reference_Attn represents the calculation result of reference attention, Softmax() represents the softmax function, and Q Back , K Back value V Back respectively represent the query, key, and value of the background features in the image redrawing network. K App and V App represent the key and value of the human appearance features in the human appearance extraction network. [;] represents the feature splicing operation, · represents the vector multiplication, and d represents the number of channels of Q Ba in the c dimension.

[0076] S4.6: After completing the feature fusion of the human appearance and the background with the help of the reference attention mechanism, the image redrawing network based on the diffusion model completes the redrawing of the human image by continuously iterating to denoise, and generates the final virtual try-on image.

[0077] S5: Output the finally generated image.

[0078] Verify the effectiveness and efficiency of the method of the present invention through the following experiments:

[0079] The present invention uses SSIM, LPIPS, FID, and PSNR as evaluation criteria:

[0080] SSIM: An index used to evaluate the image quality, which measures the similarity between two images by comparing the brightness, contrast, and structural similarity of the images.

[0081] LPIPS: Used to measure the perceptual difference between images, calculates the similarity between two images by learning the feature representation of the neural network, and can more accurately capture the differences in human visual perception.

[0082] FID: Used to evaluate the similarity between the generated image and the real image, and measures the quality of the generated image by calculating the difference in the feature distribution of the intermediate layer of the neural network.

[0083] PSNR: Peak signal-to-noise ratio, which is a method for measuring image quality. It is calculated by comparing the peak value of the denoised image with the root mean square of the noise. The higher the PSNR value, the better the quality of the denoised image.

[0084] The dataset selected for this invention is the TikTok dataset, which contains 350 single-person dance videos collected from the TikTok platform. Most of the videos contain faces and the corresponding upper parts of the faces. By extracting frames from the videos, 94,908 pairs of human body images can be obtained. Each pair of human body images consists of two human body images with the same background and human appearance but different poses. Among them, the training set includes 90,785 pairs of human body images, and the test set is 4,123 pairs of human body images. The resolution of each image is 512×512 pixels.

[0085] This invention uses the TikTok dataset. Under the same resolution condition (512×512 pixels), the results of the method of this invention are compared with methods such as FOMM, TPS, Disco, and MagicAnimate. Among these methods, FOMM and TPS are human pose transfer methods based on generative adversarial networks (GANs). Disco and MagicAnimate are the latest human pose transfer methods based on diffusion models. Among them, Disco also decouples the human appearance and the background to replace the background of the image while performing human pose transfer. With the help of the image decoupling operation, the Disco method can also avoid the mutual interference of different transformations in the background and human appearance regions to a certain extent. MagicAnimate ensures the consistency of the image content before and after pose transfer through a carefully designed content control mechanism.

[0086] The quantitative experimental results of model comparison under the TikTok dataset are as follows:

[0087] Table 1 Quantitative experimental results of this invention under the TikTok dataset

[0088] Method FID↓ SSIM↑ PSNR↑ LPIPS↓ FOMM 85.03 0.648 29.01 0.335 TPS 53.78 0.673 29.18 0.299 Disco 30.75 0.668 29.03 0.292 MagicAnimate 32.09 0.714 29.16 0.239 Ours (This invention) 30.56 0.734 29.38 0.298

[0089] According to the results in Table 1, the method of this invention is significantly better than other comparison methods in multiple evaluation metrics such as FID, SSIM, and PSNR, and also shows a competitive performance in the LPIPS metric. Compared with FOMM and TPS based on generative adversarial networks, Disco, MagicAnimate, and the method of this invention using the diffusion model as the underlying architecture show better effects.

[0090] Although both MagicAnimate and the present invention adopt a diffusion model as the core architecture and introduce a carefully designed appearance control mechanism, due to ignoring the different transformations experienced by the background and the human appearance region when dealing with human pose transfer, its performance in multiple key metrics is inferior to that of the present invention. In particular, in terms of the FID reflecting image realism and the SSIM of image similarity, the method of the present invention demonstrates its obvious advantages. It is worth mentioning that although the Disco method also adopts a decoupling strategy to reduce the mutual interference between the background and the human appearance region transformations, the present invention has achieved significantly leading results in metrics such as FID, SSIM, and PSNR through a more powerful human appearance extraction network and a carefully designed reference attention mechanism. This achievement not only proves the effectiveness of the method of the present invention but also demonstrates its superior ability in dealing with complex pose transfer tasks.

[0091] Figure 3 The visualization results of the method of the present invention and other comparative methods on the TikTok dataset are shown. In the TikTok dataset, the results of the present invention are compared with the TPS, Disco, and MagicAnimate methods. Obviously, compared with other methods, the method of the present invention performs better in terms of image consistency and precise pose control. When performing complex pose transfer, it is difficult for TPS, Disco, and MagicAnimate to generate reasonable facial and hand poses, and it is also difficult to retain the appearance details of the source person, often generating unrealistic human pose transfer images.

[0092] Figure 4 The effect diagrams of simultaneously achieving human pose transfer and background replacement by the Disco and the method of the present invention under the TikTok dataset are shown. It can be seen from the figure that although both the Disco and the method of the present invention adopt the method of separating the image background and human appearance to avoid the mutual interference between different transformations of the background and human appearance regions, thanks to the image redrawing technology and the carefully designed reference attention mechanism, the method of the present invention can effectively replace the background image while accurately transferring the human pose, generating higher-quality human images.

[0093] The present invention independently processes the human body and background parts of the image and utilizes the powerful image generation ability of the diffusion model to optimize the naturalness and detail performance of human pose transfer. This method can maintain the consistency of the background during human pose transfer, and at the same time generate high-quality images with real details, improving the realism and coherence of the visual effect. In addition, this method can also replace the background region in the image while achieving human pose transfer.

[0094] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A human body posture transfer method based on image redrawing, characterized in that: The human body and background parts of the image are processed independently, and the image generation capability of the diffusion model is used to redraw the human body image with the specified posture and appearance on the background image to generate a human body posture migration image with high consistency, thereby optimizing the naturalness and detail expression of human body posture migration; The specific steps include: S1: Separate the background and human appearance in the source image through foreground and background separation technology; S2: Use the posture encoder to extract the target human posture features to guide the source human posture transfer; S3: Obtain detailed features of the human body through the human appearance extraction network; the specific method is: S3.1: Use unet based on ResnetBlock and TransformerBlock as the human appearance extraction network; S3.2: Through the VAE encoder part of the stable diffusion, the human appearance image Mapping from pixel space to low-dimensional latent space as latent space variables for human appearance ; S3.3: Concatenate the random noise, the human appearance latent space variables and the scaled human region mask channel and input them into the human appearance extraction network; S3.4: The human body posture feature map of the human target obtained in step S2 By adding matrix elements, it is added to the output of the input layer of the human appearance extraction network and used as the input of the feature extraction module to achieve preliminary guidance of human posture and small-scale transformation of human posture; S3.5: Add the key of the self-attention layer in the TransformerBlock of the feature extraction module Sum As the extraction of fine-grained human appearance features, and input into the image redrawing network based on the diffusion model; S4: Using reference attention mechanism to guide diffusion model-based image repainting network to integrate human appearance and background features; S5: Output the final generated image.

2. The method for human body posture transfer based on image redrawing according to claim 1, characterized in that: In S1, the background and human appearance in the source image are separated using the foreground and background separation technology. The specific method is as follows: First, use the DINO model and SAM model to obtain the mask of the corresponding human area from the source human image M , and then use the mask of the human body area M To separate the source human appearance and background, get the source human appearance image and the corresponding background image .

3. The method for human posture transfer based on image redrawing according to claim 1, characterized in that: In step S2, the posture encoder is used to extract the posture features of the target human body to guide the source human body to perform posture migration. The specific process is: First, the target human posture image Input into the lightweight posture encoder to obtain the corresponding target human posture feature map , and then the target human posture feature map They are respectively input into the human appearance extraction network and the diffusion model-based image redrawing network to guide the migration of the source human posture.

4. The method for human body posture transfer based on image redrawing according to claim 1, characterized in that: The specific implementation of step S3.3 is as follows: , in, represents the input of the human appearance extraction network, represents random noise, M represents the mask of the human body area, represents the source human appearance image, express The human appearance latent space variables after being mapped to the human appearance latent space, Indicates that the final input dimension is .

5. The method for human body posture transfer based on image redrawing according to claim 1, characterized in that: The specific implementation of step S3.4 is as follows: , in, represents the input of the feature extraction module, Represents the human body posture feature map of the human target obtained in step S2 , represents the output of the input layer of the human appearance extraction network, Indicates that the final input dimension is ,in =H / 8, =W / 8, H and W are the height and width of the original input image respectively.

6. The method for human body posture transfer based on image redrawing according to claim 1, characterized in that: In step S4, the reference attention mechanism is used to guide the diffusion model-based image redrawing network to effectively fuse the human appearance and background features. The specific content is: S4.1: Use unet based on ResnetBlock and TransformerBlock as the image redrawing network based on the diffusion model; S4.2: Through the VAE encoder part of the stable diffusion, the background image Mapping from pixel space to low-dimensional latent space as background latent space variables ; S4.3: Concatenate the random noise, the human appearance latent space variables and the scaled human region mask channel as the input of the image redrawing network; S4.4: The human body posture feature map of the human target obtained in step S2 It is added to the output of the image redrawing network input layer by adding matrix elements and used as the input of the subsequent feature fusion module; S4.5: Using the reference attention mechanism, the fine-grained human appearance features extracted by the human appearance extraction network are guided into the specified area of ​​the background image to achieve the fusion of human appearance features and background features, so as to achieve the redrawing of the background image and generate the final human posture transfer image; S4.6: After completing the feature fusion of human appearance and background with the help of the reference attention mechanism, the image redrawing network based on the diffusion model completes the redrawing of the human image through continuous iterative denoising to generate the final virtual try-on image.

7. The method for human body posture transfer based on image redrawing according to claim 6, characterized in that: The specific implementation of step S4.3 is as follows: , in, represents the input of the image redrawing network, represents random noise, M The mask representing the human body region, Represents the corresponding background image, express The background latent space after mapping to the latent space, variable , indicating that the final input dimension is ,in =H / 8, =W / 8, H and W are the height and width of the original input image respectively.

8. The method for human body posture transfer based on image redrawing according to claim 6, characterized in that: The specific implementation of step S4.4 is as follows: , in, represents the input of the feature extraction module, represents the human body posture feature map of the human target obtained in step S2, represents the output of the input layer of the image redrawing network, Indicates that the final input dimension is ,in =H / 8, =W / 8, H and W are the height and width of the original input image respectively.

9. The method for human body posture transfer based on image redrawing according to claim 6, characterized in that: In step S4.5, the reference attention mechanism is calculated as follows: , in, represents the calculation result of reference attention, represents the sofmax function, , value Represent the query, key and value of the background feature in the image redrawing network, respectively. and represents the key and value of the human appearance features in the human appearance extraction network, represents the feature concatenation operation, represents vector multiplication, express The number of channel dimensions.

Citation Information

Patent Citations

  • Human body image generation method based on UV spatial transformation

    CN116071831A

  • Virtual fitting method based on diffusion model

    CN118278291A