Pose Guidance Image Synthesis Method, Device, Electronic Device and Storage Medium

By adopting a diffusion model-based method in posture-guided image synthesis, embedding multiple feature encodings to improve image quality, the problems of insufficient image fidelity and high hardware requirements in the prior art are solved, and high-quality image synthesis and resource efficiency are improved.

CN119169129BActive Publication Date: 2025-06-20GUANGZHOU ZIWEIYUN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411257994.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-06-20
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

In the existing pose-guided character image synthesis technology, the image fidelity is insufficient, the in-depth understanding of the advanced semantics of the character image is lacking, and the model parameters are large, the inference speed is slow, and the hardware requirements are high, making it difficult to implement.

Method used

Using the pose-guided image synthesis method based on the diffusion model, the pose feature map, multi-scale feature map and multi-scale refinement appearance feature generation conditional encoding is embedded through the encoder and decoder in the Unet module, and the conditional encoding of the appearance feature is gradually introduced to gradually introduce richer detailed information and ensure the consistency and consistency of appearance information.

Benefits of technology

Improve the overall quality of the generated images, reduce artifacts and color inconsistencies in the composite images, reduce the use of computing resources, and make the model easier to use in products and land.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169129B_ABST
    Figure CN119169129B_ABST
Patent Text Reader

Abstract

The present application relates to a pose-guided image synthesis method, apparatus, electronic device, and storage medium, including: extracting a pose feature map to generate a first conditional encoding and embedding it into each downsampling module; extracting a multi-scale feature map to generate a second conditional encoding and embedding it into each upsampling module; obtaining a multi-scale refined appearance feature to generate a third conditional encoding and embedding it into each downsampling module and each upsampling module; inputting a noise image into a Unet model and synthesizing a target image based on the first, second, and third conditional encodings. By embedding the first conditional encoding into the downsampling module, embedding the second conditional encoding into the upsampling module, and embedding the third conditional encoding into both the upsampling and downsampling modules, the model gradually introduces richer detail information at different stages of image reconstruction, thereby generating high-quality images with specific poses. At the same time, based on the diffusion model, the model reduces the use of computing resources while ensuring image quality and is more easily applied to products and implementation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image synthesis, and more particularly, to a pose-guided image synthesis method, apparatus, electronic device, and storage medium. Background Art

[0002] Pose-guided human image synthesis is a computer vision and deep learning technology that aims to generate new portrait images that conform to the target pose based on preset human poses and conditional images. In this process, the model not only accurately adjusts the limb postures of the human body, but also preserves and integrates visual features such as clothing details and facial features, ensuring that the identity and other specific attributes of the person remain unchanged while changing the pose. Compared with the traditional image shooting and post-editing processing workflows, pose-guided human image synthesis can quickly generate a large number of images with different poses, greatly improving work efficiency. At the same time, users can customize the target pose and scene according to their needs to generate unique images, meeting the needs of personalized image generation.

[0003] Although the pose-guided human image synthesis technology has significant advantages, it faces many technical problems. For example, although the adversarial network can generate high-quality images, the training cost is high and it is easily affected by data bias; the variational autoencoder is stable in training, but the generated images lack sufficient realism and may also lack details; the progressive diffusion model can gradually optimize the image quality, but it has high requirements for computing resources and poor generalization for rare poses. These image synthesis methods all have problems such as large model parameter quantities, slow inference speeds, and high hardware requirements, making it difficult to implement products. Summary of the Invention

[0004] The present invention aims to overcome at least one defect (insufficiency) of the above-mentioned prior art, and provides a pose-guided image synthesis method for solving the problems of insufficient realism of synthesized images and lack of in-depth understanding of high-level semantics of human images in the current technology.

[0005] According to a first aspect of the present application, there is provided a pose-guided image synthesis method. The method is based on a diffusion model, the diffusion model includes a Unet module, the Unet module includes an encoder and a decoder, the encoder includes a plurality of downsampling modules, the decoder includes an equal number of upsampling modules as the downsampling modules, and the method includes:

[0006] Obtain a source image, extract a pose feature map according to the source image, generate a first conditional encoding based on the pose feature map, and embed the first conditional encoding into each of the downsampling modules;

[0007] Extract a multi-scale feature map according to the source image, generate a second conditional encoding based on the multi-scale feature map, and embed the second conditional encoding into each of the upsampling modules;

[0008] Extract a multi-scale feature map from the source image, generate a second conditional encoding based on the multi-scale feature map, and embed the second conditional encoding into each of the upsampling modules;

[0009] Obtain multi-scale refined appearance features based on the multi-scale feature map, generate a third conditional encoding based on the multi-scale refined appearance features, and embed the third conditional encoding into each of the downsampling modules and each of the upsampling modules;

[0010] Obtain a noisy image of the source image, input the noisy image into the Unet model, and the Unet model synthesizes a target image based on the first conditional encoding, the second conditional encoding, and the third conditional encoding.

[0011] By embedding the pose feature map as the first conditional encoding into each downsampling module of the encoder, the model can generate image content that matches the given pose; by embedding the second conditional encoding generated from the multi-scale feature map into each upsampling module of the decoder, more rich detail information can be gradually introduced at different stages of image reconstruction; by embedding the third conditional encoding generated from the multi-scale refined appearance features into all the upsampling and downsampling modules of the encoder and decoder, the model can continuously consider the coherence and consistency of appearance information during the entire image synthesis process, thereby reducing problems such as artifacts and color inconsistencies in the synthesized image and improving the overall quality of the generated image.

[0012] Optionally, the diffusion model further includes a pose feature extraction module for extracting a pose feature map from the source image, and the pose feature extraction module includes a PixelUnshuffle module, a first convolutional layer, a first residual module, a second convolutional layer, a second residual module, a third convolutional layer, and a third residual module connected in sequence.

[0013] The PixelUnshuffle module can reduce the resolution of the image to reduce the computational amount of subsequent processing, and at the same time can indirectly enhance the local receptive field of the network, enabling the model to better capture the local structure and details in the image.

[0014] Optionally, each of the first residual module, the second residual module, and the third residual module includes two branches, one branch passes through a fourth convolutional layer, and the other branch passes through two fifth convolutional layers connected in series.

[0015] Optionally, the diffusion model further includes a multi-scale feature extraction module for extracting a multi-scale feature map according to the source image. The multi-scale feature extraction module includes a StemBlock module and three StageBlock modules. The StemBlock module includes a plurality of sixth convolutional layers. The StageBlock module evenly divides the number of input channels of the diffusion model into two parts to obtain a short branch Shortout and a main branch Mainout.

[0016] Optionally, the diffusion model further includes a perceptual refinement decoder for obtaining multi-scale refined appearance features based on the multi-scale feature map. The perceptual refinement decoder includes a plurality of decoding modules connected in series in sequence. Each decoding module includes a CrossAttention module and a GAU module. The obtaining of multi-scale refined appearance features based on the multi-scale feature map includes the following steps:

[0017] Performing a Flatten operation on the multi-scale feature map to obtain a one-dimensional sequence ;

[0018] Randomly initializing learnable tokens , and performing a Flatten operation on the learnable tokens to obtain a one-dimensional sequence ;

[0019] Inputting the one-dimensional sequence and the one-dimensional sequence into the perceptual refinement decoder and sequentially processing them through the cross-attention CrossAttention module and the GAU module of a plurality of decoding modules connected in series in sequence to obtain multi-scale refined appearance features.

[0020] Optionally, the attention calculation formula of the GAU module is as follows:

[0021]

[0022] where s is a preset step size, are respectively obtained by inputting the one-dimensional sequence into the GAU module and respectively mapping through learnable projection matrices , , is an activation function, is the global attention function of the GAU module.

[0023] Optionally, the obtaining of the noisy image of the source image and inputting the noisy image into the Unet model includes:

[0024] With random Gaussian noise as the initial latent noise and map the initial latent noise to latent noise at different time steps and input the latent noise into the Unet model. The calculation formula of the latent noise

[0025]

[0026] is as follows: where is the noise level coefficient at time step and is the latent noise at step and

[0027] is the random Gaussian noise sampled from the standard normal distribution.

[0028] According to the second aspect of the present application, there is provided a pose-guided image synthesis device, including:

[0029] A pose feature encoding and embedding module, configured to obtain a source image, extract a pose feature map according to the source image, generate a first conditional encoding based on the pose feature map, and embed the first conditional encoding into each of the downsampling modules.

[0030] A multi-scale feature encoding and embedding module, configured to extract a multi-scale feature map according to the source image, generate a second conditional encoding based on the multi-scale feature map, and embed the second conditional encoding into each of the upsampling modules.

[0031] A multi-scale refined appearance feature encoding and embedding module, configured to obtain multi-scale refined appearance features based on the multi-scale feature map, generate a third conditional encoding based on the multi-scale refined appearance features, and embed the third conditional encoding into each of the downsampling modules and each of the upsampling modules.

[0032] An image synthesis module, configured to obtain a noise image of the source image, input the noise image into the Unet model, and the Unet model synthesizes a target image based on the first conditional encoding, the second conditional encoding, and the third conditional encoding.

[0033] According to the third aspect of the present application, there is provided an electronic device, including:

[0034] A memory, configured to store one or more computer programs;

[0035] A processor, when the one or more computer programs are executed by the processor, implements the pose-guided image synthesis method described in the first aspect above.

[0036] According to a fourth aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the pose-guided image synthesis method described in the first aspect above when executed.

[0037] Based on any of the above aspects, the pose-guided image synthesis method, apparatus, electronic device, and computer storage medium provided by the embodiments of the present application embed the pose feature map as the first conditional encoding into each downsampling module of the encoder, enabling the model to generate high-quality images with specific poses; by generating the second conditional encoding from the multi-scale feature maps and embedding it into each upsampling module of the decoder, the model can gradually introduce richer detail information at different stages of image reconstruction. By generating the third conditional encoding from the multi-scale refined appearance features and embedding it into all the upsampling and downsampling modules of the encoder and decoder, the model can continuously consider the coherence and consistency of appearance information throughout the image synthesis process, thereby reducing problems such as artifacts and color inconsistencies in the synthesized images and further improving the quality of the generated images. The image synthesis method based on the diffusion model can reduce the use of computing resources while ensuring image quality, making the model easier to be used in products and implemented. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 A schematic application scenario diagram of the pose-guided image synthesis method provided in this embodiment.

[0040] Figure 2 A flowchart of the pose-guided image synthesis method provided in this embodiment.

[0041] Figure 3 A schematic diagram of the overall structure of the diffusion model provided in this embodiment.

[0042] Figure 4 A schematic diagram of the structure of the human pose feature extraction module provided in this embodiment.

[0043] Figure 5 A schematic diagram of the structure of the residual module provided in this embodiment.

[0044] Figure 6 Schematic diagram of the multi-scale feature extraction module provided in this embodiment.

[0045] Figure 7 Schematic diagram of the StemBlock module provided in this embodiment.

[0046] Figure 8 Schematic diagram of the CSPLayer module provided in this embodiment.

[0047] Figure 9 Schematic diagram of the pose-guided image synthesis device provided in this embodiment.

[0048] Figure 10 Schematic diagram of the electronic device provided in this embodiment. Detailed implementation manners

[0049] The accompanying drawings of this application are only for illustrative purposes and should not be construed as a limitation to this application. To better illustrate the following embodiments, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0050] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above accompanying drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0052] The inventors have found that in the existing posture-guided character image generation technology, the conditional generative adversarial network requires a lot of computing resources and is easily affected by the bias of training data, resulting in unstable image quality; the variational autoencoder generates images with low fidelity and lacks details; although the progressive conditional diffusion model can generate high-quality images, the network structure is complex, the computing resource requirements are high, and the generalization ability for rare postures is poor, and it is easy to overfit. Therefore, how to balance image quality, computing efficiency and model generalization ability, reduce the number of model parameters, improve the inference speed, and enhance the adaptability to various postures, so that the model can be applied and implemented in resource-constrained environments is a problem that needs to be solved in current research.

[0053] This embodiment provides a technical solution that can solve the above-mentioned problem. The specific implementation methods of this application are described in detail below in conjunction with the accompanying drawings.

[0054] For example, a schematic diagram of an application scenario of a posture-guided image synthesis method provided in an embodiment of the present application is provided. Figure 1 As shown, the application scenario includes at least a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has an image processing function and can also have a data transmission function for video streams and audio streams; the terminal device 200 has a streaming media playback function and can also have an image processing function.

[0055] It is understandable that the server 100 can be an independent electronic device or a cluster of multiple electronic devices; the terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a car terminal, etc., but is not limited thereto.

[0056] In one practicable manner, the server 100 and the terminal 200 may respectively execute the posture-guided image synthesis method provided in the embodiment of the present application, or, optionally, the posture-guided image synthesis method provided in the embodiment of the present application is partially executed in the server 100 and partially executed in the terminal 200.

[0057] This embodiment provides a posture-guided image synthesis method based on a diffusion model, wherein the diffusion model includes a Unet module. Figure 3 As shown, the Unet module includes an encoder and a decoder, the encoder includes several downsampling modules, and the decoder includes upsampling modules corresponding to the number of the downsampling modules.

[0058] In this embodiment, the encoder includes three downsampling layers, namely down1, down2, and down3. Through these three downsampling layers, the spatial dimension of the feature map is gradually reduced, and at the same time, the number of feature channels is correspondingly increased, enabling the model to capture more abstract and deep-level image features. The goal of the decoder is to gradually reconstruct the image from the high-level features extracted by the encoder. The decoder includes three upsampling layers, namely up1, up2, and up3. Through these three upsampling layers, the spatial dimension of the feature map is gradually restored, making the finally output image both accurate and clear. In the Unet model, by embedding additional information in the encoder and decoder, the model can understand and respond to these specific conditions. For example, by encoding specific conditions in the form of text descriptions, the model can achieve conditional image generation, greatly enhancing the flexibility and generality of the model.

[0059] As Figure 2 shown, this embodiment provides a pose-guided image synthesis method, which may include the following steps:

[0060] S1. Obtain a source image, extract a pose feature map according to the source image, generate a first conditional encoding based on the pose feature map, and embed the first conditional encoding into each of the downsampling modules.

[0061] As Figure 3 shown, the diffusion model includes a pose feature extraction module for extracting a pose feature map according to the source image. As Figure 4 shown, the human pose feature extraction module includes a PixelUnshuffle module, a first convolutional layer, a first residual module, a second convolutional layer, a second residual module, a third convolutional layer, and a third residual module connected in sequence.

[0062] Preferably, the downsampling factor of the PixelUnshuffle module is 4, the kernel size of the first convolutional layer is 1, and the kernel sizes of the second convolutional layer and the third convolutional layer are both 3 and the strides are both 2.

[0063] In the specific implementation process, the 2D bone information is input into the human pose feature extraction module, passed through the PixelUnshuffle module and the first convolutional layer, and then the first feature Feature1 is obtained through the first residual module. After passing through the first residual module, the first feature Feature1 is downsampled through the second convolutional layer, and then the second feature Feature2 is obtained through the second residual module. After passing through the second residual module, the second feature Feature2 is downsampled through the third convolutional layer, and then the third feature Feature3 is obtained through the third residual module. The first feature Feature1, the second feature Feature2, and the third feature Feature3 respectively correspond to the feature information of different scales of the human pose.

[0064] As Figure 5 shown, the residual module of this embodiment includes two branches, one of which passes through the fourth convolutional layer, and the other passes through two cascaded fifth convolutional layers. Specifically, when implemented, the outputs processed by the two branches are added together to obtain the output of the residual module.

[0065] Preferably, the convolutional kernel size of the fourth convolutional layer is 1, and the convolutional kernel size of the fifth convolutional layer is 3.

[0066] S2. Extract a multi-scale feature map from the source image, generate a second conditional encoding based on the multi-scale feature map, and embed the second conditional encoding into each of the upsampling modules.

[0067] In this embodiment, the diffusion model further includes a multi-scale feature extraction module for extracting a multi-scale feature map from the source image. The multi-scale feature extraction module adopts a CSPNeXt structure, including a StemBlock module and three StageBlock modules. Among them, the StemBlock module includes several sixth convolutional layers; the StageBlock module evenly divides the number of input channels of the diffusion model into two parts to obtain a short branch Shortout and a main branch Mainout. The number of sixth convolutional layers is preferably 3, and the convolutional kernel size of each sixth convolutional layer is 3.

[0068] Specifically, when implemented, as Figure 7 described, the output of the main branch Mainout is input into the CSPLayer module for processing, and after the output of the CSPLayer module and the output result of the short branch Shortout are concatenated by Cat, they are sequentially processed by a channel attention mechanism and a convolutional layer with a convolutional kernel of 1x1 to obtain an output value. Through the above process, the multi-scale feature extraction module extracts a multi-scale feature map Fs = [f1, f2, f3, f4] containing multi-level details from the source image, providing a coarse-to-fine appearance control basis for subsequent image generation.

[0069] As Figure 8 shown, in an alternative implementation, the CSPLayer module includes a sequentially connected convolutional layer and a depthwise separable convolutional layer. In this embodiment, the convolutional kernel size of the convolutional layer of the CSPLayer module is 3x3, and the convolutional kernel size of the depthwise separable convolutional layer is 5x5. The input value of the CSPLayer module is sequentially processed by the convolutional layer and the depthwise separable convolutional layer to obtain corresponding feature values, and the obtained feature values are added to the input value to obtain the output value of the CSPLayer module.

[0070] To improve the efficiency of image feature extraction, while maintaining the advantage of the model inference speed and making the model easy to deploy, in this embodiment, the ReLU6 function is used as the activation function of the StageBlock module. The specific formula of the ReLU6 function is as follows:

[0071]

[0072] The formula of the ReLU6 function indicates that if the input value x is negative, the output is 0; if the input value x is between 0 and 6, x is directly output; if the input value x is greater than 6, the output is 6. By restricting the maximum output value of the function to 6, the problem of gradient explosion that may occur in the deep network can be avoided, thus maintaining the numerical stability of the model. In addition, due to the simple calculation of the ReLU6 function, it can accelerate the training and inference speed of the model, enabling the module to be more efficiently deployed to embedded devices.

[0073] S3. Obtain multi-scale refined appearance features based on the multi-scale feature maps, generate the third conditional encoding based on the multi-scale refined appearance features, and embed the third conditional encoding into each of the downsampling modules and each of the upsampling modules.

[0074] In this embodiment, the diffusion model further includes a perceptual refinement decoder for obtaining multi-scale refined appearance features based on the multi-scale feature maps. The perceptual refinement decoder includes a number of decoding modules connected in series in sequence. Each decoding module includes a CrossAttention module and a GAU module;

[0075] As Figure 6 shown, in this embodiment, the perceptual refinement decoder includes 4 decoding modules, and each decoding module includes a CrossAttention module and a GAU module. Specifically, the perceptual refinement decoder adopts a cascaded architecture, and the 4 decoding modules are connected in series in sequence. Each decoding module is responsible for further refining the output of the previous decoding module to gradually improve the decoding quality, accuracy, and detail performance. Obtaining multi-scale refined appearance features based on the multi-scale feature maps includes the following steps:

[0076] Perform a Flatten operation on the multi-scale feature map F4 to obtain a one-dimensional sequence ;

[0077] Randomly initialize the learnable tokens , and perform a Flatten operation on the learnable tokens to obtain a one-dimensional sequence ;

[0078] The one-dimensional sequence and the one-dimensional sequence The input perception refinement decoder obtains multi-scale refined appearance features.

[0079] For each of the decoding modules, the input feature Z is input into the CrossAttention module to obtain a one-dimensional sequence . Specifically, the input feature Z is linearly transformed through learnable projection matrices , , (where C is the dimension of the input feature Z) to obtain the corresponding query vector Q, key vector K, and value vector V. The calculation formulas for the query vector Q, key vector K, and value vector V are as follows:

[0080]

[0081] The attention calculation formula of the CrossAttention module is as follows:

[0082]

[0083] The one-dimensional sequence output by the CrossAttention module is input into the GAU module, and through new learnable projection matrices , , , the corresponding Q, K, V are obtained, and the step size s is taken as 128. The attention calculation formula of the GAU module is as follows:

[0084]

[0085] In this embodiment, the learnable projection matrices , , can be randomly initialized before model training and optimized and updated through the backpropagation algorithm during the model training process.

[0086] S4. Obtain the noisy image of the source image, input the noisy image into the Unet model, and the Unet model synthesizes the target image based on the first conditional encoding, the second conditional encoding, and the third conditional encoding.

[0087] In this embodiment, the noisy image is generated by gradually adding Gaussian noise to the source image. Specifically, through the forward diffusion process with a set step size , using random Gaussian noise as the initial latent noise , the initial latent noise is mapped to different time steps to obtain the latent noise , the potential noise The corresponding noise image is obtained by conversion. Potential noise The calculation formula is as follows:

[0088]

[0089] in, is the time step The noise level coefficient at yes The potential noise of the step, is random Gaussian noise sampled from a standard normal distribution.

[0090] The potential noise The corresponding noise image captures the key information and features in the data through three consecutive downsampling layers, and finally outputs a highly generalized multidimensional tensor through layer-by-layer compression and abstraction; after the multidimensional tensor is input into the decoder, it passes through three consecutive upsampling layers and finally outputs a multidimensional feature tensor, which is a prediction and reconstruction of the information in the original potential noise. After obtaining the multidimensional feature tensor, the predicted noise is subjected to inverse noise mapping processing to finally obtain a synthesized image.

[0091] like Figure 9 As shown, the embodiment of the present application further provides a posture guidance image synthesis device 610. Optionally, the posture guidance image synthesis device 610 may include:

[0092] The posture feature coding and embedding module 611 is used to obtain a source image, extract a posture feature map according to the source image, generate a first conditional coding based on the posture feature map, and embed the first conditional coding into each of the downsampling modules.

[0093] The multi-scale feature coding and embedding module 612 is used to extract a multi-scale feature map according to the source image, generate a second conditional code based on the multi-scale feature map, and embed the second conditional code into each of the upsampling modules.

[0094] The multi-scale refined appearance feature encoding and embedding module 613 is used to obtain multi-scale refined appearance features based on the multi-scale feature map, generate a third conditional code based on the multi-scale refined appearance features, and embed the third conditional code into each of the downsampling modules and each of the upsampling modules.

[0095] The image synthesis module 614 is used to obtain a noise image of the source image, input the noise image into the Unet model, and synthesize the target image based on the first conditional coding, the second conditional coding and the third conditional coding by the Unet model.

[0096] It can be understood that the above device embodiments and the above method embodiments can correspond to each other. Similar descriptions of the device embodiments can refer to the method embodiments. To avoid repetition, they will not be elaborated here. A posture guidance image synthesis device provided in an embodiment of the present application can execute a posture guidance image synthesis method provided in any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. The functional modules of the posture guidance image synthesis device can be implemented in the form of hardware, can be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules.

[0097] Specifically, each step of the method embodiment of the present application can be completed by the integrated logic circuit in the hardware of the processor and / or instructions in the form of software. The steps of the posture guidance image synthesis method in combination with the embodiments of the present application can be directly embodied as being completed by the hardware-encoded processor, or completed by a combination of the hardware and software modules in the encoded processor. Optionally, the software module can be located in a random access memory, read-only memory, programmable read-only memory, flash memory, electrically erasable programmable memory, register and other storage media. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0098] An embodiment of the present application provides an electronic device 710, and its structure is as Figure 10 shown. The electronic device 710 can be the server 100 or the terminal 200 shown in this embodiment Figure 1 shown.

[0099] As Figure 10 shown, the electronic device 710 includes a memory 711, a processor 712, a communication module 713, an input / output interface 714, etc. Optionally, the memory 711, the processor 712, the communication module 713, and the input / output interface 714 can be connected and communicate through a bus 715.

[0100] The memory 711 is used to store one or more computer programs and transmit the code of the computer programs to the processor 712; when the one or more computer programs are executed by the processor 711, the posture guidance image synthesis method in the embodiments of the present application is implemented.

[0101] Optionally, the electronic device 710 may be connected to a network through the communication module 713 to communicate with other devices, such as terminals or servers, through the network to achieve data interaction. The electronic device 710 may be various forms of digital computers, such as, by way of example, desktop computers, servers, workstations, mainframe computers, or other types of computers. The electronic device 710 may also be various forms of mobile terminals, such as, by way of example, smart phones, tablet computers, wearable devices (such as helmets, glasses, watches, etc.) and other similar mobile terminals.

[0102] Optionally, the electronic device 710 may be connected to the required input / output devices, such as keyboards, display devices, etc., through the input / output interface 714. The electronic device 710 itself may have a display device, and may also externally connect other display devices through the input / output interface 714. Optionally, a storage device, such as a hard disk, etc., may also be connected through the input / output interface 714, so that the data in the electronic device 710 can be stored in the storage device, or the data in the storage device can be read, and the data in the storage device can also be stored in the memory 711. It can be understood that the input / output interface 714 may be a wired interface or a wireless interface. Depending on different actual application scenarios, the devices connected to the input / output interface 714 may be components of the electronic device 710 or external devices connected to the electronic device 710 when needed.

[0103] Optionally, the memory 711 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.

[0104] Optionally, the computer program stored in the processor 711 may be divided into one or more modules. The one or more modules are stored in the memory 711 and executed by the processor 712 to complete the method provided by the present embodiment itself. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the computer program instruction segments are used to describe the execution process of the computer program in the electronic device 710.

[0105] Optionally, the processor 712 can be various general-purpose and / or dedicated processing components with processing and computing capabilities. Some examples of the processor 712 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various dedicated artificial intelligence computing chips, various processors running machine learning model algorithms, and can also be any suitable controller, microcontroller, processor, etc. The processor 712 executes the various methods and processes of this embodiment. Exemplarily, such as a pose-guided image synthesis method according to an embodiment of the present application.

[0106] Optionally, the bus 715 can include a path for transmitting information. The bus 715 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 715 can be divided into an address bus, a data bus, a control bus, etc.

[0107] In an alternative implementation, an embodiment of the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods of the above method embodiments. Part or all of the computer program can be loaded and / or installed on the memory 711 of the electronic device 710. When the computer program is executed by the processor 712, one or more steps of a pose-guided image synthesis method according to an embodiment of the present application can be executed.

[0108] Optionally, the computer-readable storage medium can be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.

[0109] Obviously, the above embodiments of the present application are merely examples for clearly explaining the technical solutions of the present application, rather than limitations on the specific implementation manners of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present application shall be included within the protection scope of the claims of the present application.

[0110] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A posture-guided image synthesis method, the method is based on a diffusion model, the diffusion model includes a Unet module, the Unet module includes an encoder and a decoder, the encoder includes a number of downsampling modules, the decoder includes upsampling modules with the same number as the downsampling modules, characterized in that: The method comprises: Acquire a source image, extract a posture feature map according to the source image, generate a first conditional code based on the posture feature map, and embed the first conditional code into each of the downsampling modules; the posture feature map is extracted by a posture feature extraction module of a diffusion model, and the posture feature extraction module includes a PixelUnshuffle module, a first convolutional layer, a first residual module, a second convolutional layer, a second residual module, a third convolutional layer, and a third residual module connected in sequence; Extracting a multi-scale feature map according to the source image, generating a second conditional code based on the multi-scale feature map, and embedding the second conditional code into each of the upsampling modules; the multi-scale feature map is extracted by a multi-scale feature extraction module of a diffusion model, the multi-scale feature extraction module includes a StemBlock module and three StageBlock modules; the StemBlock module includes a plurality of sixth convolutional layers; the StageBlock module evenly divides the number of input channels of the multi-scale feature map into two parts to obtain a short branch Shortout and a main branch Mainout; Perform a Flatten operation on the multi-scale feature map to obtain a one-dimensional sequence ; Randomly initialize learnable labels , for the learnable marker Perform Flatten operation to obtain a one-dimensional sequence ; One-dimensional sequence and one-dimensional sequence The perceptual refinement decoder is input and processed in sequence by a cross-attention CrossAttention module and a GAU module of a plurality of decoding modules connected in series in sequence to obtain a multi-scale refined appearance feature; a third conditional code is generated based on the multi-scale refined appearance feature, and the third conditional code is embedded into each of the down-sampling modules and each of the up-sampling modules; the multi-scale refined appearance feature is obtained by a perceptual refinement decoder of a diffusion model, and the perceptual refinement decoder includes a plurality of decoding modules connected in series in sequence, and each decoding module includes a CrossAttention module and a GAU module; A noise image of the source image is obtained, and the noise image is input into the Unet module, and the Unet module synthesizes a target image based on the first conditional coding, the second conditional coding, and the third conditional coding.

2. A posture-guided image synthesis method according to claim 1, characterized in that: The first residual module, the second residual module and the third residual module each include two branches, one of which passes through the fourth convolutional layer, and the other branch passes through two fifth convolutional layers connected in series.

3. A posture-guided image synthesis method according to claim 1, characterized in that: The attention calculation formula of the GAU module is as follows: in, is the preset step size, By transforming the one-dimensional sequence Input GAU module and pass through the learnable projection matrix , , Mapping obtained, is the activation function, is the global attention function of the GAU module.

4. A posture-guided image synthesis method according to claim 1, characterized in that: The step of obtaining the noise image of the source image comprises the following steps: With random Gaussian noise As the initial potential noise , the initial potential noise Mapping to different time steps Get the potential noise , the potential noise Convert to obtain the corresponding noise image; The potential noise The calculation formula is as follows: in, is the time step The noise level coefficient at yes The potential noise of the step, is random Gaussian noise sampled from a standard normal distribution.

5. A posture-guided image synthesis device, characterized in that: The device comprises: A posture feature encoding and embedding module, used to obtain a source image, extract a posture feature map according to the source image, generate a first conditional code based on the posture feature map, and embed the first conditional code into each downsampling module; the posture feature map is extracted by a posture feature extraction module of a diffusion model, and the posture feature extraction module includes a PixelUnshuffle module, a first convolutional layer, a first residual module, a second convolutional layer, a second residual module, a third convolutional layer, and a third residual module connected in sequence; A multi-scale feature encoding and embedding module is used to extract a multi-scale feature map according to the source image, generate a second conditional code based on the multi-scale feature map, and embed the second conditional code into each upsampling module; the multi-scale feature map is extracted by a multi-scale feature extraction module of a diffusion model, and the multi-scale feature extraction module includes a StemBlock module and three StageBlock modules; the StemBlock module includes a plurality of sixth convolutional layers; the StageBlock module evenly divides the number of input channels of the multi-scale feature map into two parts to obtain a short branch Shortout and a main branch Mainout; A multi-scale refined appearance feature encoding and embedding module is used to perform a Flatten operation on the multi-scale feature map to obtain a one-dimensional sequence ; Randomly initialize learnable labels , for the learnable marker Perform Flatten operation to obtain a one-dimensional sequence ; One-dimensional sequence and one-dimensional sequence The perceptual refinement decoder is input and processed in sequence by a cross-attention CrossAttention module and a GAU module of a plurality of decoding modules connected in series in sequence to obtain a multi-scale refined appearance feature; a third conditional code is generated based on the multi-scale refined appearance feature, and the third conditional code is embedded into each of the down-sampling modules and each of the up-sampling modules; the multi-scale refined appearance feature is obtained by a perceptual refinement decoder of a diffusion model, and the perceptual refinement decoder includes a plurality of decoding modules connected in series in sequence, and each decoding module includes a CrossAttention module and a GAU module; An image synthesis module is used to obtain a noise image of the source image, input the noise image into the Unet module, and the Unet module synthesizes a target image based on the first conditional coding, the second conditional coding and the third conditional coding.

6. An electronic device, characterized in that: include: a memory for storing one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements the gesture-guided image synthesis method as described in any one of claims 1-4.

7. A computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a processor to implement the gesture-guided image synthesis method according to any one of claims 1 to 4 when executed.

Citation Information

Patent Citations

  • Human body abdominal muscle image generation method, system and device and storage medium

    CN118365723A