Zero-sample rapid object migration method based on diffusion model
Through the zero-sample fast object migration method based on diffusion model, the segmentation detection model, high-pass filtering and multi-stage high-resolution feature map, combined with identity feature vectors and cross-attention mechanism, the problems of inaccurate object migration, slow speed and poor diversity in the existing technology are solved, and high-quality and fast object migration effects are achieved.
Patent Information
- Application Number
- CN202510190812.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art cannot quickly and accurately migrate objects in the reference image to the scene image under zero sample conditions, and there is a big gap between the generated objects and the objects in the reference image in terms of shape, identity, color, etc., with poor diversity, slow training speed, and faces the problem of data scarcity.
The zero-sample fast object migration method based on the diffusion model is adopted, and the target object is extracted through preset segmentation detection model, high-pass filtering is used to extract high-frequency regions, and the encoder and decoder extract multi-level high-resolution feature maps. Combining the identity feature vector and cross-attention mechanism, multiple iterative denoising is injected into the diffusion model to generate synthetic images.
It realizes the rapid and accurate migration of objects into scene images under zero sample conditions, ensures consistency of object identity, improves the diversity and quality of generated images, reduces training time, and is suitable for image synthesis and editing tasks.
Smart Images

Figure CN120032005A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image generation, and in particular relates to a zero-sample rapid object migration method based on a diffusion model. Background Art
[0002] Image generation technology has flourished rapidly with the development of diffusion models, where humans can generate target images by providing textual prompts or other conditions. The powerful capabilities of these models also bring the potential of image editing. For example, some studies learn to edit the pose, style, or content of images through instructions, while other studies regenerate local image regions under the guidance of textual prompts.
[0003] In the prior art, two input images are provided, a reference image and a scene image. The reference image needs to include an object, and the object can be any instance. Zero-sample object migration technology refers to the ability to quickly migrate objects in the reference image to a specified position in the scene image under zero-sample conditions, that is, when the model does not need to be trained on the reference image, and to achieve seamless integration of the migrated object and the scene.
[0004] Most existing technologies are based on generative adversarial networks and are trained using an autoregressive approach. They mainly extract objects and backgrounds from reference images through segmentation, then encode them separately through a multi-layer perceptron, and map the encoded results to a generative adversarial network for fusion.
[0005] This technology cannot guarantee the accuracy of object migration. There is a large gap between the shape, identity, color, etc. of the objects in the reference image, and the diversity of the synthesized objects is poor. In addition, the existing technology requires a long training time on multiple reference images before reasoning, which greatly limits its migration speed and faces the problem of data scarcity. Summary of the invention
[0006] In order to solve the above technical problems existing in the prior art, the present invention proposes a zero-sample fast object migration method based on a diffusion model, comprising the following steps:
[0007] Step S110, obtaining a reference image and a scene image;
[0008] Step S120: input the reference image into a preset segmentation detection model to obtain a target object in the reference image;
[0009] Step S130, extracting a high-frequency region from the target object in the reference image by high-pass filtering, and aligning the extracted high-frequency region to the scene image as a scene fusion image;
[0010] Step S140, extracting a multi-level high-resolution feature map from the scene fusion map through an encoder and a decoder;
[0011] Step S150, obtaining an identity feature vector of the target object in the reference image through a pre-trained identity extraction network;
[0012] Step S160, injecting the identity feature vector of the target object into the diffusion model through a cross-attention mechanism;
[0013] Step S170, injecting the multi-level high-resolution feature map in step S140 into the decoder of the diffusion model by feature splicing, so as to achieve background preservation of the scene image and the target object;
[0014] Step S180, the diffusion model generates a feature map of the synthetic image through multiple iterations of denoising;
[0015] Step S190: The feature map of the composite image is decoded by a variational autoencoder to obtain a final composite image.
[0016] Furthermore, the reference image and the scene image in step S110 are derived from a data set, and the process of making the data set includes image filtering, and the image filtering includes the following steps:
[0017] Step S110.1, preprocessing the original image;
[0018] Step S110.2, using median filtering to detect and remove noise;
[0019] Step S110.3, evaluating the clarity of the image;
[0020] Step S110.4: Perform quality assessment.
[0021] Furthermore, the preset segmentation detection model in step S120 includes an image encoder, a text encoder and a mask decoder, wherein:
[0022] The input of the image encoder is the reference image, which is divided into image blocks of fixed size, each of which is flattened and mapped to a high-dimensional space; each image block is then linearly transformed to obtain a corresponding embedding vector; the embedding vector passes through multiple Transformer encoder layers, each of which contains a self-attention module, a feedforward neural network, a residual connection and layer normalization, and finally outputs the features of the reference image;
[0023] The input of the text encoder is a preset text. The text is first segmented, and each word or subword is converted into a corresponding word embedding. The word embedding is added to the position code, and the added word embedding passes through multiple Transformer encoder layers. The output of the last layer of encoder is globally pooled, and the pooled vector is processed through a linear layer and an activation function to finally obtain a feature vector of the text;
[0024] The input of the mask decoder is the output of the image encoder and the text encoder. Comprehensive features are generated by vector splicing. The comprehensive features are nonlinearly transformed through a feedforward neural network to extract feature representations, which are then processed through a linear layer to output a segmentation mask. The segmentation mask is further refined through a conditional random field, and finally the refined segmentation mask is used to mark the target object in the reference image.
[0025] Furthermore, in step S130, extracting a high-frequency area from the target object in the reference image by high-pass filtering includes:
[0026] The target object in the reference image is input into a high-pass filter, and through frequency analysis and transformation operations, the high-frequency components in the image are retained, while the low-frequency components are weakened or removed, and the high-frequency area in the image is extracted. The extracted high-frequency area is remapped back to the spatial domain to obtain a high-frequency feature image of the reference image;
[0027] The step of aligning the extracted high-frequency region to the scene image as a scene fusion image includes adjusting the scale, rotation and position of the image so that key points in the high-frequency feature image are aligned with corresponding positions in the scene image.
[0028] Further, step S140 includes:
[0029] The scene fusion image is first input to the encoder for multi-level processing. In the encoder, the resolution of the image is reduced layer by layer through multiple downsampling operations to generate feature layers with multiple resolutions.
[0030] The decoder performs multi-layer upsampling on the output features of the encoder and gradually restores them to high resolution. The multi-level feature information of the scene fusion graph is gradually reconstructed into multi-level high-resolution feature graphs.
[0031] Further, step S150 includes:
[0032] Step S150.1, dividing the target object in the reference image into image blocks of fixed size;
[0033] Step S150.2, the relationship between each image block position feature vector and other image block position feature vectors is calculated by the self-attention module, and the calculation result of the self-attention module is used as the global feature representation of each image block;
[0034] Step S150.3, the output of each layer of the self-attention module is subjected to a residual connection and a layer normalization operation, wherein the residual connection is used to directly add the image block position feature vector of each layer to the global feature representation of the image block, and the layer normalization operation is used to standardize the global feature representation of the image block of each layer;
[0035] Step S150.4: The global feature representation outputted from step S150.3 is transformed through two layers of linear transformation and an activation function to obtain the identity feature vector of the final target object.
[0036] Further, step S160 includes:
[0037] Step S160.1, passing the identity feature vector through two linear transformation layers respectively to generate a key matrix K and a value matrix V;
[0038] Step S160.2, extract the state vector of the diffusion model at time t, and pass the state vector through a linear transformation layer to obtain a query matrix Q;
[0039] Step S160.3, using the key matrix K, the value matrix V and the query matrix Q to perform similarity calculation, and output the self-attention weight matrix;
[0040] Step S160.4, weighted summing of the self-attention weight matrix to obtain a self-attention output matrix;
[0041] Step S160.5, introduce a multi-head self-attention module into the self-attention output matrix to obtain a multi-head self-attention output matrix.
[0042] Further, in step S170:
[0043] During the injection process, features of the same scale are extracted from the feature map of the corresponding resolution each time and concatenated with the feature map of the corresponding level of the decoder of the diffusion model.
[0044] Further, step S180 includes:
[0045] Step S180.1, setting the size of the generated noise image, using a Gaussian distribution with zero mean and unit variance to randomly sample the R, G, and B values of each pixel of the noise image, and generating a completely random initial noise image;
[0046] Step S180.2, perform successive denoising on the initialized diffusion model, and output a feature map of the synthesized image.
[0047] Further, step S190 includes:
[0048] Step S190.1, input the synthetic image feature map into the encoder of the variational autoencoder, and output a re-parameterized feature map;
[0049] Step S190.2, using the reparameterization technique to output the reparameterized latent variables on the reparameterized feature map;
[0050] Step S190.3: input the re-parameterized latent variables into the decoder for image synthesis.
[0051] The present invention solves the problem of "object migration", that is, quickly and accurately placing the target object at the desired position in the scene image, and regenerating the local area marked with a border in the scene image by using the target object as a template. This capability is of great demand in practical applications, such as image synthesis, special effects image rendering, poster production, virtual try-on and other application fields.
[0052] The present invention has high controllability in editing specific local areas of scene images, is easily extended to combinations of multiple objects, has strong diversity, can ensure the consistency of object identity, and can achieve fast object migration under zero-sample settings. The present invention can achieve high-fidelity and high-quality image generation and can become a basic solution for various image generation and editing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION
[0054] In order to better understand the purpose, technical solution and function of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings, but the present invention can be implemented in a variety of different ways as defined and covered by the claims. The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention, and the exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0055] like Figure 1 As shown, the present invention provides a zero-sample fast object migration method based on a diffusion model, comprising the following steps:
[0056] Step S110, obtaining a reference image and a scene image.
[0057] The reference image and the scene image are different images. The reference image is the image to which the object needs to be transferred, and the scene image is used to fuse the object of the reference image, that is, to transfer the object to the image. There is no restriction on the object features in the reference image, and it can be any instance in nature.
[0058] In a specific embodiment of the present invention, the reference image and the scene image are derived from a custom high-definition data set, and the process of making the data set includes image acquisition and image filtering. The image acquisition can be obtained from a network database using crawler technology, and the image filtering includes eliminating images with poor image quality, and finally obtaining a high-definition image data set.
[0059] Specifically, for the image filtering process, the first step is to preprocess the original image.
[0060] The main goal of this stage is to ensure the effectiveness of subsequent processing. The preprocessing steps include format conversion and resizing. First, all images are converted from RGB color space to grayscale images; then, the images are uniformly resized to ensure that the resolution of all images is consistent, so as to avoid errors caused by different sizes in subsequent processing.
[0061] The second step is to use median filtering to detect and remove noise.
[0062] First, set a 3×3 window and then traverse each pixel in the image. For each pixel, extract all pixel values within the window and calculate the median, then use this median to replace the pixel value at the center of the window. This process can effectively remove salt and pepper noise and preserve the edge information of the image as much as possible.
[0063] The third step is to evaluate the clarity of the image.
[0064] This process is achieved by calculating the gradient of the image. The image is convolved horizontally and vertically in the manner of the Sobel operator to obtain two gradient images. Then, a clarity image is generated by calculating the amplitude of the two gradients. In this image, the values of the clear area are higher, while the values of the blurred area are lower. By setting a threshold, the image is divided into two categories: clear and blurred. If the clarity of some images does not meet the expectations, sharpening is required. The specific steps of sharpening use a high-pass filter. First, the image is convolved with the Laplace operator to obtain an edge image. Then, the edge image is combined with the original image, and the details of the image are enhanced through the weighted combination method.
[0065] The fourth step is to conduct a quality assessment.
[0066] Peak signal-to-noise ratio (PSNR) is often used at this stage to quantify the quality of the image. When calculating PSNR, the processed image is compared to a high-quality reference image. The higher the PSNR value, the better the image quality. Similarly, an SSIM value close to 1 indicates good image quality, while a value close to 0 indicates poor image quality. After multiple processing steps, if the image quality still does not meet the standard, it will be removed. The removal process is implemented by setting a specific quality threshold. Images that are below this requirement will be marked as to be removed, and qualified images will be retained in the dataset.
[0067] Step S120: input the reference image into a preset segmentation detection model to obtain the target object in the reference image.
[0068] That is, after obtaining the reference image, a preset segmentation detection model is retrieved from the server, the reference image is input into the preset segmentation detection model, and segmentation detection is performed on the reference image through the preset segmentation detection model.
[0069] Specifically, the preset segmentation detection model consists of three parts: image encoder, text encoder and mask decoder.
[0070] The input of the image encoder is the reference image, which is divided into image blocks of fixed size. Each block is flattened and mapped to a high-dimensional space. Then each image block obtains the corresponding embedding vector through linear transformation. The linear transformation is represented by multiplying the input vector by a weight matrix. The weight matrix is a two-dimensional array whose elements are learnable parameters. By performing matrix multiplication of the original vector and the weight matrix, a new vector, namely the embedding vector, can be obtained.
[0071] The embedding vector passes through multiple Transformer encoder layers, each of which contains a self-attention module, a feedforward neural network, a residual connection, and layer normalization, and finally outputs the features of the reference image.
[0072] The input of the text encoder is a preset text, i.e., "segmentation". The text is first segmented, and each word or subword is converted into a corresponding word embedding. The word embedding is added to the position encoding to retain the position information of the word in the sentence. The added word embedding passes through multiple Transformer encoder layers, and the output of the last layer of encoder is globally pooled. The pooled vector is processed through a linear layer and an activation function to finally obtain the feature vector of the text.
[0073] The input of the mask decoder is the output of the image encoder and the text encoder, that is, the feature vector of the reference image and the feature vector of the text. A comprehensive feature is then generated by vector splicing. The comprehensive feature is nonlinearly transformed through a feedforward neural network to extract deeper feature representations, and then processed through a linear layer to output a segmentation mask. The segmentation mask is further refined through a conditional random field, and finally the refined segmentation mask is used to mark the target object in the reference image.
[0074] In this way, the target object in the reference image can be obtained through the above method.
[0075] Step S130 , extracting a high-frequency region from the target object in the reference image by high-pass filtering, and aligning the extracted high-frequency region to the scene image as a scene fusion image.
[0076] After obtaining the reference image and the scene image, the target object in the reference image is taken as the processing object and processed by high-pass filtering to extract the high-frequency area of the target object, where the high-frequency area refers to the detailed information part in the image, usually including edges, textures and other areas with faster grayscale changes, and these areas can reflect the edge contours and subtle features of the target object in image processing.
[0077] Specifically, the first step is to perform high-pass filtering by inputting the reference image target object into the high-pass filter. The high-pass filter retains the high-frequency components in the image while weakening or removing the low-frequency components through a series of frequency analysis and transformation operations.
[0078] The operation process of high-pass filtering is to first transform the image from the spatial domain to the frequency domain through the fast Fourier transform (FFT). In the frequency domain, the low-frequency information of the image is mainly concentrated in the central area of the spectrum, while the high-frequency information is distributed in the peripheral area of the spectrum. By properly shielding or weakening the center of the spectrum and retaining the high-frequency components in the periphery, high-pass filtering can be effectively implemented to extract the high-frequency area in the image. The extracted high-frequency area is remapped back to the spatial domain to obtain the high-frequency feature image of the reference image. This high-frequency feature image usually contains the edge information and detail features of the target object and is suitable for subsequent alignment operations.
[0079] The second step is alignment, that is, in order to align this high-frequency feature image to the scene image, it is necessary to perform feature matching and alignment on the high-frequency feature image of the reference image and the scene image.
[0080] The alignment process includes adjusting parameters such as image scale, rotation, and position so that the key points in the high-frequency feature image are accurately aligned with the corresponding positions in the scene image.
[0081] When realizing alignment, the key feature points in the two images are first extracted through the SIFT (Scale Invariant Feature Transform) algorithm. Then, the two images are matched based on these feature points to obtain a set of corresponding feature points. By fitting the coordinates of these corresponding feature points, the affine transformation matrix between the two images can be obtained. Using this affine transformation matrix, the high-frequency feature image of the reference image is affine transformed, including operations such as translation, scaling, and rotation, so that the positions of the feature points in the high-frequency feature image and the scene image completely coincide, thereby achieving accurate alignment.
[0082] Step S140: extracting a multi-level high-resolution feature map from the scene fusion map through an encoder and a decoder.
[0083] These feature maps include coarse-grained feature maps and fine-grained feature maps of the scene fusion map.
[0084] The scene fusion graph is first input to the encoder for multi-level processing. In the encoder, the resolution of the image is reduced layer by layer through multiple downsampling operations. In each downsampling, the downsampling ratio is first determined. If the image size needs to be reduced by half, the downsampling ratio can be set to 0.5. Then, after determining the new image size, the next step is to select the nearest neighbor interpolation algorithm for downsampling.
[0085] Traverse each pixel of the new image to determine its corresponding position in the original image. For each target pixel in the new image, find the pixel closest to the target pixel in the original image, and then assign the value of the nearest neighbor pixel directly to the corresponding pixel in the new image. Continue the above steps until all pixels in the new image are assigned values. The second step of downsampling is to perform convolution and increase the number of feature channels layer by layer. The downsampling operation gradually reduces the spatial resolution of the image, while the number of channels of the feature map gradually increases. After each downsampling, the encoder outputs the feature map at the current resolution, thereby generating feature layers of multiple resolutions.
[0086] The decoder then performs multi-layer upsampling on the output features of the encoder and gradually restores them to high resolution. The upsampling method uses bicubic interpolation to calculate the value of the target pixel by considering the surrounding 16 pixels. Its calculation process is similar to bilinear interpolation, but a cubic polynomial function is used for interpolation during calculation. The second step of upsampling is to fuse the feature map at the current resolution with the feature map of the corresponding layer of the encoder, specifically by merging feature channels in the same resolution layer. Through upsampling, the decoder gradually restores low-resolution features to higher resolutions while reducing the number of feature channels in each layer of upsampling. When the decoder restores the resolution of the feature map layer by layer, it can integrate image information of different resolutions so as to finally obtain a high-resolution feature map containing rich details. After multi-layer upsampling by the decoder, the multi-level feature information of the scene fusion map is gradually reconstructed into a multi-level high-resolution feature map, which is used as the output of step S140.
[0087] Step S150, obtaining a corresponding identity feature vector of the target object in the reference image through a pre-trained identity extraction network.
[0088] The first step in pre-training the identity extraction network is to divide the target objects in the reference image into fixed-sized image patches.
[0089] The size of these image blocks is 16x16 pixels, and each image block is then expanded into a one-dimensional vector. Each one-dimensional vector contains the local feature information of the target object in the reference image. These one-dimensional vectors are projected into the high-dimensional space of the model through a linear projection layer to form the initial image block embedding vector. Position encoding is added to each image block embedding vector to form a position-encoded image block position feature vector. This encoding provides a specific position identifier for each image block, so that the model can know the specific position of each image block in the original image.
[0090] The second step of the pre-trained identity extraction network is to calculate the relationship between the feature vector of each image patch position and the feature vectors of other image patch positions through the self-attention module to capture the global information between image patches.
[0091] Specifically, the self-attention module calculates the relationship between each image block through three matrices: query matrix Q, key matrix K, and value matrix V. The query and key matrices are used to calculate the correlation between two image blocks, while the value matrix contains the actual information of each image block. The calculation results of the self-attention module are used as the global feature representation of each image block.
[0092] The third step of pre-training the identity extraction network is to pass the output of each layer of self-attention module through residual connection and layer normalization operation.
[0093] Among them, the residual connection directly adds the image block position feature vector of each layer to the global feature representation of the image block, so that information can be directly transmitted between different layers, avoiding the gradient vanishing problem caused by the depth of the model. Layer normalization is used to standardize the global feature representation of the image block of each layer, so that the feature distribution of the model at each layer is more consistent. Specifically, in the layer normalization process, the mean and variance of the output of the layer are first calculated. For each input sample, the output feature vector of the layer is taken out, and the values of all features will be used to calculate the mean and variance. Next, the features are normalized. The calculated mean is subtracted from each feature value, and then divided by the standard deviation. This process converts the feature value into a distribution with a mean of zero and a variance of one, ensuring that the features are on the same scale. Layer normalization introduces two learnable parameters: scaling factor and bias. The scaling factor is used to adjust the scale of the normalized feature value, while the bias is used to adjust the offset of the normalized feature value. Finally, layer normalization is applied to the output of each sample, and the mean and variance are calculated independently for each sample for normalization.
[0094] The fourth step of pre-training the identity extraction network is to obtain the identity feature vector of the final target object by subjecting the feature representation output in the previous step to two layers of linear transformation and an activation function. The identity feature vector is used as the output of step S150.
[0095] Step S160: The identity feature vector of the target object is injected into the diffusion model through a cross-attention mechanism.
[0096] The first step is to inject the identity feature vector and generate a key-value matrix.
[0097] Specifically, the identity feature vectors are passed through two linear transformation layers to generate a key matrix K and a value matrix V.
[0098] The second step is to generate the query matrix.
[0099] Specifically, the state vector of the diffusion model at time t is extracted, and the state vector is passed through a linear transformation layer to obtain the query matrix Q. The size of the query matrix Q, key matrix K, and value matrix V are all N×dk, where dk is the feature dimension after projection.
[0100] The third step is to use the key-value matrix and the query matrix to calculate the similarity and output the self-attention weight matrix.
[0101] Specifically, the query matrix Q and key matrix K generated previously are used as input, and the key matrix K is transposed to obtain K^T. The query matrix Q is multiplied by the transposed key matrix K^T, and the similarity scores between the features are calculated according to the formula Q×K^T / √dk, and a similarity matrix of size N×N is output, where each element in the similarity matrix represents the correlation weight between a pair of features. Each element in the similarity matrix is divided by √dk to ensure numerical stability. The normalized similarity matrix is processed using the softmax function, and the self-attention weight matrix is output, where the correlation weight of each feature with other features in the matrix is between 0 and 1, and the sum of all weights is 1.
[0102] The fourth step is to obtain the self-attention output matrix by weighted summation of the self-attention weight matrix.
[0103] Specifically, take the self-attention weight matrix output by the third step and the value matrix V output by the first step. For each row of the weight matrix and each row of the value matrix V, multiply their elements one by one, sum the multiplication results by row, and get the weighted self-attention output matrix with a size of N×dk, where each element contains the global information of the input feature matrix.
[0104] The fifth step is to introduce a multi-head self-attention module into the self-attention output matrix to obtain the multi-head self-attention output matrix.
[0105] Specifically, the query matrix Q, key matrix K, and value matrix V are split into h different heads. At this time, in each head, the size of the Q, K, and V matrices becomes N×dk / h. For each head, repeat the following self-attention steps. The self-attention output matrices of all heads are merged into an overall feature matrix through a concatenation operation. The merged overall feature matrix is mapped back to the original feature dimension size through a linear transformation layer.
[0106] In step S170, the multi-level high-resolution feature map, that is, the output of step S140, is used as the input of this step. The high-resolution feature map is injected into the decoder of the diffusion model by feature splicing to achieve background preservation of the scene image and the target object.
[0107] The highest resolution of the multi-level high-resolution feature map is 256 × 256, the lowest resolution is 32 × 32, and it increases layer by layer to 256 × 256. During the injection process, the same-scale features are extracted from the feature map of the corresponding resolution each time and spliced with the feature map of the corresponding level of the decoder of the diffusion model.
[0108] During splicing, the resolution of the first layer of the decoder of the diffusion model is 32×32, and the number of feature channels is 256, while the corresponding feature map resolution in the high-resolution feature map is adjusted to 32×32, and the number of channels is 512. The two feature maps are spliced along the channel dimension to form a spliced feature map, and the size of the spliced feature map becomes 32×32×768.
[0109] The concatenated feature map is convolved to increase nonlinear characteristics, where the convolution kernel size is 1 × 1. Repeat the above steps, and each layer of feature map in the decoder will repeat this concatenation and convolution process.
[0110] Step S180: The diffusion model generates a feature map of the synthetic image through multiple iterations of denoising.
[0111] The first step is the denoising initialization of the diffusion model.
[0112] Specifically, the size of the generated noise image is set to 256×256×3, and a Gaussian distribution with zero mean and unit variance is used to randomly sample the R, G, and B values of each pixel in the noise image to generate a completely random initial noise image, and the total denoising steps T of the diffusion model are set.
[0113] The second step is to perform denoising on the initialized diffusion model one by one and output the feature map of the synthetic image.
[0114] Specifically, for each time step t, compare it with the noise image of the previous step, that is, the noise image generated in step t-1. If t=1, it is the initial noise image, and calculate the noise level of the current image. Use the neural network model to estimate the noise of each pixel of the current noise image of step t, and generate a noise prediction map with the same size as the current image. Subtract the noise prediction value from the current noise image to obtain a new denoised image. Adjust the amount of noise reduced in each step so that the generated new noise image is closer to the noise-free state than the noise image in the previous step, and output the feature map of the synthesized image in step T.
[0115] Step S190: The feature map of the composite image is decoded by a variational autoencoder to obtain a final composite image.
[0116] The first step is to input the synthetic image feature map into the encoder of the variational autoencoder and output the reparameterized feature map.
[0117] Specifically, a 3×3 convolution kernel and 64 output channels are used to perform a convolution operation on the feature map of the synthetic image to obtain a first convolution result. The first convolution result is batch normalized to reduce the internal covariate shift, and a ReLU activation function is used for nonlinear processing, and then a 3×3 convolution kernel is used for convolution to obtain a second convolution result. The second convolution result is batch normalized, and a ReLU activation function is used for nonlinear processing to obtain a reparameterized feature map.
[0118] The second step is to use the reparameterization technique on the reparameterized feature map to output the reparameterized latent variables.
[0119] Specifically, the parameters of the latent space, namely the mean and variance, are constructed using the reparameterized feature map. The standard normal distribution noise is randomly sampled to generate latent variables, which are multiplied by the square root of the variance and then added to the mean to output the reparameterized latent variables.
[0120] The third step is to input the re-parameterized latent variables into the decoder for image synthesis.
[0121] Specifically, the re-parameterized latent variables are input into the decoder, and the latent variables are mapped back to the higher-dimensional feature space using a fully connected layer. In this process, multiple linear layers and activation functions are used to expand the points in the latent space into the corresponding feature map. The expanded feature map is subjected to the first convolution with a 3×3 convolution kernel and 64 output channels to obtain the first decoding convolution result. The first decoding convolution result is batch normalized and ReLU activated to obtain the first decoding feature. The first decoding feature is convolved with a 3×3 convolution kernel again to obtain the second decoding convolution result, which is then batch normalized and ReLU activated. A series of upsampling operations are performed on the second decoding convolution result to gradually expand the size of the feature map to the target size of the synthetic image. The last convolution layer is used to map the decoding feature to the number of channels of the target image, and the final synthetic image is output.
[0122] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A zero-sample fast object migration method based on a diffusion model, characterized in that: The steps include: Step S110, obtaining a reference image and a scene image; Step S120: input the reference image into a preset segmentation detection model to obtain a target object in the reference image; Step S130, extracting a high-frequency region from the target object in the reference image by high-pass filtering, and aligning the extracted high-frequency region to the scene image as a scene fusion image; Step S140, extracting a multi-level high-resolution feature map from the scene fusion map through an encoder and a decoder; Step S150, obtaining an identity feature vector of the target object in the reference image through a pre-trained identity extraction network; Step S160, injecting the identity feature vector of the target object into the diffusion model through a cross-attention mechanism; Step S170, injecting the multi-level high-resolution feature map in step S140 into the decoder of the diffusion model by feature splicing, so as to achieve background preservation of the scene image and the target object; Step S180, the diffusion model generates a feature map of the synthetic image through multiple iterations of denoising; Step S190: The feature map of the composite image is decoded by a variational autoencoder to obtain a final composite image.
2. The zero-sample rapid object migration method based on a diffusion model according to claim 1, characterized in that: The reference image and the scene image in step S110 are derived from a data set, and the process of making the data set includes image filtering, and the image filtering includes the following steps: Step S110.1, preprocessing the original image; Step S110.2, using median filtering to detect and remove noise; Step S110.3, evaluating the clarity of the image; Step S110.4: Perform quality assessment.
3. The zero-sample rapid object migration method based on a diffusion model according to claim 1 or 2, characterized in that: The preset segmentation detection model described in step S120 includes an image encoder, a text encoder and a mask decoder, wherein: The input of the image encoder is the reference image, which is divided into image blocks of fixed size, each of which is flattened and mapped to a high-dimensional space; each image block is then linearly transformed to obtain a corresponding embedding vector; the embedding vector passes through multiple Transformer encoder layers, each of which contains a self-attention module, a feedforward neural network, a residual connection and layer normalization, and finally outputs the features of the reference image; The input of the text encoder is a preset text. The text is first segmented, and each word or subword is converted into a corresponding word embedding. The word embedding is added to the position code, and the added word embedding passes through multiple Transformer encoder layers. The output of the last layer of encoder is globally pooled, and the pooled vector is processed through a linear layer and an activation function to finally obtain a feature vector of the text; The input of the mask decoder is the output of the image encoder and the text encoder. Comprehensive features are generated by vector splicing. The comprehensive features are nonlinearly transformed through a feedforward neural network to extract feature representations, which are then processed through a linear layer to output a segmentation mask. The segmentation mask is further refined through a conditional random field, and finally the refined segmentation mask is used to mark the target object in the reference image.
4. The zero-sample rapid object migration method based on a diffusion model according to claim 3, characterized in that: In step S130, extracting a high-frequency area from the target object in the reference image by high-pass filtering includes: The target object in the reference image is input into a high-pass filter, and through frequency analysis and transformation operations, the high-frequency components in the image are retained, while the low-frequency components are weakened or removed, and the high-frequency area in the image is extracted. The extracted high-frequency area is remapped back to the spatial domain to obtain a high-frequency feature image of the reference image; The step of aligning the extracted high-frequency region to the scene image as a scene fusion image includes adjusting the scale, rotation and position of the image so that key points in the high-frequency feature image are aligned with corresponding positions in the scene image.
5. The zero-sample rapid object migration method based on a diffusion model according to claim 4, characterized in that: Step S140 includes: The scene fusion image is first input to the encoder for multi-level processing. In the encoder, the resolution of the image is reduced layer by layer through multiple downsampling operations to generate feature layers with multiple resolutions. The decoder performs multi-layer upsampling on the output features of the encoder and gradually restores them to high resolution. The multi-level feature information of the scene fusion graph is gradually reconstructed into multi-level high-resolution feature graphs.
6. The zero-sample rapid object migration method based on a diffusion model according to claim 5, characterized in that: Step S150 includes: Step S150.1, dividing the target object in the reference image into image blocks of fixed size; Step S150.2, the relationship between each image block position feature vector and other image block position feature vectors is calculated by the self-attention module of the pre-trained identity extraction network, and the calculation result of the self-attention module is used as the global feature representation of each image block; Step S150.3, the output of each layer of the self-attention module is subjected to a residual connection and a layer normalization operation, wherein the residual connection is used to directly add the image block position feature vector of each layer to the global feature representation of the image block, and the layer normalization operation is used to standardize the global feature representation of the image block of each layer; Step S150.4: The global feature representation outputted from step S150.3 is transformed through two layers of linear transformation and an activation function to obtain the identity feature vector of the final target object.
7. The zero-sample rapid object migration method based on a diffusion model according to claim 6, characterized in that: Step S160 includes: Step S160.1, passing the identity feature vector through two linear transformation layers respectively to generate a key matrix K and a value matrix V; Step S160.2, extract the state vector of the diffusion model at time t, and pass the state vector through a linear transformation layer to obtain a query matrix Q; Step S160.3, using the key matrix K, the value matrix V and the query matrix Q to perform similarity calculation, and output the self-attention weight matrix; Step S160.4, weighted summing of the self-attention weight matrix to obtain a self-attention output matrix; Step S160.5, introduce a multi-head self-attention module into the self-attention output matrix to obtain a multi-head self-attention output matrix.
8. The zero-sample rapid object migration method based on a diffusion model according to claim 7, characterized in that: In step S170: During the injection process, features of the same scale are extracted from the feature map of the corresponding resolution each time and concatenated with the feature map of the corresponding level of the decoder of the diffusion model.
9. The zero-sample rapid object migration method based on a diffusion model according to claim 8, characterized in that: Step S180 includes: Step S180.1, setting the size of the generated noise image, using a Gaussian distribution with zero mean and unit variance to randomly sample the R, G, and B values of each pixel of the noise image, and generating a completely random initial noise image; Step S180.2, perform successive denoising on the initialized diffusion model, and output a feature map of the synthesized image.
10. The zero-sample rapid object migration method based on a diffusion model according to claim 9, characterized in that: Step S190 includes: Step S190.1, input the synthetic image feature map into the encoder of the variational autoencoder, and output a re-parameterized feature map; Step S190.2, using the reparameterization technique to output the reparameterized latent variables on the reparameterized feature map; Step S190.3: input the re-parameterized latent variables into the decoder for image synthesis.
Citation Information
Cited By
Zero sample cross-domain diffusion segmentation method based on anatomical structure probability transmission guidance
CN121121104A
Image processing method and device
CN122493222A
An image processing method and apparatus
CN122493222B