Image completion method based on multi-scale diffusion model, medium and equipment

The multi-scale diffusion model addresses the limitations of deep learning in image inpainting by integrating semantic information and Fourier convolution to enhance the quality and coherence of inpainted images, particularly in complex scenes.

CN120318096APending Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510373677.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing deep learning methods have problems such as limitations in mask generalization, insufficient combination of local details and global structural information, and difficulty in recognition in complex semantic scenarios in the image completion task, resulting in unnatural and incoherent completion results.

Method used

Using a multi-scale diffusion model, the repair results are gradually refined by introducing semantic information fusion at high resolution and using Fourier convolution at low resolution, and combining the extraction of global and local information at multi-scale to generate high-quality complete images.

Benefits of technology

Improves the quality of image repair and the robustness of algorithms, and is suitable for complex scenes and large-scale missing areas, and the generated images are coherent and authentic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318096A_ABST
    Figure CN120318096A_ABST
Patent Text Reader

Abstract

The invention provides an image completion method based on a multi-scale diffusion model, a medium and equipment, and relates to the technical field of image completion. According to the method, firstly, a semantic information fusion step is introduced in a high-resolution complementing process, and semantic information is fused into a high-resolution repaired image, so that semantic consistency of a repaired area and a surrounding environment is ensured, and the overall visual effect of the complemented image is improved. Meanwhile, an input high-resolution incomplete image is subjected to down-sampling to generate a low-resolution incomplete image, and fast Fourier convolution is introduced in a low-resolution complementation process so as to better capture global features of the image. And then, carrying out up-sampling on the repaired image under the low resolution, and carrying out feature fusion on the repaired image under the low resolution and the repaired image under the high resolution to generate a final complemented image. Therefore, the image restoration quality and the algorithm robustness are improved by processing the image at different resolutions, gradually refining the restoration result and fusing global and local information extracted under multiple scales.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image completion technology, and particularly to an image completion method, medium, and device based on a multi-scale diffusion model. Background Art

[0002] The purpose of image completion is to fill in the missing regions in an image, and the regions to be completed need to be coordinated with other regions. This task is of great significance in many practical applications, such as photo restoration, video editing, and artistic creation. Traditional image completion mainly relies on manually designed rules and algorithms. However, these methods often perform poorly when dealing with complex scenes and large-scale missing regions.

[0003] With the advent of the big data era, deep learning methods have gradually been widely applied to the image completion task, improving the efficiency and accuracy of image completion. However, they still have limitations in aspects such as mask generalization and completion coherence, resulting in a deviation between the completed result and the expectation, which not only affects the visual effect of image completion but also limits the practicality of image completion technology in a wider range of application scenarios. Summary of the Invention

[0004] This application provides an image completion method, medium, and device based on a multi-scale diffusion model, which can effectively improve the repair quality and naturalness of images and generate high-quality and highly consistent completed images in various scenarios.

[0005] In the first aspect of the embodiments of this application, an image completion method based on a multi-scale diffusion model is provided. The method includes: Obtain a target incomplete image and text description information; wherein, the target incomplete image contains mask information, and the mask information is used to distinguish the region to be repaired in the target incomplete image, and the text description information is used to describe the physical characteristics of the missing part of the target incomplete image; Based on the text description information and a pre-set diffusion model, repair the target incomplete image at a first resolution to obtain a first repaired image; Downsample the target incomplete image, and based on Fourier transform, repair the target incomplete image at a second resolution to obtain a second repaired image; wherein, the second resolution is less than the first resolution; Upsample the second repaired image to obtain a third repaired image; Perform pixel-by-pixel superposition on the first repaired image and the third repaired image to obtain the completed image corresponding to the target incomplete image.

[0006] Optionally, the pre-set diffusion model includes a plurality of consecutive encoding layers and a plurality of decoding layers connected to the encoding layers, and there are skip connections between the symmetric encoding layers and decoding layers; Based on the text description information and a pre - set diffusion model, repair the target incomplete image at the first resolution to obtain a first repaired image, including: Encode the text description information to obtain semantic features; Successively pass through multiple consecutive encoding layers, and based on the semantic features, extract features from the target incomplete image to obtain a feature image; Successively pass through multiple decoding layers to gradually denoise the feature image within T time steps to complete the target incomplete image and obtain the first repaired image.

[0007] Optionally, any one of the encoding layers includes a self - attention layer, a cross - attention layer, and a fully - connected layer; successively passing through multiple consecutive encoding layers and extracting features from the target incomplete image based on the semantic features to obtain a feature image includes: For any one of the multiple consecutive encoding layers, through the self - attention layer, extract features from the input image based on the self - attention weights to obtain a first feature image; Through the cross - attention layer, adaptively adjust the cross - attention weights of the input image based on the semantic features, and based on the cross - attention weights, extract features from the input image to obtain a second feature image; Through the fully - connected layer, fuse the first feature image and the second feature image to obtain an output image; Wherein, the input image of the first encoding layer is the target incomplete image, the input image of the remaining encoding layers is the output image of the previous encoding layer, and the output image of the last encoding layer is the feature image.

[0008] Optionally, the cross - attention layer includes a multi - head attention layer, a first addition layer, a feed - forward network layer, and a second addition layer; adaptively adjusting the cross - attention weights of the input image based on the semantic features and extracting features from the input image based on the cross - attention weights to obtain a second feature image includes: Through the multi - head attention layer, determine the cross - attention weights from different dimensions based on the semantic features, and extract features from the input image based on the cross - attention weights to obtain image features in multiple dimensions; Through the first addition layer, fuse the image features in multiple dimensions and perform normalization processing to obtain a first process image; Through the feed - forward network layer, perform a non - linear transformation on the first process image and activate it through an activation function to obtain a second process image; Through the second addition layer, the first process image and the second process image are added and normalized to obtain the second feature image.

[0009] Optionally, the method further includes: For any dimension, determine the similarity score between the semantic feature and the initial image feature of the input image; the initial image feature is obtained by performing initial feature extraction on the input image through convolution; Determine the cross-attention weight of the input image based on the similarity score; Multiply the cross-attention weight by the semantic feature to obtain the image feature in this dimension.

[0010] Optionally, based on Fourier transform, perform repair on the target incomplete image at the second resolution to obtain a second repaired image, including: Divide the number of channels of the target incomplete image into two parts to obtain a first channel feature and a second channel feature; Perform two-layer convolution operations and a first operation on the first channel feature to obtain the local feature of the target incomplete image; the first operation includes batch normalization and non-linear activation; Perform frequency domain transformation and the first operation on the second channel feature to obtain the global feature of the target incomplete image; Concatenate the local feature and the global feature in the channel dimension to obtain the second repaired image.

[0011] Optionally, performing frequency domain transformation and the first operation on the second channel feature to obtain the global feature of the target incomplete image includes: Pass the second channel feature through one convolution and non-linear activation to obtain a first spatial domain feature; Convert the first spatial domain feature from the spatial domain to the frequency domain through two-dimensional real fast Fourier transform, and extract the frequency domain feature through one convolution and non-linear activation; Convert the frequency domain feature from the frequency domain back to the spatial domain through inverse Fourier transform to obtain a second spatial domain feature; Perform element-wise addition on the first spatial domain feature and the second spatial domain feature to obtain an addition feature; Perform the first operation on the addition feature to obtain the global feature of the target incomplete image.

[0012] Optionally, the diffusion model is trained through the following steps: Input the sample incomplete image and the sample text description information into the initial model to obtain a predicted completed image; Determine the differences between the true complete image of the incomplete sample image and the predicted completed image on each feature layer, and perform weighted summation on the differences on each feature layer to obtain a perceptual distance measurement value; Determine the cosine similarity between the predicted completed image and the sample text description information, and perform normalization processing on the cosine similarity to obtain a contrastive learning score; Based on the perceptual distance measurement value and the contrastive learning score, determine the loss function value; Based on the loss function value, optimize and adjust the initial model to obtain the diffusion model.

[0013] Based on the same inventive concept, a second aspect of the embodiments of the present application provides an image completion device based on a multi-scale diffusion model. The above device includes: An image acquisition module, configured to acquire a target incomplete image and text description information; wherein, the target incomplete image includes mask information, and the mask information is used to distinguish the area to be repaired in the target incomplete image, and the text description information is used to describe the physical characteristics of the missing part of the target incomplete image; A first repair module, configured to repair the target incomplete image at a first resolution based on the text description information and a pre-set diffusion model to obtain a first repaired image; A second repair module, configured to downsample the target incomplete image and repair the target incomplete image at a second resolution based on Fourier transform to obtain a second repaired image; wherein, the second resolution is less than the first resolution; An upsampling module, configured to upsample the second repaired image to obtain a third repaired image; An overlay module, configured to perform pixel-by-pixel overlay on the first repaired image and the third repaired image to obtain a completed image corresponding to the target incomplete image.

[0014] Based on the same inventive concept, a third aspect of the embodiments of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the image completion method based on a multi-scale diffusion model proposed in the first aspect of the present application.

[0015] Based on the same inventive concept, a fourth aspect of the embodiments of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the image completion method based on a multi-scale diffusion model proposed in the first aspect of the present application.

[0016] Compared with the prior art, the present application includes the following advantages: An image completion method based on a multi-scale diffusion model provided by an embodiment of the present application first introduces a semantic information fusion step in the high-resolution completion process, integrates semantic information into the restored high-resolution image to ensure the semantic consistency between the restored area and the surrounding environment, and improves the overall visual effect of the completed image. At the same time, the input high-resolution incomplete image is downsampled to generate a low-resolution incomplete image, and fast Fourier convolution is introduced in the low-resolution completion process to better capture the global features of the image. Then, the restored image at low resolution is upsampled and feature-fused with the restored image at high resolution to generate the final completed image. Thus, by processing the image at different resolutions, the restoration result is gradually refined, and by fusing the global and local information extracted at multiple scales, the quality of image restoration and the robustness of the algorithm are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of an image completion method based on a multi-scale diffusion model in an embodiment of the present application; Figure 2 is a logical schematic diagram of an image completion method based on a multi-scale diffusion model in an embodiment of the present application; Figure 3 is a structural schematic diagram of a pre-set diffusion model in an embodiment of the present application; Figure 4 is a structural schematic diagram of a cross-attention layer of a pre-set diffusion model in an embodiment of the present application; Figure 5 is a logical schematic diagram of a Fourier transform in an embodiment of the present application; Figure 6 is a functional module schematic diagram of an image completion device based on a multi-scale diffusion model in an embodiment of the present application; Figure 7 is a structural schematic diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0020] Deep learning methods have been widely applied to image inpainting tasks, greatly improving the efficiency and accuracy of image inpainting. However, existing deep learning methods still have the following drawbacks in image inpainting tasks: 1. There are limitations in adapting to various types of masks. Although deep learning models have powerful feature learning capabilities, when faced with irregular, complex, or even dynamically changing masks, it is difficult for the model to flexibly adjust the filling strategy, which may result in the inpainted area appearing rigid or unnatural.

[0021] 2. Although deep learning methods can automatically learn multi-level features from data, they often lack the effective combination of local details and global structure information, and are prone to the problem of "reasonable locally but distorted globally". For example, the model may generate clear details in a small area, but when integrated into the large image, it is difficult to maintain coherence, resulting in a disharmonious visual feeling in the image.

[0022] 3. Although deep learning methods can guide inpainting through semantic information, in scenes with complex semantics, it may be difficult for the model to accurately recognize and understand deep semantic structures, thus affecting the logical rationality and detail coherence of inpainting. Compared with traditional methods that simply rely on handcrafted features, although the semantic guidance of deep learning is more intelligent, it is vulnerable to data bias or noise in practical applications, resulting in the inpainting result deviating from the real scene.

[0023] In view of this, the present application proposes an image inpainting method based on a multi-scale diffusion model. First, a semantic information fusion step is introduced in the high-resolution inpainting process to integrate semantic information into the high-resolution restored image to ensure the semantic consistency between the inpainted area and the surrounding environment and improve the overall visual effect of the inpainted image. At the same time, the input high-resolution incomplete image is downsampled to generate a low-resolution incomplete image, and fast Fourier convolution is introduced in the low-resolution inpainting process to better capture the global features of the image. Then, the restored image at low resolution is upsampled and feature-fused with the restored image at high resolution to generate the final inpainted image. Thus, by processing the image at different resolutions, the restoration result is gradually refined, and by fusing the global and local information extracted at multiple scales, the quality of image restoration and the robustness of the algorithm are improved, and it has good generalization ability.

[0024] Please refer to Figure 1 ,Figure 1 is a flowchart of an image completion method based on a multi-scale diffusion model proposed in an embodiment of the present application. As Figure 1 shown, the method includes the following steps: S101: Obtain a target incomplete image and text description information; wherein, the target incomplete image contains mask information for distinguishing the area to be repaired in the target incomplete image, and the text description information is used to describe the physical characteristics of the missing part of the target incomplete image; In this embodiment, the target incomplete image refers to the image to be completed and restored this time, such as Figure 2 the image in shown; the target incomplete image contains mask information for distinguishing the area to be repaired in the target incomplete image, such as Figure 2 the image in shown; the text description information is used to describe the physical characteristics of the missing part of the target incomplete image, such as Figure 2 the word "white bear" in

[0025] Among them, the mask information can be represented by a binary image. The white area in the binary image represents the area to be repaired, while the black area represents the area that does not need to be repaired. Thus, in the subsequent image completion process, the area of image processing operations is restricted, ensuring that only the white area to be repaired is processed and avoiding interference with the black area that does not need to be modified. This not only helps to generate a more natural and accurate completion result, but also saves computing resources and improves the image processing speed.

[0026] The text description information can provide additional descriptions or explanations for the target incomplete image to be completed, highlighting the key elements or important information in the image. Adding text description information to the target incomplete image can enable the image completion algorithm to obtain richer context semantic information, and thus better understand the content and background of the image, promoting the visual effect of image completion. For example, when repairing Figure 2 the image in , words such as "bear", "white", "left hand", "holding" can be added to describe the physical characteristics such as the shape outline, color, and limb movements of the object of the missing part of the target incomplete image, enhancing the readability and comprehensibility of the image.

[0027] S102: Based on the text description information and a preset diffusion model, repair the target incomplete image at the first resolution to obtain a first repaired image.

[0028] In this embodiment, repairing the target incomplete image at the first resolution refers to the repair process for the target incomplete image with the original resolution. The preset diffusion model is constructed based on the structure of the Denoising Diffusion Probabilistic Models (abbreviated as DDPM), and is an image completion model obtained after training based on a large number of sample incomplete images and corresponding text description information.

[0029] The DDPM model includes two processes: the forward process (diffusion process) and the reverse process (inverse diffusion process). In the forward process, the model gradually adds noise to the image data, making the image data closer and closer to random noise. This process can be regarded as the process of gradually transforming the image data distribution into a simple prior distribution (usually a Gaussian distribution). In the reverse process, the model starts from simple Gaussian noise and gradually removes the noise, finally restoring the real image. This process is realized by training a neural network, which learns to predict and remove noise at each time step.

[0030] In this embodiment, it is necessary to gradually add noise to the target incomplete image first, and then input it into the preset diffusion model for gradual denoising to restore the real image. That is to say, the target incomplete image in this embodiment can be understood as an incomplete image carrying random noise itself, and needs to be completed through the preset diffusion model.

[0031] Furthermore, in the process of using the preset diffusion model for image repair in this embodiment, the text description information corresponding to the target incomplete image is also combined to ensure the semantic consistency between the repaired area and the surrounding environment by fusing the context semantic features represented by the text description information into the first repair map at the original resolution, making the repaired area more real and natural, and finally outputting a high-quality repaired image to improve the overall visual effect of the repaired image.

[0032] In this embodiment, the target incomplete image and the text description information are input into the preset diffusion model to repair the incomplete area by combining semantic information at the first resolution, and a first repair map is obtained.

[0033] S103: Downsample the target incomplete image, and repair the target incomplete image at the second resolution based on Fourier transform to obtain a second repair map; wherein, the second resolution is less than the first resolution.

[0034] In this embodiment, repairing the target incomplete image at the second resolution means repairing the downsampled target incomplete image. Since the resolution of the target incomplete image will decrease after downsampling. Therefore, compared with the repair process of the target incomplete image at the first resolution (original resolution) in step S102, the repair process at the second resolution in step S103 can be regarded as a repair at low resolution, while the repair process at the first resolution in step S102 can be regarded as a repair at high resolution.

[0035] Specifically, downsampling can be achieved by the bilinear interpolation method. Downsampling the target incomplete image helps to perform preliminary repair at low resolution, reduce the computational burden, and obtain more complete global image information.

[0036] Furthermore, in the low-resolution repair process, this embodiment is based on the Fourier transform method (i.e., Figure 2 the completion module integrating the Fourier transform in it) to process the low-resolution target incomplete image. After processing the target incomplete image by converting it from the spatial domain to the frequency domain and then back to the spatial domain, large-scale structural information can be captured to obtain the second repaired image, providing a basis for subsequent repair.

[0037] Among them, converting the target incomplete image from the spatial domain to the frequency domain can reveal the spectral information of the image, which helps to analyze the characteristics of different frequency components in the target incomplete image. For example, low frequencies usually correspond to smooth regions in the image, while high frequencies correspond to edges and details in the image. Thus, by understanding these frequency components, image completion can be more targeted to retain and restore important details in the image. At the same time, it can convert complex convolution operations in the spatial domain into simple multiplication operations in the frequency domain, thereby accelerating the calculation process and improving the efficiency of the algorithm.

[0038] S104: Upsample the second repaired image to obtain the third repaired image.

[0039] In this embodiment, by performing upsampling processing on the second repaired image, the repair result at low resolution is enhanced to high resolution to obtain the high-resolution third repaired image. That is, the resolution of the third repaired image is the same as the original resolution of the target incomplete image.

[0040] Exemplarily, upsampling can also be achieved by the bilinear interpolation method.

[0041] S105: Superimpose the first repaired image and the third repaired image pixel by pixel to obtain the completed image corresponding to the target incomplete image.

[0042] In this embodiment, by pixel-by-pixel superimposing the third repaired image obtained after repairing and upsampling at a low resolution with the first repaired image that has fused semantic information at a high resolution, it is possible to effectively coordinate the semantic information, global information, and local information of the image, making the completed image more realistic and complete.

[0043] In this embodiment, by gradually refining the repair results at different resolutions, high-quality image repair is achieved. This layer-by-layer refinement method not only improves the repair quality of the image but also enhances the robustness of the algorithm, and is applicable to repair scenarios of complex scenes and large-scale missing areas. At the same time, multi-scale processing at different resolutions helps capture the global structure of the image and restore the local details of the image, making the finally completed image coherent and realistic.

[0044] During the repair process at a high resolution, optionally, the preset diffusion model includes a plurality of consecutive encoding layers and a plurality of decoding layers connected to the encoding layers, and there are skip connections between the symmetric encoding layers and decoding layers.

[0045] Exemplarily, please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the preset diffusion model in an embodiment of the present application. As Figure 3 shown, the diffusion model adopts a symmetric U-Net architecture, including an encoder and a decoder. Among them, the encoder includes 4 consecutive encoding layers, the decoder includes 4 consecutive decoding layers, and there are skip connections between the symmetric encoding layers and decoding layers to retain more detailed information during the image completion process and improve the quality of image repair.

[0046] Based on the text description information and the preset diffusion model, performing repair on the target incomplete image at the first resolution to obtain the first repaired image, including: S201: Encoding the text description information to obtain semantic features.

[0047] In this embodiment, in order to enable the diffusion model to understand the semantic relationship between the image and the text, it is necessary to process the text description information, convert it into a vector form to obtain semantic features, and then input the semantic features into the diffusion model to assist in image repair.

[0048] Specifically, the text encoder of the CLIP model (Contrastive Language-Image Pre-Training, a multi-modal pre-trained neural network model) can be used to convert the text description information into a fixed-length vector representation to obtain the semantic features of the target incomplete image.

[0049] S202: Sequentially passing through a plurality of consecutive encoding layers to perform feature extraction on the target incomplete image based on the semantic features to obtain a feature image.

[0050] In this embodiment, by inputting the target incomplete image and semantic features into the encoder, feature extraction is performed layer by layer. Among them, the processing process of each encoding layer is the same. Feature extraction is performed on the target incomplete image based on the semantic features, and the intermediate feature image obtained by the previous encoding layer will continue to be passed to the next encoding layer for further feature extraction. Finally, the feature image is output by the last encoding layer and passed to the decoding layer for layer-by-layer restoration.

[0051] S203: Gradually denoise the feature image through multiple decoding layers within T time steps to complete the target incomplete image and obtain the first restored image.

[0052] In this embodiment, the decoding layer will gradually denoise the feature image until a complete first restored image is generated. Moreover, the intermediate restored image obtained by denoising in the previous decoding layer will be jump-fused with the intermediate feature image output by the corresponding encoding layer of this decoding layer, and then passed to the next decoding layer for further denoising restoration. Finally, the first restored image is output by the last decoding layer, realizing the restoration of image details at high resolution.

[0053] Specifically, during the process of decoding and denoising, it is assumed that the entire process is divided into T steps. In each step, the diffusion model will try to slightly reduce the noise added to the target incomplete image. Starting from the T-th step, in each step, based on the result of the previous step, the preset diffusion model is used to predict and remove a part of the noise, gradually approaching the distribution of the original image data. As the noise is gradually removed, a clear first restored image that conforms to a specific distribution is gradually restored.

[0054] It is easy to understand that for the known regions in the target incomplete image, since they are already visible, there is no need to modify them, and only the original features need to be maintained. For the unknown regions in the target incomplete image, prediction and filling need to be performed according to the surrounding known information. Therefore, an effective mechanism needs to be designed to enable the diffusion model to make full use of the information in the known regions while avoiding overfitting or generating unreasonable completion results.

[0055] In this embodiment, by converting the mask information carried by the target incomplete image into a mask matrix, the known part and the unknown part in the target incomplete image are distinguished. Specifically, assuming that the target incomplete image is represented as x and the mask matrix is represented as m, the unknown pixels in the target incomplete image can be represented as m⊙x, and the known pixels can be represented as (1−m)⊙x.

[0056] During the process of the decoder gradually denoising and completing, assuming that the completed image at the -th step is ,then the completed image It can be expressed as: ⊙ ⊙

[0057] Wherein, represents the known region in the completed image, and the known pixels in the target incomplete image can be used for sampling. represents the unknown region in the completed image, and sampling can be performed from the diffusion model. Furthermore, by combining the sampling of the known part and the unknown part, the diffusion model can effectively generate high-quality image completion results.

[0058] In addition, in this embodiment, by converting the text description information into semantic features that guide the diffusion model repair process, it is ensured that the completed part can match the text description, thereby improving the quality of image repair.

[0059] For example, Figure 3 as shown, any encoding layer includes a self-attention layer, a cross-attention layer, and a fully connected layer. The image processing process of any encoding layer in the above step S202 includes: S202-1: For any encoding layer in multiple consecutive encoding layers, through the self-attention layer, based on the self-attention weights, feature extraction is performed on the input image to obtain a first feature image.

[0060] In this embodiment, the input image of the first encoding layer is the target incomplete image, the input images of the remaining encoding layers are the output images of the previous encoding layer, and the output image of the last encoding layer is the feature image, which is used to be passed into the decoding layer for denoising diffusion recovery to complete the complement of the target incomplete image.

[0061] Inside the encoding layer, first, the self-attention layer performs weight distribution within the input image to determine the corresponding self-attention weights, so as to focus on different parts of the input image and better extract image features to obtain the first feature image. That is, in the self-attention layer, the self-attention weights are adjusted within the image, aiming to make different parts belonging to the same object have higher correlation, so as to better capture the overall structure of the object.

[0062] Exemplarily, as Figure 3 shown in the "Attention Modulation" part in the lower right corner, for the initial input image (Key), after weight distribution and feature extraction through the self-attention layer, the first feature image (A') is generated, which shows the attention degree of the diffusion model to different parts of the input image.

[0063] S202-2: Through the cross-attention layer, adaptively adjust the cross-attention weights of the input image based on semantic features, and based on the cross-attention weights, extract features from the input image to obtain a second feature image.

[0064] During the process of image inpainting, not all image regions are equally important. Some regions may be crucial for inpainting, while others are relatively less important. In this embodiment, the cross-attention layer can intelligently distinguish which regions should receive more attention and which can be slightly ignored by adaptively adjusting the cross-attention weights based on semantic features. Then, based on the cross-attention weights, features are extracted from the input image to obtain a second feature image. This enables the diffusion model to focus on the most critical features with limited computational resources, improving the processing speed and effect. That is, the cross-attention layer adjusts the weights based on the pairing between the input image and semantic features, making the text description more relevant to the corresponding regions in the image, and prompting the diffusion model to pay more attention to the regions that match the text prompt during the repair process to improve the quality of image inpainting.

[0065] Exemplarily, as Figure 3 shown in the "Attention Modulation" part in the lower left corner, for the initial input image (Key) and the input semantic features, the cross-attention layer adjusts the attention weights according to the semantic features and the image layout to ensure that the diffusion model pays more attention to the image regions related to the text description. For example, the word "yellow" guides the diffusion model to focus on the regions related to yellow in the input image, while the word "dress" guides the diffusion model to focus on the regions related to the white dress in the input image.

[0066] S202-3: Through the fully connected layer, fuse the first feature image and the second feature image to obtain the output image.

[0067] In this embodiment, the cross-attention layer strengthens the representation of specific objects by adjusting the attention weights between semantic features and the input image, and understands the objects in the input text description information and their corresponding layout positions. The self-attention layer refines the relationships between different parts within the image. This means that different parts belonging to the same object will be more closely related to each other, making the overall structure of the object more consistent and coherent. Finally, through the application of the fully connected layer, the first feature image extracted by the self-attention layer and the second feature image extracted by the cross-attention layer are fused to obtain the output image of this encoding layer. The fusion process integrates all the obtained visual and semantic information, generating an image that is both faithful to the original text description and meets the specified layout conditions. In this process, the fully connected layer not only considers the interaction between the text and the image, but also considers the relationships between the internal elements of the image, thus ensuring the quality and consistency of the output image.

[0068] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the cross-attention layer of the pre-set diffusion model in an embodiment of the present application. Optionally, the cross-attention layer includes a multi-head attention layer, a first addition layer, a feed-forward network layer, and a second addition layer. The above step S202-2 includes: S202-2-A: Through the multi-head attention layer, cross-attention weights are determined based on semantic features from different dimensions, and feature extraction is performed on the input image based on the cross-attention weights to obtain image features in multiple dimensions.

[0069] In this embodiment, the multi-head attention layer runs the attention mechanism in parallel through multiple independent attention heads to obtain the attention distributions of different sub-spaces of the input image, so as to more comprehensively capture various potential semantic associations in the image sequence, enabling the diffusion model to simultaneously focus on different parts of the input image and better understand the context information. In this process, the input image and semantic features are decomposed into queries (Q), keys (K), and values (V), and the attention weights are generated by calculating the correlations between them. Finally, the weighted sum result is used as the cross-attention weight, and feature extraction is performed on the input image based on the cross-attention weight to obtain image features in multiple dimensions.

[0070] S202-2-B: Through the first addition layer, the image features in multiple dimensions are fused and normalized to obtain the first process image.

[0071] In this embodiment, the first addition layer includes an addition operation and a layer normalization operation. First, the image features in multiple dimensions are fused, then the fused image features and the initial image features of the input image are added through the addition operation, and finally the added image features are scaled to a preset range through the layer normalization operation to obtain the first process image.

[0072] Among them, the addition operation introduces residual connection into the network, directly adding the initial image features of the input image to the transformed output image features, which can alleviate the problem of vanishing gradients in deep networks. The layer normalization operation normalizes the activation values of neurons on each mini-batch of data to ensure that the outputs of each layer have similar distribution characteristics. In this way, the addition and layer normalization operations jointly enhance the learning ability and stability of the diffusion model.

[0073] S202-2-C: Through the feed-forward network layer, perform a non-linear transformation on the first process image and activate it through an activation function to obtain the second process image.

[0074] In this embodiment, the feed-forward network layer is used to perform a non-linear transformation on the input image. It usually contains multiple neurons, and each neuron receives the input from the previous layer and generates an output through an activation function, thereby learning the complex patterns and relationships between the input and output.

[0075] S202-2-D: Through the second addition layer, add the first process image and the second process image and perform normalization processing to obtain the second feature image.

[0076] In this embodiment, the second addition layer is used to add the first process image and the second process image and perform normalization processing to obtain the second feature image. Among them, the functions and roles of the second addition layer and the first addition layer are similar, so they will not be elaborated here.

[0077] Optionally, for the multi-head attention layer, the above method further includes: For any dimension, determine the similarity score between the semantic feature and the initial image feature of the input image; the initial image feature is obtained by performing initial feature extraction on the input image through convolution; determine the cross-attention weight of the input image based on the similarity score; multiply the cross-attention weight by the semantic feature to obtain the image feature in this dimension.

[0078] As Figure 5 shown, first, denote the semantic features obtained by encoding and embedding the text description information through the CLIP model as K (Key) and V (Value), and denote the initial image features obtained by performing initial feature extraction on the input image through convolution as Q (Query).

[0079] In any dimension, the multi-head attention layer can obtain an attention weight matrix (i.e., cross-attention weight) by calculating the similarity scores between the query Q and the key K. This matrix characterizes which image features should pay more attention to the text information. Then, the cross-attention weight is multiplied by the value V in the text embedding to obtain a new feature representation, that is, the image features in this dimension are obtained.

[0080] Please refer to Figure 5 , Figure 5 is the logical schematic diagram of the Fourier transform in an embodiment of the present application. As Figure 5 shown, during the restoration process at low resolution, the above step S103 performs restoration on the target incomplete image at the second resolution based on the Fourier transform to obtain a second restored image, including: S301: Divide the number of channels of the target incomplete image into two parts to obtain a first channel feature and a second channel feature.

[0081] In this embodiment, for the target incomplete image at low resolution, the number of its channels will be divided into two parts for separate processing, denoted as the first channel feature and the second channel feature.

[0082] S302: Perform two-layer convolution operations and a first operation on the first channel feature to obtain the local feature of the target incomplete image; the first operation includes batch normalization and non-linear activation.

[0083] As Figure 5 shown, for the first channel feature, first use two convolutions with different sizes to extract features respectively, and then combine the extracted features to utilize different feature representations of the target incomplete image learned by multiple convolutions, and by combining the features extracted by these different convolutions, the network can understand the image more comprehensively, thereby improving the accuracy of image restoration.

[0084] Next, perform batch normalization operation and non-linear activation operation on the combined features to extract the local feature of the target incomplete image. Among them, through the batch normalization operation, the distribution of the combined feature values output previously can be pulled back to the standard normal distribution with a mean of 0 and a variance of 1, so that the input values of the subsequent non-linear transformation fall into a more sensitive region, avoiding the problem of gradient disappearance.

[0085] S303: Perform frequency domain transformation and a first operation on the second channel feature to obtain the global feature of the target incomplete image.

[0086] As Figure 5As shown, for the second-channel feature, first, convolution is used to extract features from the target incomplete image in the spatial domain. Meanwhile, through frequency-domain transformation, the target incomplete image is transformed to the frequency domain for feature extraction. Then, the extracted features are combined, and batch normalization and non-linear activation operations are performed on the combined features to extract the global features of the target incomplete image.

[0087] S304: Concatenate the local features and the global features along the channel dimension to obtain the second repaired image.

[0088] As Figure 5 shown, by concatenating the extracted local features and global features along the channel dimension, a feature representation containing local and global information is obtained, that is, the second repaired image.

[0089] In this embodiment, by dividing the number of channels of the target incomplete image into two equal parts, one part extracts local features through traditional convolution operations in the spatial domain, and the other part extracts global features in the frequency domain through frequency-domain transformation. Then, the two parts of features are combined and concatenated to improve the image repair effect at low resolution.

[0090] Further, the above step S303 performs frequency-domain transformation and the first operation on the second-channel feature to obtain the global feature of the target incomplete image, including: S303-1: Apply one convolution and non-linear activation to the second-channel feature to obtain the first spatial-domain feature.

[0091] As Figure 5 shown in the right half of [], during the process of extracting frequency-domain transformation features from the second-channel feature, first, one convolution and non-linear activation operations are performed on the input second-channel feature for preliminary feature extraction in the spatial domain to obtain the first spatial-domain feature.

[0092] S303-2: Convert the first spatial-domain feature from the spatial domain to the frequency domain through two-dimensional real fast Fourier transform, and apply one convolution and non-linear activation to extract the frequency-domain feature.

[0093] Next, the first spatial-domain feature is converted from the spatial domain to the frequency domain through two-dimensional real fast Fourier transform. And in the frequency domain, one more convolution and non-linear activation operations are performed to extract the features in the frequency domain to obtain the frequency-domain feature. Performing convolution in the frequency domain can reduce the number of required multiplication operations, thus accelerating the entire convolution process.

[0094] S303-3: Convert the frequency-domain feature from the frequency domain back to the spatial domain through inverse Fourier transform to obtain the second spatial-domain feature.

[0095] S303-4: Element-wise add the first spatial domain feature and the second spatial domain feature to obtain an added feature.

[0096] Then, inverse Fourier transform the frequency domain feature from the frequency domain back to the spatial domain to obtain the second spatial domain feature, add it element-wise to the first spatial domain feature that has not undergone frequency domain processing, and output it after one layer of convolution operation to fuse the features in the spatial domain and the frequency domain and improve the accuracy of image inpainting.

[0097] S303-5: Perform a first operation on the added feature to obtain the global feature of the target incomplete image.

[0098] As Figure 5 shown in the left half of, perform batch normalization operation and non-linear activation operation on the added feature obtained by frequency domain transformation to obtain the global feature of the target incomplete image.

[0099] In this embodiment, by combining frequency domain information in the convolution operation, the global and local features of the image can be effectively extracted, especially the capture of high-frequency details can be enhanced, and at the same time, the calculation process can be accelerated.

[0100] Optionally, the diffusion model is trained through the following steps: S401: Input the sample incomplete image and the sample text description information into the initial model to obtain a predicted completed image.

[0101] In this embodiment, first, the sample incomplete image for training and its corresponding real complete image and sample text description information need to be obtained. When training the diffusion model, please refer to the processing process shown in Figure 1 . By inputting the sample incomplete image and the sample text description information into the initial model and performing inpainting at high resolution, a sample first repaired image is obtained. At the same time, by downsampling the sample incomplete image and performing inpainting on the sample incomplete image at low resolution based on Fourier transform, a sample second repaired image is obtained. After upsampling the sample second repaired image, it is pixel-wise superimposed with the sample first repaired image to finally obtain the predicted completed image corresponding to the sample incomplete image.

[0102] S402: Determine the difference between the real complete image and the predicted completed image of the sample incomplete image on each feature layer, and perform weighted summation on the differences on each feature layer to obtain a perceptual distance measurement value.

[0103] The Learned Perceptual Image Patch Similarity (LPIPS) is a learning-based perceptual similarity metric designed to simulate the human visual system's perception of image differences. Its core idea is to use a pre-trained convolutional neural network to extract high-level feature representations of images. These high-level features can capture semantic information and structural details in images, thus being more in line with the human way of perceiving images.

[0104] In this embodiment, a pre-trained convolutional neural network is used to extract the high-level feature representations of the real complete image and the predicted completed image of the sample incomplete image respectively, and based on this, the differences between the two images on each layer of the feature map are calculated. These differences are weighted and summed as the final distance metric to obtain the first evaluation metric - the Learned Perceptual Image Patch Similarity.

[0105] The Learned Perceptual Image Patch Similarity can more accurately quantify the difference between the generated image and the real image, thus providing a more reliable basis for evaluating the completion quality. Specifically, the Learned Perceptual Image Patch Similarity is a non-negative value. The smaller the LPIPS value, the more similar the predicted completed image and the real complete image are perceptually, that is, the more natural the image restoration effect is and the higher the visual consistency is.

[0106] S403: Determine the cosine similarity between the predicted completed image and the sample text description information, and normalize the cosine similarity to obtain the contrastive learning score.

[0107] In this embodiment, the contrastive learning score (CLIP score) is used to measure the matching degree between the text description information and the predicted completed image. CLIP (Contrastive Language–Image Pretraining) is a model trained on a large amount of text and image data simultaneously, which can map text and images into the same vector space. The CLIPScore is obtained by calculating the cosine similarity between the text embedding of a given text prompt and the image embedding of the generated image.

[0108] Specifically, by using the image encoder and text encoder in the CLIP model to map the predicted completed image and the text description information into a high-dimensional feature space respectively, then calculate the cosine similarity between the image and the text description in the feature space. This similarity value reflects the semantic matching degree between the image and the text. Then, normalize the similarity value to make it fall within a fixed range (such as [0, 1]), that is, obtain the second evaluation metric - the contrastive learning score.

[0109] The closer the final CLIP Score value is to 1, the higher the degree of matching between the predicted and completed image and the text description information; conversely, the closer the CLIP Score value is to 0, the lower the degree of matching. CLIP Score can not only evaluate the consistency between the predicted and completed image and the text description information, but also reflect whether the predicted and completed image truly understands the input semantic information, ensuring that the generated result is not only visually realistic but also semantically accurate.

[0110] S404: Determine the loss function value based on the perceptual distance measurement value and the contrast learning score.

[0111] S405: Optimize and adjust the initial model based on the loss function value to obtain a diffusion model.

[0112] In this embodiment, the loss function value is determined based on the perceptual distance measurement value and the contrast learning score, and the initial model is optimized and adjusted based on the loss function value to evaluate the image completion ability of the initial model from aspects such as semantic feature utilization and visual effects, and the model parameters of the initial model are adjusted and optimized according to actual needs to minimize the loss function until the loss function value reaches the preset target and then stop training to obtain a trained diffusion model for image completion.

[0113] It should be noted that in this embodiment, in addition to the perceptual distance measurement value and the contrast learning score, the loss function may also include various indicators such as accuracy and recall rate to more comprehensively and objectively evaluate the quality of image completion.

[0114] Please refer to Figure 6 , based on the same inventive concept, the second aspect of the embodiments of the present application provides an image completion device based on a multi-scale diffusion model. The image completion device 600 of the multi-scale diffusion model includes: An image acquisition module 601, configured to acquire a target incomplete image and text description information; wherein, the target incomplete image includes mask information, and the mask information is used to distinguish the area to be repaired in the target incomplete image, and the text description information is used to describe the physical characteristics of the missing part of the target incomplete image; A first repair module 602, configured to repair the target incomplete image at a first resolution based on the text description information and a preset diffusion model to obtain a first repaired image; A second repair module 603, configured to downsample the target incomplete image and repair the target incomplete image at a second resolution based on Fourier transform to obtain a second repaired image; wherein, the second resolution is less than the first resolution; An upsampling module 604, configured to upsample the second repaired image to obtain a third repaired image; An overlay module 605 for pixel-by-pixel overlay of the first repaired image and the third repaired image to obtain a completed image corresponding to the target incomplete image.

[0115] Optionally, the pre-set diffusion model includes a plurality of consecutive encoding layers and a plurality of decoding layers connected to the encoding layers, and there are skip connections between symmetric encoding layers and decoding layers; The above first repair module 602 includes: A semantic feature extraction sub-module for encoding the text description information to obtain semantic features; An encoding sub-module for sequentially passing through a plurality of consecutive encoding layers to perform feature extraction on the target incomplete image based on the semantic features to obtain a feature image; A decoding sub-module for sequentially passing through a plurality of decoding layers to gradually denoise the feature image within T time steps to complete the target incomplete image to obtain a first repaired image.

[0116] Optionally, any encoding layer includes a self-attention layer, a cross-attention layer, and a fully-connected layer; the above encoding sub-module includes: A self-attention unit for, for any encoding layer in a plurality of consecutive encoding layers, performing feature extraction on the input image based on the self-attention weights through the self-attention layer to obtain a first feature image; A cross-attention unit for adaptively adjusting the cross-attention weights of the input image based on the semantic features through the cross-attention layer, and performing feature extraction on the input image based on the cross-attention weights to obtain a second feature image; A fully-connected fusion unit for fusing the first feature image and the second feature image through the fully-connected layer to obtain an output image; Wherein, the input image of the first encoding layer is the target incomplete image, the input image of the remaining encoding layers is the output image of the previous encoding layer, and the output image of the last encoding layer is the feature image.

[0117] Optionally, the cross-attention layer includes a multi-head attention layer, a first addition layer, a feed-forward network layer, and a second addition layer; the above cross-attention unit is specifically used for: Determining cross-attention weights from different dimensions based on the semantic features through the multi-head attention layer, and performing feature extraction on the input image based on the cross-attention weights to obtain image features in multiple dimensions; Fusing the image features in multiple dimensions through the first addition layer and performing normalization processing to obtain a first process image; Performing a non-linear transformation on the first process image through the feed-forward network layer and activating it through an activation function to obtain a second process image; Through the second addition layer, the first process image and the second process image are added and normalized to obtain the second feature image.

[0118] Optionally, the above-mentioned multi-head attention layer is specifically used for: For any dimension, determine the similarity score between the semantic feature and the initial image feature of the input image; the initial image feature is obtained by performing initial feature extraction on the input image through convolution; Determine the cross-attention weight of the input image based on the similarity score; Multiply the cross-attention weight by the semantic feature to obtain the image feature in the dimension.

[0119] Optionally, the above-mentioned second repair module 603 includes: A channel division sub-module for dividing the number of channels of the target incomplete image into two parts to obtain a first channel feature and a second channel feature; A local feature extraction sub-module for performing two-layer convolution operations and a first operation on the first channel feature to obtain the local feature of the target incomplete image; the first operation includes batch normalization and non-linear activation; A global feature extraction sub-module for performing a frequency domain transformation and a first operation on the second channel feature to obtain the global feature of the target incomplete image; A feature fusion sub-module for splicing the local feature and the global feature in the channel dimension to obtain the second repaired image.

[0120] Optionally, the above-mentioned global feature extraction sub-module includes: A convolution activation unit for passing the second channel feature through one convolution and non-linear activation to obtain a first spatial domain feature; A Fourier transform unit for converting the first spatial domain feature from the spatial domain to the frequency domain through a two-dimensional real fast Fourier transform, and extracting a frequency domain feature through one convolution and non-linear activation; An inverse Fourier transform unit for converting the frequency domain feature from the frequency domain back to the spatial domain through an inverse Fourier transform to obtain a second spatial domain feature; A spatial fusion unit for performing element-wise addition of the first spatial domain feature and the second spatial domain feature to obtain an addition feature; An activation unit for performing a first operation on the addition feature to obtain the global feature of the target incomplete image.

[0121] Optionally, the above-mentioned device further includes a model training module, and the model training module includes: A completion sub-module for inputting the sample incomplete image and the sample text description information into the initial model to obtain a predicted completed image; The first index calculation sub-module is used to determine the differences between the true complete image and the predicted completed image of the sample incomplete image on each feature layer, and perform weighted summation on the differences on each feature layer to obtain the perceptual distance measurement value; The second index calculation sub-module is used to determine the cosine similarity between the predicted completed image and the sample text description information, and perform normalization processing on the cosine similarity to obtain the contrast learning score; The loss function construction sub-module is used to determine the loss function value based on the perceptual distance measurement value and the contrast learning score; The model optimization sub-module is used to optimize and adjust the initial model based on the loss function value to obtain the diffusion model.

[0122] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the related parts, please refer to the partial description of the method embodiment.

[0123] In a third aspect, based on the same inventive concept, an embodiment of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the image completion method based on the multi-scale diffusion model proposed in the first aspect of the present application.

[0124] It should be noted that the specific implementation manner of the storage medium in the embodiment of the present application refers to the specific implementation manner of the image completion method based on the multi-scale diffusion model proposed in the first aspect of the present application described above, and will not be elaborated here.

[0125] In a fourth aspect, based on the same inventive concept, referring to Figure 7 , an embodiment of the present application provides an electronic device 700, including a processor 701 and a memory 702; the memory 702 stores machine-executable instructions that can be executed by the processor 701, and the processor 701 is used to execute the machine-executable instructions to implement the image completion method based on the multi-scale diffusion model proposed in the first aspect of the present application.

[0126] It should be noted that the specific implementation manner of the electronic device 700 in the embodiment of the present application refers to the specific implementation manner of the image completion method based on the multi-scale diffusion model proposed in the first aspect of the present application described above, and will not be elaborated here.

[0127] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0128] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, apparatuses, or computer program products. Therefore, the embodiments of the present application can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0130] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0132] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0133] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising said element.

[0134] The above has introduced in detail an image completion method, medium and device based on a multi-scale diffusion model provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An image completion method based on a multi-scale diffusion model, characterized in that, The method includes: Obtaining a target incomplete image and text description information; wherein, the target incomplete image contains mask information for distinguishing the area to be repaired in the target incomplete image, and the text description information is used to describe the physical characteristics of the missing part of the target incomplete image; Based on the text description information and a preset diffusion model, repairing the target incomplete image at a first resolution to obtain a first repaired image; Downsampling the target incomplete image and repairing the target incomplete image at a second resolution based on Fourier transform to obtain a second repaired image; wherein, the second resolution is less than the first resolution; Upsampling the second repaired image to obtain a third repaired image; Performing pixel-by-pixel superposition on the first repaired image and the third repaired image to obtain a completed image corresponding to the target incomplete image.

2. The method according to claim 1, characterized in that The preset diffusion model includes a plurality of consecutive encoding layers and a plurality of decoding layers connected to the encoding layers, and there are skip connections between the symmetric encoding layers and decoding layers; Based on the text description information and a preset diffusion model, repairing the target incomplete image at a first resolution to obtain a first repaired image, including: Encoding the text description information to obtain semantic features; Successively passing through a plurality of consecutive encoding layers to extract features from the target incomplete image based on the semantic features to obtain a feature image; Successively passing through a plurality of decoding layers to gradually denoise the feature image within T time steps to complete the target incomplete image to obtain the first repaired image.

3. The method according to claim 2, characterized in that, Any one of the encoding layers includes a self-attention layer, a cross-attention layer, and a fully connected layer; successively passing through a plurality of consecutive encoding layers to extract features from the target incomplete image based on the semantic features to obtain a feature image, including: For any one of the plurality of consecutive encoding layers, through the self-attention layer, extracting features from the input image based on the self-attention weights to obtain a first feature image; Through the cross-attention layer, adaptively adjusting the cross-attention weights of the input image based on the semantic features, and extracting features from the input image based on the cross-attention weights to obtain a second feature image; Through the fully connected layer, fusing the first feature image and the second feature image to obtain an output image; Wherein, the input image of the first encoding layer is the target incomplete image, the input image of the remaining encoding layers is the output image of the previous encoding layer, and the output image of the last encoding layer is the feature image.

4. The method according to claim 3, characterized in that, The cross-attention layer includes a multi-head attention layer, a first addition layer, a feed-forward network layer, and a second addition layer; adaptively adjusting the cross-attention weights of the input image based on the semantic features, and extracting features from the input image based on the cross-attention weights to obtain a second feature image, including: Through the multi-head attention layer, cross-attention weights are determined based on the semantic features from different dimensions, and based on the cross-attention weights, feature extraction is performed on the input image to obtain image features in multiple dimensions; Through the first addition layer, the image features in multiple dimensions are fused and normalized to obtain a first process image; Through the feed-forward network layer, a non-linear transformation is performed on the first process image and activated through an activation function to obtain a second process image; Through the second addition layer, the first process image and the second process image are added and normalized to obtain the second feature image.

5. The method according to claim 4, wherein The method further includes: For any dimension, determining a similarity score between the semantic feature and the initial image feature of the input image; the initial image feature is obtained by performing initial feature extraction on the input image through convolution; Determining the cross-attention weight of the input image based on the similarity score; Multiplying the cross-attention weight by the semantic feature to obtain the image feature in the dimension.

6. The method according to claim 1, characterized in that Based on the Fourier transform, performing repair on the target incomplete image at the second resolution to obtain a second repaired image, including: Dividing the number of channels of the target incomplete image into two parts to obtain a first channel feature and a second channel feature; Performing two-layer convolution operation and a first operation on the first channel feature to obtain the local feature of the target incomplete image; the first operation includes batch normalization and non-linear activation; Performing frequency domain transformation and the first operation on the second channel feature to obtain the global feature of the target incomplete image; Concatenating the local feature and the global feature in the channel dimension to obtain the second repaired image.

7. The method according to claim 6, characterized in that Performing frequency domain transformation and the first operation on the second channel feature to obtain the global feature of the target incomplete image, including: Passing the second channel feature through one convolution and non-linear activation to obtain a first spatial domain feature; Converting the first spatial domain feature from the spatial domain to the frequency domain through a two-dimensional real fast Fourier transform, and extracting a frequency domain feature through one convolution and non-linear activation; Converting the frequency domain feature from the frequency domain back to the spatial domain through an inverse Fourier transform to obtain a second spatial domain feature; Performing element-wise addition on the first spatial domain feature and the second spatial domain feature to obtain an addition feature; Performing the first operation on the addition feature to obtain the global feature of the target incomplete image.

8. The method according to claim 1, wherein The diffusion model is trained through the following steps: Inputting a sample incomplete image and sample text description information into an initial model to obtain a predicted completed image; Determining the difference between the true complete image of the sample incomplete image and the predicted completed image on each feature layer, and performing weighted summation on the differences on each feature layer to obtain a perceptual distance measurement value; Determining the cosine similarity between the predicted completed image and the sample text description information, and normalizing the cosine similarity to obtain a contrastive learning score; Based on the perceptual distance measurement value and the contrastive learning score, determining a loss function value; Based on the loss function value, the initial model is optimized and adjusted to obtain the diffusion model.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image completion method based on a multi-scale diffusion model according to any one of claims 1 to 8.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image completion method based on a multi-scale diffusion model according to any one of claims 1 to 8.

Citation Information

Cited By

  • Multi-modal data fusion processing method based on multi-stage fusion strategy

    CN120563991A

  • Image region missing content generation method and system based on three-dimensional partial differential equation

    CN121544733A