Image Inpainting Method and Related Device Generated Based on Diffusion Mode and Prior Features
By training a priori feature repair and enhancing the image model, using multi-network collaboration methods, the problem of high computational cost and inconsistent repair quality in the image repair task is solved, and efficient and high-quality image repair effect is achieved.
Patent Information
- Application Number
- CN202510570003.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-06
AI Technical Summary
Existing diffusion models are computationally cost-effective and have inconsistent repair quality in image repair tasks, making it difficult to recover missing or damaged parts of the image efficiently and with high quality.
By training a prior feature repair, multi-network collaboration of prior feature extraction network, image noise addition network, image denoising network, prior feature estimation network and mask repair network are used, and image repair is carried out by combining U-net's mask combination network and mask repair network.
It improves the accuracy and effect of image repair, can effectively and high-quality recovery of missing or damaged parts of the image, and improves the pertinence and efficiency of image repair.
Smart Images

Figure CN120088170B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image enhancement or restoration, and specifically to an image inpainting method and related device generated based on a diffusion mode and prior features. Background Art
[0002] Image inpainting is the process of filling in missing regions in a digital image. The goal of image inpainting is to maintain semantically reasonable and visually realistic content while being consistent with the rest of the image when filling in the missing regions. Therefore, image inpainting technology has received extensive attention.
[0003] Recently, diffusion models have performed well in image synthesis and inpainting tasks. However, directly applying diffusion models to image inpainting tasks faces two major challenges: on the one hand, diffusion models usually require a large number of iterations, and their high computational cost limits practical applications; on the other hand, the generation process based on the image synthesis paradigm may lead to inconsistencies with known regions, thus affecting the quality of image inpainting. Summary of the Invention
[0004] The purpose of the present invention is to provide an image inpainting method and related device generated based on a diffusion mode and prior features.
[0005] The technical solution of the present invention is as follows:
[0006] An image inpainting method generated based on a diffusion mode and prior features, including the following operations:
[0007] The image to be inpainted is processed by an image inpainting model to obtain an image inpainting result; the image inpainting model includes a training image noise addition network, a training image denoising network, a training prior feature estimation network, and a training mask inpainting network in the training prior feature inpainting enhanced image model;
[0008] The training prior feature inpainting enhanced image model is obtained by training the prior feature inpainting enhanced image model using a training data set; the prior feature inpainting enhanced image model includes: a training prior feature extraction network, an image noise addition network, an image denoising network, a prior feature estimation network, and a training mask inpainting network;
[0009] The training prior feature extraction network and the training mask inpainting network are obtained based on the training prior feature inpainting image model;
[0010] The training prior feature inpainting image model is obtained by training the prior feature inpainting image model using a training data set; the prior feature inpainting image model includes: a prior feature extraction network, a mask inpainting combined network based on U-net, and a mask inpainting network;
[0011] The training data set is formed by a number of real standard images and corresponding mask images.
[0012] In the process of training the prior feature restoration and enhancement image model, the real standard image and the corresponding mask image are processed by the prior feature extraction network and the image noise addition network to obtain a noise image; the mask image is processed by the prior feature estimation network to obtain a prior estimation feature map; the prior estimation feature map and the noise image are processed by the image denoising network and the training mask restoration network until the enhancement image loss value is less than the enhancement image loss threshold, and the training ends.
[0013] In the process of the prior feature extraction network, after splicing the real standard image and the corresponding mask image, a spliced image is obtained; the spatial information of the spliced image is rearranged to the preset channel dimension, and the rearranged features are subjected to a splicing operation to obtain a channel dimension enhanced map; the channel dimension enhanced map is successively subjected to convolution processing, LReLU activation function processing, residual block processing, average pooling processing, linear processing, and LReLU activation function processing to obtain a prior feature map for performing operations processed by the mask restoration combined network based on U-net.
[0014] The operations processed by the mask combination restoration network based on U-net are specifically as follows: the mask image and the prior feature map are processed by the mask restoration network to obtain a first mask restoration feature map; the first mask restoration feature map and the prior feature map are processed by the mask restoration network to obtain a second mask restoration feature map; the second mask restoration feature map and the prior feature map are processed by the mask restoration network to obtain a third mask restoration feature map; the third mask restoration feature map and the prior feature map are processed by the mask restoration network to obtain a fourth mask restoration feature map; after the fourth mask restoration feature map and the third mask restoration feature map are spliced, they are processed by the mask restoration network together with the prior feature map to obtain a first mask combination restoration feature map; after the first mask combination restoration feature map and the second mask restoration feature map are spliced, they are processed by the mask restoration network together with the prior feature map to obtain a second mask combination restoration feature map; after the second mask combination restoration feature map and the first mask restoration feature map are spliced, they are processed by the mask restoration network together with the prior feature map to obtain an initial restoration map for performing operations processed by the mask restoration network.
[0015] In the process of the mask restoration network, the prior feature map is injected into the initial restoration map to obtain an initial restoration map injected with the prior feature map; the initial restoration map injected with the prior feature map is subjected to global and local attention fusion processing and feature aggregation processing for several times to obtain an attention aggregation feature map; the attention aggregation feature map is subjected to convolution processing and then subjected to residual connection processing with the mask image to obtain a prior feature restoration map; the initial restoration map is the output of the mask combination restoration network based on U-net.
[0016] The operation of global and local attention fusion processing is as follows: The initial repair injection prior feature map undergoes global-local self-attention processing and global region-guided attention processing to obtain an attention fusion feature map, which is used to perform the operation of feature aggregation processing; The operation of global-local self-attention processing is as follows: The initial repair injection prior feature map is divided into horizontal windows and vertical windows, and after undergoing multi-head attention processing of non-overlapping sub-windows respectively, all the outputs are concatenated in the channel dimension to obtain an initial self-attention map; The initial attention map and the initial repair injection prior feature map undergo residual connection processing to obtain a self-attention feature map, which is used to perform the operation of global region-guided attention processing.
[0017] The operation of global region-guided attention processing is as follows: Inject the prior feature map into the self-attention feature map to obtain a self-attention injected prior feature map; Recursively use a single depthwise separable convolution to process the self-attention injected prior feature map to obtain a rough fusion map; The rough fusion map undergoes depthwise separable convolution and pixel-level convolution processing to obtain a representative feature map; Based on the key feature matrix and value feature matrix of the representative feature map, an attention matrix is obtained; Based on the query matrix of the attention matrix and the self-attention injected prior feature map, attention cross-processing is performed to obtain a cross-attention feature map; After the cross-attention feature map undergoes reshaping processing, it undergoes residual connection processing with the self-attention feature map to obtain an attention fusion feature map.
[0018] An image restoration system based on diffusion pattern and prior feature generation, which is used to implement the above-mentioned image restoration method based on diffusion pattern and prior feature generation, includes:
[0019] A training prior feature restoration image model generation module, which is used to utilize a training data set to train a prior feature restoration image model to obtain a trained prior feature restoration image model; The prior feature restoration image model includes: a prior feature extraction network, a mask restoration combined network based on U-net, and a mask restoration network; The training data set is formed by a number of real standard images and corresponding mask images;
[0020] A training prior feature restoration enhanced image model generation module, which is used to utilize a training data set to train a prior feature restoration enhanced image model to obtain a trained prior feature restoration enhanced image model; The prior feature restoration enhanced image model includes: a trained prior feature extraction network, an image noise addition network, an image denoising network, a prior feature estimation network, and a trained mask restoration network; The trained prior feature extraction network and the trained mask restoration network are obtained based on the trained prior feature restoration image model;
[0021] The image restoration result generation module is used to process the image to be restored through the image restoration model to obtain the image restoration result; the image restoration model includes a training image denoising network, a training image denoising network, a training prior feature estimation network and a training mask restoration network in the training prior feature restoration and enhancement image model.
[0022] An image restoration device based on diffusion pattern and prior feature generation comprises a processor and a memory, wherein the processor implements the above-mentioned image restoration method based on diffusion pattern and prior feature generation when executing a computer program stored in the memory.
[0023] A computer-readable storage medium is used to store a computer program, wherein when the computer program is executed by a processor, the image restoration method based on diffusion pattern and prior feature generation is implemented.
[0024] The beneficial effects of the present invention are:
[0025] The present invention provides an image restoration method based on diffusion pattern and prior feature generation. First, the prior feature restoration image model is trained using real data. The prior feature extraction network can extract effective features. The mask combination restoration network and the mask restoration network based on U-net can use the mask information to repair the image in a targeted manner, thereby improving the accuracy and effect of image restoration, so as to better restore the missing or damaged parts of the image. Then, the prior feature restoration enhanced image model is trained using the training data set. The prior feature restoration enhanced image model includes a trained prior feature extraction network, an image denoising network, an image denoising network, a prior feature estimation network and a trained mask restoration network. The image restoration details can be refined and the image restoration effect can be enhanced. Through multi-network collaboration, with the help of the prior feature correlation network and the mask restoration network, as well as the image denoising and denoising operations, the image restoration can be more accurately restored. The image is processed accurately to improve the quality of image restoration. Finally, a training image denoising network, a training image denoising network, a training prior feature estimation network and a training mask restoration network are obtained from the training prior feature restoration and enhancement image model to form an image restoration model. In the process of using this model to process the image to be restored, the training denoising and denoising networks are used to process and optimize the image, the training prior feature estimation network provides prior information, and the training mask restoration network is used to restore specific areas, thereby completing image restoration efficiently and with high quality. When applied in the field of image restoration, it can improve the pertinence, effect and efficiency of image restoration. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] By reading the detailed description of the preferred embodiment below, the scheme and advantages of the present application will become clear to those skilled in the art. The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.
[0027] In the attached picture:
[0028] Figure 1 In the embodiment, it is the image restoration effect diagram of the method of this embodiment. Detailed implementation manners
[0029] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.
[0030] This embodiment provides an image restoration method based on diffusion mode and prior features generation, including the following operations:
[0031] The image to be restored is processed by an image restoration model to obtain an image restoration result; the image restoration model includes a training image noise addition network, a training image denoising network, a training prior feature estimation network, and a training mask restoration network in the training prior feature restoration enhanced image model;
[0032] The training prior feature restoration enhanced image model is obtained by training the prior feature restoration enhanced image model using a training data set; the prior feature restoration enhanced image model includes: a training prior feature extraction network, an image noise addition network, an image denoising network, a prior feature estimation network, and a training mask restoration network;
[0033] The training prior feature extraction network and the training mask restoration network are obtained based on the training prior feature restoration image model;
[0034] The training prior feature restoration image model is obtained by training the prior feature restoration image model using a training data set; the prior feature restoration image model includes: a prior feature extraction network, a mask restoration combined network based on U-net, and a mask restoration network;
[0035] The training data set is formed by a number of real standard images and corresponding mask images.
[0036] The specific operation process of a specific embodiment is as follows.
[0037] S1. Obtain a number of real standard images and corresponding mask images to form a training data set; use the training data set to train a prior feature restoration image model to obtain a training prior feature restoration image model; the prior feature restoration image model includes: a prior feature extraction network, a mask combined restoration network based on U-net, and a mask restoration network.
[0038] Using real data to train the prior feature restoration image model, the prior feature extraction network can extract effective features, and the mask combined restoration network and the mask restoration network based on U-net can use the mask information to specifically repair the image, improving the accuracy and effect of image restoration to better restore the missing or damaged parts of the image.
[0039] Obtain a number of real standard images and corresponding mask images. Each real standard image and its corresponding mask image form a set of image sets, and all the image sets form a training dataset, which is used as the dataset for training the neural network.
[0040] Among them, the mask image is obtained by merging the real standard image and the corresponding mask map; the mask map is composed of pixels with values of 1 and 0, where 0 represents the masked area (usually the part that needs to be processed or ignored), and 1 represents the non-masked area (the valid area or the original image).
[0041] Use the training dataset to train the prior feature image restoration model to obtain the trained prior feature image restoration model. The prior feature image restoration model includes: a prior feature extraction network, a mask combination restoration network based on U-net, and a mask restoration network.
[0042] During the process of training the prior feature image restoration model, the real standard image and the corresponding mask image are processed by the prior feature extraction network, which completely retains the most important detail parts in the image and can save computing resources to obtain the prior feature map; the prior feature map and the mask image are processed by the mask combination restoration network based on U-net. The prior feature map is used as a guide to guide the fusion of local and global information, enhance the model's ability to perceive the missing area, and perform image restoration on the mask image to obtain the initial restored image; the initial restored image and the prior feature map are processed by the mask restoration network until the first training loss value is less than the first image loss threshold, and the training ends to obtain the prior feature restored image.
[0043] The loss function during training is as follows: , is the first training loss value for training the prior feature image restoration model, is the real standard image, is the prior feature restored image, is the normalization process.
[0044] The specific operations of the above prior feature extraction network processing are as follows: After splicing the real standard image and the corresponding mask image, a spliced image is obtained; the spatial information of the spliced image is rearranged to the preset channel dimension to enhance the channel dimension height, and the rearranged features are spliced to obtain a channel dimension enhanced image; the channel dimension enhanced image is successively processed by convolution, LReLU activation function, residual block, average pooling, linear processing, and LReLU activation function to obtain the prior feature map, which is used to perform the operations processed by the mask restoration combination network based on U-net.
[0045] The operations processed by the above-mentioned U-net-based mask combination repair network are specifically as follows: The mask image and the prior feature map are processed by the mask repair network to obtain the first mask repair feature map; the first mask repair feature map and the prior feature map are processed by the mask repair network to obtain the second mask repair feature map; the second mask repair feature map and the prior feature map are processed by the mask repair network to obtain the third mask repair feature map; the third mask repair feature map and the prior feature map are processed by the mask repair network to obtain the fourth mask repair feature map; after the fourth mask repair feature map and the third mask repair feature map are spliced, they are processed by the mask repair network together with the prior feature map to obtain the first mask combination repair feature map; after the first mask combination repair feature map and the second mask repair feature map are spliced, they are processed by the mask repair network together with the prior feature map to obtain the second mask combination repair feature map; after the second mask combination repair feature map and the first mask repair feature map are spliced, they are processed by the mask repair network together with the prior feature map to obtain the initial repair map.
[0046] Taking the processing of the initial repair map and the prior feature map by the mask repair network as an example, the operations of the mask repair network are specifically as follows.
[0047] Step 1: Inject the prior feature map into the initial repair map to obtain the initial repair map injected with the prior feature map. The operation of obtaining the initial repair map injected with the prior feature map can be achieved through the following formula: , is the initial repair map injected with the prior feature map, is the prior feature map, is the mask image, or the mask repair feature map, or the mask combination repair splicing feature map, ⊙ is element-wise multiplication, is layer normalization processing, , are the first linear weight and the second linear weight respectively.
[0048] Step 2: The initial repair map injected with the prior feature map undergoes several global and local attention fusion processes and feature aggregation processes, which can not only capture local features but also focus on global context information, thereby promoting more information to flow into the deep network to obtain the attention aggregation feature map.
[0049] For another example, taking the first global and local attention fusion process and feature aggregation process as an example, the specific operation details are as follows.
[0050] The operation of the global and local attention fusion process is: The initial repair map injected with the prior feature map undergoes global-local self-attention processing and global region-guided attention processing to obtain the attention fusion feature map, which is used to perform the operation of the feature aggregation process.
[0051] Among them, the operation of global-local self-attention processing is as follows: The initial repair injection prior feature map is divided into horizontal windows and vertical windows. After being processed by multi-head attention of non-overlapping sub-windows respectively, all the outputs are concatenated in the channel dimension to obtain the initial self-attention map; The initial attention map and the initial repair injection prior feature map are processed by residual connection to obtain the self-attention feature map, which is used to perform the operation of global region-guided attention processing.
[0052] The operation of global region-guided attention processing is as follows.
[0053] Step a: Inject the prior feature map into the self-attention feature map (the injection process is similar to the operation of obtaining the initial repair injection prior feature map above), to obtain the self-attention injection prior feature map; Recursively use a single depthwise separable convolution to process the self-attention injection prior feature map to compress the feature space size and obtain a rough fusion map.
[0054] Among them, the number of recursions is obtained based on the size of the self-attention injection prior feature map and the convolution stride of the depthwise separable convolution processing; The number of recursions is , is the convolution stride of the depthwise separable convolution, is a constant, is the height of the self-attention injection prior feature map.
[0055] Step b: The rough fusion map is processed by a depthwise separable convolution with a convolution stride of and a pixel-wise convolution with a convolution stride of (which can be achieved by pointwise convolution) to perform feature refinement and channel scaling, and obtain a representative feature map.
[0056] Step c: Based on the key feature matrix and value feature matrix of the representative feature map, obtain an attention matrix; Based on the attention matrix and the query matrix of the self-attention injection prior feature map, perform attention cross-processing to obtain a cross-attention feature map; After the cross-attention feature map is reshaped to reach a preset size, it is processed by residual connection with the self-attention feature map to obtain an attention fusion feature map, which is used to perform the operation of feature aggregation processing.
[0057] The operation of attention cross-processing can be realized by the following formula:
[0058] ,
[0059] ,
[0060] is the attention matrix, and the sum of the elements in each row of the attention matrix is 1, 、 They are respectively the key feature matrix and the transpose of the value feature matrix of the representative feature map. is the adjustment factor. is the query matrix for injecting the prior feature map into the self-attention. is the fusion projection matrix. is the cross-attention feature map.
[0061] The operations of the feature aggregation process are as follows: the attention fusion feature map undergoes channel feature aggregation processing to obtain the attention channel aggregation map; the attention channel aggregation map undergoes spatial neighborhood pixel feature aggregation processing to obtain the attention pixel aggregation map; the attention fusion feature map undergoes gating mechanism processing to enhance information encoding and obtain the gated aggregation map; after the gated aggregation map and the attention pixel aggregation map are multiplied element-wise, they are subjected to residual connection with the attention fusion feature map to obtain the attention aggregation feature map.
[0062] Step 3: After the attention aggregation feature map undergoes convolution processing with a convolution stride of it is subjected to residual connection processing with the initial repaired image to obtain the prior feature repaired image.
[0063] S2: Using the training dataset, train the prior feature repaired and enhanced image model to obtain the trained prior feature repaired and enhanced image model; the prior feature repaired and enhanced image model includes: the trained prior feature extraction network, the image noise addition network, the image denoising network, the prior feature estimation network, and the trained mask repair network; the trained prior feature extraction network and the trained mask repair network are obtained based on the trained prior feature repaired image model.
[0064] Using the training dataset to train the prior feature repaired and enhanced image model, which includes the trained prior feature extraction network, the image noise addition network, the image denoising network, the prior feature estimation network, and the trained mask repair network, can refine the details of image repair, enhance the image repair effect, and through the collaboration of multiple networks, with the help of the prior feature-related network and the mask repair network, as well as operations such as image noise addition and denoising, more accurately process the image and improve the quality of image repair.
[0065] Using the training dataset, train the prior feature repaired and enhanced image model to obtain the trained prior feature repaired and enhanced image model. The prior feature repaired and enhanced image model includes: the trained prior feature extraction network, the image noise addition network, the image denoising network, the prior feature estimation network, and the trained mask repair network.
[0066] Among them, the trained prior feature extraction network and the trained mask repair network are obtained based on the trained prior feature repaired image model in S1.
[0067] During the process of training the prior feature restoration and enhancement image model, the real standard image and the corresponding mask image are processed by the prior feature extraction network and the image noise addition network to obtain a noisy image; the mask image is processed by the prior feature estimation network to obtain a prior estimation feature map; the prior estimation feature map and the noisy image are processed by the image denoising network and the training mask restoration network until the enhancement image loss value is less than the enhancement image loss threshold, at which point the training ends and a prior feature restoration and enhancement map is obtained.
[0068] The enhancement image loss value is obtained through the following formula: , is the enhancement image loss value, is the mask image, is the prior feature restoration and enhancement map.
[0069] The operation to obtain the above prior estimation feature map is specifically as follows: The mask image is subjected to spatial information rearrangement to the preset channel dimension for processing, the channel dimension height is enhanced, and the rearranged features are subjected to a splicing operation to obtain a mask channel dimension enhanced map; the mask channel dimension enhanced map is successively subjected to convolution processing, LReLU activation function processing, residual block processing, average pooling processing, linear processing, and LReLU activation function processing to obtain the prior estimation feature map.
[0070] During the processing of the image noise addition network, the following is followed:
[0071] ,
[0072] ,
[0073] is the conditional probability distribution of the entire noise addition process from the initial latent variable to , is the original data sample, is the latent variable at the t th step of the noise addition process, is the latent variable at the t -1th step of the noise addition process, is the conditional Gaussian distribution of the noise addition process, is the noise variance, is the covariance structure of the noise.
[0074] During the processing of the image denoising network, the following is followed:
[0075] ,
[0076] ,
[0077] is thet The conditional probability distribution of the step is the cumulative noise attenuation factor is the latent variable at the t th step in the denoising process is the original data sample is the t latent variable at the - 1th step is the Gaussian noise mean ( ) is T the noise attenuation coefficient at time is the amount of noise removed
[0078] S3. Obtain the training image noise addition network, training image denoising network, training prior feature estimation network, and training mask repair network in the training prior feature repair and enhancement image model to form an image repair model; the image to be repaired is processed by the image repair model to obtain an image repair result
[0079] Obtain the training image noise addition network, training image denoising network, training prior feature estimation network, and training mask repair network from the training prior feature repair and enhancement image model to form an image repair model. In the process of processing the image to be repaired with this model, through the collaborative action of each network, the noise addition and denoising networks process and optimize the image, the prior feature estimation network provides prior information, and the mask repair network repairs specific regions, so as to complete image repair efficiently and with high quality
[0080] Obtain the training image noise addition network, training image denoising network, training prior feature estimation network, and training mask repair network in the training prior feature repair and enhancement image model obtained in S2 to form an image repair model; the image to be repaired is processed by the image repair model to obtain an image repair result
[0081] In the process of processing the image to be repaired by the image repair model, the mask image of the image to be repaired is processed by the training prior feature estimation network to obtain the prior feature to be processed; the prior feature to be processed and the noise feature map in the training image noise addition network (the training image noise addition network can directly extract the noise feature map according to needs after training) are processed by the training image denoising network and the training mask repair network to obtain the image repair result, see the image repair effect diagram in Figure 1
[0082] This embodiment also provides an image repair system based on diffusion mode and prior feature generation for implementing the above - mentioned image repair method based on diffusion mode and prior feature generation, including:
[0083] The training prior feature inpainting image model generation module is used to train a prior feature inpainting image model by using a training data set, and obtain the trained prior feature inpainting image model; the prior feature inpainting image model includes: a prior feature extraction network, a mask inpainting combination network based on U-net, and a mask inpainting network; the training data set is formed by a plurality of real standard images and corresponding mask images;
[0084] The training prior feature inpainting enhanced image model generation module is used to train a prior feature inpainting enhanced image model by using a training data set, and obtain the trained prior feature inpainting enhanced image model; the prior feature inpainting enhanced image model includes: a trained prior feature extraction network, an image noise addition network, an image denoising network, a prior feature estimation network, and a trained mask inpainting network; the trained prior feature extraction network and the trained mask inpainting network are obtained based on the trained prior feature inpainting image model;
[0085] The image inpainting result generation module is used to process the image to be inpainted through an image inpainting model to obtain an image inpainting result; the image inpainting model includes the trained image noise addition network, the trained image denoising network, the trained prior feature estimation network, and the trained mask inpainting network in the trained prior feature inpainting enhanced image model.
[0086] This embodiment also provides an image inpainting device based on diffusion mode and prior feature generation, including a processor and a memory. When the processor executes the computer program stored in the memory, the above-mentioned image inpainting method based on diffusion mode and prior feature generation is implemented.
[0087] This embodiment also provides a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, the above-mentioned image inpainting method based on diffusion mode and prior feature generation is implemented.
[0088] The present embodiment provides an image restoration method based on diffusion pattern and prior feature generation. First, the prior feature restoration image model is trained using real data. The prior feature extraction network can extract effective features. The mask combination restoration network and the mask restoration network based on U-net can use the mask information to repair the image in a targeted manner, thereby improving the accuracy and effect of image restoration, so as to better restore the missing or damaged parts of the image. Then, the prior feature restoration enhanced image model is trained using the training data set. The prior feature restoration enhanced image model includes a training prior feature extraction network, an image denoising network, an image denoising network, a prior feature estimation network and a training mask restoration network. The image restoration details can be refined and the image restoration effect can be enhanced. Network collaboration, with the help of prior feature correlation network and mask restoration network, as well as image denoising and denoising operations, can process images more accurately and improve the quality of image restoration; finally, the trained image denoising network, trained image denoising network, trained prior feature estimation network and trained mask restoration network are obtained from the trained prior feature restoration enhanced image model to form an image restoration model. In the process of using this model to process the image to be restored, the trained denoising and denoising networks are used to process and optimize the image, the trained prior feature estimation network provides prior information, and the trained mask restoration network is used to repair specific areas, thereby completing image restoration efficiently and with high quality; applied in the field of image restoration, it can improve the targetedness, image restoration effect and image restoration efficiency of image restoration.
Claims
1. An image inpainting method based on diffusion mode and prior feature generation, characterized in that Including the following operations: The image to be repaired is processed by an image repair model to obtain an image repair result; the image repair model includes a training image noise adding network, a training image denoising network, a training prior feature estimation network, and a training mask repair network in the training prior feature repair and enhancement image model; The training prior feature repair and enhancement image model is obtained by training the training prior feature repair and enhancement image model using a training data set; The prior feature repair and enhancement image model includes: a training prior feature extraction network, an image noise adding network, an image denoising network, a prior feature estimation network, and a training mask repair network; The training prior feature extraction network and the training mask repair network are obtained based on the training prior feature repair image model; The training prior feature repair image model is obtained by training the training prior feature repair image model using a training data set; the prior feature repair image model includes: a prior feature extraction network, a mask repair combined network based on U-net, and a mask repair network; During the processing of the prior feature extraction network, the real standard image and the corresponding mask image are spliced to obtain a spliced image; the spatial information of the spliced image is rearranged to a preset channel dimension, and the rearranged features are spliced to obtain a channel dimension enhanced map; the channel dimension enhanced map is successively processed by convolution, LReLU activation function, residual block, average pooling, linear processing, and LReLU activation function to obtain a prior feature map for performing operations processed by the mask repair combined network based on U-net; During the processing of the mask repair network, the prior feature map is injected into the initial repair map to obtain an initial repair injected prior feature map; the initial repair injected prior feature map is subjected to several global and local attention fusion processes and feature aggregation processes to obtain an attention aggregation feature map; the attention aggregation feature map is processed by convolution and then connected to the mask image through a residual connection to obtain a prior feature repair map; the initial repair map is the output of the mask combined repair network based on U-net; The training data set is formed by a number of real standard images and corresponding mask images.
2. The image inpainting method based on diffusion mode and prior features according to claim 1, wherein, During the processing of training the prior feature repair and enhancement image model, the real standard image and the corresponding mask image are processed by the training prior feature extraction network and the image noise adding network to obtain a noise image; The mask image is processed by the prior feature estimation network to obtain a prior estimation feature map; the prior estimation feature map and the noise image are processed by the image denoising network and the training mask repair network until the enhanced image loss value is less than the enhanced image loss threshold, and the training ends.
3. The image inpainting method based on diffusion pattern and prior features according to claim 1, wherein The operations processed by the mask combined repair network based on U-net are specifically as follows: The mask image and the prior feature map are processed by the mask repair network to obtain a first mask repair feature map; The first mask repair feature map and the prior feature map are processed by the mask repair network to obtain a second mask repair feature map; The second mask repair feature map and the prior feature map are processed by the mask repair network to obtain a third mask repair feature map; The third mask repair feature map and the prior feature map are processed by the mask repair network to obtain a fourth mask repair feature map; After the fourth mask repair feature map and the third mask repair feature map are concatenated, they are processed by the mask repair network together with the prior feature map to obtain the first mask combined repair feature map; After the first mask combined repair feature map and the second mask repair feature map are concatenated, they are processed by the mask repair network together with the prior feature map to obtain the second mask combined repair feature map; After the second mask combined repair feature map and the first mask repair feature map are concatenated, they are processed by the mask repair network together with the prior feature map to obtain the initial repair map, which is used to perform the operation of the mask repair network processing.
4. The image restoration method based on diffusion pattern and prior features according to claim 1, characterized in that The operation of global and local attention fusion processing is as follows: The initial repair injected prior feature map is processed by global-local self-attention and global region-guided attention to obtain the attention fusion feature map, which is used to perform the operation of feature aggregation processing; The operation of global-local self-attention processing is as follows: The initial repair injected prior feature map is divided into horizontal windows and vertical windows, and after being processed by the multi-head attention of non-overlapping sub-windows respectively, all the outputs are concatenated in the channel dimension to obtain the initial self-attention map; The initial attention map and the initial repair injected prior feature map are processed by residual connection to obtain the self-attention feature map, which is used to perform the operation of global region-guided attention processing.
5. The image restoration method based on diffusion patterns and prior features according to claim 4, wherein The operation of global region-guided attention processing is as follows: Inject the prior feature map into the self-attention feature map to obtain the self-attention injected prior feature map; Recursively use a single depthwise separable convolution to process the self-attention injected prior feature map to obtain a rough fusion map; The rough fusion map is processed by depthwise separable convolution and pixel-level convolution to obtain a representative feature map; Based on the key feature matrix and value feature matrix of the representative feature map, an attention matrix is obtained; Based on the attention matrix and the query matrix of the self-attention injected prior feature map, attention cross-processing is performed to obtain the cross-attention feature map; After the cross-attention feature map is reshaped, it is processed by residual connection with the self-attention feature map to obtain the attention fusion feature map.
6. An image inpainting system generated based on a diffusion pattern and prior features, for implementing the image inpainting method generated based on a diffusion pattern and prior features according to claim 1, characterized in that Including: Train the prior feature repair image model generation module to use the training data set to train the prior feature repair image model to obtain the trained prior feature repair image model; The prior feature repair image model includes: a prior feature extraction network, a mask repair combination network based on U-net, and a mask repair network; The training data set is formed by a number of real standard images and corresponding mask images; Train the prior feature repair enhanced image model generation module to use the training data set to train the prior feature repair enhanced image model to obtain the trained prior feature repair enhanced image model; The prior feature repair enhanced image model includes: a trained prior feature extraction network, an image noise addition network, an image denoising network, a prior feature estimation network, and a trained mask repair network; The trained prior feature extraction network and the trained mask repair network are obtained based on the trained prior feature repair image model; An image restoration result generation module is used to process the image to be restored through an image restoration model to obtain an image restoration result; the image restoration model includes a training image noise addition network, a training image denoising network, a training prior feature estimation network, and a training mask restoration network in the training prior feature restoration and enhancement image model.
7. An image inpainting device generated based on a diffusion pattern and prior features, characterized in that, It includes a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the image restoration method based on diffusion mode and prior features described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It is used to store a computer program. Among them, when the computer program is executed by a processor, it implements the image restoration method based on diffusion mode and prior features described in any one of claims 1-5.
Citation Information
Patent Citations
Real world image super-resolution method based on stable diffusion
CN118918009A
Mural image restoration method and system based on multi-stage zero-increment connection
CN119090779A