An Image Inpainting Method Based on Cross-Image Context Memory

By constructing a cross-image context memory bank, using cross-image feature information for image repair, the problem of insufficient information in a single image is solved and a better repair effect is achieved.

CN115311153BActive Publication Date: 2025-07-01TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210330146.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-07-01
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Existing image repair methods mainly rely on the context information of a single image, resulting in poor repair results in the absence of appearance and semantic information.

Method used

Cross-image context memory (CICM) Bank is constructed, feature maps of different images are stored and divided according to semantic categories, image repair is used to use cross-image context information, and missing areas are optimized through convolutional neural networks and feature similarity.

Benefits of technology

It provides richer cross-image context information, improving image repair effects, especially in rich semantic categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311153B_ABST
    Figure CN115311153B_ABST
Patent Text Reader

Abstract

The present invention relates to an image inpainting method based on cross-image context memory, comprising the following steps: establishing a CICM Bank for storing feature maps obtained from cross-images, and dividing the feature maps into different feature sets according to semantic categories; selecting a convolutional neural network for extracting image features; obtaining a rough inpainting image of a locally missing image; optimizing the rough inpainting image to obtain an optimized image Ir; updating the CICM Bank by using a ground-truth image Ig of the same size as the rough inpainting image Ic and the corresponding ground-truth semantic segmentation result Sg; and obtaining a final inpainting image I according to the rough inpainting image Ic and the optimized image Ir.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and computer vision, relates to image restoration technology, and specifically relates to an image restoration method based on cross-image context memory. Background Art

[0002] Image restoration aims to restore the missing parts of an image based on the information of the complete parts in the current image. It refers to the process of filling in the missing data in a specified area of the visual input. Extensive research on image restoration has been carried out in the applications of digital image restoration, image coding, and transmission.

[0003] Currently, most image restoration methods only use the context information in a single image. Generally, the restoration network learns the appearance representation from the image to capture the visual context information of the image region. Based on the visual context, the relevant information of the undamaged region can be propagated to the damaged region to infer its content. Liu et al. [1] and Yu et al. [2] respectively proposed partial convolution and gated convolution methods to repair the irregular missing parts in the image. Liu et al. [3] repaired the texture and structure of the image at the feature level. In addition to using the appearance representation, some of the latest methods also use a segmentation network to predict the object categories coexisting in the image to learn the high-level semantic representation. Song et al. [4] proposed a two-stage model. In the first stage, the complete semantic segmentation result is calculated, and in the second stage, the semantic segmentation result is used for image restoration. Liao et al. [5] proposed a joint learning model that uses the segmentation results of each layer in the computational decoder to constrain the final restoration result. However, these methods are all limited to a single picture. At the same time, a single incomplete image lacks appearance and semantic information, and the undamaged regions in the image may only contain very little useful information. Under such conditions, the results of these restoration methods are not satisfactory.

[0004] References:

[0005] [1]Liu G,Reda F A,Shih K J,et al.Image inpainting for irregular holesusing partial convolutions[C] / / Proceedings of the European conference oncomputer vision(ECCV).2018:85-100.

[0006] [2]Yu J,Lin Z,Yang J,et al.Free-form image inpainting with gated convolution[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2019:4471-4480.

[0007] [3]Liu H,Jiang B,Song Y,et al.Rethinking image inpainting via a mutual encoder-decoder with feature equalizations[C] / / European Conference on Computer Vision.Springer,Cham,2020:725-741.

[0008] [4]Song Y,Yang C,Shen Y,et al.Spg-net:Segmentation prediction and guidance network for image inpainting[J].arXiv preprint arXiv:1805.03356,2018.

[0009] [5]Liao L,Xiao J,Wang Z,et al.Image inpainting guided by coherence priors of semantics and textures[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2021:6539-6548. Summary of the Invention

[0010] The present invention proposes an image inpainting method based on Cross-Image Context Memory (CICM), which utilizes cross-image context information to restore missing regions in images. During the training process of the inpainting model, a CICM Bank is constructed to store the appearance and semantic representations learned from the complete contents of other image regions and related object categories. In the inpainting task, a feature set of the same category is selected from the CICM Bank to restore the damaged regions of the image. The present invention is implemented through the following technical solutions:

[0011] An image inpainting method based on cross-image context memory, comprising the following steps:

[0012] In the first step, a CICM Bank is established to store the feature maps obtained from cross-images. The feature maps are divided into different feature sets according to semantic categories, denoted as B = {F d |d = 1,…,D}, where F d = {f d,n ∈R C |0 ≤ n ≤ N} represents all the feature sets belonging to the d-th category, and f d,n represents the n-th feature in the feature set of the d-th category. Here, D represents the total number of all feature sets, and N represents the total number of features in each feature set of the d-th category;

[0013] In the second step, a convolutional neural network for extracting image features is selected;

[0014] In the third step, a rough inpainting image of the locally missing image is obtained through the convolutional neural network selected in the second step. The method is as follows: The locally missing image I m ∈R H×W×3 and its mask M ∈ R H×W are input into the convolutional neural network selected in the second step to extract features. Here, H and W respectively represent the height and width of the image and the mask, and a rough inpainting result is output, including a rough inpainting image I c for making a rough inpainting of the locally missing image and a rough semantic segmentation result S c ;

[0015] In the fourth step, the rough inpainting image is optimized to obtain an optimized image I r , and the method is as follows:

[0016] (1) The rough inpainting image I c and the rough semantic segmentation result S c are divided into an unmissing part image I u , a missing part image I t and a missing part semantic segmentation result S t according to the mask M;

[0017] (2) The missing part image I t is non-overlappingly sliced into a set of regional images , where J represents the total number of regional images in the set of regional images;

[0018] (3) The missing part semantic segmentation result S t is resized according to the regional images in the set of regional images . After resizing, each pixel of the missing part semantic segmentation result S t corresponds to an element in the set of regional images;

[0019] (4) According to the correspondence between the semantic segmentation result of the resized missing part and the regional image set of, assign a corresponding unique semantic category to each element in the regional image set ;

[0020] (5) According to the semantic category of the j-th regional image , select the feature set F with the same category from the CICM Bank h ={f h,n |n = 1,…,N};

[0021] (6) Through convolution operation, calculate the regional image feature of the regional image

[0022] (7) Calculate the feature similarity set between the regional image feature and the feature set F with the same category h ={f h,n |n = 1,…,N} where N represents the total number of feature similarities;

[0023] (8) According to the feature similarity set , calculate the optimized regional feature by weighted calculation of the feature set F h The calculation method is as follows

[0024]

[0025] (9) Pass the optimized regional feature through convolution operation to calculate the optimized regional image

[0026] (10) After passing all J elements in the regional image set through steps (6)-(7), obtain the optimized regional image set composed of J optimized regional images Stitch the optimized regional image set according to the original position to obtain the optimized image I of the missing part q , and then combine the optimized image I of the missing part q with the undamaged part image I u to obtain the optimized image I r ;

[0027] Fifthly, use the ground truth image I c with the same size as the roughly repaired image I g and the corresponding ground truth semantic segmentation result S g ​​Update to update CICM Bank;

[0028] Step 6: Based on the roughly repaired image I c and the optimized image I r , obtain the final repaired image I.

[0029] Furthermore, in (1) of the fourth step, the area of the missing part of the image I t and the area of the missing part of the semantic segmentation result S t where the retained area is larger than the corresponding white area in the mask M, the retained area of the non-missing part of the image I u will be reduced accordingly; the retained area of the non-missing part of the image I u and the retained area of the missing part of the image I t should ensure that the roughly repaired image I c can be completely pieced together.

[0030] Furthermore, the calculation method of (7) in the fourth step is as follows:

[0031] Step 1: Calculate the image features of the j-th region and the feature set F h ={f h,n} with the same category to obtain the set of feature distances

[0032]

[0033] where ‖.‖2 represents calculating the L2 norm;

[0034] Step 2: Regularize the feature distance ;

[0035] Step 3: Calculate the feature similarity by using row-wise softmax normalization on the regularized feature distance

[0036] Furthermore, the method of the fifth step is as follows:

[0037] (1) Non-overlappingly divide the ground truth image I g into a set of regional images where L represents the total number of regional images in the set of regional images;

[0038] (2) Resize the ground truth semantic segmentation result S g according to the regional images in the set of regional images After resizing, each pixel in the ground truth semantic segmentation result corresponds to an element in the set of regional images ;

[0039] (3) According to the correspondence between the resized true semantic segmentation result and the regional image set , assign a corresponding unique semantic category to each element in the regional image set ;

[0040] (4) According to the semantic category of the l-th regional image , select the feature set F h with the same category from the CICM Bank

[0041] ; (5) Through convolution operation, calculate the regional image feature f l g of the regional image

[0042] ; (6) Calculate the feature similarity set h = {f h,n} between the regional image feature and the feature set F

[0043] with similar categories, where N represents the total number of features in the feature set (7) According to the following formula, use the n-th element in the feature similarity set l g to weight the regional image feature f h,n , and merge it with the weighted n-th feature set element f h ′ ,n in the feature set to obtain the updated n-th feature set element f h,n :

[0044]

[0045] where w h,n is the weight corresponding to the feature set element f h,n , and the initial value is 1

[0046] (8) Replace the n-th feature set element f h,n and the n-th weight w h,n with the updated n-th feature set element f′ h,n and the updated n-th weight w′ h,n respectively, to obtain the updated feature set F h , and store the updated feature set F h back into the CICM Bank

[0047] Further, the calculation method of (6) in the fifth step is the same as that of (7) in the fourth step.

[0048] The advantages of the present invention are as follows:

[0049] 1. During the image restoration process of the present invention, since the feature library in the CICM Bank is extracted from different images, it provides richer cross-image context information for the restoration task.

[0050] 2. During the image restoration process of the present invention, the features in the CICM Bank are divided according to semantic categories, making the use of the CICM Bank more effective. At the same time, experiments prove that the richer the semantic categories, the finer the information contained in the CICM Bank, and the better the image restoration effect. Description of the Drawings

[0051] Figure 1 is a flowchart of an image restoration method based on cross-image context memory;

[0052] Figure 2 is a flowchart of the cross-image context memory module (CICM);

[0053] Figure 3 is the present invention Figure 2 flowcharts of the Memory Update and Memory Read modules in;

[0054] Figure 4 is the quantitative comparison of the method of the present invention with the other five existing optimal image methods. Detailed Embodiment

[0055] The technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings. Based on the technical solutions in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] First, some concepts in the present invention are elaborated. The cross-image context memory module (CICM module) contains a CICM Bank for storing feature maps obtained from cross-images. And all feature maps are divided into different feature sets according to semantic categories. The CICM Bank is represented as B = {F d |d = 1,…,D}, where F d = {f d,n ∈R C |0≤n≤N} represents all feature sets belonging to the d-th category, and f d,nIt represents the nth feature in the feature set of the dth category, where D represents the total number of all feature sets, and N represents the total number of features in the feature set of each dth category.

[0057] Secondly, the datasets used in the present invention for experimental feasibility verification are the Cityscapes dataset and the OutdoorScenes dataset respectively. Among them, the Cityscapes dataset contains a total of 5000 street view images. In the experiment, 2975 images are used for training and 1525 images are used for testing. The OutdoorScenes dataset contains 10200 natural images. In the experiment, 9900 images are used for training and 300 images are used for testing. The Cityscapes dataset and the OutdoorScenes dataset contain annotation information of 20 and 8 categories respectively. At the same time, the present invention uses five different masks for experiments, namely the central mask (the central area of the mask is missing) and four irregular masks (the irregular areas of the mask are missing). Among them, the irregular masks are divided into 0 - 20% area missing, 20 - 40% area missing, 40 - 60% area missing, and random missing. The present invention has carried out experimental verification on these five different masks respectively.

[0058] The model framework in the present invention is as Figure 1 shown. Among them, the feature extraction network uses U-Net[6] as the backbone, and skip connections are added between the encoder and the decoder. The input of the feature extraction network is the locally missing image and the mask. The locally missing image is obtained by deleting the corresponding positions of the ground truth image in the dataset according to the missing part in the mask. The output of the feature extraction network is the roughly repaired image and the roughly semantic segmentation result. The roughly repaired image, the roughly semantic segmentation result and the mask are sent into the CICM module together, and the feature maps in the CICM Bank are used to optimize the roughly repaired image to obtain the optimized image as the output of the CICM module. At the same time, the ground truth image and the ground truth semantic segmentation result in the dataset are sent into the CICM module to update the feature maps in the CICM Bank. Finally, the optimized image is combined with the roughly repaired image to obtain the final repaired image. It should be noted that the ground truth image and the ground truth semantic segmentation result are only used to update the CICM Bank during the model training process.

[0059] [6] Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation[C] / / International Conference on Medical image computing and computer-assisted intervention. Springer, Cham, 2015:234-241.

[0060] 1. Obtain a rough repaired image of the locally missing image through the network

[0061] Input the locally missing image I m ∈R H×W×3 and its mask M ∈ R H×W into the feature extraction network, where H and W represent the height and width of the image and the mask respectively. The output of the network is the rough repair result, which includes the rough repaired image I c for the locally missing image and the rough semantic segmentation result S c .

[0062] 2. Optimize the rough repaired image

[0063] Input the rough repaired image I c , the rough semantic segmentation result S c and the mask M into the CICM module to obtain the optimized image I r . The specific method is as follows:

[0064] (1) Divide the rough repaired image I c and the rough semantic segmentation result S c into the non-missing part image I u (This part retains the area corresponding to the black area of the mask M in the rough repaired image I c and discards the area corresponding to the white area of the mask M), the missing part image I t (This part discards the area corresponding to the black area of the mask M in the rough repaired image I c and retains the area corresponding to the white area of the mask M) and the missing part semantic segmentation result S t (This part discards the area corresponding to the black area of the mask M in the rough semantic segmentation result S c and retains the area corresponding to the white area of the mask M). Here, when generating the non-missing part image I u , the missing part image I t and the missing part semantic segmentation result S tWhen, in order to ensure the integrity of the missing part of the image I during step (2), t when segmenting the missing part of the image I t and the remaining area in the missing part semantic segmentation result S t may be larger than the corresponding white area in the mask M, and the remaining area in the non-missing part of the image I u will be reduced accordingly. The remaining area in the non-missing part of the image I u and the remaining area in the missing part of the image I t should ensure that the rough repaired image I c .

[0065] (2) Segment the missing part of the image I t non-overlappingly into a set of regional images where J represents the total number of regional images in the set of regional images.

[0066] (3) Resize the missing part semantic segmentation result S t according to the regional images in the set of regional images . After resizing, each pixel corresponds to an element in the set of regional images.

[0067] (4) According to the correspondence between the resized missing part semantic segmentation result and the set of regional images , assign a corresponding unique semantic category to each element in the set of regional images .

[0068] (5) According to the semantic category of the j-th regional image , select the feature set F h with the same category from the CICM Bank, where the number of the feature set is h.

[0069] (6) Through convolution operation, calculate the regional image feature of the regional image

[0070] (7) Calculate the set of feature similarities between the regional image feature h ={f h,n |n = 1,…,N} of the same category and the feature set F (see Note 1 for details), where N represents the total number of feature similarities.

[0071] (8) Weight the feature set F according to the set of feature similarities h to calculate the optimized regional feature . The calculation method is as follows

[0072] ​

[0073] (9) Calculate the optimized region features Calculate the optimized region image through convolution operation

[0074] (10) For all J elements in the region image set After steps (6)-(9), an optimized region image set composed of J optimized region images is obtained Stitch the optimized region image set according to the original position to obtain the optimized image I of the missing part q . When combining the optimized image I of the missing part q with the undamaged part image I u an optimized image I is obtained r .

[0075] Description 1: Image features of the j-th region and the feature set F with the same category h ={f h,n} Feature similarity Set Solution

[0076] Step 1: Calculate the feature of the j-th region image according to the following formula and the feature distance set between the feature set F h = {f h,n} with the same category

[0077]

[0078] where N is the total number of features in the feature set, and ‖.‖2 represents calculating the L2 norm.

[0079] Step 2: Regularize the feature distance to obtain

[0080]

[0081] Step 3: Calculate the feature similarity by using row-wise softmax normalization on the regularized feature distance

[0082]

[0083] where h>0 is a bandwidth parameter and N is the total number of features in the feature set.

[0084] 3. Update the CICM Bank

[0085] Using the rough repair image Ic True value image I of the same size g and the corresponding true value semantic segmentation result S g The specific method for updating the CICM Bank is as follows:

[0086] (1) Cut the true value image I g into a set of regional images without overlap where L represents the total number of regional images in the set of regional images.

[0087] (2) Resize the true value semantic segmentation result S g according to the regional images in the set of regional images. After resizing, each pixel in the true value semantic segmentation result corresponds to an element in the set of regional images.

[0088] (3) According to the correspondence between the resized true value semantic segmentation result and the set of regional images assign a corresponding unique semantic category to each element in the set of regional images

[0089] (3) Select the feature set F with the same category from the CICM Bank according to the semantic category of the l-th regional image h

[0090] (4) Calculate the regional image feature f of the regional image through convolution operation l g

[0091] (5) Calculate the feature similarity set between the regional image feature and the feature set F with a similar category according to the method in step (8) of "2. Optimize the rough repair image" h ={f h,n} where N represents the total number of features in the feature set and the number of feature similarities in the feature similarity set.

[0092] (6) Use the n-th element in the feature similarity set to weight the regional image feature f l g and merge it with the n-th weighted feature set element f in the feature set h,n to obtain the updated n-th feature set element f' h,n and the updated n-th weight w' h,n

[0093] ​​​​​​​

[0094] Among them, w h,n is the weight corresponding to the feature set element f h,n with an initial value of 1.

[0095] (7) Replace the updated nth feature set element f′ d,n and the updated nth weight w′ d,n with the nth feature set element f d,n and the nth weight w d,n respectively to obtain the updated feature set F d . Store the updated feature set F d back into the CICM Bank.

[0096] 4. Generation of the final repaired image I

[0097] Take the average of the rough repaired image I c and the optimized image I r to obtain the final repaired image I.

[0098] The feasibility of the method of the present invention is verified below with specific examples:

[0099] The datasets used in the experiment are the Cityscapes dataset and the Outdoor Scenes dataset. At the same time, three metrics, namely the peak signal-to-noise ratio (PSNR), the structural similarity index (SSIM), and the Frechet inception distance (FID), are used to quantitatively evaluate the repair results.

[0100] According to Figure 4 The results showing the effects of the method of the present invention and the existing optimal image repair method in repairing images under different conditions indicate that the method of the present invention has improvements in all metrics, which proves the effectiveness of the method of the present invention. At the same time, since the Cityscapes dataset has richer semantic segmentation information than the Outdoor Scenes dataset, the method of the present invention achieves a higher performance improvement on the Cityscapes dataset.

Claims

1. An image inpainting method based on cross-image context memory, comprising the following steps: First step, establish a CICM Bank to store the feature maps obtained from cross-images. The feature maps are divided into different feature sets according to semantic categories, denoted as B = {F d | d = 1, …, D}, where F d = {f d,n ∈ R C | 0 ≤ n ≤ N} represents all the feature sets belonging to the d-th category, and f d,n represents the n-th feature in the feature set of the d-th category. Here, D represents the total number of all feature sets, and N represents the total number of features in each feature set of the d-th category; The second step: Select a convolutional neural network for extracting image features; In the third step, a rough repaired image of the locally missing image is obtained through the convolutional neural network selected in the second step. The method is as follows: The locally missing image I m ∈R H×W×3 and its mask M ∈ R H×W are input into the convolutional neural network selected in the second step to extract features. Among them, H and W represent the height and width of the image and the mask respectively, and a rough repair result is output, including a rough repair image I for roughly repairing the locally missing image c and a rough semantic segmentation result S c ; Step 4: Optimize the roughly repaired image to obtain the optimized image I r , and the method is as follows: (1) Divide the roughly repaired image I c and the rough semantic segmentation result S c into the non-missing part image I u , the missing part image I t and the missing part semantic segmentation result S t ; (2) Segment the missing partial image I t into a set of regional images without overlap where J represents the total number of regional images in the set of regional images; (3)Segment the semantic segmentation result S of the missing part t According to the regional images in the set of regional images resize the regional images. After resizing, for the semantic segmentation result S of the missing part t each pixel corresponds to an element in the set of regional images; (4) According to the correspondence between the semantic segmentation result of the resized missing part and the regional image set , assign a corresponding unique semantic category to each element in the regional image set ; (5) According to the semantic category of the j-th regional image select a feature set F with the same category from the CICM Bank h ={f h,n |n = 1, …, N}; (6) Calculate the regional image through convolution operation The regional image features of (7) Calculate the regional image features and the feature set F with the same category h ={f h,n |n = 1, …, N} to obtain the feature similarity set where N represents the total number of feature similarities; (8) According to the feature similarity set weight the feature set F h to calculate the optimized region features by weighted calculation The calculation method is as follows (9) Optimize the regional features Calculate the optimized regional image through convolution operation (10) After all J elements in the regional image set have gone through steps (6)-(7), an optimized regional image set consisting of J optimized regional images is obtained The optimized regional image set is stitched together according to the original positions to obtain the optimized image I of the missing part q , and after the optimized image I of the missing part q is combined with the undamaged part of the image I u , the optimized image I is obtained r ; Step 5: Update the CICM Bank by using the ground truth image I c with the same size as the roughly repaired image I g and the corresponding ground truth semantic segmentation result S g to update the CICM Bank; Step 6: Based on the roughly repaired Image I c and the optimized Image I r , the final repaired Image I is obtained.

2. The image restoration method according to claim 1, wherein In (1) of the fourth step, the missing partial image I t and the semantic segmentation result S of the missing part t The area retained in is larger than the corresponding white area in the mask M, and the non-missing partial image I u The retained area will decrease accordingly; The non-missing part of image I u The reserved area in and the missing part of image I t The reserved area in should ensure that the roughly restored image I can be completely pieced together c .

3. The image restoration method according to claim 1, characterized in that The calculation method of (7) in the fourth step is as follows: Step 1: Calculate the image features of the j-th region and the feature set F with the same category h ={f h,n} between the feature distance sets Among them, ‖.‖2 represents calculating the L2 norm; Step 2: Regularize the feature distance ; Step 3: By means of the regularized feature distance Use row-by-row softmax normalization to calculate the feature similarity 4. The image restoration method according to claim 3, wherein The method of the fifth step is: (1) Divide the true value image I g into a set of regional images without overlapping where L represents the total number of regional images in the set of regional images; (2) Adjust the true value semantic segmentation result S g According to the regional images in the regional image set Resize it. After resizing, each pixel in the true value semantic segmentation result corresponds to an element in the regional image set ; (3) According to the correspondence between the resized true semantic segmentation result and the regional image set , assign a corresponding unique semantic category to each element in the regional image set ; (4) According to the semantic category of the l-th regional image select a feature set F with the same category from the CICM Bank h ; (5) Calculate the regional image through convolution operation The regional image feature f of l g ; (6) Calculate the regional image features and the feature set F with similar categories h ={f h,n}, the feature similarity set between them where N represents the total number of features in the feature set; Using the set of feature similarities according to the following formula the nth element in weight the region image feature f l g and merge it with the nth weighted feature set element f h,n in the feature set to obtain the updated nth feature set element f′ h,n and the updated nth weight w′ h,n : Among them, w h,n is the weight corresponding to the feature set element f h,n with an initial value of 1; (8) Replace the nth feature set element f′ h,n and the updated nth weight w′ h,n with the nth feature set element f h,n and the nth weight w h,n respectively, to obtain the updated feature set F h , and store the updated feature set F h back into the CICM Bank.

5. The image restoration method according to claim 4, wherein The calculation method of (6) in the fifth step is the same as that of (7) in the fourth step.

Citation Information

Patent Citations

  • Multi-modal feature fusion text-guided image restoration method

    CN111340122A

  • Image processing method, device and apparatus, and storage medium

    WO2020177513A1