Image inpainting method based on adversarial multi-scale and residual multi-channel spatial attention

Through the image repair method combined with the multi-channel fusion attention mechanism, the residual multi-channel spatial fusion attention encoder and the residual multi-scale spatial attention generator are used to solve the problem that image repair in the prior art is difficult to meet both semantic and visual requirements, and achieve high-quality image repair effects.

CN114782265BActive Publication Date: 2025-05-06NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210397765.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-05-06
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

Existing image repair methods are difficult to achieve high-quality repair results semantically and visually, especially with challenges in filling deep features and visual details of images.

Method used

The image repair method combined with a multi-channel fusion attention mechanism is adopted, and the residual multi-channel spatial fusion attention encoder and the residual multi-scale spatial attention generator are used to roughly repair and refined repair of the image, and the multi-scale discriminator is used for adversarial training to optimize the generation of adversarial network models.

Benefits of technology

It realizes the real-life image repair effect semantically and visually, and can effectively fill the deep features and visual details of the image, so that the repaired image remains consistent overall and locally.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782265B_ABST
    Figure CN114782265B_ABST
Patent Text Reader

Abstract

The present invention discloses an image restoration method based on adversarial multi-scale and residual multi-channel spatial attention. First, the initial damaged image is convolved to obtain a feature map, and its channels are evenly split into four feature maps, which are respectively input into the residual multi-channel spatial fusion attention generator to reconstruct and output a complete rough restoration image; secondly, the input rough restoration image is reconstructed by combining the residual multi-scale spatial attention encoder and the multi-scale decoder to obtain fine restoration images of different scales; finally, the fine restoration images of different scales are respectively input into the multi-scale discriminator to determine whether the restoration images of different scales are true, so as to determine the texture and semantic consistency of the restoration images of different scales. The present invention makes the missing area and the known area smoother, thereby obtaining a more refined and semantically consistent restoration image, so that the image maintains consistency both overall and locally.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image restoration method, and in particular to an image restoration method using a multi-channel fusion attention mechanism. Background Art

[0002] Image inpainting originally originated from an extremely primitive technique where artists repair damaged paintings to make them as close to the original as possible. In computer vision, it is achieved by filling in the missing pixels in the damaged image. Currently, this technique is widely used in many fields, such as old photo restoration, object removal, photo retouching, and text removal. Existing image inpainting methods can be divided into two categories. The first category of methods is traditional diffusion-based or patch-based texture synthesis techniques, which can only fill in the low-level features of the image. Due to the lack of high-level understanding of the image, this method cannot produce reasonable semantic results. To solve this problem, the second category of methods attempts to solve the inpainting problem through learning-based methods, by training deep convolutional networks to predict the pixels in the missing area. This method is mainly used to fill in the deep features of the image. However, although semantically relevant images can be generated, it is still challenging to obtain visually realistic results. Summary of the invention

[0003] Purpose of the invention: The purpose of the present invention is to provide an image restoration method that combines a multi-channel fusion attention mechanism to achieve restoration effects in both semantics and vision.

[0004] Technical solution: The image restoration method of the present invention comprises the following steps:

[0005] Step A: The real image I orig and mask image I mask Connect and generate the image to be repaired I input As input, the missing image feature map is obtained through convolution operation, and the channel of the missing image feature map is cut to output four new missing feature maps with the same channel size;

[0006] Step B: construct a residual multi-channel spatial fusion attention encoder, input the new missing feature maps respectively, perform feature connection, and output a rough repair image I inpainted1 ;

[0007] Step C: construct a residual multi-scale spatial attention generator and input a rough inpainted image I inpainted1 , output the repaired image I of different scales inpainted2(n) ;

[0008] Step D: Inpaint images I of different scales inpainted2(n) The missing areas of the restored images of different scales are input into the multi-scale discriminator to determine whether they are true or false;

[0009] Step E: The multi-scale discriminator and the residual multi-scale spatial attention generator train and optimize the generative adversarial network repair model according to different loss functions;

[0010] Step F, using the adversarial network repair model trained in step E, completes the repair of the image to be repaired and outputs the repaired image I inpainted2 .

[0011] Further, the specific implementation steps of step A are as follows:

[0012] Step A1: construct the image I to be repaired containing the missing area input Compared with the real image I orig Data pairs;

[0013] Step A2: The real image I orig With the mask dataset image I mask Perform element-by-element multiplication to obtain the image to be repaired I containing the missing area input ;

[0014] Step A3: image I to be repaired input Perform convolution operation to obtain the missing image feature map, perform channel cutting on the missing image feature map, and output four new missing image feature maps with the same channel size.

[0015] Further, the specific implementation steps of step B are as follows:

[0016] Step B1, based on the encoder-decoder structure, the residual block and the spatial attention mechanism are combined to construct a residual multi-channel spatial fusion attention generator;

[0017] Step B2, ensure that the encoder in the residual multi-channel spatial fusion attention generator is based on the U-Net decoder by adding residual modules and channel spatial fusion attention mechanisms layer by layer, and the multi-scale decoder is the decoder structure of U-Net;

[0018] Step B3, the residual module is processed using the ReLU activation function and the InstanceNorm normalization function;

[0019] Step B4, the convolution layer channel feature map is subjected to feature compression through an adaptive average pooling layer and an adaptive maximum pooling layer to output a feature vector of size 1×1×C; a sigmoid function is used to generate channel weights; each channel of the input image feature map of the channel attention mechanism network is multiplied by the weight value of the corresponding channel;

[0020] Step B5: Input the new missing image feature map into the spatial attention mechanism network, extract the missing area patch and calculate the attention score value, and finally fill the missing area with the context weighted by the attention value;

[0021] Step B6, the multi-scale decoder is processed using the ReLU activation function and the InstanceNorm normalization function.

[0022] Further, the specific implementation steps of step C are as follows:

[0023] Step C1, establish a generator network model, and the generator network structure uses a codec network structure based on U-Net;

[0024] Step C2: The encoder adds residual modules layer by layer based on the U-Net decoder, and fills the missing areas layer by layer in combination with the spatial attention mechanism;

[0025] Step C3, the residual module is processed using the ReLU activation function and the InstanceNorm normalization function;

[0026] Step C4, the spatial attention mechanism calculates the cosine similarity between patches inside and outside the missing area, calculates its attention score based on the semantic similarity, and fills the missing area with weighted context according to the attention score;

[0027] Step C5, the decoder is processed using the ReLU activation function and the InstanceNorm normalization function;

[0028] Step C6: The encoder outputs the restoration feature maps of different channels and performs feature connection layer by layer to obtain a new feature map after channel fusion. The feature map is decoded and output as the restoration image I of different scales. inpainted2(n) .

[0029] Further, the specific implementation steps of step D are as follows:

[0030] Step D1, the selected multi-scale discriminator includes a global discriminator and a local discriminator, wherein the multi-scale global discriminator is used to judge the global consistency of images of different scales, and the multi-scale local discriminator is used to judge the local consistency of images of different scales;

[0031] Step D2: transform the real image into the restored image I at different scales. inpainted2(n) Adjust to the same image as the repair image I inpainted2(n) Same size;

[0032] Step D3: Inpainting images I of different scales inpainted2(n) and the real image I origInput them into the global discriminator network respectively to calculate the image adversarial loss L at different scales global(n) ;

[0033] Step D4: Inpainting images I of different scales inpainted2(n) and the real image I orig The pixels in the known area are set to 0, and the repaired image and the real image that only retain the repaired area are obtained, which are input into the local discriminator network respectively to calculate the image adversarial loss L at different scales. local(n) ;

[0034] Step D5, obtain the global adversarial loss L for images of different scales global(n) and the local adversarial loss L local(n) , calculate the average loss L for each global(ave) and L local(ave) , the multi-scale discriminator loss L is the sum of the two.

[0035] Further, the specific implementation steps of step E are as follows:

[0036] Step E1: roughly repair image I inpainted1 Compared with the real image I orig Compare and find the reconstruction loss L G1 ;

[0037] Step E2: Inpainting images I of different scales inpainted2(n) and the real image I orig They are input into the global discriminator in the multi-scale discriminator respectively, and their judgment results are fed back to the residual multi-scale spatial attention generator to continue training;

[0038] Step E3: Retain only the repaired area of ​​the repaired image I at different scales inpainted2(n) and the real image I orig They are input into the local discriminators in the multi-scale discriminator respectively, and their judgment results are fed back to the residual multi-scale spatial attention generator to continue training;

[0039] Step E4, calculate the residual multi-scale spatial attention generator loss L G2 , where the loss L G Including average L1 loss and adversarial loss;

[0040] Step E5, repeat the above operations and continuously train the generative adversarial network model until the training is stable and tends to fit, and the optimal repair model is obtained.

[0041] Further, the specific implementation steps of step F are as follows:

[0042] Step F1, image I to be repaired inputPerform convolution operation to obtain the missing image feature map, perform channel cutting on the missing image feature map and output four new missing image feature maps with the same channel size;

[0043] Step F2, input the new missing image feature map into the residual multi-channel spatial fusion attention encoder structure;

[0044] Step F3, concat the padded four-channel new repaired image feature maps to obtain a complete feature map, and input it into the decoder of the residual multi-channel spatial fusion attention network to output a rough repaired image I i npa i n t e d1 ;

[0045] Step F4, input the rough repair image I inpainted1 To the residual multi-scale spatial attention encoder structure, after 6 layers of downsampling convolution, the feature maps of each convolution layer are recorded as X1, X2, X3, X4, X5, and X6 respectively. At the same time, each convolution layer is accompanied by a residual module to obtain deeper features;

[0046] Step F5, input X6 and X5 into the attention mechanism at the same time to obtain the feature map X7, input X7 and X4 into the attention mechanism at the same time to obtain the feature map X8, input X8 and X3 into the attention mechanism at the same time to obtain the feature map X9, input X9 and X2 into the attention mechanism at the same time to obtain the feature map X10, and input X10 and X1 into the attention mechanism at the same time to obtain the feature map X11;

[0047] Step F6, feature concatenate X7 and X6 to obtain UPX1, feature concatenate X8 and the feature map obtained by upsampling convolution of UPX1 to obtain UPX2, feature concatenate X9 and the feature map obtained by upsampling convolution of UPX2 to obtain UPX3, feature concatenate X10 and the feature map obtained by upsampling convolution of UPX3 to obtain UPX4, feature concatenate X11 and the feature map obtained by upsampling convolution of UPX4 to obtain UPX5, and then input UPX5 into two convolutional layers to obtain feature map UPX6;

[0048] In step F7, UPX1, UPX2, UPX3, UPX4, UPX5, and UPX6 are sequentially input into the convolutional layer for decoding to output restoration images of different scales; UPX6 is used as the final output image, and images of other scales are used to calculate the adversarial loss and the average L1 loss.

[0049] Compared with the prior art, the present invention has the following significant effects:

[0050] 1. The present invention adopts a residual multi-channel spatial fusion attention encoder. By introducing residual blocks and multiple channels, more features can be obtained, and the spatial attention mechanism can enable the model to focus on more important features;

[0051] 2. The present invention adopts a residual multi-channel spatial fusion attention encoder to obtain a rough repaired image. By introducing residual blocks and multi-scales, the model can obtain features of different scales of the image and obtain a more refined repaired image.

[0052] 3. The present invention adopts multi-scale global and local discriminators to judge the authenticity of missing areas and repaired images of different scales, and can obtain features of different scales. At the same time, it makes the missing areas and known areas smoother to obtain more refined and semantically consistent repaired images, so that the image remains consistent both overall and locally. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0054] Figure 2 It is a network model structure diagram of the present invention;

[0055] Figure 3 is a residual block structure diagram of the present invention;

[0056] Figure 4 Schematic diagram of the channel attention mechanism of the present invention;

[0057] Figure 5 Schematic diagram of the spatial attention mechanism of the present invention;

[0058] Figure 6 This is a structural diagram of the residual multi-channel spatial fusion attention encoder of the present invention;

[0059] Figure 7 It is a schematic diagram of the repair result of the present invention. DETAILED DESCRIPTION

[0060] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0061] like Figure 1The figure shows the overall flow chart of the image restoration method of the present invention. The present invention adopts a two-stage image restoration model from coarse to fine, divides the restoration model into a generator and a discriminator, and jointly optimizes the model by reconstruction loss, adversarial loss, and average L1 loss. The first stage adopts an encoder-decoder structure to restore the input missing image by constructing a residual multi-channel spatial fusion attention encoder, and outputs a rough restoration image. The second stage adopts a generative adversarial network to construct a residual multi-scale spatial attention generator to refine the input rough image, and input it into a multi-scale discriminator to determine whether the restoration image is true. First, the initial damaged image is convolved to obtain a feature map, and its channels are evenly split into four feature maps, which are respectively input into the residual multi-channel spatial fusion attention generator to reconstruct and output a complete rough restoration image. Secondly, the input rough restoration image is reconstructed by combining the residual multi-scale spatial attention encoder and the multi-scale decoder to obtain fine restoration images of different scales. Finally, the fine restoration images of different scales are respectively input into the multi-scale discriminator to determine whether the restoration images of different scales are true, so as to determine the texture and semantic consistency of the restoration image at different scales. The present invention conducts experiments on different image datasets, combined with Figure 7 As shown in the figure, good restoration effects can be achieved for both irregular random missing and regular missing areas. The restored images are excellent in both visual and semantic aspects. The implementation process includes the following steps:

[0062] Step (A): the real image I orig and mask image I mask Connect to generate the image to be repaired I input As input, the convolution operation obtains the missing image feature map, and the channel cutting of the missing image feature map is performed to output four new missing image feature maps with the same channel size;

[0063] Step (B) constructs a residual multi-channel spatial fusion attention encoder, inputs the new missing image feature maps respectively, performs feature connection, and outputs a rough repair image I inpainted1 ;

[0064] Step (C), construct a residual multi-scale spatial attention generator, input I inpainted1 , output the repaired image I of different scales inpainted2(n) ;

[0065] Step (D): the inpainted images I of different scales are inpainted2(n) The missing areas of the restored images of different scales are input into the multi-scale discriminator to determine whether they are true or false;

[0066] Step (E), the multi-scale discriminator and the residual multi-scale spatial attention generator train and optimize the generative adversarial network repair model according to different loss functions;

[0067] Step (F), using the adversarial network restoration model trained in step (E), completes the restoration of the image to be restored, and outputs the restored image I inpainted2 .

[0068] like Figure 2 As shown, in step (A), the real image I orig and mask image I mask Connect and generate the image to be repaired I input As input, the convolution operation obtains the missing image feature map, performs channel cutting on the missing image feature map, and outputs four new missing image feature maps with equal channels. The implementation process includes the following steps:

[0069] Step (A1), construct the image to be repaired containing the missing area I input Compared with the real image I orig Data pair; where the real image I orig The dataset uses CelebA-HQ face dataset, Places2 scene dataset, The four benchmark datasets are the structural dataset and the DTD texture dataset. The irregular mask dataset used at the same time comes from the Nvidia Mask Dataset proposed by the Nvidia team, which contains six types of irregular random masks with different masking areas. The regular mask data is composed of rectangular masks, which set the pixels of the rectangular area in the center of the image to 255 (white).

[0070] Step (A2): Based on the benchmark dataset (i.e., the real dataset) and the mask dataset proposed above, an image I to be repaired containing the missing area is obtained. input Compared with the real image I orig The specific operation is: the benchmark dataset image (i.e. the real image I orig ) and the mask dataset image I mask Perform element-by-element multiplication to obtain the image to be repaired I containing the missing area input ;

[0071] Step (A3), treat the repaired image I input Perform a convolution operation to obtain the missing image feature map, perform channel cutting on it and output four new missing image feature maps with the same channel size.

[0072] like Figure 3 As shown, in step (B), the implementation process of constructing the residual multi-channel spatial fusion attention network includes the following steps:

[0073] Step (B1), based on the encoder-decoder structure, the residual block and the spatial attention mechanism are combined to construct a residual multi-channel spatial fusion attention generator, and the encoder-decoder network structure is similar to the encoder-decoder network structure of U-Net;

[0074] Step (B2), wherein the encoder in the residual multi-channel spatial fusion attention generator is based on the U-Net decoder and adds residual modules and channel spatial fusion attention mechanisms layer by layer, and the multi-scale decoder is the decoder structure of U-Net;

[0075] Step (B3), wherein the residual module consists of two atrous convolutional layers with a dilation rate of 1, a convolution kernel size of 3×3, and uses the ReLU activation function and InstanceNorm normalization function to optimize the residual block model, while using spectral normalization to accelerate the model training speed;

[0076] Step (B4), such as Figure 4 As shown in the figure, the channel attention mechanism operates as follows: the convolution layer channel feature map is compressed by the adaptive average pooling layer and the adaptive maximum pooling layer to output a feature vector of size 1×1×C; it is reduced and increased in dimension by two convolution layers and the ReLU activation function to obtain a feature vector of size 1×1×C; the Sigmoid function is used to generate channel weights; each channel of the input image feature map of the channel attention mechanism network is multiplied by the weight value of the corresponding channel, so that the channel attention mechanism can focus on more important detail features;

[0077] Step (B5), such as Figure 5 As shown in Figure 1, the spatial attention mechanism operates as follows: first, the new missing image feature map is input into the spatial attention mechanism network, then the missing area patch is extracted and the attention score value is calculated, and finally the missing area is filled with the context weighted by the attention value;

[0078] In step (B6), the decoder structure includes 5 layers of upsampling convolutions, the convolution kernel size is 3×3, the step size is 1, the boundary padding is 1, and the ReLU activation function and InstanceNorm normalization function are used to optimize the decoder model.

[0079] For step (C), a residual multi-scale spatial attention generator is constructed. The specific implementation process includes the following steps:

[0080] Step (C1), establishing a generator network model, the generator network structure uses a codec network structure based on U-Net;

[0081] Step (C2), the encoder adds residual modules layer by layer based on the U-Net decoder, and fills the missing areas layer by layer in combination with the spatial attention mechanism;

[0082] Step (C3), in which the residual module consists of two dilation rate 1 hollow convolution layers, the convolution kernel size is 3×3, and the ReLU activation function and InstanceNorm normalization function are used to optimize the residual block model, and spectral normalization is used to accelerate the model training speed;

[0083] Step (C4), in which the spatial attention mechanism calculates the cosine similarity between patches inside and outside the missing area, calculates its attention score based on the semantic similarity, and fills the missing area with weighted context according to the attention score;

[0084] Step (C5), the decoder structure includes 6 layers of upsampling convolution. The upsampling convolution module uses bilinear interpolation to decode the feature map. The convolution kernel size is 3×3, the step size is 1, and the ReLU activation function and InstanceNorm normalization function are used to optimize the decoder model.

[0085] Step (C6), the encoder outputs the restoration feature maps of different channels, performs feature connection layer by layer, obtains the new restoration feature map of channel fusion, decodes the new restoration feature map of channel fusion, and outputs the restoration image I of different scales inpainted2(n) .

[0086] For step (D), the inpainted images I of different scales are inpainted2(n) The missing areas of the repaired images of different scales are input into the multi-scale discriminator to judge whether they are true or false. The specific implementation process includes the following steps:

[0087] Step (D1), the multi-scale discriminator used in the present invention includes a global discriminator and a local discriminator, wherein the multi-scale global discriminator is used to judge the global consistency of images of different scales, and the multi-scale local discriminator is used to judge the local consistency of images of different scales;

[0088] Step (D2), the real image is converted into the restored image I at different scales. inpainted2(n) Adjust to the same size M×M;

[0089] Step (D3), the inpainted images I of different scales are inpainted2(n) and the real image I orig Input them into the global discriminator network respectively to calculate the image adversarial loss L at different scales global(n) ;

[0090] Step (D4), the inpainted images I of different scales are inpainted2(n) and the real image I origThe pixels in the known area are set to 0, and the repaired image and the real image that only retain the repaired area are obtained, which are input into the local discriminator network respectively to calculate the image adversarial loss L at different scales. local(n) ;

[0091] Step (D5), obtain the global adversarial loss L for images of different scales global(n) and the local adversarial loss L local(n) , calculate the average loss L for each global(ave) and L local(ave) , the multi-scale discriminator loss L is the sum of the two; if the value of the loss function is smaller, it means that the generated image is closer to the real image, that is, the generator is better trained.

[0092] For step (E), the model discriminator and generator train and optimize the generative adversarial network repair model based on different loss functions. The specific implementation process includes the following steps:

[0093] Step (E1), roughly repair the image I inpainted1 Compared with the real image I orig Compare and find the reconstruction loss L G1 , used to optimize the residual multi-channel spatial fusion attention encoder and decoder;

[0094] Step (E2), the inpainted images I of different scales are inpainted2(n) and the real image I orig They are input into the global discriminator in the multi-scale discriminator respectively, and their judgment results are fed back to the residual multi-scale spatial attention generator to continue training;

[0095] Step (E3), only retain the repaired area of ​​the repaired image I of different scales inpainted2(n) and the real image I orig They are input into the local discriminators in the multi-scale discriminator respectively, and their judgment results are fed back to the residual multi-scale spatial attention generator to continue training;

[0096] Step (E4), calculate the residual multi-scale spatial attention generator loss L G2 The residual multi-scale spatial attention generator loss function adopts a joint loss function including average L1 loss and adversarial loss, where the average L1 loss is calculated by finding the inpainted images I at different scales. inpainted2(n) and the real image I orig , only the inpainted images of different scales in the inpainted area are retained I inpainted2(n) and the real image I orig The average L1 loss of

[0097] Step (E5), repeat the above operations and continuously train the generative adversarial network model until the training is stable and tends to fit, and the optimal repair model is obtained.

[0098] For step (F), the restoration model trained in step (E) is used to complete the restoration of the image to be restored and output the restored image I inpainted2 The specific implementation process includes the following steps:

[0099] Step (F1), treat the repaired image I input Perform convolution operation to obtain the missing image feature map, perform channel cutting on it and output four new missing image feature maps with the same channel size;

[0100] Step (F2) inputs the new missing image feature map into the residual multi-channel spatial fusion attention encoder structure so that the information of the same dimension in the encoding and decoding parts can be fused, such as Figure 6 As shown in the figure, specifically, the convolution layer of the encoding part of the network is modified, the first channel combination block is contacted with the second channel block after a residual module convolution operation, and the result of the connection is used for a channel space attention to make the network focus on important channels and fill in the missing areas, and then a downsampling convolution is performed on the channel combination block to restore the number of channels to the number of channels of the original second channel block, and the same operation is performed between the second channel block and the third channel block and between the third channel block and the fourth channel block, and finally the two new feature maps are fused and downsampled with the original feature map to keep the final output feature map at the original size;

[0101] Step (F3) concats the padded four-channel new inpainted image feature maps to obtain a complete feature map, which is input into the decoder of the residual multi-channel spatial fusion attention network to output a rough inpainted image I inpainted1 ;

[0102] Step (F4), input the rough repair image I inpainted1 To the residual multi-scale spatial attention encoder structure, after 6 layers of downsampling convolution, the convolution kernel size of each layer is 3×3, the step size is 2, the boundary padding is 1, and the number of convolution kernels obtained is 32, 64, 128, 256, 512 and 512 respectively. The output feature maps of each layer are recorded as X1, X2, X3, X4, X5, and X6 respectively. At the same time, each convolution layer is accompanied by a residual module to obtain deeper features;

[0103] Step (F5), input X6 and X5 into the attention mechanism at the same time to obtain feature map X7, input X7 and X4 into the attention mechanism at the same time to obtain feature map X8, input X8 and X3 into the attention mechanism at the same time to obtain feature map X9, input X9 and X2 into the attention mechanism at the same time to obtain feature map X10, and input X10 and X1 into the attention mechanism at the same time to obtain feature map X11;

[0104] Step (F6), feature concatenate X7 and X6 to obtain UPX1, feature concatenate X8 and the feature map obtained by upsampling convolution of UPX1 to obtain UPX2, feature concatenate X9 and the feature map obtained by upsampling convolution of UPX2 to obtain UPX3, feature concatenate X10 and the feature map obtained by upsampling convolution of UPX3 to obtain UPX4, feature concatenate X11 and the feature map obtained by upsampling convolution of UPX4 to obtain UPX5, and then input UPX5 into two convolutional layers to obtain feature map UPX6;

[0105] In step (F7), UPX1, UPX2, UPX3, UPX4, UPX5, and UPX6 are sequentially input into the convolutional layer for decoding to output restoration images of different scales; UPX6 is the final output image with a size of 256×256×3, i.e., the final restoration image, and images of other scales are used to calculate adversarial loss and average L1 loss to optimize the model.

[0106] Figure 7 It is a schematic diagram of the repair result of the present invention, and it can be concluded that:

[0107] (1) The multi-channel spatial attention used in the present invention can repair the texture details of the image on different channels and pay more attention to the repaired details;

[0108] (2) The adversarial multi-scale idea used in the present invention can restore the global and local semantic consistency of the image at different scales, so that the image restoration area is semantically consistent with the background area, and can obtain detailed information at different scales, so that the image has a better visual effect;

[0109] (3) The present invention can obtain good repaired images by repairing both regular missing regions and irregular missing regions, and is also suitable for image repair of different image data sets.

Claims

1. An image restoration method based on adversarial multi-scale and residual multi-channel spatial attention, characterized in that: The following steps are involved: Step A: The real image I orig and mask image I mask Connect and generate the image I to be repaired input As input, the missing image feature map is obtained through convolution operation, and the channel of the missing image feature map is cut to output four new missing feature maps with the same channel size; Step B: construct a residual multi-channel spatial fusion attention encoder, input the new missing feature maps respectively, perform feature connection, and output a rough repair image I inpainted1 ; Step C: construct a residual multi-scale spatial attention generator and input a rough inpainted image I inpainted1 , output the repaired image I of different scales inpainted2(n) ; Step D: Inpaint images I of different scales inpainted2(n) The missing areas of the restored images of different scales are input into the multi-scale discriminator to determine whether they are true or false; Step E: The multi-scale discriminator and the residual multi-scale spatial attention generator train and optimize the generative adversarial network repair model according to different loss functions; Step F, using the adversarial network repair model trained in step E, completes the repair of the image to be repaired and outputs the repaired image I inpainted2 .

2. The image restoration method based on adversarial multi-scale and residual multi-channel spatial attention according to claim 1 is characterized in that: The specific implementation steps of step A are as follows: Step A1: construct the image I to be repaired containing the missing area input Compared with the real image I orig Data pairs; Step A2: The real image I orig With the mask dataset image I mask Perform element-by-element multiplication to obtain the image to be repaired I containing the missing area input ; Step A3: image I to be repaired input Perform convolution operation to obtain the missing image feature map, perform channel cutting on the missing image feature map, and output four new missing image feature maps with the same channel size.

3. The image restoration method based on adversarial multi-scale and residual multi-channel spatial attention according to claim 1 is characterized in that: The specific implementation steps of step B are as follows: Step B1, based on the encoder-decoder structure, the residual block and the spatial attention mechanism are combined to construct a residual multi-channel spatial fusion attention generator; Step B2, ensure that the encoder in the residual multi-channel spatial fusion attention generator is based on the U-Net decoder by adding residual modules and channel spatial fusion attention mechanisms layer by layer, and the multi-scale decoder is the decoder structure of U-Net; Step B3, the residual module is processed using the ReLU activation function and the InstanceNorm normalization function; Step B4, the convolution layer channel feature map is subjected to feature compression through an adaptive average pooling layer and an adaptive maximum pooling layer to output a feature vector of size 1×1×C; a sigmoid function is used to generate channel weights; each channel of the input image feature map of the channel attention mechanism network is multiplied by the weight value of the corresponding channel; Step B5: Input the new missing image feature map into the spatial attention mechanism network, extract the missing area patch and calculate the attention score value, and finally fill the missing area with the context weighted by the attention value; Step B6, the multi-scale decoder is processed using the ReLU activation function and the InstanceNorm normalization function.

4. The image restoration method based on adversarial multi-scale and residual multi-channel spatial attention according to claim 1, characterized in that: The specific implementation steps of step C are as follows: Step C1, establish a generator network model, and the generator network structure uses a codec network structure based on U-Net; Step C2: The encoder adds residual modules layer by layer based on the U-Net decoder, and fills the missing areas layer by layer in combination with the spatial attention mechanism; Step C3, the residual module is processed using the ReLU activation function and the InstanceNorm normalization function; Step C4, the spatial attention mechanism calculates the cosine similarity between patches inside and outside the missing area, calculates its attention score based on the semantic similarity, and fills the missing area with weighted context according to the attention score; Step C5, the decoder is processed using the ReLU activation function and the InstanceNorm normalization function; Step C6: The encoder outputs the restoration feature maps of different channels and performs feature connection layer by layer to obtain a new feature map after channel fusion. The feature map is decoded and output as the restoration image I of different scales. inpainted2(n) .

5. The image restoration method based on adversarial multi-scale and residual multi-channel spatial attention according to claim 1, characterized in that: The specific implementation steps of step D are as follows: Step D1, the selected multi-scale discriminator includes a global discriminator and a local discriminator, wherein the multi-scale global discriminator is used to judge the global consistency of images of different scales, and the multi-scale local discriminator is used to judge the local consistency of images of different scales; Step D2: transform the real image into the restored image I at different scales. inpainted2(n) Adjust to the same image as the repair image I inpainted2(n) Same size; Step D3: Inpainting images I of different scales inpainted2(n) and the real image I orig Input them into the global discriminator network respectively to calculate the image adversarial loss L at different scales global(n) ; Step D4: Inpainting images I of different scales inpainted2(n) and the real image I orig The pixels in the known area are set to 0, and the repaired image and the real image that only retain the repaired area are obtained, which are input into the local discriminator network respectively to calculate the image adversarial loss L at different scales. local(n) ; Step D5, obtain the global adversarial loss L for images of different scales global(n) and the local adversarial loss L local(n) , calculate the average loss L for each global(ave) and L local(ave) , the multi-scale discriminator loss L is the sum of the two.

6. The image restoration method based on adversarial multi-scale and residual multi-channel spatial attention according to claim 1, characterized in that: The specific implementation steps of step E are as follows: Step E1: roughly repair image I inpainted1 Compared with the real image I orig Compare and find the reconstruction loss L G1 ; Step E2: Inpainting images I of different scales inpainted2(n) and the real image I orig They are input into the global discriminator in the multi-scale discriminator respectively, and their judgment results are fed back to the residual multi-scale spatial attention generator to continue training; Step E3: Retain only the repaired area of ​​the repaired image I at different scales inpainted2(n) and the real image I orig They are input into the local discriminators in the multi-scale discriminator respectively, and their judgment results are fed back to the residual multi-scale spatial attention generator to continue training; Step E4, calculate the residual multi-scale spatial attention generator loss L G2 , where the loss L G Including average L1 loss and adversarial loss; Step E5, repeat the above operations and continuously train the generative adversarial network model until the training is stable and tends to fit, and the optimal repair model is obtained.

7. The image restoration method based on adversarial multi-scale and residual multi-channel spatial attention according to claim 1, characterized in that: The specific implementation steps of step F are as follows: Step F1, image I to be repaired input Perform convolution operation to obtain the missing image feature map, perform channel cutting on the missing image feature map and output four new missing image feature maps with the same channel size; Step F2, input the new missing image feature map into the residual multi-channel spatial fusion attention encoder structure; Step F3, concat the padded four-channel new repaired image feature maps to obtain a complete feature map, and input it into the decoder of the residual multi-channel spatial fusion attention network to output a rough repaired image I inpainted1 ; Step F4, input the rough repair image I inpainted1 To the residual multi-scale spatial attention encoder structure, after 6 layers of downsampling convolution, the feature maps of each convolution layer are recorded as X1, X2, X3, X4, X5, and X6 respectively. At the same time, each convolution layer is accompanied by a residual module to obtain deeper features; Step F5, input X6 and X5 into the attention mechanism at the same time to obtain the feature map X7, input X7 and X4 into the attention mechanism at the same time to obtain the feature map X8, input X8 and X3 into the attention mechanism at the same time to obtain the feature map X9, input X9 and X2 into the attention mechanism at the same time to obtain the feature map X10, and input X10 and X1 into the attention mechanism at the same time to obtain the feature map X11; Step F6, feature concatenate X7 and X6 to obtain UPX1, feature concatenate X8 and the feature map obtained by upsampling convolution of UPX1 to obtain UPX2, feature concatenate X9 and the feature map obtained by upsampling convolution of UPX2 to obtain UPX3, feature concatenate X10 and the feature map obtained by upsampling convolution of UPX3 to obtain UPX4, feature concatenate X11 and the feature map obtained by upsampling convolution of UPX4 to obtain UPX5, and then input UPX5 into two convolutional layers to obtain feature map UPX6; In step F7, UPX1, UPX2, UPX3, UPX4, UPX5, and UPX6 are sequentially input into the convolutional layer for decoding to output restoration images of different scales; UPX6 is used as the final output image, and images of other scales are used to calculate the adversarial loss and the average L1 loss.

Citation Information

Patent Citations

  • Generative adversarial network image restoration method based on multi-scale texture feature branches

    CN113902630A

  • Image restoration method and system based on dense multi-scale fusion

    CN114155171A