An image restoration method based on multi-scale residual module and feature fusion

By designing an image restoration network with a multi-scale residual module and a dual-channel global attention module, the problems of insufficient attention to contextual information and insufficient feature fusion are solved, and high-quality image restoration effects are achieved, which is suitable for the restoration of works of art and historical materials.

CN119599914BActive Publication Date: 2025-09-26ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411644972.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-09-26
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

Existing image restoration algorithms based on generative adversarial networks suffer from insufficient attention to contextual information and insufficient feature fusion, resulting in insufficient texture clarity and structural integrity of the restored images.

Method used

A multi-scale residual module and a dual-channel global attention module are designed to be integrated into the feature fusion network. The image restoration network is trained through a comprehensive loss function. The multi-scale residual module is used to capture features of different scales of the image. The dual-channel global attention module is combined to enhance feature fusion and realize image restoration.

Benefits of technology

The quality of the repaired image is improved, making its texture clear, structure complete, and semantic coherent, alleviating the gradient vanishing problem. It has high computational efficiency and robustness and is suitable for complex image damage situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119599914B_ABST
    Figure CN119599914B_ABST
Patent Text Reader

Abstract

The present invention discloses an image restoration method based on a multi-scale residual module and feature fusion, which belongs to the field of image processing technology. The multi-scale residual module designed by the present invention can effectively capture the feature information of different scales of the image, from subtle texture to macroscopic structure, and improve the quality of the restored image. The feature fusion strategy of fusing detail branches and semantic branches can fully integrate features at different levels, enhance the detail performance and overall coordination of the image, and avoid the limitations that may be brought by single-scale features by fusing multi-scale features, so that the restored image is more realistic and delicate in color, texture and structure. In addition, the multi-scale residual module can alleviate the gradient vanishing problem, has high computational efficiency, and can complete the image restoration task in a shorter time. The dual-channel global attention module introduced in the design has good robustness, so that it can function stably in the face of various complex image damage situations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image restoration method based on a multi-scale residual module and feature fusion. Background Art

[0002] Image restoration, a hot topic in computer graphics and vision, has significant applications in removing occluded areas, removing specific objects, and restoring valuable historical data. Traditional restoration methods achieve good results for images with small defective areas and a high degree of similarity between the image information and the defective area. However, for images with larger defective areas and richer structural and texture information, traditional restoration algorithms often lose texture detail, resulting in unsatisfactory restoration results. With the development of deep learning theory, an increasing number of researchers have applied it to image restoration, but none of these methods fully integrate the image's contextual information to achieve images with rich textures and natural structures.

[0003] The above problems need to be solved urgently. To this end, the present invention proposes an image restoration method based on a multi-scale residual module and feature fusion. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: how to solve the problems of insufficient attention to contextual information and insufficient feature fusion in most current image restoration algorithms based on generative adversarial networks, and provide an image restoration method based on a multi-scale residual module and feature fusion, so that the restored image (i.e., the restored image) has clear texture, complete structure, and semantic coherence.

[0005] The present invention solves the above technical problems through the following technical solutions, which include the following steps:

[0006] S1: Design image restoration network structure

[0007] Design a multi-scale residual module and a dual-channel global attention module, and fuse them into a feature fusion network to obtain an image restoration network.

[0008] S2: Training the image restoration network

[0009] Use the selected dataset and construct a comprehensive loss function to train the image restoration network to obtain the trained image restoration model;

[0010] S3: Image Inpainting

[0011] The image to be repaired is input into the image repair model and the repaired image is output.

[0012] Furthermore, in the multi-scale residual module of step S1, the input feature map is input into three branches respectively:

[0013] In the first branch, Point Conv is used to increase the number of channels to three times the initial number of channels. Then, Avgpool convolution is used to extract local features, compress and integrate them. At the same time, the entire feature map is averaged to obtain global feature information. The receptive field is 1 compared to the input feature map.

[0014] In the second branch, 3×3 convolution is first used to obtain local features, and then the features output by the convolution are normalized through the BN layer, and then processed using the ReLU activation function. Finally, 3×3 convolution and BN layers are combined in sequence for processing. The receptive field is 5 compared to the input feature map.

[0015] In the third branch, 5×5 convolution is first used to obtain local features, and then the features output by the convolution are normalized through the BN layer, and then processed using the ReLU activation function. Finally, the 5×5 convolution and BN layers are combined in sequence for processing. The field of view is 9 compared to the input feature map.

[0016] The image features at different levels output by the three branches are fused by channel dimension splicing, and then Point Conv is used to compress the channel dimension of the image features. Then, Trans Conv is used to restore it to the input feature map size and output it.

[0017] Furthermore, in the dual-channel global attention module of step S1, it is divided into a channel attention mechanism and a spatial attention mechanism;

[0018] In the channel attention mechanism M C In the first branch, the feature map F is input into two branches; in the first branch, the feature map F first undergoes two layers of MLP to perform nonlinear transformation on the feature vector. The first layer of MLP maps the input feature vector to an intermediate dimension to extract a more abstract feature representation. The second layer of MLP maps the features of the intermediate dimension back to the same dimension as the number of channels, and then uses average pooling to compress the information of each channel to obtain global information at the channel level. Then, one-dimensional convolution with an adaptive convolution kernel is used to generate weights; in the second branch, the feature map F first undergoes global maximum pooling to extract aggregate features, and then a fully connected layer is used to assign different weights to the channels to build mutual dependence between channels; then the weights of the two branches are superimposed, and the sigmoid function is applied to normalize the channel attention output weight, and finally multiplied with the input feature map F to extract the feature map F with channel attention weight. mc ;

[0019] In the spatial attention mechanism M S In the feature map F mcUse two convolution layers with a convolution kernel size of 7 to fuse spatial information, and generate spatial attention weights through the sigmoid function, and finally combine them with the input feature map F mc Multiply them together to obtain a feature map F with both channel attention weights and spatial attention weights sc .

[0020] Furthermore, in the channel attention mechanism M C In the second branch, the calculation expression of the one-dimensional convolution of the adaptive convolution kernel is as follows:

[0021]

[0022] Among them, k is the convolution kernel size, C is the number of channels, γ and b are the weights for changing the number of channels and the convolution kernel, ∥ odd To find odd operations.

[0023] Furthermore, in the channel attention mechanism M C In the feature map F mc The calculation expression is as follows:

[0024] M c (F mc )=σ(W2δ(W1(AvgPool(F)))+f k×k (MaxPool(F)))F

[0025] Among them, σ is the sigmoid function, W1 is the weight of the first fully connected layer FC, W2 is the weight of the second fully connected layer FC, δ is the ReLU activation function, AvgPool is the average pooling, MaxPool is the global maximum pooling, f k×k is a convolution with a kernel of k×k.

[0026] Furthermore, in step S1, the image restoration network includes a semantic branch and a structure texture detail branch. The specific processing process is as follows: the image to be restored is input, the semantic branch first uses the dilated convolution with dilation rates of 1, 2, and 5 to expand the receptive field, and then uses the depthwise separable convolution and multi-scale residual module to extract the multi-scale features of the image, and then applies the Stem module to downsample and eliminate redundant features; the structure texture detail branch first uses 1×1 and 3×3 convolution layers to increase the number of channels of the feature map, and then uses the self-calibration convolution module to adjust the weight of the convolution kernel, and then the extracted feature map is branched and output: one of the branches introduces 3×3 convolution with a step size of 2 and average pooling , and then processed by cascaded DW convolution and multi-scale residual modules. Another branch introduces a dual-channel global attention module to capture the positional relationship between different channels and spaces, and then processes it by cascaded 3×3 convolution and multi-scale residual modules. Finally, the features extracted by the structure texture detail branch and the semantic branch are added, spliced ​​and fused in the channel dimension, and decoded by the decoder. Among them, the decoder uses transposed convolution to restore the image size, uses depthwise separable convolution to capture local details, and obtains the multi-scale features of the image through the multi-scale residual module to obtain the repaired image. The original image is the real image of the image to be repaired before it is damaged, and the image to be repaired is the damaged image.

[0027] Furthermore, in step S2, the comprehensive loss function is defined as follows:

[0028] L=λ per L per +λ pix L pix +λ TV L TV +λ adv L adv +L1

[0029] Among them, L per is the perceptual loss, L pix is the pixel reconstruction loss, L TV is the total variational loss, L adv is the adversarial loss, and L1 is the L1 loss.

[0030] Furthermore, the perceptual loss L per For comparison with the original image I gt and the repaired image I pred The difference in feature representations in the trained VGG-19 network is used to measure the perceptual similarity between them using the perceptual loss, which is calculated as follows:

[0031]

[0032] Among them, E is the expected function, N iis the number of elements in the i-th layer, is the i-th layer feature of the image in the trained VGG-19 network;

[0033] Pixel reconstruction loss L pix Used to calculate the original image I gt and the repaired image I pred To evaluate the pixel difference between the two, the calculation formula is as follows:

[0034] L pix =I gt -I pred ;

[0035] Total variational loss L TV Used to enhance the smoothness and continuity of the image. The calculation formula is as follows:

[0036]

[0037] Among them, m is the horizontal coordinate of the pixel, n is the vertical coordinate of the pixel, and p is the pixel area. To repair the pixel point of the image at coordinate (m,n);

[0038] Adversarial loss L adv By minimizing the loss of the generator G and maximizing the loss of the discriminator D, the generator G is trained to generate samples, while the discriminator D is able to accurately distinguish between real samples and generated samples. The calculation formula is as follows:

[0039]

[0040] Among them, E is the expected function, D is the discriminator, G is the generator, and I in For damaged images;

[0041] L1 loss is used to determine the restored image I pred and the original image I gt The difference between them is calculated as follows:

[0042] L1=λ hole L hole +λ valid L valid

[0043] Among them, λ hole and λ valid are the weight coefficients of pixels in the damaged area and pixels in the valid area, L hole and L valid are the loss of pixel values ​​in the damaged area and the valid area respectively.

[0044] Furthermore, L hole and L valid The calculation formula is as follows:

[0045] L hole =(HWC) -1 (1-M)(I pred -I gt )1

[0046] L val =(HWC) -1 M(I pred -I gt )1

[0047] Among them, H, W, C are the image height, width and number of channels respectively, M is the binary mask of the initial input, and is the dot product.

[0048] Compared with the prior art, the present invention has the following advantages:

[0049] 1. The multi-scale residual module designed in this paper can effectively capture feature information at different scales of the image, from subtle textures to macro structures, and improve the quality of the restored image. The feature fusion strategy of fusing detail branches and semantic branches can fully integrate features at different levels, enhance the detail expression and overall coordination of the image, and avoid the limitations of single-scale features by fusing multi-scale features, making the restored image more realistic and delicate in color, texture, and structure.

[0050] 2. The multi-scale residual module designed in the present invention can alleviate the gradient vanishing problem, has high computational efficiency, and can complete the image restoration task in a shorter time. The introduced dual-channel global attention module has good robustness, which enables it to function stably in the face of various complex image damage situations.

[0051] 3. The present invention can be used in the field of art to help repair damaged works of art, restore the original appearance of historical paintings, or creatively modify and improve existing works. In addition, the present invention has important application value in removing blocked areas, removing specific objects, and restoring precious historical materials. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 1 is a flow chart of an image restoration method based on a multi-scale residual module and feature fusion according to the present invention;

[0053] Figure 2 Schematic diagram of the structure of the image restoration network in the second embodiment of the present invention;

[0054] Figure 3 is a schematic structural diagram of a multi-scale residual module in the second embodiment of the present invention;

[0055] Figure 4Schematic diagram of the structure of the dual-channel global attention module in the second embodiment of the present invention;

[0056] Figure 5 3 is a schematic diagram of the qualitative comparison of the experimental results of the Paris Street View dataset in the fourth embodiment of the present invention, wherein (a) is an example of a real image (original image), (b) is an example of a damaged image (input image, i.e., an image to be repaired) after the original image is processed, (c) is an example of the repair result of the CE algorithm, (d) is an example of the repair result of the PEN-Net algorithm, (e) is an example of the repair result of the RFR algorithm, (f) is an example of the repair result of the MADF algorithm, and (g) is an example of the repair result of the present invention. DETAILED DESCRIPTION

[0057] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. However, the protection scope of the present invention is not limited to the following embodiment.

[0058] Example 1

[0059] like Figure 1 As shown, this embodiment provides a technical solution: an image restoration method based on a multi-scale residual module and feature fusion, comprising the following steps:

[0060] Step 1: Design the image restoration network structure;

[0061] Step 2: Train and test the image restoration network to obtain the image restoration model;

[0062] Step 3: Input the image to be repaired into the image repair model and output the repaired image to obtain the repaired image.

[0063] Example 2

[0064] In this embodiment, the specific process in step 1 is described as follows. Figure 2-4 As shown in the figure, the design of the image restoration network structure mainly includes: designing a multi-scale residual module (see Figure 3 ), design a dual-channel global attention module (see Figure 4 ), use the designed modules to build an image restoration network based on multi-scale residual modules and feature fusion (see Figure 2 ), the specific process is as follows:

[0065] (1) Design a multi-scale residual module

[0066] First, input the input feature map into three branches respectively:

[0067] In the first branch, Point Conv is used to increase the number of channels to three times the initial number of channels. Avgpool convolution is used to extract local features and compress and integrate them to reduce the size of the feature map. At the same time, the entire feature map is averaged to obtain global feature information. The receptive field is 1 compared to the input feature map.

[0068] In the second branch, 3×3 convolution is used to obtain local features, and the features output by the convolution are normalized through the BN layer to achieve effective propagation of the gradient, thereby improving the training efficiency and the convergence speed of the model; the ReLU activation function is then used to increase the expressiveness of the model, and then the 3×3 convolution and BN layers are combined in sequence to improve the training efficiency, generalization ability and robustness of the model. The receptive field is 5 compared to the input feature map.

[0069] In the third branch, a larger convolution kernel can cover more pixel areas, so a 5×5 convolution is first applied in the third branch to increase the receptive field to 9 compared to the input feature map. Since residual modules with different receptive fields can achieve feature fusion and information interaction at different levels, this paper constructs a multi-scale residual module with three receptive fields and performs channel-dimensional concatenation to fuse image features at different levels, making the restoration result more reasonable in terms of details and overall structure.

[0070] In order to fully integrate the multi-scale information of the image and maintain the balance between the overall structure and local details, PointConv is used to compress the channel dimension of the image features, and TransConv is used to restore it to the input feature map size, supplement the lost detail information, and improve the feature extraction ability of the module.

[0071] (2) Design a dual-channel global attention module

[0072] The dual-channel global attention module is divided into a channel attention mechanism and a spatial attention mechanism. After the feature map F passes through the dual-channel global attention module, the feature map F is obtained. sc , channel attention mechanism (M C ) First, the feature map F is input into the two branches. In branch one, the feature map F first undergoes a two-layer MLP to perform a nonlinear transformation on the feature vector. The first layer of MLP can map the input feature vector to an intermediate dimension to extract a more abstract feature representation. The second layer of MLP maps the intermediate dimension features back to the same dimension as the number of channels. Average pooling is then used to compress the information of each channel to obtain global information at the channel level. A one-dimensional convolution with an adaptive convolution kernel is then used to generate weights to increase the influence of each channel on the extracted features. This convolution operation reduces the amount of data while improving performance gain. The convolution kernel size is k, and the expression is:

[0073]

[0074] Among them, k is the convolution kernel size, C is the number of channels, γ and b are the weights for changing the number of channels and the convolution kernel, γ = 2, b = 1, ∥ odd To find odd operations.

[0075] In branch 2, the feature map F is first subjected to global maximum pooling (Max pooling) to extract aggregate features, and then a fully connected layer (FC) is used to assign different weights to the channels to build mutual dependence between channels. The mechanism then superimposes the weights of the two branches and applies sigmoid function normalization to solve the channel attention output weight, and finally multiplies it with the input feature map F to extract the feature map F with channel attention weights. mc , the expression is as follows:

[0076] M c (F mc )=σ(W2δ(W1(AvgPool(F)))+f k×k (MaxPool(F)))F

[0077] Among them, σ is the sigmoid function, W1 is the weight of the first fully connected layer FC, W2 is the weight of the second fully connected layer FC, δ is the ReLU activation function, AvgPool is the average pooling, MaxPool is the global maximum pooling, f k×k is a convolution with a kernel of k×k.

[0078] In the spatial attention mechanism M S In the feature map F mc Use two convolution layers with a convolution kernel size of 7 to fuse spatial information, and generate spatial attention weights through the sigmoid function, and finally add them to the input F mc Multiply the feature maps to obtain a feature map F with both channel attention weights and spatial attention weights sc , enhance the extraction of image structure and texture information.

[0079] (3) Build an image restoration network

[0080] The image inpainting network integrates a semantic branch with a structural texture detail branch. Given a damaged image to be inpainted, the semantic branch first applies dilated convolutions with dilation rates of 1, 2, and 5 to expand the receptive field, capturing a wider range of image information without increasing the number of parameters or computational overhead, thereby reducing grid artifacts. A cascade of depthwise separable convolutions (DW convolutions) and multi-scale residual modules is then constructed to extract multi-scale features, enhancing feature propagation and reuse, and improving the network's expressiveness. Finally, the Stem module is used to downsample and eliminate features with low impact and redundant information. The structural texture detail branch uses 1×1 and 3×3 convolutional layers to increase the number of channels in the feature map, enhancing the dependency of structural texture detail information and extracting image detail texture features. The ReLU activation function is then used to enhance the model's nonlinearity. To improve the model's ability to capture structural texture information, a self-calibrated convolution (SC) module is introduced to adjust the convolution kernel weights. The extracted feature map is then branched out: one branch uses a 3×3 convolution with a stride of 2 and average pooling to reduce the size of the feature map. Since average pooling has the problem of losing some important texture detail information, the model of the present invention introduces a new branch, which is weightedly fused with the features after average pooling to highlight the detail features. At the same time, the representation ability of the network is improved by cascading DW convolution and multi-scale residual modules; another branch is to improve the model's attention to and understanding of details. The designed dual-channel global attention (DCGAM) module is introduced to capture the positional relationship between different channels and spaces, enhance the cross-dimensional interaction of image features, and retain the edge texture features of the image. Finally, a cascade of 3×3 convolution and multi-scale residual modules is constructed to extract the multi-scale information of the image and improve the robustness and stability of the model. Finally, the features extracted by the structural texture detail branch and the semantic branch are added, spliced ​​and fused in the channel dimension, and decoded by the decoder. The decoder uses transposed convolution to restore the image size, adopts DW convolution to capture local details, establishes a multi-scale residual module to obtain the multi-scale features of the image, and reconstructs an image with rich and accurate texture details, that is, a repaired image is obtained.

[0081] Example 3

[0082] In this embodiment, the specific process in step 2 is described as follows:

[0083] The public dataset: Paris Street View Paris Street View dataset is trained, and the perceptual loss L is used when training the network. per , pixel reconstruction loss L pix , total variational loss L TV , adversarial loss L adv and L1 loss, where the perceptual loss L per , pixel reconstruction loss L pix , total variational loss L TV, adversarial loss L adv The and L1 loss functions are as follows:

[0084] Perceptual loss L per For comparison with the original image I gt and the repaired image I pred The perceptual loss L is the difference in feature representations in the trained VGG-19 network to weigh the perceptual similarity between them. per The calculation formula is as follows:

[0085]

[0086] Among them, E is the expected function, N i is the number of elements in the i-th layer, is the i-th layer feature in the trained VGG-19 network.

[0087] Pixel reconstruction loss L pix Used to calculate the original image I gt and the repaired image I pred To evaluate the pixel difference between the two. Pixel reconstruction loss L pix The calculation formula is as follows:

[0088] L pix =I gt -I pred

[0089] Total variational loss L TV Used to enhance the smoothness and continuity of the image. Total variation loss L TV The calculation formula is as follows:

[0090]

[0091] Among them, m is the horizontal coordinate of the pixel, n is the vertical coordinate of the pixel, and p is the pixel area. To repair the pixel point of the image at coordinate (m,n);

[0092] Adversarial loss L adv By minimizing the loss of the generator G and maximizing the loss of the discriminator D, the generator G is trained to generate realistic samples while enabling the discriminator D to accurately distinguish between real samples and generated samples. adv The calculation formula is as follows:

[0093]

[0094] Among them, E is the expected function, D is the discriminant network, G is the generating network, I in For damaged images;

[0095] L1 loss is used to determine the restored image Ipred and the original image I gt The difference between . The calculation formula of L1 loss is as follows:

[0096] L1=λ hole L hole +λ valid L valid

[0097] Among them, λ hole and λ valid are the weight coefficients of pixels in the damaged area and pixels in the valid area, L hole and L valid are the loss of pixel values ​​in the damaged area and the valid area, L hole and L valid The calculation formula is as follows:

[0098] L hole =(HWC) -1 (1-M)(I pred -I gt )1

[0099] L val =(HWC) -1 M(I pred -I gt )1

[0100] Among them, H, W, C are the image height, width and number of channels respectively, M is the binary mask of the initial input, and is the dot product.

[0101] The calculation formula of the total loss function of the network is as follows:

[0102] L=λ per L per +λ pix L pix +λ TV L TV +λ adv L adv +λ hole L hole +λ valid L valid

[0103] Among them, λ per ,λ pix ,λ TV ,λ adv are the weight coefficients of the corresponding loss terms.

[0104] Then, experiments are conducted on the obtained network structure and the experimental results are qualitatively and quantitatively analyzed.

[0105] Example 4

[0106] In this embodiment, the specific process in step three is described:

[0107] Dataset: The model was trained and tested on the Paris Street View dataset using an irregular mask dataset to verify the effectiveness of the proposed network model. The Paris Street View dataset contains 15,000 images of Paris street scenes, 14,900 of which were randomly selected as training images and 100 as test images.

[0108] Comparison methods: CE (Context encoders), a convolutional neural network model consisting of an encoder and a decoder; PEN-Net (Perceptual Encoder Network), including a pyramid context encoder and a multi-scale decoder, and also proposes an attention transfer network; RFR (Recurrent Feature Reasoning), an iterative reasoning image restoration framework RFR, and also proposes an attention module KCA suitable for the iterative restoration framework; MADF (Multi-scale Attention-based Deep Fusion), a two-stage restoration network, and also designs a bidirectional gated feature fusion module and a context feature aggregation module.

[0109] Evaluation metrics: Peak signal-to-noise ratio (PSNR), an objective measure of the difference between the original image and the restored image; structural similarity (SSIM), a measure of the similarity between two images; Mean Absolute Error (MAE), a statistical measure of the difference between the predicted value and the true value.

[0110] Experimental details: The verification experiment software environment carried out by the present invention is the pytorch1.10 framework, the CUDA version is 11.3, and the python version is 3.9. The CPU used in the computing platform is Intel Core i7-11700K, the GPU is NVIDIARTX 3070, the video memory is 8GB, the running memory is 32GB, and the system is Windows 10. The present invention adopts an end-to-end approach and uses the Adam algorithm for optimization. The parameters of the optimizer are set to β1=0.5 and β2=0.999 respectively, the initial learning rates of the discriminator and the generator are 0.0004 and 0.0001 respectively, the batch size is set to 2, and the weight hyperparameters of each part of the loss function are set to λ per=0.1,λ pix =120,λ TV =0.001,λ adv =0.009,λ hole =6,λ valid =1, using qualitative evaluation analysis and quantitative evaluation comparison as evaluation indicators.

[0111] Quantitative evaluation: As shown in Table 1, in experiments using the Paris Street View dataset at different mask ratios, the inpainted images generated by the proposed method have MAEs lower by 0.86%, 0.81%, 0.74%, and 0.61% respectively compared with the MADF (Multi-scale Attention-based Deep Fusion) method; PSNRs are improved by 2.86dB, 0.89dB, 0.38dB, and 0.02dB respectively; and SSIMs are improved by 0.6%, 0.7%, 0.7%, and 0.4% respectively.

[0112] Table 1 Comparison of quantitative experimental results on the Paris Street View dataset

[0113]

[0114] It should be noted that in the above table, bold fonts represent the optimal values ​​of each column, “↓” means the lower the better, and “↑” means the higher the better. OURS is the image restoration model in this invention.

[0115] Qualitative evaluation: Figure 5 The qualitative experimental results on the Paris Street View dataset validation images are analyzed as follows:

[0116] like Figure 5 As shown:

[0117] like Figure 5 (c) is a street view image restored by the CE algorithm, where the door and window structures have problems such as pixel blur and artifacts;

[0118] like Figure 5 (d) shows the repaired image generated by the PEN-Net algorithm. There are obvious differences in the door and window structures. When the defect ratio reaches about 30%, the window structure of the image cannot be effectively repaired.

[0119] like Figure 5 (e) shows the image restored by the RFR algorithm, which has structural distortion and artifacts at the windows and walls.

[0120] like Figure 5(f) is the image restored by the MADF algorithm, where the door frame, window and surrounding unrestored parts have unclear structural textures;

[0121] like Figure 5 As shown in (g), compared with other algorithms, the present invention can better restore the texture structure of the damaged image.

[0122] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. An image restoration method based on a multi-scale residual module and feature fusion, characterized in that: The following steps are involved: S1: Design image restoration network structure Design a multi-scale residual module and a dual-channel global attention module, and fuse them into a feature fusion network to obtain an image restoration network. S2: Training the image restoration network Use the selected dataset and construct a comprehensive loss function to train the image restoration network to obtain the trained image restoration model; S3: Image Inpainting Input the image to be repaired into the image repair model and output the repaired image; In step S1, the image restoration network includes a semantic branch and a structure texture detail branch. The specific processing process is as follows: the image to be restored is input, the semantic branch first uses the dilated convolution with dilated rates of 1, 2, and 5 to expand the receptive field, and then uses the depth-separable convolution and multi-scale residual module to extract the multi-scale features of the image, and then applies the Stem module downsampling to eliminate redundant features; the structure texture detail branch first uses 1×1 and 3×3 convolution layers to increase the number of channels of the feature map, and then uses the self-calibration convolution module to adjust the weight of the convolution kernel, and then the extracted feature map is branched and output: one of the branches introduces a 3×3 convolution with a step size of 2 and average pooling, and then passes It is processed by cascaded DW convolution and multi-scale residual modules, and another branch introduces a dual-channel global attention module to capture the positional relationship between different channels and spaces, and then processed by cascaded 3×3 convolution and multi-scale residual modules; finally, the features extracted by the structure texture detail branch and the semantic branch are added, spliced ​​and fused in the channel dimension, and decoded by the decoder; among them, the decoder uses transposed convolution to restore the image size, uses depthwise separable convolution to capture local details, and obtains the multi-scale features of the image through the multi-scale residual module to obtain the repaired image; the original image is the real image of the image to be repaired before it is damaged, and the image to be repaired is the damaged image.

2. The image restoration method based on a multi-scale residual module and feature fusion according to claim 1, characterized in that: In the multi-scale residual module of step S1, the input feature map is input into three branches respectively: In the first branch, Point Conv is used to increase the number of channels to three times the initial number of channels. Then, Avgpool convolution is used to extract local features, compress and integrate them. At the same time, the entire feature map is averaged to obtain global feature information. The receptive field is 1 compared to the input feature map. In the second branch, 3×3 convolution is first used to obtain local features, and then the features output by the convolution are normalized through the BN layer, and then processed using the ReLU activation function. Finally, 3×3 convolution and BN layers are combined in sequence for processing. The receptive field is 5 compared to the input feature map. In the third branch, 5×5 convolution is first used to obtain local features, and then the features output by the convolution are normalized through the BN layer, and then processed using the ReLU activation function. Finally, the 5×5 convolution and BN layers are combined in sequence for processing. The field of view is 9 compared to the input feature map. The image features at different levels output by the three branches are fused by channel dimension splicing, and then PointConv is used to compress the channel dimension of the image features. Then, TransConv is used to restore it to the input feature map size and output it.

3. The image restoration method based on multi-scale residual module and feature fusion according to claim 2, characterized in that: In the dual-channel global attention module of step S1, it is divided into a channel attention mechanism and a spatial attention mechanism; In the channel attention mechanism M C In the first branch, the feature map F is input into two branches; in the first branch, the feature map F first undergoes two layers of MLP to perform nonlinear transformation on the feature vector. The first layer of MLP maps the input feature vector to an intermediate dimension to extract a more abstract feature representation. The second layer of MLP maps the features of the intermediate dimension back to the same dimension as the number of channels, and then uses average pooling to compress the information of each channel to obtain global information at the channel level. Then, one-dimensional convolution with an adaptive convolution kernel is used to generate weights; in the second branch, the feature map F first undergoes global maximum pooling to extract aggregate features, and then a fully connected layer is used to assign different weights to the channels to build mutual dependence between channels; then the weights of the two branches are superimposed, and the sigmoid function is applied to normalize the channel attention output weight, and finally multiplied with the input feature map F to extract the feature map F with channel attention weight. mc ; In the spatial attention mechanism M S In the feature map F mc Use two convolution layers with a convolution kernel size of 7 to fuse spatial information, and generate spatial attention weights through the sigmoid function, and finally combine them with the input feature map F mc Multiply them together to obtain a feature map F with both channel attention weights and spatial attention weights sc .

4. The image restoration method based on a multi-scale residual module and feature fusion according to claim 3, characterized in that: In the channel attention mechanism M C In the second branch, the calculation expression of the one-dimensional convolution of the adaptive convolution kernel is as follows: Among them, k is the convolution kernel size, C is the number of channels, γ and b are the weights for changing the number of channels and the convolution kernel, || odd To find odd operations.

5. The image restoration method based on multi-scale residual module and feature fusion according to claim 4, characterized in that: In the channel attention mechanism M C In the feature map F mc The calculation expression is as follows: M c (F mc )=σ(W2δ(W1(MaxPool(F)))+f k×k (AvgPool(MLP2(F))))F Among them, σ is the sigmoid function, W1 is the weight of the first fully connected layer FC, W2 is the weight of the second fully connected layer FC, δ is the ReLU activation function, AvgPool is the average pooling, MaxPool is the global maximum pooling, f k×k is a convolution with a kernel of k×k.

6. The image restoration method based on multi-scale residual module and feature fusion according to claim 1, characterized in that: In step S2, the comprehensive loss function is defined as follows: L−λ per L per +λ pix L pix +λ TV L TV +λ adv L adv +L1 Among them, L per is the perceptual loss, L pix is the pixel reconstruction loss, L TV is the total variational loss, L adv is the adversarial loss, and L1 is the L1 loss.

7. The image restoration method based on multi-scale residual module and feature fusion according to claim 6, characterized in that: Perceptual loss L per For comparison with the original image I gt and the repaired image I pred The difference in feature representations in the trained VGG-19 network is used to measure the perceptual similarity between them using the perceptual loss, which is calculated as follows: Among them, E is the expected function, N i is the number of elements in the i-th layer, is the i-th layer feature of the image in the trained VGG-19 network; Pixel reconstruction loss L pix Used to calculate the original image I gt and the repaired image I pred To evaluate the pixel difference between the two, the calculation formula is as follows: L pix =I gt -I pred ; Total variational loss L TV Used to enhance the smoothness and continuity of the image. The calculation formula is as follows: Among them, m is the horizontal coordinate of the pixel, n is the vertical coordinate of the pixel, and p is the pixel area. To repair the pixel point of the image at coordinate (m,n); Adversarial loss L adv By minimizing the loss of the generator G and maximizing the loss of the discriminator D, the generator G is trained to generate samples, while the discriminator D is able to accurately distinguish between real samples and generated samples. The calculation formula is as follows: Among them, E is the expected function, D is the discriminator, G is the generator, and I in For damaged images; L1 loss is used to determine the restored image I pred and the original image I gt The difference between them is calculated as follows: L1−λ hole L hole +λ valid L valid Among them, λ hole and λ valid are the weight coefficients of pixels in the damaged area and pixels in the valid area, L hole and L valid are the loss of pixel values ​​in the damaged area and the valid area respectively.

8. The image restoration method based on multi-scale residual module and feature fusion according to claim 7, characterized in that: L hole and L valid The calculation formula is as follows: L hole =(HWC) -1 ||(1-M)⊙(I pred -I gt )||1 L val =(HWC) -1 ||M⊙(I pred -I gt )||1 Among them, H, W, C are the image height, width and number of channels respectively, M is the binary mask of the initial input, and ⊙ is the dot product.

Citation Information

Patent Citations

  • Image restoration model based on parallel adaptive guide network and method thereof

    CN114972062A