Multi-scale feature fusion generative adversarial network image restoration method based on structure guidance

Through the multi-scale feature fusion based on structure-guided generation adversarial network, the problem of lack of coherence between image repair areas and original images and blurred texture details in the prior art is solved, and image repair effects with higher quality and rich details are achieved.

CN120070258APending Publication Date: 2025-05-30XIAN TECH UNIV

Patent Information

Application Number
CN202411934688.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the prior art deals with large-scale pixel deletion, it is impossible to fully restore the global semantic information of the image, resulting in the lack of visual and content coherence between the generated repair area and the original image, and when the irregular and random damage to non-rectangles, the generated prediction results lack reasonable expression of the image content structure, and there are problems of distortion and blur.

Method used

A multi-scale feature fusion based on structure guidance is used to generate an adversarial network, and an adversarial network structure is generated through two stages of the structure repair network and the texture repair network. Combined with gating convolution and multi-scale feature fusion strategies, the edge information utilization rate and the accuracy of the network for relative position information are improved.

Benefits of technology

The generated repair areas are more visually and content-wise with the original image, and the texture details are clearer, which significantly improves the quality and detailed performance of image repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070258A_ABST
    Figure CN120070258A_ABST
Patent Text Reader

Abstract

The invention relates to an image restoration method of a multi-scale feature fusion generative adversarial network based on structure guidance. Comprising the following steps: constructing a generative adversarial network structure; acquiring a training image and preprocessing the acquired training image; inputting the preprocessed training image into the constructed generative adversarial network, and training the network; and step 4, inputting the to-be-restored image into the trained generative adversarial network for restoration to obtain a complete edge structure image. And finally, inputting the complete edge structure image and the to-be-restored image into the texture restoration network to restore texture details so as to obtain a final restoration result. According to the method, the accuracy of image edge structure restoration is improved, the generated restoration area and the original image are good in visual and content coherence, a highly vivid and reasonable image restoration result can be generated, and the expression in texture details is more accurate and reasonable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an image inpainting method based on a structure-guided multi-scale feature fusion generative adversarial network. Background Art

[0002] Image inpainting is a cutting-edge topic in current deep learning research and has important research value in aspects such as cultural relic protection, medical imaging, biomedical imaging, and aerospace. Image inpainting is to repair the pixels in the damaged area of an image so that the repaired image is closer to the original image. By filling the pixels in the damaged area to make it similar to the original image, the repair effect is achieved. The image inpainting task also provides a new research idea and approach to solve the existing problems such as image semantic loss, target occlusion, and content damage.

[0003] Currently, there are some problems in the field of image inpainting based on the deep neural network model architecture. When dealing with the situation of large-scale pixel loss, existing algorithms often cannot fully restore the global semantic information of the image, resulting in a lack of coherence in vision and content between the generated inpainted area and the rest of the original image. In addition, in the face of non-rectangular irregular random damages, the generated prediction results lack a reasonable expression of the image content structure, and there are certain degrees of distortion and blurring problems.

[0004] Traditional single discriminators have many limitations when dealing with complex image inpainting tasks, such as unstable training, mode collapse, and lack of effective evaluation criteria, which limit the diversity of generated samples and reduce the quality of image inpainting. The generator may only capture some data patterns, resulting in a lack of diversity in the generated results. At the same time, the single discriminator cannot effectively utilize global information, which may lead to problems such as discontinuous semantic structures and poor quality in the generated images. And it is difficult to provide an effective evaluation of the quality and diversity of output samples, which restricts the further improvement of model performance.

[0005] In the literature (Nazeri K, Ng E, Joseph T, et al. Edgeconnect: Generative image inpainting with adversarial edge learning[J]. arxiv preprint arxiv:1901.00212, 2019), partial convolution is adopted in the image feature extraction stage, which cannot truly play the guiding role of edge information. The network cannot effectively distinguish the edge pixels of the defective area, and the network cannot judge whether the position where the edge information is located is outside or inside the defective area. Therefore, it is difficult to correctly update the mask information of the next layer, weakening the learning and reasoning ability of the network. This results in low utilization rate of edge information by the model and lack of relative position information in the deep layer of the network.

[0006] In the literature (Yu J, Lin Z, Yang J, et al. Generative image inpainting with contextual attention[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018:5505-5514.), due to the weak association between the mask and the image, unnatural hole problems often occur during image inpainting, including the appearance of artifacts, the residue of the mask boundary, etc. When dealing with the situation of large-scale pixel loss, existing algorithms often cannot fully restore the global semantic information of the image, resulting in the lack of coherence in vision and content between the generated inpainted area and the rest of the original image. When facing non-rectangular irregular random damage, the generated prediction results lack a reasonable expression of the image content structure, and there are problems such as a certain degree of distortion and blurring. Summary of the Invention

[0007] The present invention aims to solve the technical problems existing in the prior art, such as low utilization rate of edge information, lack of relative position information in the deep layer of the network, lack of coherence in vision and content between the generated inpainted area and the rest of the original image, and a certain degree of distortion and blurring.

[0008] To solve the above problems, the present invention provides a structure-guided multi-scale feature fusion generative adversarial network image inpainting method, including the following steps:

[0009] Step 1: A structure-guided multi-scale feature fusion generative adversarial network is composed of a sequentially connected structure inpainting network and a texture inpainting network. Both the structure inpainting network and the texture inpainting network include a generator and a discriminator;

[0010] Step 2: Obtain training images and preprocess the obtained training images;

[0011] Step 3: Input the preprocessed training images into the structure-guided multi-scale feature fusion generative adversarial network and train the network, including the following steps:

[0012] 3.1 Edge detection stage: Extract the edge structure of the input damaged image through the HED edge detection algorithm. First, preprocess the damaged image by grayscale processing the RGB input image I input to obtain a grayscale image I gray , where the channel of the grayscale image is 1. Then, extract the edge structure of the damaged image through the Holistically-Nested Edge Detection algorithm to obtain an incomplete structure diagram E edge' .

[0013] 3.2 Structure repair stage: The incomplete structure diagram E edge′ , the mask image M, and the grayscale image I gray are used as the input E input to the structure repair network. The mask image is the shape of the damaged area, with pixel values of 1 in the missing area and 0 in the rest; the network in this stage is a generative adversarial network structure, consisting of a generator network and a discriminator network; the discriminator network consists of two spectrally-normalized Markov discriminators working at different scales. The generator and the discriminator conduct adversarial game training, and finally the generator generates a complete edge structure diagram

[0014] E input ={I gray , E edge′ , M} (1)

[0015] The generator network and the discriminator network D edg conduct adversarial training to reach the "Nash equilibrium" state, and finally the generator generates a complete edge structure diagram E edge

[0016] E edge =G edge (E input ) (2)

[0017] 3.3 Texture repair stage: The structure diagram E edge′ is a complete edge structure diagram formed by splicing the predicted structure part in the predicted edge structure image output by the network in the second stage and the undamaged part of the image in the incomplete structure diagram E edge′ in the first stage. The overall input T input of this stage is the damaged image I in and the complete edge image E edge spliced together in the channel dimension to form a joint input

[0018]

[0019] T input = {E edge , I in}} (4)

[0020] Step 4: Input the image to be repaired into the trained structure-guided multi-scale feature fusion generative adversarial network to obtain the final repaired image. Step 4: Input the image to be repaired into the trained structure-guided multi-scale feature fusion generative adversarial network to obtain the final repaired image.

[0021] Furthermore, in the above Step 1, the generator G1 in the structure repair network adopts the structure of downsampling block + gated convolutional sequence + upsampling block;

[0022] The generator G2 of the texture repair network is composed of a convolutional downsampling block, an attention mechanism module, a multi-scale dilated convolution fusion module, a multi-scale pyramid feature fusion module, and a convolutional upsampling block.

[0023] Furthermore, the combined loss of the above structure repair network consists of L1 loss, feature matching loss, and hinge loss of the generative network, as shown in Equation (13):

[0024] L = λ 1 L rec + λ 2 L Dsn + λ 3 L Fm (13)

[0025] Among them, L rec represents the pixel-level L1 reconstruction loss, represents the spectral normalization Markov discriminator loss, L fm is the feature matching loss, and λ 1 , λ 2 and λ 3 are weighted hyperparameters that balance the L1 loss, feature matching loss, and hinge loss

[0026] Furthermore, the attention mechanism module in the generator G2 of the above texture repair network is composed of a global average pooling layer, a convolutional layer, a ReLU activation function layer, a normalization layer, and a Sigmoid activation function layer.

[0027] Furthermore, the discriminators D3 and D4 of the above texture repair network are composed of two global and local discriminators with the same structure.

[0028] Furthermore, the loss function of the above texture repair network consists of a pixel reconstruction loss, a perceptual loss, a style loss, and an adversarial loss function. The total loss function L is expressed as Equation (23):

[0029] L = λ 1 L rec + λ 2 L prec + λ 3 L style + λ 4 L adv (23)

[0030] Where: L rec is the pixel reconstruction loss, L prec is the perceptual loss, L adv is the adversarial loss, and λ 1 、λ 2 、λ 3 、λ 4 are adjustment parameters.

[0031] Compared with the prior art, the advantages of the present invention are as follows:

[0032] 1. The network of the structure-guided multi-scale feature fusion generative adversarial network proposed by the present invention designs a two-stage generative adversarial network structure. In the first-stage structure repair network, by introducing a gated convolution mechanism, the utilization efficiency of the input image features is significantly improved.

[0033] The structure repair effect is optimized by embedding a gated convolution sequence in the feature extraction stage of the structure repair network. The gated convolution unit updates its filter parameters after each convolution operation, enabling the gated selection mechanism and the feature extraction process to be adaptively adjusted according to the current input features, significantly improving the utilization efficiency of the input image features. Through such a design, the generator network can effectively learn the association between the structure information of the known region and the masked region during the iterative training process, improve the utilization rate of edge information, and accurately obtain the relative position information in the deep layer of the network, so as to generate a more reasonable edge structure, and the generated repair region has good coherence with the original image in terms of vision and content.

[0034] At the same time, the single discriminator structure of the generative adversarial network is improved, and a multi-scale discriminator architecture is introduced in the structure repair network to enhance the network's ability to repair the edge structure of damaged images. The small-sized discriminator focuses on the coherence of the global structure, while the large-sized discriminator focuses on guiding the generator to generate detailed texture details. In order to improve the detail quality of the predicted image structure and stabilize the training process, the idea of spectral normalization is introduced into the discriminator network, combined with the hinge loss function, to improve the accuracy of image edge structure repair.

[0035] 2. The present invention introduces a multi-scale feature fusion strategy into the generator structure of the texture repair network and combines it with the attention mechanism to improve the detailed performance of the repair effect. The texture repair network mainly uses the multi-scale feature fusion method for deep feature extraction. By combining different convolutional kernels, it extracts and fuses multi-scale and deep-level image features. Feature maps of different sizes have different feature expression capabilities. Low-level image features can extract the basic detail information of the image without relying on the shape and spatial relationship of the image, and can effectively capture local texture detail information. In contrast, high-level image features can learn more abstract and semantic information. They can capture the overall style, texture, and higher-level structural features of the image. These features have the advantages of rotational invariance, translational invariance, and scale invariance. By combining image features of different scales, the network can effectively integrate the fine details of low-order features and the rich semantic attributes of high-order features. At the same time, since widening the network can improve the learning ability of the network, the model can better capture the complex structure and context information in the image, comprehensively utilize local detail information and overall semantic information, thereby improving the understanding ability and generalization ability of the image.

[0036] 3. The present invention applies the generative adversarial network technology to restore damaged or missing image regions. The present invention adopts a two-stage repair method, which integrates the multi-scale feature fusion strategy and the attention mechanism in the repair network, improving the visual effect and detail quality of the repair result. In the first stage, the algorithm focuses on extracting and reconstructing the edge structure information of the damaged image, ensuring the overall consistency and generalization ability of the repair result. The second stage uses this structural information to guide the final repair process, effectively improving the repair effect. The improved encoder-decoder structure and the attention mechanism are adopted, enabling the model to more accurately capture and integrate the feature information of the area to be repaired. In addition, the dual discriminator design further enhances the model's ability to distinguish global structure and local texture, and the application of spectral normalization stabilizes the training process. After being verified on multiple data sets, this method shows excellent performance and can generate highly realistic and reasonable image repair results.

[0037] 4. The present invention incorporates an improved attention mechanism into the generator structure of the texture repair network, including embedding the self-attention mechanism in the upsampling and downsampling modules to optimize the feature extraction ability of the model. The channel and pixel attention mechanisms are incorporated in the feature fusion stage to optimize the feature integration ability of the model. Introducing the attention mechanism enables the model to more efficiently integrate feature information, more accurately capture the effective features of the area to be repaired, and makes the expression of the final result in texture details more precise and reasonable.

[0038] The present invention introduces a dual discriminator structure into the discriminator structure of the texture repair network, thereby generating more realistic and accurate repair results. The global discriminator can better constrain the overall structural consistency of the repair results, and the local discriminator is used to guide the network to generate more reasonable texture details. This dual discriminator structure can improve the attention of the network to the generated images belonging to the original damaged areas in the output images of the generator, and improve the authenticity of the texture details in the generated parts. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic diagram of the network structure of the present invention;

[0040] Figure 2 are the experimental results of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to deeply explain the technical methods and their effects adopted by the present invention to achieve the established goals, the present invention will be described in detail in combination with the drawings and specific embodiments.

[0042] See Figure 1 , a method for image repair based on a structure-guided multi-scale feature fusion generative adversarial network. The overall structure diagram of the network is as shown in Figure 1 . For the input image to be repaired, first, the edge structure is extracted through the HED edge detection algorithm, and then the first-stage generative adversarial network is used to repair the edge structure to generate a complete image edge structure. The generative adversarial network in this stage is responsible for processing the high-frequency information in the image, that is, the edge structure part of the image. This part of the information is crucial for whether the overall structure of the final repair result is accurate and reasonable; then, the complete edge structure information is used as auxiliary information and input into the second-stage generative adversarial network together with the image to be repaired for the repair of the image texture details, and the repaired image is output. The generative adversarial network in this stage is responsible for processing the texture detail part of the image, and this part of the information is very important for the details and texture of the image.

[0043] Embodiment. A method for image repair based on a structure-guided multi-scale feature fusion generative adversarial network provided by the present invention, the specific steps of which include:

[0044] Step 1. A structure-guided multi-scale feature fusion generative adversarial network is composed of a sequentially connected structure repair network and a texture repair network. Both the structure repair network and the texture repair network include a generator and a discriminator, which are specifically described as follows:

[0045] (1) Structure repair network:

[0046] The generator G1 adopts a structure of downsampling block + gated convolution sequence + upsampling block. Specifically, the first layer of the generator G1 is a normalization layer, followed by two sequences of downsampling layers. The middle layer is a gated convolution sequence with residuals, which extracts features through eight gated convolutions. Finally, there are two upsampling modules.

[0047] Gated convolution is the core part of the first-stage generative adversarial network. Gated convolution determines which information should be retained or suppressed through the learned gating signal, which helps reduce the propagation of irrelevant information. Gated convolution enables the network to adaptively learn effective masks, and even after multiple feature extractions, it can still assign appropriate soft mask values to each spatial point according to the contour information and pixel position relationship. Gated convolution O y,x consists of a gated selection unit G y,x and a feature extraction unit F y,x and is composed of two parts. The implementation method is shown in formulas (10), (11), and (12), where represents the feature after downsampling.

[0048]

[0049]

[0050] O y,x = Φ(F y,x ) ⊙ σ(G y,x ) (12)

[0051] In the model calculation formula, first, the gating value g of the input feature map is calculated through g = σ(G y,x ). σ represents the sigmoid activation function, and the range of the gating value is between 0 and 1. W g is a convolution filter of learnable parameters. W m represents a multi-dilation convolution kernel used to extract features from the input image. Φ is the LeakyReLU activation function, and the gated convolution structure finally outputs the product of the feature image F y,x and the gating value.

[0052] The discriminators D1 and D2 adopt a multi-scale discriminator structure, including two spectral normalization Markov discriminator structures D1 and D2 that work at different image sizes. The specific description is as follows:

[0053] The spectral normalization Markov discriminator structure is composed of six convolutional layers stacked together. The size of the convolutional kernel in each convolutional layer is uniformly set to 5×5, and the stride is 2. Among these six convolutional layers, the number of convolutional kernels gradually increases from 64 in the first layer and reaches 256 after successive stacking. This design of increasing the number of convolutional kernels helps the network capture multi-level semantic information from low-level to high-level. Through a specific number of convolutional kernels in each layer, the network can detect and learn specific features of the input image at that level, thereby gradually building an understanding of the entire image.

[0054] The combined loss of the structure repair network consists of the L1 loss, the feature matching loss, and the hinge loss of the generator network. The combined loss function is shown in Equation (13):

[0055] L = λ 1 L rec + λ 2 L Dsn + λ 3 L Fm (13)

[0056] Among them, L rec represents the pixel-level L1 reconstruction loss, represents the spectral normalization Markov discriminator loss, and L fm is the feature matching loss. λ 1 、λ 2 and λ 3 are weighted hyperparameters that balance the L1 loss, the feature matching loss, and the hinge loss; the generator loss L G , as shown in Equation (14). The spectral normalization Markov discriminator (SN-PatchGAN) loss is shown in Equation (15):

[0057]

[0058] D sn is the Markov discriminator, the prediction structure of the generator is represented by G(z), z represents the input damaged image, and the L1 loss is used to measure the pixel-level error between the true value and the predicted value. The expression is shown in Equation (16):

[0059] L rec (x) = ||M⊙(x - F((1 - M)⊙x))|| (16)

[0060] Among them, the encoding process in the structure repair network is represented by F(·). Similar to the perceptual loss, the feature matching loss constrains the training process of the network by comparing the intermediate layer activation maps in the discriminator network. Different from the perceptual loss, the feature matching loss only compares the activation maps between specific layers in the discriminator network. The feature matching loss L FmThe calculation formula is defined as shown in Equation (17):

[0061]

[0062] Among them, the number of elements in the i-th activation function layer is represented by N i denoted as E gt represents the true complete edge structure diagram, and E pred represents the predicted edge map generated by the generation network. The activation function of the i-th selected layer in the discriminant network is denoted as

[0063] (2) Texture repair network:

[0064] The generator G2 is composed of a convolutional downsampling block, an attention mechanism block, a multi-scale dilated convolution fusion module, a multi-scale pyramid feature fusion module, and a convolutional upsampling block.

[0065] The so-called convolutional downsampling block is stacked by a convolutional layer, a normalization layer, and a ReLU activation function layer. The convolutional layer consists of a 1×1 convolution and a 3×3 convolution. Skip connections are made between the input and the output of the 3×3 convolution to fuse feature maps of different scales. Skip connections are made between the input in the first group of convolutional downsampling modules and the output of the second group of 1×1 convolutions. Skip connections are made between the output of the 3×3 convolution in the second to sixth groups of convolutional downsampling modules and the output of the second group of 1×1 convolutions in the previous group of convolutional downsampling modules. By using multiple 1×1 convolutional layers and introducing non-linear activation functions in each convolutional layer, the expression ability and non-linear mapping ability of the model are enhanced.

[0066] The so-called attention mechanism block consists of a global average pooling layer, a convolutional layer, a ReLU activation function layer, a normalization layer, and a Sigmoid activation function layer. Skip connections are made between the input of the attention mechanism module and the output of the channel attention module to supplement the input features of the pixel attention block. The output of the attention mechanism module is the feature sequence obtained by making skip connections between the feature map processed by the pixel attention block and the input feature map.

[0067] The so-called multi-scale feature pyramid fusion module and multi-scale dilated convolution fusion module are both composed of dilated convolutions with different dilation rates. Specifically:

[0068] The multi-scale dilated convolution fusion module reduces the number of channels of the input features to 64 through convolutional operations, realizing dimensionality reduction of feature representation. The dimension-reduced feature map is transmitted to three convolutional layers with different dilation rates for calculation to obtain x i (i = 1, 2, 3). Then, the feature x i (i > 1) and x i-1Add pixel by pixel, then perform feature fusion through a 3×3 convolution. The result of the fusion operation is then multiplied by the corresponding weight coefficients. The weighted output is then passed to the subsequent fusion layer, which is responsible for integrating multiple input channels to complete feature aggregation. Finally, a 1×1 convolution is used to restore the number of channels to the input channel number of the multi-scale dilated convolution fusion module and add it to the input features. Except for the last 1×1 convolution layer, instance normalization and ReLU activation function layers are performed after the convolution operation for each layer. The multi-scale fusion strategy can improve the network's ability to capture details at different scales, enabling the network to better understand and extract complex patterns and detailed information in the image, and grasp the texture details of the missing area more accurately.

[0069] The multi-scale pyramid feature fusion module uses the deep features containing context-aware information from the first three layers in front of the self-attention layer as input, and its main role is to supplement the input features of the self-attention mechanism layer. By integrating multi-scale features from different depth convolutional layers, the multi-scale pyramid feature fusion module can provide a more comprehensive feature representation, which is crucial for capturing the details and context information of the damaged area. By using convolutional kernels with different dilation rates, image features within different receptive fields can be captured, and feature maps of different scales can be obtained. After cross-channel connection, these feature maps form a comprehensive feature representation and retain the important semantic information in the original feature maps.

[0070] The so-called convolutional upsampling block is composed of a 1×1 convolution, a transposed convolution, a self-attention feature module, and a 3×3 convolution connection, corresponding to the convolutional downsampling module. There are six groups of convolutional upsampling modules in the network. The input of the upsampling module is combined by the output of the previous upsampling module and the output of the second group of 3×3 convolutional layers in the corresponding downsampling module. Skip connections are made between the downsampling module and the corresponding upsampling layer to supplement some shallow information lost during the feature extraction process to the decoding stage, enabling the decoding network to retain more original details. The input of the second group of 1×1 convolutions in the upsampling module is composed of skip connections between the output of the self-attention mechanism module and the output of the transposed convolution layer. To enhance the stability of the model and reduce the computational load, batch normalization is used after the convolutional layer to normalize the input data, stabilize the data distribution, and alleviate the problem of gradient dispersion that may occur during training. At the same time, the ReLU activation function can introduce non-linear transformation to improve the expressive ability of the model.

[0071] The discriminators D3 and D4 of the texture repair network consist of two global and local discriminators with the same structure. The two discriminators have the same structure but different inputs. The discriminator network uses convolutional layers to extract the features of the pictures. After the convolutional layers, spectral normalization and the LeakyReLU activation function are used. In addition, spectral normalization is introduced into the discriminator of the texture generation network, effectively reducing the occurrence of overfitting.

[0072] The loss function of the texture repair network consists of pixel reconstruction loss, perceptual loss, style loss, and adversarial loss function:

[0073] For the real image sample x, the encoder sampling process is represented by F(·), and M is the binary image mask. The pixel reconstruction loss L rec is used to calculate the L2 distance between the generated image and the ground truth. The L2 loss function is differentiable, which means it can be used in optimization algorithms such as gradient descent so that the network can learn to minimize this loss. The calculation formula is shown in Equation (18):

[0074]

[0075] To avoid the problem of blurry repair results caused by using a single pixel reconstruction loss, perceptual loss L prec is introduced. The perceptual loss ensures the similarity of multi-scale features between low-resolution images and high-resolution images by comparing the consistency between feature representations at different levels. The specific definition is shown in Equation (19):

[0076]

[0077] where φ i is the activation map of the i-th layer of the pre-trained network VGG19, and N i is the number of elements in the activation map of the i-th layer.

[0078] Since pixel filling is performed on the damaged area during the texture repair stage, to solve the style consistency between the generated pixels and the pixels in the non-missing area and better integrate the style in the original image into the generated area, style loss is introduced. The style loss is used to measure the L1 distance between the Gram matrices of the i-th layer deep features of the generated image and the real image. Given a feature map of size C i ×H i ×W i The expression of the style loss function is shown in Equation (20):

[0079]

[0080] where is constructed from the activation map of the i-th layer and is of size C i ×Ci The Gram matrix. The final style loss function of the texture repair network consists of the global discriminator style loss L globalstyle and the local discriminator style loss L localstyle and is composed of two parts. The specific expression is shown in Equation (21):

[0081] L style = λ g L globalstyle + λ l L localstyle (21)

[0082] where λ g and λ l are hyperparameters for balancing different losses, and the ratio is 3:2.

[0083] The adversarial loss L adv mainly focuses on the high-frequency details of the image and is calculated through the cross-entropy loss function

[52] The formula is shown in Equation (22):

[0084] L αdν = E[log(D(I in ))] + E[log(1 - D(C out )))] (22)

[0085] Finally, by jointly using the above loss functions, the total loss function L is obtained, which is specifically expressed as Equation (23):

[0086] L = λ 1 L rec + λ 2 L prec + λ 3 L style + λ 4 L adv (23)

[0087] where: L rec is the pixel reconstruction loss, L prec is the perceptual loss, L adv is the adversarial loss, and λ 1 , λ 2 , λ 3 , λ 4 are adjustment parameters.

[0088] Step 2: Obtain training images and preprocess the obtained training images, including the following steps: Randomly generate masks of any size and shape, and randomly superimpose the masks on the training images to obtain damaged training images.

[0089] In this embodiment, the Place Step 2 dataset is used to verify the repair effect of the network model, and 100,000 images are selected for training the network model. Among them, 90,000 images are selected for training the structure repair network and the texture repair network, and 10,000 image samples are selected for the test set.

[0090] Considering that the images included in the Place Step 2 dataset are all original undamaged images, while the method proposed by the present invention aims to process damaged image inputs. Therefore, in order to simulate the damage conditions in the actual application scenario and effectively train the network, a preprocessing step of superimposing masks is performed on the original training images. An irregular mask dataset is used to simulate the missing areas of the damaged images, thereby forming a damaged image dataset suitable for network training. Ensure that the generative adversarial network can learn an effective mapping from the damaged images to the original undamaged images during training, thereby improving its ability and accuracy in processing real damaged images in actual applications. During the training phase, the masks are transformed diversely, including random rotation, horizontal flipping, and vertical flipping, to improve the model's object recognition ability in different directions, increase the model's adaptability to perspective changes, and at the same time, achieve the purpose of expanding the dataset.

[0091] Step 3: Input the preprocessed training images into the structure-guided multi-scale feature fusion generative adversarial network to train the network.

[0092] During the training phase, first pre-train the structure repair network, and finely adjust the network parameters through 10,000 iterations, focusing on restoring the edge and structure information of the damaged images. Subsequently, the texture repair network undergoes pre-training for 30,000 iterations to capture and reconstruct local texture features. After that, the algorithm enters the joint training phase, where the two generative networks and the discriminative network are trained simultaneously, and the two-stage networks are alternately trained for a total of 50,000 iterations to ensure that the model can achieve the best balance between global structure consistency and local texture quality. The initial value of the learning rate is set to 10 -4 To avoid overfitting and oscillation phenomena, when the model approaches the optimal solution, the learning rate is adjusted to 10 -5 . The first-order momentum of the Adam optimizer is set to 0.5, and the second-order momentum is set to 0.9. Set the weighted hyperparameters of the structure repair network loss function to λ 1 =1, λ 2 =1, λ 3 =10, and the weighted hyperparameters in the texture repair network are respectively set to λ 1 =1, λ 2 =0.1, λ 3 =250, λ 4 =0.2.

[0093] Specifically, it includes the following steps:

[0094] 3.1 Edge detection: Extract the edge contour structure of the image to be repaired through an edge detection algorithm, obtain the edge structure information of the damaged image, and get the damaged edge structure diagram:

[0095] First, preprocess the damaged image by grayscale processing the RGB input image I input to obtain a grayscale image I gray . The grayscale image has 1 channel. Then, extract the edge structure of the damaged image through the overall nested edge detection algorithm to obtain an incomplete structure diagram E edge′ .

[0096] 3.2 Structure repair: Input the extracted incomplete edge contour structure, mask image, and image to be repaired into the structure repair network for edge structure repair to obtain a complete edge structure image.

[0097] As shown in formula (5), in the first stage, the incomplete structure diagram E edge′ extracted in step 3.1, the mask image M, and the grayscale image I gray are used as the input E input to the structure repair network. The mask image is the shape of the damaged area, with pixel values of 1 in the missing area and 0 in the remaining parts. The network in this stage is a generative adversarial network structure, consisting of a generator network and a discriminator network. The discriminator network consists of two spectrally normalized Markov discriminators working at different scales. The generator and the discriminator are trained through adversarial games, and finally, the generator can generate a complete edge structure diagram.

[0098] E input ={I gray , E edge′ , M} (1)

[0099] The generator network G edge and the discriminator network D edge are trained adversarially to reach the "Nash equilibrium" state. Finally, the generator can generate a complete edge structure diagram E edge that conforms to the semantic structure.

[0100] E edge =G edge (E input ) (2)

[0101] 3.3 Texture repair: Input the complete edge structure image and the image to be repaired into the texture repair network to restore the texture details and obtain the final repair result.

[0102] For the predicted structural part in the predicted edge structure image output in step 3.2 of the second-stage network and the incomplete structure diagram E in the first stageedge′ The complete edge structure diagram E composed of splicing the images of the undamaged parts in edge . As shown in formula (8), the overall input T of the network at this stage input is the damaged image I in and the complete edge image E edge spliced together in the channel dimension to form a combined input.

[0103]

[0104] T input ={E edge , I in} (4)

[0105] As shown in formula (9), the generation network G texture and the discriminator network D texture are trained adversarially to reach the "Nash equilibrium" state, and finally the generator can generate a repaired result E with clear texture and natural continuity result .

[0106] E result = G texture (T input ) #(9)

[0107] The Adam optimization algorithm is used to train the parameters in the structure-guided multi-scale feature fusion generative adversarial network.

[0108] The training process of the network is specifically described as follows: The training of the entire model is divided into two parts: pre-training and joint training.

[0109]

[0110] The structure repair network training algorithm is as shown in the above table. First, the discriminator network is trained, and then the discriminator is fixed and the structure generation network is trained. In each round of training, m training samples are randomly selected from the dataset, and then these samples are preprocessed. The edge structure of the samples is extracted using object detection, and then it is grayscaled. Then, the extracted edge structure diagram and grayscale image are occluded with an irregular mask to simulate the degradation process of the image. After preprocessing, m damaged edge structure diagrams and grayscale images are obtained. Subsequently, these samples are input into the structure repair network for complete edge structure prediction, and the generated complete structure diagrams will be verified for true and false values through the discriminator network. The discriminator network trains the model to update the discriminator parameters with feature matching loss, minimizing adversarial loss, and L1 reconstruction loss. Every k rounds of training of the discriminator network, the structure generation network is trained once to update the parameters of the generator.

[0111]

[0112] The texture repair network also adopts the process of first training the discriminator network and then training the generator network. The algorithm process of the texture repair network is shown in the above table. First, m complete edge structure diagrams repaired by the structure repair network and the original image with a mask are input into the generator of the texture repair network to generate m repaired images. Then, the ground truth and the images generated by the texture repair network are jointly fed into the discriminator, and the parameters of the discriminator are updated with multiple loss functions such as pixel reconstruction loss, perceptual loss, style loss, and adversarial loss. After k rounds of training of the discriminator network are completed, the m images that have completed structure repair by the structure repair network are input into the generator of the texture repair network to train and update the network parameters of the generator.

[0113] In the network training process of each stage in the joint training, it is the same as the above training process. In this process, the structure generator and the texture generator will work together to try to generate generated images that the texture discriminator cannot distinguish between true and false, while the texture discriminator tries to distinguish between real data samples and generated data samples. At the same time, the structure discriminator will be used to further optimize the output of the structure generator to make it closer to the distribution of real data. During the joint training process, the structure generator and the texture generator will also be alternately trained to help the generator in each stage focus on its respective repair tasks.

[0114] After the training is completed, the algorithm saves the parameters of the generator and the discriminator, including the training configuration, structure, weights of the network model, and the state of the optimizer, for reuse in future repair tasks.

[0115] Step 4: Input the image to be repaired into the trained structure-guided multi-scale feature fusion generative adversarial network as described above to obtain the final repaired image.

[0116] In this step, first, edge detection is performed on the damaged image to extract the edge structure of the damaged image. Then, the detected incomplete edge image is input into the structure repair network for processing to obtain a complete edge image. The texture repair network then processes the image to be repaired and the complete edge image to obtain the final repaired image. See Figure 2 , (a) original image, (b) damaged image, (c) repair, (d) complete repaired image.

[0117] Table 1 and Table 2 respectively show the evaluation metrics of the image repair results based on the Places2 and CelebA datasets.

[0118] Table 1 Evaluation Metrics of Image Repair Results Based on the Places2 Dataset

[0119]

[0120] Table 2 Evaluation Metrics for Image Inpainting Results Based on the CelebA Dataset

[0121]

[0122] In summary, an attention mechanism module, a multi-scale dilated convolution fusion module, and a multi-scale pyramid feature fusion module in the texture inpainting network are designed based on the attention mechanism and the multi-scale feature fusion strategy, optimizing the feature extraction process of the network. At the same time, the loss function of the generative adversarial network is reconstructed, and pixel reconstruction loss, perceptual loss, and style loss are introduced into the loss function of the texture inpainting network. This makes the semantics of the inpainted image more reasonable and the texture details clearer. By combining the pixel reconstruction loss, the network can learn how to reduce the difference between the generated image and the original image at the pixel level, thus directly optimizing the details and clarity of the image. The perceptual loss helps to ensure that the generated image is more visually similar to the real image. The style loss can help the network capture and replicate the style features of the original image, enhancing the consistency of the inpainting structure. In addition, the global and local double discriminator structure in the texture inpainting stage also ensures the consistency between the inpainted area and the known area, thus significantly improving the overall inpainting quality of the image. Especially when dealing with images with a relatively small proportion of damaged areas (less than 30%), the proposed method shows a more significant improvement in the inpainting effect.

[0123] For those skilled in the art, it is obvious that the present invention is not limited solely to the detailed description of the above exemplary embodiments. Without departing from the spirit and basic features of the present invention, the present invention can be implemented in other specific forms. Therefore, it is intended to include all possible variations within the meaning and scope of the equivalent elements in the claims within the scope of the present invention.

Claims

1. A structure-guided multi-scale feature fusion generative adversarial network image restoration method, characterized in that: The steps include: Step 1: A structure-guided multi-scale feature fusion generative adversarial network is formed by sequentially connecting a structure restoration network and a texture restoration network, wherein both the structure restoration network and the texture restoration network include a generator and a discriminator; Step 2: obtaining a training image and preprocessing the obtained training image; Step 3: Input the preprocessed training image into a structure-guided multi-scale feature fusion generative adversarial network to train the network, including the following steps: 3.1 Edge detection stage: The edge structure of the input damaged image is extracted by using the HED edge detection algorithm. First, the damaged image should be preprocessed and the RGB input image I input Perform grayscale processing to obtain grayscale image I gray , the channel of the grayscale image is 1, and then the edge structure of the damaged image is extracted by the overall nested edge detection algorithm to obtain the incomplete structure image E edge' ; 3.2 Structural repair stage, incomplete structural graph E extracted in the first stage edge′ Mask image M and grayscale image I gray As the input of the structure repair network E input , the mask image is the shape of the damaged area, the pixel value of the missing area is 1, and the pixel value of the rest is 0; the network at this stage is a generative adversarial network structure, which consists of a generator network and a discriminator network; the discriminator network consists of two spectral normalized Markov discriminators working at different scales, the generator and the discriminator are trained in adversarial games, and finally the generator generates a complete edge structure map HAVE BEEN input ={I gray ,HAVE BEEN edge′ ,M} (1) Generator network and discriminator network D edg Adversarial training is performed to reach the "Nash equilibrium" state, and the generator finally generates a complete edge structure graph E that conforms to the semantic structure. edge E edge =G edge (E input ) (2) 3.3 Texture restoration stage, structure diagram E edge′ The predicted structure part in the predicted edge structure image output by the second stage network and the incomplete structure image E in the first stage edge′ The complete edge structure diagram is composed of the undamaged part of the image. The overall input of the network at this stage is T input is the damaged image I in and the complete edge image E edge The joint input is concatenated in the channel dimension T input ={E edge ,I in } (4) Step 4: input the image to be repaired into the trained structure-guided multi-scale feature fusion generative adversarial network to obtain the final repaired image.

2. According to claim 1, the image restoration method based on structure-guided multi-scale feature fusion generative adversarial network is characterized by: In the step 1, the generator G1 in the structure repair network adopts a structure of downsampling block + gated convolution sequence + upsampling block; The generator G2 of the texture restoration network consists of a convolutional downsampling block, an attention mechanism module, a multi-scale dilated convolution fusion module, a multi-scale pyramid feature fusion module and a convolutional upsampling block.

3. An image restoration method based on structure-guided multi-scale feature fusion generative adversarial network according to claim 2, characterized in that: The joint loss of the structure repair network consists of L1 loss, feature matching loss and hinge loss of the generative network, as shown in formula (13): Among them, L rec represents the pixel-level L1 reconstruction loss, represents the spectral normalized Markov discriminator loss, L fm is the feature matching loss, and λ1, λ2, and λ3 are weighted hyperparameters that balance the L1 loss, feature matching loss, and hinge loss.

4. An image restoration method based on structure-guided multi-scale feature fusion generative adversarial network according to claim 2 or 3, characterized in that: The attention mechanism module in the generator G2 of the texture restoration network consists of a global average pooling layer, a convolution layer, a ReLU activation function layer, a normalization layer, and a Sigmoid activation function layer.

5. The image restoration method based on structure-guided multi-scale feature fusion generative adversarial network according to claim 4, characterized in that: The discriminators D3 and D4 of the texture inpainting network are composed of two global and local discriminators with the same structure.

6. According to the image restoration method based on structure-guided multi-scale feature fusion generative adversarial network as claimed in claim 5, the loss function of the texture restoration network consists of pixel reconstruction loss, perceptual loss, style loss and adversarial loss loss functions, and the total loss function L is expressed as formula (23): L=λ1L rec +λ2L prec +λ3L style +λ4L adv (23) in: L rec is the pixel reconstruction loss, L prec is the perceptual loss, L adv To combat the loss, λ1, λ2, λ3, and λ4 are adjustment parameters.

Citation Information

Patent Citations

  • Image restoration method based on gated convolution generative adversarial network

    CN111968053A

  • Image restoration method based on multi-scale content attention mechanism, storage medium and terminal

    CN112884669A

  • Image restoration method based on U-Net perceptual adversarial network

    CN117994139A

  • Global and local feature reconstruction network-based medical image segmentation method

    US20230274531A1

Cited By

  • Wound repair effect prediction method based on generative adversarial network

    CN120598821A

  • Virtual image restoration method and system based on image recognition

    CN121213424A

  • High-fidelity image restoration method based on mask perception attention network

    CN121414628A

  • Visible light water surface target detection method based on sea level constraint and hierarchical generative attention

    CN121640183A

  • Light-weight portrait wrinkle removal image enhancement method based on generative adversarial network

    CN121660910A