Face image inpainting method based on double-flow gate convolutional network

CN115272126BActive Publication Date: 2026-09-25CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210942055.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2026-09-25
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

这些方法在修复过程中可以生成合理的视觉结构和纹理,但由于卷积操作的不合理导致修复的结果常常生成与已知区域不一致的纹理细节

Benefits of technology

[0023]本发明采用一个带有批量归一化的双流门控卷积网络GConv-UNet,将其中UNet的每个普通卷积层替换为门控卷积层,并使用结构和纹理信息指导彼此特征的重建,尽可能生成真实的样本,生成器所生成的修复好的图像,将其和真实的图片一起送入判别器中进行判别,二者相互对抗,生成器不断的进行图像生成学习,直至判别器无法区分生成器修复好的图像与真实图像为止,提升了生成器拟合数据的能力,实现人脸图像修复,并且引入了通道级特征均衡对通道之间的关系建模来增强结构和纹理特征之间的一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272126B_ABST
    Figure CN115272126B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of image restoration, and particularly relates to a face image restoration method based on a double-flow gated convolutional network, comprising the following steps: feature generation, inputting a face image to be restored to generate complete structure features and texture features; step 200: feature reconstruction, the structure features and the texture features guiding and constraining the reconstruction of each other's features; step 300: channel-level feature balancing, modeling the relationship between channels to enhance the consistency between the structure features and the texture features, and obtaining a restored image; step 400: restoration discrimination, inputting the generator-restored image and a real picture into a discriminator for discrimination to realize face image restoration; the method proposed in the present application fully utilizes the relationship between structure and texture information to complete the guidance and constraint of each other in the structure feature and texture feature generation process, and in addition, a channel-level feature balancing method is introduced to improve the overall consistency of the restoration result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image restoration technology, specifically relating to a face image restoration method based on a dual-stream gated convolutional network. Background Technology

[0002] Image inpainting aims to reconstruct missing or damaged regions of an image based on surrounding known content, ensuring the repaired area remains consistent with the overall image. Face inpainting, as a crucial branch of this field, plays a vital role in practical applications. With the development of deep learning, this technology has achieved remarkable results in image inpainting in recent years. Deep generative methods using structural information as prior knowledge have demonstrated good performance in repairing damaged face images. These methods can generate reasonable visual structures and textures during the repair process, but improper convolutional operations often result in texture details inconsistent with the known regions. To address this issue, we propose a dual-stream gated convolutional network called GConv-UNet. During the generation of structural and texture features, it fully utilizes the relationship between structural and texture information to guide and constrain each other. Furthermore, we introduce a channel-level feature equalization method to improve the overall consistency of the repair results. Summary of the Invention

[0003] The purpose of this invention is to provide a face image restoration method based on a dual-stream gated convolutional network to address the problems existing in the background technology.

[0004] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows:

[0005] A face image inpainting method based on a two-stream gated convolutional network includes the following steps:

[0006] Step 100: Feature generation. Input the face image to be repaired. The generator adopts a two-stream network structure and generates complete structural and texture features through the encoder-decoder part of the generator.

[0007] Step 200: Feature Reconstruction, the bidirectional feature fusion module of the generator's feature fusion part, enables structural features and texture features to guide and constrain the reconstruction of each other's features;

[0008] The context feature aggregation module of the generator's feature fusion part produces more vivid details in the reconstruction results of structural and texture features by modeling the spatial dependencies.

[0009] Step 300: Channel-level feature equalization, which enhances the consistency between structural and texture features by modeling the relationship between channels, resulting in a repaired image;

[0010] Step 400: Repair and discrimination. The image repaired by the generator and the real image are input into the discriminator for discrimination. The two compete against each other. The generator continuously learns to generate images until the discriminator can no longer distinguish between the image repaired by the generator and the real image, thus realizing face image repair.

[0011] In step 100, the generator's dual-stream network uses GConv-UNet as its backbone and replaces each convolutional layer in UNet with gated convolutions to improve repair quality and color consistency.

[0012] The bidirectional feature fusion module of the generator's feature fusion section in step 200, which enables structural features and texture features to guide and constrain the reconstruction of each other's features, includes the following steps:

[0013] Step 210: The damaged face image and mask are input into the texture encoder, and the corresponding damaged edges, damaged grayscale image and mask are input into the structure encoder;

[0014] Step 220: Supplement the features of the texture encoder to the structure decoder through skip connections, and supplement the features of the structure encoder to the texture decoder, so that the texture features and structure features guide and constrain each other's feature reconstruction.

[0015] Step 300 includes the following steps:

[0016] Step 310: Introduce the SE module. The fused features are input into the SE module. The SE module then uses squeezing and excitation operations to balance the texture and structural features in the channel.

[0017] Step 310 specifically includes the following steps:

[0018] Step 311: Perform a compression operation on the original feature map to obtain global features at the channel level;

[0019] Step 312: Perform an excitation operation on the obtained global features, learn the relationship between each channel to obtain the weights of different channels, and finally multiply by the original feature map to obtain the final feature map.

[0020] The discriminator in step 400 is a two-stream discriminator with texture branch and structure branch.

[0021] The discriminator's structural branch also has an additional edge detector for edge extraction.

[0022] In step 400, adversarial loss, perceptual loss, and style loss are also introduced. The discriminator uses the loss function to distinguish between the image restored by the generator and the real image.

[0023] This invention employs a dual-stream gated convolutional network GConv-UNet with batch normalization, replacing each ordinary convolutional layer of the UNet with a gated convolutional layer. It uses structural and texture information to guide the reconstruction of features, generating samples as realistic as possible. The restored images generated by the generator are then fed into a discriminator along with real images for comparison. The two systems work against each other, with the generator continuously learning to generate images until the discriminator can no longer distinguish between the restored images and real images. This improves the generator's ability to fit data, enabling face image restoration. Furthermore, channel-level feature equalization is introduced to model the relationships between channels, enhancing the consistency between structural and texture features. Attached Figure Description

[0024] The present invention can be further illustrated by the non-limiting embodiments given in the accompanying drawings.

[0025] Figure 1 This is a schematic diagram of the main process of a face image restoration method based on a dual-stream gated convolutional network according to Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart illustrating the generator encoding-decoding part and the feature fusion part in Embodiment 1 of the present invention;

[0027] Figure 3 This is a flowchart illustrating the SE module in the channel-level feature equalization of the present invention in Embodiment 1.

[0028] Figure 4 This is a schematic diagram illustrating the effect of qualitative analysis in Embodiment 2 of the present invention;

[0029] Figure 5 This is a schematic diagram of the effect of GConv-UNet in the ablation experiment of Embodiment 2 of the present invention;

[0030] Figure 6 This is a schematic diagram illustrating the effect of channel-level feature equalization in the ablation experiment of Embodiment 2 of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0032] Example 1:

[0033] like Figure 1-3 The method for face image inpainting based on a two-stream gated convolutional network, as shown, includes the following steps:

[0034] Step 100: Feature generation. Input the face image to be repaired. The generator adopts a two-stream network structure and generates complete structural and texture features through the encoder-decoder part of the generator.

[0035] In step 100, the generator's dual-stream network uses GConv-UNet as its backbone and replaces each convolutional layer in UNet with gated convolutions to improve repair quality and color consistency.

[0036] Specifically, unknown regions in the image are considered invalid pixels, while known regions are considered valid pixels. Gated convolution helps improve restoration details and color consistency, especially for image restoration with irregular damaged areas. Furthermore, gated convolution has a flexible mask update mechanism; unlike hard gating, it can automatically learn soft masks from the data, maintaining even deep mask regions. Additionally, we add batch normalization after each gated convolutional layer to prevent gradient vanishing during training. This operation can be represented as:

[0037] Gating=∑∑W g ·I, #

[0038] Feature = ∑∑W f ·I, #

[0039]

[0040] in,

[0041] I represents the feature map; Gating represents the gate; Feature represents the convolutional feature map; W g and W f These represent different convolutional kernels; φ is the LeakyReLU activation function, and σ is the Sigmoid activation function. This indicates element-wise multiplication, so the gate value is between 0 and 1. The closer the gate value is to 1, the more effective pixels there are.

[0042] BN(*) represents batch normalization.

[0043] Step 200: Feature Reconstruction. The bidirectional feature fusion module of the generator's feature fusion section enables structural and texture features to guide and constrain the reconstruction of each other's features, including the following steps:

[0044] Step 210: The damaged face image and mask are input into the texture encoder, and the corresponding damaged edges, damaged grayscale image and mask are input into the structure encoder;

[0045] Step 220: Supplement the features of the texture encoder to the structure decoder through skip connections, and supplement the features of the structure encoder to the texture decoder, so that the texture features and structure features guide and constrain each other's feature reconstruction.

[0046] Meanwhile, the context feature aggregation module of the generator's feature fusion part produces more vivid details in the reconstruction results of structural and texture features by modeling the spatial dependencies.

[0047] Specifically, the final generated structural features are formed under the constraints of texture features, while the final generated texture features are generated under the guidance of structural features. Therefore, in order to supervise the texture and structural features generated by the two decoders, we introduce a structural loss L0. structure and texture loss L texture As a constraint, it is as follows:

[0048] L structure =BCE(E gt P s (F s ));

[0049] L texture =l1(I gt P t (F t ));

[0050] I gt For real images; E gt For the true edge; P s and P t The mapping function, composed of convolutional kernel residual blocks, transforms the structural features F s and texture features F t These are mapped to the corresponding edge maps and RGB images, respectively.

[0051] L structure Represents structural loss, used to calculate the generation of true edges E. gt The binary cross-entropy between the generated edge graph and the edge graph;

[0052] L texture Represents texture loss, used to calculate the true image I. gt The distance between the RGB image and its l1 value.

[0053] The feature fusion part fuses structural and texture features through a bidirectional feature fusion module. The context feature aggregation module produces more vivid details by modeling long-term spatial dependencies. However, this fusion operation only focuses on the consistency of spatial features and ignores the dependency of channel information. In order to improve the consistency of texture and structural information after fusion, we added a channel-level feature equalization method to the feature fusion part, as described in step 300.

[0054] Step 300: Channel-level feature equalization, which enhances the consistency between structural and texture features by modeling the relationship between channels, resulting in a repaired image;

[0055] Step 310: Introduce the SE module. The fused features are input into the SE module. The SE module balances the texture and structural features in the channel through squeezing and excitation operations.

[0056] Step 311: Perform a compression operation on the original feature map to obtain global features at the channel level;

[0057] Step 312: Perform an excitation operation on the obtained global features, learn the relationship between each channel to obtain the weights of different channels, and finally multiply by the original feature map to obtain the final feature map.

[0058] The overall structure of the SE module is as follows: Figure 3 As shown, X represents the input feature map, X' represents the output feature map, c represents the number of channels, and h×w represents the feature size.

[0059] Specifically, such as Figure 3 As shown, the squeezing operation is implemented using global pooling. During the excitation operation, two fully connected layers form a bottleneck, and the importance of each channel is predicted by the sigmoid function. In this process, the number of channels is reduced to 1 / 16 of its original size, and then restored to its original size through the fully connected layers and the ReLU function. This connection between two fully connected layers has more nonlinearity than a direct fully connected layer, better adapts to the complex inter-channel correlations, and reduces the number of computations and parameters.

[0060] Step 400: Repair and discrimination. The image repaired by the generator and the real image are input into the discriminator for discrimination. The two compete against each other. The generator continuously learns to generate images until the discriminator can no longer distinguish between the image repaired by the generator and the real image, thus realizing face image repair.

[0061] In step 400, the discriminator is a two-stream discriminator with texture branch and structure branch, and the structure branch of the discriminator also has an additional edge detector for edge extraction.

[0062] In step 400, adversarial loss, perceptual loss, and style loss are introduced. The discriminator uses the loss function to distinguish between the image restored by the generator and the real image.

[0063] Specifically, the perceptual loss L, pre-trained on ImageNet by VGG-16 prec The formula used to simulate human visual perception of image quality is shown below:

[0064]

[0065] in, Let I be the activation map of the i-th pooling layer in VGG-16, where IE represents the expectation and I is the activation map of the i-th pooling layer in VGG-16. out For the model repaired result, I gt For real images, ||x||1 represents the L1 norm of x. In practice, i∈[1,3].

[0066] Style loss L style With perceived loss L prec Similarly, the method for calculating style loss is as follows:

[0067]

[0068]

[0069] in, Let ||x||1 represent the Gram matrix corresponding to the feature map, and ||x||1 represent the L1 norm of x.

[0070] Adversarial loss is used to ensure the consistency of the reconstructed image, and its definition is as follows:

[0071]

[0072] Where IE represents expectation, I out For the model's corrected result, I gt For real images, E out E represents the edge of the model repair result. gt The edges of a real image.

[0073] In addition, we also introduce the void loss L hole and effective loss L valid These are used to calculate the l1 distance between the damaged and undamaged areas, respectively.

[0074]

[0075]

[0076] Where M represents the mask, which consists of 1 and 0, where 1 represents the known region and 0 represents the damaged region, and ||x||1 represents the L1 norm of x.

[0077] In summary, the overall loss function formula is as follows:

[0078] L total =λ prec L prec +λ style L style +λ adv L adv +λ hole L hole +λvalid L valid +λ structure L structure +λ texture L texture

[0079] Where, λ prec , λ style , λ adv , λ hole , λ valid , λ structure and λ texture These represent the calculation parameters for the corresponding losses.

[0080] The discriminator compares the generator-restored image with the real image using an overall loss function until it can no longer distinguish between the generator-restored image and the real image.

[0081] Example 2:

[0082] This invention also provides an experimental and analytical method for face image inpainting based on a two-stream gated convolutional network:

[0083] I. Experimental Conditions

[0084] The experimental setup used an NVIDIA RTX 3060 Ti GPU (8GB) and an i5-10400F 2.90GHz CPU. Initially, we trained the model using a learning rate of 0.0002, and then fine-tuned it using 0.00005. For the loss parameters, we set λprec = 0.1, λsytle = 250, λadv = 0.1, λhole = 60, λvalid = 10, λsturcture = 1, and λtexture = 1.

[0085] When training the model, we set the batch size to 3. The dataset used was the CelebA-HQ face dataset and NVIDIA irregular mask. The input images were uniformly cropped to 256*256 pixels. The entire model was implemented in PyTorch and optimized with Adam.

[0086] We qualitatively and quantitatively compared our proposed method with several state-of-the-art methods, namely DeepFillv2, EC, PRVS, MED, RFR, and CTSD.

[0087] II. Qualitative Analysis

[0088] The results of qualitative comparison with state-of-the-art methods are as follows: Figure 4 As shown:

[0089] The first column from the left shows the damaged images, and the second column from the left shows the actual images.

[0090] Using DeepFillv2, such as Figure 4 As shown in the third column from the left, DeepFillv2 produces overly smoothed content, and the semantics are only reasonable when the mask ratio is low;

[0091] Using EC, such as Figure 4 As shown in the fourth column from the left, when the mask rate is large, the result generated by EC contains a distorted structure;

[0092] Using RFRNet, such as Figure 4 As shown in the seventh column from the left, when the damaged area is small, RFRNet cannot fill in all the missing pixels, although it can generate a reasonable facial structure under a larger mask.

[0093] Using PRVS, MED and CTSDG respectively as follows Figure 4 As shown in the fifth, sixth, and eighth columns from the left, a reasonable structure can usually be produced, but the generated texture details are blurry.

[0094] Using our method, such as Figure 4 As shown in the last column from the left, our method produces more realistic visual effects in terms of structure and texture details under different masking ratios. Compared with other methods, our method produces more realistic visual results in terms of structure and texture details and is closer to the actual ground.

[0095] III. Quantitative Analysis

[0096] Experiments were conducted on the CelebA-HQ dataset, using masks with different proportions ranging from 10% to 60% to represent the size of the damaged area, and the generated results were objectively analyzed. The main evaluation metrics were PSNR, SSIM, and As shown in the table below, our method outperforms all the comparison methods across all three metrics and produces better results. For a fair comparison, none of the results repaired by the model underwent post-processing.

[0097]

[0098]

[0099] Quantitative analysis of experimental results using CelebA-HQ and NVIDIA datasets and state-of-the-art methods (↑ indicates larger is better; ↓ indicates smaller is better).

[0100] IV. Ablation Experiment

[0101] We conducted some ablation studies to verify the design effectiveness in our model, mainly validating the effectiveness of GConv-UNet, channel-level feature equalization, and normalization methods.

[0102] ①GConv-UNet

[0103] The soft gating mechanism in gated convolutional layers is more flexible during mask updates, resulting in more visually plausible details in the repaired images. To validate the effectiveness of GConv-UNet, we used a U-net variant with a partial convolutional layer backbone, keeping other conditions unchanged. Experiments demonstrate that our GConv-UNet produces high-quality results, as shown in the table below. Our model can repair more vivid details, such as... Figure 5 The images show facial organs (from left to right: the first image is a damaged image, the second is a real image, the third is a partial convolution result, and the fourth is the gated convolution result of this method).

[0104]

[0105] Quantitative analysis of gated convolution ablation experiments (↑ indicates larger is better; ↓ indicates smaller is better).

[0106] ② Channel-level feature equalization

[0107] A channel-level feature equalization method is introduced to balance the texture and structural information between channels. Quantitative comparisons are shown in the table below. The SE module can significantly improve performance when the damaged area is large. Figure 6 As shown (from left to right, the first image is the damaged image, the second is the real image, the third is a partial convolution result, and the fourth is the gated convolution result of this method), the generated results show good overall consistency, such as the texture of the eyes and mouth of the face in the fourth image from the left.

[0108]

[0109] Quantitative analysis of ablation experiments using the channel-level feature equalization method (↑ indicates larger is better; ↓ indicates smaller is better).

[0110] ③ Normalization method

[0111] Batch normalization is introduced to prevent gradient vanishing during training. Without it, the model would fail to learn image features after 200,000 training iterations. To verify the effectiveness of batch normalization in training our model, we replaced it with different normalization methods for comparison: region normalization and instance normalization. Region normalization calculates the mean and variance for damaged and undamaged portions of the image, respectively. However, region normalization is only suitable for networks where there is a clear distinction between damaged and undamaged regions. Gated convolutions can smooth image features, making it difficult for region normalization to track potentially damaged regions. Furthermore, we observe that instance normalization leads to color inconsistencies in the results.

[0112]

[0113] Quantitative analysis of batch normalized ablation experiments (↑ indicates higher is better; ↓ indicates lower is better).

[0114] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A face image inpainting method based on a two-stream gated convolutional network, characterized in that: Includes the following steps: Step 100: Feature generation. Input the face image to be repaired. The generator adopts a two-stream network structure and generates complete structural and texture features through the encoder-decoder part of the generator. In step 100, the generator's dual-stream network uses GConv-UNet as its backbone, and gated convolutions are used to replace each convolutional layer in UNet to improve the repair quality and color consistency. Batch normalization is added after each gated convolutional layer to prevent gradient vanishing during training. Step 200: Feature Reconstruction. The bidirectional feature fusion module of the generator's feature fusion section enables structural and texture features to guide and constrain the reconstruction of each other's features, including the following steps: Step 210: The damaged face image and mask are input into the texture encoder, and the corresponding damaged edges, damaged grayscale image and mask are input into the structure encoder; Step 220: Supplement the features of the texture encoder to the structure decoder through skip connections, and supplement the features of the structure encoder to the texture decoder, so that the texture features and structure features guide and constrain each other's feature reconstruction; The context feature aggregation module of the generator's feature fusion part produces more vivid details in the reconstruction results of structural and texture features by modeling the spatial dependencies. Step 300: Channel-level feature equalization, which enhances the consistency between structural and texture features by modeling the relationship between channels, resulting in a repaired image; Step 300 includes the following steps: Step 310: Introduce the SE module. The fused features are input into the SE module. The SE module then performs squeezing and excitation operations to equalize the texture and structural features in the channels, including the following steps: Step 311: Perform a compression operation on the original feature map to obtain global features at the channel level; Step 312: Perform an excitation operation on the obtained global features, learn the relationship between each channel to obtain the weights of different channels, and finally multiply by the original feature map to obtain the final feature map; Step 400: Repair and discrimination. The image repaired by the generator and the real image are input into the discriminator for discrimination. The two compete against each other. The generator continuously learns to generate images until the discriminator can no longer distinguish between the image repaired by the generator and the real image, thus realizing face image repair. The discriminator in step 400 is a two-stream discriminator with texture branch and structure branch; The discriminator's structural branch also has an additional edge detector for edge extraction; In step 400, adversarial loss, perceptual loss, and style loss are also introduced. The discriminator uses the loss function to distinguish between the image restored by the generator and the real image.