Face image local area restoration method based on prior knowledge distillation

By constructing a teacher-student network system and a dual-gated convolutional coordination attention module, the problem of insufficient error accumulation and long-distance dependency in face repair in the existing technology is solved, and high-quality face image repair is achieved.

CN120450951APending Publication Date: 2025-08-08INNER MONGOLIA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510638223.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing repair techniques based on face prior information are difficult to fully model the long-distance dependence between effective pixels, resulting in distortion of the face structure and semantic information of the repair results, and the prediction errors are easily accumulated, affecting the repair effect.

Method used

The local area repair method of face images based on prior knowledge distillation is adopted. By constructing a teacher network and a student network, using a dual-gated convolutional coordination attention module, combined with a generative adversarial network, the teacher network transmits prior information to guide students' online learning, avoiding error accumulation, and enhancing effective pixel feature extraction and semantic understanding ability.

Benefits of technology

It significantly improves the structural integrity and semantic coherence of face images after repair, narrows the gap between the repair image and the original real image, and improves the repair accuracy and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450951A_ABST
    Figure CN120450951A_ABST
Patent Text Reader

Abstract

A face image local area restoration method based on priori knowledge distillation comprises the following steps: extracting a corresponding face analysis graph from a public data set as priori information, preprocessing the face analysis graph, generating a free occlusion training set, and training a constructed teacher network containing a double-gating convolution coordination attention module, freezing parameters of the trained teacher network, and performing instructive training on the constructed student network; and finally, performing restoration processing on the face image received in real time through the trained student network. According to the method, on the basis of the face prior information, the error accumulation dilemma caused by prediction prior information deviation is ingeniously avoided, the repair accuracy is improved from the source, the long-distance dependency relationship between effective pixels is deeply mined and fully utilized, and under the condition of free shielding, the repair accuracy is improved. The repaired face image is more excellent in structural integrity and semantic coherence, and the gap between the repaired image and the original real image is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of image processing, specifically a method for repairing local areas of facial images based on prior knowledge distillation. Background Art

[0002] Existing facial restoration techniques based on prior information rely on the accuracy of predicted prior information. Prediction errors easily accumulate within the network, affecting the restoration effect. Furthermore, these methods struggle to fully model the long-range dependencies between valid pixels in the image, resulting in distortion of facial structure and semantic information in the restoration results. Summary of the Invention

[0003] In response to the shortcomings of existing technologies, such as difficulty in enhancing the ability to extract effective pixel features and understand semantics, difficulty in fully reducing interference in invalid areas, and poor performance in detail preservation and structural consistency under complex masks or high occlusion, this paper proposes a method for local area restoration of facial images based on prior knowledge distillation. Relying on facial prior information, this method cleverly avoids the error accumulation dilemma caused by bias in predicting prior information, and improves the accuracy of restoration from the source. Moreover, this method deeply explores and fully utilizes the long-distance dependencies between effective pixels, making the restored facial image perform better in free mask structural integrity and semantic coherence, significantly narrowing the gap between the restored image and the original real image, and opening up a new path for high-quality facial image restoration.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a method for local area restoration of facial images based on prior knowledge distillation. The method extracts corresponding facial parsing graphs from public datasets as prior information, preprocesses the facial parsing graphs, and generates a free occlusion training set. The generated set is then used to train a constructed teacher network containing a dual-gated convolutional coordinated attention module (DGCAM). The parameters of the trained teacher network are then frozen, and guided training is performed on the constructed student network. Finally, the trained student network is used to perform restoration processing on the facial images received in real time.

[0006] The face parsing graph is extracted using, but not limited to, a BiSeNet network.

[0007] The preprocessing mentioned above refers to cropping and aligning the face images provided in the public dataset to make them suitable for network input requirements.

[0008] The free occlusion training set is obtained by processing the high-definition face images in the public training set into 256×256 resolution face images with free occlusion. Specifically, a free occlusion of 256×256 pixels is generated by programming, and the obtained free occlusion is fused with the face image to generate a face image with free occlusion.

[0009] The teacher network includes: a face restoration generator guided by prior information and a discriminator, wherein: the face restoration generator uses prior information from the real image to restore the occluded face image to obtain a soft real image, and the discriminator performs identification processing based on the soft real image and the free occlusion image to judge the authenticity of the restoration result and optimize the restoration performance of the teacher network generator.

[0010] The face restoration generator includes: two gated convolution units, a bottleneck layer composed of four gated dilated convolution units, an encoder composed of four dual-gated convolution coordinated attention modules, and a decoder composed of four dual-gated convolution coordinated attention modules and a transposed convolution unit, wherein: the encoder and the decoder are connected via a skip connection, and after the input image is gated convolution by the first gated convolution unit, the dual-gated convolution coordinated attention module of the encoder encodes the feature information, compresses the size of the feature map, and extracts face-related information from the prior information. The gated dilated convolution unit of the bottleneck layer expands the receptive field while capturing more global information in the feature map; the decoder fuses the features extracted in the encoding stage with the features in the decoding stage through a skip connection, processes the fused features through the dual-gated convolution coordinated attention module, and then upsamples through the transposed convolution unit to gradually restore the original size of the feature map. The second gated convolution outputs the final restoration result.

[0011] The dual-gated convolution coordinated attention module includes: two gated convolution units and a coordinated attention unit, wherein: The gated convolution unit extracts valid pixels; the coordinated attention unit retains spatial encoding features along the H and W directions through horizontal and vertical pooling, respectively. It then performs learnable adaptive compression of the channel dimension through 1×1 convolution and decomposes the processed spatial and channel features into horizontal and vertical weight matrices. Dynamic spatial calibration of the feature map is achieved through element-by-element multiplication. The dual-gated convolution coordinated attention module combines gated convolution with coordinated attention, simultaneously considering inter-channel relationships and positional information, effectively extracting semantic feature information from valid pixels in occluded face images.

[0012] The student network is implemented using a face inpainting generator with the same structure as the teacher network. It is trained using input consisting solely of occluded and unoccluded face images, eliminating the need for external prior knowledge. Leveraging the prior knowledge imparted by the teacher network, the student network is able to efficiently perform image inpainting. This not only simplifies input requirements but also further mitigates the negative impact of inaccurate priors. Technical Effects

[0013] The teacher network of the present invention is based on a generative adversarial network, and the student network consists only of a generator. Through knowledge distillation, the student network learns the teacher network's ability to use prior information for image restoration. In addition, a dual-gated convolution coordinated attention module is proposed, which consists of two gated convolutions and coordinated attention. Compared with the existing technology, the present invention avoids the error transmission and accumulation caused by inaccurate prior knowledge while using prior information, thereby improving the restoration effect. The dual-gated convolution coordinated attention module enhances the network's ability to extract effective pixel features and understand semantics in undamaged areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Flowchart of the present invention;

[0015] Figure 2 It is a schematic diagram of the principle of the present invention;

[0016] Figure 3 This is a schematic diagram of the face restoration generator;

[0017] Figure 4 Schematic diagram of the double-gated convolution coordinated attention module;

[0018] Figure 5 Schematic diagram of the effect of the free mask embodiment. DETAILED DESCRIPTION

[0019] like Figure 1 As shown, this embodiment relates to a method for repairing local areas of facial images based on prior knowledge distillation, including:

[0020] Step 1: Obtain the face parsed images corresponding to the public dataset through the BiSeNet network. Preprocess the images and generate two types of training sets: free occlusion and free occlusion. Specifically,

[0021] 1.1 Download the CelebA-HQ and FFHQ datasets, process the datasets using the pre-trained weight files provided by the BiSeNet network, and obtain the corresponding face parsing images.

[0022] 1.2 By randomly generating free occlusions of 256×256 pixels, the obtained free occlusions are fused with the face image to generate an occlusion image dataset with free occlusions.

[0023] 1.3 Process the CelebA-HQ dataset, FFHQ dataset, occlusion image dataset, and corresponding occlusions to obtain images of size 256×256.

[0024] 1.4 Divide the above dataset into a training set and a test set. The training set is used to train the network model in the offline stage, and the test set is used to test the network performance.

[0025] Step 2: Input the occluded face image, the corresponding occluded image and the complete face parsed image in the training set into Figure 2 The teacher network shown in the figure completes the training of the teacher network, specifically including:

[0026] a: Prepare the occluded face image, the corresponding occluded image and the complete face parsed image, and construct Figure 3 The face restoration generator, discriminator, style loss, perceptual loss, pixel loss, and adversarial loss functions guided by the prior information shown, and a stochastic optimization method (Adam) with adaptive momentum is used to accelerate the convergence of the teacher network.

[0027] The face restoration generator includes: two gated convolution units, a bottleneck layer composed of four gated dilated convolution units, and four Figure 4 The encoder shown consists of a dual-gated convolutional coordinated attention module and a decoder consisting of four dual-gated convolutional coordinated attention modules and transposed convolution units.

[0028] The training described above uses the real loss L hard and adversarial loss L adv The loss function is obtained by weighting, and by minimizing the adversarial loss, the generator's restoration result is closer to the real image, making it impossible for the discriminator to accurately distinguish between the real image and the restored image. The loss function is specifically: L teacher =λ hard *L hard +λ adv *L adv , where: hard and λ adv is the weighting parameter.

[0029] The weighting parameter is preferably λ hard =1,λ adv =1.

[0030] The original true loss L hardLoss is a loss function obtained by weighting pixel loss, perceptual loss and style loss. Specifically, L hard (I t ,I tt )=λ1*L1(I t ,I gt )+λ percep *L percep (I t ,I gt )+λ style *L style (I t ,I gt ), where: L1 is pixel loss, L percep is the perceptual loss, L style is the style loss, λ1, λ percep and λ style are weighted parameters, I t Repair results for teacher network,I gt is the original real image.

[0031] The weighting parameters are preferably λ1=1, λ percep =0.1,λ style =250.

[0032] The pixel loss Among them: ‖·‖1 is the L1 norm, I gt and I t are the original real image and the image generated by the teacher network, M is the corresponding occluded image, and N is the number of pixels involved in the calculation.

[0033] The perceived loss This paper uses the difference between the activation features of the VGG-19 network pre-trained on ImageNet to judge the image restoration quality, where: Corresponding to the five activation layers (Relu1_1, Relu2_1, Relu3_1, Relu4_1 and Relu5_1).

[0034] The style loss This study selects the VGG-19 network activation layer that is consistent with the perceptual loss and calculates the style loss by calculating the covariance between the activation feature maps of different scales. j *W j *C j ,in: C j ×C j The Gram matrix of size is composed of activation features composition.

[0035] The adversarial loss is used to train the teacher network using the least squares generative adversarial network (LSGAN), specifically Where D is the discriminator and G is the teacher network generator. For the discriminator loss, LSGAN uses a least-squares loss to penalize the squared error between the discriminator prediction and the desired target (a = 1 for real images, b = 0 for generated data, and c = 1 for the generator). For the generator loss, LSGAN aims to minimize the squared difference between the discriminator's prediction of the generated sample and the real image.

[0036] b: Construct model-related hyperparameters such as training batch size (batch), number of training rounds (epochs), and learning rate, input the occluded face image, the corresponding occluded image, and the complete face parsed image into the teacher network for training, and perform forward propagation, loss calculation, backpropagation, and parameter update operations on the generator and discriminator in the teacher network in a loop, so that the teacher network achieves the best restoration effect under the guidance of the complete face parsed image, completing the pre-training of the teacher network.

[0037] c: Find the network parameters that optimize the teacher network repair effect, freeze the parameters of the teacher network, and prepare for the subsequent student network training.

[0038] Step 3: Input the occluded face image in the training set, the corresponding occluded image, and the complete face parsed image into the teacher network with frozen parameters to obtain the face image repaired under the guidance of the complete face parsed image, which is used as the soft real image. Input the occluded face image and occlusion map in the training set into the student network for guided training to repair the occluded image. Specifically, it includes:

[0039] a: Prepare the occluded face image, the corresponding occluded image and the complete face parsed image, construct the face restoration generator, style loss, perceptual loss, pixel loss and feature loss functions of the student network, and use the stochastic optimization method (Adam) with adaptive momentum to accelerate the convergence of the student network.

[0040] The guided training described above takes the occluded face image and the corresponding occluded image as input and uses the original true loss L hard , feature space loss L feature and soft truth loss L soft The loss function is obtained by weighting. Specifically: L student =λ hard *L hard +λ feature *L feature +λ soft *L soft , where: hard ,λfeature ,λ soft L hard 、L feature and L soft The weight of .

[0041] The weighting parameter is preferably λ hard =1,λ feature =0.5,λ soft =0.2.

[0042] The ground truth loss L hard (I s ,I gt )=λ1L1(I s ,I gt )+λ percep *L percep (I s ,I gt )+λ style *L style (I s ,I gt ), I s Fix results for student network.

[0043] The feature space loss in: and are the kth intermediate features of the teacher network and the student network, respectively, and n is the number of features. This loss function encourages the student network to imitate the intermediate features of the teacher network, ensuring that it can not only capture image details but also understand and reproduce complex facial structures and semantic information.

[0044] The soft ground truth loss L soft (I s ,I t )=λ1L1(I s ,I t )+λ percep *L percep (I s ,I t )+λ style *L style (I s ,I t ), where: L1 is pixel loss, L percep is the perceptual loss, L style is the style loss, λ1, λ percep and λ styte The soft ground truth loss measures the distance between the outputs of the teacher network and the student network, simplifies the learning objective of the student network, and provides a more effective supervision signal, which promotes the understanding and application of prior knowledge by the student network.

[0045] The weighting parameters are preferably λ1=1, λ percep =0.1,λ style =250.

[0046] b: Freeze the parameters in the teacher network generator in step 2, input the occluded face image, the corresponding occluded image and the complete face solution image into the network, obtain the teacher network repair result, and use the repair result as the soft real image to assist in training the student network.

[0047] c) The occluded face image and the corresponding occluded image are fed into the student network for forward propagation, yielding the student network's inpainted result. Simultaneously, the modules in the student network are aligned with those in the teacher network, and the feature loss between the two networks is calculated. The loss between the student network-generated image, the teacher network-generated image, and the original real image is then weighted, and backpropagation is performed to update the student parameters.

[0048] d: Perform a training cycle with the goal of optimizing the student network repair performance.

[0049] Step 4: In the online stage, the trained student network model is used to repair the occluded face image input in real time.

[0050] After Figure 5 The specific actual experiment shown in the figure is based on the specific environment settings of PyTorch 2.0.1 and NVIDIA 4070Ti GPU. The teacher network is trained with a batch size of 2 and the generator SGD initial learning rate of 1.0×10 -4 The initial learning rate of the discriminator SGD is 4.0×10 -4 The above device is run for 30 training rounds; the student network is trained with a batch size of 2 and an initial SGD learning rate of 1.0×10 -4 The above device was run for 45 training rounds. Under the condition of free occlusion of 40%-50% of the total image area, the experimental data obtained were: PSNR reached 28.96 and SSIM reached 0.9110 on the CelebA-HQ dataset; PSNR reached 27.85 and SSIM reached 0.9065 on the FFHQ dataset.

[0051] Compared with existing technologies, this method, guided by prior facial information, avoids the error accumulation caused by inaccurate prediction of prior information, improves the accuracy of the restored facial structure and semantic information, and reduces the difference between the restored image and the real image. Secondly, through a dual-gated convolutional coordinated attention module, it avoids the influence of invalid pixels while enhancing the ability to extract valid pixel features and understand semantics, thus improving the rationality of the restoration results.

[0052] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.

Claims

1. A method for local area restoration of face images based on prior knowledge distillation, characterized in that: While extracting the corresponding facial parsing graph from the public dataset as prior information, the facial parsing graph is preprocessed to generate a free occlusion training set, which is then used to train the constructed teacher network containing a dual-gated convolutional coordinated attention module. The parameters of the trained teacher network are then frozen, and the constructed student network is guided for training. Finally, the trained student network is used to repair the face images received in real time.

2. The method for local area restoration of facial images based on prior knowledge distillation according to claim 1, characterized in that: The teacher network includes: a face restoration generator guided by prior information and a discriminator, wherein: the face restoration generator uses prior information from the real image to restore the occluded face image to obtain a soft real image, and the discriminator performs identification processing based on the soft real image and the free occlusion image to judge the authenticity of the restoration result and optimize the restoration performance of the teacher network generator.

3. The method for local area restoration of facial images based on prior knowledge distillation according to claim 2, characterized in that: The face restoration generator includes: two gated convolution units, a bottleneck layer composed of four gated dilated convolution units, an encoder composed of four dual-gated convolution coordinated attention modules, and a decoder composed of four dual-gated convolution coordinated attention modules and a transposed convolution unit, wherein: the encoder and the decoder are connected via a skip connection, and after the input image is gated convolution by the first gated convolution unit, the dual-gated convolution coordinated attention module of the encoder encodes the feature information, compresses the size of the feature map, and extracts face-related information from the prior information. The gated dilated convolution unit of the bottleneck layer expands the receptive field while capturing more global information in the feature map; the decoder fuses the features extracted in the encoding stage with the features in the decoding stage through a skip connection, processes the fused features through the dual-gated convolution coordinated attention module, and then upsamples through the transposed convolution unit to gradually restore the original size of the feature map. The second gated convolution outputs the final restoration result.

4. The method for local area restoration of facial images based on prior knowledge distillation according to claim 1 or 3, characterized in that: The dual-gated convolution coordinated attention module includes: two gated convolution units and a coordinated attention unit, wherein: the gated convolution unit extracts valid pixels; the coordinated attention unit retains the spatial coding features along the H and W directions respectively through horizontal pooling and vertical pooling, and then performs learnable adaptive compression on the channel dimension through 1×1 convolution and decomposes the processed spatial and channel features into horizontal and vertical weight matrices, and realizes dynamic spatial calibration of the feature map through element-by-element multiplication. The dual-gated convolution coordinated attention module effectively extracts the semantic feature information of valid pixels in the occluded face image by combining gated convolution and coordinated attention, while taking into account the relationship between channels and position information.

5. The method for local area restoration of facial images based on prior knowledge distillation according to claim 1, characterized in that: The student network is implemented using a face restoration generator with the same structure as the teacher network.

6. The method for local area restoration of a face image based on prior knowledge distillation according to any one of claims 1 to 5, wherein: include: Step 1: Obtain the face parsed images corresponding to the public dataset through the BiSeNet network, preprocess the images and generate a free occlusion training set, specifically including: Step 2: Input the occluded face images, the corresponding occluded images, and the complete face parsed images in the training set into the teacher network to complete the training of the teacher network. Step 3: Input the occluded face image in the training set, the corresponding occluded image, and the complete face parsed image into the teacher network with frozen parameters to obtain the face image repaired under the guidance of the complete face parsed image. This image is used as the soft ground truth image, and the occluded face image and occlusion map in the training set are input into the student network for guided training to repair the occluded image. Step 4: In the online stage, the trained student network model is used to repair the occluded face image input in real time.

7. The method for local area restoration of facial images based on prior knowledge distillation according to claim 6, characterized in that: The step 1 specifically includes: 1.1 Download the CelebA-HQ and FFHQ datasets, process the datasets using the pre-trained weight files provided by the BiSeNet network, and obtain the corresponding face parsing images; 1.2 By randomly generating free occlusions of 256×256 pixels, the obtained free occlusions are fused with the face image to generate an occlusion image dataset with free occlusions; 1.3 Process the CelebA-HQ dataset, FFHQ dataset, occlusion image dataset, and corresponding occlusions to obtain images of size 256×256; 1.4 Divide the above dataset into a training set and a test set. The training set is used to train the network model in the offline stage, and the test set is used to test the network performance.

8. The method for local area restoration of facial images based on prior knowledge distillation according to claim 6, characterized in that: The step 2 specifically includes: a: Prepare the occluded face image, the corresponding occluded image, and the complete face parsed image, construct the face restoration generator, discriminator, style loss, perceptual loss, pixel loss, and adversarial loss functions guided by prior information, and use the stochastic optimization method with adaptive momentum to accelerate the convergence of the teacher network; b) Construct model-related hyperparameters such as the training batch size, number of training rounds, and learning rate. Input the occluded face image, the corresponding occluded image, and the complete face parsed image into the teacher network for training. Perform forward propagation, loss calculation, backpropagation, and parameter update operations on the generator and discriminator in the teacher network. This ensures that the teacher network achieves the best restoration effect under the guidance of the complete face parsed image, completing the pre-training of the teacher network. c) Find the network parameters that optimize the teacher network's repair effect, freeze the teacher network's parameters, and prepare for the subsequent student network training; The training described above uses the real loss L hard and adversarial loss L adv The loss function is obtained by weighting. By minimizing the adversarial loss, the generator's restoration result is made closer to the real image, making it impossible for the discriminator to accurately distinguish between the real image and the restored image. The specific loss function is: L teacher =λ hard *L hard +λ adv *L adv , where: hard and λ adv is the weighting parameter; The original true loss L hard Loss is a weighted loss function obtained by weighting pixel loss, perceptual loss and style loss. Specifically, L hard (I t ,I gt )=λ1*L1(I t ,I gt )+λ pwrcep *L percep (I t ,I gt )+λ style *L style (I t ,I gt ), where: L1 is pixel loss, L percep is the perceptual loss, L style is the style loss, λ1, λ percep and λ style are weighted parameters, I t Repair results for teacher network,I gt is the original real image; The pixel loss Among them: ‖·‖1 is the L1 norm, I gt and I t are the original real image and the image generated by the teacher network, M is the corresponding occluded image, and N is the number of pixels involved in the calculation; The perceived loss The difference between the activation features of the VGG-19 network pre-trained on ImageNet is used to judge the image restoration quality, where: Corresponding to five activation layers; The style loss in: C j ×C j The Gram matrix of size is composed of activation features composition; The adversarial loss is used to train the teacher network using the least squares generative adversarial network, specifically Among them: D is the discriminator and G is the teacher network generator.

9. The method for local area restoration of facial images based on prior knowledge distillation according to claim 6, characterized in that: The step 3 specifically includes: a: Prepare the occluded face image, the corresponding occluded image, and the complete face parsed image, construct the face restoration generator, style loss, perceptual loss, pixel loss, and feature loss functions of the student network, and use the stochastic optimization method (Adam) with adaptive momentum to accelerate the convergence of the student network; b: Freeze the parameters in the teacher network generator in step 2, input the occluded face image, the corresponding occluded image, and the complete face solution image into the network, obtain the teacher network restoration result, and use the restoration result as the soft real image to assist in training the student network; c) Input the occluded face image and the corresponding occluded image into the student network for forward propagation to obtain the student network restoration result. At the same time, align the modules in the student network with those in the teacher network, and calculate the feature loss between the two networks. Calculate the loss between the image generated by the student network, the image generated by the teacher network, and the original real image. The obtained loss is weighted, and then backpropagation is performed to update the student parameters. d: Perform a training cycle to optimize the student network repair performance; The guided training described above takes the occluded face image and the corresponding occluded image as input and uses the original true loss L hard , feature space loss L feature and soft truth loss L soft The weighted loss function is: L student =λ hard *L hard +λ feature *L feature +λ soft *L soft , where: hard ,λ feature ,λ soft L hard 、L feature and L soft The weight of The ground truth loss L hard (I s ,I gt )=λ1L1(I s ,I gt )+λ percep *L percep (I s ,I gt )+λ style *L style (i s ,I gt ), I s Fix results for student networks; The feature space loss in: and are the kth intermediate features of the teacher network and the student network respectively, and n is the number of features; The soft ground truth loss L soft (I s ,i t )=λ1L1(I s ,I t )+λ percep *L percep (I s ,I t )+λ style *K styte (I s ,I t ), where: L1 is pixel loss, L percep is the perceptual loss, K style is the style loss, λ1, λ percep and λ style are weighted parameters.