Dual-Mask Feature Attention Face Image Deblurring Method Based on Style Convolution

Through the dual-mask feature attention method based on style convolution, the semantic mask is constructed using the prior knowledge of the face and adaptively fused the identity information, solving the problem of unclear details in facial images defuzzing and improving the accuracy of face recognition.

CN117274611BActive Publication Date: 2025-07-18SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311437114.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-07-18
Estimated Expiration
2043-11-01

AI Technical Summary

Technical Problem

In the process of deblurring face images, the prior art fails to effectively utilize the prior knowledge of the face, resulting in the details of the recovered face images being not clear enough, affecting the accuracy of face recognition.

Method used

The double mask feature attention method based on style convolution is adopted, and the encoder, the style convolution double mask generator and the double mask feature attention are extracted through multi-scale feature attention, to build an accurate semantic mask and adaptively fuse identity information, accurately guide the reconstruction process.

Benefits of technology

The performance of face recognition algorithm is improved, and the clear face image is restored through rich identity details information, which improves the accuracy of face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274611B_ABST
    Figure CN117274611B_ABST
Patent Text Reader

Abstract

The present invention proposes a dual-mask feature attention face image deblurring method based on style convolution, which effectively removes the blur of face images and restores the details of face images by using a style convolution network and combining a dual-mask attention mechanism. The proposed multi-scale feature attention extraction encoder extracts identity feature information of different scales in the blurred face image, ensuring the identity consistency between the extracted features and the blurred face image. The proposed style convolution dual-mask generator analyzes the corresponding skin semantic mask and facial feature semantic mask in the features, and simultaneously obtains the corresponding deblurred features. The proposed dual-mask feature attention adaptive fusion decoder extracts the features of the masked area in the blurred face image through the attention mechanism, and adaptively fuses them with the deblurred features to enhance the details of the deblurred face image. Therefore, the proposed method can restore a single blurred face image into a clear face image to improve the performance of tasks such as face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image deblurring, face image restoration, and face image deblurring in computer vision, and particularly refers to a dual-mask feature attention face image deblurring method based on style convolution. Background Art

[0002] With the popularization of image capture devices such as smartphones, face image recognition has become a research hotspot in the field of computer vision. However, in actual scenarios, face images are interfered by factors such as camera shake or object movement, resulting in blurred imaging and reduced image quality, which greatly affects the accuracy of face recognition. Therefore, in order to avoid the reduction of face recognition rate caused by blurred face images, it is necessary to deblur face images.

[0003] Since image deblurring itself is a highly uncertain problem, the uncertainty of the face image deblurring problem is further increased due to the blurring of face detail textures and edges. However, due to the hierarchical nature of the face structure and the similar facial features, compared with other images, these similarities can provide prior knowledge for the deblurring task and help with face image deblurring. Therefore, for the face deblurring task, it is very important to utilize the prior knowledge of the face. Therefore, there are methods that use face semantic parsing maps, two-dimensional face sketch maps, or three-dimensional face maps as prior knowledge, but they all fail to consider that the prior knowledge comes from blurred images and has inaccurate problems itself, resulting in blurred details in the restored face images and affecting the performance of face recognition algorithms. Therefore, it is necessary to propose an effective face image deblurring method to avoid the unclear details of face images after deblurring caused by inaccurate prior knowledge. Summary of the Invention

[0004] The present invention overcomes the disadvantages and deficiencies in the prior art and proposes a dual-mask feature attention face image deblurring method based on style convolution. This method can fully extract the identity information in the blurred face image, and at the same time effectively construct an accurate semantic mask according to the identity information and perform adaptive fusion to accurately guide the use of identity information in the reconstruction process, so that the reconstructed face image has richer identity detail information and improves the performance of face recognition algorithms.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A dual-mask feature attention face image deblurring method based on style convolution, the method includes two processes: network training and inference:

[0007] The network training process includes the following steps:.

[0008] Step 1, the blurred face image (i.e.,B ) Input into the multi-scale feature attention extraction encoder (i.e., E ) to generate the corresponding blurred face information feature map (i.e., E ( B ));

[0009] Step 2: Input the blurred face information feature map into the style convolutional double mask generator D m to predict the preliminary de-blurred face image feature map and its corresponding skin semantic mask and facial feature semantic mask;

[0010] Step 3: Input the blurred face information feature map, skin semantic mask, facial feature semantic mask, and preliminary de-blurred face image feature map into the double mask feature attention adaptive fusion decoder D a and decode to obtain the corresponding clear face image;

[0011] Step 4: Input the reconstructed clear face image R' and the real clear face image R into the discriminator D dis to judge the probability that the image is a real clear image and construct the total objective function, which is specifically expressed as: , where L adv represents the generative adversarial objective function, L rec represents the reconstruction objective function, L id represents the identity consistency objective function, L fea represents the texture feature consistency objective function;

[0012] Step 5: Use the clear face image and the corresponding blurred face image dataset and the total objective function to train the multi-scale feature extraction encoder, style convolutional semantic mask generator, double mask feature attention adaptive fusion decoder, and discriminator, and use the Adam optimizer to update the network weight gradients.

[0013] The specific process of the multi-scale feature attention extraction encoder E in the above Step 1 is as follows: Input the blurred face image into the feature attention extraction module to obtain the preliminary blurred face image features. At the same time, downsample the blurred face image to the same size as the blurred face image feature map, obtain the original blurred face image feature map through convolution, and finally fuse the two through product operation and summation operation to obtain the blurred face information feature map.

[0014] Further, the style convolutional double mask generator in the second step D m The specific process is as follows:

[0015] a. Input the blurred face information feature map E ( B ) into the style convolutional decoder to obtain a preliminary de-blurred face image feature map, and the process is expressed as: G = S[E(B)] , where G represents the preliminary de-blurred face image feature map, S represents the style convolutional decoder;

[0016] b. Input the preliminary de-blurred face image feature map G into the double mask generator to predict the skin semantic mask corresponding to the preliminary de-blurred face image (i.e., M' s ) and the facial feature semantic mask (i.e., M' c ), and the process is expressed as: , , where D ms represents the skin semantic generator, D mc represents the facial feature semantic generator, and the bilingual semantic mask prediction loss function is: {L}_{mask}=-\left [ {{M}_{s}log{M}^{'}_{s}+\left ( {1-{M}_{s}} \right )log\left ( {1-{M}^{'}_{s}} \right )} \right ]-\left [ {{M}_{c}log{M}^{'}_{c}+\left ( {1-{M}_{c}} \right )log\left ( {1-{M}^{'}_{c}} \right )} \right ] , where M s represents the true skin semantic mask, M' s represents the predicted skin semantic mask, M c represents the true facial feature semantic mask, M' c represents the predicted facial feature semantic mask, log represents the logarithmic function.

[0017] Further, the specific process of the double mask feature attention adaptive fusion decoder in the third step is as follows:

[0018] a. First, input the blurred face information feature map, the skin semantic mask, the facial feature semantic mask, and the preliminary de-blurred face image feature map into the attention feature extraction module. The specific process is as follows:

[0019] 1). Multiply the skin semantic mask M' s and the facial feature semantic mask M' c separately with the preliminary de-blurred face image feature map G to obtain the face images in the corresponding regions of the masks. The process is expressed as: , , where G s represents the face image corresponding to the skin mask region, G c represents the face image corresponding to the facial feature mask region;

[0020] 2). Multiply the face image corresponding to the skin mask region G s and the face image corresponding to the facial feature mask region G c separately with the preliminary de-blurred face image feature map G and obtain the face image attention maps in the mask regions through the activation function. The process is expressed as: , , where A s represents the face image attention map of the skin mask region, A c represents the face image attention map of the facial feature mask region, and SoftMax represents the activation function;

[0021] 3). Multiply the face image attention map of the skin mask region A s and the face image attention map of the facial feature mask region A c separately with the blurred face information feature map E ( B ) to increase the proportion weight of the face information feature map in the mask regions and obtain the masked attention face information feature map. The process is expressed as: , , where F s represents the skin masked attention face information feature map, F c represents the facial feature masked attention face information feature map;

[0022] 4). The skin masked attention face information feature map Fs Combined with the facial feature map of the five - sense organ mask, the masked attention facial information feature map F c is obtained. The process is described as follows: F , where represents the combination method by channel; cat

[0023] b. Input the masked attention facial information feature map, the facial information feature map, and the preliminary de - blurred facial image feature map into the adaptive fusion module. The specific process is as follows:

[0024] 1). Input the masked attention facial information feature map F and the preliminary de - blurred facial image feature map G as well as the blurred facial information feature map E ( B ) into the convolutional layer, and after passing through the activation function, the weight F for measuring the masked attention facial information feature map E ( B ) and the blurred facial information feature map w is obtained. The process is described as: w = Sigmoid{conv[F, G, E(B)]} , where Sigmoid represents the activation function, conv represents the convolutional layer;

[0025] 2). Combine the masked attention facial information feature map F and the blurred facial information feature map E ( B ) according to the weight w , sum it with the preliminary de - blurred facial image feature map G , and after passing through the convolutional layer, a clear de - blurred facial image map is obtained. The process is described as: R' = conv[G + w * F+(1 - w)*E(B)] , where conv represents the convolutional layer, R' represents the reconstructed clear facial image.

[0026] ​The described network inference process includes the following steps: First, input the blurred face image into the trained encoder and feature decoder to output the blurred face information feature map; Second, use the style convolutional double mask generator to obtain the preliminary de-blurred face image feature map based on the input blurred face information feature map, and obtain the corresponding face semantic mask map; Finally, input the preliminary de-blurred face image feature map, the blurred face information feature map, and the double face semantic mask map into the double mask feature attention adaptive fusion decoder to obtain the corresponding clear face image.

[0027] The described multi-scale attention feature extraction encoder E is composed of a convolutional layer, a pooling layer, a normalization layer, and a downsampling layer, and fuses the feature maps at the same scale through product and summation operations. Input the blurred image into the encoder, and the resolution of the blurred face information feature map output by the encoder is 1 / 16 of the input image;

[0028] The described style convolutional double mask generator D m is composed of a style convolutional layer and an upsampling layer. Input the blurred face information feature map into D m to output the predicted preliminary de-blurred face feature map, the skin semantic mask, and the facial feature semantic mask;

[0029] The described double mask attention feature adaptive fusion decoder D a is composed of a style convolutional layer and an ordinary convolutional layer. Input the skin semantic mask, the facial feature semantic mask, the blurred face information feature map, and the preliminary de-blurred face information feature map into D a to output the reconstructed clear face image;

[0030] The generative adversarial objective function in the described network training process L adv is expressed as {L}_{adv}={E}_{R'}log\left \{{1+exp\left [ {-{D}_{dis}\left ( {R'} \right )} \right ]} \right \} , where E represents the expected value, R' represents the reconstructed face image, log represents the logarithmic function, exp represents the exponential function, D dis represents the discriminator; the reconstruction objective function L rec is expressed as , where ||·||1 represents the mean absolute error; the identity consistency objective function L id is expressed as {L}_{id}=1-\cos {\left [ {{f}_{arc}\left ( {R'} \right ),{f}_{arc}\left ( {R} \right )} \right ]} , where cos (·) represents the cosine function f arc (·) represents the face recognition model algorithm; the texture feature consistency objective function L fea is expressed as , where ||·||2 represents the mean square error

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows: by fully extracting the identity information at different scales in the blurred face image through multi-scale feature attention extraction and encoding, and at the same time using the style convolutional double-mask generator to effectively construct an accurate semantic mask according to the identity information, and adaptively fusing the identity information and the semantic mask through the double-mask feature attention adaptive fusion decoder, so as to accurately guide the network to use the identity information in the reconstruction process, making the reconstructed face image have richer identity detail information BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings

[0033] Figure 1 is a training and inference schematic diagram of the style convolutional double-mask feature attention face image deblurring method of the present invention

[0034] Figure 2 is a structural diagram of the multi-scale feature attention extraction encoder of the present invention

[0035] Figure 3 is a structural diagram of the style convolutional double-mask generator of the present invention

[0036] Figure 4 is a structural diagram of the double-mask feature adaptive fusion decoder of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following will describe in detail a typical embodiment of the face image deblurring method based on style convolution and dual-mask feature attention of the present invention, and further specifically describe this method. It is necessary to point out here that the following embodiments are only used to further illustrate this method and cannot be understood as limiting the protection scope of this method. Those skilled in the art make some non-essential improvements and adjustments to this method according to the content of this method, and still fall within the protection scope of the present invention.

[0038] The present invention proposes a face image deblurring method based on style convolution and dual-mask feature attention. As Figure 1 shown, this method includes two processes: network training and inference.

[0039] The network training process includes the following steps:

[0040] Step 1: Input the blurred face image into the multi-scale feature attention extraction encoder (i.e., E ) to generate the corresponding face information feature map (i.e., E ( B )). This step is reflected in Figure 2 the multi-scale feature attention extraction encoder structure diagram. In this step, the blurred face image is input into the feature attention extraction module to obtain the preliminary blurred face image features. At the same time, the blurred face image is downsampled to the same size as the blurred face image feature map, and the original blurred face image feature map is obtained through convolution. Finally, the two are fused through product operation and summation operation to obtain the blurred face information feature map;

[0041] Step 2: Input the blurred face information feature map into the style convolution dual-mask generator D m to predict the preliminary deblurred face image feature map and its corresponding skin semantic mask and facial feature semantic mask. This step is reflected in Figure 3 the style convolution dual-mask generator structure diagram. In this step, the output blurred face information feature map E ( B ) of the multi-scale feature attention extraction encoding is used as the original input, and is decoded through the style convolution decoder to obtain the corresponding preliminary deblurred face image feature map G , where the style convolution decoder is composed of a style convolution layer and an upsampling layer; while obtaining the preliminary deblurred face image feature map G , the preliminary deblurred face image feature map is input into the dual-mask generator to predict the skin semantic mask (i.e., M' s ) and the facial feature semantic mask (i.e., M' c ) corresponding to the preliminary deblurred face image. The process is expressed as: , , where D ms is represented as a skin semantic generator, D mc is represented as a facial feature semantic generator. According to the semantic mask loss function, it is determined whether the predicted semantic mask conforms to the true semantic mask. The bilingual semantic mask prediction loss function is: {L}_{mask}=-\left [ {{M}_{s}log{M}^{'}_{s}+\left ( {1-{M}_{s}} \right )log\left ( {1-{M}^{'}_{s}} \right )} \right ]-\left [ {{M}_{c}log{M}^{'}_{c}+\left ( {1-{M}_{c}} \right )log\left ( {1-{M}^{'}_{c}} \right )} \right ] , where M s represents the true skin semantic mask, M' s represents the predicted skin semantic mask, M c represents the true facial feature semantic mask, M' c represents the predicted facial feature semantic mask, log is represented as a logarithmic function;

[0042] Step 3: Input the blurred face information feature map, skin semantic mask, facial feature semantic mask, and the preliminary de-blurred face image feature map into the dual-mask feature attention adaptive fusion decoder D a and decode to obtain the corresponding clear face image. This step is reflected in Figure 4 the structural diagram of the dual-mask feature attention adaptive fusion decoder, which is mainly divided into two operations:

[0043] Operation 1: First, input the face information feature map, skin semantic mask, facial feature semantic mask, and the preliminary de-blurred face image feature map into the attention feature extraction module. The specific process is as follows:

[0044] 1). Multiply the skin semantic mask M' s and the facial feature semantic mask M' c separately with the preliminary de-blurred face image feature map G to obtain the face image in the corresponding region of the mask. The process is expressed as: , , where G c represents the face image corresponding to the facial feature mask region, G s represents the face image corresponding to the skin mask region;

[0045] 2), Multiply the face image corresponding to the skin mask region G s and the face image corresponding to the facial feature mask region G c with the preliminary deblurred face image feature map G respectively, and obtain the face image attention map of the mask region through the activation function. The process is expressed as: , , where A s represents the face image attention map of the skin mask region, A c represents the face image attention map of the facial feature mask region, and SoftMax represents the activation function;

[0046] 3), Multiply the face image attention map of the skin mask region A s and the face image attention map of the facial feature mask region A c with the blurred face information feature map E ( B ) respectively, to increase the proportion weight of the face information feature map in the mask region part, and obtain the masked attention face information feature map. The process is expressed as: , , where F s represents the masked attention face information feature map of the skin mask, F c represents the masked attention face information feature map of the facial feature mask;

[0047] 4), Combine the masked attention face information feature map of the skin mask F s with the masked attention face information feature map of the facial feature mask F c to obtain the masked attention face information feature map F , and the process is expressed as: , where cat represents the combination method by channels.

[0048] Operation 2: Input the masked attention face information feature map, the face information feature map, and the preliminary deblurred face image feature map into the adaptive fusion module. The specific process is as follows:

[0049] 1) Mask the facial information feature map F and the initial deblurred face image feature map G And the fuzzy face information feature map E ( B ) is input into the convolutional layer and then activated to obtain the feature map of the measured mask attention face information. F and fuzzy face information feature map E ( B ) w , the process is expressed as: w=Sigmoid\left \{{conv\left [ {F,G,E\left ( {B} \right )} \right ]} \right \} ,in Sigmoid Denoted as activation function, conv Represented as a convolutional layer;

[0050] 2) Mask the facial information feature map F and face information feature map E ( B ) By weight w Combined with the initial deblurred face image feature map G The sum is obtained through the convolutional layer to obtain a clear deblurred face image. The process is described as follows: {R}^{'}=conv\left [ {G+w\ast F+\left ( {1-w} \right )\ast E\left ( {B} \right )} \right ] ,in conv Represented as a convolutional layer, R' Represented as a reconstructed clear face image;

[0051] Step 4: Reconstruct the clear face image R' and real clear face images R Input to the discriminator D dis In the above example, the probability that the image is a real clear image is judged, and the overall objective function is constructed, which is specifically expressed as: ,in L adv Expressed as the generative adversarial objective function, L rec Expressed as the reconstruction objective function, L id Expressed as the identity consistent objective function, L fea It is expressed as a texture feature consistent objective function;

[0052] Step 5: Use the clear face images, the corresponding blurred face image datasets, and the overall objective function to train the multi-scale feature extraction encoder, the style convolutional semantic mask generator, the dual-mask feature attention adaptive fusion decoder, and the discriminator, and use the Adam optimizer with an initial learning rate of 0.001 to update the network weight gradients.

[0053] The multi-scale attention feature extraction encoder E is composed of convolutional layers, pooling layers, normalization layers, and downsampling layers, and fuses the feature maps at the same scale through product and summation operations. Input the blurred image into the encoder, and the resolution of the blurred face information feature map output by the encoder is 1 / 16 of the input image;

[0054] The style convolutional dual-mask generator D m is composed of style convolutional layers and an upsampling layer. Input the blurred face information feature map into D m to output the predicted preliminary de-blurred face feature map, skin semantic mask, and facial feature semantic mask;

[0055] The dual-mask attention feature adaptive fusion decoder D a is composed of a style convolutional layer and a normal convolutional layer. Input the skin semantic mask, facial feature semantic mask, blurred face information feature map, and preliminary de-blurred face information feature map into D a to output the reconstructed clear face image.

[0056] The generative adversarial objective function in the network training process L adv is expressed as {L}_{adv}={E}_{R'}log\left \{{1+exp\left [ {-{D}_{dis}\left ( {R'} \right )} \right ]} \right \} , where E represents the expected value, R' represents the reconstructed face image, log represents the logarithmic function, exp represents the exponential function, D dis represents the discriminator; the reconstruction objective function L rec is expressed as , where ||·||1 represents the mean absolute error; the identity consistency objective function L idExpressed as {L}_{id}=1-\cos {\left [ {{f}_{arc}\left ( {R'} \right ),{f}_{arc}\left ( {R} \right )} \right ]} , where cos (·) represents the cosine function, f arc (·) represents the face recognition model algorithm; the texture feature consistent objective function L fea is expressed as , where ||·||2 represents the mean square error.

[0057] The network inference process described above includes the following steps: First, input the blurred face image into the trained encoder and feature decoder to output the blurred face information feature map; Second, use the style convolutional double-mask generator to obtain the preliminary de-blurred face image feature map based on the input blurred face information feature map, and obtain the corresponding face semantic mask map; Finally, input the preliminary de-blurred face image feature map, the blurred face information feature map, and the double-face semantic mask map into the double-mask feature attention adaptive fusion decoder to obtain the corresponding clear face image.

Claims

1. A face image deblurring method based on a style convolutional double mask, characterized in that It includes two processes: network training and inference: The specific description of the network training process is as follows: Step 1: Input the blurred face image into the multi-scale feature attention extraction encoder E to generate the corresponding blurred face information feature map E ( B ) The specific process is as follows: Input the blurred face image into the feature attention extraction module to obtain the preliminary blurred face image features. At the same time, downsample the blurred face image to the same size as the blurred face image feature map, and obtain the original blurred face image feature map through convolution. Finally, fuse the two through product operation and summation operation to obtain the blurred face information feature map; Step 2: Input the blurred face information feature map into the style convolutional double mask generator D m to predict the preliminary de-blurred face image feature map and its corresponding skin semantic mask and facial feature semantic mask. The specific process is as follows: a. Input the blurred face information feature map E ( B ) into the style convolutional decoder to obtain a preliminary de-blurred face image feature map. The process is described as follows: , where G represents the preliminary de-blurred face image feature map, S represents the style convolutional decoder; b. Input the preliminarily deblurred facial image feature map G into a dual-mask generator to predict the skin semantic mask (i.e., M' s ) and the facial feature semantic mask (i.e., M' c ) corresponding to the preliminarily deblurred facial image. The process is described as follows: , , where D ms represents the skin semantic generator, D mc represents the facial feature semantic generator. The dual-semantic mask prediction loss function is: , where M s represents the true skin semantic mask, M' s represents the predicted skin semantic mask, M c represents the true facial feature semantic mask, M' c represents the predicted facial feature semantic mask, log represents the logarithmic function; Step 3: Input the blurred face information feature map, skin semantic mask, facial feature semantic mask, and preliminary de-blurred face image feature map into the dual-mask feature attention adaptive fusion decoder D a and decode to obtain the corresponding clear face image. The specific process is as follows: a. First, input the face information feature map, skin semantic mask, facial feature semantic mask, and the preliminary deblurred face image feature map into the attention feature extraction module. The specific process is as follows: 1), Multiply the skin semantic mask M' s with the facial feature mask M' c respectively by the preliminary deblurred face image feature map G to obtain the face image in the corresponding masked area. The process is described as: , where G s represents the face image corresponding to the skin mask area, G c represents the face image corresponding to the facial feature mask area; 2), the face image corresponding to the skin mask area G s and the face image corresponding to the facial feature mask area G c are respectively multiplied by the preliminary de-blurred face image feature map G , and the face image attention map of the mask area is obtained through the activation function. The process is expressed as: , , where A s represents the face image attention map of the skin mask area, A c represents the face image attention map of the facial feature mask area, and softMax represents the activation function; 3), the face image attention map of the skin mask region A s and the face image attention map of the facial feature mask region A c are respectively multiplied by the face information feature map E ( B ) to increase the proportion weight of the face information feature map in the mask region part, and a masked attention face information feature map is obtained. The process is expressed as: , , where F s represents the skin mask attention face information feature map, F c represents the facial feature mask attention face information feature map; 4), combine the skin mask attention face information feature map F s with the facial feature mask attention face information feature map F c to obtain the mask attention face information feature map F , and its process is expressed as: , where cat represents the combination method by channels; b. Input the masked attention face information feature map, face information feature map, and the preliminary deblurred face image feature map into the adaptive fusion module. The specific process is as follows: 1), input the masked attention face information feature map F and the preliminary de-blurred face image feature map G as well as the face information feature map E ( B ) into the convolutional layer, and obtain the weighted masked attention face information feature map F and the face information feature map E ( B ) through the activation function. The process is expressed as: w , where , Sigmoid represents the activation function, conv represents the convolutional layer; 2), the masked attention face information feature map F and the face information feature map E ( B ) are combined according to weights w and summed with the preliminary deblurred face image feature map G to obtain a clear face image deblurring map through a convolutional layer. The process is expressed as: , where conv represents the convolutional layer, R' represents the reconstructed clear face image; Step 4: Input the reconstructed clear face image R' and the real clear face image R into the discriminator D dis to determine the probability that the image is a real clear image, and construct the total objective function, which is specifically expressed as: , where L adv is expressed as a generative adversarial objective function, L rec is expressed as a reconstruction objective function, L id is expressed as an identity consistency objective function, L fea is expressed as a texture feature consistency objective function; Step 5: Train the multi-scale feature extraction encoder, style convolutional semantic mask generator, dual-mask feature attention adaptive fusion decoder, and discriminator using the clear face image, corresponding blurred face image dataset, and the total objective function; The specific description of the network inference process is as follows: Step 1: Input the blurred face image into the trained encoder and feature decoder to output the blurred face information feature map; Step 2: Use the style convolutional dual-mask generator to obtain the preliminary deblurred face image feature map based on the input blurred face information feature map and obtain the corresponding face semantic mask map; Step 3: Input the preliminary deblurred face image feature map, blurred face information feature map, and dual-face semantic mask map into the dual-mask feature attention adaptive fusion decoder to obtain the corresponding clear face image.

2. The face image deblurring method based on style convolution double masks according to claim 1, wherein The multi-scale attention feature extraction encoder in Step 1 of the network training process E It is composed of a convolutional layer, a pooling layer, a normalization layer, and a downsampling layer. The feature maps at the same scale are fused through product and summation operations. The blurred image is input into the encoder, and the resolution of the blurred face information feature map output by the encoder is 1 / 16 of the input image. The style convolutional double-mask generator in Step 2 of the network training process D m It is composed of a style convolutional layer and an upsampling layer. The blurred face information feature map is input into D m to output the predicted preliminary de-blurred face feature map, skin semantic mask, and facial feature semantic mask. The double-mask attention feature adaptive fusion decoder in Step 3 of the network training process D a It is composed of a style convolutional layer and a common convolutional layer. The skin semantic mask, facial feature semantic mask, blurred face information feature map, and preliminary de-blurred face information feature map are input into D a to output the reconstructed clear face image. The generative adversarial objective function in Step 4 of the network training process L adv is expressed as , where E represents the expected value, R' represents the reconstructed face image, log represents the logarithmic function, exp represents the exponential function, D dis represents the discriminator; the reconstruction objective function in Step 4 of the network training process L rec is expressed as , where ||·||1 represents the mean absolute error; the identity consistency objective function in Step 4 of the network training process L id is expressed as , where cos (·) represents the cosine function, f arc (·) represents the face recognition model algorithm; the texture feature consistency objective function in Step 4 of the network training process L fea is expressed as , where ||·||2 represents the mean square error.

Citation Information

Patent Citations

  • Portrait restoration method and device, electronic equipment and computer storage medium

    CN112330574A

  • Oil painting generation method and device, computer equipment and storage medium

    CN112734874A