A Mask Occluded Face Detection and Recognition Method Incorporating Attention Mechanism

By integrating attention mechanism and data enhancement technology in mask mask face recognition, combined with facial organ attention mechanism and knowledge distillation, the existing algorithms have solved the problems of low recognition rate and large model parameters in mask mask occlusion, and efficient mask face detection and recognition are achieved.

CN115497139BActive Publication Date: 2025-06-27SOUTHWEAT UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211188477.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-06-27
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

The existing mask mask mask recognition algorithm based on deep learning has problems such as lack of data sets, large model parameters, low recognition rate, and inability to detect whether faces wear masks.

Method used

Using the method of fusion attention mechanism, virtual masks are added through data augmentation technology, facial features are extracted and fused, facial organ attention mechanism is used to focus on facial organs that are not blocked by masks, and the model parameters are compressed through knowledge distillation.

Benefits of technology

It realizes efficient face detection and recognition under mask occlusion, reduces the number of model parameters, improves the recognition rate, and has wide applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 220928162509
    Figure 220928162509
  • Figure 220928162703
    Figure 220928162703
  • Figure 220928162708
    Figure 220928162708
Patent Text Reader

Abstract

The present invention provides a method for detecting and recognizing a face occluded by a mask by integrating an attention mechanism. First, the network improves the Swin Transformer for face feature extraction. Second, a face organ attention mechanism (FOA) is proposed to enable the model to focus on the face organs not occluded by the mask. Then, to address the problem of insufficient current face-occluded-by-mask datasets, a data augmentation method for adding mask occlusion by generating three-dimensional face meshes is proposed. Finally, to address the problem of a large number of model parameters, a method for compressing the model using knowledge distillation is proposed. This method balances speed and accuracy well, achieves excellent performance in detecting and recognizing faces occluded by masks, and has wide applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and more specifically, to a method for detecting and recognizing a face occluded by a mask by integrating an attention mechanism. Background Art

[0002] Face recognition (FR) has been an active research topic in the field of computer vision for the past few decades. Although significant progress has been made in FR technology during this period, there are still many difficulties to be solved in actual application scenarios. Wearing a mask will hide some facial features and prevent the face recognition system from making a correct decision.

[0003] For a long time, local feature methods and shallow feature learning have been the focus of face recognition research. It wasn't until the birth of FaceNet in 2015 that the research focus of face recognition shifted to the deep learning direction. Currently, the most advanced methods based on deep learning, such as ArcFace and CosFace, have achieved an accuracy rate of over 99.5% on the LFW dataset, and deep learning has achieved great success in face recognition research. However, the methods based on deep learning still cannot solve the influence brought by uncontrollable environmental factors such as environmental illumination, face pose, and partial occlusion. Among them, facial occlusion is one of the most challenging problems in face recognition algorithms. Some previous studies have dealt with a series of occlusion scenarios with a relatively small occlusion area, such as glasses and facial ornaments. However, the mask causes the absence of two inherent facial structures, the nose and the mouth, and also brings more noise to the position information of facial key points. Half of the key facial features are hidden. Compared with other occluders, mask occlusion poses a greater challenge to face recognition algorithms. Therefore, face recognition with mask occlusion is a difficult point in occlusion face recognition. Summary of the Invention

[0004] Currently, the existing deep learning-based occlusion face recognition algorithms still have the following problems: lack of datasets; large number of model parameters, which is not convenient for deployment on embedded and mobile devices; low recognition rate; and inability to detect whether a face is wearing a mask. To address the above problems, the present invention provides a method for detecting and recognizing a face occluded by a mask by integrating an attention mechanism, which mainly includes five parts: The first part is to perform data augmentation processing on the unoccluded face dataset by adding virtual masks; the second part is to extract and fuse features of the face image; the third part is to focus on local organs of the extracted face features; the fourth part is to use knowledge distillation to compress the number of model parameters; the fifth part is network training and testing to detect whether a face is wearing a mask and identify the face identity.

[0005] The first part includes two steps:

[0006] Step 1: Download the publicly available face recognition dataset. At this time, the dataset contains only normal, unoccluded faces. Next, use the face key point detection algorithm to extract 468 key points of the face, and select the coordinates of five key points from them for affine transformation to align and crop the face to obtain the original samples;

[0007] Step 2: Establish a data enhancer that adds virtual mask occlusion. This data enhancer obtains the 468 face key point coordinates obtained in Step 1, filters the lower half of the key points that will be occluded by the mask according to the index from the 468 face key point coordinates, and divides the face into multiple meshes by performing Delaunay triangulation based on these key points; perform triangulation on various styles of masks to obtain meshes corresponding to the positions of the original sample faces; map the masks to the meshes at the corresponding positions of the face by affine transformation grid by grid, and finally input the face after data enhancement into the network as training samples;

[0008] The second part includes two steps:

[0009] Step 3: Input the training samples obtained in Step 2 into the improved Swin Transformer backbone for feature extraction to obtain the preliminary extracted feature map of the face I ;

[0010] Step 4: Pass the preliminary extracted feature map obtained in Step 3 I into the subsequent Face Organ Attention (FOA) mechanism to focus on the face organs not occluded by the mask. See the third part for details;

[0011] The third part includes four steps:

[0012] Step 5: Convert the preliminary extracted feature map in Step 4 I ∈ℝ (H×W)×C to a three-dimensional feature map G ∈ℝ H×W×C and then set the pooling kernel scale to K h , and the stride to S h to perform average pooling along the horizontal direction; set the pooling kernel scale to K w , and the stride to S w to perform average pooling along the vertical direction to obtain the concentrated features W avg ∈ℝ 1×H / Windowsize×C 、 H avg ∈ℝ H / Windowsize×1×C where Windowsizerepresents the number of windows into which the feature map will be divided. In the present invention, ∈ℝ H×W×C is used to describe the scale size of the feature map. H, W, and C respectively represent the height, width, and number of channels of the feature map;

[0013] Step 6: Concatenate the concentrated features obtained in Step 5 W avg and H avg to obtain M . Set a hyperparameter r such that M obtains a feature layer M 1 after 2D convolution of 1×1. Next, insert a BN layer and a GELU activation function to obtain a feature layer M 2 . At this time, M 2 simultaneously has the feature concentration of the input feature G on the x-axis and y-axis;

[0014] Step 7: Divide and transpose the M 2 mixed with spatial position information in Step 6, and then change it back to the number of channels c after passing through 2D convolution of 1×1 again. The parameters of these two feature layers represent the spatial weights. Finally, multiply the corresponding position elements of W′, H′ and W′, H′ with G to obtain G′ , that is, superimpose the spatial weights on the input feature layer;

[0015] Step 8: Add FOA attention to the Transformer Layer, and use the subsequent joint loss function to supervise the model to adaptively adjust the window weights;

[0016] The fourth part includes two steps:

[0017] Step 9: Train a teacher model with a large number of parameters (embed_dim is 96), and then calculate the cosine distance between the output of this teacher model and the student model (embed_dim is 48) to obtain the cosine loss to guide the output feature vector of the student model to approximate the output feature vector of the teacher model;

[0018] Step 10: Add the cosine loss obtained in Step 9 during the training of the student model to guide the output feature vector of the student model to approximate the output feature vector of the teacher model;

[0019] The fifth part includes two steps:

[0020] Step 11: Debug the hyperparameters of the network structure from Step 3 to Step 10, and obtain the final teacher model and student model;

[0021] Step 12: Input the test set into the training model in Step 11 to perform the detection and recognition of occluded faces.

[0022] The present invention provides a method for detecting and recognizing masked faces by integrating an attention mechanism. First, the network improves Swin Transformer for face feature extraction. Second, a face organ attention mechanism FOA is proposed to make the model focus on the face organs not occluded by the mask. Then, aiming at the problem of insufficient current masked face datasets, a data augmentation method using 3D face meshes to generate masked data is proposed. Finally, aiming at the problem of a large number of model parameters, a method for compressing the model using knowledge distillation is proposed. This method better balances speed and accuracy, achieves excellent performance in detecting and recognizing masked faces, and has wide applicability. Description of the Drawings

[0023] Figure 1 It is a data augmenter diagram of the present invention;

[0024] Figure 2 It is an overall network structure diagram of the present invention;

[0025] Figure 3 It is a face organ attention mechanism diagram of the present invention;

[0026] Figure 4 It is a knowledge distillation structure diagram of the present invention;

[0027] Figure 5 It is a result diagram of detecting and recognizing occluded faces using the present invention. Detailed Embodiments

[0028] To better understand the present invention, the method for detecting and recognizing masked faces by integrating an attention mechanism of the present invention will be described in more detail below in conjunction with specific embodiments. In the following descriptions, the detailed descriptions of the current prior art may dilute the subject matter of the present invention, and these descriptions will be ignored here.

[0029] Step 1: Download a publicly available face recognition dataset. At this time, the dataset contains only normal unoccluded faces. Next, use a face key point detection algorithm to extract 468 key points of the face, and screen out the coordinates of five key points: the left eye, the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth for affine transformation. After aligning and cropping the face, obtain 101 training set samples;

[0030] Step 2: Establish a data augmenter for adding virtual mask occlusion. The data augmenter is asFigure 1 As shown, the data enhancer obtains the 468 facial key-point coordinates obtained in step 1, screens the lower half of the key points that will be blocked by the mask according to the index from the 468 facial key-point coordinates, and performs Delaunay triangulation based on these key points to divide the face into multiple grids; the masks of various styles are also triangulated to obtain grids corresponding to the positions of the faces; the masks are mapped to the grids at the corresponding positions of the faces through affine transformation grid by grid. Finally, the data enhancer generates four types of faces, namely, faces without wearing a mask (without any processing), faces wearing a mask correctly, faces wearing a mask but showing the nose, and faces wearing a mask but showing the nose and mouth, in a ratio of 3:1:1:1 for the input face 102, and inputs them into the network for training;

[0031] Figure 2 It is the specific network model diagram of the mask-covered face detection and recognition method integrating the attention mechanism of the present invention. In this implementation scheme, it is carried out according to the following steps:

[0032] Step 3: The network takes out the face pictures 201 that are not blocked by the mask from the dataset. Next, the faces are aligned through step 1, and then the data enhancer is used to generate four types of faces, namely, faces without wearing a mask (without any processing), faces wearing a mask correctly, faces wearing a mask but showing the nose, and faces wearing a mask but showing the nose and mouth 202. Next, the data-augmented face images 202 are input into the improved Swin Transformer backbone feature extraction network to obtain the preliminary extraction feature map of the face I ;

[0033] Step 4: The improved Swin Transformer backbone feature extraction network is used to perform preliminary feature extraction from 202 to obtain the preliminary extraction feature map of the face I , and the specific implementation is as follows:

[0034] Step 4-1: The data enhancer enables the model to load face pictures of people wearing masks with different degrees of occlusion and face pictures of people not wearing masks from the dataset at the same time. Next, the face pictures 202 are input into the backbone feature extraction network. The input scale of the original Swin Transformer is G ∈ℝ 224×224×3 , however, in the field of face recognition, the network can fully extract the required features from the face tensor of F ∈ℝ 112×112×3 scale. A larger-scale design will lead to performance redundancy in the network. Therefore, the network keeps the window size Windowsize as 7, and changes the number of Transformer blocks from [2, 2, 6, 2] to [2, 6, 2] to ensure that the scale size of the Transformer block before the fully connected layer is L ∈ℝ7×7×4C , the present invention uses ∈ℝ H×W×C symbols to describe the scale size of the feature map, where H, W, and C represent the height, width, and number of channels of the feature map respectively;

[0035] Step 4-2, since the Transformer requires the input to be a token vector, and face images are all three-dimensional tensors, the image 202 needs to be input into the Patch Partition 203 for block processing. First, the face image is divided into (4, 4) non-overlapping blocks 209. Next, use Linear Embedding 204 to expand these blocks along the Channel direction. The scale of the input face image is F ∈ℝ 112×112×3 , then the scale of the obtained feature map after expansion is S ∈ℝ 112 / 4×112 / 4×48 , next, use Conv2D to adjust the number of channels of this feature map from 48 to embed_dim. The present invention sets embed_dim to 96. The scale of the feature map at this time is S′ ∈ℝ 112 / 4×112 / 4×96 , then flatten the feature map along the width and height dimensions, and its scale also becomes E ∈ℝ 784×96 , next, input the encoded two-dimensional tensor into the RDSTL 205;

[0036] Step 4-3, the RDSTL is as Figure 2 shown in the lower right part, where l and z l respectively represent the output features of the (S)W-MSA module and the MLP module in the th STB; W-MSA and SW-MSA respectively represent the conventional multi-head self-attention module and the sliding window multi-head self-attention module. And the STBs always exist in pairs. The difference is that W-MSA is used for odd numbers and SW-MSA is used for even numbers. Therefore, the number of STBs in the RDSTL is always an even number. The calculation of the RDSTL is shown in formula (1);

[0037] (1)

[0038] Step 5, after the self-attention calculation in the RDSTL, it is also necessary to input it into the face organ attention mechanism (Face Organ Attention, FOA) 210 proposed by the present invention to focus on the unoccluded face organs. The specific structure of the FOA is as Figure 3 shown. The FOA converts the input two-dimensional feature map I ∈ℝ (H×W)×C 301 into a three-dimensional feature mapG ∈ℝ H×W×C Set the pooling kernel scale to K h , and the stride to S h Perform average pooling along the horizontal direction, with the pooling kernel scale being K w , and the stride being S w Perform average pooling along the vertical direction, where:

[0039] (2)

[0040] Obtain the concentrated features W avg ∈ℝ 1×H / Windowsize×C , H avg ∈ℝ H / Windowsize×1×C . Next, aggregate the features of the two concentrated feature layers;

[0041] Step 6, first concatenate the concentrated features W avg and H avg . Since the dimensions between the features W avg and H avg do not match, the width and height dimensions of the feature W avg need to be transposed and then concatenated with H avg to obtain the feature layer M . Set a hyperparameter r such that M obtains the feature layer M 1 after 2D convolution of 1×1, and its number of channels changes from c to c / r. In the present invention, r = 32 and M 1 's number of channels shall not be less than 8. Next, insert a BN layer and a GELU activation function to obtain the feature layer M 2 . At this time, M 2 simultaneously has the feature concentration of the input feature G on the x-axis and y-axis. Therefore, the spatial information of the input feature G is interacted;

[0042] Step 7, split and transpose the M 2 mixed with spatial position information, and then change it back to the number of channels c after passing through 2D convolution of 1×1W′, H′ , the parameters of these two feature layers represent the spatial weights. Finally, W′ , H′ and G are multiplied by the corresponding elements of the matrix to obtain G′ , that is, the spatial weights are superimposed on the input feature layer. Therefore, G the spatial weights beneficial to the recognition task in

[0043] are increased. In the figure, ⊙ represents the pairwise multiplication of the corresponding elements of the matrix and the input matrix:

[0044] (3)

[0045] Step 8, as Figure 2 shown, after the feature map of the face is extracted by the backbone feature extraction network, the model uses a fully connected layer to obtain a face feature vector of size 18816, and then the face feature vector of length 512 extracted by the face identity classifier 207 x After that, the present invention uses ArcFace as the loss function to map the face feature vector x onto the hypersphere and compress the cosine distance of the same face feature vector x and expand the cosine distance of different face feature vectors x :

[0046] (4)

[0047] Among them N is the number of samples (number of face images), n is the number of classes (number of face identities), s is the radius of the hypersphere, θ is the weight W and the face feature vector x The angle between them, ArcFace further increases the cosine interval between different face features by adding a margin θ on this angle m . This can make the features learned by the model have stronger discriminative ability. The ArcFace loss function is obtained by calculating the cross entropy between ArcFace and the face labels:

[0048] (5)

[0049] In step 9, the mask wearing classifier 208 extracts a feature vector of length 4 x′ , and the values of this vector after Softmax calculation respectively correspond to the probabilities of a face without a mask, correctly wearing a mask, showing the nose, and showing the nose and mouth. By x′ calculating the cross entropy with the mask wearing label, the Mask loss function is obtained, and it is combined with the ArcFace joint auxiliary model to identify whether a face is correctly wearing a mask:

[0050] (6)

[0051] In step 10, during the training process of the face recognition model, only labels a, b, and c are input to supervise the model to learn face features. However, these labels cannot reflect the similarity degree between each face. Therefore, the present invention uses the idea of knowledge distillation to make up for the problem that the label information cannot characterize the face similarity, enriches the information contained in the labels, and compresses the number of parameters of the model. The method of knowledge distillation is as Figure 4 shown. First, a teacher model with a larger number of parameters (embed_dim is 96) is trained, and then the cosine distance between the output of this teacher model and the output of the student model (embed_dim is 48) is calculated to obtain L Face to guide the output feature vector of the student model to approximate the output feature vector of the teacher model. Finally, when training the student model, L Face is added to guide the output feature vector of the student model to approximate the output feature vector of the teacher model. The present invention sets L Face with a weight of 100:

[0052] (7)

[0053] where and are respectively 512-dimensional face feature vectors output by the teacher model and the student model, and ϵ is set to a very small value of 1e-8 to avoid division by zero.

[0054] In step 11, the minimum batch size is set to 64, the total number of epochs is 20, and the Adam optimizer with patience of 4 and an initial learning rate of is used to train the model according to steps 3 to 10.

[0055] Step 12: The method is verified on the unoccluded face dataset LFW, the virtual mask occluded face recognition dataset MLFW, and the real mask occluded dataset MWHN respectively. The accuracy of the teacher model on the LFW test set is 99.62%, on the virtual mask occluded test set MLFW is 99.10%, and on the real mask occluded dataset MWHN is 85.60%. The accuracy of the student model after using knowledge distillation on the LFW test set is 99.56%, on the virtual mask occluded test set MLFW is 99.03%, and on the real mask occluded dataset MWHN is 83.96%. The number of parameters of the model is reduced from 70.27MB of the teacher network to 26.85MB, which fully proves the effectiveness of this method. The final occlusion face detection and recognition effect of the network is as Figure 5 shown.

[0056] The present invention provides a method for mask-occluded face detection and recognition that integrates an attention mechanism. First, the network improves Swin Transformer for face feature extraction; second, a face organ attention mechanism FOA is proposed to make the model focus on the face organs not occluded by the mask; then, aiming at the problem of insufficient current mask-occluded face datasets, a data augmentation method using 3D face meshes to generate data with added mask occlusion is proposed. Finally, aiming at the problem of a large number of model parameters, a method of compressing the model using knowledge distillation is proposed. This method better balances speed and accuracy, achieves excellent performance in mask-occluded face detection and recognition, and has wide applicability.

[0057] Although the above describes the illustrative specific embodiments of the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

Claims

1. A method for detecting and recognizing a face occluded by a mask that integrates an attention mechanism, characterized in that, Adopt the face organ attention mechanism FOA to make the model focus on the face organs not blocked by the mask, including five parts: data augmentation processing of adding virtual masks to the unobstructed face dataset, feature extraction and fusion of face images, local organ attention to the extracted face features, using knowledge distillation to compress the number of model parameters, and network training and testing. The first part includes two steps: Step 1: Download the publicly available face recognition dataset. At this time, the dataset contains only normal unobstructed faces. Next, use the face key point detection algorithm to extract 468 key points of the face, and select the coordinates of five key points from them for affine transformation to align and crop the face to obtain the original sample. Step 2: Establish a data augmenter for adding virtual mask occlusion. This data augmenter obtains the 468 face key point coordinates obtained in Step 1, filters the lower half of the key points that will be blocked by the mask according to the index from the 468 face key point coordinates, and divides the face into multiple grids according to these key points by Delaunay triangulation; triangulate various styles of masks to obtain grids corresponding to the positions of the original sample faces; map the masks to the grids at the corresponding positions of the faces by affine transformation grid by grid, and finally input the face after data augmentation into the network as a training sample. The second part includes two steps: Step 3: Input the training sample obtained in Step 2 into the improved Swin Transformer backbone for feature extraction to obtain the preliminary extracted feature map I of the face. Step 4: Pass the preliminary extracted feature map obtained in Step 3 into the subsequent face organ attention mechanism FOA to focus on the face organs not blocked by the mask. See the third part for details. The third part includes four steps: Step 5: Convert the preliminary extracted feature map in Step 3 into a three-dimensional feature map and then set the pooling kernel scale to K h , with a stride of S h Perform average pooling along the horizontal direction with a pooling kernel scale of K w , with a stride of S w Perform average pooling along the vertical direction to obtain the concentrated features respectively where Windowsize represents the number of windows into which the feature map will be divided, and the symbol is used to describe the scale size of the feature map. H, W, and C represent the height, width, and number of channels of the feature map respectively; Step 6: Concatenate the concentrated features W avg and H avg to obtain M. Set a hyperparameter r such that after M passes through a 2D convolution of 1×1, the feature layer M 1 is obtained. Next, insert a BN layer and a GELU activation function to obtain the feature layer M 2 . At this time, M 2 simultaneously has the concentrated features of the input feature G on the x-axis and y-axis; Step 7: Split and transpose M with spatial location information mixed in Step 6, and then pass it through a 1×1 2D convolution again to change it back to W' and H' with the number of channels being c. The parameters of these two feature layers represent the spatial weights. Finally, multiply the corresponding elements of W' and H' with the elements of the G matrix to obtain G', that is, superimpose the spatial weights on the input feature layer; 2 ​ Step 8: Add the FOA attention to the TransformerLayer and use the subsequent joint loss function to supervise the model to adaptively adjust the window weights. The fourth part includes two steps: Step 9: Train a teacher model with a large number of parameters, where embed_dim is 96. Then, calculate the cosine distance between the output of this teacher model and the student model with embed_dim of 48 to obtain the cosine loss to guide the output feature vector of the student model to approximate the output feature vector of the teacher model. Step 10: Add the cosine loss obtained in Step 9 during the training of the student model to guide the output feature vector of the student model to approximate the output feature vector of the teacher model. The fifth part includes two steps: Step 11, debug the hyperparameters of the network structure from Step 3 to Step 10. Among them, set the minimum batch size to 64, the total number of epochs to 20, and use the Adam optimizer with a patience of 4 and an initial learning rate of 10 to train the model according to Steps 3 to 10, and obtain the final teacher model and student model; -3 ​ Step 12: Input the test set into the training model in Step 11 and verify the method on the unobstructed face dataset LFW, the virtual mask occlusion face recognition dataset MLFW, and the real mask occlusion dataset MWHN respectively.

2. The method for detecting and recognizing a face blocked by a mask integrating an attention mechanism according to claim 1, wherein, In Step 5, the face organ attention mechanism FOA is used to focus on the face organs not blocked by the mask.

3. A method for detecting and recognizing a face blocked by a mask by integrating an attention mechanism according to claim 1, characterized in that, In Step 9, knowledge distillation is used to compress the number of model parameters.

Citation Information

Patent Citations

  • Mask-wearing face recognition method and device, electronic equipment and storage medium

    CN112200154A

  • Knowledge distillation network-based mask face shielding recognition method and device, and equipment

    CN113343898A