A mask face recognition method based on dual-branch optimization of local block attention
Through the method based on the optimization of the double branch of attention based on local block attention, the problem of facial features occlusion affecting the performance of the recognition system due to wearing a mask is solved, and efficient mask facial recognition is achieved, and the accuracy of universal facial recognition is maintained.
Patent Information
- Application Number
- CN202111500297.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-09
AI Technical Summary
Wearing a mask causes facial features to be blocked on a large scale, affecting the accuracy of the recognition system. It is difficult for the existing technology to effectively solve this problem.
The method based on local block attention dual branch optimization is adopted to simulate the face image of the mask through face key point detection and affine transformation, and build the local block attention module, and embed it into the deep convolution network to optimize the feature extraction with the dual branch structure.
The performance of mask facial recognition is improved, the attention to the characteristics of the upper half of the unblocked area is enhanced, and the overall perception of the face contour is maintained, achieving efficient mask facial recognition.
Smart Images

Figure CN114120426B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face recognition, and in particular to a mask face recognition method based on local block attention dual-branch optimization. Background Art
[0002] Biometric recognition technology uses the physiological or behavioral characteristics of organisms to identify individuals. Among them, face recognition technology, which has been relatively mature, has been applied in many important fields. However, partial occlusion of the face is one of the challenging problems in practical open scenarios, which will cause the loss of identity information and greatly reduce the accuracy of the recognition system. The commonly used processing method is to use the face quality estimation module to filter partially occluded faces. Wearing a mask is a common requirement in hospitals, factories, and public places under respiratory infectious epidemics. It will cause a large area of the lower half of the face to be occluded and the facial features to be covered. Directly identifying the face of a person wearing a mask cannot be solved using general face recognition methods. Therefore, studying face recognition while wearing a mask is a challenging and urgent problem to be solved.
[0003] The face recognition system mainly includes the following important modules: Camera acquisition: The face image is captured by the acquisition device, which includes monocular, binocular, and depth cameras. Face images of different qualities and types can be obtained according to the acquisition device; Face detection: Use the face detection network to locate the face frame, capture the face area and set it to a pre-defined appropriate size; Key point positioning and correction: Detect the face key points, calibrate the face image using affine transformation, and improve subsequent recognition performance; Feature extraction: Extract effective features to reduce the dimensionality of the original image pattern space, combine with the currently popular deep learning technology, train the face feature extraction network, and extract highly cohesive face deep features from the face image; Feature matching: Make decision classifications based on the extracted features, and the classification space is based on the feature extraction method in the previous step, and divides the decision recognition results by metric functions and thresholds.
[0004] Among them, the most critical part of the mask face recognition task is to train a feature extraction model that can extract discriminative features from the face image occluded by the mask. The patent of this invention proposes an effective method based on the local block attention dual-branch optimization face feature extraction model to improve the performance of mask face recognition.
[0005] Existing technical solutions:
[0006] Methods for optimizing general face recognition systems include patent CN112818901A for optimizing system logic and patent CN112200108A for optimizing model structure and training data.
[0007] The above-mentioned specific patent reference documents and related literature are:
[0008] 1) "A method for face recognition with masks based on eye attention mechanism", patent number CN112818901A, after detecting the face image, first perform mask detection to determine whether the person is wearing a mask, then extract the features of the face with or without a mask according to whether the person is wearing a mask, and then perform feature comparison. This method optimizes the logic of the face recognition system, extracting different features according to whether the person is wearing a mask, but does not optimize the feature extraction model accordingly, and the intra-class aggregation of the extracted features is reduced due to the interference of mask occlusion.
[0009] 2) "A method for face recognition with masks", patent number CN112200108A, is to embed the CBAM attention module on the basis of the ResNet101 deep feature extraction model, optimize the model structure and perform data amplification at the same time, mix normal face images and simulated mask image data for model training, and finally perform distillation learning to obtain a lightweight model. This method modifies the model structure to a general attention module CBAM, and does not optimize for the characteristics of masks blocking fixed areas, which limits the performance of the feature extraction model. Summary of the invention
[0010] In order to solve the technical problem that the performance of the recognition system is affected by the large-area occlusion of facial features when wearing a mask, especially the large intra-class variance between masked faces and normal faces, which brings a series of complex and challenging technical difficulties to the recognition system, the purpose of the present invention is to provide a mask face recognition method based on local block attention dual-branch optimization. The method simulates the face image of a person wearing a mask after facial key point detection, optimizes the convolutional neural network in a dual-branch manner, and maintains the overall perception of the face contour while enhancing the attention of the face feature extraction network to the upper unobstructed area, thereby realizing efficient mask face recognition.
[0011] The purpose of the present invention is achieved through the following technical solutions:
[0012] A mask face recognition method based on local block attention dual-branch optimization, comprising:
[0013] Step 10: collecting face RGB image data;
[0014] Step 20 estimates the face deflection direction based on dlib face key point detection;
[0015] Step 30: Based on the detected facial key points and the estimated facial deflection direction, the mask image is affine transformed to fit the face to simulate the mask occlusion image;
[0016] Step 40 constructs a local block attention module;
[0017] Step 50 embeds the local block attention module into the deep convolutional network, uses the deep convolutional network as the backbone of the face feature extraction module to extract the global features of the face, and derives local branches to extract local features;
[0018] Step 60 performs mask face recognition by matching the global features of the global branch of the deep convolutional network.
[0019] Compared with the prior art, one or more embodiments of the present invention may have the following advantages:
[0020] The technical solution of the present invention optimizes the convolutional neural network in a dual-branch manner, while enhancing the focus of the face feature extraction network on the upper unobstructed area, maintaining the overall perception of the face contour, and realizing efficient mask face recognition. The feature extraction model improves the performance of mask face recognition while maintaining the accuracy of general face recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of the mask face recognition method based on the optimization of the local block attention dual branch;
[0022] Figure 2 It is a schematic diagram of the structure of the local block attention module;
[0023] Figure 3 It is the overall framework diagram of dual-branch joint optimization. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below in conjunction with embodiments and drawings.
[0025] like Figure 1 As shown in FIG. 1 , a mask face recognition method process based on local block attention dual-branch optimization is shown, which includes the following steps:
[0026] Step 10: collecting face RGB image data;
[0027] Collect normal unobstructed face RGB image data, including multiple identity IDs, each identity contains at least 5 images, and construct a normal face image dataset.
[0028] Step 20 estimates the face deflection direction based on dlib face key point detection;
[0029] While using dlib to locate and detect the face and crop the face image, key point detection is performed to estimate the deflection direction of the face.
[0030] Step 30: Based on the detected facial key points and the estimated facial deflection direction, the mask image is affine transformed to fit the face to simulate the mask occlusion image;
[0031] According to the detected facial key points and the estimated face deflection direction, the mask image of the front face is affine transformed to the corresponding deflection direction. According to the facial contour and the position of the key points of the nose tip and mouth corners, the mask image is fit to the unobstructed face image to simulate the face image data of people wearing masks. A mixed mask face dataset is formed with the collected normal face image dataset, which includes both face data of people wearing and without masks.
[0032] Step 40 constructs a local block attention module;
[0033] A local block attention module is constructed to be embedded in the deep convolutional network. As a local branch, it focuses on extracting features of the upper half of the face, enhancing the model's feature extraction of the upper half of the face to obtain facial feature encoding that is robust to mask occlusion. Figure 2 shown).
[0034] Specifically, the feature map of the middle layer of the cropped convolutional network is taken out of the upper feature map as the input of the local block attention module. On the one hand, the upper feature map is directly input into the convolution layer, regularization layer, activation layer, feature mapping layer and weighting layer in sequence for feature extraction; on the other hand, the upper feature map is divided into 16 local blocks in a sliding window manner, and the local feature blocks are input into the convolution layer, regularization layer, activation layer, feature mapping layer and weighting layer in sequence for feature extraction. Finally, multiple block codes are spliced together to become a high-dimensional feature, which is input into the fully connected layer to reduce the dimension into a 512-dimensional local code as a local feature representation.
[0035] Step 50 embeds the local block attention module into the deep convolutional network, uses the deep convolutional network as the backbone of the face feature extraction module to extract the global features of the face, and derives local branches to extract local features;
[0036] The overall framework of the face deep feature extraction network with dual-branch joint optimization is as follows Figure 3As shown in the figure, the ResNet50 deep convolutional network is used as the backbone of the face feature extraction module to extract global face features. After the constructed local block attention is embedded in the second residual block of the backbone network, the local branch is introduced, and the feature map of the middle layer is used as input to extract local features, guiding the shared network to focus on specific areas for feature extraction, focusing on facial features such as the periorbital area and forehead that are not blocked by the mask. The entire network forms a dual-branch structure, and the global and local features output by the global and local branches are connected to independent classification heads, and jointly optimized on the mixed mask face dataset. The loss function is the weighted dual-branch arcface classification loss:
[0037] L total =L global +λ*L local
[0038]
[0039]
[0040] Among them, cosθ j is the angle between the class center vector and the face depth feature, m is the angle margin, and λ is the weight coefficient of the local branch, which is used to balance the influence of the global branch and the local branch on the shared network during training.
[0041] Step 60 performs mask face recognition by matching the global features of the global branch of the deep convolutional network.
[0042] During the recognition process, the masked face image or face image is input into the deep feature extraction network, and only the global features are used for feature matching and identity recognition. The local block attention has increased the backbone network's attention to the upper half of the face features during the dual-branch joint optimization, and obtained a face feature representation that is robust to mask occlusion.
[0043] The local block attention module structure is a training method that embeds the module into the main deep convolutional network and performs dual-branch optimization. The above two technical keys can enhance the attention of the face feature extraction network to the upper unobstructed area while maintaining the overall perception of the face contour, achieving efficient mask face recognition. The feature extraction model improves the performance of mask face recognition while maintaining the accuracy of general face recognition.
[0044] The method of simulating masks in the above embodiment can be replaced by other mapping methods. The core is to build a training set of mixed normal face images and simulated mask face images for subsequent training of the local block attention model. The backbone network in the feature extraction model can replace other deep convolutional networks, such as ResNet50 and ResNet101.
[0045] Although the embodiments disclosed in the present invention are as above, the above contents are only embodiments adopted for facilitating the understanding of the present invention and are not intended to limit the present invention. Any technician in the technical field to which the present invention belongs can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the present invention, but the patent protection scope of the present invention shall still be subject to the scope defined in the attached claims.
Claims
1. A mask face recognition method based on local block attention dual-branch optimization, It is characterized in that The method comprises the following steps: Step 10: collecting face RGB image data; Step 20 estimates the face deflection direction based on dlib face key point detection; Step 30: Based on the detected facial key points and the estimated facial deflection direction, the mask image is affine transformed to fit the face to simulate the mask occlusion image; Step 40 constructs a local block attention module; Step 50 embeds the local block attention module into the deep convolutional network, uses the deep convolutional network as the backbone of the face feature extraction module to extract the global features of the face, and derives local branches to extract local features; Step 60 performs mask face recognition by matching the global features of the global branch of the deep convolutional network; The step 40 specifically includes: cropping the middle layer feature map of the convolutional network, taking out the upper half feature map as the input of the local block attention module, and directly extracting features from the upper half feature map and extracting features by block mode, splicing multiple block codes to form a high-dimensional feature, and inputting it to the fully connected layer to reduce the dimension into a local code as a local feature; The direct extraction of the upper half feature map includes: inputting the upper half feature map into a convolution layer, a regularization layer, an activation layer, a feature mapping layer and a weighted layer in sequence for feature extraction; The block-based feature extraction includes: dividing the upper half feature map into 16 local blocks in a sliding window manner, and the local feature blocks are sequentially input into the convolution layer, regularization layer, activation layer, feature mapping layer and weighted layer for feature extraction; The step 50 specifically includes: using the deep convolutional network as the backbone of the face feature extraction module to extract the global features of the face; after embedding the local block attention into the second residual block of the backbone network, a local branch is introduced, and the feature map of the middle layer is used as input to extract the local features. The entire network forms a dual-branch feature, and the global features and local features output by the global and local branches are connected to independent classification heads, and joint optimization is performed on the mixed mask face data set. The loss function is a weighted dual-branch arcface classification loss: L total =L global +λ·K local Among them, θ j is the angle between the class center vector and the face depth feature, m is the angle margin, and λ is the weighting coefficient of the local branch, which is used to balance the influence of the global branch and the local branch on the shared network during training.
2. The mask face recognition method based on local block attention dual branch optimization according to claim 1, It is characterized in that The step 10 specifically includes: collecting normal unobstructed human face RGB image data, including multiple identity IDs, each identity containing at least 5 images, and constructing a normal human face image dataset.
3. The mask face recognition method based on local block attention dual branch optimization according to claim 1, It is characterized in that In step 20, dlib is used to locate and detect the face and crop the face image while performing key point detection.
4. The mask face recognition method based on local block attention dual branch optimization according to claim 1, It is characterized in that The step 30 specifically includes performing an affine transformation on the mask image of the front face to transform it to the corresponding deflection direction, fitting the mask image to the unobstructed face image according to the facial contour and the key point positions of the nose tip and mouth corners, simulating the face image data of people wearing masks, and forming a mixed mask face data set with the collected normal face image data set.
5. The mask face recognition method based on local block attention dual branch optimization according to claim 1, It is characterized in that The step 60 specifically includes: during the recognition process, inputting a mask face image or a face image into a deep feature extraction network, and using global features for feature matching to identify the identity; the local block attention improves the backbone network's attention to the upper face features during the dual-branch joint optimization, and obtains a face feature representation that is robust to mask occlusion.
Citation Information
Patent Citations
Mask-wearing face recognition method based on eye attention mechanism
CN112818901A
Multi-scale face age estimation method and system embedded with high-order information
CN111814611A
Mask face recognition method
CN112200108A