Face recognition method and system
Through the feature extraction method that integrates multi-scale blocked CS-LBP features and weighted PCA features, combined with feature enhancement networks and face recognition networks, the problem of low accuracy of occluding face recognition in the prior art is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202310352487.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-04
AI Technical Summary
The prior art has problems in facial recognition that the data set of masking faces is not rich enough, the model training is difficult, the stability is not high, and the loss function is greatly affected. Especially when face recognition and low-resolution image processing are used in facial recognition, the accuracy rate is low.
A feature extraction operator that fuses multi-scale blocked CS-LBP features and weighted PCA features is adopted to extract features of the face images to be recognized through the feature extraction network, and the feature enhancement network is used to enhance the features of the non-occluded area, and to recognize them through the face recognition network.
It improves the accuracy of masking face recognition, enhances the robustness and classification ability of face features at low resolution, reduces interference from environmental factors, and improves the accuracy of recognition results.
Smart Images

Figure CN116486452B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of face recognition, and particularly to face recognition methods and systems. Background Art
[0002] Face recognition technology has become increasingly common in daily life. For example, face recognition technology is required in scenarios such as road monitoring, mobile phone unlocking, and security inspections. However, if the face is blocked by a mask, the accuracy of face recognition technology will decrease.
[0003] In the prior art, the face recognition problem based on occlusion is mainly divided into two research ideas, namely highlighting the face area in the image and weakening the non-face background area in the image, and trying to expand the dataset used for model training and testing to improve the recognition effect.
[0004] The prior art has a significant effect on occluded face recognition. However, for face recognition of people wearing masks intercepted in the scenario of crowd passage, the occluded face dataset is still not rich enough, and the model of face recognition technology has problems such as large training difficulty, low stability, and being greatly affected by the loss function. At the same time, since the mask-blocked image is intercepted by a monitor, the resolution of the intercepted image by the monitor is low and the size is small, resulting in low face recognition accuracy. Summary of the Invention
[0005] This application provides a face recognition method and system for improving the accuracy of occluded face recognition.
[0006] In a first aspect, this application provides a face recognition method, which includes:
[0007] Obtain a face image to be recognized, where the face image to be recognized includes an occluded area and a non-occluded area;
[0008] Extract features from the face image to be recognized through a feature extraction network in the trained occluded face recognition model to obtain a first feature map;
[0009] Enhance the features of the non-occluded area in the first feature map through a feature enhancement network in the trained occluded face recognition model to obtain a second feature map;
[0010] Recognize the second feature map through a face recognition network in the trained occluded face recognition model to obtain the face recognition result of the face image to be recognized.
[0011] With the above technical solution, after obtaining the occluded face image to be recognized, the face image is subjected to feature extraction through the feature extraction network, the features of the non-occluded area are enhanced through the feature enhancement network, and the face recognition result is obtained through the face recognition network. The feature extraction network can enhance the robustness and classification ability of face features at low resolution, and the features of the non-occluded area of the face are more obvious. The feature enhancement network weights the features of the non-occluded area, so that the accuracy of the face recognition result is higher.
[0012] Preferably, the step of extracting the first feature map from the face image to be recognized through the feature extraction network in the trained occluded face recognition model specifically includes:
[0013] The face image to be recognized is divided into blocks to obtain a number of sub-blocks;
[0014] CS-LBP features are extracted from the sub-blocks to obtain the CS-LBP features of the sub-blocks;
[0015] Histogram statistics are performed on the CS-LBP features of the several sub-blocks to obtain image features;
[0016] The image features are subjected to weighted PCA dimensionality reduction processing to obtain the first feature map.
[0017] With the above technical solution, the face image to be recognized is divided into blocks, and the obtained sub-blocks are subjected to CS-LBP feature extraction. By extracting local textures, the robustness and classification ability of face features at low resolution are enhanced. Histogram statistics are performed on the CS-LBP features of the sub-blocks to obtain image features. The image features are subjected to dimensionality reduction processing through weighted PCA, which can effectively remove redundant information and reduce the interference caused by environmental factors. Thus, the problem that the resolution of the monitored intercepted image is low and the size is small, resulting in a low face recognition accuracy, is effectively solved, and the face recognition accuracy is further improved.
[0018] Preferably, the step of dividing the face image to be recognized into blocks to obtain a number of sub-blocks specifically includes:
[0019] The face image to be recognized is downsampled in a three-layer scale space through a Gaussian pyramid to obtain a sampled image;
[0020] The sampled image is divided into 2 2b (2 b ·2 b ) blocks to obtain 2 2b (2 b ·2 b ) sub-blocks, where b is the block division level.
[0021] With the above technical solution, after downsampling the face image to be recognized through a Gaussian pyramid in a three-layer scale space to obtain a sampled image, the sampled image is divided into 2 2b (2 b ·2 b ) blocks. Using an appropriate number of block levels can fully express the facial information in the local area represented by the block image and reduce noise during the processing.
[0022] Preferably, the step of performing CS-LBP feature extraction on the sub-blocks specifically includes:
[0023] Obtain the gray value of the central pixel point of the sub-block and the gray values of the surrounding pixel points, compare the gray value of the central pixel point with the gray values of the surrounding pixel points to obtain the CS-LBP coding value of the sub-block; the CS-LBP coding value of the sub-block is:
[0024]
[0025] where CS-LBP p,R,ε is the CS-LBP coding value of the sub-block, p is the number of pixel sampling points, g i is the pixel gray value, S(X) is the result of comparing the pixel gray values, and the expression of the S(X) is:
[0026]
[0027] With the above technical solution, obtain the gray value of the central pixel point of the sub-block and the gray values of the surrounding pixel points, compare the gray value of the central pixel point with the gray values of the surrounding pixel points to obtain the CS-LBP coding value of the sub-block, so as to capture the edge and prominent texture information in the image, and the comparison threshold of this solution is set smaller, making the CS-LBP feature more robust on the planar image.
[0028] Preferably, the step of performing weighted PCA dimensionality reduction processing on the image features specifically includes:
[0029] Convert the image features into a covariance matrix, and calculate the eigenvalues and eigenvectors of the covariance matrix;
[0030] Based on a preset reduced dimension, a preset weighted matrix, and the eigenvalues and eigenvectors of the covariance matrix, obtain a weighted mapping transformation matrix;
[0031] According to the weighted mapping transformation matrix and the eigenvectors of the covariance matrix, obtain a first feature map with a principal component projection matrix.
[0032] Adopting the above technical solution, the image features are converted into a covariance matrix, the eigenvalues and eigenvectors of the covariance matrix are calculated, a weighted mapping transformation matrix is obtained based on the preset dimension reduction, the preset weighted matrix, and the eigenvalues and eigenvectors of the covariance matrix, and a first feature map with a principal component projection matrix is obtained according to the weighted mapping transformation matrix and the eigenvectors of the covariance matrix; when the block series b is relatively large, the dimension of the face image features is relatively high and contains a lot of redundant information, and the face image is prone to cause relatively large differences among samples of the same class when some objective conditions change. This kind of difference has a greater impact on the principal components corresponding to larger eigenvalues, that is, the principal components corresponding to larger eigenvalues are more susceptible to environmental factors (such as illumination, pose changes, etc.). In this application, weighted PCA dimensionality reduction processing is performed on the image features, which can reduce the feature dimension and reduce the interference caused by environmental factors.
[0033] Preferably, the step of enhancing the features of the non-occluded area in the first feature map to obtain a second feature map through the feature enhancement network in the trained occluded face recognition model specifically includes:
[0034] Inputting the first feature map into the multi-attention module of the feature enhancement network to obtain the output result of the attention module, where the multi-attention module includes a channel attention module, a spatial attention module, and a global attention module;
[0035] Inputting the output result of the attention module into the feature vector extraction module of the feature enhancement network to obtain the second feature map.
[0036] Adopting the above technical solution, inputting the first feature map into the multi-attention module of the feature enhancement network to obtain the output result of the attention module, and inputting the output result of the attention module into the feature vector extraction module of the feature enhancement network to obtain the second feature map. The multi-attention module includes a channel attention module, a spatial attention module, and a global attention module. Introducing the spatial attention module makes full use of the information in the space of the occluded face, such as the texture information of the eyes, eyebrows, and forehead; at the same time, the feature maps generated by the global attention module, the spatial attention module, and the channel attention module are superimposed and fused, thereby enhancing the feature weights of the non-occluded area.
[0037] Preferably, the step of inputting the output result of the attention module into the feature vector extraction module of the feature enhancement network to obtain the second feature map specifically includes:
[0038] Inputting the output result of the attention module into the first depthwise separable convolutional layer to obtain the output result of the first depthwise separable convolutional layer;
[0039] Input the output result of the first depthwise separable convolutional layer into the bottleneck layer to obtain the output result of the bottleneck layer;
[0040] Input the output result of the bottleneck layer into the convolutional layer to obtain the output result of the convolutional layer;
[0041] Input the output result of the convolutional layer into the second depthwise separable convolutional layer to obtain the second feature map.
[0042] With the above technical solution, the feature enhancement network further includes a feature vector extraction module. Input the output result of the attention module into the feature vector extraction module. The feature vector extraction module uses a lightweight MobileNetV2 sub-network, which requires fewer samples, can retain more feature information, and improves the representation ability of the network.
[0043] Preferably, before obtaining the face image to be recognized, it further includes:
[0044] Train the initial occluded face recognition model based on the labeled sample images to obtain the trained occluded face recognition model. The loss function for training the initial occluded face recognition model is the ArcFace loss function.
[0045] With the above technical solution, before recognizing the face image to be recognized, train the initial occluded face recognition model to obtain the trained occluded face recognition model, and recognize the face image to be recognized through the trained occluded face recognition model; use the ArcFace loss function to train the initial occluded face recognition model. The ArcFace loss function can obtain stable performance without combining with other loss functions and can easily converge to any training dataset.
[0046] In a second aspect, the present application provides a face recognition system, which includes:
[0047] An image acquisition module for acquiring a face image to be recognized, where the face image to be recognized includes an occluded area and a non-occluded area; a feature extraction module for extracting features from the face image to be recognized through the feature extraction network in the trained occluded face recognition model to obtain a first feature map;
[0048] A feature enhancement module for enhancing the features of the non-occluded area in the first feature map through the feature enhancement network in the trained occluded face recognition model to obtain a second feature map;
[0049] A face recognition module for recognizing the second feature map through the face recognition network in the trained occluded face recognition model to obtain the face recognition result of the face image to be recognized.
[0050] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0051] 1. The present application uses a feature extraction operator that fuses multi-scale block CS-LBP features and weighted PCA features to extract the features of the image to be processed. By enhancing the robustness and classification ability of face features at low resolution through local texture extraction, and effectively removing redundant information and reducing interference caused by environmental factors through the weighted PCA method, it effectively solves the problem that the resolution of the monitored and intercepted images is low and the size is small, resulting in low face recognition accuracy, and thus improves the face recognition accuracy.
[0052] 2. The present application introduces spatial attention, making full use of partial information on the occluded face, such as the texture information of the eyes, eyebrows, and forehead. At the same time, the feature maps generated by the global attention module, spatial attention module, and channel attention module are superimposed and fused to enhance the feature weights of the non-occluded area, thereby increasing the recognition accuracy of the occluded face. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a schematic flowchart of the face recognition method in the embodiments of the present application;
[0054] Figure 2 is a schematic flowchart of the first feature map acquisition step in the embodiments of the present application;
[0055] Figure 3 is a schematic diagram for understanding the Gaussian pyramid in the embodiments of the present application;
[0056] Figure 4 is a schematic diagram for understanding the channel attention module in the embodiments of the present application;
[0057] Figure 5 is a schematic diagram for understanding the spatial attention module in the embodiments of the present application;
[0058] Figure 6 is a schematic diagram for understanding the global attention module in the embodiments of the present application;
[0059] Figure 7 is a schematic diagram for understanding the feature enhancement network in the embodiments of the present application;
[0060] Figure 8 is a schematic diagram for understanding the occluded face recognition model in the embodiments of the present application;
[0061] Figure 9 is a schematic diagram of the modules of the face recognition system in the embodiments of the present application.
[0062] Description of the reference numerals: 1. Image acquisition module; 2. Feature extraction module; 3. Feature enhancement module; 4. Face recognition module. Detailed implementation manners
[0063] The terms used in the following embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and appended claims of this application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term " / and" used in this application refers to any or all possible combinations including one or more of the listed items.
[0064] Hereinafter, the terms "first" and "second" are only for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0065] Embodiments of this application disclose a face recognition method.
[0066] Referring to Figure 1 , Figure 1 is a schematic flowchart of the face recognition method in the embodiments of this application. The specific steps of this method include:
[0067] S10: Obtain a face image to be recognized, where the face image to be recognized includes an occluded area and a non-occluded area.
[0068] The face image to be recognized is an occluded face image, and the face image to be recognized includes two areas, namely an occluded area and a non-occluded area. For example, the occluded area may be the mouth area and part of the nose area, and the non-occluded area may be other areas of the face.
[0069] Specifically, the steps for obtaining the face image to be recognized may be: intercept an initial occluded face image wearing a mask in a crowd passage scenario through a camera device, intercept the face image in the initial occluded face image through face detection, adjust the face angle in the initial occluded face image through face positioning, convert the initial occluded face image into an image of a specific size through image normalization, and grayscale the initial occluded face image through grayscale processing, so as to obtain the face image to be recognized.
[0070] S20: Extract features from the face image to be recognized through the feature extraction network in the trained occluded face recognition model to obtain a first feature map.
[0071] The trained occluded face recognition model includes a feature extraction network. After inputting the obtained face image to be recognized into the feature extraction network in the trained occluded face recognition model, a first feature map can be obtained.
[0072] Referring to Figure 2 , Figure 2 is a schematic flowchart of the first feature map acquisition step in an embodiment of the present application. S20 specifically includes S201 - S203.
[0073] S201: Divide the face image to be recognized into blocks to obtain a number of sub - blocks.
[0074] Among them, specifically dividing the face image to be recognized into blocks includes: performing downsampling on the face image to be recognized in three - layer scale space through a Gaussian pyramid to obtain a sampled image, and dividing the sampled image into 2 2b (2 b ·2 b ) blocks to obtain 2 2b (2 b ·2 b ) sub - blocks, where b is the block - dividing level.
[0075] The Gaussian pyramid continuously reduces the size of the face image to be recognized through Gaussian blur filtering and downsampling, so as to obtain images with multiple resolutions in the Gaussian pyramid, that is, the scales of the face images to be recognized in different layers of the Gaussian pyramid are different.
[0076] Referring to Figure 3 , Figure 3 is a schematic diagram for understanding the Gaussian pyramid in an embodiment of the present application. σ is the scale - space coordinate. In an embodiment of the present application, it is necessary to perform downsampling on the face image to be recognized and select three groups of images in three - layer scale space, that is, select the sampled images in the first octave (Octave 1), the second octave (Octave 2), and the third octave (Octave 3).
[0077] Specifically, divide the obtained sampled image into blocks. After the block - dividing is completed, all the sampled images include 2 2b (2 b ·2 b ) sub - blocks. In this embodiment, the obtained 2 2b (2 b ·2 b ) sub - blocks are non - overlapping sub - blocks, that is, each sub - block does not overlap with each other, thereby reducing the computational amount in the subsequent process.
[0078] S202: Perform CS - LBP feature extraction on the sub - blocks to obtain the CS - LBP features of the sub - blocks, and perform histogram statistics on the CS - LBP features of a number of sub - blocks to obtain image features.
[0079] In the embodiment of the present application, the Center-Symmetric Local Binary Pattern (CS-LBP) is an algorithm for extracting local texture features of a sampled image. When using this algorithm to extract features, the expression of the features on the planar image has good robustness, which can effectively describe the effect of the face image and provide reliable guarantee for subsequent classification operations.
[0080] Among them, the specific steps for extracting CS-LBP features from a sub-block include: obtaining the grayscale value of each pixel point in the sub-block, sequentially selecting pixel points as the central pixel point in the order of pixel point numbers, and taking the pixel points adjacent to the central pixel point as the surrounding pixel points. Compare the grayscale value of the central pixel point with the grayscale values of its adjacent surrounding pixel points in turn. If the comparison result is greater than the comparison threshold, record it as 1, otherwise as 0. Arrange the obtained 0 / 1 through a preset standard arrangement order to obtain a binary number, and use this binary number as the CS-LBP coding value of the sub-block pixel point. The CS-LBP coding values of each pixel point in the sub-block constitute the CS-LBP features of the sub-block.
[0081] The specific calculation process of the CS-LBP coding value is as follows:
[0082]
[0083] Among them, CS-LBP p,R,ε is the CS-LBP coding value, p is the number of pixel points in a sub-block, i is the number of the pixel point, and the number of the pixel point reflects the position of the pixel point in the sub-block. g i represents the pixel grayscale value at this position, S(X) is the comparison result of the pixel point grayscale value, and the expression of S(X) is:
[0084]
[0085] Among them, ε is the comparison threshold, ε is a constant, and its value is small.
[0086] After obtaining the CS-LBP features of the sub-blocks, perform histogram statistics on the CS-LBP features of each sub-block, thereby obtaining sub-blocks described by histograms. Each sub-block together forms an image, and the obtained image composed of several histograms is the image feature.
[0087] S203: Perform weighted PCA dimensionality reduction processing on the image feature to obtain the first feature map.
[0088] Principal components analysis (PCA) is a method for converting data from a high-dimensional space to a low-dimensional space to facilitate data analysis.
[0089] Specifically, convert the image features into a covariance matrix, calculate the eigenvalues and eigenvectors of the covariance matrix, sort the eigenvalues in descending order, and obtain a weighted mapping transformation matrix based on a preset dimensionality reduction, a preset weighted matrix, and the eigenvalues and eigenvectors of the covariance matrix. Then, obtain a first feature map with a principal component projection matrix according to the weighted mapping transformation matrix and the eigenvectors of the covariance matrix.
[0090] For example, when the data is reduced from n dimensions to k dimensions, the specific processing steps are as follows: convert the image features into a covariance matrix, calculate the eigenvalues and eigenvectors of the covariance matrix, sort several eigenvalues in descending order, select the eigenvectors corresponding to the first k eigenvalues in the sorting order, introduce a weighted matrix to form a weighted mapping transformation matrix W, and calculate the coordinates y of the original feature x in the mapping space, where y = W T x, and obtain a principal component projection matrix. Consider the image corresponding to the principal component projection matrix as the first feature map.
[0091] In this embodiment, during the block processing, when the block level b is set relatively large, that is, the number of sub-blocks is too large, the feature dimension of the face image will be relatively high. Moreover, when some objective conditions change, the face image is likely to cause a relatively large difference among samples of the same class. This difference has a greater impact on the principal components corresponding to larger eigenvalues, that is, the principal components corresponding to larger eigenvalues are more susceptible to environmental factors (such as illumination, pose changes, etc.). To reduce the feature dimension and reduce the interference caused by environmental factors, the present application performs weighted PCA dimensionality reduction processing on the image features.
[0092] S30: Through the feature enhancement network in the trained occluded face recognition model, enhance the features of the non-occluded regions in the first feature map to obtain a second feature map.
[0093] The feature enhancement network includes a multi-attention module. The multi-attention module includes a channel attention module, a spatial attention module, and a global attention module, which are used to enhance the feature weights of the non-occluded regions.
[0094] Refer to Figure 4 , Figure 4It is a schematic diagram for understanding the channel attention module in the embodiments of the present application. The channel attention module (Channel Attention Module) performs max pooling (Max Pool) and average pooling (Avg Pool) on the input first feature map (input feature F) in the spatial dimension respectively to obtain the max pooling result and the average pooling result in the spatial dimension. The max pooling result and the average pooling result in the spatial dimension are input into a multi-layer perceptron (Multi-Layer Perception, abbreviated as MLP) to obtain the MPL output result. The two MPL output results are added together and activated using the Sigmoid function to obtain the output result of the channel attention module (Channel Attention Mc). The channel attention mechanism focuses on the channels of the feature map and realizes attention by assigning different weights to different channels, that is, it can determine which features in the feature map need attention through the channel attention mechanism.
[0095] Refer to Figure 5 , Figure 5 It is a schematic diagram for understanding the spatial attention module in the embodiments of the present application. The spatial attention module (Spatial Attention Module) is connected in series with the channel attention module. The spatial attention module performs max pooling and average pooling on the feature map (Channel-refined feature F) output by the channel attention module in the channel dimension to obtain the max pooling result and the average pooling result in the channel dimension. The max pooling result and the average pooling result in the spatial dimension are concatenated according to the channels. The concatenated result is subjected to a convolution operation (conv layer) and activated using the Sigmoid function to obtain the output result of the channel attention module (Spatial Attention Ms). The channel attention mechanism focuses on the spatial positions of the feature map and can determine where the features in the feature map need attention through the spatial attention mechanism.
[0096] Refer to Figure 6 , Figure 6 It is a schematic diagram for understanding the global attention module in the embodiments of the present application. The global attention module is connected in parallel with the cascaded spatial attention module and channel attention module. The global attention module consists of multiple convolutional layers and activation layers. The specific network architecture of the global attention can refer to Figure 5 , where C represents the number of channels of the feature map, H represents the height of the feature map, W represents the width of the feature map, conv 1×1 is a convolutional layer for performing 1*1 convolution operation, and ReLU is an activation layer for performing linear correction. The global attention module extracts global image information and weights the features at all positions of the image.
[0097] Refer to Figure 7 ,Figure 7 It is a schematic diagram for understanding the feature enhancement network in an embodiment of the present application. The feature enhancement network further includes a feature vector extraction module. The network architecture of the feature vector extraction module is a MobileNetV2 network. The feature vector extraction module includes a first depthwise separable convolutional layer, a bottleneck layer, a convolutional layer, and a second depthwise separable convolutional layer.
[0098] The steps of obtaining the second feature map are specifically as follows: Input the output result of the multi-attention module into the first depthwise separable convolutional layer through the feature vector extraction module to obtain the output result of the first depthwise separable convolutional layer; Input the output result of the first depthwise separable convolutional layer into the 10-layer bottleneck layer, and use the 10-layer bottleneck layer to extract features, and raise the channel dimension of the feature map to 128 to obtain the output result of the bottleneck layer; Input the output result of the bottleneck layer into the convolutional layer, and integrate features through a 1×1 convolutional layer to increase the dimension to 512 to obtain the output result of the convolutional layer; Input the output result of the convolutional layer into the second depthwise separable convolutional layer, and use the second depthwise separable convolutional layer to retain more information to obtain the second feature map.
[0099] Specifically, after obtaining the first feature map, input the first feature map into the multi-attention module in the feature enhancement network to enhance the feature weights of the non-occluded regions, input the output result of the multi-attention module into the feature vector extraction module, and extract the face feature vector of the face image through the feature vector extraction module to obtain the second feature map.
[0100] S40: Recognize the second feature map through the face recognition network in the trained occluded face recognition model to obtain the face recognition result of the face image to be recognized.
[0101] Specifically, after obtaining the second feature map, input the second feature map into the face recognition network in the trained occluded face recognition model. In the face recognition network, calculate the feature distance between the face feature vector of the second feature map and the feature vector of the face image in the preset feature database, screen out the smallest feature distance, compare the obtained smallest feature distance with the preset feature distance threshold. If the smallest feature distance is less than the preset feature distance threshold, then use the identity of the face image corresponding to the smallest feature distance as the face recognition result of the face image to be recognized.
[0102] The method for obtaining the trained occluded face recognition model is as follows: construct an initial occluded face recognition model, which includes an initial feature extraction network, an initial feature enhancement network, and an initial face recognition network; obtain a labeled sample set, which contains a large number of sample images with confirmed identities, input the sample images into the initial occluded face recognition model, preprocess the sample images through the initial feature extraction network, calculate the feature vectors of the sample images through the initial feature enhancement network, obtain the recognition result through the initial face recognition network, calculate the probability of correct recognition according to the recognition result and the labeled identity information, calculate the loss function value according to the probability and the formula of the preset loss function, and readjust the parameters of the initial occluded face recognition model according to the loss function value until the loss function converges, so as to obtain the trained occluded face recognition model. In the embodiments of the present application, the loss function used for training is the ArcFace loss function.
[0103] In summary, referring to Figure 8 , Figure 8 is a schematic diagram for understanding the trained occluded face recognition model in the embodiments of the present application. The trained occluded face recognition model includes a feature extraction network, a feature enhancement network, and a face recognition network. The feature extraction network divides the face image to be recognized into blocks to obtain a number of sub-blocks; performs CS-LBP feature extraction on the sub-blocks to obtain the CS-LBP features of the sub-blocks, performs histogram statistics on the CS-LBP features of the number of sub-blocks to obtain image features; performs weighted PCA dimensionality reduction processing on the image features to obtain a first feature map. The feature enhancement network enhances the feature weights of the non-occluded regions of the first feature map; calculates the feature vectors of the first feature map to obtain a second feature map. The face recognition network obtains the face recognition result according to the second feature map and the face images in the preset feature database.
[0104] The implementation principle of the face recognition method is as follows: obtain a face image to be recognized, which includes an occluded region and a non-occluded region, perform feature extraction on the face image to be recognized through the feature extraction network in the trained occluded face recognition model to obtain a first feature map, enhance the features of the non-occluded regions in the first feature map through the feature enhancement network in the trained occluded face recognition model to obtain a second feature map, and perform recognition on the second feature map through the face recognition network in the trained occluded face recognition model to obtain the face recognition result of the face image to be recognized. In the present application, the feature extraction network can enhance the robustness and classification ability of face features at low resolutions, and the feature enhancement network weights the features of the non-occluded regions, so that the accuracy of the face recognition result is higher.
[0105] The embodiments of the present application also disclose a face recognition system. Referring to Figure 9, the face recognition system includes: an image acquisition module 1, a feature extraction module 2, a feature enhancement module 3, and a face recognition module 4.
[0106] The image acquisition module 1 is used to acquire a face image to be recognized, and the face image to be recognized includes an occluded area and a non-occluded area;
[0107] The feature extraction module 2 is used to extract features from the face image to be recognized through the feature extraction network in the trained occluded face recognition model to obtain a first feature map;
[0108] The feature enhancement module 3 is used to enhance the features of the non-occluded area in the first feature map through the feature enhancement network in the trained occluded face recognition model to obtain a second feature map;
[0109] The face recognition module 4 is used to recognize the second feature map through the face recognition network in the trained occluded face recognition model to obtain the face recognition result of the face image to be recognized.
[0110] It should be noted that when the system provided in the above embodiment realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the face recognition system and the face recognition method embodiment provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0111] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0112] In the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".
[0113] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0114] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes various media that can store program codes such as ROM or random access memory RAM, magnetic disks, or optical discs.
[0115] The above are all the preferred embodiments of the present application. The protection scope of the present application is not limited thereby. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A face recognition method, characterized in that, it includes: Obtain a face image to be recognized, where the face image to be recognized includes an occluded area and a non-occluded area; Through the feature extraction network in the trained occluded face recognition model, perform feature extraction on the face image to be recognized to obtain a first feature map; Through the feature enhancement network in the trained occluded face recognition model, enhance the features of the non-occluded area in the first feature map to obtain a second feature map; Through the face recognition network in the trained occluded face recognition model, perform recognition on the second feature map to obtain the face recognition result of the face image to be recognized; The step of enhancing the features of the non-occluded area in the first feature map through the feature enhancement network in the trained occluded face recognition model to obtain a second feature map specifically includes: inputting the first feature map into the multi-attention module of the feature enhancement network to obtain the output result of the attention module, and the multi-attention module includes a channel attention module, a spatial attention module, and a global attention module; Input the output result of the attention module into the feature vector extraction module of the feature enhancement network to obtain the second feature map; The step of inputting the output result of the attention module into the feature vector extraction module of the feature enhancement network to obtain the second feature map specifically includes: inputting the output result of the attention module into the first depthwise separable convolutional layer to obtain the output result of the first depthwise separable convolutional layer; inputting the output result of the first depthwise separable convolutional layer into the bottleneck layer to obtain the output result of the bottleneck layer; inputting the output result of the bottleneck layer into the convolutional layer to obtain the output result of the convolutional layer; inputting the output result of the convolutional layer into the second depthwise separable convolutional layer to obtain the second feature map.
2. The face recognition method according to claim 1, characterized in that, The step of performing feature extraction on the face image to be recognized through the feature extraction network in the trained occluded face recognition model to obtain a first feature map specifically includes: performing block processing on the face image to be recognized to obtain a number of sub-blocks; Perform CS-LBP feature extraction on the sub-blocks to obtain the CS-LBP features of the sub-blocks; Perform histogram statistics on the CS-LBP features of the number of sub-blocks to obtain image features; Perform weighted PCA dimensionality reduction processing on the image features to obtain the first feature map.
3. The face recognition method according to claim 2, characterized in that, The step of performing block processing on the face image to be recognized to obtain a number of sub-blocks specifically includes: Perform downsampling on the face image to be recognized through a Gaussian pyramid in a three-layer scale space to obtain a sampled image; Divide the sampled image into 2 2b (2 b ·2 b ) blocks to obtain 2 2b (2 b ·2 b ) sub-blocks, where b is the block division level.
4. The face recognition method according to claim 2, characterized in that, The step of performing CS-LBP feature extraction on the sub-blocks specifically includes: Obtain the gray value of the central pixel point of the sub-block and the gray values of the surrounding pixel points, compare the gray value of the central pixel point with the gray values of the surrounding pixel points, and obtain the CS-LBP coding value of the sub-block; the CS-LBP coding value of the sub-block is: Among them, CS-LBP p,R,ε is the CS-LBP coding value of the sub-block, p is the number of pixel sampling points, and g i is the pixel gray value, S(X) is the comparison result of the pixel point gray value, and the expression of the S(X) is: The ε is a comparison threshold value.
5. The face recognition method according to claim 2, characterized in that the step of performing weighted PCA dimensionality reduction processing on the image features specifically includes: Convert the image features into a covariance matrix, and calculate the eigenvalues and eigenvectors of the covariance matrix; Obtain a weighted mapping transformation matrix based on a preset reduced dimension, a preset weighted matrix, and the eigenvalues and eigenvectors of the covariance matrix; Obtain a first feature map with a principal component projection matrix according to the weighted mapping transformation matrix and the eigenvectors of the covariance matrix.
6. The face recognition method according to claim 1, before obtaining the face image to be recognized, further includes: Train an initial occluded face recognition model based on the labeled sample images to obtain a trained occluded face recognition model, and the loss function for training the initial occluded face recognition model is the ArcFace loss function.
7. A system based on the face recognition method according to any one of claims 1-6, characterized in that the system includes: An image acquisition module (1) for acquiring a face image to be recognized, where the face image to be recognized includes an occluded area and a non-occluded area; A feature extraction module (2) for extracting features from the face image to be recognized through the feature extraction network in the trained occluded face recognition model to obtain a first feature map; A feature enhancement module (3) for enhancing the features of the non-occluded area in the first feature map through the feature enhancement network in the trained occluded face recognition model to obtain a second feature map; The face recognition module (4) is configured to recognize the face recognition result of the face image to be recognized by using the face recognition network in the trained occluded face recognition model; the step of enhancing the features of the non-occluded region in the first feature map to obtain the second feature map by using the feature enhancement network in the trained occluded face recognition model specifically includes: inputting the first feature map into the multi-attention module of the feature enhancement network to obtain the output result of the attention module, where the multi-attention module includes a channel attention module, a spatial attention module, and a global attention module; inputting the output result of the attention module into the feature vector extraction module of the feature enhancement network to obtain the second feature map; the step of inputting the output result of the attention module into the feature vector extraction module of the feature enhancement network to obtain the second feature map specifically includes: inputting the output result of the attention module into the first depthwise separable convolutional layer to obtain the output result of the first depthwise separable convolutional layer; inputting the output result of the first depthwise separable convolutional layer into the bottleneck layer to obtain the output result of the bottleneck layer; inputting the output result of the bottleneck layer into the convolutional layer to obtain the output result of the convolutional layer; and inputting the output result of the convolutional layer into the second depthwise separable convolutional layer to obtain the second feature map.
Citation Information
Patent Citations
Face expression identification method based on partially shielded image
CN105825183A
Uncertain facial expression recognition method based on multi-attention fusion Transform architecture
CN113963422A