Face recognition method and system suitable for outdoor electronic instrument
By combining the improved MobileViT and ArcFace methods, the robustness and accuracy issues of face recognition in complex outdoor environments have been resolved. This enables efficient recognition under conditions of strong light, backlight, and large-angle deflection, making it suitable for outdoor electronic instruments.
Patent Information
- Application Number
- CN202511117571.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Existing facial recognition methods are unstable in complex outdoor environments, especially under conditions of strong light, backlight, rain, fog, posture changes, and limited device computing power. They have low recognition accuracy and high false positive rate, and cannot meet the practicality and stability requirements of outdoor environments.
A face detection model employing an improved MobileViT structure, combined with illumination adaptive modulation and pose guidance mechanisms, achieves face region detection, alignment, and recognition through a dual-channel spatial attention mechanism and an improved ArcFace loss function. Furthermore, it enhances robustness and accuracy by incorporating pose enhancement and feature reconstruction.
It significantly improves the robustness and recognition accuracy of face detection under strong light, backlight, and large-angle deflection conditions, reduces the false recognition rate, meets the recognition needs of outdoor environments, and has a lightweight network structure suitable for embedded devices, improving processing efficiency and the reliability of recognition results.
Smart Images

Figure CN120976991A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of facial recognition technology, and in particular to a facial recognition method and system suitable for outdoor electronic instruments. Background Technology
[0002] Against the backdrop of the rapid development of facial recognition technology, deep learning-based algorithms have been widely applied in scenarios such as security monitoring, mobile terminals, and identity verification. However, most existing mainstream facial recognition methods rely on stable and controlled lighting and shooting environments, and are mainly geared towards recognition tasks in indoor scenes or under conditions where image quality is controllable. When applied to outdoor electronic instruments, such as drones, outdoor monitoring, and intelligent inspection terminals, traditional methods are significantly insufficient in terms of facial detection accuracy, alignment accuracy, and recognition robustness due to multiple factors such as strong light, backlight, rain and fog, posture deflection, and device computing power.
[0003] Current methods attempt to introduce lightweight networks and attention mechanisms, but under uneven lighting and multi-pose interference, it is still difficult to guarantee recognition stability. Especially with low-quality image input, the recognition error is large and the false judgment rate is high, which cannot meet the dual requirements of practicality and stability in outdoor environments. Therefore, there is an urgent need for a face recognition method for complex outdoor environments that combines multi-source image processing, adaptive lighting compensation, highly robust feature extraction and recognition optimization mechanisms to achieve a face recognition system with stronger adaptability to environmental changes. Summary of the Invention
[0004] One objective of this invention is to propose a face recognition method and system suitable for outdoor electronic instruments. This invention can provide an efficient and scientific optimization solution for outdoor face recognition, bringing significant technical value and economic benefits to practical applications.
[0005] A face recognition method for outdoor electronic instruments according to an embodiment of the present invention includes the following steps: S1. Collect image data in an outdoor environment, including RGB images and infrared images, and preprocess the original input image to obtain an illumination feature vector; S2. Using a face detection model based on the improved MobileViT structure, face region detection and face key point localization are performed on the illumination-enhanced image, and a face detection result set containing face bounding boxes and coordinates of facial key points is output. S3. Perform affine transformation on the face detection result set to align faces and obtain standardized aligned face images; S4. Input the standardized aligned face image into a face recognition model based on the ResNet100 structure. The face recognition model embeds a dual-channel spatial attention mechanism to extract attention maps. The attention maps are fused to obtain weighting coefficients, and the face feature vector is output. S5. Perform pose-guided feature reconstruction processing on the face feature vector, extract head angle information through the pose estimation module, and use it to reweight and fuse the face feature vector to obtain a pose-enhanced face feature vector. S6. The improved ArcFace loss function is used to normalize the features and classify the identity of the pose-enhanced face feature vector, and the face recognition result is output. S7. Calculate the final recognition score based on face detection confidence and feature matching score, and output the final identity recognition result and confidence score through a fusion scoring mechanism.
[0006] Optionally, S1 includes the following steps: S11. Simultaneously acquire RGB and infrared images through the image acquisition module configured on the outdoor electronic instrument; S12. Perform histogram equalization on each channel of the RGB image and infrared image respectively, and merge the processed three-channel image and single-channel image into an enhanced RGB image; S13. Based on the enhanced RGB image, perform illumination estimation to extract illumination feature vectors, and process the illumination feature vectors... Normalization is performed.
[0007] Optionally, S2 includes the following steps: S21, Face detection model uses illumination feature vectors As input, the feature map is obtained after passing through the initial convolutional downsampling layer. ; S22, Feature map Input up to a multi-stage improved MobileViT block, which contains the following modules: Lightweight depthwise separable convolutional modules for extracting local spatial features , ; Local residual enhancement module, based on local spatial features Perform cross-layer feature overlay and channel attention modulation to output enhanced features. The specific calculation process is as follows: First, local spatial features Global average pooling is performed, and channel weights are generated through two fully connected layers. ,Right now: ; in, This is the weight matrix. Image height, Image width, It is the ReLU activation function. For the Sigmoid function; Then channel attention is applied to the local spatial features. Weighted and fused features are then obtained. ,Right now: ; By employing a local residual enhancement module, the robustness of feature representation under occlusion or complex background conditions is enhanced. The explicit facial region guidance module enhances features based on predefined or pre-trained facial key region heatmaps. Perform pixel-by-pixel weighting to output region-guided feature maps. ; Region-guided feature map Divide into local blocks, and flatten each local block into a vector. The input is a low-rank attention Transformer module for attention modeling, and the output is an enhanced representation. ; All enhanced representations Reconstructed into enhanced feature maps ; An illumination modulation function is used to generate a feature map after illumination modulation based on the illumination intensity matrix. ,Right now: ; in, This is the light modulation intensity control coefficient. This is the illumination intensity matrix, representing the pixel positions of the input image. Local light intensity at the location, Represents the light intensity matrix The average pixel value of the entire image. Represents the light intensity matrix The standard deviation of the face detection model, compared to the original MobileViT's lack of awareness of lighting changes, is significantly improved by adopting a lighting modulation function and embedding a Transformer. This allows the face detection model to automatically adjust its feature representation in extreme environments such as strong light and backlight, thus significantly improving the robustness and adaptability of actual outdoor detection. S23. Feature map after illumination modulation Multi-scale regression and classification are performed to obtain a set of face bounding boxes. , It outputs the corresponding set of facial key points based on an independent keypoint regression head. ; S24. The candidate detection results are simplified by using confidence threshold filtering, and the face detection result set is output. .
[0008] Optionally, S3 includes the following steps: S31, Based on the set of face detection results using key point coordinates The center points of the two eyes are used as the reference points for the affine transformation, and a standard reference coordinate set is defined. , used to construct standard geometric models of faces; S32. Calculate the affine transformation matrix. The affine transformation matrix satisfy: ; in, Let j be the coordinates of the actual key point. The coordinates of the standard key point corresponding to the coordinates of the j-th actual key point; Based on affine matrix Perform geometric transformations on the face regions in the original image to generate an aligned face image. ; S33. Align the generated face image. The images are uniformly scaled to the preset input size to obtain standardized and aligned images. .
[0009] Optionally, S4 includes the following steps: S41. Standardize and align the image. The input is fed into a face recognition model built on the ResNet100 architecture, which includes a backbone network, a dual-channel spatial attention module, and an output embedding head module. S42. The backbone network of the face recognition model performs multi-scale hierarchical feature extraction on the input standardized and aligned image through multiple residual blocks to obtain a deep feature tensor. ,in, , The spatial dimensions of the feature map. Number of channels; S43, convert the depth feature tensor The input is fed into the dual-channel spatial attention module, where two attention mechanisms are used for computation: The first channel is for extracting salient regions based on gradient response and calculating an edge-guided attention map. ; The second channel generates a structure-sensitive attention map based on regional weighting of facial region heatmaps. ; S44. Merge the edge-guided attention map and the structure-sensitive attention map into the final spatial attention map. : ; S45, on depth feature tensors Perform a weighted operation to obtain the enhanced feature tensor. : ; S46, Enhance the feature tensor The input is fed into a global average pooling layer and an embedding mapping layer to obtain a face feature vector. ,in, For feature dimensions.
[0010] Optionally, S5 includes the following steps: S51, convert the facial feature vector and standardized aligned images Input to the attitude estimation module; In the pose estimation module, a multi-branch regression network is used to convolve the image and output a three-dimensional face pose angle vector. S52. Normalize the face pose angle vector to obtain the normalized pose vector. ; S53. Normalize the attitude vector The input is fed into the pose-guided feature enhancement module to generate a weight vector. ; S54, Facial Feature Vectors With weight vector Element-level fusion is performed to obtain pose-enhanced facial feature vectors. .
[0011] Optionally, S6 includes the following steps: S61. Facial feature vectors with enhanced pose Input is fed into the recognition and classification module with an improved ArcFace loss function; S62, Represent the set of facial feature vectors of all registered users as follows: ,in, For the first The unit-normalized class center vectors of each identity satisfy: ; S63. Calculate the input facial feature vector. Cosine similarity with all category centers: ; in, Indicates the input features and the first The angle between the class center vectors; S64, Angle corresponding to the real category label Add angle interval This yields an improved ArcFace discriminant boundary. ,in, The set angular interval constant; S65. Scale all similarities to obtain a normalized classification score. : ; in, This is the scaling factor. Index for real categories; S66. Normalize the classification scores for all categories. Input the Softmax function to calculate the predicted probability; S67. Based on the maximum probability classification criterion, output the predicted identity category. : ; in, This indicates that when the face feature vector is The identity category is number The probability of each identity.
[0012] Optionally, S7 includes the following steps: S71. Use face detection confidence as the first scoring factor and feature matching score as the second scoring factor. S72. Construct the final fusion scoring function : ; in, , representing the weight coefficient of each factor, Confidence for face detection Feature matching score; S73. Set the recognition confidence threshold. ,when At that time, confirm that the output recognition result is the predicted identity category. Otherwise, output the "Unrecognized" flag.
[0013] According to an embodiment of the present invention, a face recognition system suitable for outdoor electronic instruments includes the following modules: The image acquisition module is used to simultaneously acquire RGB and infrared images, perform image channel separation, time alignment and size normalization, and generate standard image data pairs. The face detection module is used to input enhanced images into the face detection model of the improved MobileViT structure, and generates bounding boxes and key point sets through local attention enhancement and illumination modulation mechanisms; The face alignment module is used to calculate the affine transformation matrix based on facial key points and perform geometric alignment and standardization processing on the face region in the image; The feature extraction module is used to input aligned face images into the ResNet100 network and combine a dual-channel spatial attention mechanism to extract normalized face feature vectors. The pose enhancement module is used to estimate the 3D pose angles of the face, construct the pose guidance weight vector, and reweight and fuse the face feature vectors to enhance multi-pose robustness. The identity classification module is used to perform normalized matching and identity prediction on pose enhancement features based on the improved ArcFace loss function, and output classification scores and identity labels. The scoring fusion module is used to integrate face detection confidence, matching similarity and illumination adaptability scores to calculate the final recognition score and determine the recognition validity based on the threshold. The results output module is used to output the final identified identity and confidence score.
[0014] The beneficial effects of this invention are: 1. This invention addresses the shortcomings of existing technologies in recognizing faces under complex environments by introducing an improved detection network with light-adaptive MobileViT and an ArcFace recognition model with a fusion of posture guidance mechanism. This enables accurate detection and robust recognition of faces under conditions of strong light, backlight, low illumination, and large-angle deflection.
[0015] 2. This invention utilizes image illumination estimation and feature reweighting to effectively improve the positioning accuracy of the detection module under non-ideal lighting conditions. By using a pose-guided feature fusion structure to align and enhance facial features at different angles, it significantly improves recognition accuracy and consistency.
[0016] 3. This invention adopts a lightweight network structure design, taking into account the resource constraints and response speed requirements of embedded devices. While ensuring recognition accuracy, it improves processing efficiency and is suitable for practical deployment in outdoor terminal devices. Through a confidence scoring fusion mechanism, it enhances the credibility assessment capability of recognition results and further reduces the false recognition rate and false negative rate. The overall system is superior to existing methods in terms of environmental adaptability, recognition robustness, and computational efficiency, and has good engineering feasibility and promotion value. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the face recognition method in this invention; Figure 2 This is a schematic diagram of the improved MobileViT structure in this invention; Figure 3 This is a schematic diagram of the face recognition system in this invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] Example 1 Current outdoor face recognition technologies generally face problems such as unstable detection, low recognition accuracy, and poor adaptability to strong light, backlight, and occlusion in complex environments. Especially when deployed on mobile terminals or low-power devices, traditional models struggle to balance accuracy and efficiency, severely limiting recognition performance. Therefore, this embodiment proposes a face recognition method suitable for outdoor electronic instruments. Combining an improved MobileViT detection network with light adaptation and a pose-guided face recognition algorithm, it exhibits higher robustness and recognition accuracy in scenarios involving strong light interference, multiple pose shifts, and edge occlusion. Compared to existing recognition systems that rely on single feature extraction or fixed lighting models, this invention not only improves adaptability to various environments but also boasts advantages such as lightweight structure and edge deployment, making it suitable for various practical outdoor scenarios such as security patrols, smart terminals, and emergency response. Figure 1 and Figure 2 As shown, the method includes the following steps: S1. Acquire image data from an outdoor environment, including RGB and infrared images, and preprocess the original input image to obtain an illumination feature vector. This includes the following steps: S11. Simultaneously acquire RGB and infrared images through the image acquisition module configured on the outdoor electronic instrument; S12. Perform histogram equalization on each channel of the RGB image and infrared image respectively, and merge the processed three-channel image and single-channel image into an enhanced RGB image; S13. Based on the enhanced RGB image, perform illumination estimation to extract illumination feature vectors, and process the illumination feature vectors... Normalization is performed.
[0020] This invention employs an improved MobileViT structure for feature extraction and detection, combined with an adaptive illumination modulation mechanism, enabling the system to stably detect facial regions and accurately extract key facial features even under strong illumination interference or complex background conditions, thus enhancing the detection module's adaptability to complex scenes.
[0021] S2. Using a face detection model based on an improved MobileViT structure, perform face region detection and facial landmark localization on the illuminated image, outputting a face detection result set containing face bounding boxes and coordinates of facial landmarks. This includes the following steps: S21, Face detection model uses illumination feature vectors As input, the feature map is obtained after passing through the initial convolutional downsampling layer. ; S22, Feature map Input up to a multi-stage improved MobileViT block, see reference. Figure 2 The MobileViT block contains the following modules: Lightweight depthwise separable convolutional modules for extracting local spatial features , ; Local residual enhancement module, based on local spatial features Perform cross-layer feature overlay and channel attention modulation to output enhanced features. The specific calculation process is as follows: First, local spatial features Global average pooling is performed, and channel weights are generated through two fully connected layers. ,Right now: ; in, This is the weight matrix. Image height, Image width, It is the ReLU activation function. For the Sigmoid function; Then channel attention is applied to the local spatial features. Weighting and merging: ; By employing a local residual enhancement module, the robustness of feature representation under occlusion or complex background conditions is enhanced. The explicit facial region guidance module enhances features based on predefined or pre-trained facial key region heatmaps. Perform pixel-by-pixel weighting to output region-guided feature maps. ; Region-guided feature map Divide into local blocks, and flatten each local block into a vector. The input is a low-rank attention Transformer module for attention modeling, and the output is an enhanced representation. ; All enhanced representations Reconstructed into enhanced feature maps ; An illumination modulation function is used to generate a feature map after illumination modulation based on the illumination intensity matrix. ,Right now: ; in, This is the light modulation intensity control coefficient. This is the illumination intensity matrix, representing the pixel positions of the input image. Local light intensity at the location, Represents the light intensity matrix The average pixel value of the entire image. Represents the light intensity matrix The standard deviation of the face detection model, compared to the original MobileViT's lack of awareness of lighting changes, is significantly improved by adopting a lighting modulation function and embedding a Transformer. This allows the face detection model to automatically adjust its feature representation in extreme environments such as strong light and backlight, thus significantly improving the robustness and adaptability of actual outdoor detection. S23. Feature map after illumination modulation Multi-scale regression and classification are performed to obtain a set of face bounding boxes. , It outputs the corresponding set of facial key points based on an independent keypoint regression head. ; S24. The candidate detection results are simplified by using confidence threshold filtering, and the face detection result set is output. .
[0022] This invention achieves geometric alignment through affine transformation based on key points, which can effectively eliminate the differences in angle and pose of faces during shooting, improve the accuracy and consistency of subsequent feature extraction and recognition, and lay a standardized input foundation for deep feature learning.
[0023] S3. Perform affine transformation on the face detection result set to align faces and obtain standardized aligned face images. This includes the following steps: S31, Based on the set of face detection results using key point coordinates The center points of the two eyes are used as the reference points for the affine transformation, and a standard reference coordinate set is defined. , used to construct standard geometric models of faces; S32. Calculate the affine transformation matrix. The affine transformation matrix satisfy: ; in, Let j be the coordinates of the actual key point. The coordinates of the standard key point corresponding to the coordinates of the j-th actual key point; Based on affine matrix Perform geometric transformations on the face regions in the original image to generate an aligned face image. ; S33. Align the generated face image. The images are uniformly scaled to the preset input size to obtain standardized and aligned images. .
[0024] This invention employs a dual-channel spatial attention mechanism to weight features in the facial region, effectively enhancing the expressive power of key regions (such as eyes, bridge of the nose, and mouth), and improving feature discriminativeness and robustness under complex backgrounds and noise interference.
[0025] S4. Input the standardized aligned face image into a face recognition model based on the ResNet100 structure. The face recognition model embeds a dual-channel spatial attention mechanism to extract an attention map. Weighting coefficients are obtained by fusing the attention maps, and a face feature vector is output. This includes the following steps: S41. Standardize and align the image. The input is fed into a face recognition model built on the ResNet100 architecture, which includes a backbone network, a dual-channel spatial attention module, and an output embedding head module. S42. The backbone network of the face recognition model performs multi-scale hierarchical feature extraction on the input standardized and aligned image through multiple residual blocks to obtain a deep feature tensor. ,in, , The spatial dimensions of the feature map. Number of channels; S43, convert the depth feature tensor The input is fed into the dual-channel spatial attention module, where two attention mechanisms are used for computation: The first channel is for extracting salient regions based on gradient response and calculating an edge-guided attention map. ; The second channel generates a structure-sensitive attention map based on regional weighting of facial region heatmaps. ; S44. Merge the edge-guided attention map and the structure-sensitive attention map into the final spatial attention map. : ; S45, on depth feature tensors Perform a weighted operation to obtain the enhanced feature tensor. : ; S46, Enhance the feature tensor The input is fed into a global average pooling layer and an embedding mapping layer to obtain a face feature vector. ,in, For feature dimensions.
[0026] This invention estimates head posture and constructs a posture-guided feature fusion method, which can compensate for feature distortion caused by large-angle deflection or side profile, enhance the system's recognition ability under multi-angle and multi-view conditions, and reduce the false recognition rate caused by posture deviation.
[0027] S5. Perform pose-guided feature reconstruction processing on the facial feature vector. Extract head angle information through the pose estimation module and use it to reweight and fuse the facial feature vector to obtain a pose-enhanced facial feature vector. This includes the following steps: S51, convert the facial feature vector and standardized aligned images Input to the attitude estimation module; In the pose estimation module, a multi-branch regression network is used to convolve the image and output a three-dimensional face pose angle vector. S52. Normalize the face pose angle vector to obtain the normalized pose vector. ; S53. Normalize the attitude vector The input is fed into the pose-guided feature enhancement module to generate a weight vector. ; S54, Facial Feature Vectors With weight vector Element-level fusion is performed to obtain pose-enhanced facial feature vectors. .
[0028] This invention utilizes an improved ArcFace loss function to enhance the ability to distinguish inter-class boundaries and achieves high-precision identity classification based on feature normalization, enabling face recognition to maintain high stability and scalability in complex environments and improving overall recognition accuracy.
[0029] S6. The improved ArcFace loss function is used to normalize the features of the pose-enhanced face feature vector and perform identity classification, outputting the face recognition result. This includes the following steps: S61. Facial feature vectors with enhanced pose Input is fed into the recognition and classification module with an improved ArcFace loss function; S62, Represent the set of facial feature vectors of all registered users as follows: ,in, For the first The unit-normalized class center vectors of each identity satisfy: ; S63. Calculate the input facial feature vector. Cosine similarity with all category centers: ; in, Indicates the input features and the first The angle between the class center vectors; S64, Angle corresponding to the real category label Add angle interval This yields an improved ArcFace discriminant boundary. ,in, The set angular interval constant; S65. Scale all similarities to obtain a normalized classification score. : ; in, This is the scaling factor. Index for real categories; S66. Normalize the classification scores for all categories. Input the Softmax function to calculate the predicted probability; S67. Based on the maximum probability classification criterion, output the predicted identity category. : ; in, This indicates that when the face feature vector is The identity category is number The probability of each identity.
[0030] This invention constructs a multi-factor fusion scoring model by comprehensively considering detection confidence and matching similarity, thereby achieving dynamic evaluation and reasonable output of the credibility of the recognition results, significantly reducing the false judgment rate, and improving the reliability and security of the system in practical applications.
[0031] S7. Calculate the final recognition score based on face detection confidence and feature matching score, and output the final identity recognition result and confidence score through a fusion scoring mechanism. This includes the following steps: S71. Use face detection confidence as the first scoring factor and feature matching score as the second scoring factor. S72. Construct the final fusion scoring function : ; in, , representing the weight coefficient of each factor, Confidence for face detection Feature matching score; S73. Set the recognition confidence threshold. ,when At that time, confirm that the output recognition result is the predicted identity category. Otherwise, output the "Unrecognized" flag.
[0032] Example 2 Based on the face recognition method proposed in Example 1, this example provides a face recognition system suitable for outdoor electronic instruments, which is used to implement the face recognition method. It consists of an image acquisition module, a face detection module, a face alignment module, a feature extraction module, a pose enhancement module, an identity classification module, a scoring fusion module, and a result output module.
[0033] The image acquisition module is used to simultaneously acquire RGB and infrared images, perform image channel separation, time alignment, and size normalization, and generate standard image data pairs; the face detection module is used to input the enhanced image into the improved MobileViT structure face detection model, and generate bounding boxes and key point sets through local attention enhancement and illumination modulation mechanisms; the face alignment module is used to calculate the affine transformation matrix based on the face key points, and perform geometric alignment and normalization processing on the face regions in the image.
[0034] The feature extraction module inputs aligned face images into a ResNet100 network and combines a dual-channel spatial attention mechanism to extract normalized face feature vectors. The pose enhancement module estimates the 3D pose angles of the face, constructs a pose-guided weight vector, and reweights and fuses the face feature vectors to enhance multi-pose robustness. The identity classification module performs normalized matching and identity prediction on the pose-enhanced features based on an improved ArcFace loss function, outputting a classification score and identity label. The scoring fusion module integrates face detection confidence, matching similarity, and illumination adaptability scores to calculate the final recognition score and determines the recognition validity based on a threshold. The result output module outputs the final recognized identity and confidence score.
[0035] The face recognition system proposed in this embodiment adopts a lightweight network structure design, which takes into account the resource limitations and response speed requirements of embedded devices. While ensuring recognition accuracy, it improves processing efficiency and is suitable for actual deployment in outdoor terminal devices.
[0036] Example 3 In this embodiment, in the construction project of an intelligent patrol system in a coastal city, the R&D team deployed a batch of outdoor electronic instrument terminals with facial recognition function to quickly identify target personnel in the moving crowd in outdoor mobile scenarios. These outdoor electronic instruments are installed on the top of electric patrol vehicles or in related equipment, and have requirements such as low power consumption, high real-time performance, and all-weather operation. Typical scenarios include direct sunlight, dense crowds, night duty, and severe weather (such as fog, haze, and high humidity).
[0037] Traditional facial recognition methods suffer from significant problems in such application scenarios. For example, under strong direct sunlight at midday, overexposure leads to the loss of information in key facial areas, significantly reducing recognition accuracy. At night or in backlit environments, insufficient image brightness or light interference prevents the detection module from accurately identifying facial regions. Furthermore, during dynamic, moving photography, frequent changes in facial angles, such as side profiles and large tilt angles, further exacerbate the recognition difficulties. Initial testing showed that the existing system's accuracy dropped below 72% and the false alarm rate rose to 18% when the facial deflection angle exceeded 25°.
[0038] To address the aforementioned issues, the project team introduced the method proposed in this invention, integrating the improved MobileViT detection model and the ArcFace+PGFE recognition module into the terminal system. The system first utilizes dual cameras to acquire RGB and infrared images in parallel, and performs image preprocessing through time synchronization and illumination estimation to ensure consistent image quality under different lighting conditions. Subsequently, it achieves accurate face localization through an illumination-adaptive face detection model, and performs standardized alignment of the face based on key point coordinates. The extracted features are enhanced by a dual-channel spatial attention mechanism before being input into the pose guidance module for multi-view reconstruction, effectively enhancing the recognition capability for faces with large angle deviations.
[0039] In the specific test, the team selected a certain area of a city and conducted a 7-day on-site deployment test. The pre-arranged target personnel mingled in the crowd and collected a total of about 86,000 sample images, covering all day and night and various weather conditions. The system identified and compared the images by connecting to a pre-established database and compared and evaluated them with traditional CNN-based systems.
[0040] The results show that in bright daylight, the recognition accuracy of this system improved from 81.3% to 94.6% compared to the comparison system, and from 68.7% to 89.2% in low-light conditions at night. Especially when the face deflection angle is greater than 30°, the recognition accuracy remains above 87.5%, outperforming the traditional system by more than 15 percentage points. Meanwhile, the average recognition time is controlled within 176ms, meeting the real-time response requirements of mobile terminals.
[0041] By integrating the scoring mechanism, the system also effectively reduced the false recognition rate and repeated alarms. In tests in high-density areas (such as subway entrances and shopping mall entrances), the false recognition rate dropped from 9.3% in the original system to 3.2%, achieving accurate, safe, and efficient outdoor facial recognition.
[0042] Table 1. Recognition performance data under different lighting and pose conditions. The above embodiments and test data demonstrate that the face recognition method of the present invention has good outdoor adaptability, recognition accuracy and response speed. It can effectively solve the problem of unstable recognition performance of existing technologies in complex scenarios such as strong light, low light and posture deviation, and has significant practicality and promotion value.
[0043] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A face recognition method suitable for outdoor electronic instruments, characterized in that, Includes the following steps: S1. Collect image data in an outdoor environment, including RGB images and infrared images, and preprocess the original input image to obtain an illumination feature vector; S2. Using a face detection model based on the improved MobileViT structure, face region detection and face key point localization are performed on the illumination-enhanced image, and a face detection result set containing face bounding boxes and coordinates of facial key points is output. S3. Perform affine transformation on the face detection result set to align faces and obtain standardized aligned face images; S4. Input the standardized aligned face image into a face recognition model based on the ResNet100 structure. The face recognition model embeds a dual-channel spatial attention mechanism to extract attention maps. The attention maps are fused to obtain weighting coefficients, and the face feature vector is output. S5. Perform pose-guided feature reconstruction processing on the face feature vector, extract head angle information through the pose estimation module, and use it to reweight and fuse the face feature vector to obtain a pose-enhanced face feature vector. S6. The improved ArcFace loss function is used to normalize the features and classify the identity of the pose-enhanced face feature vector, and the face recognition result is output. S7. Calculate the final recognition score based on face detection confidence and feature matching score, and output the final identity recognition result and confidence score through a fusion scoring mechanism.
2. The face recognition method according to claim 1, characterized in that, S1 includes the following steps: S11. Simultaneously acquire RGB and infrared images through the image acquisition module configured on the outdoor electronic instrument; S12. Perform histogram equalization on each channel of the RGB image and infrared image respectively, and merge the processed three-channel image and single-channel image into an enhanced RGB image; S13. Based on the enhanced RGB image, perform illumination estimation to extract illumination feature vectors, and process the illumination feature vectors... Normalization is performed.
3. The face recognition method according to claim 1, characterized in that, S2 includes the following steps: S21, The face detection model uses illumination feature vectors As input, the feature map is obtained after passing through the initial convolutional downsampling layer. ; S22, Feature map Input up to a multi-stage improved MobileViT block, which contains the following modules: Lightweight depthwise separable convolutional modules for extracting local spatial features , ; Local residual enhancement module, based on local spatial features Perform cross-layer feature overlay and channel attention modulation to output enhanced features. ; The explicit facial region guidance module enhances features based on predefined or pre-trained facial key region heatmaps. Perform pixel-by-pixel weighting to output region-guided feature maps. ; Region-guided feature map Divide into local blocks, and flatten each local block into a vector. The input is a low-rank attention Transformer module for attention modeling, and the output is an enhanced representation. ; All enhanced representations Reconstructed into enhanced feature maps ; An illumination modulation function is used to generate a feature map after illumination modulation based on the illumination intensity matrix. ,Right now: ; in, This is the light modulation intensity control coefficient. This is the illumination intensity matrix, representing the pixel positions of the input image. Local light intensity at the location, Represents the light intensity matrix The average pixel value of the entire image. Represents the light intensity matrix Standard deviation; S23. Feature map after illumination modulation Multi-scale regression and classification are performed to obtain a set of face bounding boxes. , It outputs the corresponding set of facial key points based on an independent keypoint regression head. ; S24. The candidate detection results are simplified by using confidence threshold filtering, and the face detection result set is output. .
4. The face recognition method according to claim 1, characterized in that, S3 includes the following steps: S31, Based on the set of face detection results using key point coordinates The center points of the two eyes are used as the reference points for the affine transformation, and a standard reference coordinate set is defined. , used to construct standard geometric models of faces; S32. Calculate the affine transformation matrix. Based on affine matrices Perform geometric transformations on the face regions in the original image to generate an aligned face image. ; S33. Align the generated face image. The images are uniformly scaled to the preset input size to obtain standardized and aligned images. .
5. The face recognition method according to claim 1, characterized in that, S4 includes the following steps: S41. Standardize and align the image. The input is fed into a face recognition model built on the ResNet100 architecture, which includes a backbone network, a dual-channel spatial attention module, and an output embedding head module. S42. The backbone network of the face recognition model performs multi-scale hierarchical feature extraction on the input standardized and aligned image through multiple residual blocks to obtain a deep feature tensor. ,in, , The spatial dimensions of the feature map. Number of channels; S43, convert the depth feature tensor The input is fed into the dual-channel spatial attention module, where two attention mechanisms are used for computation: The first channel is for extracting salient regions based on gradient response and calculating an edge-guided attention map. ; The second channel generates a structure-sensitive attention map based on regional weighting of facial region heatmaps. ; S44. Merge the edge-guided attention map and the structure-sensitive attention map into the final spatial attention map. ; S45, on depth feature tensors Perform a weighted operation to obtain the enhanced feature tensor. ; S46, Enhance the feature tensor The input is fed into a global average pooling layer and an embedding mapping layer to obtain a face feature vector. ,in, For feature dimensions.
6. The face recognition method according to claim 1, characterized in that, S5 includes the following steps: S51, convert the facial feature vector and standardized aligned images Input to the attitude estimation module; In the pose estimation module, a multi-branch regression network is used to convolve the image and output a three-dimensional face pose angle vector. S52. Normalize the face pose angle vector to obtain the normalized pose vector. ; S53. Normalize the attitude vector The input is fed into the pose-guided feature enhancement module to generate a weight vector. ; S54, Facial Feature Vectors With weight vector Element-level fusion is performed to obtain pose-enhanced facial feature vectors. .
7. The face recognition method according to claim 1, characterized in that, S6 includes the following steps: S61. Facial feature vectors with enhanced pose Input is fed into the recognition and classification module with an improved ArcFace loss function; S62, Represent the set of facial feature vectors of all registered users as follows: ,in, For the first Unit-normalized category center vector for each identity; S63. Calculate the input facial feature vector. Cosine similarity with all category centers; S64, Angle corresponding to the real category label Add angle interval This yields an improved ArcFace discriminant boundary. ,in, The set angular interval constant; S65. Scale all similarities to obtain a normalized classification score. ; S66. Normalize the classification scores for all categories. Input the Softmax function to calculate the predicted probability; S67. Based on the maximum probability classification criterion, output the predicted identity category. .
8. The face recognition method according to claim 1, characterized in that, S7 includes the following steps: S71. Use face detection confidence as the first scoring factor and feature matching score as the second scoring factor. S72. Construct the final fusion scoring function ; S73. Set the recognition confidence threshold. ,when At that time, confirm that the output recognition result is the predicted identity category. Otherwise, output the "Unrecognized" flag.
9. A face recognition system suitable for outdoor electronic instruments, performing the face recognition method according to any one of claims 1 to 8, characterized in that, Includes the following modules: The image acquisition module is used to simultaneously acquire RGB and infrared images, perform image channel separation, time alignment and size normalization, and generate standard image data pairs. The face detection module is used to input enhanced images into the face detection model of the improved MobileViT structure, and generates bounding boxes and key point sets through local attention enhancement and illumination modulation mechanisms; The face alignment module is used to calculate the affine transformation matrix based on facial key points and perform geometric alignment and standardization processing on the face region in the image; The feature extraction module is used to input aligned face images into the ResNet100 network and combine a dual-channel spatial attention mechanism to extract normalized face feature vectors. The pose enhancement module is used to estimate the 3D pose angles of the face, construct the pose guidance weight vector, and reweight and fuse the face feature vectors to enhance multi-pose robustness. The identity classification module is used to perform normalized matching and identity prediction on pose enhancement features based on the improved ArcFace loss function, and output classification scores and identity labels. The scoring fusion module is used to integrate face detection confidence, matching similarity and illumination adaptability scores to calculate the final recognition score and determine the recognition validity based on the threshold. The results output module is used to output the final identified identity and confidence score.
Citation Information
Patent Citations
Face recognition method of attitude robust
CN101763503A
Multi-pose facial recognition method based on infrared camera
CN110135361A
Animal husbandry image recognition method and device based on deep learning
CN116452792A
High-definition monitoring camera tracking method and system for multi-angle face detection
CN120279064A
Cited By
Image quality enhancement method and system based on face recognition
CN122289006A
Face recognition based image quality enhancement method and system
CN122289006B