Face recognition method and system for dynamic environment
Through multi-spectral imaging equipment and multi-scale spatiotemporal fusion feature extraction network, combined with dynamic noise reduction, non-uniform light compensation and domain adaptive alignment technology, the recognition accuracy and real-time problems of traditional face recognition technology in dynamic environments are solved, and the face recognition effect with high accuracy, robustness and real-time is achieved.
Patent Information
- Application Number
- CN202510289367.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Traditional facial recognition technology faces problems such as lighting changes, posture and expression changes, occlusion and background noise in dynamic environments, making it difficult to meet the actual application requirements of recognition accuracy and real-time.
Multi-spectral imaging equipment is used to collect face video streams in visible light and near-infrared bands, combine dynamic noise reduction, distortion correction and non-uniform light compensation models, and build a multi-scale spatiotemporal fusion feature extraction network, embed a bidirectional optical flow estimation module and channel attention mechanism, generate virtual samples and feature maps through domain adaptive alignment algorithms, design an online incremental feature update mechanism, deploy a heterogeneous graph neural network for multimodal decision fusion, and use a hierarchical verification architecture to complete identity recognition.
It significantly improves the accuracy, robustness and real-time performance of face recognition in dynamic environments, can effectively adapt to the complex environment of lighting, posture, expression changes, occlusion and background noise, and improves the reliability and application prospects of recognition.
Smart Images

Figure CN120220208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face recognition, and specifically to a face recognition method and system for dynamic environments. Background Art
[0002] In today's digital age, face recognition technology, as a key technology in the field of biometrics, is widely used in many scenarios such as security monitoring, access control systems, and intelligent payment. However, traditional face recognition methods face severe challenges in dynamic environments and are difficult to meet the requirements of practical applications.
[0003] The complex and variable lighting conditions are one of the important problems. In outdoor environments, from early morning to evening, the light intensity changes greatly, and in different weather conditions, such as sunny, cloudy, rainy days, etc., the nature of the light is also completely different. In addition, the lighting in indoor environments is not evenly distributed either. In places such as large shopping malls and warehouses, there are obvious bright and dark areas. In low-light areas, the face images obtained by traditional face recognition algorithms are often blurred, and a large amount of detail information is lost, making feature extraction extremely difficult, and thus seriously affecting the recognition accuracy. When directly irradiated by strong light, the image is prone to overexposure, which will also damage the key features of the face and lead to recognition failure.
[0004] The dynamic changes in face poses and expressions also pose great obstacles to recognition. In actual scenarios, people's head poses are constantly changing, such as nodding, shaking the head, tilting the head, etc., which makes the angle and position of the face in the image constantly changing. Moreover, the rapid changes in micro-expressions cannot be ignored. These subtle expression changes may contain important emotional information, but due to their rapid and small amplitude changes, it is very difficult for traditional algorithms to accurately capture and analyze these dynamic features. When the face pose and expression change relatively violently, the recognition performance of traditional face recognition systems will drop significantly and cannot accurately identify the identity.
[0005] Occlusion situations are also common in practical applications. Partial occlusions, such as wearing masks, glasses, hats, etc., will directly cover the key parts of the face, making the recognition algorithms based on complete face features unable to work properly. In some special scenarios, such as medical places and industrial environments, people often need to wear protective equipment, which further increases the difficulty of recognition. At the same time, the interference of background noise will also affect face recognition. Complex background patterns, interlaced light and shadow, etc., may all interfere with the accurate extraction of face features by the algorithm and lead to misrecognition.
[0006] Facing the above problems, although some existing improvement methods have alleviated some difficulties to a certain extent, there are still many deficiencies. Some methods address the lighting problem by using simple brightness adjustment algorithms, which cannot fundamentally solve the problem of non-uniform lighting and have poor effects in complex lighting scenarios. When dealing with pose and expression changes, some algorithms increase the feature dimension, but this brings about a sharp increase in computational complexity, resulting in poor real-time performance of the system and making it difficult to meet the requirements of rapid recognition in practical applications. For occlusion and background noise problems, existing occlusion processing algorithms often rely on specific occlusion models and have poor generality. Once encountering new occlusion situations, they cannot effectively respond. Therefore, it is urgent to develop a method and system for stable, accurate, and efficient face recognition in dynamic environments. This can not only promote the in-depth application of face recognition technology in more fields but also provide more powerful guarantees for social security and convenience. Summary of the Invention
[0007] The purpose of the present invention is to provide a face recognition method and system for dynamic environments to solve the problems raised in the above background technology.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A face recognition method for dynamic environments, the method comprising:
[0009] Step 1: Collect a face video stream containing visible light and near-infrared bands through a multispectral imaging device, perform dynamic noise reduction and distortion correction on the original data, use an adaptive smoothing algorithm based on spatio-temporal domain joint filtering to eliminate motion blur, and enhance the brightness of low-illumination areas through a non-uniform illumination compensation model;
[0010] Step 2: Construct a multi-scale spatio-temporal fusion feature extraction network, embed a bidirectional optical flow estimation module in a three-dimensional convolutional neural network, extract the dynamic spatio-temporal features of pose changes and micro-expressions in the face sequence, and use a channel attention mechanism to dynamically weight and fuse cross-modal features;
[0011] Step 3: Establish an environmental perturbation simulation generation model, use a conditional generative adversarial network to generate virtual samples with different illumination intensities, occlusion ratios, and background noises, and map the generated samples and real-scene feature distributions to a shared latent space through a domain adaptation alignment algorithm;
[0012] Step 4: Design an online incremental feature update mechanism, collect scene change data in real time based on a sliding window strategy, use a contrastive self-supervised learning framework to dynamically optimize the parameters of the feature encoder, and prevent historical feature degradation through an elastic weight consolidation algorithm;
[0013] Step 5: Deploy a heterogeneous graph neural network for multi-modal decision fusion. Construct visible light features, infrared thermal imaging features, and depth sensor point cloud features as heterogeneous nodes, and use an edge attention mechanism to model cross-modal feature correlations to generate a robust joint representation.
[0014] Step 6: Complete identity recognition using a hierarchical verification architecture. First, calculate the face feature similarity threshold through a lightweight siamese network, and then use a triplet verification module based on hypersphere metric learning to perform a secondary discrimination on the suspected matching results.
[0015] Preferably, in step 1, the non-uniform illumination compensation model uses an improved Retinex decomposition method to decompose the image into a reflection component and an illumination component, uses a variational energy function to constrain the smoothness of the illumination component, and predicts the illumination compensation coefficient matrix through a depthwise separable convolutional network.
[0016] Preferably, in step 2, the bidirectional optical flow estimation module uses a recurrent all-to-all field transform (RAFT) algorithm framework, fuses ConvLSTM units in the feature pyramid to capture temporal dependencies, and uses a robust optimization strategy based on the Huber loss for optical flow residual calculation.
[0017] Preferably, in step 3, the domain adaptation alignment algorithm uses the Wasserstein distance metric to measure the distribution difference between the generated samples and the real samples, adds a spectral normalization layer to the generator network, and constructs an adversarial training mechanism through a gradient reversal layer.
[0018] Preferably, in step 4, the contrastive self-supervised learning framework uses a momentum contrast memory bank mechanism, introduces a temporal continuity constraint when constructing positive and negative sample queues, and uses an information noise contrastive estimation loss function based on cosine similarity for feature similarity measurement.
[0019] Preferably, in step 5, the edge attention mechanism is implemented using a multi-head graph attention network, defines the inter-modal correlation degree as a learnable parameter, and fuses a gated recurrent unit in the graph convolution operation to model the temporal evolution law of cross-modal features.
[0020] Preferably, in step 6, the hypersphere metric learning uses a dynamic radius adjustment strategy, optimizes the inter-class margin through an adaptive radius scaling algorithm during the training stage, and uses an angular distance-based soft margin classifier during the verification stage.
[0021] Preferably, the adaptive radius scaling algorithm constructs a differentiable radius prediction network, takes the feature distribution compactness as the input condition, outputs the hypersphere radius adjustment coefficient for each class, and the loss function uses a weighted combination of the margin loss and the radius regularization term.
[0022] Preferably, the multispectral imaging device includes a polarization imaging unit. Before step 1, a polarization state analysis sub-step is added. The Stokes vector decomposition algorithm is used to separate the specular reflection component, and a depolarization feature map is constructed as an auxiliary input channel.
[0023] Preferably, the present invention also includes a face recognition system for a dynamic environment, including:
[0024] Data acquisition and preprocessing module: Collect a face video stream containing visible light and near-infrared bands through a multispectral imaging device. Use a dynamic noise reduction and distortion correction unit to process the original data. The adaptive smoothing algorithm based on spatio-temporal domain joint filtering is used to eliminate motion blur, and the non-uniform illumination compensation model is used to enhance the brightness of low-illumination areas;
[0025] Multi-scale spatio-temporal fusion feature extraction module: Construct a multi-scale spatio-temporal fusion feature extraction network, embed a bidirectional optical flow estimation module in a three-dimensional convolutional neural network to extract the dynamic spatio-temporal features of pose changes and micro-expressions in the face sequence, and use a channel attention mechanism to dynamically weight and fuse cross-modal features;
[0026] Virtual sample generation and domain adaptation module: Establish an environmental perturbation simulation generation model, use a conditional generative adversarial network to generate virtual samples with different illumination intensities, occlusion ratios, and background noises, and map the generated samples and real scene feature distributions to a shared latent space through a domain adaptation alignment algorithm;
[0027] Online incremental feature update module: Design an online incremental feature update mechanism, collect scene change data in real time based on a sliding window strategy, use a contrastive self-supervised learning framework to dynamically optimize the parameters of the feature encoder, and use an elastic weight consolidation algorithm to prevent historical feature degradation;
[0028] Multi-modal decision fusion module: Deploy a heterogeneous graph neural network for multi-modal decision fusion, construct visible light features, infrared thermal imaging features, and depth sensor point cloud features as heterogeneous nodes, use an edge attention mechanism to model the cross-modal feature correlation, and generate a robust joint representation;
[0029] Hierarchical verification module: Use a hierarchical verification architecture to complete identity recognition. First, calculate the face feature similarity threshold through a lightweight Siamese network, and then use a triplet verification module based on hypersphere metric learning to perform a secondary discrimination on the suspected matching results.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] In the data acquisition and preprocessing stage, the multispectral imaging device captures a face video stream containing visible light and near-infrared bands. Combining an adaptive smoothing algorithm for dynamic noise reduction, distortion correction, and spatio-temporal joint filtering, as well as a non-uniform illumination compensation model, it can effectively improve the image quality. Through an improved Retinex decomposition method for non-uniform illumination compensation, the reflection component and illumination component are accurately separated. Using a variational energy function, smoothness constraints, and a depthwise separable convolutional network to predict the compensation coefficient matrix, even in complex illumination environments, the details of the face can be clearly presented, providing a high-quality image basis for subsequent feature extraction and greatly improving the recognition accuracy in low-light and non-uniform illumination scenarios.
[0032] The construction of the multi-scale spatio-temporal fusion feature extraction network is a major highlight. Embedding a bidirectional optical flow estimation module in a three-dimensional convolutional neural network, adopting the Recurrent All Pairs Field Transforms (RAFT) algorithm framework, and combining ConvLSTM units to capture temporal dependencies. The optical flow residual calculation uses a robustness optimization strategy based on the Huber loss, which can accurately extract dynamic spatio-temporal features such as pose changes and micro-expressions in the face sequence. At the same time, the channel attention mechanism dynamically weights and fuses cross-modal features, fully exploiting the value of different modal features, making the system more accurate in capturing and analyzing face dynamic features, enhancing the adaptability to different pose and expression changes, and effectively reducing recognition errors caused by pose and expression changes.
[0033] The application of the environmental perturbation simulation generation model and the domain adaptation alignment algorithm greatly enhances the robustness of the system. Using a conditional generative adversarial network to generate virtual samples with different illumination intensities, occlusion ratios, and background noises, and through a domain adaptation alignment algorithm based on the Wasserstein distance metric, mapping the generated samples and the real-scene feature distributions to a shared latent space, enabling the system to learn features in various complex environments, simulating various adverse factors in the real scene, enhancing the model's adaptability to different environments, and reducing the interference of environmental factors on the recognition results.
[0034] The online incremental feature update mechanism ensures the real-time performance and accuracy of the system. Based on the sliding window strategy, it real-time collects scene change data, adopts a contrastive self-supervised learning framework with a momentum contrast memory bank mechanism to dynamically optimize the parameters of the feature encoder, and combines the elastic weight consolidation algorithm to prevent the degradation of historical features. It can timely adapt to the dynamic changes of the scene, continuously update and optimize the model, and ensure a high recognition accuracy under new environmental conditions.
[0035] The multi-modal decision fusion module constructs heterogeneous nodes from visible light features, infrared thermal imaging features, and depth sensor point cloud features by deploying a heterogeneous graph neural network. It uses an edge attention mechanism implemented by a multi-head graph attention network to model cross-modal feature correlations and generate a robust joint representation. By fully leveraging the complementarity of multi-modal information, it identifies faces from multiple dimensions, further improving the accuracy and reliability of recognition and reducing the false recognition rate.
[0036] The hierarchical verification architecture improves the accuracy and efficiency of recognition. First, a lightweight siamese network is used to calculate the face feature similarity threshold, and then a triplet verification module based on hypersphere metric learning is used to perform a secondary discrimination on the suspected matching results. The hypersphere metric learning adopts a dynamic radius adjustment strategy, optimizing the inter-class margin through an adaptive radius scaling algorithm during the training phase and using an angular distance-based soft margin classifier during the verification phase, which can quickly and accurately determine the face identity, reduce unnecessary computational complexity, and improve the overall recognition efficiency of the system.
[0037] The present patent technology innovatively designs multiple links from data acquisition, feature extraction, environmental adaptation, feature update, decision fusion to identity verification, effectively solving the problems faced by face recognition in dynamic environments such as illumination, pose and expression changes, occlusion, and background noise, significantly improving the accuracy, robustness, real-time performance, and reliability of recognition, and having broad application prospects and important social and economic value. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the working principle diagram of the face recognition method described in the present invention;
[0039] Figure 2 is the working principle diagram of the non-uniform illumination compensation model;
[0040] Figure 3 is the working principle diagram of the bidirectional optical flow estimation module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] Please refer to Figures 1-3 , the present invention provides a technical solution: a face recognition method for a dynamic environment, the method comprising:
[0043] Step 1: Collect a face video stream containing visible light and near-infrared bands through a multispectral imaging device. After collecting the original data, use dynamic noise reduction and distortion correction techniques to process the original data. Adopt an adaptive smoothing algorithm based on spatio-temporal domain joint filtering to eliminate motion blur, and enhance the brightness of low-illumination areas through a non-uniform illumination compensation model.
[0044] Step 2: Construct a multi-scale spatio-temporal fusion feature extraction network. Embed a bidirectional optical flow estimation module in a three-dimensional convolutional neural network to extract the dynamic spatio-temporal features of pose changes and micro-expressions in the face sequence, and adopt a channel attention mechanism to dynamically weight and fuse cross-modal features.
[0045] Step 3: Establish an environmental perturbation simulation generation model. Use a conditional generative adversarial network to generate virtual samples with different illumination intensities, occlusion ratios, and background noises, and map the generated samples and real-scene feature distributions to a shared latent space through a domain adaptation alignment algorithm.
[0046] Step 4: Design an online incremental feature update mechanism. Based on the sliding window strategy, collect scene change data in real time, adopt a contrastive self-supervised learning framework to dynamically optimize the parameters of the feature encoder, and prevent historical feature degradation through an elastic weight consolidation algorithm.
[0047] Step 5: Deploy a heterogeneous graph neural network for multi-modal decision fusion. Construct visible light features, infrared thermal imaging features, and depth sensor point cloud features as heterogeneous nodes, and adopt an edge attention mechanism to model the cross-modal feature correlation to generate a robust joint representation.
[0048] Step 6: Complete identity recognition using a hierarchical verification architecture. First, calculate the face feature similarity threshold through a lightweight siamese network, and then use a triplet verification module based on hypersphere metric learning to perform a secondary discrimination on the suspected matching results.
[0049] The present invention will be further described below in conjunction with Embodiments 1 to 5:
[0050] Embodiment 1:
[0051] In Step 1, the non-uniform illumination compensation model adopts an improved Retinex decomposition method. The Retinex theory believes that an image can be decomposed into a reflection component R and an illumination component L, that is, image I = R·L. When the traditional Retinex decomposition method processes the illumination component, it may not be able to well balance the relationship between illumination smoothness and detail preservation.
[0052] This embodiment uses a variational energy function to constrain the smoothness of the illumination component. The variational energy function E(L) can be expressed as:
[0053]
[0054] where λ1 and λ2 are weight parameters, denotes the gradient operator. The first term is used to constrain the smoothness of the illumination component, so that the illumination change is not too drastic; the second term ensures that the product of the decomposed reflectance component and the illumination component is as close as possible to the original image.
[0055] The illumination compensation coefficient matrix is predicted by a depthwise separable convolutional network. The depthwise separable convolutional network consists of a depthwise convolutional layer and a pointwise convolutional layer. The depthwise convolutional layer is responsible for independently convolving the features of each channel, and the pointwise convolutional layer is used to fuse the channel information. Compared with traditional convolution, depthwise separable convolution can greatly reduce the computational amount and improve the training efficiency of the model. During the training process, a large number of images containing different illumination conditions are used as samples and input into the depthwise separable convolutional network. The network outputs the illumination compensation coefficient matrix, which is used to adjust the illumination component of the original image, so as to enhance the brightness of the low-illumination area and improve the overall quality of the image.
[0056] Example 2:
[0057] In step 2, the bidirectional optical flow estimation module adopts the Recurrent All Pairs Field Transforms (RAFT) algorithm framework. The RAFT algorithm is based on the idea of feature matching and fuses ConvLSTM units in the feature pyramid to capture temporal dependencies. The feature pyramid can capture image features at different scales simultaneously. Small-scale features contain more detailed information, and large-scale features contain more global information. By fusing features at different scales, the image content can be described more comprehensively.
[0058] The ConvLSTM unit has unique advantages when processing time series data. It can remember the information at the previous moment and update according to the current input, so as to effectively capture the temporal dependencies in the face sequence, such as the continuous change of face pose, the dynamic process of micro-expression, etc.
[0059] In terms of optical flow residual calculation, a robustness optimization strategy based on the Huber loss is adopted. The Huber loss function is defined as:
[0060]
[0061] where x is the error of optical flow estimation and δ is a hyperparameter. Compared with the traditional mean squared error loss, the Huber loss is similar to the mean squared error loss when the error is small and can converge quickly; when the error is large, its growth rate is slower than the mean squared error loss, and it is more robust to outliers, which can effectively improve the accuracy of optical flow estimation and further enhance the ability to capture dynamic features.
[0062] Example 3:
[0063] In step 3, the domain adaptation alignment algorithm uses the Wasserstein distance metric to generate the distribution difference between the generated samples and the real samples. The Wasserstein distance W(P,Q) can measure the difference between two probability distributions P and Q, and its definition is:
[0064]
[0065] where Π(P,Q) is the set of all joint distributions of P and Q, denotes the expectation, and ‖·‖ denotes the distance metric. By minimizing the Wasserstein distance, the distribution of the generated samples can be made as close as possible to the distribution of the real samples.
[0066] A spectral normalization layer is added to the generator network. Spectral normalization makes the training of the generator network more stable by normalizing the weight matrix. Specifically, for the weight matrix W, the normalized weight matrix is: where σ(W) is the spectral norm of the weight matrix W.
[0067] An adversarial training mechanism is constructed through a gradient reversal layer. The gradient reversal layer is an identity mapping during forward propagation and multiplies the gradient by -1 during backward propagation. During the training process, the generator network generates virtual samples, and the discriminator network determines whether the samples are real samples or generated samples. Through the gradient reversal layer, the generator network learns how to generate virtual samples closer to the real samples, thereby realizing the mapping of the generated samples and the real scene feature distribution to the shared latent space and enhancing the adaptability of the virtual samples to the real scene.
[0068] Example 4:
[0069] In step 4, the contrastive self-supervised learning framework adopts a momentum contrast memory bank mechanism. The momentum contrast memory bank maintains a queue that stores historical features. During the training process, the features of the current sample are compared and learned with the features in the queue. When constructing the positive and negative sample queues, a temporal continuity constraint is introduced, that is, samples at adjacent times are more likely to belong to the same category. Samples at adjacent times are used as positive samples, and other samples are used as negative samples.
[0070] The feature similarity metric uses an information noise contrastive estimation loss function based on cosine similarity. The cosine similarity sim(x,y) is defined as: where x and y are two feature vectors. The information noise contrastive estimation loss function L can be expressed as:
[0071]
[0072] Where x is the current sample feature, q is the positive sample feature, k is the sample feature in the negative sample queue, K is the size of the negative sample queue, and τ is the temperature hyperparameter. By minimizing this loss function, the model can learn more discriminative feature representations. At the same time, combined with the elastic weight consolidation algorithm, when updating the parameters of the feature encoder, important historical weights are protected to prevent the degradation of historical features, ensuring that the model does not lose the useful information learned previously while continuously updating features.
[0073] Example 5:
[0074] In step 5, the edge attention mechanism is implemented using a multi-head graph attention network. The multi-head graph attention network can capture the correlations between cross-modal features from different perspectives by calculating attention weights in parallel for multiple heads. Define the inter-modal correlation degree as a learnable parameter, and obtain the importance between different modalities through training. The gated recurrent unit is fused in the graph convolution operation to model the temporal evolution law of cross-modal features. The gated recurrent unit can adaptively update cross-modal features according to the current input and the previous state, better capturing the changes of features over time.
[0075] In step 6, the hypersphere metric learning adopts a dynamic radius adjustment strategy. In the training stage, the inter-class margin is optimized through an adaptive radius scaling algorithm. The adaptive radius scaling algorithm constructs a differentiable radius prediction network, takes the feature distribution compactness as the input condition, and outputs the hypersphere radius adjustment coefficient for each category. The loss function adopts a weighted combination of the margin loss and the radius regularization term. The margin loss L margin can be expressed as:
[0076] L margin = max(0, m + d(a, p) - d(a, n))
[0077] Where m is the preset margin threshold, d(a, p) is the distance between the anchor sample and the positive sample, and d(a, n) is the distance between the anchor sample and the negative sample. The radius regularization term is used to prevent the hypersphere radius from being too large or too small, ensuring the stability of the model. In the validation stage, a soft margin classifier based on angular distance is adopted. The angular distance can better measure the similarity between features, and the soft margin classifier can alleviate the overfitting problem to a certain extent and improve the accuracy of identity recognition.
[0078] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0079] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A face recognition method for a dynamic environment, characterized in that: The following steps are involved: Step 1: Use multispectral imaging equipment to collect face video streams containing visible light and near-infrared bands, perform dynamic noise reduction and distortion correction on the original data, use an adaptive smoothing algorithm based on spatiotemporal joint filtering to eliminate motion blur, and use a non-uniform illumination compensation model to enhance the brightness of low-illuminance areas; Step 2: Construct a multi-scale spatiotemporal fusion feature extraction network, embed a bidirectional optical flow estimation module in the three-dimensional convolutional neural network, extract the dynamic spatiotemporal features of posture changes and micro-expressions in the face sequence, and use the channel attention mechanism to dynamically weighted fuse cross-modal features; Step 3: Establish an environmental disturbance simulation generation model, use the conditional generative adversarial network to generate virtual samples with different light intensities, occlusion ratios, and background noises, and map the generated samples and the real scene feature distribution to a shared latent space through a domain adaptive alignment algorithm; Step 4: Design an online incremental feature update mechanism to collect scene change data in real time based on a sliding window strategy, dynamically optimize feature encoder parameters using a contrastive self-supervised learning framework, and prevent historical feature degradation through an elastic weight solidification algorithm; Step 5: Deploy heterogeneous graph neural networks for multimodal decision fusion, construct visible light features, infrared thermal imaging features, and depth sensor point cloud features as heterogeneous nodes, use edge attention mechanism to model cross-modal feature correlation, and generate robust joint representation; Step 6: Use a hierarchical verification architecture to complete identity recognition. First, calculate the facial feature similarity threshold through a lightweight twin network, and then use a triplet verification module based on hypersphere metric learning to perform secondary judgment on the suspected matching results.
2. The face recognition method for dynamic environment according to claim 1, characterized in that: The non-uniform illumination compensation model in step 1 adopts an improved Retinex decomposition method to decompose the image into a reflection component and an illumination component, uses a variational energy function to constrain the smoothness of the illumination component, and predicts the illumination compensation coefficient matrix through a deep separable convolutional network.
3. The face recognition method for dynamic environment according to claim 1, characterized in that: The bidirectional optical flow estimation module in step 2 adopts the RAFT algorithm framework, integrates the ConvLSTM unit in the feature pyramid to capture the temporal dependency, and the optical flow residual calculation adopts the robustness optimization strategy based on Huber loss.
4. The face recognition method for dynamic environment according to claim 1, characterized in that: In step 3, the domain adaptive alignment algorithm uses the Wasserstein distance to measure the distribution difference between the generated samples and the real samples, adds a spectral normalization layer to the generator network, and constructs an adversarial training mechanism through a gradient reversal layer.
5. The face recognition method for dynamic environment according to claim 1, characterized in that: In the step 4, the comparative self-supervised learning framework adopts a momentum comparative memory bank mechanism, introduces time continuity constraints when constructing positive and negative sample queues, and the feature similarity measurement adopts an information noise comparative estimation loss function based on cosine similarity.
6. The face recognition method for dynamic environment according to claim 1, characterized in that: The edge attention mechanism in step 5 is implemented using a multi-head graph attention network, the inter-modal correlation is defined as a learnable parameter, and the gated recurrent unit is integrated in the graph convolution operation to model the temporal evolution of cross-modal features.
7. The face recognition method for dynamic environment according to claim 1, characterized in that: In step 6, the hypersphere metric learning adopts a dynamic radius adjustment strategy, optimizes the inter-class interval through an adaptive radius scaling algorithm in the training phase, and adopts a soft boundary classifier based on angular distance in the verification phase.
8. The face recognition method for dynamic environment according to claim 7, characterized in that: The adaptive radius scaling algorithm constructs a differentiable radius prediction network, takes the feature distribution compactness as an input condition, outputs the hypersphere radius adjustment coefficient of each category, and the loss function adopts a weighted combination of margin loss and radius regularization term.
9. The face recognition method for dynamic environment according to claim 1, characterized in that: The multi-spectral imaging device includes a polarized light imaging unit, and a polarization state analysis sub-step is added before step 1, a Stokes vector decomposition algorithm is used to separate the specular reflection component, and a depolarization feature map is constructed as an auxiliary input channel.
10. A face recognition system for dynamic environments, characterized in that: include: Data acquisition and preprocessing module: collects face video streams containing visible light and near-infrared bands through multispectral imaging equipment, processes the original data using dynamic noise reduction and distortion correction units, eliminates motion blur using an adaptive smoothing algorithm based on spatiotemporal joint filtering, and enhances the brightness of low-illuminance areas with the help of a non-uniform illumination compensation model; Multi-scale spatiotemporal fusion feature extraction module: Construct a multi-scale spatiotemporal fusion feature extraction network, embed a bidirectional optical flow estimation module in the three-dimensional convolutional neural network to extract the dynamic spatiotemporal features of posture changes and micro-expressions in face sequences, and use the channel attention mechanism to dynamically weighted fuse cross-modal features; Virtual sample generation and domain adaptation module: establish an environmental perturbation simulation generation model, use conditional generative adversarial networks to generate virtual samples with different light intensities, occlusion ratios, and background noises, and map the generated samples and real scene feature distributions to a shared latent space through a domain adaptive alignment algorithm; Online incremental feature update module: Design an online incremental feature update mechanism, collect scene change data in real time based on the sliding window strategy, use a contrastive self-supervised learning framework to dynamically optimize feature encoder parameters, and use an elastic weight solidification algorithm to prevent historical feature degradation; Multimodal decision fusion module: deploys heterogeneous graph neural networks for multimodal decision fusion, constructs visible light features, infrared thermal imaging features, and depth sensor point cloud features into heterogeneous nodes, uses edge attention mechanism to model cross-modal feature correlation, and generates robust joint representation; Hierarchical verification module: A hierarchical verification architecture is used to complete identity recognition. The facial feature similarity threshold is first calculated through a lightweight twin network, and then a triplet verification module based on hypersphere metric learning is used to perform secondary judgment on the suspected matching results.
Citation Information
Patent Citations
Aircraft detection and tracking method based on multi-scale self-adaption and side domain attention
CN113792631A
Virtual image generation and fusion method for face recognition
CN115171190A
Face recognition system based on machine learning
CN117115881A
Traffic target detection method and system based on cross-modal cross attention mechanism
CN117173399A
Video transmission semantic communication system
CN117579840A
Cited By
Target identification method and system based on deep learning
CN120876834A
A target recognition method and system based on deep learning
CN120876834B
Gate identity recognition acceleration method based on deep learning
CN121524929A
A Deep Learning-Based Method for Accelerating Gate Identity Recognition
CN121524929B
Micro-expression recognition method and system based on balanced adaptive grouped sampling
CN121768057A