Face recognition system based on computer vision and application method thereof
By introducing fine-grained occlusion classification, multimodal feature fusion, cross-modal adversarial identification and compensation verification modules into the face recognition system, the problems of occlusion and forgery attacks are solved, and the robustness and security of the identification system are improved.
Patent Information
- Application Number
- CN202510622356.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing computer vision-based facial recognition technology is difficult to deal with the incomplete information caused by occlusion in high-density public places and special environments, and there is a risk of forgery attacks, resulting in insufficient reliability and security of the identification system.
A face recognition system based on computer vision is adopted, including a fine-grained occlusion classification module, a multimodal feature fusion module, a cross-modal adversarial identification module and a compensation verification module. The occlusion type is analyzed through depth data and polarized light parameters, and infrared and acoustic sensor data are dynamically activated. Feature extraction and weight allocation are combined with spatiotemporal convolutional layers and a cross-modal attention mechanism to generate a predicted facial occlusion area feature map, and cross-modal distribution similarity calculation and forgery attack verification are performed.
It significantly improves the robustness and anti-counterfeiting capabilities of the face recognition system in the case of occlusion, enhances the security and credibility of the recognition results, and can achieve high-precision and high-security real-time recognition in complex environments.
Smart Images

Figure CN120148092A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face recognition, and specifically to a face recognition system based on computer vision and its application method. Background Art
[0002] With the rapid development of fields such as smart cities, public security, and financial payment, face recognition technology, as the core means of biometric recognition, has been widely used in scenarios such as identity verification, security monitoring, and personalized services. However, in high-density public places and special environments, users often wear masks, hats, scarves and other occluders, resulting in incomplete face information, which seriously reduces the reliability of the recognition system. At the same time, the threats of face data abuse and forgery attacks are increasing day by day. How to achieve high-precision and high-security real-time recognition in complex occlusion scenarios has become the key challenge restricting the development of the industry.
[0003] Current face recognition technologies based on computer vision cannot distinguish the physical materials of occluders, resulting in the same processing strategy for hard occlusion and soft occlusion, with low compensation accuracy. Single-modal data is difficult to penetrate the occlusion area, and relying on manually designed features is vulnerable to interference from illumination changes and pose offsets. Existing multi-modal solutions often use fixed-weight fusion and cannot dynamically allocate sensor resources according to the occlusion type, resulting in wasted computing power and excessive load on edge devices, and it is difficult to capture the dynamic physiological characteristics of the occlusion area. Existing defense means rely on a single modality and are easily bypassed by multi-modal forgery attacks. Centralized training requires uploading original face data, posing a risk of privacy leakage. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a face recognition system based on computer vision and its application method, which solves the problems in the above background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A face recognition system based on computer vision, including the following modules: a fine-grained occlusion classification module, a multi-modal feature fusion module, a cross-modal adversarial discrimination module, and a compensation verification module; the fine-grained occlusion classification module is used to analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable area and the occlusion area; the multi-modal feature fusion module is used to dynamically activate the infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and the acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; the cross-modal adversarial discrimination module is used to calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated; the compensation verification module is used to receive a multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensation distortion, and at the same time combine the confidence score and the attack mark to eliminate the low-confidence compensation data and attack frames, and output a complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification.
[0006] Further, the specific process of analyzing the depth data and polarization light parameters of the input image, classifying the occlusion types through the material reflection characteristics, and generating an occlusion heat map is as follows: fuse the depth data and polarization light parameters of the input image, divide the physical distance levels of the occlusion area based on the depth information, combine the polarization angle distribution difference and the material reflectivity to distinguish between hard occlusion and soft occlusion types, generate a pixel-level occlusion heat map and label the material label, where the hard occlusion area is marked as a completely unrecognizable area, and the soft occlusion area is marked as a partially compensable area.
[0007] Further, the specific process of extracting the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and the acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer is as follows: According to the multi-modal spatio-temporal convolutional kernel, continuously slide along the time dimension and synchronously capture the dynamic temperature gradient change of the infrared thermal distribution and the vibration frequency time series correlation of the acoustic spectrum; eliminate the environmental noise interference through cross-modal feature cross-verification, and generate an encoded feature vector that fuses the spatio-temporal correlation of infrared and acoustic waves, which is used to characterize the coordination of the physiological activities and physical movements in the occlusion area.
[0008] Furthermore, the specific process of dynamically allocating the contribution weights of infrared and acoustic wave features based on the cross-modal attention mechanism to generate the predicted facial occlusion region feature map is as follows: Using the material type label of the heat map as the conditional vector, concatenating it with the multi-modal fusion feature vector, and inputting it into the attention network; Dynamically allocating the contribution weights of the infrared heat distribution and the acoustic wave spectrum according to the occlusion type label and the modal saliency of the fusion feature vector, performing channel-level weighted fusion on the infrared heat distribution temporal features and the acoustic wave spectrum temporal features to generate a compensated feature vector for the occlusion region; Inputting the weighted compensated feature vector into the generator network, restoring the facial details of the occlusion region through multi-level upsampling to generate the predicted facial occlusion region feature map; And imposing cross-modal consistency constraints to ensure the alignment of the predicted features with the visible light features in the non-occluded region in the semantic space.
[0009] Furthermore, the specific process of calculating the cross-modal distribution similarity between the generated predicted facial occlusion region feature map and the visible light features in the non-occluded region to output a confidence score is as follows: Mapping the generated predicted occlusion region feature map and the visible light features in the non-occluded region to the shared embedding space, calculating the cosine similarity between the two through the contrastive learning algorithm, and combining the confidence decay factor of the occlusion region to generate a normalized confidence score, where the confidence score is used to characterize the alignment degree between the generated features and the real facial features in the semantic space.
[0010] Furthermore, when the acoustic wave sensor detects a voice trigger event, the synchronization between the temperature change of the target region in the infrared image and the acoustic wave spectrum is verified. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated. The specific process is as follows: Aligning the peak of the acoustic wave spectrum and the infrared temperature temporal curve according to the voice trigger timestamp, calculating the temporal matching degree between the two through the dynamic time warping algorithm. If the matching degree is lower than the adaptive threshold, it is determined as a cross-modal forgery attack, generating an attack mark and triggering an alarm signal, and at the same time freezing the feature fusion weights of the current frame to block the spread of the attack.
[0011] Furthermore, the specific process of receiving a multi-frame prediction feature sequence and analyzing the difference rate of the occlusion region features of adjacent frames through a sliding window to mark compensation distortion is as follows: Intercepting a frame interval of a fixed length in the continuous multi-frame prediction feature sequence, calculating the Euclidean distance difference rate of the occlusion region features of adjacent frames frame by frame, and determining the compensation distortion through the adaptive threshold. Specifically, it includes: Performing pixel-by-pixel comparison on the occlusion region feature maps of each pair of adjacent frames in the window, calculating the average value of the Euclidean distance as the difference rate; If the occlusion type is hard occlusion, the difference rate threshold is automatically reduced to enhance the sensitivity to small feature jumps; If the difference rate exceeds the adaptive threshold of the current window, mark this frame interval as compensation distortion and generate a distortion type label.
[0012] Furthermore, the specific process of combining the confidence score with the attack flag to eliminate low-confidence compensation data and attack frames and output the complete facial feature vector and risk level through spatio-temporal coherence verification and cross-modal logic verification is as follows: conduct a logical correlation analysis on the low-confidence compensation data, attack flags, and temporal distortion labels. If multiple types of abnormal signals exist within the same frame interval, it is marked as a composite anomaly; the low-confidence compensation data refers to the feature data with a confidence score lower than the preset confidence score threshold; dynamically allocate weights according to the type of anomaly and the area of the occluded region, and calculate the comprehensive risk score; if the risk score reaches the high-risk level, only output the feature vector of the unoccluded region and trigger a warning; if the risk score is at the medium-risk level, output the complete feature vector but restrict the downstream sensitive operation permissions; if the risk score is at the low-risk level, directly output the complete feature vector and conduct identity verification.
[0013] A computer vision-based face recognition method includes the following steps: S1. Analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable region and the occluded region; S2. Dynamically activate the infrared and acoustic sensor data according to the occlusion type labels of the heat map, extract the temperature change patterns and vibration frequency correlations in the infrared thermal distribution time series data and acoustic frequency spectrum time series data of the occluded region through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occluded region feature map; S3. Calculate the cross-modal distribution similarity between the generated predicted facial occluded region feature map and the visible light features of the unoccluded region to output a confidence score, and verify the synchronization of the temperature change in the target region of the infrared image and the acoustic frequency spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a spoofing attack and an attack flag is generated; S4. Receive a multi-frame prediction feature sequence and analyze the difference rate of the occluded region features of adjacent frames through a sliding window to mark the compensation distortion. At the same time, combine the confidence score with the attack flag to eliminate low-confidence compensation data and attack frames, and output the complete facial feature vector and risk level through spatio-temporal coherence verification and cross-modal logic verification.
[0014] The present invention has the following beneficial effects:
[0015] (1) The computer vision-based face recognition system significantly improves the robustness of the face recognition system under occlusion through the collaborative action of the fine-grained occlusion classification module and the multi-modal feature fusion module. The fine-grained occlusion classification module utilizes depth data and polarization light parameters to achieve precise classification of occlusions of different materials and generate a pixel-level occlusion heat map, thereby accurately calibrating the recognizable area and the occlusion area and improving the detection accuracy of the occlusion area. The multi-modal feature fusion module dynamically activates the infrared and acoustic sensors based on the occlusion type label, extracts the temporal features in the infrared heat distribution and acoustic spectrum through the spatio-temporal convolutional layer, and adaptively adjusts the feature weights through the cross-modal attention mechanism, enabling the system to effectively compensate for the information loss in the occlusion area, thereby improving the integrity and stability of face feature extraction.
[0016] (2) The computer vision-based face recognition method enhances the anti-counterfeiting ability and data credibility of the face recognition system through the design of the cross-modal adversarial discrimination module and the compensation verification module. The cross-modal adversarial discrimination module combines visible light, infrared, and acoustic data to calculate the distribution similarity of cross-modal features, and when a voice trigger event is detected by the acoustic sensor, it verifies whether there is a forgery attack through the synchrony of the temperature change in the infrared image and the acoustic spectrum, thereby effectively improving the system's detection ability for high-fidelity spoofing attacks. The compensation verification module analyzes the change rate of the occlusion area features of adjacent frames through a sliding window to identify compensation distortion, and combines the confidence score and attack flag to eliminate low-confidence or deceptive data, and finally outputs a complete facial feature vector and risk level after spatio-temporal consistency verification and cross-modal logical verification, improving the security and credibility of the recognition result.
[0017] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of the computer vision-based face recognition system of the present invention.
[0019] Figure 2 It is a detailed flowchart of the compensation verification module of the present invention
[0020] Figure 3 It is a flowchart of the computer vision-based face recognition method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The embodiments of the present application solve the problems of decreased recognition accuracy and insufficient security caused by occlusion, environmental interference, and forgery attacks in the face recognition process through the computer vision-based face recognition system and its application method.
[0022] The overall idea of the solution in the embodiments of the present application is as follows:
[0023] Analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable areas and occlusion areas.
[0024] Dynamically activate the infrared and acoustic sensor data according to the occlusion type labels of the heat map, extract the temporal data of the infrared thermal distribution and the correlation of the vibration frequency in the temporal data of the acoustic spectrum in the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map.
[0025] Calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a spoofing attack and an attack mark is generated.
[0026] Receive a multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark compensation for distortion. At the same time, combine the confidence score and the attack mark to eliminate low-confidence compensation data and attack frames, and output a complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification.
[0027] Please refer to Figure 1, an embodiment of the present invention provides a technical solution: a face recognition system based on computer vision, including the following modules: a fine-grained occlusion classification module, a multi-modal feature fusion module, a cross-modal adversarial discrimination module, and a compensation verification module; the fine-grained occlusion classification module is used to analyze the depth data and polarized light parameters of the input image, classify the occlusion type through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable area and the occlusion area; the multi-modal feature fusion module is used to dynamically activate the infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; the cross-modal adversarial discrimination module is used to calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated; the compensation verification module is used to receive a multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensation distortion. At the same time, combined with the confidence score and the attack mark, the low-confidence compensation data and attack frames are excluded, and a complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification are output.
[0028] In this implementation, a highly robust face recognition system for complex occlusion scenarios of the present invention has its core in solving the key problems of occlusion recognition, feature compensation, attack defense, and dynamic stability through multi-module collaboration. The working principles and technical advantages are described in the following sub-modules. The system includes four core modules, forming a closed-loop processing chain of "occlusion classification, multi-modal fusion, attack defense, and compensation verification": Input: visible light image (RGB), depth map (TOF), polarized light data, infrared thermal imaging, acoustic spectrum; Processing: Gradually transmit intermediate results such as occlusion heat maps, predicted feature maps, confidence scores, etc.; Output: complete facial feature vectors (for identity verification) and risk level labels (for permission control). Fine-grained occlusion classification module: Distinguish occlusion types (hard / soft occlusion) and generate heat maps. Technical implementation: Input depth map (physical distance) and polarized light parameters (material reflection characteristics), and determine the type of occluder (such as mask = hard occlusion, scarf = soft occlusion) through a material classification network (such as a ResNet variant); Output pixel-level heat maps, marking recognizable areas (such as forehead, eyebrow bone) and occlusion areas (such as the lower half of the face covered by the mask). Break through the limitations of traditional RGB texture analysis, and combine depth and polarized light to achieve accurate material discrimination (such as distinguishing a plastic mask from a cloth scarf). Multi-modal feature fusion module: Dynamically select sensor data and generate compensated features for the occlusion area. Technical implementation: Dynamically activate sensors according to the occlusion labels of the heat maps: for hard occlusion (such as a mask), focus on infrared thermal distribution (detecting the temperature rise of the nose tip during breathing); for soft occlusion (such as a scarf), focus on acoustic spectrum (analyzing the correlation of voice vibrations); Extract the temporal correlation of temperature changes and vibration frequencies through spatio-temporal convolution (such as the synchronization of breathing frequency and acoustic spectrum); The cross-modal attention mechanism dynamically allocates sensor weights to generate a predicted feature map of the occlusion area (such as the chin contour under the mask). Avoid wasting resources with fixed fusion rules and improve edge computing efficiency; Capture the physical correlation between physiological activities (breathing, speech) and occlusion compensation. Cross-modal adversarial discrimination module function: Verify the credibility of the generated features and defend against forgery attacks. Technical implementation: Cross-modal similarity calculation: Map the generated features and the unoccluded area of visible light to a shared semantic space and calculate the cosine similarity (such as the anatomical continuity between the generated chin under the mask and the real forehead structure); Attack defense: Verify the acoustic-infrared synchronization when the voice is triggered (such as the lip temperature rise during pronunciation needs to match the acoustic peak within 0.5 seconds) to determine asynchronous attacks (such as recorded playback + heated mask). Semantic alignment constraint: Solve the problem of comparability of heterogeneous data through a shared embedding space; Multi-modal attack interception: Combine acoustic and infrared data to defend against highly realistic attacks. Compensation verification module function: Eliminate abnormal data and output complete features and risk levels.Technical implementation: Temporal coherence verification: Analyze the difference rate of multi-frame features through a sliding window (such as the jump of the chin contour caused by head turning), and mark and compensate for distortion; Risk grading: Integrate low-confidence data, attack marks, and temporal distortion to output low / medium / high risk levels (such as high risk prohibits payment, low risk for normal verification). Dynamic threshold adjustment: Adaptively set the difference rate threshold according to the occlusion type and environmental noise; Grading permission control: Balance security and user experience. Through the full-link innovation of physical-level occlusion classification, dynamic multi-modal fusion, cross-modal attack defense, and temporal risk control, the present invention solves the accuracy, security, and real-time performance of face recognition in complex occlusion scenarios, and at the same time meets privacy compliance requirements through lightweight design and federated learning.
[0029] Specifically, the specific process of analyzing the depth data and polarization light parameters of the input image and classifying the occlusion type through the material reflection characteristics to generate an occlusion heat map is as follows: Integrate the depth data and polarization light parameters of the input image, divide the physical distance levels of the occlusion area based on the depth information, combine the polarization angle distribution difference and the material reflectivity to distinguish between hard occlusion and soft occlusion types, and generate a pixel-level occlusion heat map and label the material tags, where the hard occlusion area is marked as a completely unrecognizable area, and the soft occlusion area is marked as a partially compensable area.
[0030] In this implementation, integrate depth data and polarization light parameters Depth data: Use an RGB-D camera or other depth sensors to obtain the three-dimensional information of the scene, and the depth value of each pixel represents its physical distance from the camera. Polarization light parameters: Obtain the light intensity distribution at different polarization angles through a polarization camera, calculate the polarization angle distribution difference and the material reflectivity, and use them to judge the scattering and transmission of light on the surface. Divide the physical distance levels of the occlusion area: According to the depth data, stratify the objects in the scene by distance, for example: close range (0 - 50 cm), medium range (50 - 100 cm), long range (more than 100 cm). Occlusions at close range have a greater impact on the face, while long-range occlusions may interfere less with recognition. Distinguish between hard occlusion and soft occlusion: Hard occlusion (non-compensable): If the depth data of an area changes suddenly and the polarization angle changes violently, it indicates that this area may be a solid object (such as a hand, a mask, a wall, etc.), and it is marked as a completely unrecognizable area. Soft occlusion (partially compensable): If the polarization light parameters show a large transmittance, it indicates that this area may be a semi-transparent material (such as glass, gauze, fog, etc.), and it is marked as a partially compensable area. Generate a pixel-level occlusion heat map, combine the depth data and polarization light information, calculate the occlusion type pixel by pixel, and generate an occlusion heat map. The hard occlusion area is marked as high risk (red), indicating that this area cannot be recognized; the soft occlusion area is marked as low risk (yellow), indicating that it can be compensated by combining other sensors.
[0031] Specifically, the specific process of extracting the correlation between the temperature change pattern and the vibration frequency in the infrared thermal distribution time series data and the acoustic spectrum time series data of the occluded area through the spatio-temporal convolutional layer is as follows: According to the multi-modal spatio-temporal convolution kernel, continuously slide along the time dimension and synchronously capture the dynamic temperature gradient change of the infrared thermal distribution and the vibration frequency time series correlation of the acoustic spectrum; eliminate environmental noise interference through cross-modal feature cross-validation, and generate an encoded feature vector that fuses the spatio-temporal correlation of infrared and acoustic waves, which is used to characterize the coordination of physiological activities and physical movements in the occluded area.
[0032] In this implementation scheme, in order to extract the correlation between the temperature change pattern and the vibration frequency from the occluded area, a spatio-temporal convolutional layer is used to jointly process the infrared thermal distribution data and the acoustic spectrum data to capture the dynamic features that change over time. The specific process is as follows: Based on spatio-temporal convolution for dynamic feature extraction, assume that the input infrared thermal distribution in the time series data is , where t represents the time frame, represents the spatial coordinates. Assume that the input acoustic spectrum data is , where f represents the frequency component and t represents the time frame. Define a three-dimensional spatio-temporal convolution kernel , where is the time window size, and m, n are the spatial dimensions or spectral dimensions. Adopt a sliding window strategy on the time axis, with a step size of to calculate the local temperature gradient for the infrared thermal data : ; where, represents the local temperature gradient, which measures the dynamic change of the infrared thermal distribution. Extract the time correlation features of the acoustic wave vibration frequency: perform a short-time Fourier transform (STFT) on the input acoustic spectrum to obtain the frequency-time distribution matrix: ; use the same spatio-temporal convolution kernel to perform spatio-temporal convolution on the acoustic spectrum: ; where, represents the vibration frequency change pattern on the time axis. Cross-modal feature cross-validation to eliminate environmental noise: Calculate the correlation between the infrared thermal data and the acoustic spectrum data to eliminate environmental noise: ; where, is the mean value, is the standard deviation. Only retain the data with a correlation higher than the threshold , and finally obtain the fused spatio-temporal correlation encoded feature vector: ; This feature vector is used to describe the coordination of physiological activities (temperature change) and physical movements (vibration frequency) in the occluded area, and improve the accuracy of recognition.
[0033] Specifically, based on the cross-modal attention mechanism, the contribution weights of infrared and acoustic features are dynamically allocated, and the specific process of generating the predicted facial occlusion region feature map is as follows: Using the material type label of the heat map as a conditional vector, concatenating it with the multi-modal fusion feature vector, and inputting it into the attention network; According to the occlusion type label and the modal saliency of the fusion feature vector, dynamically allocate the contribution weights of the infrared heat distribution and the acoustic spectrum, perform channel-level weighted fusion on the infrared heat distribution temporal features and the acoustic spectrum temporal features to generate a compensated feature vector for the occlusion region; Input the weighted compensated feature vector into the generator network, restore the facial details of the occlusion region through multi-level upsampling, and generate the predicted facial occlusion region feature map; And impose cross-modal consistency constraints to ensure the alignment of the predicted features with the visible light features in the unoccluded region in the semantic space.
[0034] In this implementation, in order to reasonably utilize the information of different modalities, a cross-modal attention mechanism is adopted to dynamically adjust the feature weights of the infrared heat distribution and the acoustic spectrum. The specific process is as follows: Construct the input features, the infrared heat distribution feature vector is , the acoustic spectrum feature vector is , and the occlusion type label is . Form the concatenated vector: ; where serves as a conditional vector, providing prior information on the occlusion material type. Calculate the cross-modal attention weights, and calculate the contribution weights of infrared and acoustic waves through a trainable attention network : ; where the attention mechanism adopts softmax normalization: ; where and are weight parameters learned by the neural network. Fuse the feature vectors, and perform channel-level weighted fusion using the calculated weights: ; This fused feature vector retains the multi-modal information of the occlusion region and enhances the features helpful for face reconstruction. Input the generator network to restore the facial details of the occlusion region, and use the generator network in the generative adversarial network (GAN) to complete the facial feature complementation: ; where is the predicted facial occlusion region feature map. Impose cross-modal consistency constraints. To ensure the alignment of the generated occlusion region features with the visible light features in the unoccluded region in the semantic space, a cross-modal consistency loss is introduced. Finally, the facial occlusion region feature map generated by this process can more accurately restore the facial details, maintain the feature consistency with the unoccluded region, and improve the reliability of the system for face recognition in the case of occlusion.
[0035] Specifically, the specific process of calculating the cross-modal distribution similarity between the generated predicted facial occlusion region feature map and the visible light features of the unoccluded region to output a confidence score is as follows: Map the generated predicted occlusion region feature map and the visible light features of the unoccluded region to a shared embedding space, calculate the cosine similarity between the two through a contrastive learning algorithm, and combine the confidence decay factor of the occlusion region to generate a normalized confidence score, which is used to characterize the alignment degree between the generated features and the real facial features in the semantic space.
[0036] In this implementation, by calculating the cross-modal distribution similarity between the generated facial occlusion region feature map and the visible light features of the unoccluded region, the correspondence between the generated features and the real facial features is evaluated. The specific process is as follows: Mapping to the shared embedding space. To compare the generated facial occlusion region feature map and the visible light features of the unoccluded region, we first map them to a shared embedding space to ensure that features of different modalities can be compared in the same semantic space. The generated facial occlusion region feature map is , and the visible light features of the occlusion region are . Use a mapping network to transform these two feature maps into the shared embedding space: , ; where is the mapping function of the shared embedding space. Calculate the cross-modal similarity: To measure the similarity between the two features, the cosine similarity is used to calculate the correlation between the two embedding space vectors and . The cosine similarity formula is: ; where represents the dot product of the vectors, and represents the norm (magnitude) of the vector. The range of this similarity value is [-1, 1], and the larger the value, the higher the similarity between the two features. Introduction of the confidence decay factor. Based on the cross-modal similarity calculation, the confidence decay factor of the occlusion region is combined to adjust the final confidence score. Set the confidence decay factor to for adjusting the calculation result according to the difficulty or uncertainty of the occlusion region: ; where ranges from [0, 1] and is used to simulate the influence of uncertainty. Normalize the confidence score. The final confidence score will be standardized through a normalization operation for easier comparison between different occlusion regions. Set the maximum and minimum confidence scores to and respectively, and the normalization process is: ; The finally obtained normalized confidence score ConfidenceScore^ ranges from 0 to 1, indicating the alignment degree between the generated features and the real facial features in the semantic space.
[0037] Specifically, when the acoustic wave sensor detects a voice trigger event, the synchronization between the temperature change in the target area of the infrared image and the acoustic wave spectrum is verified. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated. The specific process is as follows: Align the peak of the acoustic wave spectrum and the infrared temperature time series curve according to the voice trigger timestamp, calculate the temporal matching degree between the two through the dynamic time warping algorithm. If the matching degree is lower than the adaptive threshold, it is determined as a cross-modal forgery attack, an attack mark is generated and an alarm signal is triggered. At the same time, freeze the feature fusion weight of the current frame to block the attack propagation.
[0038] In this implementation scheme, in this experimental scheme, in order to detect cross-modal forgery attacks, we use the temporal synchronization between the acoustic wave spectrum and the infrared image for verification. If the temporal deviation exceeds the preset threshold, it is determined as a forgery attack and a corresponding alarm signal is sent. The specific process is as follows: When the acoustic wave sensor detects a voice trigger event, we first obtain the trigger timestamp , and align the temperature change time series data in the infrared image according to this timestamp. Let the acoustic wave spectrum data be , and the temperature time series of the infrared image be , where is time, is the frequency component of the acoustic wave, is the spatial coordinate of the infrared image. Synchronize the infrared image and the acoustic wave spectrum according to the timestamp to obtain the aligned temperature data and the acoustic wave spectrum data . Calculate the spatio-temporal matching degree: Dynamic Time Warping (DTW) In order to accurately measure the synchronization between the infrared image temperature time series and the acoustic wave spectrum time series, the Dynamic Time Warping (DTW) algorithm is used for time series alignment and similarity calculation. The purpose of the DTW algorithm is to find the optimal matching path between two time series data, so as to evaluate their temporal consistency. Set the distance metric function of the dynamic time warping algorithm, and calculate the optimal matching path on the time axis: ; Obtain the matching degree through dynamic programming, and measure the temporal synchronization between the two. Compare with the preset value to determine the forgery attack If the temporal matching degree is lower than the adaptive threshold , it means that the synchronization between the infrared image temperature change and the acoustic wave spectrum is insufficient, and there may be a forgery attack. Set the forgery attack mark as , and the judgment condition is: Trigger an alarm signal and freeze the feature fusion weights of the current frame to block the spread of the attack.
[0039] Please refer to Figure 2 , specifically, the specific process of receiving a multi-frame prediction feature sequence and analyzing the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensated distortion is as follows: intercept a frame interval of a fixed length from the continuous multi-frame prediction feature sequence, calculate the Euclidean distance difference rate of the occlusion area features of adjacent frames frame by frame, and determine the compensated distortion through an adaptive threshold, specifically including: performing pixel-by-pixel comparison on the occlusion area feature maps of each pair of adjacent frames within the window, and calculating the average value of the Euclidean distance as the difference rate; if the occlusion type is hard occlusion, the difference rate threshold is automatically reduced to enhance the sensitivity to small feature jumps; if the difference rate exceeds the adaptive threshold of the current window, mark this frame interval as compensated distortion and generate a distortion type label.
[0040] In this implementation scheme, this implementation scheme aims to analyze the prediction feature sequence of continuous multi-frames through a sliding window, calculate the difference rate of the occlusion area features of adjacent frames, and judge whether there is compensated distortion according to the difference rate. The specific process is as follows: intercept a frame interval of a fixed length Intercept a frame interval of a fixed length from the continuous multi-frame prediction feature sequence, and set the window length to , and this window contains frame data, and calculate the difference rate between adjacent frames frame by frame. Calculate the Euclidean distance difference rate For the occlusion area feature maps of each pair of adjacent frames within the window and , perform pixel-by-pixel comparison, calculate the Euclidean distance difference between them, and the formula is: ; where, and respectively represent the occlusion area feature values at pixel (i,j) in frame t and frame t+1. After calculating the Euclidean distance of each pair of adjacent frames, the difference rate between each pair of adjacent frames is obtained. Calculate the average value of the difference rate After calculating the difference rate of all adjacent frames within the window, the average difference rate DiffRate is obtained: where, is the size of the window, and DiffRate represents the average difference rate of adjacent frames within the window. Judge whether the difference rate threshold needs to be adjusted according to the occlusion type: Hard occlusion: For the occlusion area of hard occlusion, the difference rate threshold will be automatically reduced to enhance the sensitivity to small feature jumps. Set the threshold under hard occlusion to ; where, is a sensitivity factor specific to hard occlusion. If the difference rate within the window exceeds the threshold , then mark this frame interval as compensated distortion. If , a distortion type label is generated, indicating that there may be compensation distortion in the features of the current frame interval.
[0041] Specifically, the specific process of combining the confidence score and the attack flag to eliminate low-confidence compensation data and attack frames, and output the complete facial feature vector and risk level through spatio-temporal coherence verification and cross-modal logic verification is as follows: conduct a logical correlation analysis on the low-confidence compensation data, attack flag, and temporal distortion label. If there are multiple types of abnormal signals in the same frame interval, it is marked as a compound anomaly; low-confidence compensation data refers to feature data with a confidence score lower than the preset confidence score threshold; dynamically allocate weights according to the anomaly type and the occluded area, and calculate the comprehensive risk score; if the risk score reaches the high-risk level, only output the feature vector of the unoccluded area and trigger a warning; if the risk score is at the medium-risk level, output the complete feature vector but restrict the permissions of downstream sensitive operations; if the risk score is at the low-risk level, directly output the complete feature vector and perform identity verification.
[0042] In this implementation plan, we combine the confidence score and the attack flag to conduct a logical analysis on the low-confidence compensation data, attack frames, and distortion labels, and then determine the risk level of the facial feature vector. The specific process is as follows: Elimination of low-confidence compensation data and attack flags. For low-confidence compensation data and attack frames, first judge according to the confidence score whether it is lower than the preset threshold . If it is lower than this threshold, it is marked as low-confidence compensation data and eliminated. Let the low-confidence data be , and its determination condition is: ; at the same time, if the frame contains an attack flag , the facial feature data of this frame will also be eliminated to prevent incorrect data from flowing into the system. Logical correlation analysis of temporal distortion labels. For the temporal distortion label , conduct a logical analysis and combine whether the features of the current frame conform to normal spatio-temporal coherence. If there are multiple types of abnormal signals (such as low-confidence compensation data and attack flags) within the same frame time interval, then judge that this interval is a compound anomaly. At this time, further processing is required: if a compound anomaly is found, generate a compound anomaly label and mark this frame as high risk. Dynamic weight allocation for anomaly type and occluded area. Based on the anomaly type (low confidence, attack flag, temporal distortion) and the area of the occluded area , dynamically adjust the weight of this feature and calculate the comprehensive risk score of this frame : ; among them, is the weight corresponding to the anomaly type, is the area of each anomaly type. Risk level determination and output; according to the calculated comprehensive risk score , a threshold is set to determine the risk level: High risk: If , the feature vector of the unoccluded area is output and a warning is issued. Medium risk: If , the complete feature vector is output, but the permissions for downstream sensitive operations are restricted. Low risk: If , the complete feature vector is output and identity verification is performed. Among them, and are the thresholds for high risk and medium risk respectively. Through the above process, the system can effectively judge various types of abnormal signals and perform appropriate feature data processing to ensure the accuracy and security of the output facial feature vector.
[0043] Please refer to Figure 3 , a face recognition method based on computer vision, including the following steps: S1. Analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable area and the occlusion area; S2. Dynamically activate the infrared and acoustic sensor data according to the occlusion type labels of the heat map, extract the temperature change patterns and vibration frequency correlations in the infrared thermal distribution time series data and the acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; S3. Calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated; S4. Receive multiple frames of predicted feature sequences and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark compensation distortion, and at the same time combine the confidence score and the attack mark to eliminate low-confidence compensation data and attack frames, and output a complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification.
[0044] In this implementation scheme: S1. In this step, by combining depth data and polarization light parameters, the material reflection characteristics in the image are analyzed to identify occluded areas. This method is more adaptable than traditional illumination- and color-based analyses and can accurately identify occluded areas in different environments (such as low-light or special light source conditions). The generation of the occlusion heatmap can effectively visualize the occlusion information in different regions of the image, providing clear guidance for subsequent analysis. Marking the recognizable areas and occluded areas can provide accurate input data for subsequent processing stages (such as occlusion compensation and cross-modal data fusion), which is one of the core innovations of this method. S2. Dynamically activating sensor data is an innovative feature that flexibly determines whether to activate infrared and acoustic sensors according to the occlusion type label. This strategy enables the system to select the most effective sensor data for processing in different occlusion scenarios, saving computing resources and improving recognition accuracy. The spatio-temporal convolutional layer can capture spatio-temporal information in dynamic changes when analyzing the temperature change patterns and vibration frequency correlations in temporal data, greatly improving the system's adaptability to dynamic occlusions. The use of the cross-modal attention mechanism is an innovation in this scheme. It can automatically allocate the contribution weights of infrared and acoustic data in multi-modal data, enabling different sensor data to complement each other and enhancing the accurate prediction of occluded areas in complex environments. This mechanism significantly improves the system's adaptability and recognition accuracy. S3. The calculation of cross-modal distribution similarity provides a comprehensive analysis of multi-modal information for recognition by comparing the predicted facial occlusion area feature map with the visible light features of the non-occluded area. In this way, the system can improve the recognition accuracy of real facial features through the correlation of different data sources. The synchronization check of the temperature changes in the acoustic sensor and the infrared image is innovative. Especially when facing spoofing attacks, potential spoofed images can be detected in a timely manner through spatio-temporal deviation analysis, ensuring the security of the system. This mechanism provides strong protection against spoofing attacks and improves the reliability of the system. The dynamic detection is added to the verification process to ensure real-time response under sound-triggered events, which is particularly important for live detection and dynamic face recognition.
[0045] S4. In this step, by analyzing the difference rate of the occlusion area features of adjacent frames through a sliding window, the processing methods for different occlusion types can be dynamically adjusted, and the compensation distortion can be marked and corrected in real time to ensure the consistency of facial features in each frame. This method can effectively correct short-term dynamic errors and improve recognition accuracy. The combination of confidence score and attack marking further enhances the security and recognition accuracy of the system by actively eliminating low-confidence data and attack frames. This design reduces misidentifications caused by data noise or attacks. The combination of the output complete facial feature vector and the risk level can adjust the system response according to different situations (such as low risk or high risk), improving the intelligence and application flexibility of the recognition system.
[0046] In summary, the present application has at least the following effects:
[0047] A face recognition system based on computer vision and its application method, by combining multi-modal information such as depth data, polarization light parameters, infrared and acoustic sensor data, can more accurately identify and compensate for occluded areas, improving the face recognition accuracy of the system in complex environments. By using spatio-temporal convolutional layers and cross-modal attention mechanisms, the contribution weights of infrared and acoustic features are dynamically adjusted to effectively handle the interference of occluded areas and spoofing attacks, thereby improving the ability to prevent attacks. Through sliding window analysis and dynamic time warping algorithms, the system can monitor and correct the facial features of each frame in real time, quickly respond to changes in occluded areas and environmental changes, and ensure the stability and continuity of the face recognition process. The system combines confidence scoring, attack marking, and risk assessment to achieve dynamic feedback based on different risk levels, can effectively manage different security levels and operation permissions, and improve the intelligence and security of the system. By using cross-modal distribution similarity calculation, the synergy between different sensor data is enhanced, making face recognition in complex environments more accurate and reliable, and avoiding the limitations of single-modal data.
[0048] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0049] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0050] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the function specified in the process(es) Figure 1 step(s) or block(s) Figure 1 or blocks specified in the step(s) or block(s).
[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in the process(es) Figure 1 step(s) or block(s) Figure 1 or blocks specified in the step(s) or block(s).
[0052] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0053] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A face recognition system based on computer vision, characterized in that: It includes the following modules: fine-grained occlusion classification module, multimodal feature fusion module, cross-modal adversarial identification module, and compensation verification module; The fine-grained occlusion classification module is used to analyze the depth data and polarized light parameters of the input image, classify the occlusion type according to the material reflection characteristics, generate an occlusion heat map, and mark the identifiable area and the occluded area; The multimodal feature fusion module is used to dynamically activate infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared heat distribution time series data and the acoustic spectrum time series data of the occlusion area through the spatiotemporal convolution layer, and dynamically allocate the contribution weights of infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; The cross-modal adversarial identification module is used to calculate the cross-modal distribution similarity between the predicted facial occlusion area feature map and the visible light feature of the unoccluded area to output a confidence score, and to verify the synchronization of the temperature change of the target area in the infrared image and the sound wave spectrum when the sound wave sensor detects a voice trigger event. If the spatiotemporal deviation exceeds a preset threshold, it is determined to be a counterfeit attack and an attack mark is generated; The compensation verification module is used to receive a multi-frame prediction feature sequence and analyze the difference rate of the occluded area features of adjacent frames through a sliding window to mark the compensation distortion, and at the same time, combine the confidence score and the attack mark to eliminate low-confidence compensation data and attack frames, and output a complete facial feature vector and risk level that pass the spatiotemporal coherence check and cross-modal logic verification.
2. The computer vision-based face recognition system according to claim 1, characterized in that: The specific process of analyzing the depth data and polarization parameters of the input image, classifying the occlusion type by the material reflection characteristics, and generating the occlusion heat map is as follows: The depth data and polarization light parameters of the input image are integrated, and the physical distance levels of the occluded area are divided based on the depth information. The hard occlusion and soft occlusion types are distinguished by combining the difference in polarization angle distribution and material reflectivity. A pixel-level occlusion heat map is generated and the material labels are annotated. The hard occlusion area is marked as a completely unrecognizable area, and the soft occlusion area is marked as a partially compensable area.
3. The computer vision-based face recognition system according to claim 1, characterized in that: The specific process of extracting the correlation between the infrared heat distribution time series data of the blocked area and the temperature change pattern and vibration frequency in the acoustic spectrum time series data through the spatiotemporal convolution layer is as follows: According to the multimodal spatiotemporal convolution kernel, it continuously slides along the time dimension and synchronously captures the temporal correlation between the dynamic temperature gradient changes of infrared heat distribution and the vibration frequency of the acoustic spectrum; Through cross-modal feature cross-validation, the interference of environmental noise is eliminated, and a coded feature vector that integrates the spatiotemporal correlation of infrared and sound waves is generated to characterize the synergy between physiological activities and physical movements in the occluded area.
4. The computer vision-based face recognition system according to claim 3, characterized in that: The specific process of dynamically allocating the contribution weights of infrared and acoustic features based on the cross-modal attention mechanism and generating the predicted facial occlusion area feature map is as follows: The material type label of the heat map is used as a conditional vector, concatenated with the multimodal fusion feature vector, and input into the attention network; According to the modal significance of the occlusion type label and the fused feature vector, the contribution weights of the infrared thermal distribution and the acoustic spectrum are dynamically allocated, and the temporal features of the infrared thermal distribution and the acoustic spectrum are weightedly fused at the channel level to generate the compensation feature vector of the occluded area. The weighted compensated feature vector is input into the generator network, and the facial details of the occluded area are restored through multi-level upsampling to generate a predicted facial occluded area feature map; A cross-modal consistency constraint is imposed to ensure the alignment of the predicted features with the visible light features of the unoccluded area in the semantic space.
5. The computer vision-based face recognition system according to claim 4, characterized in that: The specific process of calculating the cross-modal distribution similarity between the predicted facial occlusion area feature map and the visible light feature of the unoccluded area to output the confidence score is as follows: The generated predicted occluded area feature map and the visible light feature map of the unoccluded area are mapped to a shared embedding space, and the cosine similarity between the two is calculated through a contrastive learning algorithm. Combined with the confidence attenuation factor of the occluded area, a normalized confidence score is generated. The confidence score is used to characterize the degree of alignment between the generated features and the real facial features in the semantic space.
6. The computer vision-based face recognition system according to claim 5, characterized in that: When the acoustic wave sensor detects a voice trigger event, the synchronization between the temperature change of the target area in the infrared image and the acoustic wave spectrum is verified. If the temporal and spatial deviation exceeds the preset threshold, it is determined to be a forged attack and an attack mark is generated. The specific process is as follows: The sound wave spectrum peak and the infrared temperature timing curve are aligned according to the voice trigger timestamp, and the timing matching degree between the two is calculated through the dynamic time warping algorithm. If the matching degree is lower than the adaptive threshold, it is judged as a cross-modal forgery attack, an attack mark is generated and an alarm signal is triggered. At the same time, the feature fusion weight of the current frame is frozen to block the attack propagation.
7. The computer vision-based face recognition system according to claim 6, characterized in that: The specific process of receiving multi-frame prediction feature sequences and analyzing the difference rate of occluded area features of adjacent frames through a sliding window to mark compensation distortion is as follows: A fixed-length frame interval is intercepted from a continuous multi-frame prediction feature sequence, and the Euclidean distance difference rate of the occluded area features of adjacent frames is calculated frame by frame. The compensation distortion is determined by an adaptive threshold, which includes: The feature maps of the occluded area of each pair of adjacent frames in the window are compared pixel by pixel, and the average value of the Euclidean distance is calculated as the difference rate; If the occlusion type is hard occlusion, the difference rate threshold is automatically lowered to enhance the sensitivity to small feature jumps; If the difference rate exceeds the adaptive threshold of the current window, the frame interval is marked as compensation distortion and a distortion type label is generated.
8. The computer vision-based face recognition system according to claim 7, characterized in that: The specific process of combining confidence scores with attack marks to eliminate low-confidence compensation data and attack frames and outputting complete facial feature vectors and risk levels that have passed spatiotemporal coherence verification and cross-modal logic verification is as follows: Perform logical correlation analysis on low-confidence compensation data, attack markers, and timing distortion labels. If there are multiple types of abnormal signals in the same frame interval, they are marked as composite abnormalities. The low confidence compensation data represents feature data whose confidence score is lower than a preset confidence score threshold; Dynamically assign weights based on anomaly type and blocked area to calculate comprehensive risk scores; If the risk score reaches a high risk level, only the feature vector of the unobstructed area is output and a warning is triggered; If the risk score is at a medium risk level, the complete feature vector is output but the downstream sensitive operation permissions are restricted; If the risk score is a low risk level, the complete feature vector is directly output and identity verification is performed.
9. A computer vision-based face recognition method, applied to a computer vision-based face recognition system according to any one of claims 1 to 8, characterized in that: The following steps are involved: S1. Analyze the depth data and polarization parameters of the input image, classify the occlusion type by the material reflection characteristics, generate an occlusion heat map, and mark the identifiable area and the occluded area; S2. Dynamically activate infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared heat distribution time series data and acoustic spectrum time series data of the occluded area through the spatiotemporal convolution layer, and dynamically allocate the contribution weights of infrared and acoustic features based on the cross-modal attention mechanism to generate the predicted facial occlusion area feature map; S3. Calculate the cross-modal distribution similarity between the predicted facial occlusion area feature map and the visible light feature of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change of the target area in the infrared image and the sound wave spectrum when the sound wave sensor detects a voice trigger event. If the spatiotemporal deviation exceeds a preset threshold, it is determined to be a forged attack and an attack mark is generated; S4. Receive multi-frame predicted feature sequences and analyze the difference rate of occluded area features of adjacent frames through a sliding window to mark compensation distortion. At the same time, combine confidence scores and attack marks to eliminate low-confidence compensation data and attack frames, and output a complete facial feature vector and risk level that has passed spatiotemporal consistency verification and cross-modal logic verification.
Citation Information
Patent Citations
AI-based face recognition verification management method, system and equipment and storage medium
CN115810214A
Mask wearing detection method based on polarization imaging AI identification
CN115909457A
Face authentication including occlusion detection based on material data extracted from images
CN118805170A
Dual face counterfeiting active defense method based on frequency limitation
CN119274241A
Cited By
Bridge underwater disease detection system and method based on array imaging
CN120490291A
Urban low-altitude unmanned aerial vehicle target identification method and system
CN120544086A
Face recognition system based on face mask detection model
CN120853235A
Face recognition detection method and device, computer equipment and program product
CN121096006A
Adaptive vehicle sensor attack detection method based on multi-modal fusion
CN122204557A