Computer Vision-Based Face Recognition System and Its Application Method
Through fine-grained occlusion classification, multimodal feature fusion and cross-modal adversarial identification modules, the problem of reduced reliability of the recognition system and forged attacks caused by occlusion is solved, and high-precision and high-security face recognition is achieved.
Patent Information
- Application Number
- CN202510622356.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing computer vision-based facial recognition technology has reduced the reliability of the recognition system due to occlusion in high-density public places and special environments, and it is difficult to defend against multimodal forgery attacks, which poses a risk of privacy leakage.
The fine-grained occlusion classification module, multimodal feature fusion module, cross-modal adversarial identification module and compensation verification module are used to analyze the occlusion type through depth data and polarized light parameters, and the infrared and acoustic sensors are dynamically activated. The features are extracted in combination with the cross-modal attention mechanism and the spatiotemporal convolutional layer to generate the occlusion area feature map, and the forgery attack is checked when the acoustic sensor detects a voice trigger event.
It improves the robustness and anti-counterfeiting capabilities of the face recognition system in the case of occlusion, enhances the security and credibility of the recognition results, effectively defends against high-simulation spoofing attacks, and ensures privacy protection.
Smart Images

Figure CN120148092B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face recognition, and in particular to a face recognition system based on computer vision and its application method. Background Art
[0002] With the rapid development of fields such as smart cities, public security, and financial payment, face recognition technology, as the core means of biometric recognition, has been widely used in scenarios such as identity verification, security monitoring, and personalized services. However, in high-density public places and special environments, users often wear masks, hats, scarves and other occluders, resulting in incomplete face information, which seriously reduces the reliability of the recognition system. At the same time, the threats of face data abuse and forgery attacks are increasing day by day. How to achieve high-precision and high-security real-time recognition in complex occlusion scenarios has become a key challenge restricting the development of the industry.
[0003] Current face recognition technology based on computer vision cannot distinguish the physical materials of occluders, resulting in the same processing strategy for hard occlusion and soft occlusion, and the compensation accuracy is low. Single-modal data is difficult to penetrate the occlusion area, and relying on manually designed features is vulnerable to interference from light changes and pose offsets. Existing multi-modal solutions often use fixed-weight fusion and cannot dynamically allocate sensor resources according to the occlusion type, resulting in wasted computing power and excessive load on edge devices, and it is difficult to capture the dynamic physiological characteristics of the occlusion area. Existing defense means rely on a single modality and are easily bypassed by multi-modal forgery attacks. Centralized training requires uploading original face data, which poses a risk of privacy leakage. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a face recognition system based on computer vision and its application method, which solves the problems in the above background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A face recognition system based on computer vision, including the following modules: a fine-grained occlusion classification module, a multi-modal feature fusion module, a cross-modal adversarial discrimination module, and a compensation verification module; The fine-grained occlusion classification module is used to analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable area and the occlusion area; The multi-modal feature fusion module is used to dynamically activate the infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; The cross-modal adversarial discrimination module is used to calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change and the acoustic spectrum in the target area of the infrared image when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated; The compensation verification module is used to receive a multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensation distortion. At the same time, combined with the confidence score and the attack mark, the low-confidence compensation data and attack frames are eliminated, and a complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification are output.
[0006] Further, the specific process of analyzing the depth data and polarization light parameters of the input image, classifying the occlusion types through the material reflection characteristics, and generating an occlusion heat map is as follows: Fuse the depth data and polarization light parameters of the input image, divide the physical distance levels of the occlusion area based on the depth information, combine the polarization angle distribution difference and the material reflectivity to distinguish between hard occlusion and soft occlusion types, generate a pixel-level occlusion heat map and label the material label, where the hard occlusion area is marked as a completely unrecognizable area, and the soft occlusion area is marked as a partially compensable area.
[0007] Further, the specific process of extracting the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer is as follows: According to the multi-modal spatio-temporal convolutional kernel, continuously slide along the time dimension and synchronously capture the dynamic temperature gradient change of the infrared thermal distribution and the vibration frequency time series correlation of the acoustic spectrum; Eliminate the environmental noise interference through cross-modal feature cross-verification, and generate an encoded feature vector that fuses the spatio-temporal correlation of infrared and acoustic waves, which is used to characterize the synergy of physiological activities and physical movements in the occlusion area.
[0008] Further, the specific process of dynamically allocating the contribution weights of infrared and acoustic wave features based on the cross-modal attention mechanism to generate the predicted facial occlusion region feature map is as follows: Using the material type label of the heat map as a conditional vector, concatenating it with the multi-modal fusion feature vector, and inputting it into the attention network; Dynamically allocating the contribution weights of the infrared heat distribution and the acoustic wave spectrum according to the occlusion type label and the modal saliency of the fusion feature vector, performing channel-level weighted fusion on the infrared heat distribution temporal features and the acoustic wave spectrum temporal features to generate a compensated feature vector for the occlusion region; Inputting the weighted compensated feature vector into the generator network, restoring the facial details of the occlusion region through multi-level upsampling to generate the predicted facial occlusion region feature map; And imposing cross-modal consistency constraints to ensure the alignment of the predicted features with the visible light features of the unoccluded region in the semantic space.
[0009] Further, the specific process of calculating the cross-modal distribution similarity between the generated predicted facial occlusion region feature map and the visible light features of the unoccluded region to output a confidence score is as follows: Mapping the generated predicted occlusion region feature map and the visible light features of the unoccluded region to a shared embedding space, calculating the cosine similarity between the two through a contrastive learning algorithm, and combining the confidence decay factor of the occlusion region to generate a normalized confidence score, where the confidence score is used to characterize the alignment degree between the generated features and the real facial features in the semantic space.
[0010] Further, when the acoustic wave sensor detects a voice trigger event, the synchronization between the temperature change of the target region in the infrared image and the acoustic wave spectrum is verified. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated. The specific process is as follows: Aligning the peak of the acoustic wave spectrum and the infrared temperature temporal curve according to the voice trigger timestamp, calculating the temporal matching degree between the two through the dynamic time warping algorithm. If the matching degree is lower than the adaptive threshold, it is determined as a cross-modal forgery attack, generating an attack mark and triggering an alarm signal, and at the same time freezing the feature fusion weight of the current frame to block the attack propagation.
[0011] Further, the specific process of receiving a multi-frame prediction feature sequence and analyzing the difference rate of the occlusion region features of adjacent frames through a sliding window to mark compensation distortion is as follows: Intercepting a frame interval of a fixed length in the continuous multi-frame prediction feature sequence, calculating the Euclidean distance difference rate of the occlusion region features of adjacent frames frame by frame, and determining the compensation distortion through an adaptive threshold, which specifically includes: Performing pixel-by-pixel comparison on the occlusion region feature maps of each pair of adjacent frames within the window, calculating the average value of the Euclidean distance as the difference rate; If the occlusion type is hard occlusion, the difference rate threshold is automatically reduced to enhance the sensitivity to tiny feature jumps; If the difference rate exceeds the adaptive threshold of the current window, mark this frame interval as compensation distortion and generate a distortion type label.
[0012] Furthermore, the specific process of combining the confidence score with the attack flag to eliminate low-confidence compensation data and attack frames and output the complete facial feature vector and risk level through spatio-temporal coherence verification and cross-modal logic verification is as follows: perform logical correlation analysis on the low-confidence compensation data, attack flag, and temporal distortion label. If multiple types of abnormal signals exist within the same frame interval, it is marked as a composite anomaly; the low-confidence compensation data refers to the feature data with a confidence score lower than the preset confidence score threshold; dynamically allocate weights according to the type of anomaly and the area of the occluded region, and calculate the comprehensive risk score; if the risk score reaches the high-risk level, only output the feature vector of the unoccluded region and trigger a warning; if the risk score is at the medium-risk level, output the complete feature vector but restrict the downstream sensitive operation permissions; if the risk score is at the low-risk level, directly output the complete feature vector and perform identity verification.
[0013] A computer vision-based face recognition method includes the following steps: S1. Analyze the depth data and polarization light parameters of the input image, classify the occlusion type through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable region and the occluded region; S2. Dynamically activate the infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and acoustic spectrum time series data of the occluded region through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occluded region feature map; S3. Calculate the cross-modal distribution similarity between the generated predicted facial occluded region feature map and the visible light features of the unoccluded region to output the confidence score, and verify the synchronization of the temperature change in the target region of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a spoofing attack and an attack flag is generated; S4. Receive multiple frames of predicted feature sequences and analyze the difference rate of the occluded region features of adjacent frames through a sliding window to mark compensation distortion. At the same time, combine the confidence score with the attack flag to eliminate low-confidence compensation data and attack frames, and output the complete facial feature vector and risk level through spatio-temporal coherence verification and cross-modal logic verification.
[0014] The present invention has the following beneficial effects:
[0015] (1) The face recognition system based on computer vision significantly improves the robustness of the face recognition system under occlusion through the collaborative action of the fine-grained occlusion classification module and the multi-modal feature fusion module. The fine-grained occlusion classification module utilizes depth data and polarization light parameters to achieve precise classification of occlusions of different materials and generate a pixel-level occlusion heat map, thereby accurately calibrating the recognizable area and the occlusion area and improving the detection accuracy of the occlusion area. The multi-modal feature fusion module dynamically activates the infrared and acoustic sensors based on the occlusion type label, extracts the temporal features in the infrared heat distribution and the acoustic spectrum through the spatio-temporal convolutional layer, and adaptively adjusts the feature weights in combination with the cross-modal attention mechanism, enabling the system to effectively compensate for the information loss in the occlusion area, thereby improving the integrity and stability of face feature extraction.
[0016] (2) The face recognition method based on computer vision enhances the anti-counterfeiting ability and data credibility of the face recognition system through the design of the cross-modal adversarial discrimination module and the compensation verification module. The cross-modal adversarial discrimination module combines visible light, infrared, and acoustic data to calculate the distribution similarity of cross-modal features, and when the acoustic sensor detects a voice trigger event, it verifies whether there is a forgery attack through the synchrony of the temperature change of the infrared image and the acoustic spectrum, thereby effectively improving the system's detection ability for high-fidelity spoofing attacks. The compensation verification module analyzes the change rate of the occlusion area features of adjacent frames through a sliding window to identify compensation distortion, and combines the confidence score and the attack flag to eliminate low-confidence or deceptive data, and finally outputs a complete facial feature vector and risk level after spatio-temporal consistency verification and cross-modal logic verification, improving the security and credibility of the recognition result.
[0017] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of the face recognition system based on computer vision of the present invention.
[0019] Figure 2 It is a detailed flowchart of the compensation verification module of the present invention
[0020] Figure 3 It is a flowchart of the face recognition method based on computer vision of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0021] The embodiments of the present application solve the problems of decreased recognition accuracy and insufficient security caused by occlusion, environmental interference, and forgery attacks in the face recognition process through the face recognition system based on computer vision and its application method.
[0022] The general idea of the solution in the embodiments of the present application is as follows:
[0023] Analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable areas and occlusion areas.
[0024] Dynamically activate the infrared and acoustic sensor data according to the occlusion type labels of the heat map, extract the temporal data of the infrared thermal distribution and the correlation of the vibration frequency in the temporal data of the acoustic spectrum in the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map.
[0025] Calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a spoofing attack and an attack mark is generated.
[0026] Receive a multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark compensation for distortion. At the same time, combine the confidence score and the attack mark to eliminate low-confidence compensation data and attack frames, and output a complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification.
[0027] Please refer to Figure 1, an embodiment of the present invention provides a technical solution: a face recognition system based on computer vision, including the following modules: a fine-grained occlusion classification module, a multi-modal feature fusion module, a cross-modal adversarial discrimination module, and a compensation verification module; the fine-grained occlusion classification module is used to analyze the depth data and polarization light parameters of the input image, classify the occlusion type through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable area and the occlusion area; the multi-modal feature fusion module is used to dynamically activate the infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and the acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; the cross-modal adversarial discrimination module is used to calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds a preset threshold, it is determined as a forgery attack and an attack mark is generated; the compensation verification module is used to receive a multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensation distortion. At the same time, combined with the confidence score and the attack mark, the low-confidence compensation data and the attack frames are eliminated, and a complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification are output.
[0028] In this implementation, a highly robust face recognition system for complex occlusion scenarios of the present invention has its core in solving the key problems of occlusion recognition, feature compensation, attack defense, and dynamic stability through multi-module collaboration. The working principles and technical advantages are described in the following sub-modules. The system includes four core modules, forming a closed-loop processing chain of "occlusion classification, multi-modal fusion, attack defense, and compensation verification": Input: visible light images (RGB), depth maps (TOF), polarization data, infrared thermal images, acoustic spectra; Processing: Gradually transfer intermediate results such as occlusion heat maps, predicted feature maps, confidence scores, etc.; Output: complete facial feature vectors (for identity verification) and risk level labels (for permission control). Fine-grained occlusion classification module: Distinguish occlusion types (hard / soft occlusion) and generate heat maps. Technical implementation: Input depth maps (physical distance) and polarization parameters (material reflection characteristics), and determine the type of occluder (such as a mask = hard occlusion, a scarf = soft occlusion) through a material classification network (such as a ResNet variant); Output pixel-level heat maps, marking recognizable areas (such as the forehead, eyebrow bones) and occlusion areas (such as the lower half of the face covered by a mask). Break through the limitations of traditional RGB texture analysis, and combine depth and polarization to achieve accurate material discrimination (such as distinguishing a plastic mask from a cloth scarf). Multi-modal feature fusion module: Dynamically select sensor data and generate compensated features for the occlusion area. Technical implementation: Dynamically activate sensors according to the occlusion labels of the heat map: hard occlusion (such as a mask) to focus on infrared thermal distribution (detecting the temperature rise of the tip of the nose during breathing); soft occlusion (such as a scarf) to focus on acoustic spectra (analyzing the correlation of voice vibrations); Extract the temporal correlation of temperature changes and vibration frequencies through spatio-temporal convolution (such as the synchronization of breathing frequency and acoustic spectra); The cross-modal attention mechanism dynamically allocates sensor weights to generate a predicted feature map for the occlusion area (such as the contour of the chin under the mask). Avoid wasting resources with fixed fusion rules and improve the efficiency of edge computing; Capture the physical correlation between physiological activities (breathing, speech) and occlusion compensation. Cross-modal adversarial discrimination module function: Verify the credibility of the generated features and defend against forgery attacks. Technical implementation: Cross-modal similarity calculation: Map the generated features and the unoccluded area of visible light to a shared semantic space and calculate the cosine similarity (such as the anatomical continuity between the generated chin under the mask and the real forehead structure); Attack defense: Verify the acoustic-infrared synchronization when the voice is triggered (such as the temperature rise of the lips during pronunciation needs to match the peak of the acoustic wave within 0.5 seconds) to determine asynchronous attacks (such as recording playback + heated mask). Semantic alignment constraint: Solve the problem of comparability of heterogeneous data through a shared embedding space; Multi-modal attack interception: Combine acoustic and infrared data to defend against highly realistic attacks. Compensation verification module function: Eliminate abnormal data and output complete features and risk levels.Technical implementation: Temporal coherence verification: Analyze the difference rate of multi-frame features through a sliding window (e.g., the jumping of the chin contour caused by head turning), and mark and compensate for distortion; Risk grading: Integrate low-confidence data, attack marks, and temporal distortion to output low / medium / high risk levels (e.g., high risk prohibits payment, low risk conducts normal verification). Dynamic threshold adjustment: Adaptively set the difference rate threshold according to the occlusion type and environmental noise; Hierarchical permission control: Balance security and user experience. Through the full-link innovation of physical-level occlusion classification, dynamic multi-modal fusion, cross-modal attack defense, and temporal risk control, the present invention solves the accuracy, security, and real-time performance of face recognition in complex occlusion scenarios, and at the same time meets privacy compliance requirements through lightweight design and federated learning.
[0029] Specifically, the specific process of analyzing the depth data and polarization light parameters of the input image and classifying the occlusion type through the material reflection characteristics to generate an occlusion heat map is as follows: Integrate the depth data and polarization light parameters of the input image, divide the physical distance levels of the occlusion area based on the depth information, combine the polarization angle distribution difference and the material reflectivity to distinguish between hard occlusion and soft occlusion types, and generate a pixel-level occlusion heat map and label the material tags, where the hard occlusion area is marked as a completely unrecognizable area, and the soft occlusion area is marked as a partially compensable area.
[0030] In this implementation, integrate depth data and polarization light parameters. Depth data: Use an RGB-D camera or other depth sensors to obtain the three-dimensional information of the scene, and the depth value of each pixel represents its physical distance from the camera. Polarization light parameters: Obtain the light intensity distribution at different polarization angles through a polarization camera, calculate the polarization angle distribution difference and the material reflectivity, and use them to judge the scattering and transmission of light on the surface. Divide the physical distance levels of the occlusion area: According to the depth data, stratify the objects in the scene by distance, for example: close range (0 - 50 cm), medium range (50 - 100 cm), long range (more than 100 cm). Occlusions at close range have a greater impact on the face, while long-range occlusions may interfere less with recognition. Distinguish between hard occlusion and soft occlusion: Hard occlusion (non-compensable): If the depth data of an area changes suddenly and the polarization angle changes violently, it indicates that this area may be a solid object (such as a hand, a mask, a wall, etc.), and it is marked as a completely unrecognizable area. Soft occlusion (partially compensable): If the polarization light parameters show a large transmittance, it indicates that this area may be a semi-transparent material (such as glass, gauze, fog, etc.), and it is marked as a partially compensable area. Generate a pixel-level occlusion heat map, combine the depth data and polarization light information, calculate the occlusion type pixel by pixel, and generate an occlusion heat map. The hard occlusion area is marked as high risk (red), indicating that this area cannot be recognized; the soft occlusion area is marked as low risk (yellow), indicating that it can be compensated by combining other sensors.
[0031] Specifically, the specific process of extracting the correlation between the temperature change pattern and the vibration frequency in the infrared thermal distribution time series data and the acoustic wave spectrum time series data of the occluded area through the spatio-temporal convolutional layer is as follows: According to the multi-modal spatio-temporal convolution kernel, continuously slide along the time dimension and synchronously capture the dynamic temperature gradient change of the infrared thermal distribution and the vibration frequency time series correlation of the acoustic wave spectrum; Eliminate environmental noise interference through cross-modal feature cross-validation, and generate an encoded feature vector that fuses the spatio-temporal correlation of infrared and acoustic waves, which is used to characterize the synergy between the physiological activities and physical movements in the occluded area.
[0032] In this implementation scheme, in order to extract the correlation between the temperature change pattern and the vibration frequency from the occluded area, a spatio-temporal convolutional layer is used to jointly process the infrared thermal distribution data and the acoustic wave spectrum data to capture the dynamic features that change over time. The specific process is as follows: Based on spatio-temporal convolution for dynamic feature extraction, assume that the input infrared thermal distribution in the time series data is , where t represents the time frame, represents the spatial coordinates. Assume that the input acoustic wave spectrum data is , where f represents the frequency component and t represents the time frame. Define a three-dimensional spatio-temporal convolution kernel , where is the time window size, and m, n are the spatial dimensions or spectral dimensions. Adopt a sliding window strategy on the time axis, with a step size of to calculate the local temperature gradient for the infrared thermal data : ; where, represents the local temperature gradient, which measures the dynamic change of the infrared thermal distribution. Extract the time correlation features of the acoustic wave vibration frequency: Perform a short-time Fourier transform (STFT) on the input acoustic wave spectrum to obtain the frequency-time distribution matrix: ; Use the same spatio-temporal convolution kernel to perform spatio-temporal convolution on the acoustic wave spectrum: ; where, represents the vibration frequency change pattern on the time axis. Cross-modal feature cross-validation to eliminate environmental noise: Calculate the correlation between the infrared thermal data and the acoustic wave spectrum data to eliminate environmental noise: ; where, is the mean value, is the standard deviation. Only retain the data with a correlation higher than the threshold , and finally obtain the fused spatio-temporal correlation encoded feature vector: ; This feature vector is used to describe the synergy between the physiological activities (temperature changes) and physical movements (vibration frequencies) in the occluded area, and improve the accuracy of recognition.
[0033] Specifically, based on the cross-modal attention mechanism, the contribution weights of infrared and acoustic wave features are dynamically allocated, and the specific process of generating the predicted facial occlusion area feature map is as follows: Using the material type label of the heat map as the conditional vector, concatenating it with the multi-modal fusion feature vector, and inputting it into the attention network; According to the occlusion type label and the modal saliency of the fusion feature vector, the contribution weights of the infrared heat distribution and the acoustic wave spectrum are dynamically allocated, and the channel-level weighted fusion of the infrared heat distribution temporal feature and the acoustic wave spectrum temporal feature is performed to generate the compensated feature vector of the occlusion area; Inputting the weighted compensated feature vector into the generator network, restoring the facial details of the occlusion area through multi-level upsampling, and generating the predicted facial occlusion area feature map; And imposing cross-modal consistency constraints to ensure the alignment of the predicted features with the visible light features of the unoccluded area in the semantic space.
[0034] In this implementation, in order to reasonably utilize the information of different modalities, a cross-modal attention mechanism is adopted to dynamically adjust the feature weights of the infrared heat distribution and the acoustic wave spectrum. The specific process is as follows: Construct the input features, the infrared heat distribution feature vector is , the acoustic wave spectrum feature vector is , and the occlusion type label is . Form the concatenated vector: ; where serves as the conditional vector, providing the prior information of the occlusion material type. Calculate the cross-modal attention weights, and calculate the contribution weights of infrared and acoustic waves through a trainable attention network : ; where the attention mechanism adopts softmax normalization: ; where and are the weight parameters learned by the neural network. Fuse the feature vectors, and perform channel-level weighted fusion using the calculated weights: ; This fused feature vector retains the multi-modal information of the occlusion area and enhances the features helpful for face reconstruction. Input the generator network to restore the facial details of the occlusion area, and use the generator network in the generative adversarial network (GAN) to complete the facial feature complementation: ; where is the predicted facial occlusion area feature map. Impose cross-modal consistency constraints. To ensure the alignment of the generated occlusion area features with the visible light features of the unoccluded area in the semantic space, a cross-modal consistency loss is introduced. Finally, the facial occlusion area feature map generated by this process can more accurately restore the facial details, maintain the feature consistency with the unoccluded area, and improve the reliability of the system for face recognition in the case of occlusion.
[0035] Specifically, the specific process of calculating and generating the cross-modal distribution similarity between the predicted facial occlusion region feature map and the visible light features of the unoccluded region to output the confidence score is as follows: Map the generated predicted occlusion region feature map and the visible light features of the unoccluded region to a shared embedding space, calculate the cosine similarity between the two through a contrastive learning algorithm, and combine the confidence decay factor of the occlusion region to generate a normalized confidence score, which is used to characterize the alignment degree between the generated features and the real facial features in the semantic space.
[0036] In this implementation, by calculating the cross-modal distribution similarity between the generated facial occlusion region feature map and the visible light features of the unoccluded region, the correspondence between the generated features and the real facial features is evaluated. The specific process is as follows: Map to the shared embedding space. To compare the generated facial occlusion region feature map and the visible light features of the unoccluded region, we first map them to a shared embedding space to ensure that features of different modalities can be compared in the same semantic space. The generated facial occlusion region feature map is , and the visible light features of the occlusion region are . Use a mapping network to transform these two feature maps into the shared embedding space: , ; where is the mapping function of the shared embedding space. Calculate the cross-modal similarity: To measure the similarity between the two features, the cosine similarity is used to calculate the correlation between the two embedding space vectors and . The cosine similarity formula is: ; where represents the dot product of the vectors, and represents the norm (modulus length) of the vector. The range of this similarity value is [-1, 1], and the larger the value, the higher the similarity between the two features. Introduction of the confidence decay factor. On the basis of calculating the cross-modal similarity, the confidence decay factor of the occlusion region is combined to adjust the final confidence score. Set the confidence decay factor as , which is used to adjust the calculation result according to the difficulty or uncertainty of the occlusion region: ; where ranges from [0, 1] and is used to simulate the influence of uncertainty. Normalize the confidence score. The final confidence score will be standardized through a normalization operation for easy comparison between different occlusion regions. Set the maximum and minimum confidence scores as and respectively, and the normalization process is: ; The finally obtained normalized confidence score ConfidenceScore^ ranges from 0 to 1, indicating the alignment degree between the generated features and the real facial features in the semantic space.
[0037] Specifically, when the acoustic wave sensor detects a voice trigger event, the synchronization between the temperature change in the target area of the infrared image and the acoustic wave spectrum is verified. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and the specific process of generating an attack mark is as follows: Align the peak of the acoustic wave spectrum with the time series curve of the infrared temperature according to the voice trigger timestamp, calculate the temporal matching degree of the two through the dynamic time warping algorithm. If the matching degree is lower than the adaptive threshold, it is determined as a cross-modal forgery attack, generate an attack mark and trigger an alarm signal, and at the same time freeze the feature fusion weight of the current frame to block the spread of the attack.
[0038] In this implementation plan, in this experimental plan, in order to detect cross-modal forgery attacks, we use the temporal synchronization between the acoustic wave spectrum and the infrared image for verification. If the temporal deviation exceeds the preset threshold, it is determined as a forgery attack and the corresponding alarm signal is sent. The specific process is as follows: Align the data according to the voice trigger timestamp. When the acoustic wave sensor detects a voice trigger event, we first obtain the trigger timestamp , and align the temporal data of the temperature change in the infrared image according to this timestamp. Let the acoustic wave spectrum data be , and the temperature time series of the infrared image be , where is time, is the frequency component of the acoustic wave, is the spatial coordinate of the infrared image. Synchronize the infrared image and the acoustic wave spectrum according to the timestamp to obtain the aligned temperature data and the acoustic wave spectrum data . Calculate the spatio-temporal matching degree: Dynamic Time Warping (DTW) In order to accurately measure the synchronization between the temperature time series of the infrared image and the acoustic wave spectrum time series, the Dynamic Time Warping (DTW) algorithm is used for temporal alignment and similarity calculation. The purpose of the DTW algorithm is to find the optimal matching path between two time series data, so as to evaluate their temporal consistency. Set the distance metric function of the dynamic time warping algorithm, and calculate the optimal matching path on the time axis: ; Obtain the matching degree through dynamic programming, and measure the temporal synchronization of the two. Compare with the preset value to determine the forgery attack. If the temporal matching degree is lower than the adaptive threshold , it means that the synchronization between the temperature change of the infrared image and the acoustic wave spectrum is insufficient, and there may be a forgery attack. Set the forgery attack mark as , and the judgment condition is: Trigger an alarm signal and freeze the feature fusion weights of the current frame to block the spread of the attack.
[0039] Please refer to Figure 2 , specifically, the specific process of receiving a multi-frame prediction feature sequence and analyzing the difference rate of the occlusion area features of adjacent frames through a sliding window to mark compensation distortion is as follows: intercept a frame interval of a fixed length from the continuous multi-frame prediction feature sequence, calculate the Euclidean distance difference rate of the occlusion area features of adjacent frames frame by frame, and determine compensation distortion through an adaptive threshold, specifically including: performing pixel-by-pixel comparison on the occlusion area feature maps of each pair of adjacent frames within the window, and calculating the average value of the Euclidean distance as the difference rate; if the occlusion type is hard occlusion, the difference rate threshold is automatically reduced to enhance the sensitivity to small feature jumps; if the difference rate exceeds the adaptive threshold of the current window, mark this frame interval as compensation distortion and generate a distortion type label.
[0040] In this implementation scheme, this implementation scheme aims to analyze the prediction feature sequence of continuous multi-frames through a sliding window, calculate the difference rate of the occlusion area features of adjacent frames, and judge whether there is compensation distortion according to the difference rate. The specific process is as follows: intercept a frame interval of a fixed length Intercept a frame interval of a fixed length from the continuous multi-frame prediction feature sequence, and set the window length to , and this window contains frame data, and calculate the difference rate between adjacent frames frame by frame. Calculate the Euclidean distance difference rate For the occlusion area feature maps of each pair of adjacent frames within the window and , perform pixel-by-pixel comparison, calculate the Euclidean distance difference between them, and the formula is: ; where and respectively represent the occlusion area feature values at pixel (i, j) in frame t and frame t+1. After calculating the Euclidean distance of each pair of adjacent frames, the difference rate between each pair of adjacent frames is obtained. Calculate the average value of the difference rate After calculating the difference rate of all adjacent frames within the window, the average difference rate DiffRate is obtained: where is the size of the window, and DiffRate represents the average difference rate of adjacent frames within the window. Determine whether the difference rate threshold needs to be adjusted according to the occlusion type: Hard occlusion: For the occlusion area of hard occlusion, the difference rate threshold will be automatically reduced to enhance the sensitivity to small feature jumps. Set the threshold under hard occlusion to ; where is a sensitivity factor specific to hard occlusion. If the difference rate within the window exceeds the threshold , then mark this frame interval as compensation distortion. If , a distortion type label is generated, indicating that there may be compensation distortion in the features of the current frame interval.
[0041] Specifically, the specific process of combining the confidence score and the attack flag to eliminate low-confidence compensation data and attack frames, and output the complete facial feature vector and risk level through spatio-temporal coherence verification and cross-modal logic verification is as follows: conduct a logical correlation analysis on the low-confidence compensation data, attack flag, and temporal distortion label. If there are multiple types of abnormal signals in the same frame interval, it is marked as a composite anomaly; low-confidence compensation data refers to feature data with a confidence score lower than the preset confidence score threshold; dynamically allocate weights according to the anomaly type and the occluded area area, and calculate the comprehensive risk score; if the risk score reaches the high-risk level, only output the feature vector of the unoccluded area and trigger a warning; if the risk score is at the medium-risk level, output the complete feature vector but restrict the downstream sensitive operation permissions; if the risk score is at the low-risk level, directly output the complete feature vector and perform identity verification.
[0042] In this implementation plan, we combine the confidence score and the attack flag to conduct a logical analysis on the low-confidence compensation data, attack frames, and distortion labels, and then determine the risk level of the facial feature vector. The specific process is as follows: Elimination of low-confidence compensation data and attack flags. For low-confidence compensation data and attack frames, first judge according to the confidence score whether it is lower than the preset threshold . If it is lower than this threshold, it is marked as low-confidence compensation data and eliminated. Let the low-confidence data be , and its judgment condition is: ; at the same time, if the frame contains an attack flag , the facial feature data of this frame will also be eliminated to prevent incorrect data from flowing into the system. Logical correlation analysis of temporal distortion labels. For the temporal distortion label , conduct a logical analysis, combined with whether the features of the current frame conform to normal spatio-temporal coherence. If there are multiple types of abnormal signals (such as low-confidence compensation data and attack flags) within the same frame time interval, then judge that this interval is a composite anomaly. At this time, further processing is required: if a composite anomaly is found, generate a composite anomaly label and mark this frame as high risk. Dynamic weight allocation of anomaly type and occluded area area. Based on the anomaly type (low confidence, attack flag, temporal distortion) and the area of the occluded area , dynamically adjust the weight of this feature and calculate the comprehensive risk score of this frame : ; among them, is the weight corresponding to the anomaly type, is the area of each anomaly type. Risk level determination and output; according to the calculated comprehensive risk score , set thresholds to determine the risk level: High risk: If , then output the feature vector of the unoccluded area and issue a warning. Medium risk: If , then output the complete feature vector, but restrict the permissions of downstream sensitive operations. Low risk: If , then output the complete feature vector and perform identity verification. Among them, and are the thresholds for high risk and medium risk respectively. Through the above process, the system can effectively judge various types of abnormal signals and perform appropriate feature data processing to ensure the accuracy and security of the output facial feature vector.
[0043] Please refer to Figure 3 , a computer vision-based face recognition method, including the following steps: S1. Analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable area and the occlusion area; S2. Dynamically activate the infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; S3. Calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a spoofing attack and an attack mark is generated; S4. Receive multiple frames of predicted feature sequences and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensation distortion. At the same time, combine the confidence score and the attack mark to eliminate the low-confidence compensation data and attack frames, and output the complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification.
[0044] In this implementation scheme: S1. In this step, by combining depth data and polarization light parameters, the material reflection characteristics in the image are analyzed to identify occluded areas. This method is more adaptable than traditional illumination- and color-based analyses and can accurately identify occluded areas in different environments (such as low-light or special light source conditions). The generation of an occlusion heatmap can effectively visualize the occlusion information in different regions of the image, providing clear guidance for subsequent analysis. Marking the recognizable areas and occluded areas can provide accurate input data for subsequent processing stages (such as occlusion compensation and cross-modal data fusion), which is one of the core innovations of this method. S2. Dynamically activating sensor data is an innovative feature that flexibly determines whether to activate infrared and acoustic sensors according to the occlusion type label. This strategy enables the system to select the most effective sensor data for processing in different occlusion scenarios, saving computing resources and improving recognition accuracy. When analyzing the temperature change pattern and vibration frequency correlation in time-series data, the spatio-temporal convolutional layer can capture spatio-temporal information in dynamic changes, greatly improving the system's adaptability to dynamic occlusion. The use of a cross-modal attention mechanism is an innovation in this scheme. It can automatically allocate the contribution weights of infrared and acoustic data in multi-modal data, enabling different sensor data to complement each other and enhancing the accurate prediction of occluded areas in complex environments. This mechanism significantly improves the system's adaptability and recognition accuracy. S3. Calculating the cross-modal distribution similarity by comparing the feature map of the predicted facial occluded area with the visible light features of the non-occluded area provides a comprehensive analysis of multi-modal information for recognition. In this way, the system can improve the recognition accuracy of real facial features through the correlation of different data sources. Verifying the synchronization of the acoustic sensor and the temperature change in the infrared image is innovative. Especially when facing spoofing attacks, potential spoofed images can be detected in a timely manner through spatio-temporal deviation analysis, ensuring the security of the system. This mechanism provides strong protection against spoofing attacks and improves the reliability of the system. The dynamic detection is added to the verification process to ensure real-time response under sound-triggered events, which is particularly important for live detection and dynamic face recognition.
[0045] S4. In this step, by analyzing the difference rate of the occluded area features in adjacent frames through a sliding window, the processing methods for different occlusion types can be dynamically adjusted, and the compensation distortion can be marked and corrected in real time to ensure the consistency of facial features in each frame. This method can effectively correct short-term dynamic errors and improve recognition accuracy. Combining the confidence score with the attack mark further enhances the security and recognition accuracy of the system by actively eliminating low-confidence data and attack frames. This design reduces misidentifications caused by data noise or attacks. Combining the output complete facial feature vector with the risk level can adjust the system response according to different situations (such as low risk or high risk), improving the intelligence and application flexibility of the recognition system.
[0046] In summary, the present application has at least the following effects:
[0047] A face recognition system based on computer vision and its application method can more accurately identify and compensate for occluded areas by combining multi-modal information such as depth data, polarization light parameters, infrared and acoustic sensor data, improving the face recognition accuracy of the system in complex environments. By using spatio-temporal convolutional layers and cross-modal attention mechanisms, the contribution weights of infrared and acoustic features are dynamically adjusted to effectively handle the interference of occluded areas and spoofing attacks, thereby improving the ability to prevent attacks. Through sliding window analysis and dynamic time warping algorithms, the system can monitor and correct the facial features of each frame in real time, quickly respond to changes in occluded areas and environmental changes, ensuring the stability and continuity of the face recognition process. The system combines confidence scoring, attack marking, and risk assessment to achieve dynamic feedback based on different risk levels, effectively managing different security levels and operation permissions, and enhancing the intelligence and security of the system. By using cross-modal distribution similarity calculation, the synergy between different sensor data is enhanced, making face recognition in complex environments more accurate and reliable, and avoiding the limitations of single-modal data.
[0048] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0049] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0050] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in the flowchart(s) Figure 1 a flowchart or flowcharts and / or block(s) Figure 1 a block or blocks.
[0051] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in the flowchart(s) Figure 1 a flowchart or flowcharts and / or block(s) Figure 1 a block or blocks.
[0052] Although the preferred embodiments of the present invention have been described, additional changes and modifications to these embodiments can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0053] It is obvious that those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A computer vision-based face recognition system, characterized in that, It includes the following modules: fine-grained occlusion classification module, multi-modal feature fusion module, cross-modal adversarial discrimination module, and compensation verification module; The fine-grained occlusion classification module is used to analyze the depth data and polarized light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable area and the occlusion area; The multi-modal feature fusion module is used to dynamically activate the infrared and acoustic sensor data according to the occlusion type label of the heat map, extract the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and the acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; The cross-modal adversarial discrimination module is used to calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated; The compensation verification module is used to receive the multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensation distortion. At the same time, it combines the confidence score and the attack mark to eliminate the low-confidence compensation data and attack frames, and outputs the complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification.
2. The face recognition system based on computer vision according to claim 1, wherein: The specific process of analyzing the depth data and polarized light parameters of the input image, classifying the occlusion types through the material reflection characteristics, and generating an occlusion heat map is as follows: Fuse the depth data and polarized light parameters of the input image, divide the physical distance levels of the occlusion area based on the depth information, combine the polarization angle distribution difference and the material reflectivity to distinguish the hard occlusion and soft occlusion types, generate a pixel-level occlusion heat map and label the material label, where the hard occlusion area is marked as a completely unrecognizable area, and the soft occlusion area is marked as a partially compensable area.
3. The face recognition system based on computer vision according to claim 1, wherein: The specific process of extracting the temperature change pattern and vibration frequency correlation in the infrared thermal distribution time series data and the acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer is as follows: According to the multi-modal spatio-temporal convolution kernel, continuously slide along the time dimension and synchronously capture the dynamic temperature gradient change of the infrared thermal distribution and the vibration frequency time series correlation of the acoustic spectrum; Eliminate the environmental noise interference through cross-modal feature cross-verification, and generate an encoded feature vector that fuses the spatio-temporal correlation of infrared and acoustic waves, which is used to characterize the coordination of physiological activities and physical movements in the occlusion area.
4. The face recognition system based on computer vision according to claim 3, characterized in that: The specific process of dynamically allocating the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map is as follows: Take the material type label of the heat map as a conditional vector, splice it with the multi-modal fusion feature vector, and input it into the attention network; According to the modal salience of the occlusion type label and the fused feature vector, dynamically allocate the contribution weights of the infrared thermal distribution and the acoustic spectrum, perform channel-level weighted fusion on the temporal features of the infrared thermal distribution and the temporal features of the acoustic spectrum, and generate a compensated feature vector for the occluded area; Input the weighted compensated feature vector into the generator network, restore the facial details of the occluded area through multi-level upsampling, and generate a predicted facial occluded area feature map; And impose a cross-modal consistency constraint to ensure the alignment of the predicted features and the visible light features of the non-occluded area in the semantic space.
5. The face recognition system based on computer vision according to claim 4, characterized in that: The specific process of calculating the cross-modal distribution similarity between the generated predicted facial occluded area feature map and the visible light features of the non-occluded area to output the confidence score is as follows: Map the generated predicted occluded area feature map and the visible light features of the non-occluded area to the shared embedding space, calculate the cosine similarity between the two through the contrastive learning algorithm, and combine the confidence decay factor of the occluded area to generate a normalized confidence score, where the confidence score is used to characterize the alignment degree between the generated features and the real facial features in the semantic space.
6. The computer vision-based face recognition system according to claim 5, characterized in that: When the acoustic wave sensor detects a voice trigger event, verify the synchronization between the temperature change in the target area of the infrared image and the acoustic spectrum. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a forgery attack and an attack mark is generated. The specific process is as follows: Align the peak of the acoustic spectrum and the infrared temperature time series curve according to the voice trigger timestamp, calculate the temporal matching degree between the two through the dynamic time warping algorithm. If the matching degree is lower than the adaptive threshold, it is determined as a cross-modal forgery attack, generate an attack mark and trigger an alarm signal, and at the same time freeze the feature fusion weight of the current frame to block the attack propagation.
7. The face recognition system based on computer vision according to claim 6, characterized in that: The specific process of receiving a multi-frame prediction feature sequence and analyzing the difference rate of the occluded area features of adjacent frames through a sliding window to mark compensation distortion is as follows: Intercept a fixed-length frame interval in the continuous multi-frame prediction feature sequence, calculate the Euclidean distance difference rate of the occluded area features of adjacent frames frame by frame, and determine the compensation distortion through the adaptive threshold, specifically including: Perform pixel-by-pixel comparison on the occluded area feature maps of each pair of adjacent frames in the window, and calculate the average value of the Euclidean distance as the difference rate; If the occlusion type is hard occlusion, the difference rate threshold is automatically reduced to enhance the sensitivity to small feature jumps; If the difference rate exceeds the adaptive threshold of the current window, mark this frame interval as compensation distortion and generate a distortion type label.
8. The face recognition system based on computer vision according to claim 7, wherein: The specific process of combining the confidence score and the attack mark to eliminate the low-confidence compensation data and attack frames, and output the complete facial feature vector and risk level verified by spatio-temporal coherence and cross-modal logic is as follows: Perform logical correlation analysis on the low-confidence compensation data, attack marks, and temporal distortion labels. If there are multiple types of abnormal signals in the same frame interval, mark it as a compound abnormality; The low-confidence compensation data refers to the feature data whose confidence score is lower than the preset confidence score threshold; Dynamically allocate weights according to the abnormal type and the occluded area area, and calculate the comprehensive risk score; If the risk score reaches the high-risk level, only output the feature vector of the non-occluded area and trigger a warning; If the risk score is in the medium risk level, output the complete feature vector but restrict the permissions for downstream sensitive operations; If the risk score is in the low risk level, directly output the complete feature vector and perform identity verification.
9. A computer vision-based face recognition method, applied to the computer vision-based face recognition system according to any one of claims 1-8, characterized in that, It includes the following steps: S1. Analyze the depth data and polarization light parameters of the input image, classify the occlusion types through the material reflection characteristics, generate an occlusion heat map, and mark the recognizable areas and occlusion areas; S2. Dynamically activate the infrared and acoustic sensor data according to the occlusion type labels of the heat map, extract the temperature change patterns and vibration frequency correlations in the infrared thermal distribution time series data and acoustic spectrum time series data of the occlusion area through the spatio-temporal convolutional layer, and dynamically allocate the contribution weights of the infrared and acoustic features based on the cross-modal attention mechanism to generate a predicted facial occlusion area feature map; S3. Calculate the cross-modal distribution similarity between the generated predicted facial occlusion area feature map and the visible light features of the unoccluded area to output a confidence score, and verify the synchronization of the temperature change in the target area of the infrared image and the acoustic spectrum when the acoustic sensor detects a voice trigger event. If the spatio-temporal deviation exceeds the preset threshold, it is determined as a spoofing attack and an attack mark is generated; S4. Receive the multi-frame prediction feature sequence and analyze the difference rate of the occlusion area features of adjacent frames through a sliding window to mark the compensation distortion. At the same time, combine the confidence score and the attack mark to eliminate the low-confidence compensation data and attack frames, and output the complete facial feature vector and risk level that pass the spatio-temporal coherence verification and cross-modal logic verification.
Citation Information
Patent Citations
AI-based face recognition verification management method, system and equipment and storage medium
CN115810214A
Mask wearing detection method based on polarization imaging AI identification
CN115909457A