A composite interference lightweight face verification method, system and device based on risk calibration factor decomposition
Patent Information
- Application Number
- CN202610979492.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]为了解决现有人脸验证方法中干扰线索向身份嵌入泄漏不可控、且训练目标未对目标低误识率分位点附近的冒名得分尾部直接施加约束,从而在复合干扰条件下验证可靠性下降、且单纯依靠骨干网络扩容无法兼顾边缘部署需求的问题,本发明提出一种基于风险校准因子分解的复合干扰轻量级人脸验证方法、系统和装置,具体如下:
[0084]1. 本发明首次在轻量级人脸验证管线的“反射感知归一化—因子分解表征—低误识率尾部校准”三个关键阶段统一施加同一干扰-身份分解先验,彻底改变了现有方法中各阶段约束相互割裂、干扰线索沿管线被静默携带的本质问题。在固定参数预算(约参数、约
)下,本发明在涵盖侧面姿态、真实口罩遮挡、非配合采集与反射增广四类条件的复合干扰评测集上将复合干扰综合得分(CDS)由同容量基线的84.38提升至89.05,提升幅度达4.67个百分点(经Bonferroni校正后
),证明了所述跨阶段统一约束机制的有效性。
Smart Images

Figure CN122821601A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and artificial intelligence, specifically to a lightweight face verification method and system based on risk calibration factor decomposition for composite interference. It can be applied to edge resource-constrained scenarios such as access control, video surveillance, and mobile terminal identity verification to perform highly reliable, low-false-recognition-rate one-to-one face verification on face image data under composite interference conditions such as specular reflection, side or large-angle yaw attitude, and obstruction by medical or protective masks. Background Technology
[0002] With the rapid development of deep learning, facial verification has been widely deployed in business scenarios such as access control, intelligent security, and mobile terminal identity verification. However, when facial verification systems migrate from cooperative to non-cooperative acquisition environments, the acquired facial images are often superimposed with multiple interference factors, such as specular reflections, side or large-angle yaw postures, and occlusion by medical or protective masks. In such complex interference scenarios, the interference components are very likely to leak into the identity embedding space as false recognition clues through the feature extraction channel, thereby raising the tail of the impersonation score near the target false recognition rate, resulting in a significant decrease in the true recognition rate (TAR) at the fixed false recognition rate operating point.
[0003] Existing research mainly follows four independent paths: The first path is the design of angle-separated loss functions, such as enhancing intra-class compactness and inter-class separability based on multiplicative, additive, and angle-separated methods; the second path is robust representation learning specific to interference types, such as part decomposition and mask reconstruction for occlusion, 3D alignment and frontalization for pose, and bi-branch decoupling representation for illumination; the third path is compact structures and knowledge distillation for edge deployment, such as lightweight backbones based on depthwise separable convolutions and adaptive distillation based on teacher features; and the fourth path is quality perception and fairness calibration, such as quality adaptive margin based on feature magnitude and threshold correction based on explicit quality scores.
[0004] In practice, the four routes mentioned above are often optimized independently, introducing constraints only for local parts of the face verification pipeline. This leads to two fundamental defects that have not yet been fully resolved: First, interfering cues are not explicitly excluded from identity embedding, causing interfering information to be carried into the verification stage as implicit features; second, the training phase lacks direct targeting to reduce false recognition rates (such as...). Explicit constraints on the tail of spoofing scores (at the level of classification) rely solely on angular interval loss for uniform supervision of all samples, making it difficult to constrain the small but decisive tail of spoofing scores. Furthermore, simply increasing the backbone network capacity cannot simultaneously alleviate these two types of shortcomings and is not feasible in scenarios with limited edge resources. In summary, there is an urgent need to propose a method that can uniformly impose constraints at both the representation and decision levels, and significantly improve the reliability of face verification under combined interference conditions with a fixed and compact parameter budget. Summary of the Invention
[0005] To address the problems in existing face verification methods, such as uncontrollable leakage of interference clues into identity embedding, lack of direct constraints on the tail of imposter scores near low false recognition rate points in the training target leading to decreased verification reliability under combined interference conditions, and inability to meet edge deployment requirements by simply relying on backbone network expansion, this invention proposes a lightweight face verification method, system, and device based on risk calibration factor decomposition under combined interference, as detailed below:
[0006] This invention provides a lightweight face verification method based on risk calibration factor decomposition for composite interference, comprising:
[0007] Acquire visible light face image data;
[0008] The face image data is subjected to reflectance-sensory normalization preprocessing to obtain a normalized face image with suppressed specular highlight components; the reflectance-sensory normalization preprocessing is a deterministic operation without learnable parameters.
[0009] The normalized face image is extracted using a pre-set compact backbone network to obtain a mid-level face feature map. During the training phase, a pose-mask decoupling branch is attached to generate two parallel compensation branches for the mid-level face feature map in the pose structure direction and the occlusion compensation direction, respectively. The two parallel compensation branches and the original mid-level features are then fused with gating to obtain the fused mid-level face features.
[0010] The fused mid-level face features are decomposed into identity codes and interference codes using identity projection matrices and interference projection matrices. Then, linear independence constraints and nonlinear independence constraints represented by the Hilbert-Schmidt Independence Criterion (HSIC) are jointly applied in a small batch, along with counterfactual consistency constraints on the identity codes under three types of identity-preserving interventions: reflection, pose, and mask, to obtain the factor-decomposed identity embedding.
[0011] Based on the identity embedding, the scores of same-identity pairs and different-identity pairs are calculated, and a softplus tail risk calibration term for target-oriented low false recognition rate quantiles is constructed. The identity embedding is then optimized in a supervised manner using a curriculum-based angular interval recognition head to obtain a trained face verification model.
[0012] During the deployment phase, the pose-mask decoupling branch and all auxiliary branches related to training are structurally stripped away, leaving only a single-branch inference path consisting of reflection perception normalization preprocessing, a compact backbone network, and a recognition head; the face image data to be verified is propagated forward along the single-branch inference path to obtain the identity embedding and complete one-to-one face verification.
[0013] Wherein, both the identity code and the interference code are low-dimensional vectors, and the identity projection matrix and the interference projection matrix are jointly optimized during the training phase, and the identity embedding is obtained by L2 normalization of the identity code.
[0014] The facial image data includes one or more of the following: facial image data acquired under cooperative acquisition conditions, facial image data acquired under non-cooperative acquisition conditions, facial image data containing specular reflection highlight components, facial image data containing side or large-angle yaw postures, facial image data containing components obscured by medical masks or protective masks, and continuous facial video frame data acquired by access control terminals or video surveillance equipment.
[0015] The step of performing reflectance-sensory normalization preprocessing on the face image data to obtain a normalized face image with suppressed specular highlight components includes:
[0016] The face image data is subjected to face detection and five-point keypoint alignment to obtain the aligned face image. ;
[0017] The aligned face image The following deterministic operations are performed sequentially to obtain the soft specular map. : (i) will (ii) Convert to grayscale and perform contrast-limited adaptive histogram equalization; (iii) obtain a binary specular mask by using the 92nd percentile of the pixel intensity of the equalized image as a threshold; (iv) perform Gaussian blurring on the binary specular mask to obtain the soft specular image with values in the range [0,1]. ;
[0018] Perform specular highlight suppression on the aligned face image using the following formula to obtain the normalized face image. :
[0019] ;
[0020] in, This is for specular highlight attenuation gain; For normalization bias; It is a numerical stability constant; preferably, , , The reflection perception normalization preprocessing does not contain learnable parameters and remains consistent during the training and inference phases.
[0021] The process of attaching a pose-mask decoupling branch during the training phase generates two parallel compensation branches for the mid-layer facial feature map in both the pose structure direction and the occlusion compensation direction. These two parallel compensation branches are then fused with the original mid-layer features through gating to obtain the fused mid-layer facial features, including:
[0022] The face mid-layer feature map output by the intermediate layer of the compact backbone network Simultaneous input of parallel attitude compensation sub-branch and occlusion compensation sub-branch ;
[0023] The attitude compensation sub-branch With the occlusion compensation sub-branch Each includes a depth-separable residual block, and the depth-separable residual blocks are as follows: Depthwise separable convolution, batch normalization, ReLU activation, Pointwise convolution and batch normalization are performed to output pose compensation features. With occlusion compensation features ;
[0024] Will , , After stitching along the channel dimension, the data is input into the fusion module. The fusion module include Convolution, batch normalization, ReLU activation, and a learnable scalar gate with an initial value of 0.1 are applied; the output of the fusion module is then combined with the original mid-layer features through the learnable scalar gate. Residual mixing yields the fused mid-layer facial features: ;
[0025] The posture-mask decoupling branch is a training-specific branch that is structurally removed during the deployment phase; the total number of parameters in the posture-mask decoupling branch does not exceed the total number of parameters in the compact backbone network. Preferably, the compact backbone network is a MobileFaceNet-shaped residual bottleneck network with a width factor of approximately 0.9, and the input resolution of the compact backbone network is [missing information]. The total number of parameters of the compact backbone network does not exceed The number of floating-point operations does not exceed .
[0026] The fused mid-level facial features are decomposed into identity codes and interference codes using an identity projection matrix and an interference projection matrix. Furthermore, linear independence constraints and nonlinear independence constraints represented by the Hilbert-Schmidt independence criterion are jointly applied within a small batch, along with counterfactual consistency constraints on the identity codes under three types of identity-preserving interventions: reflection, pose, and mask. This yields the factor-decomposed identity embedding, including:
[0027] Through identity projection matrix With interference projection matrix The fused mid-layer facial features Projected as identification code and interference code respectively: , The identity code With the interference code All are 128-dimensional vectors;
[0028] In a small batch, the identity code and the interference code are stacked into a matrix. and The factorization loss is constructed using the following formula:
[0029] ;
[0030] in, This represents the linear correlation suppression term between the identity code and the interference code; Indicates identity code and interference attributes The nonlinear correlation suppression term between them is based on the Hilbert-Schmidt independence criterion; The weighting coefficients of the nonlinear correlation suppression term; preferably, The interference attribute Labeling attributes for three types of interference factors: reflection, attitude, and mask;
[0031] Constructing a system that includes reflection intervention Posture intervention With mask intervention Identity Preservation Intervention Set The reflection intervention is achieved by randomly attaching a specular highlight template with random opacity in the range of [0.3, 0.8] to the aligned face image; the pose intervention is achieved by using a three-dimensional deformation model at a yaw angle of [0.3, 0.8]. The face within the interval is synthesized by frontal projection; the mask intervention is synthesized by attaching a medical mask, cloth mask or protective mask template to the aligned face image; the above three types of intervention are applied independently with preset probabilities during the training phase and do not change the identity label of the sample;
[0032] Apply counterfactual consistency constraints to the identity code:
[0033] ;
[0034] in, To train a dedicated two-layer projection head, the hidden layer of the projection head has a dimension of 256 and includes a layer normalization operation; the counterfactual consistency constraint is applied between the original sample and the identity code projection of the sample after any intervention. Distance monitoring;
[0035] Build an auxiliary interference detector Apply interference reproducibility monitoring to the interference code:
[0036] ;
[0037] in, Represents the binary cross-entropy; Indicating intervention The system generates interference attribute labels corresponding to the samples; during the training phase, the auxiliary interference identifyr ensures that the interference code maintains the identifiability of the interference information, and avoids the identity code and interference code from collapsing into constant solutions at the same time.
[0038] The process involves calculating scores for same-identity pairs and different-identity pairs based on the identity embedding, constructing a softplus tail risk calibration term for low-false-recognition-rate quantiles; and performing supervised optimization of the identity embedding using a curriculum-based angular interval recognition head to obtain a trained face verification model, including:
[0039] Within a small batch, construct a set of scores for same-identity pairs from the identity embeddings based on identity tags. Scoring set with different identities ;
[0040] Estimated target false recognition rate The corresponding scores for opposite identities quantiles The estimation employs a hybrid estimation method combining a first-in-first-out (FIFO) fraudulent scoring pool and an average operating index; wherein the capacity of the FIFO fraudulent scoring pool is... The average decay coefficient of the operating index is 0.95; preferably, , ;
[0041] Constructing a softplus tail risk calibration term for target low false recognition rate quantiles:
[0042] ;
[0043] in, This indicates the score suppression margin of different identities; This indicates that the same identity maintains a margin for scoring; This represents the balance coefficient between the same-identity preservation term and the different-identity suppression term; preferably, The first term in the above formula applies to scores exceeding the tail threshold. For dissimilar identity pairs, the scores of the dissimilar identity pairs are suppressed downwards; the second term applies to scores below a threshold. To avoid excessive narrowing of scores for pairs of individuals with the same identity;
[0044] A curriculum-based angular interval recognition head is constructed, and its training objective is:
[0045] Apply an exponential moving average to the cosine similarity of the target categories: ,in The coefficient is the exponential moving average coefficient; preferably, ;
[0046] To satisfy Hard negative sample category Modulate its logit as follows:
[0047] ;
[0048] Construct the identity loss term as follows:
[0049] ;
[0050] in, Represents the Sigmoid restricted mapping function. This represents the Sigmoid activation function; Indicates the angular scale factor; Indicates angular interval; preferably, The logit modulation operation for the hard negative samples is performed only during the training phase.
[0051] Constructing efficiency regularization terms:
[0052] ;
[0053] in, This represents the student identity embedding obtained from the compact backbone network; This indicates the teacher identity embedding obtained from the frozen teacher network; Indicates the first Scale parameters for batch normalization; Indicates channel sparse weights; preferably, The teacher network is an IResNet-100 backbone network trained using additive angular margin loss.
[0054] The total loss function is constructed by weighted summing of the above loss terms according to the following formula, and then jointly optimized based on the gradient backpropagation algorithm for the compact backbone network, the posture-mask decoupling branch, the identity projection matrix, the interference projection matrix, the projection head, the auxiliary interference recognizer, and the curriculum-based angle interval recognition head:
[0055] ;
[0056] in, , , , , These represent the weights of the factorization loss, counterfactual consistency loss, tail risk calibration loss, interference reproducibility loss, and efficiency regularization term, respectively; preferably, .
[0057] During the deployment phase, the pose-mask decoupling branch and all training-related auxiliary branches are structurally stripped away, retaining only a single-branch inference path consisting of reflection-aware normalization preprocessing, a compact backbone network, and a recognition head, including:
[0058] During the deployment phase, the posture-mask decoupling branch, all auxiliary branches in the identity projection matrix except for those necessary for identity embedding projection, the interference projection matrix, the projection head, the auxiliary interference recognizer, the portion of the curriculum-based angle interval recognition head except for those necessary for forward embedding, and all branches related to the teacher network are removed from the computation graph.
[0059] The face image data to be verified is sequentially processed through the reflection perception normalization preprocessing, the compact backbone network, and the recognition head to output a 128-dimensional identity embedding; and the identity embedding is then L2 normalized.
[0060] Calculate the cosine similarity between the identity embedding of the face image data to be verified and the target identity embedding, and compare the cosine similarity with a preset deployment threshold to obtain a one-to-one face verification result; wherein, the deployment threshold is based on the target false recognition rate. Scoring of the heteroidentity pairs The estimated values for quantiles were determined.
[0061] Based on the same concept, the present invention also provides a lightweight face verification system based on risk calibration factor decomposition and composite interference, comprising:
[0062] The data acquisition module is used to acquire facial image data;
[0063] The reflection perception normalization module is used to perform reflection perception normalization preprocessing on the face image data to obtain a normalized face image with suppressed specular highlight components; the reflection perception normalization preprocessing is a deterministic operation without learnable parameters.
[0064] The interference factor decoupling module is used to extract features from the normalized face image using a preset compact backbone network to obtain a mid-level face feature map. During the training phase, a pose-mask decoupling branch is attached to generate two parallel compensation branches for the mid-level face feature map in the pose structure direction and the occlusion compensation direction, respectively. The two parallel compensation branches and the original mid-level features are then fused together through gating to obtain the fused mid-level face features.
[0065] The factor decomposition and counterfactual consistency module is used to decompose the fused mid-level face features into identity codes and interference codes through identity projection matrices and interference projection matrices; and jointly apply linear independence constraints and nonlinear independence constraints represented by the Hilbert-Schmidt independence criterion, as well as counterfactual consistency constraints on the identity codes under three types of identity preservation interventions: reflection, pose, and mask, to obtain the identity embedding after factor decomposition.
[0066] The tail risk calibration and recognition head module is used to calculate the same identity pair score and different identity pair score based on the identity embedding, and construct a softplus tail risk calibration item for target low false recognition rate quantiles; and combine the course-based angular interval recognition head to perform supervised optimization of the identity embedding to obtain the trained face verification model.
[0067] The inference verification module is used to structurally strip the pose-mask decoupling branch and all auxiliary branches related to training during the deployment phase, retaining only a single-branch inference path consisting of reflection perception normalization preprocessing, compact backbone network and recognition head, and performing forward propagation on the face image data to be verified to obtain identity embedding and complete one-to-one face verification.
[0068] Wherein, both the identity code and the interference code are low-dimensional vectors, and the identity projection matrix and the interference projection matrix are jointly optimized during the training phase, and the identity embedding is obtained by L2 normalization of the identity code.
[0069] The factorization and counterfactual consistency module includes:
[0070] The factorization submodule is used to project the fused mid-level face features using the identity projection matrix and the interference projection matrix to obtain the identity code and the interference code; and to construct a linear correlation suppression term between the identity code and the interference code and a nonlinear independence suppression term between the identity code and the interference attribute in a small batch.
[0071] The intervention generation submodule is used to independently synthesize the intervened samples based on the aligned face images according to three types of identity-preserving interventions: reflection, pose, and mask, while maintaining their identity labels unchanged;
[0072] The counterfactual consistency submodule is used to apply an L2 distance consistency constraint to the identity code after training a dedicated projection head;
[0073] The interference reproducibility submodule is used to attach an auxiliary interference identifier to the interference code and apply binary cross-entropy supervision.
[0074] The projection matrix optimization submodule is used to jointly optimize the identity projection matrix and the interference projection matrix during the backpropagation process, and send the identity code to the downstream identification head module after L2 normalization.
[0075] Based on the same concept, the present invention also provides a lightweight face verification device based on risk calibration factor decomposition of composite interference, characterized in that it includes: a visual sensing device and a face verification host computer.
[0076] The visual sensing device is a camera and auxiliary light source used in scenarios such as access control, video surveillance, and mobile terminal identity verification to collect facial image data.
[0077] The face verification host computer includes a control backend and a frontend interface;
[0078] The front-end interface is used for human-computer interaction to collect user input of algorithm parameters and training parameters of the lightweight face verification method based on risk calibration factor decomposition and composite interference, as well as the working parameters and instructions of the visual sensing device, and sends them to the control backend. It also visualizes the calculation process and verification results of the face verification. The algorithm parameters include, but are not limited to, the image processing methods, network model calculation formulas and specific formula parameter values mentioned in the lightweight face verification method based on risk calibration factor decomposition and composite interference, and model training parameters. The working parameters of the visual sensing device are the camera's frame rate, exposure time, and the color temperature and brightness of the auxiliary light source. The working instructions of the visual sensing device are the trigger times of the camera and the auxiliary light source.
[0079] The control backend is equipped with a memory and a processor. The memory stores the program, and the processor converts the working parameters and instructions of the front-end visual sensing device into hardware timing command signals and outputs them to the camera and auxiliary light source. It also loads the program to execute the method steps described above, adjusts the model parameters according to the algorithm formula input from the front-end interface, and realizes the inference learning training and deployment of inference applications for lightweight face verification based on risk calibration factor decomposition and composite interference.
[0080] Based on the same inventive concept, this invention also provides a lightweight face verification system based on risk calibration factor decomposition, comprising a data acquisition module, a reflection perception normalization module, an interference factor decoupling module, a factor decomposition and counterfactual consistency module, a tail risk calibration and recognition head module, and an inference verification module, the functions of each module corresponding one-to-one with the method described above.
[0081] In another aspect, the present invention also provides an electronic device, including at least one processor and a memory; the memory and the processor are connected via a bus; the memory is used to store one or more programs; when the one or more programs are executed by the at least one processor, a lightweight face verification method based on risk calibration factor decomposition based on composite interference is implemented as described above.
[0082] In another aspect, the present invention also provides a computer-readable storage medium having an executable program stored thereon, which, when executed, implements the aforementioned lightweight face verification method based on risk calibration factor decomposition of composite interference.
[0083] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0084] 1. This invention, for the first time, applies the same interference-identity decomposition prior to all three key stages of a lightweight face verification pipeline: "reflection perception normalization—factor decomposition representation—low false recognition rate tail calibration." This fundamentally changes the existing methods' inherent problems of fragmented constraints at each stage and the silent carrying of interference cues along the pipeline. Under a fixed parameter budget (approximately...), Parameters, approx. Under these conditions, this invention improves the Composite Disturbance Score (CDS) from 84.38 to 89.05 on a composite interference evaluation set covering four conditions: side posture, actual mask occlusion, non-cooperative acquisition, and reflection augmentation. This improvement represents a 4.67 percentage point increase (after Bonferroni correction). This demonstrates the effectiveness of the proposed cross-stage unified constraint mechanism.
[0085] 2. The reflection perception normalization preprocessing is based on contrast-limited adaptive histogram equalization, 92nd percentile threshold binarization, and... Gaussian blur is used to obtain a soft highlight image, and then... Specular highlight suppression is performed without any learnable parameters throughout the process. Compared with learning-based mirror removal networks, it can significantly reduce the contamination of identity features by highlights without increasing inference overhead, and maintains complete consistency between the training and inference stages, avoiding distribution shifts during training and deployment.
[0086] 3. The posture-mask decoupling branch is a dedicated training module, with a parameter cost only about one-third that of the backbone network. However, it can explicitly separate pose structure cues and occlusion compensation cues from the interference and reinject mid-level features using a gated fusion method. Since this branch is structurally stripped during the deployment phase, its inference path is completely consistent with the basic backbone, thus improving feature robustness under complex interference conditions without increasing inference latency or floating-point computation.
[0087] 4. The factorization constraint based on the Hilbert-Schmidt independence criterion simultaneously suppresses both linear and nonlinear correlations between the identity code and interference attributes. Compared with decoupling methods based on domain adversarial gradient inversion, the factorization constraint maintains the objective function as fully differentiable, avoiding optimization instability issues in minimax games. Compared with decoupling methods based on mutual information estimation, the factorization constraint has lower sample complexity. Combined with counterfactual consistency constraints on the identity code under three types of identity-preserving interventions (reflection, posture, and mask) and the binary cross-entropy supervision of the interference code by the auxiliary interference identifier, these three elements form a closed loop of "bidirectional independence—consistency constraint—interference reproduction," ensuring that interference information is routed to the interference code while the identity code remains stable.
[0088] 5. The softplus tail risk calibration term for target-oriented low false recognition rate quantiles is the first to explicitly anchor the focus of the training target to the target false recognition rate (e.g., Near the different identity score quantiles corresponding to the (level), a smooth hinge-type suppression is applied to different identity scores exceeding the tail threshold, while same identity scores below the retention threshold are retained, directly shaping the score distribution near the working point; compared with alternatives based on focus-based low false recognition rate mining, the average accuracy is 0.83 percentage points higher. This mechanism compensates for the inherent deficiency of angular interval loss in constraining tail samples.
[0089] 6. The course-based angular interval recognition head adaptively adjusts the logit modulation intensity of hard negative samples by using the exponential moving average of the cosine of the target class, in conjunction with the Sigmoid restricted mapping function. By limiting the effective cosine dynamic range, the most difficult negative samples are suppressed in the early stage of training, and the gradients are concentrated on the boundary samples in the later stage of training. The recognition head can form a co-optimization with the tail risk calibration term to further reduce the equal error rate (EER) near the operating point.
[0090] 7. The method described herein is applicable to the NVIDIA RTX 3090 platform, FP32, and batch size. , Under input, the single forward median delay is approximately Zero-shot transfer results on TinyFace and RFW demonstrate that the robustness gain does not impair cross-domain generalization ability, with a cross-race standard deviation of 1.74, comparable to the heavyweight model. In summary, this invention achieves both representation-level interference leakage suppression and decision-level low-false-rate tail shaping under a fixed and compact parameter budget, making it particularly suitable for edge deployment scenarios such as access control, video surveillance, and mobile terminal identity verification. Attached Figure Description
[0091] Figure 1 A schematic diagram of the overall process of a lightweight face verification method based on risk calibration factor decomposition for the present invention.
[0092] Figure 2 This is a schematic diagram of the internal structure and data flow of the reflection sensing normalization module in the method provided by the present invention;
[0093] Figure 3 This is a schematic diagram of the internal structure and data flow of the posture-mask decoupling module (PMD) in the method provided by the present invention;
[0094] Figure 4 This is a schematic diagram of the internal structure and data flow of identity-interference factor decomposition and counterfactual consistency constraints in the method provided by the present invention;
[0095] Figure 5 A schematic diagram of the internal structure and data flow of the low false recognition rate tail risk calibration and curriculum-based angle interval recognition head provided by the present invention;
[0096] Figure 6 A schematic diagram of the structural composition of a lightweight face verification system based on risk calibration factor decomposition for the present invention.
[0097] Figure 7 A schematic diagram of the structure of an electronic device provided by the present invention;
[0098] Figure 8 This is a schematic diagram of a lightweight face verification device based on risk calibration factor decomposition for the present invention. Detailed Implementation
[0099] This invention proposes a lightweight face verification method, system, device, and medium based on risk calibration factor decomposition and composite interference. The specific embodiments of this invention will be further described in detail below with reference to the accompanying drawings.
[0100] Example 1: Lightweight Face Verification Method with Composite Interference
[0101] Embodiment 1 of this invention provides a lightweight face verification method based on risk calibration factor decomposition and composite interference. The overall process is illustrated in the following diagram. Figure 1As shown, it includes:
[0102] Step 1: Acquire visible light face image data;
[0103] Step 2: Perform face detection and five-point keypoint alignment on the face image data, and perform reflection-sensing normalization preprocessing on the aligned face image to obtain a normalized face image with suppressed specular highlight components. ;
[0104] Step 3: Utilize the pre-set compact backbone network to... Feature extraction is performed to obtain the mid-level feature map of the face. During the training phase, a pose-mask decoupling module (PoseMaskDecouple, or PMD) is used to obtain the fused mid-level facial features. ;
[0105] Step 4: Through the identity projection matrix With interference projection matrix Will Decomposed into identity code With interference code Furthermore, linear independence constraints, Hilbert-Schmidt independence criteria, nonlinear independence constraints, counterfactual consistency constraints, and auxiliary interference identification supervision are applied within small batches.
[0106] Step 5: Calculate the score sets of same-identity pairs and different-identity pairs based on identity embedding, and construct a softplus tail risk calibration item for target low false recognition rate quantiles. Supervised optimization of the identity embedding is performed by combining a course-based angle interval recognition head;
[0107] Step 6: Sum the above loss terms with weights to form the total loss function. Joint optimization is performed based on the gradient backpropagation algorithm; after training convergence, the posture-mask decoupling branch and all training-specific branches are structurally stripped away during the deployment phase, leaving only a single-branch inference path consisting of reflection perception normalization preprocessing, compact backbone network and recognition head.
[0108] Step 7: Propagate the face image data to be verified forward along the single-branch inference path to obtain a 128-dimensional identity embedding; after L2 normalization of the identity embedding, calculate the cosine similarity with the target identity embedding, and obtain a one-to-one face verification result according to the preset deployment threshold.
[0109] Generally, the facial image data in step 1 may include one or more of the following: facial image data acquired under cooperative acquisition conditions, facial image data acquired under non-cooperative acquisition conditions, facial image data containing specular reflection highlight components, facial image data containing side profile or large-angle yaw postures, facial image data containing components obscured by medical masks or protective masks, and continuous facial video frame data acquired by access control terminals or video surveillance equipment. This data may come from a single sensor or from a multi-sensor fusion acquisition system.
[0110] 1. Reflection Aware Normalization Module (ReflAwareNorm, corresponding to step 2)
[0111] To eliminate false activations caused by specular reflection highlights in the identity feature channel, this invention inserts a reflection perception normalization module before the backbone network. The internal structure of this module and the data flow are as follows: Figure 2 As shown, its characteristics lie in complete determinism and the absence of learnable parameters. For a face image x after face detection and five-point keypoint alignment, the specific execution steps are as follows:
[0112] (a) Grayscale conversion and contrast-limited adaptive histogram equalization: Convert to grayscale And divided into blocks according to the shear limit of 2.0. right Perform Limit Contrast Adaptive Histogram Equalization (CLAHE) to obtain ;
[0113] (b) 92nd percentile thresholding: Binarization is performed on the pixel-by-pixel intensity at the 92nd percentile as the threshold to obtain a binary specular mask. ;
[0114] (c) Gaussian blur to obtain a soft specular map: for Apply standard deviation Gaussian blur is used to obtain a soft specular image. ;
[0115] (d) Specular highlight suppression: The normalized face image is obtained by the following formula. :
[0116] ;
[0117] in, , , The parameter values are obtained by balancing identity-preserving brightness retention and specular suppression intensity, and remain completely consistent during the training and inference phases;
[0118] (e) Identity-preserving augmentation (training phase only): as described in Based on this, random augmentation is performed using four identity-preserving augmentation strategies: reflection, posture, mask, and combination. The augmentation probabilities are 0.16, 0.14, 0.06, and 0.03, respectively. The augmentation strategies do not change the identity label of the sample.
[0119] Since the entire process described above does not contain any learnable parameters, the reflection perception normalization module is executed directly as the first segment of a single-branch inference path during deployment, without introducing additional parameters or floating-point operation overhead, and there is no training-deployment distribution offset.
[0120] II. Posture-Mask Decoupling Module (PMD, corresponding to step 3)
[0121] To decouple strong interferences such as attitude and occlusion without increasing the backbone network parameter budget, this invention attaches an attitude-mask decoupling module to the middle layer of the compact backbone network. This module is a dedicated training branch, and its internal structure is as follows: Figure 3 As shown. Let the face mid-layer feature map output by the backbone mid-layer be... The specific structure and execution steps of the module are as follows:
[0122] (a) Attitude compensation subbranch : Execute sequentially Depthwise separable convolution (number of groups = stride Batch normalization, ReLU activation, Pointwise convolution (output channels = Batch normalization, output pose compensation features ;
[0123] (b) Occlusion Compensation Subbranch :and The structures are the same but the parameters are independent, and the output occlusion compensation features are the same. ;
[0124] (c) Fusion Module :Will , , The intermediate feature with 3C channels is obtained by concatenating along the channel dimension, and then the following steps are performed sequentially. Convolution (input channels = 3C, output channels = C), batch normalization, ReLU activation; and gating with a learnable scalar initialized to 0.1. fused output Compared with the original middle layer features according to Residual mixing is performed to obtain the fused mid-layer facial features. ;
[0125] (d) Training-Specific Decoupling from Structure: All parameters of the pose-mask decoupling module are optimized only during the training phase. During deployment, the output of the fusion module is directly bypassed, and the learnable scalar gating is zeroed out. The inference path is optimized by merging operators to obtain the same inference path as the basic backbone. Therefore, the posture-mask decoupling module does not introduce any additional parameters or floating-point operation overhead during the deployment phase.
[0126] In a typical implementation, the number of backbone feature channels (With a width factor of approximately 0.9), the total number of parameters in the posture-mask decoupling module is approximately It only accounts for a small percentage of the total parameters of the backbone network. about.
[0127] III. Identity-Distractor Decomposition and Counterfactual Consistency (corresponding to step 4)
[0128] To explicitly eliminate interfering cues at the representational level, this invention incorporates mid-level facial features after fusion. An identity-interference dual-projection structure is constructed on top of this, and its internal data flow is as follows: Figure 4 As shown. Let the small batch size be... The specific execution steps are as follows:
[0129] (a) Bi-projection decomposition: through identity projection matrix With interference projection matrix Will Projected as identity codes respectively With interference code ; obtained by stacking in small batches ;
[0130] (b) Linear independence constraint: for and Apply To suppress the second-order correlation between the identity code and the interference code;
[0131] (c) Nonlinear independence constraints: for Interference properties Imposing nonlinear independence constraints based on the Hilbert-Schmidt independence criterion Specifically, HSIC uses Gaussian kernel estimation and is calculated using the following formula:
[0132] ;
[0133] in, and They represent respectively to and Gaussian kernel matrix; It is a centered matrix; preferably, Compared with decoupling methods based on domain adversarial gradient reversal, the nonlinear independence constraint keeps the objective function completely differentiable, avoiding optimization instability problems in minimax games.
[0134] The factorization loss is obtained as follows:
[0135] ;
[0136] Among them, the interference attribute Labeling attributes for three types of interference factors: reflection, attitude, and mask;
[0137] (d) Counterfactual consistency constraints: Constructing constraints that include reflection interventions Posture intervention With mask intervention Identity Preservation Intervention Set The reflection intervention involves randomly attaching a specular highlight template with random opacity in the range of [0.3, 0.8] to... The above is obtained; the attitude intervention generates the yaw angle based on the three-dimensional deformation model. The composite view is obtained; the mask intervention is achieved by attaching a medical mask, cloth mask, or protective mask template to... The above is synthesized. A trained dedicated projection head is applied to the identity code. After Distance consistency constraint:
[0138] ;
[0139] in, The image is a raw, visible light image of a human face. The normalized face image is input into the network after reflection perception normalization preprocessing (i.e., specular highlight suppression). , These are the encoding mappings for the identity code and the interference code, respectively. This is a projection head used only during the training phase. To be applied to normalized face images Identity preservation intervention, collection It consists of reflex intervention, posture intervention, and mask intervention. For intervention Corresponding reflection / attitude / mask interference annotation attributes.
[0140] The projection head It is a two-layer multilayer perceptron with a hidden layer dimension of 256 and layer normalization, and is used only during the training phase;
[0141] (e) Interference reproducibility monitoring: Attaching an auxiliary interference identifier For interference codes Output three binary probabilities, corresponding to reflection, attitude, and mask attributes, respectively, and apply binary cross-entropy supervision as follows:
[0142] ;
[0143] The auxiliary interference identifier ensures that the interference code will not collapse into a constant solution simultaneously with the identity code, thereby maintaining the functional separation between identity and interference.
[0144] IV. Low False Recognition Rate Tail Risk Calibration and Curriculum-Based Angle Spacing Recognition Head (corresponding to Step 5)
[0145] To explicitly shape the score tail near the target misrecognition rate at the decision-making level, this invention constructs a joint optimization structure of tail risk calibration and curriculum-based angular interval recognition head on top of identity embedding. The structure is as follows: Figure 5 As shown, the specific steps are as follows:
[0146] (a) For the identity code L2 normalization is performed to obtain identity embedding ;
[0147] (b) Construct score sets for each identity pair within a small batch, based on the identity label. Scoring set with different identities ;
[0148] (c) Estimating the target false recognition rate The corresponding scores for opposite identities quantiles A first-in-first-out (FIFO) scoring pool (capacity) is used. A hybrid estimation method combining the exponential moving average (attenuation coefficient 0.95) with the exponential moving average coefficient of the course-based angular interval; the attenuation coefficient is related to the exponential moving average coefficient of the course-based angular interval. Independent of each other; preferably ;
[0149] (d) Constructing the softplus tail risk calibration item:
[0150] ;
[0151] in, Indicates the suppression margin of scores for different identities; This indicates that the score should maintain a margin for the same identity. This represents the balance coefficient between the same-identity preservation term and the different-identity suppression term; preferably... ;
[0152] (e) Curriculum-based angular interval recognition head: cosine of target category Applying an exponential moving average ;in For satisfying Hard negative sample category Logit Modulation is performed. Early training phase. The modulation function is relatively small, and it suppresses the most difficult negative samples; in the later stages of training Increase the modulation center of gravity and shift it to the boundary samples, thereby completing the implicit curriculum from easy to difficult.
[0153] (f) Construct the identity loss term as follows, where , It is the Sigmoid activation function. , :
[0154] ;
[0155] (g) Efficiency regularization term: Constructing the efficiency regularization term for teacher-student distillation and channel sparse composite. ;in Obtained from the frozen IResNet-100 / additive angular-spaced teacher network. For the first Scale parameters for batch normalization; optimal selection ;
[0156] (h) Total Loss Function and Joint Optimization:
[0157] ;
[0158] Preferred weights The total loss function is jointly optimized using the gradient backpropagation algorithm on the compact backbone network, the posture-mask decoupling branch, the identity / interference projection matrix, the projection head, the auxiliary interference recognizer, and the curriculum-based angle interval recognition head.
[0159] V. Training Implementation Details and Deployment Forms (corresponding to steps 6 and 7)
[0160] In a typical training implementation, the compact backbone network employs a MobileFaceNet-shaped residual bottleneck structure with a width factor of approximately 0.9, and the input resolution is... Outputs a 128-dimensional L2-normalized identity embedding with approximately [number of parameters]. The floating-point operation volume is approximately Training uses a large-scale, publicly available face dataset (such as a cleaned version of MS1MV2) with deduplicated identities, a batch size of 64, and an initial learning rate of... And decays according to cosine; weighted decay ; FP16 mixed precision was used; EMA weight decay was 0.995; training lasted for 24 rounds.
[0161] After training, a structural decoupling operation is performed during the deployment phase: the learnable scalar gating of the pose-mask decoupling module is zeroed and its output is bypassed, equivalently removing the module from the computation graph; all auxiliary branches in the identity / interference projection matrix are removed except for the components necessary for identity embedding projection; the training-dedicated projection head is removed. The auxiliary interference identification h is removed from the computation graph; all branches related to the teacher network are removed. The final single-branch inference path is: reflection perception normalization preprocessing → compact backbone network → identification head → output 128-dimensional L2 normalized identity embedding.
[0162] In the inference phase, the face image data to be verified is propagated forward along the single-branch inference path to obtain the identity embedding. Then, the cosine similarity is calculated with the target identity embedding and compared with a preset deployment threshold to obtain a one-to-one face verification result. This is achieved on the NVIDIA RTX 3090 platform, FP32, and with a batch size of... , Under the given input, the single forward median delay of the method is approximately It can meet the real-time requirements of scenarios such as access control and video surveillance.
[0163] In summary, this invention significantly improves the reliability of face verification under composite interference conditions by uniformly applying interference-identity decomposition priors in three stages: reflection perception normalization preprocessing, identity-interference factor decomposition, and low false recognition rate tail calibration, and by combining it with a dedicated training auxiliary branch for structural stripping.
[0164] Example 2: Lightweight Face Verification System with Composite Interference
[0165] Based on the same inventive concept, Embodiment 2 of this invention also provides a lightweight face verification system based on risk calibration factor decomposition and composite interference, as shown in the schematic diagram below. Figure 6 As shown, it includes:
[0166] Data acquisition module: used to acquire face image data; the face image data may include one or more of the following visible light face image data: face image data acquired under cooperative acquisition conditions, face image data acquired under non-cooperative acquisition conditions, face image data containing specular reflection highlight components, face image data containing side or large-angle yaw postures, face image data containing components obscured by medical masks or protective masks, and continuous face video frame data acquired by access control terminals or video surveillance equipment.
[0167] A reflection-aware normalization module is used to perform face detection and five-point keypoint alignment on the face image data, and to perform reflection-aware normalization preprocessing without learnable parameters on the aligned face image. This module further includes:
[0168] Face detection and alignment submodule: performs face detection, key point regression, and five-point affine transformation alignment operations;
[0169] The Grayscale Conversion and CLAHE submodule converts the aligned face image to grayscale and performs contrast-limited adaptive histogram equalization.
[0170] Soft specular map generation submodule: Binarization based on 92 percentile threshold and Gaussian blur generates soft specular images ;
[0171] Specular highlight suppression submodule: Press Obtain a normalized face image;
[0172] Identity Preservation Augmentation Submodule (Training Phase Only): Randomly augments normalized face images using four strategies: reflection, pose, masking, and combination.
[0173] Interference factor decoupling module: used to extract features from the normalized face image using a preset compact backbone network to obtain a mid-level face feature map, and to attach a pose-mask decoupling branch during the training phase. This module further includes:
[0174] Compact backbone network submodule: Construct a low-parameter mid-level face feature extractor based on depthwise separable convolution;
[0175] Attitude compensation sub-branch: depth-separable residual block, outputs attitude compensation features. ;
[0176] Occlusion compensation sub-branch: Depth-separable residual block, output occlusion compensation features. ;
[0177] Gated fusion sub-branch: via Convolution, batch normalization, ReLU, and learnable scalar gating will , , The fusion yields the mid-level facial features. ;
[0178] Structural stripping submodule: During the deployment phase, the output of the posture-mask decoupling branch is bypassed and the learnable scalar gating is set to zero, which is equivalent to removing the branch from the inference graph.
[0179] Factorization and Counterfactual Consistency Module: This module further includes:
[0180] Factorization submodule: via identity projection matrix With interference projection matrix The merged mid-level facial features are decomposed into an identity code. With interference code ;
[0181] Linear independence constraint submodule: Based on Frobenius norm pairs Apply suppression;
[0182] HSIC Nonlinear Independence Constraint Submodule: Based on the Hilbert-Schmidt Independence Criterion Interference properties Apply nonlinear independence constraints;
[0183] Intervention generation submodule: Generates post-intervention samples based on three types of identity-preserving interventions: reflective template fitting, 3D deformation model rotation, and mask template fitting;
[0184] Counterfactual consistency submodule: trained by a dedicated projection head Apply to the identity code Distance consistency constraint;
[0185] Interference Reproducibility Submodule: Connects to an auxiliary interference identifier Apply binary cross-entropy monitoring to the interference code;
[0186] Projection Matrix Optimization Submodule: Jointly optimizes the identity projection matrix and interference projection matrix during backpropagation.
[0187] Tail Risk Calibration and Identification Head Module: This module further includes:
[0188] The score set construction submodule constructs score sets for each identity pair within a mini-batch, based on the identity tag. Scoring set with different identities ;
[0189] Tail-partial location estimation submodule: Estimating target false recognition rate based on a hybrid approach of first-in-first-out imposter score pool and running index averaging. Corresponding quantiles ;
[0190] Softplus Tail Risk Calibration Submodule: Press Apply smooth hinge-type suppression to the tail scores of different identities and apply maintenance to the scores of the same identity;
[0191] Course-based angular interval recognition head module: modulates the logit of hard negative samples based on the exponential moving average of the target class cosine;
[0192] Efficiency Regularization Submodule: Performs teacher-student identity embedding distillation and batch normalization scaling parameters. Sparse penalty;
[0193] Joint optimization submodule: Performs end-to-end joint optimization of parameters of each branch based on the gradient backpropagation algorithm.
[0194] Inference Verification Module: Used to perform structured stripping and single-branch inference during the deployment phase. This module further includes:
[0195] The structural stripping submodule: performs the stripping of training-specific structures such as the attitude-mask decoupling branch, unnecessary components in the identity / interference projection matrix, projection head, auxiliary interference recognizer, and teacher network;
[0196] Forward inference submodule: The face image data to be verified is propagated forward along a single-branch path of reflection perception normalization → compact backbone network → recognition head;
[0197] L2 normalization submodule: performs L2 normalization on identity embedding;
[0198] Cosine similarity comparison submodule: Calculates the cosine similarity between the identity embedding to be verified and the target identity embedding and compares it with the preset deployment threshold, outputting a one-to-one face verification result.
[0199] Example 3: Electronic Equipment
[0200] like Figure 7 As shown, Embodiment 3 of the present invention provides an electronic device, which may be a computer device, an embedded AI computing device, a smart access control terminal, a smart mobile device, an edge video surveillance front-end device, etc. The electronic device in this embodiment includes a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory is used to store execution programs (instructions), and the processor is used to execute the instructions stored in the memory, thereby implementing all the steps of the composite interference lightweight face verification method described in Embodiment 1.
[0201] The processor may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In particular, for access control terminals, mobile terminals, and video surveillance front-end units, low-power GPUs or SoC chips with NPUs are preferred to achieve real-time face verification.
[0202] Example 4: Readable storage medium
[0203] Based on the same inventive concept, Embodiment 4 of the present invention also provides a readable storage medium, specifically an electronic device readable storage medium, which is used to store programs and data. The storage medium may include a built-in storage medium of the electronic device, or it may include an extended storage medium supported by the electronic device. The storage space stores one or more instructions suitable for loading and execution by a processor, and these instructions may be one or more executable programs (including program code). The storage medium may be a high-speed RAM memory or a non-volatile memory (such as flash memory, disk storage, etc.). Loading and executing the instructions stored in the storage medium by the processor can implement all the steps of the composite interference lightweight face verification method described in Embodiment 1.
[0204] Example 5: Face Verification Device
[0205] Based on the same inventive concept, this application also provides a lightweight face verification device based on risk calibration factor decomposition and composite interference, such as... Figure 8 As shown, it includes a visual sensing device and a face verification host computer.
[0206] The visual sensing device is a camera and auxiliary light source used in scenarios such as access control, video surveillance, and mobile terminal identity verification to collect facial image data; the facial verification host computer includes a control backend and a frontend interface.
[0207] The front-end interface is used for human-computer interaction to collect user input of algorithm parameters and training parameters of the lightweight face verification method based on risk calibration factor decomposition and composite interference, as well as the working parameters and instructions of the visual sensing device, and sends them to the control backend. It also visualizes the calculation process and verification results of the face verification. The algorithm parameters include, but are not limited to, the image processing methods, network model calculation formulas and specific formula parameter values mentioned in the lightweight face verification method based on risk calibration factor decomposition and composite interference, and model training parameters. The working parameters of the visual sensing device are the camera's frame rate, exposure time, and the color temperature and brightness of the auxiliary light source. The working instructions of the visual sensing device are the trigger times of the camera and the auxiliary light source.
[0208] The control backend is equipped with a memory and a processor. The memory stores the program, and the processor converts the working parameters and instructions of the front-end visual sensing device into hardware timing command signals and outputs them to the camera and auxiliary light source. It also loads the program to execute the above method steps, adjusts the model parameters according to the algorithm formula input by the front-end interface, and realizes the inference learning training and deployment of inference application for composite interference lightweight face verification based on risk calibration factor decomposition.
[0209] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code.
[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit its scope of protection. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading the present invention, they can still make various changes, modifications or equivalent substitutions to the specific implementation of the application, but these changes, modifications or equivalent substitutions are all within the scope of protection of the claims pending approval.
Claims
1. A lightweight face verification method based on risk calibration factor decomposition and composite interference, characterized in that, include: Acquire visible light face image data; The face image data is preprocessed by reflection perception normalization to obtain a normalized face image with suppressed specular highlight components. The reflection sensing normalization preprocessing is a deterministic operation without learnable parameters; The normalized face image is extracted using a pre-set compact backbone network to obtain a mid-level face feature map. During the training phase, a pose-mask decoupling branch is attached to generate two parallel compensation branches for the mid-level face feature map in the pose structure direction and the occlusion compensation direction, respectively. The two parallel compensation branches and the original mid-level features are then fused with gating to obtain the fused mid-level face features. The fused mid-level face features are decomposed into identity codes and interference codes using identity projection matrices and interference projection matrices. Then, linear independence constraints and nonlinear independence constraints represented by the Hilbert-Schmidt Independence Criterion (HSIC) are jointly applied in a small batch, along with counterfactual consistency constraints on the identity codes under three types of identity-preserving interventions: reflection, pose, and mask, to obtain the factor-decomposed identity embedding. Based on the identity embedding, the scores of same-identity pairs and different-identity pairs are calculated, and a softplus tail risk calibration term for target-oriented low false recognition rate quantiles is constructed. The identity embedding is then optimized in a supervised manner using a curriculum-based angular interval recognition head to obtain a trained face verification model. During the deployment phase, the pose-mask decoupling branch and all auxiliary branches related to training are structurally stripped away, leaving only a single-branch inference path consisting of reflection perception normalization preprocessing, a compact backbone network, and a recognition head; the face image data to be verified is propagated forward along the single-branch inference path to obtain the identity embedding and complete one-to-one face verification. Wherein, both the identity code and the interference code are low-dimensional vectors, and the identity projection matrix and the interference projection matrix are jointly optimized during the training phase, and the identity embedding is obtained by L2 normalization of the identity code.
2. The lightweight face verification method based on risk calibration factor decomposition for composite interference as described in claim 1, characterized in that, The facial image data includes one or more of the following: facial image data acquired under cooperative acquisition conditions, facial image data acquired under non-cooperative acquisition conditions, facial image data containing specular reflection highlight components, facial image data containing side or large-angle yaw postures, facial image data containing components obscured by medical masks or protective masks, and continuous facial video frame data acquired by access control terminals or video surveillance equipment.
3. The lightweight face verification method based on risk calibration factor decomposition for composite interference as described in claim 1, characterized in that, The step of performing reflectance-sensory normalization preprocessing on the face image data to obtain a normalized face image with suppressed specular highlight components includes: The face image data is subjected to face detection and five-point keypoint alignment to obtain the aligned face image. ; The aligned face image The following deterministic operations are performed sequentially to obtain the soft specular map. : (i) will (ii) Convert to grayscale and perform contrast-limited adaptive histogram equalization; (iii) obtain a binary specular mask by using the 92nd percentile of the pixel intensity of the equalized image as a threshold; (iv) perform Gaussian blurring on the binary specular mask to obtain the soft specular image with values in the range [0,1]. ; Perform specular highlight suppression on the aligned face image using the following formula to obtain the normalized face image. : ; in, This is for specular highlight attenuation gain; For normalization bias; It is a numerical stability constant; preferably, , , The reflection perception normalization preprocessing does not contain learnable parameters and remains consistent during the training and inference phases.
4. The lightweight face verification method based on risk calibration factor decomposition for composite interference as described in claim 1, characterized in that, The process of attaching a pose-mask decoupling branch during the training phase generates two parallel compensation branches for the mid-layer facial feature map in both the pose structure direction and the occlusion compensation direction. These two parallel compensation branches are then fused with the original mid-layer features through gating to obtain the fused mid-layer facial features, including: The face mid-layer feature map output by the intermediate layer of the compact backbone network Simultaneous input of parallel attitude compensation sub-branch and occlusion compensation sub-branch ; The attitude compensation sub-branch With the occlusion compensation sub-branch Each includes a depth-separable residual block, and the depth-separable residual blocks are as follows: Depthwise separable convolution, batch normalization, ReLU activation, Pointwise convolution and batch normalization are performed to output pose compensation features. With occlusion compensation features ; Will , , After stitching along the channel dimension, the data is input into the fusion module. The fusion module include Convolution, batch normalization, ReLU activation, and a learnable scalar gate with an initial value of 0.1 are applied; the output of the fusion module is then combined with the original mid-layer features through the learnable scalar gate. Residual mixing yields the fused mid-layer facial features: ; The posture-mask decoupling branch is a training-specific branch that is structurally removed during the deployment phase; the total number of parameters in the posture-mask decoupling branch does not exceed the total number of parameters in the compact backbone network. Preferably, the compact backbone network is a MobileFaceNet-shaped residual bottleneck network with a width factor of approximately 0.9, and the input resolution of the compact backbone network is [missing information]. The total number of parameters of the compact backbone network does not exceed The number of floating-point operations does not exceed .
5. The lightweight face verification method based on risk calibration factor decomposition for composite interference as described in claim 1, characterized in that, The fused mid-level facial features are decomposed into identity codes and interference codes using an identity projection matrix and an interference projection matrix. Furthermore, linear independence constraints and nonlinear independence constraints represented by the Hilbert-Schmidt independence criterion are jointly applied within a small batch, along with counterfactual consistency constraints on the identity codes under three types of identity-preserving interventions: reflection, pose, and mask. This yields the factor-decomposed identity embedding, including: Through identity projection matrix With interference projection matrix The fused mid-layer facial features Projected as identification code and interference code respectively: , The identity code With the interference code All are 128-dimensional vectors; In a small batch, the identity code and the interference code are stacked into a matrix. and The factorization loss is constructed using the following formula: ; in, This represents the linear correlation suppression term between the identity code and the interference code; Indicates identity code and interference attributes The nonlinear correlation suppression term between them is based on the Hilbert-Schmidt independence criterion; The weighting coefficients of the nonlinear correlation suppression term; preferably, The interference attribute Labeling attributes for three types of interference factors: reflection, attitude, and mask; Constructing a system that includes reflection intervention Posture intervention With mask intervention Identity Preservation Intervention Set The reflection intervention is achieved by randomly attaching a specular highlight template with random opacity in the range of [0.3, 0.8] to the aligned face image; the pose intervention is achieved by using a three-dimensional deformation model at a yaw angle of [0.3, 0.8]. The face within the interval is synthesized by frontal projection; the mask intervention is synthesized by attaching a medical mask, cloth mask or protective mask template to the aligned face image; the above three types of intervention are applied independently with preset probabilities during the training phase and do not change the identity label of the sample; Apply counterfactual consistency constraints to the identity code: ; in, To train a dedicated two-layer projection head, the hidden layer of the projection head has a dimension of 256 and includes a layer normalization operation; the counterfactual consistency constraint is applied between the original sample and the identity code projection of the sample after any intervention. Distance monitoring; Build an auxiliary interference detector Apply interference reproducibility monitoring to the interference code: ; in, Represents the binary cross-entropy; Indicating intervention The system generates interference attribute labels corresponding to the samples; during the training phase, the auxiliary interference identifyr ensures that the interference code maintains the identifiability of the interference information, and avoids the identity code and interference code from collapsing into constant solutions at the same time.
6. The lightweight face verification method based on risk calibration factor decomposition for composite interference as described in claim 1, characterized in that, The same identity pair score and different identity pair score are calculated based on the identity embedding, and a softplus tail risk calibration item is constructed for the target low false recognition rate quantile. The identity embedding is then subjected to supervised optimization using a curriculum-based angular interval recognition head to obtain a trained face verification model, including: Within a small batch, construct a set of scores for same-identity pairs from the identity embeddings based on identity tags. Scoring set with different identities ; Estimated target false recognition rate The corresponding scores for opposite identities quantiles The estimation employs a hybrid estimation method combining a first-in-first-out (FIFO) fraudulent scoring pool and an average operating index; wherein the capacity of the FIFO fraudulent scoring pool is... The average decay coefficient of the operating index is 0.95; preferably, , ; Constructing a softplus tail risk calibration term for target low false recognition rate quantiles: ; in, This indicates the score suppression margin of different identities; This indicates that the same identity maintains a margin for scoring; This represents the balance coefficient between the same-identity preservation term and the different-identity suppression term; preferably, The first term in the above formula applies to scores exceeding the tail threshold. For dissimilar identity pairs, the scores of the dissimilar identity pairs are suppressed downwards; the second term applies to scores below a threshold. To avoid excessive narrowing of scores for pairs of individuals with the same identity; A curriculum-based angular interval recognition head is constructed, and its training objective is: Apply an exponential moving average to the cosine similarity of the target categories: ,in The coefficient is the exponential moving average coefficient; preferably, ; To satisfy Hard negative sample category Modulate its logit as follows: ; Construct the identity loss term as follows: ; in, Represents the Sigmoid restricted mapping function. This represents the Sigmoid activation function; Indicates the angular scale factor; Indicates angular interval; preferably, The logit modulation operation for the hard negative samples is performed only during the training phase. Constructing efficiency regularization terms: ; in, This represents the student identity embedding obtained from the compact backbone network; This indicates the teacher identity embedding obtained from the frozen teacher network; Indicates the first Scale parameters for batch normalization; Indicates channel sparse weights; preferably, The teacher network is an IResNet-100 backbone network trained using additive angular margin loss. The total loss function is constructed by weighted summing of the above loss terms according to the following formula, and then jointly optimized based on the gradient backpropagation algorithm for the compact backbone network, the posture-mask decoupling branch, the identity projection matrix, the interference projection matrix, the projection head, the auxiliary interference recognizer, and the curriculum-based angle interval recognition head: ; in, , , , , These represent the weights of the factorization loss, counterfactual consistency loss, tail risk calibration loss, interference reproducibility loss, and efficiency regularization term, respectively; preferably, .
7. The lightweight face verification method based on risk calibration factor decomposition for composite interference as described in claim 6, characterized in that, During the deployment phase, the pose-mask decoupling branch and all training-related auxiliary branches are structurally stripped away, retaining only a single-branch inference path consisting of reflection-aware normalization preprocessing, a compact backbone network, and a recognition head, including: During the deployment phase, the posture-mask decoupling branch, all auxiliary branches in the identity projection matrix except for those necessary for identity embedding projection, the interference projection matrix, the projection head, the auxiliary interference recognizer, the portion of the curriculum-based angle interval recognition head except for those necessary for forward embedding, and all branches related to the teacher network are removed from the computation graph. The face image data to be verified is sequentially processed through the reflection perception normalization preprocessing, the compact backbone network, and the recognition head to output a 128-dimensional identity embedding; and the identity embedding is then L2 normalized. Calculate the cosine similarity between the identity embedding of the face image data to be verified and the target identity embedding, and compare the cosine similarity with a preset deployment threshold to obtain a one-to-one face verification result; wherein, the deployment threshold is based on the target false recognition rate. Scoring of the heteroidentity pairs The estimated values for quantiles were determined.
8. A lightweight face verification system based on risk calibration factor decomposition and composite interference, characterized in that, include: The data acquisition module is used to acquire facial image data; The reflection perception normalization module is used to perform reflection perception normalization preprocessing on the face image data to obtain a normalized face image with suppressed specular highlight components; the reflection perception normalization preprocessing is a deterministic operation without learnable parameters. The interference factor decoupling module is used to extract features from the normalized face image using a preset compact backbone network to obtain a mid-level face feature map. During the training phase, a pose-mask decoupling branch is attached to generate two parallel compensation branches for the mid-level face feature map in the pose structure direction and the occlusion compensation direction, respectively. The two parallel compensation branches and the original mid-level features are then fused together through gating to obtain the fused mid-level face features. The factor decomposition and counterfactual consistency module is used to decompose the fused mid-level face features into identity codes and interference codes through identity projection matrices and interference projection matrices; and jointly apply linear independence constraints and nonlinear independence constraints represented by the Hilbert-Schmidt independence criterion, as well as counterfactual consistency constraints on the identity codes under three types of identity preservation interventions: reflection, pose, and mask, to obtain the identity embedding after factor decomposition. The tail risk calibration and recognition head module is used to calculate the same identity pair score and different identity pair score based on the identity embedding, and construct a softplus tail risk calibration item for target low false recognition rate quantiles; and combine the course-based angular interval recognition head to perform supervised optimization of the identity embedding to obtain the trained face verification model. The inference verification module is used to structurally strip the pose-mask decoupling branch and all auxiliary branches related to training during the deployment phase, retaining only a single-branch inference path consisting of reflection perception normalization preprocessing, compact backbone network and recognition head, and performing forward propagation on the face image data to be verified to obtain identity embedding and complete one-to-one face verification. Wherein, both the identity code and the interference code are low-dimensional vectors, and the identity projection matrix and the interference projection matrix are jointly optimized during the training phase, and the identity embedding is obtained by L2 normalization of the identity code.
9. A lightweight face verification system based on risk calibration factor decomposition for composite interference as described in claim 8, characterized in that, The factorization and counterfactual consistency module includes: The factorization submodule is used to project the fused mid-level face features using the identity projection matrix and the interference projection matrix to obtain the identity code and the interference code; and to construct a linear correlation suppression term between the identity code and the interference code and a nonlinear independence suppression term between the identity code and the interference attribute in a small batch. The intervention generation submodule is used to independently synthesize the intervened samples based on the aligned face images according to three types of identity-preserving interventions: reflection, pose, and mask, while maintaining their identity labels unchanged; The counterfactual consistency submodule is used to apply an L2 distance consistency constraint to the identity code after training a dedicated projection head; The interference reproducibility submodule is used to attach an auxiliary interference identifier to the interference code and apply binary cross-entropy supervision. The projection matrix optimization submodule is used to jointly optimize the identity projection matrix and the interference projection matrix during the backpropagation process, and send the identity code to the downstream identification head module after L2 normalization.
10. A lightweight face verification device based on risk calibration factor decomposition and composite interference, characterized in that, include: Visual sensing devices and facial recognition host computer; The visual sensing device is a camera and auxiliary light source used in scenarios such as access control, video surveillance, and mobile terminal identity verification to collect facial image data. The face verification host computer includes a control backend and a frontend interface; The front-end interface is used for human-computer interaction to collect the algorithm parameters and training parameters of the composite interference lightweight face verification method based on risk calibration factor decomposition, the working parameters and working instructions of the visual sensing device, and send them to the control backend, as well as to visually display the calculation process and verification results of face verification. The algorithm parameters include, but are not limited to, the image processing methods, network model calculation formulas and specific formula parameter values, and model training parameters mentioned in the lightweight face verification method based on risk calibration factor decomposition; the operating parameters of the visual sensing device are the camera's acquisition frame rate, exposure time, and the color temperature and brightness of the auxiliary light source; the operating instructions of the visual sensing device are the trigger times of the camera and the auxiliary light source. The control backend is equipped with a memory and a processor. The memory stores the program, and the processor converts the working parameters and instructions of the front-end visual sensing device into hardware timing command signals and outputs them to the camera and auxiliary light source. It also loads the program to execute the method steps as described in any one of claims 1-8, and adjusts the model parameters according to the algorithm formula input by the front-end interface to realize the inference learning training and deployment of inference application for composite interference lightweight face verification based on risk calibration factor decomposition.