A multi-modal biometric feature fusion method, apparatus, device and storage medium
Patent Information
- Application Number
- CN202611140184.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0005](1)缺乏跨模态注意力交互机制,可见光与红外特征之间信息交互不充分,融合特征判别力不足;
[0039]本申请提供的多模态生物特征融合方法,采用双向跨模态注意力机制,同时计算“可见光→近红外”和“近红外→可见光”两个方向的注意力权重,使两种模态特征实现深度双向语义交互;引入稀疏注意力机制过滤无效特征交互,规避传统注意力平方级计算缺陷,在保证识别精度不下降的前提下,有效降低模块计算量与推理延迟,提升模型运行效率;依托双向跨模态交互与动态门控权重,摆脱固定融合权重无法适配多变光照场景的缺陷,模型可自适应适配不同光照环境,最大化发挥可见光与红外模态的互补优势。本申请所提方法能够自适应完成双模态特征互补融合,在降低计算开销、保证推理效率的同时,显著提升复杂光照环境下的识别精度,且具备极强的网络兼容性与任务拓展性,能够高效适配各类双模态视觉识别任务。
Smart Images

Figure CN122842166A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and deep learning technology, and in particular to a method, apparatus, device and storage medium for multimodal biometric fusion. Background Technology
[0002] With the rapid development of deep learning technology, facial recognition-based attendance and access control systems have been widely deployed in enterprises, residential communities, transportation hubs, and other locations. Existing attendance and access control machines already possess a high recognition rate and a low false positive rate in typical usage scenarios, enabling them to accurately complete identity authentication tasks in most situations.
[0003] However, in large-scale deployments, attendance and access control systems still face significant technical challenges. As the number of registered user identities increases and the database expands, the performance of facial recognition algorithms degrades significantly, manifested in frequent false positives and rejections. The root cause of these problems lies in the fact that as the database expands, the distribution of different identity categories in the feature space becomes increasingly dense, resulting in smaller inter-class distances and larger intra-class differences. Traditional single-modal facial recognition methods struggle to maintain a low false positive rate while simultaneously maintaining a low rejection rate.
[0004] To address the aforementioned issues, multimodal biometric fusion technology, which integrates image information from complementary modalities such as visible light and infrared, enhances the model's feature discrimination capabilities under complex lighting conditions and large-scale databases. This has become an important technical approach to improve the recognition performance of attendance and access control systems. However, existing multimodal biometric fusion-based facial recognition technologies for attendance and access control systems mainly suffer from the following shortcomings:
[0005] (1) Lack of cross-modal attention interaction mechanism, insufficient information interaction between visible light and infrared features, and insufficient discriminative power of fused features;
[0006] (2) Most existing technologies use fixed weights or direct splicing to fuse multimodal features, lacking dynamic adaptive fusion capabilities and unable to adjust modal contribution weights according to scene changes;
[0007] (3) Some methods have too high computational complexity and are difficult to support the real-time recognition requirements under large-scale deployment. Summary of the Invention
[0008] To address the aforementioned issues, this application provides a method, apparatus, device, and storage medium for multimodal biometric fusion.
[0009] In view of this, the first aspect of this application provides a multimodal biometric fusion method, comprising:
[0010] The visible light features and near-infrared features extracted from the visible light image and near-infrared image, respectively, are input into a shared linear projection layer for query, key, and value triple mapping, to obtain the query, key, and value of the visible light features and the query, key, and value of the near-infrared features, respectively.
[0011] Based on visible light and near-infrared features, the query and key calculations are performed to determine the first cross-modal attention weight in the near-infrared direction for visible light queries and the second cross-modal attention weight in the visible light direction for near-infrared queries.
[0012] The first cross-modal attention weight and the second cross-modal attention weight are respectively subjected to sparsification processing to obtain the first sparse attention weight and the second sparse attention weight; the value of the near-infrared feature is weighted by the first sparse attention weight to obtain the first cross-modal feature, and the value of the visible light feature is weighted by the second sparse attention weight to obtain the second cross-modal feature;
[0013] The first cross-modal feature and the second cross-modal feature are concatenated and input into a gating network. The gating network adaptively learns to obtain dynamic gating weights. The first cross-modal feature and the second cross-modal feature are then fused using the gating weights to obtain multimodal fusion features.
[0014] Optionally, the method further includes:
[0015] The multimodal fusion features are linearly projected, and the projected fusion features are normalized and nonlinearly transformed. The nonlinearly transformed fusion features are then residually connected with the multimodal fusion features to obtain the final multimodal fusion features.
[0016] Optionally, the first cross-modal attention weights are sparsified to obtain first sparse attention weights, including:
[0017] Based on the first cross-modal attention weights, an importance score is calculated for each attention head;
[0018] The importance scores are sorted in descending order, and the first cross-modal attention weights corresponding to the top K positions are retained. The attention weights of the remaining positions are reset to zero to obtain the first sparse attention weights.
[0019] Optionally, the visible light features extracted from the visible light image are input into a shared linear projection layer for query, key, and value triple mapping to obtain the query, key, and value of the visible light features, including:
[0020] The visible light features extracted from the visible light image are input into a shared linear projection layer and linearly mapped. During the mapping process, the channel dimension is expanded by 3 times to obtain the visible light feature tensor.
[0021] The visible light feature tensor is split according to the channel dimension to obtain the query, key, and value of the visible light feature.
[0022] Optionally, the step of weighting the near-infrared feature values using the first sparse attention weight to obtain the first cross-modal feature, and weighting the visible light feature values using the second sparse attention weight to obtain the second cross-modal feature, further includes:
[0023] Normalization and regularization operations are performed on the first sparse attention weight and the second sparse attention weight, respectively.
[0024] Optionally, the gated network includes two fully connected layers;
[0025] The first fully connected layer is used to compress the channels of the spliced features and remove redundant features from the compressed spliced features using the ReLU activation function;
[0026] The second fully connected layer is used to map the features output by the first fully connected layer into a single-channel weight map, and normalize the single-channel weight map using the Sigmoid function to obtain dynamic gating weights.
[0027] A second aspect of this application provides a multimodal biometric fusion device, comprising:
[0028] The linear projection unit is used to input the visible light features and near-infrared features extracted from the visible light image and near-infrared image respectively into the shared linear projection layer and perform query, key, and value triple mapping to obtain the query, key, and value of the visible light features and the query, key, and value of the near-infrared features respectively.
[0029] The attention weight calculation unit is used to calculate the first cross-modal attention weight in the near-infrared direction for visible light queries and the second cross-modal attention weight in the visible light direction for near-infrared queries based on the query keys of visible light features and near-infrared features.
[0030] A cross-modal feature aggregation unit is used to perform sparsification processing on the first cross-modal attention weight and the second cross-modal attention weight respectively to obtain a first sparse attention weight and a second sparse attention weight; the value of the near-infrared feature is weighted by the first sparse attention weight to obtain a first cross-modal feature; and the value of the visible light feature is weighted by the second sparse attention weight to obtain a second cross-modal feature.
[0031] The feature fusion unit is used to concatenate the first cross-modal feature and the second cross-modal feature and input them into the gating network. The gating network adaptively learns to obtain dynamic gating weights, and the first cross-modal feature and the second cross-modal feature are fused using the gating weights to obtain multimodal fused features.
[0032] Optionally, the device further includes:
[0033] The projection and regularization unit is used to perform linear projection on the multimodal fusion features, normalize and nonlinearly transform the projected fusion features, and perform residual connection between the nonlinearly transformed fusion features and the multimodal fusion features to obtain the final multimodal fusion features.
[0034] A third aspect of this application provides an electronic device, the device including a processor and a memory;
[0035] The memory is used to store program code and transmit the program code to the processor;
[0036] The processor is used to execute any one of the multimodal biometric fusion methods described in the first aspect according to the instructions in the program code.
[0037] A fourth aspect of this application provides a computer-readable storage medium for storing program code that, when executed by a processor, implements the multimodal biometric fusion method described in any of the first aspects.
[0038] As can be seen from the above technical solutions, this application has the following advantages:
[0039] The multimodal biometric fusion method provided in this application employs a bidirectional cross-modal attention mechanism, simultaneously calculating attention weights in both the "visible light → near-infrared" and "near-infrared → visible light" directions, enabling deep bidirectional semantic interaction between the two modalities. A sparse attention mechanism is introduced to filter out invalid feature interactions, avoiding the shortcomings of traditional quadratic attention calculations. While maintaining recognition accuracy, this effectively reduces module computation and inference latency, improving model efficiency. Relying on bidirectional cross-modal interaction and dynamic gating weights, the method overcomes the limitation of fixed fusion weights being unable to adapt to varying lighting conditions. The model can adaptively adapt to different lighting environments, maximizing the complementary advantages of the visible light and infrared modalities. The proposed method can adaptively complete the complementary fusion of bimodal features, significantly improving recognition accuracy in complex lighting environments while reducing computational overhead and ensuring inference efficiency. It also possesses strong network compatibility and task scalability, efficiently adapting to various bimodal visual recognition tasks. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A flowchart illustrating a multimodal biometric fusion method provided in this application embodiment;
[0042] Figure 2 Comparative experimental results of the baseline method and the multimodal biometric fusion method provided in the embodiments of this application;
[0043] Figure 3 This is a schematic diagram of a multimodal biometric fusion device provided in an embodiment of this application. Detailed Implementation
[0044] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0045] For easier understanding, please refer to Figure 1 This application provides a multimodal biometric fusion method, including:
[0046] Step 110: Input the visible light features and near-infrared features extracted from the visible light image and near-infrared image respectively into the shared linear projection layer and perform query, key, and value triple mapping to obtain the query, key, and value of the visible light features and the query, key, and value of the near-infrared features respectively.
[0047] Feature extraction can be performed on visible light and near-infrared images using a backbone network (such as a Transformer network) to obtain visible light feature x and near-infrared feature y. Both visible light feature x and near-infrared feature y have shapes [B, L, C], where B is the batch size, L = W × H is the spatial sequence length, W is the width of the input image, H is the height of the input image, and C is the channel embedding dimension. Preferably, in this embodiment, C = 512 dimensions and L = 49.
[0048] The visible light feature x and the near-infrared feature y are respectively input into the parameter-shared linear projection layer self.qkv (weights W∈R).C×3C Bias b∈R 3C The channel dimension is expanded from C to 3C. The 3C channels are split into 3 groups (corresponding to query Q, key K, and value V respectively), and each group is further divided into H attention heads, each with a depth of d. h =C / H (Preferably, H=8, then d h =64). The transformation process of visible light feature x is [B, L, C] → reshape → [B, L, 3, H, d]. h ] →permute(2, 0, 3, 1, 4) → [3, B, H, L, d h The transformation process for near-infrared feature y is similar to that for visible light feature x:
[0049] qkv x =self.qkv(x).reshape(B,L,3,H,d h ).permute(2,0,3,1,4);
[0050] q x , k x ,v x =qkv x [0],qkv x [1],qkv x [2];
[0051] qkv y =self.qkv(y).reshape(B,L,3,H,d h ).permute(2,0,3,1,4);
[0052] q y , k y , v y =qkv y [0],qkv y [1],qkv y [2];
[0053] qkv x ,qkv y These are the triplet fusion tensors of visible light feature x and near-infrared feature y after linear projection and dimensional rearrangement, respectively; q x , k x ,v x These represent the query, key, and value of the visible light feature x, respectively, and q. y , k y , v yThese represent the query, key, and value of the near-infrared feature y, respectively; `self.qkv = nn.Linear(dim, dim * 3, bias=True);`; `permute(2, 0, 3, 1, 4)` adjusts the original "batch-length-type-head-depth" data arrangement to "type-batch-head-length-depth," which facilitates the separation of query Q, key K, and value V, and also adapts to subsequent multi-head parallel matrix multiplication operations. Using multi-head decomposition allows the model to calculate the correlation between modalities in different subspaces, enhancing its ability to capture diverse semantic patterns.
[0054] This application embodiment sets up eight parallel attention heads, which decompose the original high-dimensional features into multiple independent feature subspaces. Each attention head independently learns a set of cross-modal feature associations, capturing feature information of different dimensions of the target. The features of multiple subspaces are learned in parallel and then fused, which effectively improves the model's ability to express the differentiated features of the target and enriches the feature representation dimensions.
[0055] Step 120: Calculate the first cross-modal attention weight in the near-infrared direction for visible light query and the second cross-modal attention weight in the visible light direction for near-infrared query based on the query keys of visible light and near-infrared features;
[0056] The first cross-modal attention weight attn in the near-infrared direction of the visible light query is calculated based on the query and key of visible light feature x and near-infrared feature y. xy And the second cross-modal attention weight attn in the near-infrared query visible light direction. yx :
[0057] ;
[0058] ;
[0059] In the formula, `attn` is a scaling factor used to scale attention scores to prevent gradient vanishing. xy ,attn yx The first and second cross-modal attention weights are of shape [B, H, L, L]. This bidirectional attention ensures that the two modalities can fully capture the semantic relationships between them, enabling bidirectional information interaction.
[0060] Step 130: Sparsify the first cross-modal attention weight and the second cross-modal attention weight to obtain the first sparse attention weight and the second sparse attention weight; weight the near-infrared feature value with the first sparse attention weight to obtain the first cross-modal feature, and weight the visible light feature value with the second sparse attention weight to obtain the second cross-modal feature.
[0061] Attention weights attn∈R B×H×L×L Each attention head is processed separately, and attn is attn. xy or attn yx Fix the i-th head:
[0062] Take the tensor atn[:,i,:,:] of this attention head, with shape [B,L,L]; calculate the mean of the second component (i.e., dim=1) to obtain the global importance score of each query position under this attention head, with shape [B,L]:
[0063] ;
[0064] Calculate the quantity to be retained The ratio represents the sparsity ratio; preferably, ratio = 0.5. The `torch.topk` operation can be used on the last dimension of `scores` to return the top K indices `topk_indices∈R` with the largest values. B×K , that is, topk_indices = torch.topk(scores, K, dim=1).
[0065] Initialize an all-zero tensor sparse_attn∈R B×H×L×L For each sample b and each head h in the batch, only the columns of the selected rows (topk_indices) in the original attn are copied to the corresponding positions in sparse_attn, and the remaining rows are set to zero. This row index preservation means that only the K most important query positions can participate in subsequent value aggregation, greatly filtering out interference from background or low-discrimination regions. This is understandable, as the first cross-modal attention weight attn... xy The corresponding first sparse attention weight is sparse_attn xy Second cross-modal attention weights attn yx The corresponding second sparse attention weight is sparse_attn yx .
[0066] The obtained sparse attention weights sparse_attn are normalized in the last dimension (dim=-1) so that the sum of the weights in each row is 1:
[0067] sparse_attn xy = sparse_attn xy .softmax(dim=-1);
[0068] sparse_attnyx = sparse_attn yx .softmax(dim=-1);
[0069] The normalized sparse attention weights are then regularized by randomly deactivating them via drop, enhancing generalization ability. The shape of the regularized sparse attention weights remains unchanged, maintaining [B, H, L, L].
[0070] sparse_attn xy = nn.Dropout(sparse_attn xy );
[0071] sparse_attn yx = nn.Dropout(sparse_attn yx );
[0072] By weighting the values of visible light features and near-infrared features using regularized sparse attention weights, cross-modal features are obtained, all with shapes [B, H, L, d]. h ]:
[0073] The value v of the near-infrared feature is obtained by applying the first sparse attention weight after regularization. y Weighting is performed to obtain the first cross-modal feature x. cross Visible light fused with infrared features: x cross =sparse_attn xy ·v y ;
[0074] The values of the visible light features are weighted by the regularized second sparse attention weights to obtain the second cross-modal feature y. cross That is, infrared fusion of visible light features: y cross =sparse_attn yx ·v x ;
[0075] Then, perform `transpose(1, 2)` on the cross-modal features (including the first and second cross-modal features) to change the dimensions to [B, L, H, d]. h Then, execute `reshape(B, L, C)` to merge the multi-head subspace back into the original C-dimensional embedding space. At this point, x... cross With y cross The shapes are all [B, L, C].
[0076] Traditional full attention mechanisms suffer from quadratic-level computational redundancy, and numerous interactions at invalid feature locations increase inference time. This application introduces a sparse attention optimization strategy. A fixed sparsity ratio is set, and the global importance score is applied to the attention matrix of each attention head. A Top-K algorithm is used to select the most relevant key feature locations, retaining only the attention weights of these key locations while setting the weights of all other redundant locations to zero. This approach significantly reduces invalid computations and lowers the overall computational complexity of the model, thereby improving inference speed, with almost no loss of effective fusion information.
[0077] Step 140: The first cross-modal feature and the second cross-modal feature are concatenated and input into the gating network. The gating network adaptively learns to obtain dynamic gating weights. The first cross-modal feature and the second cross-modal feature are fused using the gating weights to obtain multimodal fusion features.
[0078] The first cross-modal features and the second cross-modal features are concatenated:
[0079] ;
[0080] The shape of the splicing feature gate_input is [B, L, 2C].
[0081] The concatenated feature `gate_input` is input into the gating network `self.gate`. The gating network is a two-layer fully connected structure. The first fully connected layer `nn.Linear(2C, C)` maps the dimension from 2C to C and is activated by the ReLU activation function. The second fully connected layer `nn.Linear(C,1)` maps the dimension from C to 1 and is activated by the Sigmoid activation function. The output is the gating weight `gate_weight`, whose value is between 0 and 1.
[0082] ;
[0083] In the formula, σ represents the Sigmoid function, W1, b1 and W2, b2 are the weights and biases of the two fully connected layers, respectively; the first linear layer realizes channel dimensionality reduction and cross-modal correlation mining, and the second linear layer maps the features to single-channel weight coefficients; ReLU introduces nonlinear fitting ability to improve the gating network's ability to fit complex lighting inputs; the terminal Sigmoid activation function strictly constrains the output weights within the range of 0 to 1 to ensure that the weights conform to the numerical specifications of weighted fusion.
[0084] Two cross-modal features are dynamically fused based on gating weights:
[0085] ;
[0086] In the formula, This indicates element-wise multiplication.
[0087] Traditional bimodal fusion methods often employ static fusion approaches such as direct addition, concatenation summation, and fixed-weighting. All input samples share the same set of fusion weights, making them unsuitable for complex scenarios like sudden changes in lighting, low light, strong light, and backlighting. For instance, under normal lighting, visible light facial texture details are more discriminative, while in low light, visible light features have higher noise levels, while infrared features are more stable. Fixed weights cannot dynamically balance the advantages of each modality, easily leading to the dilution of effective features and the retention of redundant noise. To address this, this application designs a lightweight two-level fully connected gated fusion network that adaptively generates dynamic fusion coefficients based on data, achieving scene-aware intelligent weighted fusion of bimodal features.
[0088] As a further improvement, embodiments of this application also include:
[0089] Step 150: Perform linear projection on the multimodal fusion features, normalize and nonlinearly transform the projected fusion features, and perform residual connection between the nonlinearly transformed fusion features and the multimodal fusion features to obtain the final multimodal fusion features.
[0090] The features output by gated fusion are obtained by weighted concatenation and fusion of two cross-modal features. This results in redundant coupling and a chaotic feature distribution, while gated weighting introduces feature amplitude shifts. To address these issues, this application adds an nn.Linear projection layer to reconstruct the channel dimensions and calibrate the fused global multimodal features, decoupling redundant information between channels, unifying the feature distribution, and ensuring the multimodal fusion features adapt to the input distribution of subsequent networks. A proj_drop random deactivation layer is used to regularize the multimodal fusion features, suppressing overfitting risks from the attention fusion and gated weighting stages, and preventing the model from over-relying on local attention and gate weights. This layer only performs linear feature mapping without changing the feature dimensions, ensuring complete alignment of module input and output dimensions.
[0091] Specifically, linear projection and regularization are applied to the multimodal fusion features:
[0092] fused_feature = self.proj(fused_feature);
[0093] fused_feature = self.proj_drop(fused_feature);
[0094] In the formula, self.proj is the linear projection layer nn.Linear(dim, dim), and self.proj_drop is the Dropout layer.
[0095] Then, the regularized multimodal fused feature is subjected to layer normalization, multilayer perceptron (MLP) transformation, and residual connection to obtain the final multimodal fused feature:
[0096] output_feature = fused_feature + self.mlp(self.norm(fused_feature));
[0097] In the formula, self.norm = nn.LayerNorm(dim), self.mlp = nn.Sequential(nn.Linear(dim, dim * 4),nn.GELU( ),nn.Linear(dim * 4, dim));
[0098] LayerNorm stabilizes the distribution of MLP inputs, accelerates training convergence, and reduces internal covariate bias. As a feature extractor, the MLP performs independent nonlinear transformations and information mixing on features at each location through dimensionality increase and decrease, as well as nonlinear activation (GELU). Residual connections allow gradients to flow directly through the network, solving the gradient vanishing problem in deep networks and enabling deeper model training. It ensures that even if the MLP learns a "zero mapping," the output will not be worse than the input (identity mapping path), and the final output features maintain the same dimensions [B, L, C] as the input features.
[0099] The final multimodal fusion features can be directly fed into downstream network layers. For example, they can be fed into classification heads (such as face classification heads or object detection classification heads) and iteratively trained based on the loss between the classification categories output by the classification heads and the actual category labels.
[0100] To verify the effectiveness of the method in this application, visible light features and near-infrared features of the face were used as examples. Comparative experiments were conducted based on the MS1MV3 public face dataset. All models were trained using the ArcFace loss function, with 12 training rounds. No additional diversity constraint loss function was added to eliminate the interference of additional loss terms on the evaluation of the fusion effect.
[0101] The experiment uses the false rejection rate (FRR) under the ROC curve (Receiver Operating Characteristic Curve) as the evaluation metric, comparing the performance of the baseline method (without cross-modal attention fusion) and the proposed scheme under different false alarm rate (FAR) thresholds. The experimental results of the baseline method and the proposed scheme are as follows: Figure 2 As shown. Regarding the equal error rate (EER) metric, this scheme reduced the false alarm rate (FRR) from 0.3303% to 0.3245%, achieving a slight but stable improvement. The advantages of this scheme are particularly significant under extremely low false alarm rate conditions: at FAR=1e-4, the FRR decreased from 2.2432% to 1.9964%; at FAR=1e-5, the FRR decreased significantly from 7.0825% to 4.8621%, a relative reduction of 31.35%; at FAR=1e-6, the FRR decreased from 15.2034% to 7.9425%, a reduction of nearly 50%; and under the extreme condition of FAR=0, the FRR decreased from 22.7890% to 12.1579%, a similarly significant reduction.
[0102] The above data fully demonstrates that in high-security identification scenarios (such as payment authentication, access control and security) where extremely low false alarm rates are strictly required, this solution can significantly reduce the probability of false rejection and significantly improve user experience and system availability.
[0103] This application establishes attention dependencies in both the "visible light → near-infrared" and "near-infrared → visible light" directions through a bidirectional cross-modal attention mechanism, enabling features from the two modalities to interact fully and compensating for the incomplete information in unidirectional fusion. Combined with a multi-head attention design, the model can capture semantic associations between modalities in parallel in multiple different feature subspaces, further enhancing the discriminative power of the features.
[0104] This application also introduces a sparse attention mechanism, which calculates an importance score for each query position, retains the attention weights for only the Top-K key positions, and sets the rest to zero. When the feature sequence length is L, this reduces the computational complexity of attention from the standard O(L^2)^2. 2 The computational complexity is reduced to O(K·L), significantly decreasing floating-point operations and memory usage. This design enables the solution to achieve higher inference speeds in practical deployments, making it particularly suitable for edge computing scenarios with high real-time requirements, such as security monitoring and autonomous driving.
[0105] Traditional methods often use fixed weighting coefficients for modality fusion, which cannot adapt to dynamic changes in the input scene. This solution introduces a learnable gating network that can automatically output gating weights in the (0, 1) interval based on the input features, achieving dynamic weighted fusion per sample and even per spatial location. This data-driven adaptive fusion mechanism enables the model to maintain optimal feature representation in complex and ever-changing environments, significantly enhancing the model's scene robustness.
[0106] Furthermore, this application employs layer normalization, MLP, and residual connections to ensure that gradients can flow directly from deep layers back to shallow layers through the identity mapping path, fundamentally solving the gradient vanishing problem in deep networks and enabling the model to stack more fusion modules without training degradation.
[0107] Meanwhile, layer normalization standardizes the feature channels of each sample, stabilizes the distribution of MLP input, accelerates training convergence, and reduces sensitivity to the initial value of the learning rate.
[0108] Please refer to Figure 3 This application also provides a multimodal biometric fusion device, comprising:
[0109] The linear projection unit 210 is used to input the visible light features and near-infrared features extracted from the visible light image and near-infrared image respectively into the shared linear projection layer and perform query, key, and value triple mapping to obtain the query, key, and value of the visible light features and the query, key, and value of the near-infrared features respectively.
[0110] Attention weight calculation unit 220 is used to calculate the first cross-modal attention weight of visible light query in the near-infrared direction and the second cross-modal attention weight of near-infrared query in the visible light direction based on the query and key of visible light features and near-infrared features.
[0111] The cross-modal feature aggregation unit 230 is used to perform sparsification processing on the first cross-modal attention weight and the second cross-modal attention weight respectively to obtain the first sparse attention weight and the second sparse attention weight; the value of the near-infrared feature is weighted by the first sparse attention weight to obtain the first cross-modal feature, and the value of the visible light feature is weighted by the second sparse attention weight to obtain the second cross-modal feature.
[0112] The feature fusion unit 240 is used to concatenate the first cross-modal feature and the second cross-modal feature and input them into the gating network. The gating network adaptively learns to obtain dynamic gating weights, and the first cross-modal feature and the second cross-modal feature are fused through the gating weights to obtain multimodal fused features.
[0113] As a further improvement, the device also includes:
[0114] The projection and regularization unit is used to perform linear projection on the multimodal fusion features, normalize and nonlinearly transform the projected fusion features, and perform residual connection between the nonlinearly transformed fusion features and the multimodal fusion features to obtain the final multimodal fusion features.
[0115] This application also provides an electronic device, which includes a processor and a memory;
[0116] The memory is used to store program code and transfer the program code to the processor;
[0117] The processor is used to execute the multimodal biometric fusion method in the aforementioned method embodiments according to the instructions in the program code.
[0118] This application also provides a computer-readable storage medium for storing program code, which, when executed by a processor, implements the multimodal biometric fusion method described in the foregoing method embodiments.
[0119] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0120] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.
[0121] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0123] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0124] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0125] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the methods described in the various embodiments of this application through a computer device (which may be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0126] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multimodal biometric fusion method, characterized in that, include: The visible light features and near-infrared features extracted from the visible light image and near-infrared image, respectively, are input into a shared linear projection layer and mapped to query, key, and value triples to obtain the query, key, and value of the visible light features and the query, key, and value of the near-infrared features, respectively. Based on visible light and near-infrared features, the query and key calculations are performed to determine the first cross-modal attention weight in the near-infrared direction for visible light queries and the second cross-modal attention weight in the visible light direction for near-infrared queries. The first cross-modal attention weights and the second cross-modal attention weights are sparsified to obtain the first sparse attention weights and the second sparse attention weights. The first cross-modal feature is obtained by weighting the values of the near-infrared feature using the first sparse attention weight, and the second cross-modal feature is obtained by weighting the values of the visible light feature using the second sparse attention weight. The first cross-modal feature and the second cross-modal feature are concatenated and input into a gating network. The gating network adaptively learns to obtain dynamic gating weights. The first cross-modal feature and the second cross-modal feature are fused using the gating weights to obtain multimodal fusion features.
2. The multimodal biometric fusion method according to claim 1, characterized in that, The method further includes: The multimodal fusion features are linearly projected, and the projected fusion features are normalized and nonlinearly transformed. The nonlinearly transformed fusion features are then residually connected with the multimodal fusion features to obtain the final multimodal fusion features.
3. The multimodal biometric fusion method according to claim 1, characterized in that, The first cross-modal attention weights are sparsified to obtain the first sparse attention weights, including: Based on the first cross-modal attention weights, an importance score is calculated for each attention head; The importance scores are sorted in descending order, and the first cross-modal attention weights corresponding to the top K positions are retained. The attention weights of the remaining positions are reset to zero to obtain the first sparse attention weights.
4. The multimodal biometric fusion method according to claim 1, characterized in that, Visible light features extracted from visible light images are input into a shared linear projection layer for query, key, and value triple mapping to obtain the query, key, and value of the visible light features, including: The visible light features extracted from the visible light image are input into a shared linear projection layer and linearly mapped. During the mapping process, the channel dimension is expanded by 3 times to obtain the visible light feature tensor. The visible light feature tensor is split according to the channel dimension to obtain the query, key, and value of the visible light feature.
5. The multimodal biometric fusion method according to claim 1, characterized in that, The process of weighting the near-infrared feature values using the first sparse attention weight to obtain the first cross-modal feature, and weighting the visible light feature values using the second sparse attention weight to obtain the second cross-modal feature, further includes: Normalization and regularization operations are performed on the first sparse attention weight and the second sparse attention weight, respectively.
6. The multimodal biometric fusion method according to claim 1, characterized in that, The gated network includes two fully connected layers; The first fully connected layer is used to compress the channels of the spliced features and remove redundant features from the compressed spliced features using the ReLU activation function; The second fully connected layer is used to map the features output by the first fully connected layer into a single-channel weight map, and normalize the single-channel weight map using the Sigmoid function to obtain dynamic gating weights.
7. A multimodal biometric fusion device, characterized in that, include: The linear projection unit is used to input the visible light features and near-infrared features extracted from the visible light image and near-infrared image respectively into the shared linear projection layer and perform query, key, and value triple mapping to obtain the query, key, and value of the visible light features and the query, key, and value of the near-infrared features respectively. The attention weight calculation unit is used to calculate the first cross-modal attention weight in the near-infrared direction for visible light queries and the second cross-modal attention weight in the visible light direction for near-infrared queries based on the query keys of visible light features and near-infrared features. A cross-modal feature aggregation unit is used to perform sparsification processing on the first cross-modal attention weight and the second cross-modal attention weight respectively to obtain the first sparse attention weight and the second sparse attention weight; The first cross-modal feature is obtained by weighting the values of the near-infrared feature using the first sparse attention weight, and the second cross-modal feature is obtained by weighting the values of the visible light feature using the second sparse attention weight. The feature fusion unit is used to concatenate the first cross-modal feature and the second cross-modal feature and input them into the gating network. The gating network adaptively learns to obtain dynamic gating weights, and the first cross-modal feature and the second cross-modal feature are fused using the gating weights to obtain multimodal fused features.
8. The multimodal biometric fusion device according to claim 7, characterized in that, Also includes: The projection and regularization unit is used to perform linear projection on the multimodal fusion features, normalize and nonlinearly transform the projected fusion features, and perform residual connection between the nonlinearly transformed fusion features and the multimodal fusion features to obtain the final multimodal fusion features.
9. An electronic device, characterized in that, The device includes a processor and a memory; The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the multimodal biometric fusion method according to any one of claims 1-6 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code, which, when executed by a processor, implements the multimodal biometric fusion method according to any one of claims 1-6.