Face multimode image feature collaborative retrieval method

Through the methods of multimodal feature decoupling layer, collaborative perception dynamic fusion layer and cross-modal attitude metric learning layer, the problems of poor feature integration effect and insufficient stability in multimodal face image retrieval are solved, and efficient and accurate retrieval in complex scenarios are achieved.

CN120544253APending Publication Date: 2025-08-26CHONGQING UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510671308.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art has problems in the retrieval of multimodal face images that have poor feature integration effect, poor feature space consistency, and insufficient retrieval stability and accuracy. In particular, it is difficult to effectively distinguish the semantics of modality-specific information from the cross-modal shared by cross-modal in complex scenarios, and there is a lack of an effective cross-modal collaborative processing mechanism.

Method used

The multimodal feature decoupling layer, collaborative perception dynamic fusion layer and cross-modal attitude metric learning layer are used to extract modal exclusive features and shared identity features through the multimodal feature decoupling layer, and the collaborative perception dynamic fusion layer allocates modal weights according to the scene and compensates for occlusion features. The measurement optimization is performed in combination with the cross-modal attitude metric learning layer, and finally the matching result is output.

Benefits of technology

It realizes more efficient feature integration, enhances the consistency of feature space, ensures stable and accurate completion of search tasks in complex scenarios, breaks through the existing technology's collaborative processing bottleneck, and improves the robustness and accuracy of cross-modal retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544253A_ABST
    Figure CN120544253A_ABST
Patent Text Reader

Abstract

The invention provides a face multimode image feature collaborative retrieval method, and aims to solve the defects of the prior art in the aspects of multimode feature processing, complex scene adaptability and the like. The method comprises the following steps: firstly, acquiring multi-modal data such as visible light, infrared and depth images, pre-processing the multi-modal data, and inputting the pre-processed multi-modal data into a retrieval model comprising a multi-modal feature decoupling layer, a collaborative perception dynamic fusion layer and a cross-modal attitude measurement learning layer; wherein the decoupling layer separates modal exclusive features from shared identity features, eliminates semantic differences while keeping modal uniqueness, and enhances feature space consistency; when shielding is detected, the dynamic fusion layer allocates modal weights according to scenes, compensates shielding features and generates fusion features; and finally, outputting a result through metric learning optimization in combination with a hierarchical strategy. According to the scheme, through feature decoupling and dynamic fusion, the bottleneck of a traditional method under feature integration and complex scenes is broken through, and the retrieval accuracy and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face image retrieval, and specifically provides a method for collaborative retrieval of multimodal face image features. Background Art

[0002] Facial image retrieval, a key technology in computer vision and biometric recognition, plays a critical role in numerous scenarios, including security, access control, and identity authentication. With the continuous advancement of multi-sensor technology, cross-modal retrieval technology based on multimodal facial images, such as visible light, infrared, and depth, has gradually become a focus of industry attention. This technology aims to improve recognition performance in complex scenarios by integrating the advantages of information from different modalities. However, current research and application in this field still face numerous challenges. In the implementation process, the processing and fusion of multimodal features have limitations. Traditional methods struggle to effectively distinguish between modality-specific information and cross-modal shared semantics, and they are unable to fully address inter-modal differences, resulting in poor feature integration and poor feature space consistency. Furthermore, in complex and changing application scenarios, such as those with occlusion and suboptimal lighting conditions, existing technologies struggle to ensure retrieval stability and accuracy, lacking effective cross-modal collaborative processing mechanisms. In addition, existing methods do not fully utilize the prior knowledge of facial biometrics, and fail to effectively separate identity features from non-identity attribute features, resulting in insufficient generalization capabilities of the model in different scenarios and objects, affecting the overall performance and practical application effects of multimodal face retrieval technology.

[0003] Therefore, a new collaborative retrieval method of multimodal facial image features is urgently needed to solve the above problems. Summary of the Invention

[0004] In order to overcome the above-mentioned defects, the present invention is proposed to provide a solution or partial solution to the problems of poor feature integration effect, poor feature space consistency, and retrieval stability and accuracy.

[0005] The present invention provides a collaborative retrieval method for multimodal image features of faces, comprising: acquiring multimodal data, wherein the multimodal data includes visible light images, infrared images, and depth images; preprocessing the multimodal data, and inputting the preprocessed data into a constructed retrieval model to obtain retrieval results; wherein the retrieval model includes a multimodal feature decoupling layer, a collaborative perception dynamic fusion layer, and a cross-modal metric learning layer; inputting the preprocessed data into the constructed retrieval model to obtain retrieval results, specifically comprising: inputting the preprocessed multimodal data into the multimodal feature decoupling layer to extract modality-specific features and shared identity features; determining whether there is an occlusion area, and if an occlusion area is detected, using the collaborative perception dynamic fusion layer to assign modal weights according to the scene and compensate for the occlusion features to generate fusion features; performing metric optimization on the fusion features through the cross-modal metric learning layer, and outputting matching results in combination with a hierarchical retrieval strategy.

[0006] In a technical solution of the above-mentioned collaborative retrieval method for multimodal image features of faces, the multimodal feature decoupling layer includes: a dual-branch decoupling structure, which includes a modality-specific branch and a semantic sharing branch; the process of inputting the preprocessed multimodal data into the multimodal feature decoupling layer to extract modality-specific features and shared identity features includes: inputting the preprocessed multimodal data in parallel into the dual-branch decoupling structure of the multimodal feature decoupling layer; in the modality-specific branch, processing each modality data separately to obtain the exclusive features of each modality; in the semantic sharing branch, mapping the common semantic information of the shallow features of each modality to a shared space through a joint embedding mechanism to generate a shared identity feature vector that is consistent across modalities.

[0007] In one technical solution of the above-mentioned collaborative retrieval method for multimodal facial image features, in the modality-specific branch, each modality data is processed separately to obtain the specific features of each modality. The process includes: the preprocessed visible light image is subjected to the first four layers of convolution of ResNet18 to extract shallow features, and then the key area response is enhanced through the spatial attention module to output texture features; two-stream network parallel processing: global average pooling of the infrared image is performed to extract the overall temperature distribution vector, that is, the global feature; the nose bridge and eye socket hot areas are segmented through U-Net to output local features; global features are fused with local features; the depth image is input into the PointNet++ network, and three-dimensional geometric features are extracted through a hierarchical feature aggregation mechanism.

[0008] In one technical solution of the above-mentioned collaborative retrieval method for multimodal image features of faces, a collaborative perception dynamic fusion layer is used to assign modal weights according to the scene and compensate for occlusion features to generate fusion features. Specifically, the following steps are performed: a scene classifier is used to determine the current scene type and output the scene results; based on the scene results, weights are assigned to the specific features of each modality through a gating mechanism to form weighted modal features; and cross-modal priors are used to complement the visible light texture features of the occluded area to generate corrected fusion features.

[0009] In one technical solution of the above-mentioned collaborative retrieval method for multimodal facial image features, a scene classifier is used to determine the current scene type. The process of outputting the scene results includes: extracting image features through the MobileNetV3 backbone network, outputting the scene type probability distribution through the classification head, and taking the category corresponding to the maximum probability value as the scene type; at the same time, the model converts the scene classification results into initial reliability scores of each modality through the reliability mapping layer, and combines the predefined scene-modality offset rules to adjust the initial scores and normalize them, and finally outputs the scene type label and the reliability scores of each modality.

[0010] In one technical solution of the above-mentioned collaborative retrieval method of multimodal image features of faces, the process of using cross-modal priors to complement the visible light texture features of the occluded area and generate corrected fusion features includes: for the occluded area, extracting infrared features or depth features as cross-modal prior information, splicing the cross-modal prior information with the visible light feature map of the occluded area, and using the input conditions to generate an adversarial network; the adversarial network learns the mapping relationship between cross-modalities based on the prior information, and generates a visible light texture that semantically matches the occluded area; the facial structure prior loss function designed based on facial anatomical laws constrains the visible light texture based on facial anatomical laws, calculates the difference between the visible light texture and the facial structure prior, and the difference is back-propagated as a loss signal to the generator and discriminator of the adversarial network to drive the iterative update of the model parameters.

[0011] In one technical solution of the above-mentioned collaborative retrieval method of multimodal image features of faces, the fusion features are metrically optimized through a cross-modal metric learning layer, and the matching results are output in combination with a hierarchical retrieval strategy, specifically including: inputting the fusion features and the shared identity features into the cross-modal metric learning layer, optimizing the feature space through an improved triplet loss function so that the intra-class distance of multimodal samples with the same identity is smaller than the inter-class distance; filtering non-relevant samples based on the shared identity features to generate a candidate sample set; calculating the modality-specific feature weighted distance of the candidate samples according to the scene dynamic weight, sorting them in combination with the shared feature distance, and outputting the final retrieval results.

[0012] In a technical solution of the above-mentioned collaborative retrieval method of multimodal image features of faces, the construction process of the improved triplet loss function includes: each anchor sample matches at least one cross-modal positive sample and one heteromodal negative sample to construct a mixed modal triplet set; according to the modality of the anchor point and the positive sample, the adaptive margin value is calculated; according to the optimized loss function formula, the weighted loss of each triplet is calculated, and the losses of all triplets are summed to obtain the final loss value.

[0013] In one technical solution of the above-mentioned face multimodal image feature collaborative retrieval method, the process of calculating the adaptive margin value according to the modalities of the anchor point and the positive sample includes: calculating the feature distribution difference d between different modal pairs mod ; Adaptive Margin formula: m adapt =m0+γ·d mod , where m0 is the base Margin value, γ is the scaling factor, and d mod is the difference between the modes.

[0014] In one technical solution of the above-mentioned collaborative retrieval method for multimodal facial image features, the process of preprocessing the multimodal data includes: visible light images: equalizing the lighting and normalizing the skin color through the CLAHE algorithm; infrared images: removing noise using median filtering and normalizing based on the facial hot zone prior, where the facial hot zone prior is that the temperature of the nose bridge is greater than that of the cheek; depth images: removing outliers using voxel filtering and converting into a point cloud with a standard posture through a facial alignment algorithm.

[0015] The beneficial effects of the collaborative retrieval method for multimodal facial image features provided by the present invention are as follows: in the multimodal feature processing mechanism, the multimodal feature decoupling layer not only fully retains the unique information of each modality by separating the modality-specific features from the shared identity features, but also effectively eliminates the semantic differences between the modalities. Compared with traditional methods, it achieves more efficient feature integration, significantly enhances the consistency of the feature space, and lays an accurate feature foundation for cross-modal retrieval. In addition, for complex and changing application scenarios, the collaborative perception dynamic fusion layer can detect occluded areas in real time and dynamically adjust the weights of different modalities. This mechanism breaks through the bottleneck of the lack of collaborative processing capabilities of existing technologies, ensuring that retrieval tasks can still be completed stably and accurately in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The disclosure of the present invention will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. Furthermore, similar numbers in the drawings represent similar components, wherein:

[0017] Figure 1It is a flowchart of the main steps of inputting preprocessed data into a constructed retrieval model to obtain retrieval results according to an embodiment of the present invention.

[0018] Figure 2 3 is a schematic structural diagram of a multimodal feature decoupling layer according to an embodiment of the present invention. DETAILED DESCRIPTION

[0019] Some embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0020] Example 1

[0021] like Figure 1 As shown, a collaborative retrieval method for multimodal facial image features in an embodiment of the present invention mainly includes the following steps S1-S2.

[0022] Step S1: Acquire multimodal data, where the multimodal data includes a visible light image, an infrared image, and a depth image.

[0023] In this embodiment, the visible light image can present visual information such as the color, texture, and expression of the human face; the infrared image uses the thermal radiation characteristics of the object to form an image, and can capture facial features that cannot be presented by visible light at night or in low-light environments; the depth image records the distance information from each point on the surface of the face to the camera, constructing the three-dimensional spatial structure of the face.

[0024] Step S2: preprocessing the multimodal data and inputting the preprocessed data into the constructed retrieval model to obtain retrieval results;

[0025] The retrieval model includes a multimodal feature decoupling layer, a collaborative perception dynamic fusion layer, and a cross-modal attitude metric learning layer;

[0026] like Figure 1 As shown, step S2, inputting the preprocessed data into the constructed retrieval model, and obtaining the retrieval results specifically includes: step S21, inputting the preprocessed multimodal data into the multimodal feature decoupling layer to extract modality-specific features and shared identity features; step S22, judging whether there is an occlusion area. If an occlusion area is detected, the collaborative perception dynamic fusion layer is used to assign modal weights according to the scene and compensate for the occlusion features to generate fusion features; step S23, the fusion features are measured and optimized through the cross-modality metric learning layer, and the matching results are output in combination with the hierarchical retrieval strategy.

[0027] In this embodiment, modality-specific features refer to features determined by the unique physical properties or imaging characteristics of data from a specific modality. For example: visible light images: color (RGB values), texture details (such as skin pores and wrinkles), light intensity, expression changes and other optical characteristics; infrared images: thermal radiation distribution (facial temperature differences, such as the heat dissipation characteristics of the nose and forehead), stable thermal signals in day and night environments (not affected by visible light); depth images: three-dimensional spatial coordinates (X / Y / Z axis distance), facial surface geometry (such as nose bridge height, cheek convexity), and spatial depth relationships.

[0028] Shared identity features are features that carry "identity semantics" across cross-modal data. They are independent of the specific modality and reflect only the essential identity attributes of a face. Examples include geometric features such as face shape (oval, square), facial feature proportions (eye spacing, nasolabial angle), and facial contour curves; topological features include the relative positions of facial features (e.g., eyes below the eyebrows, nose centered), and the spatial relationships between key facial features (corners of the eyes, nose tip, corners of the mouth).

[0029] In one embodiment, the multimodal feature decoupling layer includes: a dual-branch decoupling structure, the dual-branch decoupling structure includes a modality-specific branch and a semantic sharing branch; step S21, the process of inputting the preprocessed multimodal data into the multimodal feature decoupling layer to extract modality-specific features and shared identity features includes: inputting the preprocessed multimodal data in parallel into the dual-branch decoupling structure of the multimodal feature decoupling layer; in the modality-specific branch, processing each modality data separately to obtain the exclusive features of each modality; in the semantic sharing branch, mapping the common semantic information of the shallow features of each modality to the shared space through a joint embedding mechanism to generate a shared identity feature vector that is consistent across modalities.

[0030] In this embodiment, modality-specific branch: an independent feature extraction path is designed for each modality data. For example, a dedicated convolutional layer or Transformer structure can be used to capture modality-specific attributes.

[0031] Semantic sharing branch: First, the early features of each modality (such as the low- and middle-level features of CNN) are spliced ​​or weightedly fused to obtain a hybrid feature representation. The hybrid features are then projected into a shared semantic space through a nonlinear transformation, with the following constraints: the different modal features of the same identity are close in distance in the shared space; the features of different identities are far apart in the shared space. Use contrastive learning or mutual information maximization methods to ensure that the shared feature vector maintains a consistent identity expression for all modalities. In this embodiment, shared space mapping is used to resolve the semantic gap between visible light and infrared and depth images, thereby improving the accuracy of cross-modal retrieval; in occluded or low-quality modal scenarios, shared features can be used to collaboratively compensate with other modality-specific features to improve robustness.

[0032] In an optional embodiment, in the modality-specific branch, the process of processing each modality data separately to obtain the specific features of each modality includes:

[0033] The preprocessed visible light image is convolved through the first four layers of ResNet18 to extract shallow features, and then the key area response is enhanced through the spatial attention module to output texture features. Compared with traditional convolutional networks, it can focus on key facial areas (such as facial features) more accurately, enhance the discriminative ability of texture features, avoid redundant interference of global features, and retain clearer local details, especially in low-resolution or blurred scenes.

[0034] Parallel processing of two-stream networks: Global average pooling of infrared images is performed to extract the overall temperature distribution vector, i.e., global features; U-Net is used to segment the nose bridge and eye socket hot zones and output local features; and global features are fused with local features. In this embodiment, a two-stream network is used to fuse the global temperature distribution with local features of hot zones such as the nose bridge and eye sockets, thus breaking through the limitations of traditional single global features. Global features retain the overall pattern of body temperature distribution, while local hot zone segmentation captures differentiated information of facial physiological features (such as blood vessel distribution), improving identity differentiation across lighting and camouflage scenarios.

[0035] The depth image is fed into the PointNet++ network, where 3D geometric features are extracted through a hierarchical feature aggregation mechanism. This embodiment uses PointNet++'s hierarchical aggregation mechanism to process depth data. Compared to traditional voxelization or projection methods, this allows for direct extraction of multi-scale 3D geometric features (such as surface curvature and spatial positional relationships) from point cloud structures, avoiding information loss caused by data structure conversion. This makes it more suitable for 3D feature modeling in scenes with complex occlusion or posture changes.

[0036] In one embodiment, a collaborative perception dynamic fusion layer is used to assign modal weights according to the scene and compensate for occlusion features to generate fusion features, specifically including: using a scene classifier to determine the current scene type and outputting the scene result; according to the scene result, weights are assigned to the specific features of each modality through a gating mechanism to form weighted modal features; and cross-modal priors are used to complement the visible light texture features of the occluded area to generate a corrected fusion feature.

[0037] In this embodiment, the scene classifier adopts a lightweight neural network or a multi-layer perceptron.

[0038] Scene types include: normal lighting with no occlusion, low lighting, partial occlusion (e.g., masks, sunglasses), severe occlusion (e.g., hands covering the face), and dynamic scenes (e.g., motion blur).

[0039] Dynamic gating mechanism: wi = σ(W · [F_shared||s] + b), where wi is the weight of the i-th modality (between 0 and 1), σ is the Sigmoid activation function, F_shared is the shared identity feature vector, s is the scene embedding vector output by the scene classifier, and W and b are learnable parameters.

[0040] Dynamic weight allocation strategy:

[0041] A. Scene Adaptation: For example, in low-light scenarios, the infrared modality weight is increased and the visible light weight is decreased. In mask-occluded scenarios, the depth modality weight is increased and the visible light weight of the occluded area is decreased.

[0042] B. Modality complementarity: If the quality of a modality is low (such as serious loss of depth image), its weight is reduced and compensation from other modalities is increased.

[0043] The formula for weighted modal feature F_w·eighted_i is:

[0044] F_w·eighted_i=wi×F_specific_i, where F_specific_i is the exclusive feature of the i-th modality.

[0045] Furthermore, occlusion detection can be performed by comparing visible light and depth images. If the visible light region exists but the depth value is missing, it is considered an occlusion. Alternatively, a pre-trained face segmentation model (such as BiSeNet) can be used to identify occlusions such as masks and sunglasses.

[0046] In an optional embodiment, a scene classifier is used to determine the current scene type, and the process of outputting the scene result includes: extracting image features through the MobileNetV3 backbone network, outputting the scene type probability distribution through the classification head, and taking the category corresponding to the maximum probability value as the scene type (such as low light probability 0.85, normal scene probability 0.15, and taking the category "low light" corresponding to the maximum probability value as the scene type label); at the same time, the model converts the scene classification results into initial reliability scores of each modality through a reliability mapping layer (such as a fully connected layer), and combines the predefined scene-modality offset rules to adjust the initial scores and normalize them, and finally outputs the scene type labels and the reliability scores of each modality.

[0047] Furthermore, MobileNetV3 is used as a feature extractor, using depthwise separable convolution and SE attention mechanisms to reduce the number of parameters while retaining the ability to capture multi-scale features. For example, low-level features can identify the edge texture of masks, while high-level features can determine the overall light intensity.

[0048] The classification head is a global average pooling + fully connected layer.

[0049] The predefined scene-modality offset rule can be designed as follows: in low-light scenes, the infrared modality score increases by 0.2 and the visible light modality score decreases by 0.3; in mask occlusion scenes, the depth modality score increases by 0.15 and the score of the visible light blocked area decreases by 0.5.

[0050] In an optional embodiment, the process of using the cross-modal prior to complete the visible light texture features of the occluded area to generate the corrected fusion features includes:

[0051] For occluded areas, such as the mouth covered by a mask or the eyes obscured by sunglasses, the thermal radiation characteristics of the infrared image or the 3D geometric features of the depth image are used as cross-modal priors. For example, the infrared modality can capture the temperature distribution of unobstructed areas (such as the heat zone at the corner of the eye), while the depth modality can provide the spatial structure of the occluded area (such as the height of the nose bridge). This prior information is concatenated with the feature map of the occluded area in the visible light image (such as a low-resolution blurred texture) in the channel dimension to form an input condition containing multimodal cues (such as visible light features + infrared thermal features + deep structural features). This provides a basis for cross-modal semantic association restoration in the Generative Adversarial Network (GAN). The GAN then uses this concatenated multimodal feature as input to learn cross-modal mapping relationships. The generator infers the plausible texture of the occluded area in visible light based on the infrared / depth prior information (such as generating the lip contour under a mask based on the deep structure), while the discriminator distinguishes the difference between the generated texture and the real texture. Through adversarial training, the generated visible light texture is semantically matched to the context of the occluded area (such as unoccluded cheeks and eyebrows), for example, generating skin texture consistent with facial skin color or shadows that conform to the lighting direction.

[0052] Prior knowledge based on facial anatomy (such as the relative positions of facial features and facial symmetry) is introduced, and a loss function is designed to constrain the generated visible light texture. Specifically, the difference between the generated texture and the facial structure prior is calculated: key point constraint: ensures that the positions of key points such as the corners of the eyes and the tip of the nose in the generated texture are consistent with the prior model; symmetry constraint: requires that the generated texture be approximately symmetrical on both sides (such as matching the curvature of the left and right cheeks); semantic consistency constraint: extracts features through a pre-trained VGG network to ensure that the high-level semantics of the generated texture (such as face shape category) is consistent with the real face. These differences are back-propagated to the generator and discriminator as loss signals, driving the iterative update of the model parameters to avoid generating unreasonable textures (such as distorted facial features or abnormal skin color), and ultimately outputting corrected fusion features that conform to anatomical laws.

[0053] This method uses cross-modal priors to compensate for the information loss in visible light occlusion areas, and combines facial structure constraints to ensure the rationality of generated textures, breaking through the limitations of traditional single-modal restoration (such as interpolation methods) that are prone to blurring or semantic errors.

[0054] In one embodiment, the fusion features are metrically optimized through a cross-modal metric learning layer, and matching results are output in combination with a hierarchical search strategy, specifically including:

[0055] A. Input the fusion features and the shared identity features into the cross-modal metric learning layer, optimize the feature space through the improved triplet loss function, so that the intra-class distance of multimodal samples with the same identity is smaller than the inter-class distance; filter non-related samples based on shared identity features (such as face shape and facial features ratio) to generate a candidate sample set, which can quickly exclude obviously irrelevant samples. Specifically, calculate the cosine distance of the shared identity features of the sample to be retrieved and all samples in the database, set a threshold (such as 0.6) to filter out samples with a distance less than the threshold, and generate a candidate sample set (usually compressed to 5%-10% of the original database). This can greatly reduce the amount of subsequent calculations, and is especially suitable for large-scale database scenarios of millions.

[0056] B. Calculate the modality-specific feature weighted distance of candidate samples based on the scene dynamic weight, sort them based on the shared feature distance, and output the final retrieval results.

[0057] In this embodiment, first, the reliability scores of each modality are normalized and used as weights. Then, for each candidate sample, the single modality-specific feature distance between the sample to be retrieved and the database sample is calculated, and the distances of each modality are weighted and summed according to the scene dynamic weight to obtain the total modality-specific feature distance d specific-total ; Use cosine distance or L2 distance to calculate the shared identity feature distance d between the sample to be retrieved and the candidate sample shared Finally, the total distance of modality-specific features and the shared feature distance are combined to form the final retrieval distance d final .

[0058] d final =α·d shared +(1-α)·d sepcific-total

[0059] Where α is a balance coefficient (usually set to 0.6-0.8) that adjusts the importance of identity semantics and modal details, prioritizing identity consistency. When α = 1, only shared feature distance is relied upon (suitable for extreme occlusion scenarios); when α = 0, only modal feature distance is relied upon (suitable for ideal scenarios without occlusion).

[0060] In one embodiment, the construction process of the improved triple loss function includes: each anchor sample is matched with at least one cross-modal positive sample and one heteromodal negative sample to construct a set of mixed-modal triples; according to the modality of the anchor point and the positive sample, an adaptive margin value is calculated; according to the optimized loss function formula, the weighted loss of each triple is calculated, and the losses of all triplets are summed to obtain the final loss value. In this embodiment, adaptive adjustment of modal weights is introduced to apply smaller weights to the feature distance calculation of low-quality modalities (such as visible light in occluded scenes) to avoid noise interference in the optimization direction.

[0061] In one embodiment, the process of calculating the adaptive margin value according to the modalities of the anchor point and the positive sample includes: calculating the feature distribution difference d between different modality pairs mod ;

[0062] Adaptive Margin formula: m adapt =m0+γ·d mod , where m0 is the base Margin value, γ is the scaling factor, and d mod is the difference between the modes.

[0063] Furthermore, the improved triplet loss function is:

[0064]

[0065] Among them, A represents the anchor point, P represents the cross-modal positive sample, N represents the heteromodal negative sample, d(A, P) and d(A, P) are the distances between the anchor point and the positive sample and the negative sample respectively, m adapt It is the margin value that is dynamically adjusted based on the modal pair of the anchor point and the positive sample; core constraint: d(A, P)+m adapt <d(A,N);

[0066] Combined with the modality-specific feature weights output by the collaborative perception dynamic fusion layer, the loss function is further optimized:

[0067] L weighted_triplet =∑ (A,P,N)∈T α mod(A) α mod(P) max(0,d(A,P)-d(A,N)+m adapt )

[0068] Among them, (A, P, N) represents a triple, T is the set of all triples, α mod (A) and α mod (P) are the weights of the corresponding modalities of the anchor point A and the positive sample P respectively.

[0069] In one embodiment, the process of preprocessing the multimodal data includes: visible light images: equalizing illumination and normalizing skin color through the CLAHE algorithm; infrared images: removing noise through median filtering and normalizing based on a facial hot zone prior, where the facial hot zone prior is that the bridge of the nose temperature is greater than the cheek temperature; depth images: removing outliers through voxel filtering and converting into a point cloud with a standard posture through a facial alignment algorithm.

[0070] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the original technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. A method for collaborative retrieval of multimodal face image features, characterized in that: include: Acquiring multimodal data, the multimodal data including visible light images, infrared images, and depth images; Preprocessing the multimodal data and inputting the preprocessed data into the constructed retrieval model to obtain retrieval results; The retrieval model includes a multimodal feature decoupling layer, a collaborative perception dynamic fusion layer, and a cross-modal attitude metric learning layer; The preprocessed data is input into the constructed retrieval model to obtain the retrieval results, which specifically include: inputting the preprocessed multimodal data into the multimodal feature decoupling layer to extract modality-specific features and shared identity features; judging whether there is an occluded area. If an occluded area is detected, the collaborative perception dynamic fusion layer is used to assign modal weights according to the scene and compensate for the occlusion features to generate fusion features; the fusion features are measured and optimized through the cross-modal attitude metric learning layer, and the matching results are output in combination with the hierarchical retrieval strategy.

2. The method according to claim 1, characterized in that The multimodal feature decoupling layer includes: a dual-branch decoupling structure, wherein the dual-branch decoupling structure includes a modality-specific branch and a semantic sharing branch; The process of inputting the pre-processed multimodal data into the multimodal feature decoupling layer to extract modality-specific features and shared identity features includes: The preprocessed multimodal data are input in parallel into the dual-branch decoupling structure of the multimodal feature decoupling layer; In the modality-specific branch, each modality data is processed separately to obtain the unique features of each modality; In the semantic sharing branch, the common semantic information of shallow features of each modality is mapped to the shared space through a joint embedding mechanism to generate a shared identity feature vector that is consistent across modalities.

3. The method according to claim 2, characterized in that In the modality-specific branch, the process of processing each modality data separately to obtain the unique features of each modality includes: The preprocessed visible light image is convolved through the first four layers of ResNet18 to extract shallow features, and then the key area response is enhanced through the spatial attention module to output texture features; Two-stream network parallel processing: Global average pooling of infrared images is performed to extract the overall temperature distribution vector, i.e., global features; U-Net is used to segment the nose bridge and eye socket hot spots and output local features; and the global features are fused with the local features. The depth image is input into the PointNet++ network, and 3D geometric features are extracted through a hierarchical feature aggregation mechanism.

4. The method according to claim 1, wherein The collaborative perception dynamic fusion layer is used to assign modal weights according to the scene and compensate for occlusion features. The generated fusion features include: Use the scene classifier to determine the current scene type and output the scene result; According to the scene results, weights are assigned to the specific features of each modality through a gating mechanism to form weighted modal features; The cross-modal prior is used to complement the visible light texture features of the occluded area and generate the corrected fusion features.

5. The method according to claim 4, characterized in that The process of using the scene classifier to determine the current scene type and outputting the scene results includes: Image features are extracted through the MobileNetV3 backbone network, and the classification head outputs the scene type probability distribution, taking the category corresponding to the maximum probability value as the scene type; at the same time, the model converts the scene classification results into the initial reliability scores of each modality through the reliability mapping layer, and combines the predefined scene-modality offset rules to adjust the initial scores and normalize them, and finally outputs the scene type label and the reliability scores of each modality.

6. The method according to claim 5, characterized in that The process of using cross-modal priors to complement the visible light texture features of the occluded area and generate the corrected fusion features includes: For the occluded area, infrared features or depth features are extracted as cross-modal prior information. This cross-modal prior information is concatenated with the visible light feature map of the occluded area and used as input to generate an adversarial network. The adversarial network uses this prior information as a condition to learn the mapping relationship between the two modalities and generate a visible light texture that semantically matches the occluded area. The facial structure prior loss function designed based on facial anatomical laws constrains the visible light texture based on facial anatomical laws, calculates the difference between the visible light texture and the facial structure prior, and the difference is back-propagated as a loss signal to the generator and discriminator of the adversarial network to drive the iterative update of the model parameters.

7. The method according to claim 4, characterized in that The fusion features are metrically optimized through the cross-modal attitude metric learning layer, and the matching results are output in combination with the hierarchical retrieval strategy, specifically including: Inputting the fused features and the shared identity features into a cross-modal metric learning layer, optimizing the feature space by using an improved triplet loss function so that the intra-class distance of multimodal samples with the same identity is smaller than the inter-class distance; Non-relevant samples are filtered based on shared identity features to generate a candidate sample set; the modality-specific feature weighted distance of the candidate samples is calculated according to the scene dynamic weight, and they are sorted based on the shared feature distance to output the final retrieval results.

8. The method according to claim 7, characterized in that The construction process of the improved triplet loss function includes: Each anchor sample matches at least one cross-modal positive sample and one heteromodal negative sample to construct a set of mixed-modal triples; Calculate the adaptive margin value based on the modality of the anchor point and the positive sample; According to the optimized loss function formula, the weighted loss of each triple is calculated, and the losses of all triplets are summed to obtain the final loss value.

9. The method according to claim 8, wherein the process of calculating the adaptive margin value according to the modality of the anchor point and the positive sample comprises: Calculate the feature distribution difference d between different modality pairs mod ; Adaptive Margin formula: m adapt =m0+γ·d mod , where m0 is the base Margin value, γ is the scaling factor, and d mod is the difference between the modes.

10. The method according to any one of claims 1 to 9, characterized in that: The process of preprocessing the multimodal data includes: Visible light images: Light equalization and skin color normalization are achieved through the CLAHE algorithm; Infrared image: Use median filtering to remove noise and perform normalization based on the facial hotspot prior, where the temperature of the nose bridge is greater than that of the cheek. Depth image: Use voxel filtering to remove outliers and convert it into a point cloud with a standard pose through a face alignment algorithm.

Citation Information

Cited By

  • Method and device for eliminating neck and shoulder infrared thermal image interference and medium

    CN121147152A

  • A method, device and medium for excluding interference of infrared thermal image of neck and shoulder

    CN121147152B

  • Unmanned aerial vehicle inspection system multi-modal data fusion and intelligent analysis platform and method for wind power plant

    CN121479645A

  • Gate identity recognition acceleration method based on deep learning

    CN121524929A