Multi-modal target re-identification method based on modal perception graph reasoning
By constructing a model perceptual graph and designing a selective graph node exchange strategy in multimodal target recognition, the problem of the difference in local feature quality in different modes is solved, and the missing modal information is restored through the missing modal graph inference strategy, achieving more efficient information interaction and recognition performance.
Patent Information
- Application Number
- CN202510171687.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
The existing multimodal target re-identification technology ignores the quality differences between local features of different modalities, resulting in the introduction of noise in features and the reduction of recognition effects, and lacks effective solutions to the problem of missing modalities.
A multimodal target re-identification method based on modal perceptual graph inference is proposed. By constructing a modal perceptual graph and designing a selective graph node exchange strategy, the structural relationship between local features is captured and the impact of low-quality features is reduced. At the same time, the missing modal graph inference strategy is introduced to restore the missing modal information.
It effectively improves the information interaction and feature representation quality in multimodal target re-identification, can restore missing information in the absence of modality, significantly improves recognition performance, and surpasses the existing state-of-the-art methods.
Smart Images

Figure CN120107733A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision and multimodal target re-identification technology, and in particular to a multimodal target re-identification method based on modal perception graph reasoning. Background Art
[0002] Multimodal target re-identification is a task that aims to identify the identity of the same target by combining multiple data modalities. Common multimodal data in this task include RGB images, near-infrared images (NIR) and thermal infrared images (TIR). Combining multimodal data for re-identification tasks can effectively overcome the limitations of single-modal data affected by the environment. The development of multimodal target re-identification plays an important role in monitoring security, intelligent transportation and unmanned driving.
[0003] At present, there has been significant progress in the relevant research in the field of multimodal target re-identification. Many studies have provided more comprehensive modal interactions by fusing multiple data sources to improve the re-identification performance. Most of these methods adopt a local multimodal fusion strategy, which effectively integrates the advantageous information from different modalities by progressively fusing different spectral features (such as RGB, NIR, and TIR). One of the methods proposed a progressive fusion network, which aims to efficiently fuse different spectral features and improve the accuracy of target re-identification. Another method introduced a relationship-based embedding module to embed global information into fine-grained local features, enhancing the feature representation capability in multimodal tasks. In addition, there are some existing methods that perform multimodal re-identification through target-centric tokens, and use specific modules to select and aggregate multimodal features. Other methods propose a token arrangement module that can effectively align multispectral images and use global tokens to perceive local information of other modalities.
[0004] Although these methods have achieved good results in extracting and utilizing local information, a common problem faced by existing technologies is that they often ignore the quality differences between local features of different modalities. For example, some methods divide features into multiple parts by random segmentation or based on target parts, but these strategies may lead to the introduction of noise in the features, thus affecting the final recognition effect. In order to overcome these technical problems, future research may need to consider the quality differences of local features of each modality more accurately to avoid noise problems caused by inappropriate feature division.
[0005] At the same time, graph-based methods have also been widely used in target re-identification, which can model the relationship between single spectral images and promote the improvement of representation. For example, a study proposed a graph-based character signature that integrates detailed character descriptions and visual features into the graph. Another method proposed a graph convolutional network, which first constructed an information content structure to capture local information and then input it into the graph convolutional network for prediction, achieving good results. Another method aims to solve the problem of single-modal pedestrian re-identification and provides an effective solution that can effectively integrate local, global and structured feature representations. In addition, there are methods that propose to use adaptive graph attention convolution to learn the contribution matrix of local information; there are also studies that propose a polymorphic mask wavelet graph convolutional network to decouple the content and degradation features of cross-modal images; some studies propose edge weight embedding graph convolutional networks to embed human joints and bones into the feature representation of target re-identification.
[0006] Although these methods have achieved good results in unimodal object re-identification, they mainly emphasize the interaction between local and global features, while ignoring the quality difference of local features and the complementary information across modalities. In addition, in real-world scenarios, the problem of missing modalities is common and often unavoidable. Some studies have proposed image reconstruction and feature reconstruction methods to compensate for the missing spectral images. Among them, some methods use zero padding to reconstruct information.
[0007] However, existing works mainly focus on reconstruction at the basic feature or image level, while ignoring the structural relationships and dependencies between missing modality features. Summary of the invention
[0008] Purpose of the invention: The purpose of the present invention is to address the deficiencies in the prior art and to provide a multimodal target re-identification method based on modal-aware graph reasoning. Graph reasoning is used to perform modal interaction and solve the missing problem, taking into account the global and local features in multimodal target re-identification. The modal-aware graph reasoning network MGRNet of the present invention can effectively improve information interaction and restore missing modalities in multimodal target re-identification.
[0009] Technical solution: A multimodal target re-identification method based on modal perception graph reasoning of the present invention constructs a modal perception graph reasoning network MGRNet, uses the modal perception graph reasoning network MGRNet to realize the information interaction of multimodal images, and restores the missing modality in multimodal target re-identification; corresponding modal image I (m) The specific execution process after inputting the modality perception graph reasoning network MGRNet is as follows:
[0010] Step 1: Use a multi-branch backbone network based on visual transformer ViTs to perform initial feature X on each modality image. (m) Extraction of X(m) Represents the mth modal image I (m) The initial feature representation of ; m = {N, R, T} represents three different modalities of near infrared image NIR, RGB image and thermal infrared image TIR respectively;
[0011] Step 2: The modal interaction graph reasoning strategy GRMI considers both local and global features, captures local details and mitigates the impact of low-quality local features; specifically, it includes the following:
[0012] Step 2.1: To obtain local multimodal information and learn the structural relationship between local features, we conduct modal perception map learning, that is, we construct a modal perception map for each modality image separately. The modal perception map of the mth modality image is G(V (m) , E (m) ), V (m) represents the nodes in the modality perception graph, E (m) Represents the edge in the modal perception graph; this can fully encode the local features of multimodal information; the local feature here is each patch, and the global feature is the combination of all patches; and outputs the adjacency matrix A of the local feature relationship;
[0013] Step 2.2: To further promote the interaction between the obtained local features and reduce the impact of low-quality local features between modalities, a selective graph node exchange strategy is designed, that is, global tokens are introduced to select the K minimum values of the edges in each graph using the Top-K method, low-quality patches are found, and the local feature nodes after the three modal enhancements are obtained.
[0014] Step 2.3: Local perception graph reasoning: A multi-layer local perception graph reasoning module is used to learn the local patch representation of multimodal data to obtain better quality node patches and obtain the local patch representation F. l (m) ;.
[0015] That is, local features are obtained;
[0016] Step 3: To obtain better global information, we increase the interaction between global and local information and capture richer node representations. We use a global-aware multi-head attention module to aggregate rich local information from different regions (different patches) into global features.
[0017] For missing modal images, the missing modal graph inference strategy GRMM is used to complete the missing modal information, reduce the differences between multimodal data, and restore the missing modal information;
[0018] For the constructed modality-aware graph reasoning network MGRNet, a comprehensive loss function L is used for training optimization.
[0019] In order to capture the unique information of each modality, a multi-branch backbone network is used when extracting the initial features in step 1, so as to extract rich modal features; at the same time, in order to retain the specific information of each modality, the multi-branch backbone network is a non-shared structure;
[0020] The initial features obtained The expression is:
[0021]
[0022] in, represents the input mth modality image, and They represent patch tokens and class tokens of the input image respectively; P = 128 represents the number of local tokens, and D is the dimension of the embedded token; the subscript l represents local, which is the patch token of the input image, indicating local; the subscript g represents global, which is the class token of the input image, indicating global.
[0023] Furthermore, the modal perception map in the learning process of the modal perception map in step 2 is G(V (m) , E (m) ), node V (m) Represents the local feature set of the image, that is Represents the feature set of all patches;
[0024] The edge E of the graph (m) Connect different patches using the adjacency matrix A (m) To represent the structural relationship between local features;
[0025] Since edge learning is dynamic, relationships are constructed by calculating the Euclidean distance:
[0026]
[0027] The Euclidean distance between the i-th and j-th patch token feature vectors of the m-th modality is used to build the relationship between different patches;
[0028] in and Represent the feature vectors of the i-th and j-th token respectively;
[0029] To learn more effective graphs, two learnable parameters α and t are introduced to calculate The formula is as follows:
[0030] σ represents the sigmoid activation function;
[0031] The adjacency matrix of the ith and jth patch tokens of the mth modality is used to represent the structural relationship between local features. α is the offset and t is the scaling factor, which is mainly used to optimize the structure of the graph during the learning process.
[0032] Furthermore, the specific method for performing selective graph node exchange in step 2 is as follows:
[0033] First, the Euclidean distance between the global feature and each local feature is calculated, and the relationship strength W between the global feature and the local feature is calculated. (m) , W (m) The calculation formula is as follows:
[0034]
[0035] Then the W of all modes (m) Composition:
[0036] cdist() represents Euclidean distance calculation. Represents the strength of the relationship between the Pth local feature and the global feature of the m-mode;
[0037] Then, use W (m) The low-quality nodes obtained in the first step are filtered out. This operation helps to obtain patches of lower quality. Since the local features in different modalities have different performances and qualities, the patches between the three modal images of RGB, NIR and TIR will be exchanged according to their respective qualities. For example, the low-quality nodes of the RGB modality will be replaced by the mean features of NIR and TIR. Finally, the updated local feature nodes are obtained.
[0038] Here, we replace the inferior patches with better patches from other modalities while considering the correlation between patches; the updated local feature nodes The definition is as follows:
[0039]
[0040]
[0041]
[0042] It means using the features of RGB modality to exchange the poor quality features of NIR modality; It means using the features of RGB modality to exchange the poor quality features of TIR modality; It means using the mean of NIR and TIR features to exchange the poor features of RGB modality.
[0043] To avoid exchanging low-quality patches, before replacing low-quality patches with better patches from other modalities, it is necessary to first determine whether the swapped patch is a low-quality patch. If it is a low-quality patch, reinitialize the current patch to an all-zero matrix, and learn the feature expression of the current patch through the local neighbors of the current patch; if it is not a low-quality patch, swap the patch in the above manner. Pre-judgment not only helps to learn information from other modalities, but also reduces the impact of low-quality local features by considering local and global features.
[0044] Furthermore, after completing the selective graph node exchange in step 2.2, for the obtained enhanced modality-aware graph node representation, a multi-layer local-aware graph reasoning module is used to learn the local patch representation of multimodal data to obtain better quality node patches;
[0045] The message propagation rules are defined as follows:
[0046]
[0047]
[0048] In the above formula, is a learnable transformation matrix, Abbreviated as Indicates how the next layer of graph convolution is updated. represents a layer of graph convolution, represents the last layer of graph convolution, represents the adjacency matrix of the ith and jth patch tokens of the mth modality, yes A general term for .
[0049] Furthermore, step 3 uses the global perception multi-head attention module to increase the interaction between global and local information. First, we assume Represents the patches obtained through the above local perception graph reasoning network;
[0050] Then, a linear projection layer is applied to obtain the global labeled query matrix And for local tokens Different linear projection layers are used to obtain the key matrix K(m) Sum value matrix V (m) ;
[0051] Next, we use H heads of multi-head attention to aggregate local information from different patches into a global feature representation. This interaction operation is defined in the hth head as follows:
[0052]
[0053] in is the scaling factor; It represents the output features obtained by the h-th head, where the value of h ranges from 1 to H. To aggregate the information of all heads, a concatenation operation is used to obtain a new class label:
[0054]
[0055] Finally, the features of all modes Splice to get the fused feature Z.
[0056] Furthermore, if a mode is missing, the missing mode graph inference strategy GRMM is used to compensate for the missing modal information and reduce the difference between the modes. That is, feature reconstruction is first applied to enhance the feature representation, and then structure reconstruction is used to learn the relationship between the corresponding tags. The specific method is as follows:
[0057] Assuming that there is only RGB modality, we use the RGB modality to recover the labels of NIR and TIR modalities. First, we use the Euclidean distance to calculate the original structural relationship between the true labels of each modality:
[0058]
[0059]
[0060] in and Represent the feature vectors of the i-th and j-th token respectively;
[0061] Next, construct the dynamically reconstructed structural relationship A R2N , As shown below:
[0062] A R2N =1-σ(t N D (R) ),
[0063] A R2T =1-σ(t T D (R) ),
[0064] where t N and tT are two learnable hyperparameters;
[0065] Then, a hierarchical GCN is used to propagate information to recover the feature H R2N and H R2T . This recovery process is defined as:
[0066]
[0067] in Indicates the number of the GCN layer,
[0068] and is the trainable transformation matrix of each layer, m = {N, T};
[0069] The outputs after passing through L layers of GCN are Abbreviated as H R2N , H R2T ;
[0070] This GRMM strategy not only solves the limitation of the receptive field of the CNN architecture by performing convolution on the graph structure, but also effectively solves the missing modality problem in the re-identification task.
[0071] Furthermore, the expression of the comprehensive loss function L is:
[0072]
[0073] Among them, L MM , L ce , L 3m They are GRMM loss, cross entropy loss, and multimodal margin loss, respectively. is a balancing hyperparameter;
[0074] GRMM loss L MM The calculation formula is: L MM =L R +L N +L T ;
[0075] L R , L N and L represent the reconstruction losses of RGB, NIR, and TIR modalities, respectively;
[0076] In order to promote the feature reconstruction capability, the restored features are constrained by the mean square error (MSE) loss, using the token features and their structural relationships, which is defined as:
[0077]
[0078]
[0079]
[0080] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0081] 1. This paper proposes a Modality-aware Graph Reasoning Network (MGRNet), which can jointly capture the important structural relationships and complementary information between different modalities, thereby solving the modality interaction and missing problems in multimodal object ReID.
[0082] 2. The present invention introduces multiple modal perception graphs to incorporate the structural information of local features, and designs a selective graph node exchange operation to effectively alleviate the impact of low-quality local features.
[0083] 3.MGRNet itself has the ability to reconstruct missing modal features based on the structural relationship between modalities and minimize the differences between modalities.
[0084] 4. Experimental results on four benchmark datasets show that the proposed MGRNet surpasses the existing state-of-the-art methods in the multimodal object ReID task. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 It is a schematic diagram of the training process of the present invention;
[0086] Figure 2 It is a schematic diagram of the overall network architecture of the present invention;
[0087] Figure 3 A schematic diagram of an embodiment is shown. DETAILED DESCRIPTION
[0088] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.
[0089] like Figure 1 As shown in (a) in the figure, for the global-based existing technical methods, current multimodal methods often focus on global feature extraction and use the complementary information between different modalities, which cannot mine local clues for multimodal re-identification tasks. For the part-based existing multimodal methods, simply dividing the features into multiple parts by random partitioning or extracting the location of the target part according to the visual Transformer (ViT) fails to consider the quality difference between local features of different modalities and is ineffective for low-quality local features.
[0090] like Figure 1As shown in (b), the multimodal target re-identification method based on modal perception graph reasoning of the present invention proposes a new modal perception graph reasoning network (MGRNet), which can effectively improve information interaction and restore missing modalities in multimodal target re-identification. Specifically, a modal perception graph is first constructed to promote the extraction of important local details by modeling the relationship between patches; then a selective graph node exchange operation is designed to reduce the impact of low-quality local features by considering local and global information at the same time, thereby enhancing discriminative information. The exchange of local features of the present invention can solve problems related to low-quality features, thereby improving the effectiveness of modal representation; finally, the exchanged modal perception graph is input into the local perception graph reasoning module to realize multimodal information propagation and generate reliable feature representation.
[0091] The key to the graph reasoning method of the present invention is to restore the features of the missing modalities by exploiting the structural relationship of the modalities, thereby minimizing the unnecessary differences between different modalities; this restoration loss aims to reduce the gap between the reconstructed representation and the true representation and minimize the difference between the modalities. In the test phase, it can be directly used to generate the features of the missing modalities, such as Figure 1 (b) as shown.
[0092] In general, the MGRNet network model of the present invention can capture rich and fine-grained local features on the basis of considering local and global information, while reducing the impact of low-quality local tokens, promoting the interaction of target information, and using graph reasoning to solve the modal interaction and missing problems in multimodal target re-identification, and has achieved excellent performance on common multimodal target re-identification datasets.
[0093] like Figure 2 As shown, the multimodal target re-identification method based on modal perception graph reasoning of this embodiment constructs a modal perception graph reasoning network MGRNet, and uses the modal perception graph reasoning network MGRNet to realize information interaction of multimodal images and restore missing modalities in multimodal target re-identification.
[0094] Corresponding modality image I (m) The specific execution process after inputting the modality perception graph reasoning network MGRNet is as follows:
[0095] Step 1: Use a multi-branch backbone network based on visual transformer ViTs to perform initial feature X on each modality image. (m) Extraction of X (m) Represents the mth modal image I (m) The initial feature representation of ; m = {N, R, T} represents the near infrared image NIR, RGB image and thermal infrared TIR image modalities respectively;
[0096] Step 2: The modal interaction graph reasoning strategy GRMI considers both local and global features, captures local details and mitigates the impact of low-quality local features; specifically, it includes the following:
[0097] Step 2.1: Modal perception map learning, that is, constructing a modal perception map for each modality image separately. The modal perception map of the mth modality image is G(V (m) , E (m) ), V (m) represents the nodes in the modality perception graph, E (m) represents the edges in the modality-aware graph;
[0098] Step 2.2, selective graph node exchange, that is, introduce global tokens and use the Top-K method to select the K minimum values of the edges in each graph, find low-quality patches, and obtain the local feature nodes after the three modal enhancement updates.
[0099] Step 2.3: Local perception graph reasoning: A multi-layer local perception graph reasoning module is used to learn the local patch representation of multimodal data to obtain better quality node patches and obtain the local patch representation F. l (m) ;
[0100] Step 3: Use the global-aware multi-head attention module to aggregate the rich local information from different regions into global features;
[0101] For missing modal images, the missing modal graph inference strategy GRMM is used to complete the missing modal information, reduce the differences between multimodal data, and restore the missing modal information;
[0102] For the modality perception graph reasoning network MGRNet constructed above, the comprehensive loss function L is used for training optimization. During the training process, Figure 1 As shown in (a), due to the quality differences of local features of different modes, the data is first segmented to obtain more detailed local information, and then the information is exchanged. Figure 1 As shown in (b), when the thermal infrared image (TIR) is missing in the test phase, the graph reasoning trained by the constraints of modal and structural information is used to restore the features, combining the existing RGB and near-infrared features (nodes) and their relationships (edges).
[0103] The multi-branch backbone network used in step 1 of this embodiment to extract the initial features is non-shared;
[0104] The initial features obtained The expression is:
[0105]
[0106] in, represents the input mth modality image, and Represent the patch tokens and class tokens of the input image respectively.
[0107] For step 2, the modal perception map in the process of modal perception map learning is G(V (m) , E (m) ), node V (m) Represents the local feature set of the corresponding modality image;
[0108] use Represents the feature set of all patches, P is the number of local tokens;
[0109] The edge E of the graph (m) Connect different patches and use the adjacency matrix A (m) To represent the structural relationship between local features;
[0110] Since edge E (m) The learning is dynamic, building relationships by calculating the Euclidean distance:
[0111]
[0112] represents the Euclidean distance between the i-th and j-th patch token feature vectors of the m-th modality; and Represent the feature vectors of the i-th and j-th token respectively;
[0113] To learn more effective graphs, two learnable parameters α and t are introduced to calculate The formula is as follows:
[0114] α represents the sigmoid activation function;
[0115] Represents the adjacency matrix between the ith and jth patch tokens of the mth modality.
[0116] In this embodiment, the specific method for performing selective graph node exchange in step 2 is as follows:
[0117] First, the Euclidean distance between the global feature and each local feature is calculated, and the relationship strength W between the global feature and the local feature is calculated. (m) , W(m) The calculation formula is as follows:
[0118]
[0119] Then the W of all modes (m) Composition:
[0120] cdist() represents Euclidean distance calculation. Represents the strength of the relationship between the Pth local feature and the global feature of the m-mode;
[0121] Then, use W (m) The inferior nodes obtained in the first step are filtered out to obtain the updated local feature nodes; here, the inferior patches are replaced by better patches of other modes, and the correlation between patches is considered, so the updated local feature nodes The definition is as follows:
[0122]
[0123]
[0124]
[0125] It means using the features of RGB modality to exchange the poor quality features of NIR modality; It means using the features of RGB modality to exchange the poor quality features of TIR modality; It means using the mean of NIR and TIR features to exchange the poor features of RGB modality.
[0126] In this embodiment, before using a better patch of other modalities to replace a low-quality patch, it is necessary to determine whether the exchanged patch is a low-quality patch. If it is determined to be a low-quality patch, the current patch is reinitialized to be an all-zero matrix, and the feature expression of the current patch is learned through the local neighbors of the current patch; if it is determined not to be a low-quality patch, the patch is exchanged.
[0127] In this embodiment, after completing the selective graph node exchange in step 2.2, for the obtained enhanced modality-aware graph node representation, a multi-layer local-aware graph reasoning module is used to learn the local patch representation of multimodal data to obtain better quality node patches;
[0128] The message propagation rules are defined as follows:
[0129]
[0130]
[0131]
[0132] In the above formula, is a learnable transformation matrix, Abbreviated as Indicates how the next layer of graph convolution is updated. represents a layer of graph convolution, Represents the last layer of graph convolution.
[0133] In this embodiment, when step 3 uses the global perception multi-head attention module to increase the interaction between global and local information, first assume that Represents the patches obtained through the above local perception graph reasoning network;
[0134] Then, a linear projection layer is applied to obtain the global labeled query matrix And for local tokens Different linear projection layers are used to obtain the key matrix K (m) Sum value matrix V (m) ;
[0135] Next, we use H heads of multi-head attention to aggregate local information from different patches into a global feature representation. This interaction operation is defined in the hth head as follows:
[0136]
[0137] in is the scaling factor;
[0138] represents the output features obtained by the h-th head;
[0139] To aggregate the information of all headers, a concatenation operation is used to obtain a new class tag:
[0140]
[0141] Finally, the features of all modes Splice to get the fused feature Z.
[0142] In this embodiment, if a mode is missing, the missing mode graph inference strategy GRMM is used to compensate for the missing mode information and reduce the difference between the modes. That is, feature reconstruction is first applied to enhance the feature representation, and then structure reconstruction is used to learn the relationship between the corresponding tags. The specific method is:
[0143] Assuming that there is only RGB modality, we use the RGB modality to recover the labels of NIR and TIR modalities. First, we use the Euclidean distance to calculate the original structural relationship between the true labels of each modality:
[0144]
[0145]
[0146] in and Represent the feature vectors of the i-th and j-th token respectively;
[0147] Next, construct the dynamically reconstructed structural relationship A R2N , As shown below:
[0148] A R2N =1-σ(t N D (R) ),
[0149] A R2T =1-σ(t T D (R) ),
[0150] Among them, t N and t T are two learnable hyperparameters;
[0151] Then, a hierarchical GCN is used to propagate information to recover the feature H R2N and H R2T ; The recovery process is defined as:
[0152]
[0153] in Indicates the number of the GCN layer,
[0154] and is the trainable transformation matrix of each layer, m = {N, T};
[0155] The outputs after passing through L layers of GCN are Abbreviated as H R2N , H R2T .
[0156] The expression of the comprehensive loss function L in this embodiment is:
[0157]
[0158] Among them, L MM , L ce , L 3mThey are GRMM loss, cross entropy loss and multimodal margin loss, respectively. is a balancing hyperparameter;
[0159] GRMM loss L MM The calculation formula is: L MM =L R +L N +L T ;
[0160] L R , L N and L T They represent the reconstruction losses of RGB, NIR, and TIR modalities, respectively;
[0161] In order to promote the feature reconstruction capability, the restored features are constrained by the mean square error (MSE) loss, using the token features and their structural relationships, which is defined as:
[0162]
[0163]
[0164]
[0165] Example
[0166] In order to verify the technical effect of the present invention, four common multimodal re-identification datasets are used here, including two pedestrian re-identification datasets (RGBNT201, Market1501-MM) and two vehicle re-identification datasets (RGBNT100, MSVR310).
[0167] This example uses mean average precision (mAP) and cumulative matching features (CMC) as evaluation indicators on all used datasets; higher mAP and CMC values indicate better model performance, as shown in Tables 1 and 2.
[0168] Table 1 Experimental data of person re-identification datasets RGBNT201 and Market1501-MM
[0169]
[0170] Table 2 Experimental data of vehicle re-identification datasets RGBNT100 and MSVR310
[0171]
[0172] The technical solution of the present invention is also comprehensively compared with other current methods. The experimental results show that the existing CNN-based method has lower performance. The TOP-ReID method achieves good performance by combining the global information of each modality with the local information of all modalities. However, the TOP-ReID method regards all local patches as equal and fails to distinguish patches of different qualities.
[0173] In contrast, the network model of the present invention can effectively focus on the quality differences of patches and extract more discriminative features. In particular, on the RGBNT201 dataset, the present invention surpasses the second place in multiple indicators, specifically by 6.1%, 5.9%, 6.2% and 5.8% in mAP, rank-1, rank-5 and rank-10 respectively. For the Market1501-MM dataset, the network model of the present invention surpasses the second best model in mAP and rank-1, by 2.4% and 1.6% respectively. .
[0174] In addition, for the two vehicle datasets, the proposed method consistently performs well, further demonstrating its effectiveness in different scenarios by considering the quality differences of local features and reducing the impact of low-quality local features.
[0175] To further evaluate the missing modal graph reasoning strategy GRMM of the present invention, this embodiment simulates different modal missing scenarios and displays the results, as shown in Table 3.
[0176] Despite the lack of certain modalities, the network model of the present invention always performs well; the lack of RGB modality has the greatest impact among all modalities because RGB modality is consistent with the dominant modality. At the same time, even in the absence of near-infrared modality, the present invention still outperforms single-modal methods and some multimodal methods. In short, using graph reasoning can effectively reconstruct the missing modal information, thereby providing reliable multimodal data for object re-identification.
[0177] Table 3 Effects of missing modal graph reasoning strategies for different modal missing scenarios
[0178]
[0179] In Table 3, M() indicates missing mode.
[0180] like Figure 3As shown, the visualization diagram of this embodiment shows the visualization results of the non-modal interaction graph reasoning (GRMI) and the modal interaction graph reasoning (GRMI) model using the gradient weighted class activation mapping (Grad-CAM). These visualizations show the ability of the network model MGRNet of the present invention to effectively capture the relevant areas of the input image. Compared with the solution that does not adopt the GRMI strategy, the present invention can capture more key areas while significantly reducing the impact of irrelevant areas or noise areas.
[0181] To further verify the feasibility of the technical solution of the present invention, this embodiment implements the network model MGRNet through PyTorch and runs it on an RTX 4090 GPU. For all data sets, the images are adjusted to: for the pedestrian ReID data set, the image size is 256×128×3 pixels; for the vehicle ReID data set, the image size is 128×256×3 pixels. In addition, common data enhancement techniques such as random erasing, random flipping and padding are used in the training process; the batch size is set to 64, and each batch contains 4 different identities. The training cycle of the network model MGRNet is 80 epochs, using a stochastic gradient descent (SGD) optimizer, a momentum coefficient of 0.9, and a weight decay of 0.0001. The initial learning rate is set to 0.0066, and a learning rate warm-up strategy combined with cosine decay is adopted.
Claims
1. A multimodal target re-identification method based on modal perception graph reasoning, characterized in that: Construct a modality-aware graph reasoning network MGRNet, use the modality-aware graph reasoning network MGRNet to realize the information interaction of multimodal images and restore the missing modality in multimodal target re-identification; corresponding modality image I (m) The specific execution process after inputting the modality perception graph reasoning network MGRNet is as follows: Step 1: Use a multi-branch backbone network based on visual transformer ViTs to perform initial feature X on each modality image. (m) Extraction of X (m) Represents the mth modal image I (m) The initial feature representation of ; m = {N, R, T} represents the near infrared image NIR, RGB image and thermal infrared TIR image modalities respectively; Step 2: The modal interaction graph reasoning strategy GRMI considers both local and global features, captures local details and mitigates the impact of low-quality local features; specifically, it includes the following: Step 2.1: Modal perception map learning, that is, constructing a modal perception map for each modality image separately. The modal perception map of the mth modality image is G(V (m) ,E (m) ), V (m) represents the nodes in the modality perception graph, E (m) represents the edges in the modality-aware graph; Step 2.2, selective graph node exchange, that is, introduce global tokens and use the Top-K method to select the K minimum values of the edges in each graph, find low-quality patches, and obtain the local feature nodes after the three modal enhancement updates. Step 2.3: Local perception graph reasoning: A multi-layer local perception graph reasoning module is used to learn the local patch representation of multimodal data to obtain better quality node patches and obtain the local patch representation F. l (m) ; Step 3: Use the global-aware multi-head attention module to aggregate the rich local information from different regions into global features; For missing modal images, the missing modal graph inference strategy GRMM is used to complete the missing modal information, reduce the differences between multimodal data, and restore the missing modal information; For the above-constructed modality-aware graph reasoning network MGRNet, a comprehensive loss function L is used for training optimization.
2. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1 is characterized in that: The multi-branch backbone network used in step 1 to extract the initial features is non-shared; The initial features obtained The expression is: in, represents the input mth modality image, and They represent the patch tokens and class tokens of the input image respectively. The subscript l represents local, which is the patch token of the input image and represents the local part; the subscript g represents global, which is the class token of the input image and represents the global part.
3. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1 is characterized in that: For step 2, the modal perception map in the process of modal perception map learning is G(V (m) ,E (m) ), node V (m) Represents the local feature set of the corresponding modality image; use Represents the feature set of all patches, P is the number of local tokens; The edge E of the graph (m) Connect different patches and use the adjacency matrix A (m) To represent the structural relationship between local features; Since edge E (m) The learning is dynamic, building relationships by calculating the Euclidean distance: represents the Euclidean distance between the i-th and j-th patch token feature vectors of the m-th modality; and Represent the feature vectors of the i-th and j-th token respectively; To learn more effective graphs, two learnable parameters α and t are introduced to calculate The formula is as follows: σ represents the sigmoid activation function; Represents the adjacency matrix between the i-th and j-th patch tokens of the m-th modality.
4. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1 is characterized in that: The specific method for selective graph node exchange in step 2 is: First, the Euclidean distance between the global feature and each local feature is calculated, and the relationship strength W between the global feature and the local feature is calculated. (m) , W (m) The calculation formula is as follows: Then the W of all modes (m) Composition: cdist() represents Euclidean distance calculation. Represents the strength of the relationship between the Pth local feature and the global feature of the m-mode; Then, use W (m) The inferior nodes obtained in the first step are filtered out to obtain the updated local feature nodes; here, the inferior patches are replaced by better patches of other modes, and the correlation between patches is considered, so the updated local feature nodes The definition is as follows: It means using the features of RGB modality to exchange the poor quality features of NIR modality; It means using the features of RGB modality to exchange the poor quality features of TIR modality; It means using the mean of NIR and TIR features to exchange the poor features of RGB modality.
5. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1 or 4, characterized in that: Before replacing a low-quality patch with a better patch of another modality, it is necessary to determine whether the swapped patch is a low-quality patch. If it is a low-quality patch, the current patch is reinitialized to an all-zero matrix, and the feature expression of the current patch is learned through the local neighbors of the current patch. If it is determined not to be a low-quality patch, the patch is exchanged.
6. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1 is characterized in that: After completing the selective graph node exchange in step 2.2, for the obtained enhanced modality-aware graph node representation, a multi-layer local-aware graph reasoning module is used to learn the local patch representation of multimodal data to obtain better quality node patches; The message propagation rules are defined as follows: In the above formula, is a learnable transformation matrix, Abbreviated as Indicates how the next layer of graph convolution is updated. represents a layer of graph convolution, represents the last layer of graph convolution, represents the adjacency matrix of the ith and jth patch tokens of the mth modality, yes A general term for .
7. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1 is characterized in that: Step 3: When using the global-aware multi-head attention module to increase the interaction between global and local information, we first assume that Represents the patches obtained through the above local perception graph reasoning network; Then, a linear projection layer is applied to obtain the global labeled query matrix And for local tokens Different linear projection layers are used to obtain the key matrix K (m) Sum value matrix V (m) ; Next, we use H heads of multi-head attention to aggregate local information from different patches into a global feature representation. This interaction operation is defined in the hth head as follows: in is the scaling factor; represents the output features obtained by the h-th head; To aggregate the information of all headers, a concatenation operation is used to obtain a new class tag: Finally, the features of all modes Splice to get the fused feature Z.
8. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1, characterized in that: If mode missing occurs, the missing mode graph inference strategy GRMM is used to compensate for the missing mode information and reduce the difference between modes. That is, feature reconstruction is first applied to enhance feature representation, and then structure reconstruction is used to learn the relationship between corresponding tags. The specific method is as follows: Assuming that there is only RGB modality, we use the RGB modality to recover the labels of NIR and TIR modalities. First, we use the Euclidean distance to calculate the original structural relationship between the true labels of each modality: in and Represent the feature vectors of the i-th and j-th token respectively; Next, construct the dynamically reconstructed structural relationship As shown below: A R2N =1-σ(t N D (R) ), A R2T =1-σ(t T D (R) ), Among them, t N and t T are two learnable hyperparameters; Then, a hierarchical GCN is used to propagate information to recover the feature H R2N and H R2T ; The recovery process is defined as: in Indicates the number of the GCN layer, and is the trainable transformation matrix of each layer, m = {N, T}; The outputs after passing through L layers of GCN are Abbreviated as H R2N , H R2T .
9. The multimodal target re-identification method based on modal perception graph reasoning according to claim 1, characterized in that: The expression of the comprehensive loss function L is: Among them, L MM , L ce , L 3m They are GRMM loss, cross entropy loss and multimodal margin loss, respectively. is a balancing hyperparameter; GRMM loss L MM The calculation formula is: L MM =L R +L N +L T ; L R , L N and L T Represent the reconstruction losses of RGB, NIR, and TIR modalities respectively; In order to promote the feature reconstruction capability, the restored features are constrained by the mean square error (MSE) loss, using the token features and their structural relationships, which is defined as:
Citation Information
Cited By
Multi-modal vehicle re-identification method and system based on flow generation type multi-expert fusion
CN121616900A
Multi-modal vehicle re-identification method and system based on flow generation multi-expert fusion
CN121616900B