Multi-modal pedestrian re-identification method based on complementary data enhancement
By enhancing the multimodal pedestrian re-identification method with dual-color space data and modal reconstruction processing, the problem of modal interaction calculation overhead and low sample quality is solved, and efficient pedestrian re-identification effect is achieved.
Patent Information
- Application Number
- CN202510703178.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing multimodal pedestrian re-identification method will introduce additional computing overhead when dealing with modal interactions, and in practical applications, the problems of low sample quality and missing specific modalities are more prominent, resulting in a decline in model discrimination ability.
A multimodal pedestrian recognition method based on complementary data enhancement is adopted. By enhancing the initial modal image data with double color spatial data, global semantic information and local detail information are extracted, and the trained modal reconstruction module is used to generate predictive features of missing modalities in the absence of modality, and finally construct multimodal image features to realize pedestrian recognition.
This method effectively solves the problem of modal incompleteness, reduces the calculation overhead in the inference stage, improves the response speed and retrieval accuracy in actual applications, and is suitable for more general pedestrian re-identification scenarios.
Smart Images

Figure CN120236301A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multi-modal person re-identification method based on complementary data augmentation. Background Art
[0002] Person Re-identification (Person ReID) is a technology used to retrieve and identify the same person under different cameras, and has been widely applied in the fields of security, surveillance, and intelligent transportation.
[0003] Different from the traditional person re-identification task based on a single visible light image, multi-modal person re-identification aims to achieve robust re-identification by introducing multiple complementary modal images for each person. This person re-identification method helps to process more complex lighting scenarios, greatly assisting the traditional person re-identification task and solving its application limitations. In addition, the popularization of various types of cameras (such as various infrared and RGB cameras) makes multi-modal person re-identification possible and has attracted more attention in recent years. Therefore, due to its strong complementary advantages, multi-modal person re-identification has great potential application value in the field of intelligent surveillance systems.
[0004] In the multi-modal person re-identification task, visible light, thermal infrared (TI), and near infrared (NI) modal images are widely adopted due to their modal complementarity and the ability to meet the requirements of all-weather and full-scene person recognition. Among them, the visible light modal image is implemented through RGB images. When the lighting conditions are good, RGB images can provide rich color and texture information; NI images are not affected by lighting and can provide clear edge information; while TI images can distinguish people from the surrounding environment using temperature and are not disturbed by complex environments. Therefore, how to make full use of the complementary information of different modalities is the key to multi-modal person re-identification. Multi-modal person re-identification is very different from traditional cross-modal person re-identification. Traditional cross-modal person re-identification focuses on reducing modal differences and learning modal-shared features. In contrast, multi-modal person re-identification focuses on effective modal fusion to absorb the complementary information in multi-modal data, thereby improving the distinguishability of people.
[0005] Current multi-modal pedestrian re-identification methods design various interaction modules to process multi-modal data in order to make full use of the complementary information between modalities. However, these methods are not only limited by the backbone architectures used, but also introduce additional computational overhead during the inference stage due to modality interaction. In addition, existing contrastive learning-based multi-modal alignment methods can achieve global feature alignment between modalities, but their forced alignment strategy requires that the unimodal representations of the same identity (ID) be strictly consistent in the embedding space. This strict constraint will over-suppress modality-specific features (such as texture details) during the processing. In fine-grained classification tasks such as pedestrian re-identification, where subtle differences need to be distinguished, this phenomenon will significantly reduce the discriminative ability of the model. At the same time, compared with unimodal data, annotating multi-modal data is more challenging and the number of samples is usually limited. Due to equipment failures, environmental limitations, or sensor failures, the sample quality may be relatively low, and even some modalities may be missing. These problems pose challenges to the practical application of pedestrian re-identification technology. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a multi-modal pedestrian re-identification method based on complementary data augmentation. The technical problems to be solved by the present invention are realized through the following technical solutions: An embodiment of the present invention provides a multi-modal pedestrian re-identification method based on complementary data augmentation, including the steps of: Perform dual color space data augmentation on the initial modal image data to obtain initial modal enhanced images, where the initial modal image data includes at least one of RGB images, TI images, and NI images; Use a trained feature extraction module to extract global semantic information and local detail information from the initial modal enhanced images to obtain initial modal image features; When the initial modal image data includes one or two of RGB images, TI images, and NI images, based on the initial modal image features, use different modal reconstructors in the trained modal reconstruction module to generate predicted features of the missing modality to obtain multi-modal image features; When the initial modal image data includes all three of RGB images, TI images, and NI images, form multi-modal image features from the initial modal image features; Connect the class labels of the same identity and the same sample in the multi-modal image features to obtain positive samples, and obtain the pedestrian re-identification result.
[0007] In an embodiment of the present invention, the training methods of the feature extraction module and the modal reconstruction module include the steps of: Perform dual color space data augmentation on the multi-modal training data to obtain multi-modal augmented training data, where the multi-modal training data includes RGB images, TI images, and NI images; Use the feature extraction module to extract global semantic information and local detail information from the multi-modal augmented training data to obtain multi-modal image training features; Use the multi-modal image training features to train the modal reconstruction module, and combine the total loss of significant feature reconstruction to constrain the modal reconstruction module, where the total loss of significant feature reconstruction includes the total loss of generating NI and TI modalities from the RGB modality, the total loss of generating RGB and TI modalities from the NI modality, and the total loss of generating RGB and NI modalities from the TI modality; Select the class labels of the same identity and the same sample from the multi-modal image training features and connect them to obtain training positive samples, and randomly select the class labels of different identity samples or the same identity but different samples and connect them to obtain training negative samples; Input the training positive samples and the training negative samples into a binary classifier for discrimination; Combine the predicted labels output by the binary classifier and use the total loss of modal-aware soft alignment to constrain the binary classifier and the feature extraction module, where the total loss of modal-aware soft alignment includes the first cross-entropy loss and the cross-modal hard triplet loss; Input the training positive samples into a linear classification layer for classification of different identities; Use the second cross-entropy loss and the triplet loss to constrain the linear classification layer and the feature extraction module; Repeat the iteration until the total loss converges to obtain the trained feature extraction module and the trained modal reconstruction module, where the total loss includes the total loss of significant feature reconstruction, the first cross-entropy loss, the cross-modal hard triplet loss, the second cross-entropy loss, and the triplet loss.
[0008] Compared with the prior art, the beneficial effects of the present invention are: A multi-modal pedestrian re-identification method based on complementary data augmentation provided by the present invention first performs dual-color space data augmentation on the initial modal images and extracts features from the initial modal images, achieving data expansion and data quality improvement. Then, in the case of missing modalities, different modal reconstructors in the trained modal reconstruction module are used to generate predicted features of the missing modalities, effectively solving the problem of incomplete modalities. Finally, the multi-modal image features are used to construct positive samples to obtain the pedestrian re-identification results, promoting the collaborative perception of multi-modal features by multiple networks, enhancing the shared features, and solving the problems that modal interaction will introduce additional computational overhead in the inference stage and low sample quality and specific modal missing in practical applications. Therefore, this method does not require any interaction module, and in the case of missing modalities, it will not spread the noise in the reconstructed features to other modalities, thereby improving the response speed and retrieval accuracy in practical applications and being applicable to more general pedestrian re-identification scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 FIG. is a schematic flowchart of a multi-modal pedestrian re-identification method based on complementary data augmentation provided by an embodiment of the present invention; Figure 2 FIG. is a schematic diagram of dual-color space data augmentation provided by an embodiment of the present invention; Figure 3 FIG. is a schematic diagram of extracting different modal features using the ViT-B / 16 network in different modalities provided by an embodiment of the present invention; Figure 4 FIG. is a schematic structural diagram of a modal reconstructor provided by an embodiment of the present invention; Figure 5 FIG. is a schematic flowchart of a training method for a feature extraction module and a modal reconstruction module provided by an embodiment of the present invention; Figure 6 FIG. is a schematic diagram of constructing positive and negative samples provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0011] Embodiment 1 Please refer to Figure 1 , Figure 1 FIG. is a schematic flowchart of a multi-modal pedestrian re-identification method based on complementary data augmentation provided by an embodiment of the present invention.
[0012] The multi-modal pedestrian re-identification method based on complementary data augmentation in this embodiment includes the steps: S11. Perform dual-color space data augmentation on the initial modal image data to obtain the initial modal enhanced image.
[0013] Please refer to Figure 2 , Figure 2 which is a schematic diagram of dual - color - space data augmentation provided by an embodiment of the present invention. Specifically, the dual - color space includes the RGB space and the HSV space, and the dual - color - space data augmentation includes two parts: rotation of the initial - modality image blocks and RGB image enhancement. The initial - modality image data includes at least one of RGB images, TI images, and NI images.
[0014] Step S11 specifically includes: S111. Rotate the image blocks of the initial - modality image data to obtain the rotated image of the initial - modality image. Specifically, it includes the following steps: S1111. Initialize the probability for images with the same identity and the same sample in the initial - modality image data. Among them, the same identity means the same ID, for example, the same pedestrian; the same sample means the same time and the same space. Taking the initial - modality image data including RGB images, TI images, and NI images as an example, images with the same identity and the same sample refer to images in RGB images, TI images, and NI images with the same ID and the same time and space.
[0015] S1112. When the probability is less than or equal to a preset value, retain the images with the same identity and the same sample. Specifically, the preset value can be 0.5, that is, when the probability , retain the images with the same identity and the same sample.
[0016] When the probability is greater than the preset value, randomly initialize the position and size of the image block, cut out the image block from the same position of each modality image according to the position and size of the image block, and then perform cyclic replacement on the image blocks of different modalities to obtain the rotated image of each initial - modality image. Specifically, it includes the following steps: 1) When the probability is , randomly initialize the position and size
[0017] of the image block. Specifically, when the probability is , randomly generate the row coordinate and column coordinate of the starting point of the image block, and determine the size of the image block. Among them, the image block is a rectangular image block, and the position can be the row coordinate and column coordinate of the upper - left vertex of the rectangular image block, and the size refers to that both the length and width of the rectangular image block are The row coordinate and column coordinates need to satisfy and , is the image height, is the image width. The size of the image patch takes a random value within , is a fixed value, determines the maximum size of the image patch, is the image size.
[0018] 2) Cut out image patches from the same positions of each modal image according to the position and size of the image patches.
[0019] Specifically, taking the initial modal image data including RGB images, TI images and NI images as an example, according to the position and size of the image patches, for the images of the same identity and the same sample in the RGB images, TI images and NI images, cut out image patches with a size of at the same position, to obtain RGB image patches, TI image patches and NI image patches.
[0020] 3) Perform cyclic replacement on the image patches of different modalities to obtain the rotated images of each initial modal image.
[0021] Specifically, when the initial modal image data includes RGB images, TI images and NI images, perform cyclic replacement on the RGB image patches, TI image patches and NI image patches; for example, replace the RGB image patches with TI image patches, replace the TI image patches with NI image patches, and replace the NI image patches with RGB image patches; or, replace the RGB image patches with NI image patches, replace the NI image patches with TI image patches, and replace the TI image patches with RGB image patches, to obtain the rotated images of the RGB images, TI images and NI images.
[0022] When the initial modal image data includes two of the RGB images, TI images and NI images, replace the image patches of the two modalities to obtain the rotated images of the two modal images. For example, replace the RGB image patches with TI image patches, and replace the TI image patches with RGB image patches, to obtain the rotated images of the RGB images and TI images.
[0023] When the initial modal image data includes any one of the RGB images, TI images and NI images, no replacement is performed, and the rotated image of the initial modal image is directly output. For example, if the initial modal image data includes RGB images, the RGB images are directly output as the rotated images of the RGB images.
[0024] S112. When the initial modal image data includes an RGB image, perform saturation and brightness enhancement on the rotated image of the RGB image to obtain an enhanced RGB image. ; When the initial modal image data includes a TI image, use the rotated image of the TI image as the enhanced TI image. ; When the initial modal image data includes an NI image, use the rotated image of the NI image as the enhanced NI image. .
[0025] It can be understood that when the initial modal image data includes an RGB image, enhance the rotated image of the RGB image; when the initial modal image data does not include an RGB image, directly output the rotated image as the enhanced image.
[0026] Specifically, performing saturation and brightness enhancement on the rotated image of the RGB image to obtain an enhanced RGB image includes the steps of: 1) Convert the rotated image of the RGB image to the HSV space to obtain a converted image.
[0027] 2) Calculate the saturation and brightness of the converted image. Among them, the calculation formulas for saturation and brightness are: ; ; Among them, is the saturation, is the brightness, is the maximum pixel value of the R, G, and B channels, is the minimum pixel value of the R, G, and B channels.
[0028] 3) Perform saturation enhancement and brightness enhancement on the converted image to obtain an SV enhanced image. Among them, the enhancement formulas for saturation and brightness are: ; Among them, is the enhanced saturation, is the enhanced brightness.
[0029] 4) Convert the SV enhanced image to the RGB space to obtain an enhanced RGB image. .
[0030] Exemplarily, the initial modal image data input in step S1 is an image of 256×128×3 that has been size-normalized, and the size of the output initial modal enhanced image remains unchanged.
[0031] S12. Use the trained feature extraction module to extract global semantic information and local detail information from the initial modal enhanced image to obtain initial modal image features.
[0032] Specifically, in different modalities, the trained ViT-B / 16 network is used as the backbone network to extract global semantic information and local detail information from the initial modality enhanced image, obtaining the initial modality image features, where the initial modality image features include class tokens representing the global semantic information of the image.
[0033] Please refer to Figure 3 , Figure 3 which is a schematic diagram of extracting different modality features using the ViT-B / 16 network in different modalities provided by the embodiments of the present invention. Taking the initial modality image data including RGB images, TI images, and NI images as an example, step S12 includes: Using the trained first ViT-B / 16 network to extract global semantic information and local detail information from the enhanced RGB image to obtain RGB modality image features: ; wherein, is the RGB modality image feature, is the enhanced RGB image, is the first ViT-B / 16 network in the RGB modality.
[0034] Using the trained second ViT-B / 16 network to extract global semantic information and local detail information from the enhanced TI image to obtain TI modality image features: ; wherein, is the TI modality image feature, is the enhanced TI image, is the second ViT-B / 16 network in the TI modality.
[0035] Using the trained third ViT-B / 16 network to extract global semantic information and local detail information from the enhanced NI image to obtain NI modality image features: ; wherein, is the NI modality image feature, is the enhanced NI image, is the third ViT-B / 16 network in the NI modality.
[0036] Exemplarily, the RGB modality image feature , the TI modality image feature , and the NI modality image feature The dimensions are all 768×129. Each modal image feature includes a class token representing the global semantic information of the image, and the dimension of the class token is 768×1.
[0037] S13. When the initial modal image data includes one or two of RGB images, TI images, and NI images, based on the initial modal image features, use different modal reconstructors in the trained modal reconstruction module to generate predicted features of the missing modality, and obtain multi-modal image features; when the initial modal image data includes all three of RGB images, TI images, and NI images, the multi-modal image features are composed of the initial modal image features.
[0038] Specifically, the modal reconstruction module includes a first modal reconstructor that generates TI modal image features from RGB modal image features a second modal reconstructor that generates NI modal image features from RGB modal image features a third modal reconstructor that generates RGB modal image features from TI modal image features a fourth modal reconstructor that generates NI modal image features from TI modal image features a fifth modal reconstructor that generates RGB modal image features from NI modal image features and a sixth modal reconstructor that generates TI modal image features from NI modal image features .
[0039] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a modal reconstructor provided by an embodiment of the present invention. The first modal reconstructor, the second modal reconstructor, the third modal reconstructor, the fourth modal reconstructor, the fifth modal reconstructor, and the sixth modal reconstructor have the same structure, and all include: The first layer normalization layer is used to perform layer normalization processing on the input initial modal image features to standardize the input features and obtain the first layer normalized features; The multi-head self-attention module is used to process the first layer normalized features using the multi-head self-attention mechanism to learn how to reconstruct another modal image feature from one modal image feature and obtain the multi-head self-attention output features; The first residual fusion module is used to add and fuse the input initial modal image features and the multi-head self-attention output features to achieve residual fusion and obtain the fused features; The second layer normalization layer is used to perform layer normalization processing on the fused features to ensure the stability of the input to the subsequent multi-layer perceptron and obtain the second layer normalized features; The multi-layer perceptron is used to generate preliminary reconstruction features based on the second layer normalized features; The second residual fusion module is used to add and fuse the fusion feature and the preliminary reconstruction feature to achieve residual fusion, and obtain the predicted feature of the missing modality.
[0040] Exemplarily, the predicted feature output by each modality reconstructor is consistent with the modality image feature of the input in dimension, both being 768×129.
[0041] Specifically, when the initial modality image data includes any one of RGB images, TI images, and NI images, that is, when two modalities are missing from the initial modality image data, the modality image feature of this type is input into the corresponding modality reconstructor to generate the predicted features of the other two missing modalities, thereby obtaining the image features of RGB, TI, and NI modalities.
[0042] The TI modality image feature generated from the RGB modality image feature and the NI modality image feature are: ; .
[0043] The RGB modality image feature generated from the TI modality image feature and the NI modality image feature are: ; .
[0044] The RGB modality image feature generated from the NI modality image feature and the TI modality image feature are: ; .
[0045] When the initial modality image data includes two of RGB images, TI images, and NI images, that is, when one modality is missing from the initial modality image data, the two modality image features are respectively input into the corresponding modality reconstructors to generate the features of the missing modality, and then the generated features are averaged as the predicted feature of the missing modality. For example, when the initial modality image data includes RGB images and TI images and the NI image is missing, the RGB modality image feature is input into the second reconstructor to generate the NI modality image feature, the TI modality image feature is input into the fourth reconstructor to generate the NI modality image feature, and the two generated NI modality image features are averaged to obtain the predicted feature of the NI modality.
[0046] In this embodiment, when one modality is missing, the features of the missing modality are jointly constructed using two modalities. The features of the two available modalities can provide more comprehensive information, make up for the limitations of a single modality, and realize the utilization of complementary information. The joint use of the features of the two modalities can reduce noise interference, improve the reconstruction quality, and have good noise resistance. Through multi-modal collaborative reconstruction, it is ensured that the model still maintains high precision when some modalities are missing, and has good robustness.
[0047] When the initial modality image data includes three of RGB images, TI images, and NI images, the multi-modal image features are composed of the initial modality image features, that is, the multi-modal image features are directly composed of the RGB modality image features, TI modality image features, and NI modality image features.
[0048] S14. Connect the class labels of the same identity and the same sample in the multi-modal image features to obtain positive samples, and obtain the pedestrian re-identification result.
[0049] Specifically, connect the class labels of the same identity and the same sample in the RGB modality, TI modality, and NI modality image features to obtain positive samples, and obtain the pedestrian re-identification result. Among them, the same sample refers to having the same time and space.
[0050] Exemplarily, the dimension of the class label of the three modality image features is 768, and the dimension represented by the final positive sample is 768×3 = 2304.
[0051] The multi-modal pedestrian re-identification method of this embodiment first performs dual-color space data augmentation on the initial modality image and extracts features from the initial modality image, realizing data expansion and data quality improvement. Then, in the case of modality missing, different modality reconstructors in the trained modality reconstruction module are used to generate predicted features of the missing modality, effectively solving the problem of modality incompleteness. Finally, positive samples are constructed using multi-modal image features to obtain the pedestrian re-identification result, promoting the collaborative perception of multi-modal features by multiple networks, enhancing shared features, and solving the problems that modality interaction will introduce additional computational overhead in the inference stage and low sample quality and specific modality missing in practical applications. Therefore, this method does not require any interaction module, and in the case of modality missing, it will not spread the noise in the reconstructed features to other modalities, thereby improving the response speed and retrieval accuracy in practical applications and being applicable to more general pedestrian re-identification scenarios.
[0052] The method of this embodiment significantly reduces the inference time by independently extracting single-modal features and directly splicing, while maintaining high precision, achieving a better balance between efficiency and performance.
[0053] Please refer to Figure 5 , Figure 5Schematic diagram of the training method for a feature extraction module and a modality reconstruction module provided by an embodiment of the present invention. The training method includes the steps: S501. Perform dual-color space data augmentation on the multi-modal training data to obtain multi-modal augmented training data, where the multi-modal training data includes RGB images, TI images, and NI images.
[0054] S502. Use the feature extraction module to extract global semantic information and local detail information from the multi-modal augmented training data to obtain multi-modal image training features.
[0055] For the specific execution steps of steps S501 and S502, please refer to the above steps S11 and S12.
[0056] S503. Use the multi-modal image training features to train the modality reconstruction module, and combine the total loss of significant feature reconstruction to constrain the modality reconstruction module, where the total loss of significant feature reconstruction includes the total loss of generating NI modality and TI modality from RGB modality, the total loss of generating RGB modality and TI modality from NI modality, and the total loss of generating RGB modality and NI modality from TI modality.
[0057] Specifically, for each embedded image patch in the feature map generated by the modality reconstructor, use the mean squared error loss function (MSE) to impose pixel-level constraints: ; ; ; ; ; ; where is the mean squared error loss of generating NI modality from RGB modality is the number of embedded image patches in the generated modality features, is the number of embedded image patches in the generated modality features, is the NI modality image feature is the th embedded image patch, is the NI modality feature map generated from RGB modality is the th embedded image patch; is the mean squared error loss of generating TI modality from RGB modality is the mean squared error loss, is the TI modality image feature is the th embedded image patch, The TI modal feature map generated from the RGB modality The th embedded image patch; The mean squared error loss of the RGB modality generated from the NI modality where is the RGB modality image feature The th embedded image patch, and is the RGB modal feature map generated from the NI modality The th embedded image patch; The mean squared error loss of the TI modality generated from the NI modality where is the TI modal feature map generated from the NI modality The th embedded image patch; The mean squared error loss of the RGB modality generated from the TI modality where is the RGB modal feature map generated from the TI modality The th image patch; The mean squared error loss of the NI modality generated from the TI modality where is the NI modal feature map generated from the TI modality The , , , is the enhanced RGB image, is the first ViT-B / 16 network in the RGB modality, is the enhanced TI image, is the second ViT-B / 16 network in the TI modality, is the enhanced NI image, is the third ViT-B / 16 network in the NI modality; , , , , , , is the first modal reconstructor, is the second modal reconstructor, is the third modal reconstructor, is the fourth modal reconstructor, is the fifth modal reconstructor, is the sixth modal reconstructor.
[0058] Furthermore, class labels are extracted from the feature maps generated by the modality reconstructor and the input multi-modal image training features, and a focal frequency loss function (FFL) is constructed using the class labels to constrain the modality reconstructor: ; ; ; ; ; ; where is the focal frequency loss, is , , , , or , is the dimension of the class label, is the th dimension spectral weight matrix of the class label, , is the fast Fourier transform of the th dimension of the class label in representing RGB, NI or TI, is the fast Fourier transform of the th dimension of the class label in
[0059] In summary, the total loss of significant feature reconstruction is: ; ; ; ; where is the total loss of significant feature reconstruction, is the total loss of generating NI modality and TI modality from RGB modality, is the total loss of generating RGB modality and TI modality from NI modality, is the total loss of generating RGB modality and NI modality from TI modality.
[0060] It can be understood that the modality reconstruction module learns to reconstruct the features of other modalities from the features of one modality through the complete multi-modal data during the training phase; when a modality is missing during the test phase, it can generate the features of the missing modality based on the existing modalities.
[0061] S504. Select the class labels of the same identity and the same sample from the multi-modal image training features, connect them to obtain training positive samples, and randomly select the class labels of different identity samples or different samples of the same identity to connect and obtain training negative samples.
[0062] Specifically, connect the class labels of the same identity and the same sample in the three-modal image training features to obtain training positive samples, and assign the first pseudo-label 1 to the training positive samples: ; Among them, represents the label of the th training positive sample, is the class label of the th sample in the RGB-modal image training features, is the class label of the th sample in the NI-modal image training features, is the class label of the th sample in the TI-modal image training features.
[0063] Randomly select an equal number of class labels of samples belonging to different identities or different samples of the same identity from the three-modal image training features, connect them to obtain training negative samples, and assign the second pseudo-label 0 to the training negative samples: ; Among them, represents the label of the th training negative sample, is the class label of the th sample in the RGB-modal image training features, is the class label of the th sample in the NI-modal image training features, is the class label of the th sample in the TI-modal image training features, , At least two of the values in
[0064] Here, the same sample refers to a sample with the same time and the same space; different identity samples refer to samples with different identities; different samples of the same identity refer to samples with different times or different spaces within the same identity.
[0065] Please refer to Figure 6 , Figure 6Schematic diagram for constructing positive and negative samples provided by an embodiment of the present invention. Among them, from left to right in the picture are the training features of RGB modality images, the training features of NI modality images, and the training features of TI modality images. In the figure, the class labels of the same identity, i.e., ID#1, and the same sample, i.e., sample#1, are selected from the training features of the three modality images and connected to obtain the training positive sample. The class labels of the same identity, i.e., ID#1, the sample#2 is selected from the training features of the RGB modality image, and the sample#1 is selected from the training features of the TI modality image and the NI modality image and connected to obtain the training negative sample; or, the sample#1 is selected from the training features of the three modality images, ID#2 is selected from the training features of the RGB modality image, and ID#1 is selected from the training features of the TI modality image and the NI modality image and connected to obtain the training negative sample.
[0066] Exemplarily, the dimension of the class label of the three modality image features is 768, and the dimension of the final training positive sample and the training negative sample representation is 768×3 = 2304.
[0067] S505. Input the training positive sample and the training negative sample into a binary classifier for discrimination.
[0068] Specifically, the binary classifier adopts a binary classification linear layer. The training positive sample and the training negative sample are discriminated in the binary classification linear layer, and the predicted label of the sample is output.
[0069] S506. Combining the predicted labels output by the binary classifier, use the modality-aware soft alignment total loss to constrain the binary classifier and the feature extraction module, where the modality-aware soft alignment total loss includes the first cross-entropy loss and the cross-modal hard triplet loss.
[0070] Specifically, use the first cross-entropy loss as the classification loss to guide sample classification: ; where is the first cross-entropy loss, is the number of samples in the batch, is the true label of the th sample, and
[0071] Please refer to Figure 6 again, Figure 6The construction methods of positive and negative samples will lead to a problem that non-corresponding samples with the same identity in a batch may be marked as negative samples, which may cause the distance between cross-modal non-corresponding samples (positive samples pos) with the same identity to be greater than or equal to the distance between cross-modal samples with different identities (negative samples neg) under the guidance of classification loss, that is, the distance between different samples with the same ID is greater than the distance between different IDs. Therefore, a cross-modal hard triplet loss is adopted to impose constraints to prevent the above problems from occurring in the batch, ensuring that the feature distances of different samples with the same identity are as small as possible, while the feature distances of different identities are as large as possible. The cross-modal hard triplet loss function includes: ; ; ; Among them, is the RGB cross-modal hard triplet loss, is the NI cross-modal hard triplet loss, is the TI cross-modal hard triplet loss, is the margin parameter, represents hard negative mining for negative samples , represents the distance metric function, is the feature of the th anchor sample, is the feature of the hardest positive sample is the feature of the hardest negative sample
[0072] In summary, the total loss of the modality-aware soft alignment module is: ; Among them, is the weight parameter.
[0073] S507. Input the training positive samples into the linear classification layer for classification of different identities.
[0074] S508. Use the second cross-entropy loss and the triplet loss to constrain the linear classification layer and the feature extraction module.
[0075] Specifically, use the second cross-entropy loss and the triplet loss to constrain the training positive samples, solve the feature distribution differences between different samples, and make the feature distances of different samples with the same identity closer and the feature distances of different identities farther.
[0076] Specifically, the second cross-entropy loss is: ; Among them, is the second cross-entropy loss, is the number of samples in a batch, is the number of identity IDs, is the th sample in the constructed multi-modal training positive samples, is the th sample's true label indicating whether it belongs to a certain identity , is the predicted label for the th sample belonging to a certain identity .
[0077] The triplet loss is: ; Among them, is the triplet loss, is the th identity in the constructed multi-modal training positive samples, is the hard sample mining for the negative samples of the th identity, is the anchor sample of the th identity in the constructed multi-modal training positive samples, is the positive sample of the th identity in the constructed multi-modal training positive samples, is the negative sample of the th identity in the constructed multi-modal training positive samples, represents the anchor sample of the triplet, represents the positive sample of the triplet, represents the negative sample of the triplet, is the square of the Euclidean distance, is the margin parameter.
[0078] S509. Repeat the iteration until the total loss converges to obtain the trained feature extraction module and the trained modal reconstruction module. Among them, the total loss includes the significant feature reconstruction total loss, the first cross-entropy loss, the cross-modal hard triplet loss, the second cross-entropy loss, and the triplet loss.
[0079] Specifically, the total loss is: .
[0080] Furthermore, during the training, gradually determine the parameters of the maximum size of the image patch , the boundary parameter in the cross-modal hard triplet loss , and the weight parameter in the total loss of the modal perception soft alignment module through the validation set.
[0081] Parameter The method for determining the value is as follows: Fix the value to 0.01, use the mAP, Rank-1, Rank-5, and Rank-10 of the RGBNT201 dataset as indicators, set different values, execute step S501, observe the levels of the mAP, Rank-1, Rank-5, and Rank-10 indicators under different values, and select the value when the mAP, Rank-1, Rank-5, and Rank-10 indicators are the highest as the final result. In this embodiment, according to the experimental results, select .
[0082] Boundary parameter The method for determining the value is as follows: Fix the value to 0.5, use the mAP, Rank-1, Rank-5, and Rank-10 of the RGBNT201 dataset as indicators, set different values, execute steps S501 - S506, observe the levels of the mAP, Rank-1, Rank-5, and Rank-10 indicators under different values, and select the value when the mAP, Rank-1, Rank-5, and Rank-10 indicators are the highest as the final result. In this embodiment, according to the experimental results, select .
[0083] Weight parameter The method for determining the value is as follows: Fix the value to 0.5, the value to 0.3, use the mAP, Rank-1, Rank-5, and Rank-10 of the RGBNT201 dataset as indicators, set different values, execute steps S501 - S506, observe the levels of the mAP, Rank-1, Rank-5, and Rank-10 indicators under different values, and select the value when the mAP, Rank-1, Rank-5, and Rank-10 indicators are the highest as the final result. In this embodiment, according to the experimental results, select .
[0084] In the training method of this embodiment, the mean square error loss function and the focal frequency loss function are used to guide the reconstruction of significant features, effectively solving the problem of incomplete modality; based on the total loss of the modality-aware soft alignment module, it promotes the collaborative perception of multi-modal features by multiple networks and enhances the shared features.
[0085] In this embodiment, the multi-modal pedestrian re-identification method based on complementary data augmentation is further verified through the following simulations.
[0086] Simulation database: In this embodiment, evaluations and verifications are carried out on two multi-modal pedestrian re-identification datasets, RGBNT201 and MARKET1501_RGBNT. RGBNT201 is the first multi-modal pedestrian re-identification dataset, with 141 identities in its training set, 30 identities in its validation set, and 30 identities in its test set. MARKET1501_RGBNT is an extension of the Market1501 dataset, incorporating multi-modal information. This dataset has 750 identities for training and 751 identities for testing.
[0087] Evaluation criteria: The existing re-identification methods are compared with the method of this embodiment. The cumulative match characteristics (CMC) of the first-hit rate (RankR, R = 1, 5, 10, expressed as a percentage) and the mean average precision (mAP, expressed as a percentage) are used as evaluation metrics. Among them, the existing re-identification methods are respectively: multi-spectral vehicle re-identification method (HAMNet), robust multi-modal pedestrian re-identification (PFNet), dynamic enhancement network: method for partial multi-modal pedestrian re-identification (DENet), multi-modal pedestrian re-identification method based on low-rank fusion network (LRFNet), method for enhancing modal-specific representations in multi-modal pedestrian re-identification (IEEE), pedestrian re-identification method based on multi-modal consistency collaborative auxiliary training (MMCF), key marker screening: diversity feature selection mechanism in multi-modal object re-identification (EDITOR), application of the representation selective coupling method based on marker sparsification in multi-spectral object re-identification (RSCNet), heterogeneous test-time training: dynamic adaptation method for multi-modal pedestrian re-identification (HTT), multi-spectral object re-identification method based on marker permutation (TOP-REID), multi-modal object re-identification method based on the visual Mamba model (MambaReID), low-rank multi-scale multi-modal fusion pedestrian re-identification method based on RGB-NI-TI (LRMM). Please refer to Tables 1 and 2 for the evaluation results.
[0088] Table 1 Multi-modal pedestrian re-identification accuracy (RGBNT201 dataset)
[0089] Table 2 Multi-modal pedestrian re-identification accuracy (MARKET1501_RGBNT dataset)
[0090] As can be seen from Table 1 and Table 2, the method of this embodiment has achieved excellent performance on both the RGBNT201 and MARKET1501_RGBNT datasets. On the RGBNT201 dataset, the mAP reaches 74.6%, at least 2.3% higher than the existing best method TOP-REID; the Rank-1 accuracy is 77.6%, 1.0% higher than the existing best method TOP-REID. On the MARKET1501_RGBNT dataset, the mAP reaches 82.5%, 2.1% higher than the existing best method TOP-REID, and the Rank-1 accuracy is 92.7%, 0.5% higher than the existing best method TOP-REID.
[0091] The method of this embodiment significantly improves the robustness and inference efficiency of multi-modal ReID, and solves problems such as redundant modal interaction, rigid feature alignment, and low-quality modal noise propagation in existing methods. Through the HSV space enhancement and RGB-NI-TI cross-modal image block rotation strategy, the complementary information of multi-source data is fully exploited; combined with the global feature reconstruction with frequency-domain focal loss constraint and cross-modal hard sample triplet alignment, efficient feature fusion and noise suppression are achieved, significantly improving the cross-modal feature alignment efficiency and anti-noise interference ability. This method is applicable to all-weather intelligent monitoring systems under low-light, complex occlusion, and multi-modal missing scenarios, providing a high-precision and low-latency solution for practical deployment.
[0092] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A multi-modal pedestrian re-identification method based on complementary data augmentation, characterized in that, Including the steps: Performing dual-color space data augmentation on the initial modal image data to obtain an initial modal enhanced image, where the initial modal image data includes at least one of an RGB image, a TI image, and an NI image; Using the trained feature extraction module to extract global semantic information and local detail information from the initial modal enhanced image to obtain initial modal image features; When the initial modal image data includes one or two of an RGB image, a TI image, and an NI image, based on the initial modal image features, using different modal reconstructors in the trained modal reconstruction module to generate prediction features of the missing modality to obtain multi-modal image features; When the initial modal image data includes all three of an RGB image, a TI image, and an NI image, forming multi-modal image features from the initial modal image features; Connecting the class labels of the same identity and the same sample in the multi-modal image features to obtain positive samples to obtain the person re-identification result.
2. The multimodal pedestrian re-identification method based on complementary data augmentation according to claim 1, wherein, Performing dual-color space data augmentation on the initial modal image data to obtain an initial modal enhanced image, including: Performing image block rotation on the initial modal image data to obtain a rotated image of the initial modal image; When the initial modal image data includes an RGB image, performing saturation and brightness enhancement on the rotated image of the RGB image to obtain an enhanced RGB image; when the initial modal image data includes a TI image, using the rotated image of the TI image as the enhanced TI image; when the initial modal image data includes an NI image, using the rotated image of the NI image as the enhanced NI image.
3. The multi-modal pedestrian re-identification method based on complementary data augmentation according to claim 2, wherein Performing image block rotation on the initial modal image data to obtain a rotated image of the initial modal image, including: Initialize the probability for images of the same identity and the same sample in the initial modal image data ; When the probability is less than or equal to a preset value, retain the images of the same identity and the same sample; When the probability is greater than a preset value, randomly initialize the position and size of the image patch, cut out the image patch from the same position of each modal image according to the position and size of the image patch, and then perform cyclic replacement on the image patches of different modalities to obtain the rotated image of each initial modal image.
4. The multimodal pedestrian re-identification method based on complementary data augmentation according to claim 2, wherein Performing saturation and brightness enhancement on the rotated image of the RGB image to obtain an enhanced RGB image, including: Converting the rotated image of the RGB image to the HSV space to obtain a converted image; Calculating the saturation and brightness of the converted image, where the calculation formulas for the saturation and the brightness are: ; ; Among them, is the saturation, is the brightness, is the maximum pixel value of the R, G, and B channels, is the minimum pixel value of the R, G, and B channels; Performing saturation enhancement and brightness enhancement on the converted image to obtain an SV enhanced image, where the enhancement formulas for the saturation and the brightness are: ; Among them, is the enhanced saturation, is the enhanced brightness; Converting the SV enhanced image to the RGB space to obtain the enhanced RGB image.
5. The multimodal pedestrian re-identification method based on complementary data augmentation according to claim 1, wherein Using the trained feature extraction module to extract global semantic information and local detail information from the initial modal enhanced image to obtain initial modal image features, including: Under different modalities, using the trained ViT-B / 16 network to extract global semantic information and local detail information from the initial modal enhanced image to obtain initial modal image features, where the initial modal image features include class tokens representing the global semantic information of the image.
6. The multimodal pedestrian re-identification method based on complementary data augmentation according to claim 1, wherein The modal reconstruction module includes a first modal reconstructor for generating TI modal image features from RGB modal image features, a second modal reconstructor for generating NI modal image features from RGB modal image features, a third modal reconstructor for generating RGB modal image features from TI modal image features, a fourth modal reconstructor for generating NI modal image features from TI modal image features, a fifth modal reconstructor for generating RGB modal image features from NI modal image features, and a sixth modal reconstructor for generating TI modal image features from NI modal image features.
7. The multi-modal pedestrian re-identification method based on complementary data augmentation according to claim 6, wherein The first modal reconstructor, the second modal reconstructor, the third modal reconstructor, the fourth modal reconstructor, the fifth modal reconstructor, and the sixth modal reconstructor have the same structure, and all include: A first layer normalization layer for performing layer normalization processing on the input initial modal image features to obtain first layer normalized features; A multi-head self-attention module for processing the first layer normalized features using the multi-head self-attention mechanism to obtain multi-head self-attention output features; A first residual fusion module for adding and fusing the input initial modal image features and the multi-head self-attention output features to obtain fused features; A second layer normalization layer for performing layer normalization processing on the fused features to obtain second layer normalized features; A multi-layer perceptron for generating preliminary reconstruction features according to the second layer normalized features; A second residual fusion module for adding and fusing the fused features and the preliminary reconstruction features to obtain predicted features of the missing modality.
8. The multi-modal pedestrian re-identification method based on complementary data augmentation according to claim 1, wherein, The training method of the feature extraction module and the modal reconstruction module includes the steps of: Performing dual-color space data augmentation on multi-modal training data to obtain multi-modal augmented training data, where the multi-modal training data includes RGB images, TI images, and NI images; Using the feature extraction module to extract global semantic information and local detail information from the multi-modal augmented training data to obtain multi-modal image training features; Using the multi-modal image training features to train the modal reconstruction module, and constraining the modal reconstruction module in combination with the total loss of significant feature reconstruction, where the total loss of significant feature reconstruction includes the total loss of generating NI modal and TI modal from RGB modal, the total loss of generating RGB modal and TI modal from NI modal, and the total loss of generating RGB modal and NI modal from TI modal; Selecting the class labels of the same identity and the same sample from the multi-modal image training features and connecting them to obtain training positive samples, and randomly selecting the class labels of different identity samples or the same identity but different samples and connecting them to obtain training negative samples; Inputting the training positive samples and the training negative samples into a binary classifier for discrimination; Combining the predicted labels output by the binary classifier, and constraining the binary classifier and the feature extraction module using the total loss of modal-aware soft alignment, where the total loss of modal-aware soft alignment includes the first cross-entropy loss and the cross-modal hard triplet loss; Inputting the training positive samples into a linear classification layer for classification of different identities; Constraint the linear classification layer and the feature extraction module using the second cross-entropy loss and the triplet loss; Iterate repeatedly until the total loss converges to obtain the trained feature extraction module and the trained modality reconstruction module, where the total loss includes the significant feature reconstruction total loss, the first cross-entropy loss, the cross-modal hard triplet loss, the second cross-entropy loss, and the triplet loss.
9. The multi-modal pedestrian re-identification method based on complementary data augmentation according to claim 8, wherein The significant feature reconstruction total loss is: ; ; ; ; Among them, is the total loss for reconstructing significant features, is the total loss for generating the NI modality and the TI modality from the RGB modality, is the total loss for generating the RGB modality and the TI modality from the NI modality, is the total loss for generating the RGB modality and the NI modality from the TI modality; ; ; ; ; ; ; Among them, is the mean squared error loss for generating the NI modality from the RGB modality . is the number of embedded image patches in the generated modality features, is the NI modality image feature of the th embedded image patch, is the NI modality feature map generated from the RGB modality of the th embedded image patch; is the mean squared error loss for generating the TI modality from the RGB modality . is the TI modality image feature of the th embedded image patch, is the TI modality feature map generated from the RGB modality of the th embedded image patch; is the mean squared error loss for generating the RGB modality from the NI modality . is the RGB modality image feature of the th embedded image patch, is the RGB modality feature map generated from the NI modality of the th embedded image patch; is the mean squared error loss for generating the TI modality from the NI modality . is the TI modality feature map generated from the NI modality of the th embedded image patch; is the mean squared error loss for generating the RGB modality from the TI modality . is the RGB modality feature map generated from the TI modality of the th image patch; is the mean squared error loss for generating the NI modality from the TI modality . is the NI modality feature map generated from the TI modality of the th embedded image patch; , , , is the enhanced RGB image, is the first ViT-B / 16 network in the RGB modality, is the enhanced TI image, It is the second ViT-B / 16 network in the TI mode, It is the enhanced NI image, It is the third ViT-B / 16 network in the NI mode; , , , , , , It is the first mode reconstructor, It is the second mode reconstructor, It is the third mode reconstructor, It is the fourth mode reconstructor, It is the fifth mode reconstructor, It is the sixth mode reconstructor; ; ; ; ; ; ; Among them, is the focus frequency loss, is , , , , or , is the dimension of the class label, is the th spectral weight matrix of the dimension of the class label, , is the fast Fourier transform of the th dimension of the class label in represents RGB, NI or TI, is the fast Fourier transform of the th dimension of the class label in The total loss of the modality-aware soft alignment module is: ; ; ; ; ; Among them, is the total loss of the modal perception soft alignment module, is the weight parameter, is the first cross-entropy loss, is the number of samples in the batch, is the true label of the th sample, is the predicted label of the th sample output by the binary classifier, is the RGB cross-modal hard triplet loss, is the NI cross-modal hard triplet loss, is the boundary parameter, indicates hard negative sample mining, indicates the distance metric function, is the feature of the th anchor sample, is the feature of the hardest positive sample ; is the feature of the hardest negative sample ; The second cross-entropy loss is: ; Among them, is the second cross-entropy loss, is the number of samples in the batch, is the number of identities, is the th sample in the constructed multi-modal training positive samples, is the th sample's true label indicating whether it belongs to a certain identity , is the predicted label for the th sample indicating whether it belongs to a certain identity ; The triplet loss is: ; Among them, is the triplet loss, is the th identity in the multi-modal positive samples constructed during the training process, is to perform hard sample mining on the negative samples of the th identity, is the anchor sample of the th identity in the constructed multi-modal training positive samples, is the positive sample of the th identity in the constructed multi-modal training positive samples, is the negative sample of the th identity in the constructed multi-modal training positive samples, is the anchor sample of the triplet, is the positive sample of the triplet, is the negative sample of the triplet, is the square of the Euclidean distance, is the margin parameter; The total loss is: 。 10. The multimodal pedestrian re-identification method based on complementary data augmentation according to claim 8, wherein, Select the class labels of the same identity and the same sample from the multi-modal image training features and connect them to obtain training positive samples, and randomly select the class labels of different identity samples or different samples of the same identity for connection to obtain training negative samples, including: Connect the class labels of the same identity and the same sample in the multi-modal image training features to obtain training positive samples, and assign a first pseudo-label to the training positive samples: ; Among them, represents the label of the th training positive sample, is the class label of the th sample in the training features of RGB modality images, is the class label of the th sample in the training features of NI modality images, is the class label of the th sample in the training features of TI modality images; Randomly select the class labels of different identity samples or different samples of the same identity from the multi-modal image training features for connection to obtain training negative samples, and assign a second pseudo-label to the training negative samples: ; Among them, is the label of the th training negative sample, is the class label of the th sample in the training features of RGB modality images, is the class label of the th sample in the training features of NI modality images, is the class label of the th sample in the training features of TI modality images, , At least two of the values are different.
Citation Information
Patent Citations
Cross-modal pedestrian re-identification method in combination with local threshold binarization image
CN113723236A
Novel multi-modal fusion pedestrian re-identification algorithm
CN114694089A
Incomplete multi-modal medical image learning method
CN117218453A
Cross-modal pedestrian re-identification method based on auxiliary modal enhancement and multi-scale feature fusion
CN117994822A
Incomplete multi-mode pedestrian re-identification method and system
CN118570878A
Cited By
Multi-modal target re-identification method and device, electronic equipment and medium
CN122416206A