Cross-modal pedestrian re-identification method based on dynamic redundancy suppression

By using dynamic redundancy suppression technology in the cross-modal pedestrian re-identification method to adjust and fuse visible light and infrared image features, the problem of inaccurate identification of the prior art at night or under insufficient light is solved, and more accurate pedestrian identity feature extraction and recognition is achieved.

CN119919962AActive Publication Date: 2025-05-02CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411816586.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-02
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Existing cross-modal pedestrian re-identification methods are difficult to provide accurate pedestrian appearance information at night or under insufficient lighting conditions, and are difficult to deal with feature differences in complex scenarios.

Method used

The cross-modal pedestrian re-identification method based on dynamic redundancy suppression is adopted, and the low-level features of visible light and infrared images are extracted through two independent convolutional processes. Combined with the inner modal and cross-modal feature learners, the proportion of different modal feature information is dynamically adjusted, auxiliary modal features are generated, and redundant information is eliminated through the dynamic redundancy suppression module to build a reasonable shared feature space.

Benefits of technology

It effectively alleviates the internal and external differences of visible infrared modes, realizes the full exploration of pedestrian identity characteristics, and improves the recognition accuracy in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919962A_ABST
    Figure CN119919962A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal pedestrian re-identification method based on dynamic redundancy suppression. The cross-modal pedestrian re-identification method comprises the following steps: preprocessing a retrieval target and a visible light pedestrian image and an infrared pedestrian image of a matching database; inputting to the optimized dynamic redundancy suppression cross-modal model to obtain recognition features, and obtaining a recognition matching result according to the similarity between the recognition features of the retrieval target and the recognition features of the matching database; the dynamic redundancy suppression cross-modal model firstly carries out convolution processing on an input image to obtain visible light features and infrared features, channel attention information is calculated through two internal modal feature learners and a cross-modal feature learner, and fusion is carried out to obtain auxiliary modal features; and carrying out global average pooling and batch normalization through a weight sharing feature encoder to obtain identification features. According to the method, the difference inside and outside the visible light infrared mode is effectively relieved, and the pedestrian identity features are fully mined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a pedestrian re-recognition method. Background Art

[0002] The goal of pedestrian re-identification is to find specific pedestrians in the image gallery taken by different cameras at different times. This task has a wide range of applications in video surveillance, intelligent security and other fields. Previous recognition methods mainly achieved the matching of pedestrian images captured by visible light cameras. However, at night or under insufficient lighting conditions, visible light cameras cannot provide accurate information about the appearance of pedestrians. Therefore, the prior art proposes a cross-modal (visible light infrared) pedestrian re-identification (VI-ReID) method.

[0003] One of the existing methods is to use a two-stream network to learn multimodal shared features, mapping visible light and infrared features into a common embedding space, which alleviates the modal differences to a certain extent. Due to the significant difference in modalities, it is difficult for this method to directly project cross-modal images into a common feature space. Another method relies on modality translation technology, which uses generative adversarial network (GAN) technology to convert visible light images into infrared images or vice versa, thereby reducing the difference between modalities. However, these methods usually introduce additional image noise and significant computational costs, making them difficult to apply in practical scenarios.

[0004] Using the auxiliary modality approach, visible, infrared, and auxiliary modality features are compactly integrated in a common space to alleviate the differences between visible and infrared modalities, but there is still the problem of difficulty in handling features in complex scenes. Summary of the invention

[0005] In view of the defects of the above-mentioned prior art, the present invention provides a cross-modal pedestrian re-identification method based on dynamic redundancy suppression, which makes full use of modal-specific features to construct a more reasonable shared feature space, while strengthening information interaction between different levels, eliminating the influence of redundant information, effectively bridging the differences within and between modalities, and fully mining the identity information of pedestrians, solving the problem that previous methods were difficult to handle in complex scenes.

[0006] The technical solution of the present invention is as follows:

[0007] A cross-modal person re-identification method based on dynamic redundancy suppression includes the following steps:

[0008] Step S1, preprocessing the visible light pedestrian image and infrared pedestrian image of the retrieval target and the matching database;

[0009] Step S2, the pre-processed visible light pedestrian image and infrared pedestrian image are input into the optimized dynamic redundancy suppression cross-modal model to obtain recognition features;

[0010] Step S3, obtaining an identification and matching result based on the similarity between the retrieval target and the identification features of the matching database;

[0011] Among them, the dynamic redundancy suppression cross-modal model first performs convolution processing on the input image to obtain visible light features and infrared features; the visible light features are passed through a visible light intramodal feature learner to obtain visible light image features, and the infrared features are passed through an infrared intramodal feature learner to obtain infrared image features; the visible light image features and the infrared image features are calculated in the cross-modal feature learner for channel attention information and fused to obtain auxiliary modal features; the visible light image features, the infrared image features and the auxiliary modal features are respectively passed through a weight-shared feature encoder and then globally averaged pooled and batch normalized to obtain the recognition features;

[0012] The feature encoder includes a four-stage residual network module of ResNet50, and a dynamic redundancy suppression module is set after the first three residual network modules to perform channel-space feature changes and feature fusion on the features of the previous stage and the current stage.

[0013] Furthermore, the dynamic redundancy suppression module performs channel-space feature changes and feature fusion on the features of the previous stage and the current stage, including using three convolution layers get f l represents the characteristics of the previous stage of the current stage, f h Represents the current stage features, and then calculates the channel similarity matrix M through matrix multiplication c , and then through the convolutional layer ε w Restore and add the matrix to f h Add to get the output feature f′ of the feature encoder h .

[0014] Further,

[0015] Furthermore, the convolutional layer ε w The convolution kernel is 1×1.

[0016] Furthermore, the process of obtaining the visible light image features through the visible light intramodal feature learner is as follows:

[0017] F vis =Conv 1×1 (Conv 3×3 (x vis )+Conv 5×5 (x vis )+Conv 7×7 (xvis )),

[0018] x vis is the visible light feature, Conv represents convolution, F vis is the visible light image feature;

[0019] The process of obtaining the infrared image features by the infrared features through the infrared intramodal feature learner is as follows:

[0020] F ir =Conv 1×1 (Conv 3×3 (x ir )+Conv 5×5 (x ir )+Conv 7×7 (x ir )),

[0021] x ir is the infrared feature, F ir Infrared image features.

[0022] Furthermore, the process of calculating the channel attention information and fusing the auxiliary modality features in the cross-modal feature learner is:

[0023]

[0024] where F vis is the visible light image feature, F ir is the infrared image feature, GAP is the global average pooling, FC is the fully connected, v Fully connected on the visible light image feature side, FC i is the full connection on the feature side of the infrared image, F aux Auxiliary modal features.

[0025] Furthermore, the dynamic redundancy suppression cross-modal model first performs convolution processing on the input image to obtain visible light features and infrared features. The pre-processed visible light pedestrian image and infrared pedestrian image are extracted through the Stage 0 stage of the independent dual-stream Resnet50 respectively.

[0026] Furthermore, the optimization process in the optimized dynamic redundancy suppression cross-modal model is to use the recognition features and the real pedestrian labels to calculate the loss and perform reverse propagation according to the loss value, and the loss calculation includes the trimodal identity center loss.

[0027]

[0028] Where α is the threshold boundary parameter, represents the identity center of the p-th (p≠q)th person's auxiliary modality, represents the identity center of the p-th person’s visible light modality, represents the identity center of the visible light modality of the qth person, represents the identity center of the p-th individual infrared modality, Represents the identity center of the qth person’s infrared modality. max(·) represents the distance between the anchor point and the maximum positive sample, and min(·) represents the distance between the anchor point and the minimum negative sample.

[0029] Furthermore, the calculation loss is For identity loss, is the triplet loss, and λ1 is the weight parameter.

[0030] Furthermore, the preprocessing includes random cropping, horizontal flipping and random erasing.

[0031] Compared with the prior art, the present invention has the following advantages:

[0032] The dynamic redundancy suppression cross-modal model in the present invention utilizes two independent convolution processes to extract low-level features of visible light images and infrared images; then, two endomodal feature learners are combined with a cross-modal feature learner to dynamically adjust the proportion of specific feature information of different modalities while strengthening modality-specific features, thereby generating effective auxiliary modal features; then, by constructing a dynamic redundancy suppression module, the dynamic redundancy suppression module is embedded into the backbone network to further eliminate the influence of auxiliary modal redundant information, and the three modal features are input into a feature encoder with weight sharing for feature extraction and fusion. In the process of model optimization, identity loss and trimodal identity center loss are calculated based on pedestrian identity features and modal features and optimized according to the loss function. By adopting the present invention, the internal and external differences of visible light infrared modalities are effectively alleviated, and the full mining of pedestrian identity features is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a flow chart of the cross-modal pedestrian re-identification method based on dynamic redundancy suppression of the present invention.

[0034] Figure 2 Schematic diagram of the dynamic redundancy suppression cross-modal model of the present invention.

[0035] Figure 3 Schematic diagram of the intramodal feature learner of the present invention.

[0036] Figure 4 Schematic diagram of the cross-modal feature learner of the present invention.

[0037] Figure 5 Schematic diagram of the tri-modal identity center loss function of the present invention. DETAILED DESCRIPTION

[0038] The present invention will be further described below in conjunction with the embodiments, but are not intended to limit the present invention.

[0039] Please combine Figures 1 to 4 As shown, the cross-modal pedestrian re-identification method based on dynamic redundancy suppression according to an embodiment of the present invention comprises the following steps:

[0040] Step S1, preprocessing the visible light pedestrian image and infrared pedestrian image of the retrieval target and the matching database;

[0041] Step S2, the pre-processed visible light pedestrian image and infrared pedestrian image are input into the optimized dynamic redundancy suppression cross-modal model to obtain recognition features;

[0042] Step S3, obtaining the recognition matching result based on the similarity of the recognition features of the search target and the matching database, specifically calculating the cosine similarity between the pedestrian target and the pedestrian search database features, and sorting the results in descending order to obtain Rank-1 as the best matching result.

[0043] The following is an explanation of the dynamic redundancy suppression cross-modal model in conjunction with the model optimization process. When training and optimizing the dynamic redundancy suppression cross-modal model, the visible light images, infrared images and identity labels of pedestrians are obtained from the existing data sets SYSU-MM01 and RegDB, and they are divided into training sets, validation sets and test sets. The training and images are preprocessed by horizontal flipping, random erasing and other preprocessing operations. At the same time, random cropping is used to obtain images of size 288*144, and the cropped images are randomly flipped horizontally.

[0044] The visible light pedestrian image and infrared pedestrian image input to the dynamic redundancy suppression cross-modal model are extracted through the Stage 0 of the independent dual-stream Resnet50. The visible light features are obtained by the visible light intramodal feature learner, and the infrared features are obtained by the infrared intramodal feature learner.

[0045] The following process is performed in the intramodal feature learner:

[0046] F vis =Conv 1×1 (Conv 3×3 (x vis )+Conv 5×5 (x vis )+Conv 7×7 (x vis )),

[0047] F ir =Conv 1×1 (Conv3×3 (x ir )+Conv 5×5 (x ir )+Conv 7×7 (x ir )),

[0048] x vis is the visible light characteristic, x ir is the infrared feature, Conv represents convolution, F vis is the visible light image feature obtained, F ir is the infrared image feature obtained.

[0049] The following process is performed in the cross-modal feature learner,

[0050]

[0051] Among them, F vis is the visible light image feature, F ir is the infrared image feature, GAP is the global average pooling, FC is the fully connected, v Fully connected on the visible light image feature side, FC i is the full connection on the feature side of the infrared image, represents the attention weight of the visible light modality channel, represents the infrared modality channel attention weight, F aux Auxiliary modal features obtained by the cross-modal feature learner.

[0052] The features of the three modes, namely visible light features, infrared features and auxiliary modality features, are input into the weight-shared feature encoder for processing to build a more reasonable shared feature space. By extracting effective information layer by layer and filtering redundant information, a wider range of contextual information is captured, the influence of redundant information at the current stage is eliminated, and the final output features are ensured to be more refined.

[0053] Please combine Figure 2 As shown in the figure, the feature encoder includes a four-stage residual network module of ResNet50, and a dynamic redundancy suppression module is set after the first three residual network modules to perform channel-space feature changes and feature fusion on the features of the previous stage and the current stage. The dynamic redundancy suppression module processes the features as follows:

[0054]

[0055] That is, the channel-space feature changes and feature fusion of the previous stage features and the current stage features include the use of three 1×1 convolutional layers get f l represents the characteristics of the previous stage of the current stage, fh Represents the current stage features, and then calculates the channel similarity matrix M through matrix multiplication c , and then through a 1×1 convolutional layer ε w Restore and add the matrix to f h Add to get the output feature f′ of the feature encoder h .

[0056] After global average pooling and batch normalization of the feature results output by the feature encoder, the loss value is calculated according to the pedestrian label and back-propagated for model optimization. Assume that there are P×K images in each small batch, where P is the number of pedestrian identities and K is the number of images contained in each identity. The center of pedestrians with the same identity in the same modality is calculated as follows:

[0057]

[0058] in, Represents the pedestrian feature vector representation obtained after global average pooling of the K-th image.

[0059] Then the loss is calculated with the real pedestrian label and the loss value is reversed, where the three-modal identity center loss and the overall loss are calculated as follows:

[0060]

[0061]

[0062] Where α is the threshold boundary parameter, represents the identity center of the p-th (p≠q)th person's auxiliary modality, represents the identity center of the p-th person’s visible light modality, represents the identity center of the visible light modality of the qth person, represents the identity center of the p-th individual infrared modality, Represents the identity center of the qth person’s infrared modality. max(·) represents the distance between the anchor point and the maximum positive sample, and min(·) represents the distance between the anchor point and the minimum negative sample. Indicates identity loss, represents the triplet loss, and λ1 is the weight parameter. The goal is to narrow the distance between the auxiliary modality identity center and the same identity center of visible light / infrared, and push the distance between the auxiliary modality identity center and the different identity centers of visible light / infrared, so as to suppress cross-modal changes while ensuring the discriminability of features.

[0063] Comparing the method of the present invention with other prior art methods, the results are as follows:

[0064] The comparative experimental results of SYSU-MM01 are shown in Table 1:

[0065] Table 1 Accuracy comparison of different methods on the SYSU-MM01 dataset

[0066]

[0067] The RegDB comparison experiment results are shown in Table 2:

[0068] Table 2 Comparison of accuracy of different methods on RegDB dataset

[0069]

[0070]

[0071] It can be seen from the results of the above two tables that the proposed method has achieved good performance on the two mainstream cross-modal pedestrian re-identification datasets SYSU-MM01 and RegDB.

Claims

1. A cross-modal person re-identification method based on dynamic redundancy suppression, characterized in that: The following steps are involved: Step S1, preprocessing the visible light pedestrian image and infrared pedestrian image of the retrieval target and the matching database; Step S2, the pre-processed visible light pedestrian image and infrared pedestrian image are input into the optimized dynamic redundancy suppression cross-modal model to obtain recognition features; Step S3, obtaining an identification and matching result based on the similarity between the retrieval target and the identification features of the matching database; Among them, the dynamic redundancy suppression cross-modal model first performs convolution processing on the input image to obtain visible light features and infrared features; the visible light features are passed through a visible light intramodal feature learner to obtain visible light image features, and the infrared features are passed through an infrared intramodal feature learner to obtain infrared image features; the visible light image features and the infrared image features are calculated in the cross-modal feature learner for channel attention information and fused to obtain auxiliary modal features; the visible light image features, the infrared image features and the auxiliary modal features are respectively passed through a weight-shared feature encoder and then globally averaged pooled and batch normalized to obtain the recognition features; The feature encoder includes a four-stage residual network module of ResNet50, and a dynamic redundancy suppression module is set after the first three residual network modules to perform channel-space feature changes and feature fusion on the features of the previous stage and the current stage.

2. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 1 is characterized in that: The dynamic redundancy suppression module performs channel-space feature changes and feature fusion on the features of the previous stage and the current stage, including using three convolutional layers get f l represents the characteristics of the previous stage of the current stage, f h Represents the current stage features, and then calculates the channel similarity matrix M through matrix multiplication c , and then through the convolutional layer ε w Restore and add the matrix to f h Add to get the output feature f′ of the feature encoder h .

3. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 2 is characterized in that:

4. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 2 is characterized in that: The convolutional layer ε w The convolution kernel is 1×1.

5. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 1, characterized in that: The process of obtaining the visible light image features by the visible light features through the visible light intramodal feature learner is as follows: F vis =Conv 1×1 (Conv 3×3 (x vis )+Conv 5×5 (x vis )+Conv 7×7 (x vis )), x vis is the visible light feature, Conv represents convolution, F vis is the visible light image feature; The process of obtaining the infrared image features by the infrared features through the infrared intramodal feature learner is as follows: F ir =Conv 1×1 (Conv 3×3 (x ir )+Conv 5×5 (x ir )+Conv 7×7 (x ir )), x ir is the infrared feature, F ir Infrared image features.

6. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 1, characterized in that: The process of calculating channel attention information and fusing auxiliary modal features in the cross-modal feature learner is: where F vis is the visible light image feature, F ir is the infrared image feature, GAP is the global average pooling, FC is the fully connected, v Fully connected on the visible light image feature side, FC i is the full connection on the feature side of the infrared image, F aux Auxiliary modal features.

7. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 1, characterized in that: The dynamic redundancy suppression cross-modal model first performs convolution processing on the input image to obtain visible light features and infrared features. The pre-processed visible light pedestrian image and infrared pedestrian image are extracted through the Stage 0 stage of the independent dual-stream Resnet50 respectively.

8. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 1, characterized in that: The optimization process in the optimized dynamic redundancy suppression cross-modal model is to use the recognition features and the real pedestrian labels to calculate the loss and perform reverse propagation according to the loss value, and the loss calculation includes the trimodal identity center loss Where α is the threshold boundary parameter, represents the identity center of the p-th (p≠q)th person's auxiliary modality, represents the identity center of the p-th person’s visible light modality, represents the identity center of the visible light modality of the qth person, represents the identity center of the p-th individual infrared modality, Represents the identity center of the qth person’s infrared modality. max(·) represents the distance between the anchor point and the maximum positive sample, and min(·) represents the distance between the anchor point and the minimum negative sample.

9. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 8, characterized in that: The calculated loss is For identity loss, is the triplet loss, and λ1 is the weight parameter.

10. The cross-modal person re-identification method based on dynamic redundancy suppression according to claim 1, characterized in that: The preprocessing includes random cropping, horizontal flipping and random erasing.

Citation Information

Patent Citations

  • Pedestrian re-identification method, device and equipment and storage medium

    CN109711316A

  • Near infrared-visible light cross-modal double-current pedestrian re-identification method and system

    CN114220124A

  • Visible light-infrared pedestrian re-identification method based on multi-stage auxiliary learning

    CN117576729A

  • Cross-modal pedestrian re-identification method for modal enhancement and compensation

    CN117746467A

  • Cross-modal pedestrian re-identification method based on auxiliary modal enhancement and multi-scale feature fusion

    CN117994822A