Cross-modal pedestrian re-identification method based on intermediate feature guidance

Image features are extracted through the ResNet50 network and feature mining module, combined with intermediate feature fusion and loss function training network, the problem of modal gap in cross-modal pedestrian re-identification is solved, and a higher recognition accuracy is achieved.

CN120340070AActive Publication Date: 2025-07-18NANTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510506441.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing cross-modal pedestrian re-identification method fails to effectively alleviate the modal gap between visible light and infrared images, resulting in a decrease in recognition accuracy, especially in the case of poor lighting conditions.

Method used

Image features are extracted by ResNet50 network and feature mining module, and the network is trained by combining intermediate feature fusion and three-modal triple loss function, central deviation loss function and id loss function to alleviate modal differences and improve recognition accuracy.

Benefits of technology

It significantly improves the accuracy of cross-modal pedestrian re-identification and improves the recognition ability of the model under different lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340070A_ABST
    Figure CN120340070A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of computer vision, and particularly relates to a cross-modal pedestrian re-identification method based on intermediate feature guidance. According to the method, ResNet50 serves as a network of a backbone network, a novel feature mining and fusion module is provided, according to the phenomenon that cross-modal pedestrian re-identification multi-modal data deviation is large, the mode difference is reduced by means of an intermediate mode, a three-mode triple loss function and a center deviation loss function are provided, and the cross-modal pedestrian re-identification multi-modal data deviation is improved. And a common id loss function is combined to jointly train the network, so that the recognition accuracy of the model is effectively improved. According to the method, the mode difference is relieved by utilizing the middle feature containing the two original modes at the same time, and the identification accuracy of the model is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a cross-modal pedestrian re-identification method guided by intermediate features. Background Art

[0002] With the rapid development of social economy, public safety has become increasingly prominent. Nowadays, surveillance devices have been widely used in all aspects of daily life. However, the application of surveillance information is still limited to a relatively basic level. In emergency situations, manually viewing surveillance videos to identify and retrieve suspects is a time-consuming and laborious task, especially in complex shooting environments, and the accuracy of human eye resolution is relatively low.

[0003] Therefore, the industry has begun to actively explore the pedestrian re-identification (Person Re-identification, Re-ID) technology, which has become an important part of the intelligent security field. Pedestrian re-identification is a technology aimed at identifying specific pedestrians from a large number of pedestrian images captured by different cameras to meet various search needs. The task of person re-identification involves matching images or videos of relevant individuals captured by multiple non-overlapping cameras. Due to its wide range of practical applications, such as criminal investigation, danger warning, unmanned supermarket management, missing person rescue and many other fields, this task has become crucial. Pedestrian re-identification has become a hot topic in the scientific research field and has received extensive attention.

[0004] Early pedestrian re-identification technologies were limited to processing visible light pedestrian images. However, in practical applications, the single-modal pedestrian re-identification technology alone cannot meet the diverse application scenario requirements. For example, some criminals often commit crimes at night or choose locations with poor lighting conditions to hide or escape. In low-light conditions, visible light cameras are difficult to capture clear pedestrian images. Therefore, the accuracy of single-modal pedestrian re-identification technology drops significantly. To solve this problem, surveillance devices have been upgraded. The new generation of surveillance devices are equipped with both visible light and infrared cameras, and can automatically switch the shooting mode according to the lighting conditions, thus overcoming the limitation of lighting conditions on identification. This technology upgrade makes the surveillance system more adaptable to different environments and situations, and greatly improves the reliability of identification.

[0005] However, this upgrade also brings new challenges, that is, it is necessary to process two different modal images simultaneously, which greatly increases the complexity of the original single-modal pedestrian re-identification task. Therefore, in order to ensure that the pedestrian re-identification technology can play a role in various scenarios, the academic community has begun to study the "visible light-infrared" cross-modal pedestrian re-identification task (Visible-Infrared Person Re-Identification, VI-ReID).

[0006] Visible-Infrared Person Re-Identification (V-I ReID) is a more challenging variant of the Person Re-Identification (ReID) task, involving image matching across visible light (RGB) and infrared (IR) cameras. The challenge of V-I ReID lies in the need to match individuals with significant appearance differences between two different modalities. For example, visible light images capture color information, while infrared images provide information about body temperature and are less affected by changes in lighting conditions. State-of-the-art V-I ReID systems typically use backbone models such as convolutional neural networks (CNNs) or vision transformers (ViTs) to extract features from visible light and infrared images, and then match these features before cross-modal feature fusion.

[0007] The SYSU-MM01 dataset was created in the 2017 paper "RGB-infrared cross-modality person re-identification". SYSU-MM01 is the first dataset specifically designed for the cross-modal pedestrian re-identification task. The dataset includes pedestrian images captured by indoor and outdoor visible light cameras, as well as images captured by infrared cameras. The training set contains 395 different pedestrian identities, including 22,258 visible light images and 11,909 infrared images. The test set consists of images of 96 different pedestrians and can be evaluated in two different modes: the All-Search mode and the Indoor-Search mode.

[0008] In recent years, many cross-modal pedestrian re-identification methods have been proposed.

[0009] For example, the paper "Visible-infrared person re-identification using privileged intermediate information." noticed that the modality gap was too large, so it narrowed the modality gap by alignment. It used multiple modules to extract fine-grained features, enhanced cross-modal discrimination ability, and then aligned the features between different modalities. The modality elimination module removed modality-specific information, making the model focus more on the shared identity features across modalities. The center clustering loss optimized identity learning and detailed feature learning, improving feature discrimination. As Figure 9 shown.

[0010] Compared with the method in this paper, although it notices the modality gap between different modalities, it only uses the original two modalities for cross-modal information alignment, which is difficult to achieve and is not conducive to mining the common information of pedestrians in different modalities, greatly affecting its accuracy. Existing pedestrian re-identification methods ignore the modality gap between different modalities, which is not conducive to mining the common information of pedestrians in different modalities and will greatly affect the accuracy of cross-modal pedestrian re-identification. Summary of the Invention

[0011] The purpose of the present invention is to overcome the shortcomings of the existing technology and propose a cross-modal pedestrian re-identification method guided by intermediate features, so as to improve the accuracy of model recognition.

[0012] The technical solution adopted by the method of the present invention is as follows: a cross-modal pedestrian re-identification method guided by intermediate features, including the following steps:

[0013] Step 1: Preprocess P*K visible light images and infrared images, and enter Step 2;

[0014] Step 2: Input the P*K preprocessed visible light images and P*K infrared images obtained in Step 1 into the feature extraction network, and enter Step 3;

[0015] Step 3: Use the ResNet50 network and the feature mining module to generate informative image features based on the images input in Step 2, and enter Step 4;

[0016] Step 4: Dot-multiply and fuse the image features of the two branches to form intermediate features, and input the intermediate features together with the features of the original two branches into the feature fusion module to obtain new intermediate features, and enter Step 5;

[0017] Step 5: Input the image features and intermediate features of the two branches obtained in Step 3 and Step 4 into the loss function, including the id loss, the three-modal triplet loss function and the center deviation loss function, so as to train a network with good performance, and enter Step 6;

[0018] Step 6: If the specified number of training rounds is reached, proceed to Step 7; otherwise, continue with the training and return to Step 1;

[0019] Step 7: End.

[0020] Further, the steps for extracting the preliminary features of the pictures in Step 2 are as follows:

[0021] Step 2-1: Input the visible light image into the feature extraction network to extract its preliminary features;

[0022] Step 2-2: Input the infrared image into the feature extraction network to extract its preliminary features;

[0023] Step 2-3: Output features.

[0024] Furthermore, in Step 2, the first layer of the Residual Network is used for feature extraction to fully perceive the shallow information of the image, and the feature extraction network parameters for visible light images and infrared images are not shared.

[0025] Furthermore, the steps for extracting image features in Step 3 are as follows:

[0026] Step 3-1: Input the features extracted in Step 2 into a network composed of the convolutional layers of the last four layers of ResNet50 and the attention block of the feature mining module to obtain information-rich image features;

[0027] Step 3-2: Perform MAX-pooling on the features in Step 3-1 and proceed to Step 3-3;

[0028] Step 3-3: Output the pooled features of the image.

[0029] Furthermore, in Step 3, in the feature mining module, a spatial attention module and a dimensional attention module are introduced. The spatial attention module focuses on the dependence between local features and spatial positions to enhance the perception ability of the target area; while the dimensional attention module focuses on channel information, adjusts the importance of different feature channels, enables the network to dynamically allocate weights, and improves the expression ability of key features;

[0030] Input the input feature f0 into two different attention modules respectively. Utilize spatial attention and dimensional attention to obtain attention features with different information emphases, then perform a dot product operation with the input feature, and perform activation and batch normalization; then perform a dot product operation on the features of the two branches to obtain the feature f1 that fuses spatial and dimensional attention;

[0031] For the obtained feature f1, perform global max pooling and global average pooling respectively, and perform a concatenation operation on the features to obtain a feature vector with a dimension of 2. To further enhance the discriminative ability of the model, perform a matrix multiplication with the feature that fuses spatial and dimensional attention, so that the output feature can fully integrate the information from different attention mechanisms, and finally obtain a more abundant representative feature f2 with the same dimension as the original input feature;

[0032] Finally, during the training process of the deep network, the problems of gradient disappearance and gradient explosion will lead to unstable model convergence; to avoid gradient descent, a residual connection mechanism is introduced in the final feature generation process, that is, add f2 and the original input feature element by element.

[0033] Further, in step 4, the intermediate feature and the features of the original two branches are jointly input into the feature fusion module to obtain a new intermediate feature, achieving the purpose of deeply fusing the two-modal features in the intermediate feature. The specific steps are as follows:

[0034] Step 4-1: Multiply the visible light image feature F vi extracted in step 3 with the transposed intermediate feature F mid by matrix multiplication, then after relu activation, multiply again with the visible light image feature F vi by matrix multiplication, and then perform bitwise addition with the intermediate feature F mid to obtain F fused1 ;

[0035] Step 4-2: Input the infrared image feature F in extracted in step 3 into the SelfAttention network to obtain the feature F in1 ; Multiply the feature F in1 with the transposed feature F fused1 by matrix multiplication, then after relu activation, multiply again with the feature F in1 by matrix multiplication, and then perform bitwise addition with the intermediate feature F fused1 to obtain the feature F fused2 ;

[0036] Step 4-3: Input the feature F fused2 into the SelfAttention network to obtain the output feature F fused ;

[0037] Step 4-4: Perform BN normalization on the output feature F fused .

[0038] Further, in step 5, the id loss function is used to jointly constrain the model. The specific steps are as follows:

[0039] Step 5-1: Pool the features extracted in step 3;

[0040] Step 5-2: Perform BN normalization on the features extracted in step 5-1;

[0041] Step 5-3: Pass the normalized features in step 5-2 through a fully connected layer to reduce the dimension of the information while obtaining the dimension information corresponding to the category, and enter step 5-4;

[0042] Step 5-4: Then input the features into the id loss function;

[0043] Step 5-5: Calculate the gradient using the losses in 4-4 and 5-4, perform backpropagation, and update the model parameters;

[0044] During training, P identities are selected from the training set, and K pedestrian images are randomly selected from each identity. There are 2*P*K pedestrian images in each Batch. The identity loss for the pedestrian re-identification task is shown in the following formula:

[0045]

[0046] Among them, y i,k represents whether the identity of the i-th image is k, N represents the total number of pedestrian categories in the dataset, and p i,k Indicates the probability that the identity of the i-th image is k.

[0047] Furthermore, the trimodal triplet loss function in step 5 uses the image features of the intermediate modality as anchor points, and respectively brings the farthest images of the same identity closer and pushes the closest images of different identities farther away for the two original modality images;

[0048] For the extracted trimodal data, it is necessary to make full use of the generated intermediate modal features to achieve a more efficient recognition of pedestrian identities. For this purpose, a trimodal triplet loss is proposed specifically. The common difficult triplet loss is a loss function widely used in pedestrian re-identification tasks. It brings the farthest positive pairs closer and pushes the closest negative pairs farther away. In simple terms, the image features of pedestrians of the same type taken by different cameras are closer, and the image features of different pedestrians are pushed farther away.

[0049] The existing difficult cross-modal triplet loss, namely the difficult bimodal triplet loss;

[0050] The same shape represents the same category, that is, pictures of different modes of the same pedestrian identity. Different shapes represent pictures of different categories, that is, pictures of different pedestrian identities. For each picture in each batch, the distance between it and the picture is calculated, and the positive sample farthest from the sample and the negative sample closest to the sample are selected to calculate the loss, as shown in formula (2):

[0051]

[0052] Where A is the anchor sample, (A,p) is the sample with the same identity category as the anchor sample, and (A,n) is the sample with a different identity category from the anchor sample. is the threshold parameter, Calculate the Euclidean distance, representing the feature f i With f i The Euclidean distance is calculated as follows:

[0053]

[0054] The three-modal triplet loss function uses the image features of the middle modality as the anchor points, and for the two original modality images, it pulls the farthest images of the same identity closer and pushes the closest images of different identities farther away;

[0055] For each intermediate feature sample in each batch, it will be used as an anchor sample, and the loss will be calculated with the farthest positive sample and the closest negative sample in the two original modality images to this sample, as shown in Equation (4):

[0056]

[0057] In the formula, A is the anchor sample, (A, p) is the positive sample, (A, n) is the negative sample, and ε is the threshold parameter.

[0058] Furthermore, for the center deviation loss function in step 5, considering the extremely strong expression ability of visible light images and the pre-training of the ResNet50 network parameters with visible light images, which will cause the model to have a strong tendency towards visible light images. To avoid the influence of this factor on the middle modality, a center deviation loss function will be designed for the feature centers of the three modalities for constraint;

[0059] All first find the feature centers for each category of pedestrians in the three modalities, and constrain the middle feature to be at the center of the two original features, and also constrain the distance from the middle feature to the two original features to be equal;

[0060] For each image with a different identity ID in each batch, the feature center in this category will be calculated separately, and then the visible light feature F vi and the infrared feature F i will be obtained, and the distance will be calculated with the middle feature center, and the formulas are as shown in Equations (5) and (6):

[0061]

[0062]

[0063] In the formula, Batch i and Batch vi represent the visible light feature and the infrared feature in the batch respectively, K represents the number of images of each sample in each modality in this batch, P represents the number of sample categories in this batch, c final_i represents the center of the i-th intermediate feature, and c mid_i is the center of the visible light feature center and the infrared feature center of the i-th category;

[0064] Then the total network loss is expressed as shown in Equation (7):

[0065] L all =Lid +λ1L Tri_Tri_Loss +λ2L Center_Bias_Loss (7)

[0066] Among them, λ1 and λ2 are hyperparameters, which are set separately in the experiment.

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] (1) The present invention is a network with ResNet50 as the backbone network, and a novel feature mining and fusion module is proposed. According to the phenomenon of large multi-modal data deviation in cross-modal person re-identification, it is proposed to use the intermediate modality to narrow the modality gap, and a three-modal triplet loss function and a central deviation loss function are proposed to jointly train the network with the common id loss function, effectively improving the recognition accuracy of the model.

[0069] (2) The present invention proposes to use the intermediate features that simultaneously contain two original modalities to alleviate the modality difference, effectively improving the recognition accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.

[0071] Figure 1 It is a pedestrian image in the color modality and the infrared modality in the SYSU-MM01 dataset of the present invention;

[0072] Figure 2 It is a schematic diagram of the person re-identification network framework of the present invention;

[0073] Figure 3 It is a flowchart of the training stage of the present invention;

[0074] Figure 4 It is a structural diagram of the feature fusion module of the present invention;

[0075] Figure 5 It is a schematic diagram of the cross-modal triplet of the present invention;

[0076] Figure 6 It is a schematic diagram of the three-modal triplet loss function of the present invention;

[0077] Figure 7 It is a schematic diagram of the modality deviation of the present invention;

[0078] Figure 8 It is a structural diagram of the feature mining module of the present invention;

[0079] Figure 9 It is a model schematic diagram of the comparison method of the present invention. Detailed Implementation Modes

[0080] The following will further explain and illustrate the present invention in detail with reference to the accompanying drawings, so that those skilled in the art can understand the present invention more deeply and be able to implement it. However, the following is only for explaining the present invention by referring to examples and does not limit the present invention.

[0081] The present invention proposes a cross-modal pedestrian re-identification method guided by intermediate features. The framework diagram of this network is as Figure 2 shown.

[0082] The training flow chart of the network proposed by the present invention is as Figure 3 shown. This training process is carried out in the form of mini-batch training. Each batch will randomly select P pedestrians, and randomly select K visible light and infrared images for each of these pedestrians. Then there are a total of 2P*K pedestrian images. Next, taking the input of one image as an example, the training process will be introduced as follows:

[0083] Step 1: Perform data preprocessing on P*K visible light images and infrared images, and enter Step 2.

[0084] Step 2: Input the P*K preprocessed visible light images and P*K infrared images obtained in Step 1 into the feature extraction network, and enter Step 3;

[0085] Step 3: Use the ResNet50 network and the feature mining module that adds the third layer and the fourth layer of the ResNet50 network to generate informative image features according to the images input in Step 2, and enter Step 4;

[0086] Step 4: Perform dot product fusion on the image features of the two branches to form intermediate features, and input the intermediate features together with the original feature images of the two branches into the feature fusion module to obtain new intermediate features, and enter Step 5;

[0087] Step 5: Input the image features and intermediate features of the two branches obtained in Steps 3 and 4 into the loss function, including the id loss, the tri-modal triplet loss function, and the center deviation loss function. Thus, a network with good performance is trained, and enter Step 6;

[0088] Step 6: If the specified number of training rounds is reached, then perform Step 7; otherwise, continue to complete the training and return to Step 1;

[0089] Step 7: End.

[0090] For the images input in Step 1, three data preprocessing methods of random cropping, horizontal flipping, and erasing are all considered.

[0091] The steps of extracting the initial features of the image in step 2 of the training process of the present invention are as follows:

[0092] Step 2-1: Input the visible light image into the feature extraction network to extract its initial features;

[0093] Step 2-2: Input the infrared image into the feature extraction network to extract its initial features;

[0094] Step 2-3: Output the features.

[0095] In step 2, the present invention uses the first layer network of the Residual Network to extract features, fully perceiving the shallow information of the image, and the feature extraction network parameters of the visible light image and the infrared image are not shared.

[0096] The steps of extracting the image features in step 3 of the training process of the present invention are as follows:

[0097] Step 3-1: Input the features extracted in step 2 into a network composed of the convolutional layers of the last four layers of ResNet50 and the attention block of the feature mining module to obtain rich-information image features;

[0098] Step 3-2: Perform MAX-pooling on the features in step 3-1 and enter step 3-3;

[0099] Step 3-3: Output the pooled features of the image.

[0100] In step 3, the present invention uses the Residual Network and the feature mining module to extract features. Different from the feature extraction network in step 2, the parameters of the network in this stage are shared. In the feature mining module proposed by the present invention, the present invention introduces Spatial Attention and Channel Attention to make full use of the feature information and improve the discrimination ability and generalization performance of the model. Its structure is as Figure 8 shown.

[0101] The Spatial Attention Module (SAM) focuses on the dependence between local features and spatial positions to enhance the perception ability of the target area. At the same time, the Channel Attention Module (DAM) focuses on the channel information, adjusts the importance of different feature channels, enabling the network to dynamically allocate weights and improve the expression ability of key features.

[0102] The input feature f0 is respectively input into two different attention modules. Using spatial attention and dimensional attention, attention features with different information emphases are obtained, and then a dot product operation is performed with the input feature, followed by activation and batch normalization. Then, the features of the two branches are subjected to a dot product operation to obtain the feature f1 that fuses spatial and dimensional attention.

[0103] For the obtained feature f1, global max pooling and global average pooling are respectively performed, and a concatenation operation is performed on the features to obtain a feature vector with a dimension of 2. To further enhance the discriminative ability of the model, a matrix multiplication is performed with the feature that fuses spatial and dimensional attention, so that the output feature can fully integrate the information from different attention mechanisms, and finally a more abundant representative feature f2 with the same dimension as the original input feature is obtained.

[0104] Finally, during the training process of the deep network, the problems of gradient disappearance and gradient explosion may lead to unstable model convergence. To avoid gradient descent, a residual connection mechanism is introduced in the final feature generation process, that is, f2 and the original input feature are added element by element.

[0105] In step 4, the feature fusion module proposed by the present invention has the structure as Figure 4 shown. The intermediate feature and the features of the original two branches are jointly input into the feature fusion module to obtain a new intermediate feature, achieving the purpose of deeply fusing the two-modal features of the intermediate feature. The specific steps are as follows.

[0106] Step 4-1: Perform matrix multiplication on the visible light image feature F vi extracted in step 3 and the transposed intermediate feature F mid , then after relu activation, perform matrix multiplication with the visible light image feature F vi again, and then perform bitwise addition with the intermediate feature F mid to obtain F fused1 ;

[0107] Step 4-2: Input the infrared image feature F in extracted in step 3 into the self-attention network to obtain the feature F in1 . Perform matrix multiplication on the feature F in1 and the transposed feature F fused1 , then after relu activation, perform matrix multiplication with the feature F in1 again, and then perform bitwise addition with the intermediate feature F fused1 to obtain the feature F fused2 ;

[0108] Step 4-3: Input the feature F fused2 into the self-attention network to obtain the output feature F fused ;

[0109] Step 4-4: Output feature F fused Perform BN normalization.

[0110] In step 5, the present invention uses the id loss function to jointly constrain the model, and the specific steps are as follows.

[0111] Step 5-1: Pool the features extracted in step 3;

[0112] Step 5-2: Perform BN normalization on the features extracted in step 5-1.

[0113] Step 5-3: Pass the normalized features in step 5-2 through the fully connected layer to reduce the dimension of the information and obtain the dimensional information corresponding to the category, and then proceed to step 5-4;

[0114] Step 5-4: Input the features into the id loss function;

[0115] Step 5-5: Using the losses of 4-4 and 5-4, calculate the gradient, back propagate, and update the model parameters.

[0116] Specifically, during the training, the ID loss selects P identities in the training set, randomly selects K pedestrian images from each identity, and each batch contains 2*P*K pedestrian images. The identity loss for the pedestrian re-identification task is shown in the following formula.

[0117]

[0118] Among them, y i,k represents whether the identity of the i-th image is k, N represents the total number of pedestrian categories in the dataset, and p i,k Indicates the probability that the identity of the i-th image is k

[0119] In step 5, the present invention uses the trimodal triplet loss function and the center deviation loss function to jointly constrain the model.

[0120] For the extracted trimodal data, it is necessary to make full use of the generated intermediate modal features to achieve a more efficient recognition of pedestrian identities. In this regard, a trimodal triplet loss is proposed specifically. The common difficult triplet loss is a loss function widely used in pedestrian re-identification tasks. It brings the farthest positive pairs closer and pushes the closest negative pairs farther away. Simply put, the image features of pedestrians of the same type taken by different cameras are closer, and the image features of different pedestrians are pushed farther away.

[0121] The existing difficult cross-modal triplet loss, namely the difficult bimodal triplet loss, is shown in Figure 2. Figure 5 shown.

[0122] In the example figure, the green arrow represents pulling closer the features to increase their similarity, while the red arrow represents pushing away the features to decrease their similarity. The same shape represents the same category, that is, images of different modalities of the same pedestrian's identity. Different shapes represent different categories, that is, images of different pedestrians' identities. For each image in each batch, the distance between it and other images will be calculated, and the positive sample farthest from this sample and the negative sample closest to this sample will be selected to calculate the loss, as shown in Equation (2):

[0123]

[0124] In the formula, A is the anchor sample, (A, p) is the sample with the same identity category as the anchor sample, and (A, n) is the sample with a different identity category from the anchor sample. is the threshold parameter, and is used to calculate the Euclidean distance, representing the feature f i and f i The Euclidean distance between them is calculated by the formula as shown in Equation (3):

[0125]

[0126] Compared with the traditional cross-modal triplet loss, directly pulling closer and pushing away the features of the two modalities will affect the effect due to the large modality differences. The triplet loss function of the three modalities takes the feature of the middle modality image as the anchor, and respectively pulls closer the farthest image with the same identity and pushes away the closest image with different identities for the two original modality images. Its schematic diagram is as shown in Figure 6 shown.

[0127] For each middle feature sample in each batch, it will be used as the anchor sample, and the loss will be calculated by taking the positive sample farthest from this sample and the negative sample closest to this sample in the two original modality images, as shown in Equation (4):

[0128]

[0129] In the formula, A is the anchor sample, (A, p) is the positive sample, (A, n) is the negative sample, and ε is the threshold parameter.

[0130] Considering the extremely strong expressive ability of visible light images and the pre-training of the ResNet50 network parameters with visible light images, which will cause the model to have a strong visible light image tendency. In order to avoid the middle modality being affected by this factor. A center deviation loss function will be designed for the feature centers of the three modalities for constraint. The concept of modality deviation is as shown in Figure 7 shown.

[0131] In the figure, each dotted circle represents the same person in the same modality, and the dark pattern represents the feature center of this person in this modality for this batch. First, the feature centers of each category of pedestrians in the three modalities are calculated, and the intermediate feature is constrained to be at the center of the two original features. Further, the distance from the intermediate feature to the two original features is also constrained to be equal.

[0132] For each image with a different identity ID in each batch, the feature center in this category is calculated separately, and then the visible light feature F vi and the infrared feature F i of the feature centers are calculated, and the distance between them and the intermediate feature center is calculated. The formulas are shown in Formulas (5) and (6) as follows:

[0133]

[0134] In the formula, Batch i and Batch vi represent the visible light feature and the infrared feature in the batch respectively, K represents the number of images of each sample in each modality in this batch, P represents the number of sample categories in this batch, c final_i represents the center of the i-th intermediate feature, and c mid_i is the center of the visible light feature center and the infrared feature center of the i-th category.

[0135] Then the total network loss can be expressed as shown in Formula (7).

[0136] L all = L id + λ1L Tri_Tri_Loss + λ2L Center_Bias_Loss (7)

[0137] Among them, λ1 and λ2 are hyperparameters, which are set separately in the experiment.

[0138] The test process of the present invention is as follows:

[0139] Step S1: Input the query set and the gallery set, and enter Step S2;

[0140] Step S2: Use the model obtained in the training process to extract the features of all pedestrian images in the query set and the gallery set input in Step S1, and enter Step S3;

[0141] Step S3: Calculate the similarity between the query set features and the gallery set features, and enter Step S4;

[0142] Step S4: Obtain the matching results corresponding to each pedestrian image in the query set according to the similarity level, and proceed to Step S5;

[0143] Step S5: End.

[0144] In Step S1 of the test process, the query set represents the set of pedestrian images to be queried, while the gallery set represents the set of pedestrian images to be matched with the query set.

[0145] The similarity calculation method in Step S3 of the test process is dot product similarity.

[0146] Embodiment 1:

[0147] In this embodiment, Mini-batch Gradient Descent is adopted to update the model parameters of the present invention, that is, in each gradient descent, a small batch of samples is randomly selected for parameter update. The ResNet50 network is selected as the basic framework, and the parameters of ResNet50 are pre-trained on ImageNet. The stride of the last bottleneck block of the teacher network is set to 1. In terms of image preprocessing, three data augmentation methods are considered, including random cropping, horizontal flipping, and erasing.

[0148] In the experiment, the image size is uniformly set to a specific size. In addition, the momentum parameter is set to 0.9, and the initial learning rate is set to 0.01. In the first 10 epochs, a warm-up learning rate strategy is adopted, and then at the 20th and 50th epochs, the learning rates are adjusted to 0.01 and 0.001 respectively, and a total of 120 rounds of training are carried out. In the test stage, the pedestrian features processed by the batch normalization (BN) layer are used for retrieval. In this experiment, the deep learning framework pytorch1.8 is adopted, and an NVIDIA 3090 graphics card is used to accelerate the training process. The selection of these parameters and settings aims to optimize the performance and training efficiency of the model.

[0149] As Figure 1 shown, this embodiment will utilize the SYSU-MM01 dataset to complete the pedestrian re-identification task and test the performance of the model to evaluate the method in this paper. The comparison results are shown in Table 1.

[0150] Table 1 Performance comparison of the method of the present invention and other methods on the SYSU-MM01 dataset

[0151]

[0152]

[0153] In the experiment on the SYSU-MM01 dataset, among the comparison methods, the SFANet method only spectrally reduces the feature gap and transforms it into a data form with a slightly smaller modal gap. The experimental results show that the effect of this method is limited. In comparison, the method in this chapter has higher recognition accuracy. In the full search mode, Rank-1 and mAP increased by 2.99% and 5.29% respectively; in the indoor search mode, Rank-1 and mAP increased by 4.16% and 0.38% respectively. It can be seen that the method in this chapter uses a three-stream network structure to extract intermediate features, achieve a more prominent effect of modal feature pulling, optimize the feature measurement stage, and significantly improve the overall effect.

[0154] Compared with the PMWGCN method that uses noise filtering, it attempts to filter cross-modal differences as noise, but the infrared modality data itself has less information. The method in this chapter first extracts modality-specific features and retains as much valuable pedestrian feature information as possible in the visible light modality. Then, it passes through a feature extraction network that integrates the attention mechanism to extract closer modal features and alleviate cross-modal differences. Therefore, its effect is slightly lower than this method. In the full search mode, Rank-1 and mAP increased by 1.91% and 1.24% respectively; in the indoor search mode, Rank-1 and mAP increased by 3.12% and 5.04% respectively.

[0155] Embodiment 2:

[0156] In large shopping malls, multiple surveillance cameras are usually deployed to improve consumer experience, optimize mall management, and ensure public safety. However, the mall environment has the characteristics of large changes in lighting, such as strong light at the entrance, complex changes in indoor lighting, and low-light areas in some corners. Traditional single visible light cameras are difficult to meet the needs of all-weather pedestrian tracking. Therefore, infrared cameras and visible light cameras can be combined, and cross-modal pedestrian re-identification technology can be used to achieve accurate pedestrian matching and behavior analysis.

[0157] First, the system collects video streams from surveillance cameras in different areas of the mall and extracts pedestrian images through a pedestrian detection algorithm. These images are accompanied by camera numbers and timestamps for subsequent analysis.

[0158] These visible light images and infrared images are then fed into a cross-modal recognition model. The core task of the model is to recognize the image of the same customer under different lighting conditions and match its trajectory across the entire venue. For example, if a customer is captured by a visible light camera at the entrance when entering a mall, and then recorded by an infrared camera in a low-light area or parking lot, the model can accurately match these images to achieve cross-modal re-identification.

[0159] Next, the system analyzes the customer's movement trajectory based on the matching results and examines their shopping habits, such as their frequent areas, staying time, and frequently visited stores. This information can be used to optimize the store layout, adjust product placement strategies, and provide personalized recommendations. In addition, in terms of detecting abnormal behavior, this technology can also be used to identify individuals who loiter for a long time, stay abnormally, or may pose a theft risk, and the mall management can intervene in a timely manner.

[0160] Ultimately, the cross-modal pedestrian re-identification technology not only improves the intelligent management level of the mall but also enhances public safety, ensuring a more targeted and secure shopping experience for consumers.

[0161] The specific implementation solutions described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation solutions of the present invention and are not intended to limit the scope of the present invention. Any equivalent changes and modifications made by those skilled in the art without departing from the concept and principles of the present invention shall fall within the scope of protection of the present invention.

Claims

1. A cross-modal pedestrian re-identification method based on intermediate feature guidance, characterized in that It includes the following steps: Step 1: Preprocess P*K visible light images and infrared images, and proceed to Step 2; Step 2: Input the P*K preprocessed visible light images and P*K infrared images obtained in Step 1 into the feature extraction network, and proceed to Step 3; Step 3: Use the ResNet50 network and the feature mining module to generate informative image features based on the images input in Step 2, and proceed to Step 4; Step 4: Perform dot product fusion on the image features of the two branches to form intermediate features, and input the intermediate features together with the features of the original two branches into the feature fusion module to obtain new intermediate features, and proceed to Step 5; Step 5: Input the image features and intermediate features of the two branches obtained in Step 3 and Step 4 into the loss function, including the id loss, the three-modal triplet loss function, and the center deviation loss function, so as to train a network with good performance, and proceed to Step 6; Step 6: If the specified number of training rounds is reached, proceed to Step 7, otherwise continue the training and return to Step 1; Step 7: End.

2. The cross-modal pedestrian re-identification method based on intermediate feature guidance according to claim 1, wherein, The steps for extracting the preliminary image features in Step 2 are as follows: Step 2-1: Input the visible light image into the feature extraction network to extract its preliminary features; Step 2-2: Input the infrared image into the feature extraction network to extract its preliminary features; Step 2-3: Output the features.

3. A cross-modal pedestrian re-identification method based on intermediate feature guidance according to claim 2, characterized in that In Step 2, the first layer network of the Residual Network is used for feature extraction to fully perceive the shallow information of the image, and the feature extraction network parameters of the visible light image and the infrared image are not shared.

4. A cross-modal pedestrian re-identification method based on intermediate feature guidance according to claim 3, characterized in that The steps for extracting the image features in Step 3 are as follows: Step 3-1: Input the features extracted in Step 2 into the network composed of the convolutional layers of the last four layers of ResNet50 and the attention block of the feature mining module to obtain informative image features; Step 3-2: Perform MAX-pooling on the features in Step 3-1, and proceed to Step 3-3; Step 3-3: Output the pooled features of the image.

5. A cross-modal pedestrian re-identification method based on intermediate feature guidance according to claim 4, characterized in that In Step 3, in the feature mining module, a spatial attention module and a dimensional attention module are introduced. The spatial attention module focuses on the dependence between local features and spatial positions to enhance the perception ability of the target area; while the dimensional attention module focuses on channel information, adjusts the importance of different feature channels, enabling the network to dynamically allocate weights and improve the expression ability of key features; Input the input feature f0 into the two different attention modules respectively, use spatial attention and dimensional attention to obtain attention features with different information emphases, then perform dot product operations with the input feature, and perform activation and batch normalization; then perform dot product operations on the features of the two branches to obtain the feature f1 that fuses spatial and dimensional attention. For the obtained feature f1, global maximum pooling and global average pooling are performed respectively, and the features are concatenated to obtain a feature vector with a dimension of 2. In order to further enhance the discriminative ability of the model, matrix multiplication is performed with the features of the fusion space and dimensional attention, so that the output features can fully integrate the information from different attention mechanisms, and finally obtain the representation feature f2 with the same dimension as the original input feature but richer; Finally, during deep network training, the gradient vanishing and gradient exploding problems can lead to unstable model convergence. In order to avoid gradient descent, a residual connection mechanism is introduced in the final feature generation process, that is, f2 is added element by element to the original input feature.

6. The cross-modal pedestrian re-identification method based on intermediate feature guidance according to claim 5, characterized in that In step 4, the intermediate features are input into the feature fusion module together with the features of the original two branches to obtain new intermediate features, so as to achieve the purpose of deeply fusing the intermediate features with the features of the two modalities. The specific steps are as follows: Step 4-1: Multiply the visible light image feature F extracted in Step 3 vi by the transposed intermediate feature F mid to perform matrix multiplication, then after relu activation, multiply by the visible light image feature F again vi to perform matrix multiplication, and then perform bitwise addition with the intermediate feature F mid to obtain F fused1 ; Step 4-2: Input the infrared image feature F extracted in Step 3 in into the SelfAttention network to obtain the feature F in1 ; Multiply the feature F in1 by the transposed feature F fused1 through matrix multiplication, then after relu activation, multiply by the feature F in1 again through matrix multiplication, and then perform bitwise addition with the intermediate feature F fused1 to obtain the feature F fused2 ; Step 4-3: Input feature F fused2 into the SelfAttention network to obtain the output feature F fused ; Step 4-4: Perform BN normalization on the output feature F fused ​ 7. A cross-modal pedestrian re-identification method based on intermediate feature guidance according to claim 6, characterized in that In step 5, the id loss function is used to jointly constrain the model. The specific steps are as follows: Step 5-1: Pool the features extracted in step 3; Step 5-2: Perform BN normalization on the features extracted in step 5-1; Step 5-3: Pass the normalized features in step 5-2 through the fully connected layer to reduce the dimension of the information and obtain the dimensional information corresponding to the category, and then proceed to step 5-4; Step 5-4: Input the features into the id loss function; Step 5-5: Use the losses of 4-4 and 5-4 to calculate the gradient, back propagate, and update the model parameters; During training, P identities are selected from the training set, and K pedestrian images are randomly selected from each identity. There are 2*P*K pedestrian images in each Batch. The identity loss for the pedestrian re-identification task is shown in the following formula: where y i,k represents whether the identity of the i-th image is k, N represents the total number of pedestrian categories in the dataset, and p i,k represents the probability that the identity of the i-th image is k.

8. An intermediate feature-guided cross-modal pedestrian re-identification method according to claim 7, characterized in that The trimodal triplet loss function in step 5 is to use the image features of the intermediate modality as anchor points, and to bring the farthest images of the same identity closer and push the closest images of different identities farther away for the two original modal images; For the extracted trimodal data, it is necessary to make full use of the generated intermediate modal features to achieve a more efficient recognition of pedestrian identities. For this purpose, a trimodal triplet loss is proposed specifically. The common difficult triplet loss is a loss function widely used in pedestrian re-identification tasks. It brings the farthest positive pairs closer and pushes the closest negative pairs farther away. In simple terms, the image features of pedestrians of the same type taken by different cameras are brought closer, and the image features of different pedestrians are pushed farther away. The existing difficult cross-modal triplet loss, namely the difficult bimodal triplet loss; The same shape represents the same category, that is, pictures of different modes of the same pedestrian identity. Different shapes represent pictures of different categories, that is, pictures of different pedestrian identities. For each picture in each batch, the distance between it and the picture is calculated, and the positive sample farthest from the sample and the negative sample closest to the sample are selected to calculate the loss, as shown in formula (2): Where, A is the anchor sample, (A, p) is the sample with the same identity category as the anchor sample, and (A, n) is the sample with a different identity category from the anchor sample. is the threshold parameter, and is used to calculate the Euclidean distance, representing the feature f i and f i The Euclidean distance, and its formula calculation is as shown in Equation (3): The three-modal triplet loss function uses the image features of the middle modality as the anchor points, pulling the farthest images of the same identity closer for the two original modality images respectively, and pushing the nearest images of different identities farther away; For each intermediate feature sample in each batch, it will be used as an anchor sample, and the loss is calculated with the farthest positive sample and the nearest negative sample in the two original modality images to this sample, as shown in Equation (4): In the formula, A is the anchor sample, (A, p) is the positive sample, (A, n) is the negative sample, and ε is the threshold parameter.

9. A cross-modal pedestrian re-identification method based on intermediate feature guidance according to claim 8, characterized in that, Regarding the central deviation loss function in Step 5, considering the extremely strong expressive ability of visible light images and the fact that the ResNet50 network parameter values are pre-trained with visible light images, which will cause the model to have a strong visible light image tendency. To avoid the middle modality being affected by this factor, a central deviation loss function will be designed for the feature centers of the three modalities for constraint; All first find the feature centers for each category of pedestrians in the three modalities, and constrain the middle feature to be at the center of the two original features, and also constrain the distance from the middle feature to the two original features to be equal; For each different identity ID image in each batch, the feature center in this category is calculated separately, and then the visible light feature F is obtained. vi and the infrared feature F i of the feature center, and the distance between it and the intermediate feature center is calculated. The formulas are shown in Eqs. (5) and (6) as follows: where Batch i and Batch vi represent the visible light feature and the infrared feature in the batch respectively, K represents the number of pictures of each sample in each modality in the batch, P represents the number of sample categories in the batch, c final_i represents the center of the i-th intermediate feature, and c mid_i is the center of the visible light feature center and the infrared feature center representing the i-th category; Then the total network loss is expressed as shown in Equation (7): L all = L id + λ1L Tri_Tri_Loss + λ2L Center_Bias_Loss (7) Among them, λ1 and λ2 are hyperparameters, which are set separately in the experiment.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on intermediate modal parameter sharing and feature learning

    CN115731574A

  • Cross-modal pedestrian re-identification method, electronic equipment and medium

    CN116645690A

  • Cross-modal pedestrian re-identification method based on feature collaborative attention

    CN117727066A

  • Cross-modal pedestrian re-identification method based on cross-dimensional interactive attention

    CN117746457A

  • Cross-modal pedestrian re-identification method based on three-modal collaborative learning

    CN117975556A