A Cross-modal Person Re-identification Method Based on Convolutional Neural Network

By introducing multi-scale feature correspondence modules and joint loss functions in cross-modal pedestrian recognition, the problem of corresponding details between modes when pedestrian pose changes is solved, and a more efficient pedestrian recognition effect is achieved.

CN114627500BActive Publication Date: 2025-06-17ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210230686.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-10
Publication Date
2025-06-17
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

In cross-modal settings, pedestrian re-identification faces challenges, especially when pedestrian posture changes greatly, it is difficult for the prior art to effectively capture the corresponding details between modals.

Method used

The cross-modal pedestrian recognition method based on convolutional neural network is adopted. By introducing multi-scale feature correspondence modules, the features of infrared and sunlight mode images are obtained, the feature correspondence relationship is calculated, and the network is trained through joint loss functions to extract and reconstruct features to improve the recognition effect.

Benefits of technology

This method improves the pedestrian recognition effect and overcomes the problem of identifying corresponding details between modals. Especially when the pedestrian posture changes greatly, it can still maintain a good re-identification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627500B_ABST
    Figure CN114627500B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-modal pedestrian re-identification method based on a neural network, obtains a cross-modal training data set with identity annotations, each training sample in the training data set includes an infrared modal image and a daylight modal image corresponding to an identity, inputs the training sample into a network model built based on Resnet‑50, obtains multi-scale image features through a branch network, and calculates the feature correspondence between modalities thereon, fully mining the common features of modalities at different scales. A joint loss function is constructed to filter out identity-distinguishing features in the common features of modalities. The present invention combines global and local features as the representation of pedestrians, and achieves good results in the task of cross-modal pedestrian re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a cross-modal pedestrian re-identification method based on a convolutional neural network. Background Art

[0002] ReID is a basic problem in image retrieval, and its purpose is to match the target image in the query set to the image in the gallery set captured by different cameras. This is a challenge due to varying shooting perspectives, target morphologies, lighting, and backgrounds. Most existing methods currently focus on the target ReID problem captured by visible light cameras, that is, the single-modal ReID problem. However, in some scenes with insufficient lighting (such as at night, in a dimly lit indoor environment), we need to capture pedestrian images with an infrared camera. Therefore, in this cross-modal setting, the ReID problem becomes extremely challenging, which is essentially a cross-modal retrieval problem.

[0003] For cross-modal pedestrian re-identification, the mainstream technical solutions include feature learning methods that bridge the gap between RGB and IR images through feature alignment and methods that eliminate modal differences or feature disentanglement through generative adversarial networks. The mainstream algorithms for feature learning, such as the Two-stream series, directly learn features by adding some operations to the two-stream network. The algorithms have high accuracy and speed, but when the appearance of pedestrians changes greatly, the ability to capture details is not strong. The method of generative adversarial networks aims to directly generate images of another modality or disentangle modality-independent features using the network. However, due to the existence of a large number of modality-related features, the quality of image generation is not high, and it takes a huge amount of time. Summary of the Invention

[0004] The purpose of this application is to provide a cross-modal pedestrian re-identification method based on a convolutional neural network, which introduces a multi-scale feature correspondence module in the existing technical solution to overcome the problem of finding corresponding details between modalities when the pedestrian posture changes greatly.

[0005] To achieve the above purpose, the technical solution of this application is as follows:

[0006] A cross-modal pedestrian re-identification method based on a convolutional neural network, comprising:

[0007] Obtain a cross-modal training data set with identity annotations, and each training sample in the training data set includes an infrared modality image and a daylight modality image corresponding to an identity;

[0008] Input the training samples into the network model built based on Resnet-50. Denote the feature map output by the first residual block in the third residual layer of the Resnet-50 as F3. Feed the feature map F3 into three branches for separate processing to obtain the feature maps f g 、f l1 、f l2 、f l3 、f l4 、f l5 , including:

[0009] The first branch includes the remaining residual blocks in the third residual layer of Resnet-50 and the fourth residual layer, and extracts the global feature map f g ;

[0010] The second branch includes the remaining residual blocks in the third residual layer of Resnet-50 and the fourth residual layer, and obtains the local feature maps f l1 、f l2 ;

[0011] The third branch includes the remaining residual blocks in the third residual layer of Resnet-50 and the fourth residual layer, and obtains the local feature maps f l3 、f l4 、f l5 ;

[0012] Calculate the feature correspondence relationships between the infrared modality and daylight modality feature maps F3, f l1 、f l2 、f l3 、f l4 、f l5 ;

[0013] Perform feature reconstruction on the infrared modality and daylight modality feature maps F3, f l1 、f l2 、f l3 、f l4 、f l5 to obtain the reconstructed feature maps

[0014] Construct a joint loss function, and calculate the joint loss according to the infrared modality and daylight modality feature maps f g 、F3、f l1 、 f l2 、f l3 、f l4 、f l5 and the reconstructed feature maps and perform backpropagation to update the network parameters of the network model;

[0015] Use the trained network model to extract the features of the query image, compare them with the features of the images in the database, and identify the identity of the pedestrians in the query image.

[0016] Further, the fourth residual layer of the first branch has downsampling.

[0017] Further, calculate the feature correspondence between the infrared modality and daylight modality feature maps F3, f l1 、f l2 、f l3 、 f l4 、f l5 The calculation formula is as follows:

[0018] C(i, j) = f RGB (i) T ·f IR (j)

[0019] where, f RGB (i) and f IR (j) respectively represent the position feature vectors of the daylight modality feature map and the infrared modality feature map. i represents the position i of the daylight modality feature map, j represents the position j of the infrared modality feature map, and C(i, j) represents the position feature correspondence.

[0020] Further, perform feature reconstruction on the infrared modality and daylight modality feature maps F3, f l1 、f l2 、f l3 、 f l4 、f l5 to obtain the reconstructed feature map. The reconstruction formula is as follows:

[0021]

[0022]

[0023] M RGB (i) = |f RGB (i)|

[0024] M IR (j) = |f IR (j)|

[0025]

[0026]

[0027] where, f RGB (i) and f IR(j) represents the position feature vectors of the daylight modality feature map and the infrared modality feature map, i represents the i-th position of the daylight modality feature map, j represents the j-th position of the infrared modality feature map, and M RGB represents the response intensity at all positions on the daylight modality feature map, MIN represents taking the minimum value, MAX represents taking the maximum value, and M IR represents the response intensity at all positions on the infrared modality feature map, and M RGB (i) represents the response intensity at the i-th position on the daylight modality feature map, and M IR (j) represents the response intensity at the j-th position on the infrared modality feature map, represents the position feature vector at the i-th position of the reconstructed daylight modality feature map, represents the position feature vector at the j-th position of the reconstructed infrared modality feature map.

[0028] Furthermore, the formula of the joint loss function is as follows:

[0029]

[0030] Among them, represents the identity loss function, represents the triplet loss function, and the identity loss function and the triplet loss function calculate the loss for the global features respectively, and the global features are obtained by performing GeM pooling and fully connected dimensionality reduction operations on the global feature map f g ;

[0031] The represents the SmoothAP loss function, and the SmoothAP loss function calculates the loss for the local features and the local reconstruction features respectively. The local features are obtained by performing GeM pooling and fully connected dimensionality reduction operations on the local feature maps f l1 、f l2 、f l3 、f l4 、f l5 , and the local reconstruction features are obtained by performing GeM pooling and fully connected dimensionality reduction operations on the local reconstruction feature ;

[0032] The represents the dense triplet loss function, and the dense triplet loss function calculates the loss for the reconstruction feature map respectively.

[0033] A cross-modal pedestrian re-identification method based on a convolutional neural network proposed in this application. First, multi-scale feature extraction enables the network to focus on the detailed information of pedestrians and overcome the information loss caused by convolutional downsampling. Second, the feature correspondence operation can alleviate the modal differences and the feature misalignment problem caused by pedestrian pose changes. Finally, the proposed joint loss function imposes appropriate constraints on features at different levels, enabling the network to discover discriminative modality-shared features. The technical solution of this application improves the pedestrian recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a flowchart of the cross-modal pedestrian re-identification method based on a convolutional neural network in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further describes this application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0036] A cross-modal pedestrian re-identification method based on a convolutional neural network proposed in this application. Specifically, first, global and local feature maps are extracted, and then feature correspondences are calculated at the global and local levels respectively. Finally, a joint loss function is introduced to train features at different levels with different loss functions, guiding the network to retain features with identity information in the modality-shared features extracted.

[0037] In one embodiment, as Figure 1 shown, a cross-modal pedestrian re-identification method based on a convolutional neural network includes:

[0038] Step S1, obtain a cross-modal training data set with identity annotations, and each training sample in the training data set includes an infrared modality image and a daylight modality image corresponding to an identity.

[0039] To train a neural network, it is first necessary to obtain a training data set. In this embodiment, a training data set with identity annotations is read and randomly sampled and divided into batches according to the identities of the pedestrians in the images. For example, each batch contains 8 identities. Each training sample in this embodiment includes 4 daylight modality images (RGB images) and 4 infrared images (IR images) of an identity.

[0040] Step S2, input the training samples into a network model constructed based on Resnet-50, denote the feature map output by the first residual block in the third residual layer of the Resnet-50 as F3, and send the feature map F3 to 3 branches for processing respectively to obtain feature maps f g 、fl1 , f l2 , f l3 , f l4 , f l5 .

[0041] This step is used to obtain multi-scale feature maps. The backbone network uses a two-stream Resnet50. The ResNet50 model mainly consists of a shallow convolutional block layer0 and four residual convolutional layers layer1, layer2, layer3, and layer4. In layer0, the parameters of the network are specific to each modality, and all subsequent modules share the parameters.

[0042] The first residual blocks of layer1, layer2, and layer 3 extract the feature F3 as the backbone and extend three branches backward from it.

[0043] The first branch includes the remaining residual blocks of the third residual layer of Resnet-50 and the fourth residual layer, and extracts the global feature map f g ;

[0044] The second branch includes the remaining residual blocks of the third residual layer of Resnet-50 and the fourth residual layer, and obtains the local feature map f through vertical uniform slicing l1 , f l2 ;

[0045] The third branch includes the remaining residual blocks of the third residual layer of Resnet-50 and the fourth residual layer, and obtains the local feature map f through vertical uniform slicing l3 , f l4 , f l5 .

[0046] Specifically, the first branch is used to extract the global feature map, which consists of the last three blocks of layer3 and layer4 with downsampling. The network structures of the second and third branches are the same as the first one. The only difference is that layer4 without downsampling is used to retain details for the extraction of local feature maps. The output feature maps of the second and third branches are bisected and trisected vertically respectively to obtain the local feature maps f l1 , f l2 , f l3 , f l4 , f l5 . This application uses branches to extract features at different levels, which is conducive to discovering the corresponding relationships of features of different sizes.

[0047] It should be noted that the above operations are performed on the infrared modality images and daylight modality images respectively to obtain the global feature maps and local feature maps under different modalities.

[0048] To facilitate the calculation of the loss function in subsequent steps, for the global feature map and the local feature map, this application also performs GeM pooling operations (generalized-mean pooling) and dimensionality reduction operations respectively, converting the feature maps into feature vectors.

[0049] For the global feature map, after layer4, this application does not use the commonly used max pooling, but instead uses GeM pooling (generalized-mean pooling) to convert the output into a one-dimensional feature vector, and then uses a fully connected layer behind it to reduce the dimension to 256 for easy connection with local features. Finally, the same pooling and dimensionality reduction operations are performed on the local feature map to obtain the local feature vector.

[0050] Step S3: Calculate the feature correspondence between the infrared modality and the daylight modality feature maps F3, f l1 、f l2 、f l3 、 f l4 、f l5 respectively.

[0051] In this step, for the multi-scale feature maps F3, f l1 、f l2 、f l3 、f l4 、f l5 calculate the feature correspondence. Essentially, feature correspondence is also a problem of finding the common features of the target object in different images, which is exactly the main problem of cross-modal person re-identification.

[0052] The problem of the appearance change of pedestrians and the differences between modalities can be solved by establishing the feature correspondence between pedestrians in two modalities. In the training stage, by finding the feature correspondence between modalities, the network can learn to discover common features.

[0053] In this embodiment, the feature cosine similarity is used to represent the feature similarity. Let f IR ∈R c×h×w and f RGB ∈R c×h×w represent the feature maps of the IR and RGB images respectively. Each position feature vector is represented by f RGB / IR (i)∈R c . Calculate the feature correspondence C∈R hw×hw , and the formula is as follows:

[0054] C(i, j) = f RGB (i) T ·f IR (j)

[0055] where f RGB(i) and f IR (j) respectively represent the position feature vectors of the daylight modality feature map and the infrared modality feature map. i represents the position i of the daylight modality feature map, j represents the position j of the infrared modality feature map, and C(i, j) represents the corresponding relationship of position features. All C(i, j) together constitute the feature correspondence relationship C between modalities.

[0056] Use the above formula to calculate the feature relationships between the infrared modality and the daylight modality feature map F3, between the infrared modality and the daylight modality feature map f l1 between, between the infrared modality and the daylight modality feature map f l2 between, between the infrared modality and the daylight modality feature map f l3 between, between the infrared modality and the daylight modality feature map f l4 between, and between the infrared modality and the daylight modality feature map f l5 between.

[0057] According to the above formula, calculate the global feature similarity for a pair of cross-modal image features F3 of the same identity, and find the corresponding significant features. For f of the same identity l1 ~f l5 calculate the local feature similarity, and find the corresponding detailed features. For the multi-scale feature maps F3, f l1 , f l2 , f l3 , f l4 , f l5 calculate the corresponding relationships of the features respectively, so as to capture the modality-common features of different sizes.

[0058] Step S4: Perform feature reconstruction on the infrared modality and the daylight modality feature maps F3, f l1 , f l2 , f l3 , f l4 , f l5 to obtain the reconstructed feature maps

[0059] Referring to the way in the generative adversarial network to guide network learning according to the quality of the reconstructed map, this embodiment also reconstructs the feature map according to the feature correspondence relationship, and guides the network to find the feature correspondence according to the reconstruction quality. Due to the existence of modality-related information such as the background, directly restoring the features will definitely be affected. Therefore, a mask is used to filter out the modality-related information. Taking the RGB image as an example, assuming that in the re-identification task, the response of the useful modality-unrelated information is greater than that of the modality-related information. So use the modulus of each position feature vector f RGB (i) ∈ R c as the response intensity, and the formula is as follows:

[0060] M RGB (i) = |f RGB (i)|

[0061]

[0062] The above formula filters out modality-related information using a Mask, and the feature reconstruction formula for the RGB image feature map is:

[0063]

[0064] Similarly, the reconstruction formula for the infrared modality image can also be obtained:

[0065]

[0066] M IR (j) = |f IR (j)|

[0067]

[0068] Among them, f RGB (i) and f IR (j) respectively represent the position feature vectors of the daylight modality feature map and the infrared modality feature map. i represents the i position of the daylight modality feature map, j represents the j position of the infrared modality feature map, M RGB represents the response intensity at all positions on the daylight modality feature map, MIN represents taking the minimum value, MAX represents taking the maximum value, M IR represents the response intensity at all positions on the infrared modality feature map, M RGB (i) represents the response intensity at the i position on the daylight modality feature map, M IR (j) represents the response intensity at the j position on the infrared modality feature map, represents the position feature vector of the i position of the reconstructed daylight modality feature map, represents the position feature vector of the j position of the reconstructed infrared modality feature map.

[0069] It should be noted that in this embodiment, the above formula is used to perform feature reconstruction on the feature maps F3, f l1 , f2, f l3 , f l4 , f l5 to obtain the reconstructed feature map The multi-scale feature correspondence enables the network to focus on the modality-common detailed features, and still maintain a good re-identification effect when the pedestrian's pose changes.

[0070] Step S5, construct a joint loss function, according to the infrared modality and daylight modality feature maps fg , F3, f l1 , f l2 , f l3 , f l4 , f l5 and the reconstructed feature map Calculate the combined loss, perform backpropagation, and update the network parameters of the network model.

[0071] In this embodiment, a combined loss function is constructed, including improving the quality of the network reconstructed feature map and finding identity discriminative features in modality-independent features. It consists of an identity loss function (ID loss), a triplet loss function (Triplet loss), a SmoothAP loss function (SmoothAP loss), and a dense triplet loss function (Dense triplet loss) These four loss functions are combined. They are added together with different weights to obtain the final objective function, and the formula is as follows:[[]]

[0072]

[0073] Next, each loss will be described in detail. These loss functions are relatively mature technologies in the art. This application adopts these loss functions, and the specific calculations of how the loss functions are applied to this application will not be elaborated here.

[0074] represents the identity loss function,[[]] represents the triplet loss function, and the identity loss function and the triplet loss function calculate the loss for the global features respectively. The global features are obtained by performing GeM pooling and fully connected dimensionality reduction operations on the global feature map f g

[0075] In the pedestrian re-identification task, the identity loss function ID loss can learn discriminative features and at the same time reduce the intra-class distance. ID Loss in the multi-classification task is often considered for training. The ID loss formula is:[[]]

[0076]

[0077] The above identity loss function calculates for the global feature f i . The global feature fi is obtained by performing GeM pooling and fully connected dimensionality reduction operations on the global feature map f g , and its corresponding label is y i . The global feature fi is input into the classifier for classification and recognition. C is the number of pedestrian identities, that is, the total number of categories classified by the classifier. w k ​represents the weight of the kth class of the classifier, N is the batch size, Indicates the yth i The weight of the class, where T represents the transpose. ID loss can quickly make the features of the same class similar during training and complete a basic clustering task, but for modal differences, excessive pursuit of the feature's ability to represent the ID may lead the network to focus on information that is specific to the target but lacks modal universality, such as pedestrian clothing color, posture, etc. Therefore, this application does not adopt the setting of applying ID loss to both global and local variables commonly used in re-identification, but only applies ID loss to a global feature obtained by dimensionality reduction in the first branch. This can guide the network to perform a coarser ID clustering, but there is no need to over-pursue the detailed information that can represent the ID.

[0078] The triplet loss function uses a threshold to limit the relative distance between positive and negative samples to achieve the purpose of shortening the distance within a class and increasing the distance between classes. Its combination with ID loss has achieved good results in pedestrian ReID tasks. The triplet loss function formula is as follows:

[0079]

[0080] Specifically, an input triplet includes a pair of positive sample pairs and a pair of negative sample pairs. The three images are named anchor a, positive sample p, and negative sample n. Image a and image p are a pair of positive sample pairs, and image a and image n are a pair of negative sample pairs. Represent the features of anthor, positive and negative samples respectively. The triplet of difficult sample mining is to limit the relative distance between the farthest positive sample and the closest positive and negative samples. P represents the number of classes in the batch, and k represents the number of images of each class in the batch. The triple loss of difficult sample mining enhances the robustness of metric learning and further improves the performance. It should be noted that the features calculated in the formula are also obtained by filtering the global feature map f g It is obtained by performing GeM pooling and fully connected dimensionality reduction operations.

[0081] In this embodiment Represents the SmoothAP loss function, the SmoothAP loss function The loss is calculated for local features and local reconstruction features respectively. The local features are obtained by reconstructing the local feature map f l1 、f l2 、f l3 、f l4 、f l5It is obtained by performing GeM pooling and fully connected dimensionality reduction operations. The local reconstruction feature is obtained by performing GeM pooling and fully connected dimensionality reduction operations on the local reconstruction feature It is obtained by performing GeM pooling and fully connected dimensionality reduction operations.

[0082] mAP is a commonly used evaluation metric in the ReID task. However, due to the discrete sorting function involved in its calculation process, it cannot be used as an objective function to guide network learning. SmoothAP smooths the sorting process of query image search through the sigmoid function to approximate the calculation of AP. Specifically, the calculation formula of AP is as follows:

[0083]

[0084] S P represents the samples of the same class as instance i (positive class, images belonging to the same class as the query image). S Ω represents all samples. R(i, S P ) represents the ranking of instance i in S P . R(i, S Ω ) represents the ranking of instance i among all images. |S p | represents the number of positive class images. Expand the ranking function:

[0085]

[0086] I{·} represents the indicator function. D ij represents the difference in similarity between the query image and instances j and i respectively. The cosine distance is used to represent the similarity. If D ij > 0, it means that instance j is closer to the query image. Obviously, the numerator and denominator respectively represent the similarity rankings of instance i among positive classes and all images. Since the indicator function I{·} is not differentiable, the sigmoid function is used to approximate the indicator function. The formula is as follows:

[0087]

[0088] τ controls the accuracy of the sigmoid approximate indicator function. The lower τ is, the better the reduction degree. The approximate formula of AP is:

[0089]

[0090] To be consistent with other loss functions, 1 - AP is used as the final objective function:

[0091]

[0092] N is the batch size. Different from metric-based loss functions such as contrastive loss and triplet loss, SmoothAP can directly measure the quality of sorting. In this application, the SmoothAP function is used to train the local features obtained from the second and third branches and the local features obtained after cross-modal restoration, enabling the network to focus on the discriminative features shared between the two modalities.

[0093] ID Loss and Triplet Loss are used to reduce the intra-class distance and increase the inter-class distance in the early stage. SmoothAPLoss retains those discriminative detailed features by constraining the local feature selection.

[0094] As described in this embodiment represents the dense triplet loss function, and the dense triplet loss function calculates the loss for the reconstructed feature maps respectively.

[0095] To solve the problem of feature occlusion caused by the environment or pose, the dense triplet loss function is adopted. It first calculates the modality-shared mask to filter out the occluded features. Then, taking the L2 distance of the feature maps as the metric, the triplet loss function is calculated. This helps the network learn the shared features with discriminative ability. Taking IR-to-RGB as an example, the formula for the shared mask is:

[0096]

[0097] Let the feature map of the original image be the anchor, the RGB feature map restored from the infrared feature map of the same class be the positive, and the RGB feature map restored from the infrared feature map of different classes be the negative. The formula for the dense triplet loss function is:

[0098]

[0099] d + (i), d - (i) represent the L2 distances between the anchor and the positive and negative feature maps respectively, and α is the margin value.

[0100] In this embodiment, the network is trained with the joint loss function. The training samples are trained batch by batch. The joint loss is calculated for each batch, and backpropagation is performed to update the network parameters of the network model. The training samples are cycled 80 times to obtain the final network model.

[0101] Step S6: Use the trained network model to extract the features of the query image, compare them with the features of the images in the database, and identify the identity of the pedestrians in the query image.

[0102] The trained network model extracts features from each image in the query image and the images in the database. The multi-scale feature maps f g 、f l1 、f l2 、f l3 、f l4 、 f l5 After GeM pooling and dimensionality reduction, they are concatenated along the channel dimension as the final features of the pedestrians. The Euclidean distance between the features is used as the feature similarity metric to calculate the similarity between the features of the images in the query and the features of the images in the gallery, and the re-identification results are obtained by sorting according to the similarity.

[0103] The above embodiments only represent the implementation modes of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A cross-modal pedestrian re-identification method based on a convolutional neural network, characterized in that, The cross-modal person re-identification method based on a convolutional neural network includes: Obtain a cross-modal training data set with identity annotations, where each training sample in the training data set includes an infrared modality image and a daylight modality image corresponding to an identity; Input the training samples into the network model constructed based on Resnet-50. Denote the feature map output by the first residual block in the third residual layer of the Resnet-50 as F3. The feature map F3 is fed into three branches for processing respectively to obtain the feature maps f g , f1, f l2 , f l3 , f l4 , f l5 , including: The first branch includes the residual blocks remaining from the third residual layer of Resnet-50 and the fourth residual layer, and extracts the global feature map f g ; The second branch includes the remaining residual blocks of the third residual layer and the fourth residual layer of Resnet-50, and local feature maps f l1 , f l2 ; The third branch includes the residual blocks remaining from the third residual layer of Resnet-50 and the fourth residual layer, and local feature maps f are obtained by vertically and evenly slicing. l3 , f l4 , f l5 ; Calculate the feature correspondence between the infrared modality and daylight modality feature maps F3, f l1 , f l2 , f l3 , f l4 , f l5 ; Perform feature reconstruction on the infrared mode and daylight mode feature maps F3, f l1 , f l2 , f l3 , f l4 , f l5 to obtain the reconstructed feature map Construct a joint loss function, and calculate the joint loss based on the infrared modality and daylight modality feature maps f g 、F3、f l1 、f l2 、f l3 、f l4 、f l5 and the reconstructed feature map Calculate the joint loss, perform backpropagation, and update the network parameters of the network model; Use the trained network model to extract the features of the query image, compare them with the features of the images in the database, and identify the identity of the pedestrian in the query image; Among them, the infrared modality and the daylight modality feature maps F3, f l1 、f l2 、f l3 、f l4 、f l5 are subjected to feature reconstruction to obtain a reconstructed feature map, and the reconstruction formula is as follows: M RGB (i) = |f RGB (i)| M IR (j) = |f IR (j)| Among them, f RGB (i) and f IR (j) respectively represent the position feature vectors of the daylight modality feature map and the infrared modality feature map. i represents the position i of the daylight modality feature map, j represents the position j of the infrared modality feature map, M RGB represents the response intensity at all positions on the daylight modality feature map, MIN represents taking the minimum value, MAX represents taking the maximum value, M IR represents the response intensity at all positions on the infrared modality feature map, M RGB (i) represents the response intensity at position i on the daylight modality feature map, M IR (j) represents the response intensity at position j on the infrared modality feature map, represents the position feature vector of the reconstructed daylight modality feature map at position i, represents the position feature vector of the reconstructed infrared modality feature map at position j.

2. The cross-modal pedestrian re-identification method based on a convolutional neural network according to claim 1, characterized in that, The fourth residual layer of the first branch has downsampling.

3. The cross-modal pedestrian re-identification method based on a convolutional neural network according to claim 1, characterized in that, The feature correspondence relationships among the calculated infrared mode and sunlight mode feature maps F3, f l1 , f l2 , f l3 , f l4 , f l5 are as follows. The calculation formula is as follows: C(i,j) = f RGB (i) T ·f IR (j) Among them, f RGB (i) and f IR (j) respectively represent the position feature vectors of the daylight modality feature map and the infrared modality feature map. i represents the position i of the daylight modality feature map, j represents the position j of the infrared modality feature map, and C(i, j) represents the corresponding relationship of the position features.

4. The cross-modal pedestrian re-identification method based on a convolutional neural network as claimed in claim 1, characterized in that, The formula of the joint loss function is as follows: Among them, represents the identity loss function, represents the triplet loss function, and the identity loss function and the triplet loss function calculate losses for the global features respectively, and the global features are obtained by performing GeM pooling and fully connected dimensionality reduction operations on the global feature map f g ; The represents the SmoothAP loss function, and the SmoothAP loss function calculates the loss for the local features and the local reconstruction features respectively. The local features are obtained by performing GeM pooling and fully connected dimensionality reduction operations on the local feature maps f l1 、f l2 、f l3 、f l4 、f l5 . The local reconstruction features are obtained by performing GeM pooling and fully connected dimensionality reduction operations on the local reconstruction feature ; The represents a dense triplet loss function, and the dense triplet loss function calculates losses for the reconstructed feature maps respectively.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on difficult quintuple

    CN111597876A

  • Cross-modal pedestrian re-identification method and system based on double-flow convolutional neural network

    CN111931637A