A Cross-modal Person Re-identification Method Based on Multi-feature Joint Progressive Learning
Through the multi-feature combined with a gradual learning network, the problem of mode differences in cross-modal pedestrian re-identification is solved, and the recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202311430114.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-10-31
AI Technical Summary
In the cross-modal pedestrian re-identification task, the measurement function has data bias, and the measurement and representation learning are not suitable, making it difficult to alleviate the modal differences.
Using a method based on multi-feature joint progressive learning, a multi-feature joint progressive learning network is constructed, including a dual-stream Resnet-50 network, a contour feature enhancement module and a batch adaptive module, gradually learning features from coarse-grained to fine-grained to reduce modal differences.
It effectively alleviates modal differences, improves the accuracy and robustness of cross-modal pedestrian re-identification, and ensures that the model's feature learning is more consistent and effective under different modes.
Smart Images

Figure CN117315720B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a cross-modal pedestrian re-identification method based on multi-feature joint progressive learning. Background Art
[0002] Given a query image captured by a visible light or infrared camera, the purpose of cross-modal pedestrian re-identification (visible-infrared person re-identification, VI-ReID) is to retrieve pedestrians with the same identity in an infrared or visible light library. With the popularization of surveillance cameras and the introduction of relevant datasets, the VI-ReID task has received a great deal of attention from researchers.
[0003] As an image retrieval task, VI-ReID faces two challenges: data bias existing in the metric function and the mismatch problem between metric and representation learning.
[0004] To narrow the intra-class distance and modal differences, current methods can be divided into two categories: pixel-based and feature-based. Pixel-based [AlignGan, JSIA] methods aim to generate fake images with the coloring style of the corresponding modality through a generative model, and the fake images are used to assist image retrieval. However, due to the large modal differences, style transfer requires a complex network structure, and this approach has not achieved good performance.
[0005] The current mainstream approach is to design a metric function and train the model to learn co-modal features by restricting the feature distances between modalities. MMN, AGW, and FBP-AL are based on the classic triplet strategy, restricting the relative distances between the same-class modalities and between classes to alleviate modal differences. However, these methods only optimize a pair of samples each time, lack robustness, and are prone to falling into sub-optimal results. HCT and MPNAet propose to constrain the distance from sample points to the modal center, which alleviates the above problems to a certain extent. However, the modal center is based on the feature mean of the current batch, and it is still impossible to avoid the data bias brought by small batches. In addition, the metric function evaluates features from the perspective of feature distance, while representation learning evaluates features from the perspective of identity expression, and this difference will affect model training. Summary of the Invention
[0006] The purpose of the present invention is to provide a cross-modal pedestrian re-identification method based on multi-feature joint progressive learning, which helps the multi-feature joint progressive learning network to learn from coarse-grained to fine-grained from the perspective of diverse features, effectively alleviating modal differences.
[0007] To achieve the above purpose, the technical solution provided by the present invention is as follows: The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning includes:
[0008] Given an image set D = {Xvis , X ir} Divide into a visible light image set X according to the modality vis and an infrared image set X ir ;
[0009] Construct a multi - feature joint progressive learning network. The multi - feature joint progressive learning network includes a two - stream Resnet - 50 network, a contour feature enhancement module, and a batch adaptive module. The two - stream Resnet - 50 network contains two branches and a softmax function. The two branches are respectively used to extract the features of visible light images and infrared images;
[0010] Input the image set D into the contour feature enhancement module and the two - stream Resnet - 50 network respectively to obtain the original image feature map and the contour feature map. Obtain the perceptual contour loss according to the original image feature map and the contour feature map, and calculate the overall loss in the first stage;
[0011] Grayscale the visible light images in the image set D to obtain the corresponding grayscale images, and use the grayscale image set X gray and the infrared image set X ir as the input of the two - stream Resnet - 50 network to calculate the overall loss in the second stage;
[0012] Input the image set D into the two - stream Resnet - 50 network, and input the obtained feature sequence into the batch adaptive module to calculate the overall loss in the third stage;
[0013] Jointly optimize the overall losses of the three stages, and optimize the parameters of the multi - feature joint progressive learning network through backpropagation;
[0014] Given a pedestrian query image and an image gallery as the input of the trained multi - feature joint progressive learning network, calculate the similarity between the query image and all images in the image gallery using the output features to complete the re - identification process.
[0015] Furthermore, the two branches of the two - stream Resnet - 50 network both include a sub - network with independent parameters and a sub - network with shared parameters. The sub - networks with independent parameters of the two branches are respectively denoted as E V , E IR , which are composed of the first layer of each branch of the two - stream Resnet - 50 network. The sub - network with shared parameters is denoted as E S , which is composed of the remaining layers of each branch of the two - stream Resnet - 50 network.
[0016] Further, inputting the image set D into the contour feature enhancement module and the dual-stream ResNet-50 network respectively to obtain the original image feature map and the contour feature map, and obtaining the perceptual contour loss according to the original image feature map and the contour feature map, and calculating the overall loss in the first stage, including:
[0017] Input the image set D into the sub-network E with independent parameters of the dual-stream ResNet-50 network V 、E IR , and obtain the feature map C1;
[0018] Input the image set D into the Sobel convolutional layer of the contour feature enhancement module to obtain the feature map C2;
[0019] Perform element-wise addition on C1 and C2, and the element-wise added feature map is trained through the first to fourth blocks of the VGG-16 network in the contour feature enhancement module. Let ψ = {ψ1, ψ2, ψ3, ψ4} represent the four blocks of VGG-16;
[0020] For the finally obtained total perceptual contour loss, it can be expressed by the formula:
[0021]
[0022] Wherein, Represents the input The original image feature map obtained after passing through the i-th block of the VGG-16 network, i ∈ [1, 4], Represents the original image feature map extracted from the visible light modality feature map after passing through the (i - 1)-th block of the VGG-16 network, Represents the input The contour feature map obtained after passing through the i-th block of the VGG-16 network, Represents the contour feature map extracted from the visible light modality feature map after passing through the (i - 1)-th block of the VGG-16 network, Represents the input The original image feature map obtained after passing through the i-th block of the VGG-16 network, Represents the original image feature map extracted from the infrared modality feature map after passing through the (i - 1)-th block of the VGG-16 network, Represents the input The contour feature map obtained after passing through the i-th block of the VGG-16 network, Represents the contour feature map extracted from the infrared modality feature map after passing through the (i - 1)-th block of the VGG-16 network;
[0023] According to the original image feature map And the contour feature map Obtain the perceptual contour loss, which can be calculated and expressed by the following formula:
[0024]
[0025] Among them, represents the perceptual contour loss, C represents the loss calculation for the current channel, and the original image feature map and the contour feature map have the shape of C i ×H i ×W i , C i represents the total number of channels, H i represents the height, and W i represents the width. m ∈ {vis, ir} represents two modalities of visible light and infrared rays. represents the calculation and summation of the loss for each channel in the features of each modality. L δ represents the Huber loss, which can be expressed by the formula:
[0026]
[0027] Among them, δ represents the hyperparameter;
[0028] The overall loss in the first stage can be expressed by the following formula:
[0029]
[0030] Among them, represents the id loss, which calculates the loss for the features obtained by processing the input images {X vis , X ir} through the multi-feature joint progressive learning network. The formula is as follows:
[0031]
[0032] Among them, P is the number of categories in the image set, y represents the category label of the pedestrian, P a is the predicted probability of category a, ε is a small constant, q a represents the actual label distribution probability, and p a is obtained by processing the feature vector after processing the original image feature map through the softmax function;
[0033] The triplet loss between modalities calculates the loss for the features obtained by processing the input images {X vis , X ir} through the multi-feature joint progressive learning network, and can be expressed by the formula:
[0034]
[0035] Among them, represents the triplet loss calculated from positive sample pairs from the VIS and IR modalities and negative sample pairs from the VIS modality, represents the triplet loss calculated from positive sample pairs from the IR and VIS modalities and negative sample pairs from the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality. and can be expressed by the formula:
[0036]
[0037]
[0038] Among them, M represents randomly selecting M pedestrians each iteration, D(V b , i c ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the c-th pedestrian in the IR image, D(V b , I k ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the k-th pedestrian in the IR image, D(I b , V c ) represents the Euclidean distance between the b-th pedestrian in the IR image and the c-th pedestrian in the VIS image, D(I b , V k ) represents the Euclidean distance between the b-th pedestrian in the IR image and the k-th pedestrian in the VIS image. b is the same as c, representing the same pedestrian; k is different from b and c, representing different pedestrians, y b represents the class label of the b-th pedestrian in the current modality, y c represents the class label of the same pedestrian in a different modality from y b , y k represents the class label of a pedestrian different from y b and y c . ξ is the margin parameter, [z] + = max(z, 0).
[0039] Furthermore, inputting the image set D into the Sobel convolutional layer of the contour feature enhancement module to obtain the feature map C2 includes:
[0040] By inputting the image set D into the Sobel convolutional layer of the contour feature enhancement module, the horizontal and vertical image grayscale values are obtained. The formula is as follows:
[0041]
[0042] Among them, A represents the original image in the image set D, and G x represents the horizontal image grayscale value, and G y represents the vertical image grayscale value;
[0043] At each point of the image, by combining the horizontal and vertical image grayscale values, the approximate gradient of each point is obtained through the following formula:
[0044]
[0045] The feature map C2 is obtained according to the approximate gradient.
[0046] Furthermore, the visible light images in the image set D are grayscaled to obtain the corresponding grayscale images, and the grayscale image set X gray and the infrared image set X ir are used as the input of the two-stream Resnet-50 network to calculate the overall loss in the second stage, including:
[0047] The visible light images in the image set D are grayscaled to obtain grayscale images;
[0048] The grayscale image set x gray and the infrared image set X ir are input into the two-stream Resnet-50 network, and the human grayscale features of the grayscale images are output and the overall loss in the second stage is calculated It can be expressed by the following formula:
[0049]
[0050] Among them, represents the heterocentric cluster loss, represents the id loss, represents the triplet loss;
[0051] The heterocentric cluster loss calculates the loss of the features obtained by the multi-feature joint progressive learning network for the input images {X gray , X ir}, and the formula is as follows:
[0052]
[0053] Among them, N is the total number of and in the training stage batch, f() is the output before the fully connected layer of the two-stream Resnet-50 network, mean[A - B] is the mean operation of the difference between A and B, represents the grayscale image generated by the visible light modality, represents the input image of the infrared modality;
[0054] ID loss Calculate the loss for the features obtained from the input images {X gray , X ir} through the multi - feature joint progressive learning network. The formula is as follows:
[0055]
[0056] Among them, P is the number of categories in the image set, y represents the category label of the pedestrian, p a is the predicted probability of category a, ε is a small constant, q a represents the actual label distribution probability, p a is obtained through the softmax function based on the human grayscale feature;
[0057] Triplet loss between modalities Calculate the loss for the features obtained from the input images {X gray , X ir} processed by the multi - feature joint progressive learning network. It can be expressed by the formula:
[0058]
[0059] Among them, represents the triplet loss calculated from the positive sample pairs of the input from the VIS and IR modalities and the negative sample pairs from the VIS modality, represents the triplet loss calculated from the positive sample pairs of the input from the IR and VIS modalities and the negative sample pairs from the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality, and can be expressed by the formula:
[0060]
[0061]
[0062] Among them, M represents randomly selecting M pedestrians in each iteration. D(V b , I c ) represents the Euclidean distance between the b - th pedestrian in the VIS image and the c - th pedestrian in the IR image. D(V b , I k ) represents the Euclidean distance between the b - th pedestrian in the VIS image and the k - th pedestrian in the IR image. D(I b , V c ) represents the Euclidean distance between the b - th pedestrian in the IR image and the c - th pedestrian in the VIS image. D(I b , V k) represents the Euclidean distance between the b-th pedestrian in the IR image and the k-th pedestrian in the VIS image. b is the same as c, representing the same pedestrian; k is different from b and c, representing different pedestrians, y b represents the class label of the b-th pedestrian in the current modality, y c represents the same as y b the class label of the same pedestrian in different modalities, y k represents the same as y b and y c the class labels of different pedestrians, ξ is the interval parameter, [z] + = max(z, 0).
[0063] Furthermore, taking the image set D as the input of the two-stream Resnet-50 network, and inputting the obtained feature map into the batch adaptive module to calculate the overall loss in the third stage, including:
[0064] Inputting the image set D into the two-stream Resnet-50 network to obtain the feature sequence Y, where Y ∈ R E×F , E and F respectively represent the length and dimension of the feature sequence, and replace the dimension of the feature sequence Y:
[0065]
[0066] Taking input into the batch adaptive module to obtain The batch adaptive module can be expressed as:
[0067]
[0068]
[0069] where, MHSA represents multi-head attention, FF represents a fully connected layer, LN1, LN2 represent linear layers;
[0070] Restoring the dimension of the feature sequence :
[0071]
[0072] Calculating the loss of the feature sequence after passing through the batch adaptive module to obtain the overall loss in the third stage It can be expressed by the following formula:
[0073]
[0074] where, represents the heterocentric cluster loss, represents the id loss, represents the triplet loss;
[0075] Eccentric cluster loss For the features obtained from the input images {X vis , X ir} through the multi-feature joint progressive learning network, the loss is calculated as follows:
[0076]
[0077] where N is the total number of and in the training stage batch, f() is the output before the fully connected layer of the multi-feature joint progressive learning network, mean[A - B] is the mean operation of the difference between A and B, represents the feature sequence Y' generated by the visible light modality, represents the feature sequence Y' generated by the infrared modality;
[0078] id loss For the features obtained from the input images {X vis , X ir} through the multi-feature joint progressive learning network, the loss is calculated as follows:
[0079]
[0080] where P is the number of categories in the image set, y represents the category label of the pedestrian, p a is the predicted probability of category a, ε is a small constant, q a represents the actual label distribution probability, p a is obtained by processing the feature sequence Y' through the softmax function;
[0081] Inter-modal triplet loss For the features obtained from the input images {X vis , X ir} after being processed by the network, the loss can be calculated using the formula:
[0082]
[0083] where, represents the triplet loss calculated from the positive sample pairs of the input from the VIS and IR modalities and the negative sample pairs from the VIS modality, represents the triplet loss calculated from the positive sample pairs of the input from the IR and VIS modalities and the negative sample pairs from the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality, and can be expressed by the formula:
[0084]
[0085]
[0086] Among them, M represents the random selection of M pedestrians in each iteration, D(V b ,I c ) represents the Euclidean distance between the bth pedestrian in the VIS image and the cth pedestrian in the IR image, D(V b ,I k ) represents the Euclidean distance between the bth pedestrian in the VIS image and the kth pedestrian in the IR image, D(I b ,V c ) represents the Euclidean distance between the bth pedestrian in the IR image and the cth pedestrian in the VIS image, D(I b ,V k ) represents the Euclidean distance between the bth pedestrian in the IR image and the kth pedestrian in the VIS image. b and c are the same, indicating the same pedestrian; k is different from b and c, indicating different pedestrians. b Indicates the category label of the bth pedestrian in the current mode, y c Indicates that b Category labels of the same person in different modalities, y k Indicates that b and c The category labels of different pedestrians, ξ is the interval parameter, [z] + =max(z,0).
[0087] Furthermore, the overall loss of the three stages of the joint optimization includes:
[0088] The overall loss of the three stages is obtained by adding the overall loss of the first stage, the overall loss of the second stage and the overall loss of the third stage.
[0089] Furthermore, the re-identification process includes:
[0090] Given a query image in visible light modality and an infrared image library, or given a query image in infrared modality and a visible light image library;
[0091] The given query image and image library are input into the multi-feature joint progressive learning network, and the two-stream Resnet-50 network outputs the character features of the query image and the images in the image library;
[0092] The feature similarity between the query image and all images in the image library is calculated, and the e images in the image library with the highest feature similarity to the query image are taken as candidate images, which are recorded as the re-identification results obtained by the multi-feature joint progressive learning network.
[0093] Compared with the prior art, the significant advantages of the present invention are as follows: more diverse features are learned, such as grayscale features, contour features of people, and features of the original image. Compared with the existing methods that only focus on modality-shared features, while better focusing on modality-shared features, the present invention will not ignore modality-specific features. This can help the multi-feature joint progressive learning network to progressively learn features from coarse-grained to fine-grained, thereby effectively alleviating modality differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 is a flowchart of the method of the present invention;
[0095] Figure 2 is a schematic structural diagram of the multi-feature joint progressive learning network of the present invention;
[0096] Figure 3 is a schematic diagram of the contour feature enhancement module in the multi-feature joint progressive learning network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0097] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0098] In one embodiment, as Figure 1 shown, a cross-modal pedestrian re-identification method based on multi-feature joint progressive learning is proposed, including:
[0099] Step S1: Divide the data set D = {X vis , X ir} into a visible light set X vis and an infrared set X ir .
[0100] In this example, the input images of the model are divided into a visible light set X vis and an infrared set X ir to facilitate the subsequent processing of the two-stream network.
[0101] Step S2: Construct a multi-feature joint progressive learning network.
[0102] The multi-feature joint progressive learning network includes a two-stream Resnet-50 network, a contour feature enhancement module, and a batch adaptive module. The two-stream Resnet-50 network serves as the backbone network of the multi-feature joint progressive learning network.
[0103] The two-stream Resnet-50 network applicable in this example includes two branches and a softmax function. The two branches respectively process the visible light set X vis and the infrared set Xir Each branch contains a sub-network with independent parameters, denoted as E V or E IR , which is implemented by the first layer of each branch of the two-stream Resnet-50 network and a sub-network E with shared parameters S , which is implemented by the remaining layers of each branch of the two-stream Resnet-50 network. E V and E IR with independent parameters is to adapt to modal differences. E S The shared parameters are designed to extract modality-independent features of pedestrians.
[0104] Step S3: Extract the human contour features and calculate the loss.
[0105] As Figure 2 shown, in the first training stage of this example, the image set is used as the input of the designed contour feature enhancement module and the two-stream Resnet-50 network to obtain the original image feature map and the contour feature map The contour features belong to shallow features, contain more detailed information, and are not affected by modal differences. To improve the expression ability of the network, this example progressively learns diverse feature representations. The specific steps are as follows:
[0106] As Figure 3 shown, in this example, {X vis , X ir} is passed through E V and E IR to obtain the feature map C1. Then, {X vis , X ir} is passed through the Sobel convolutional layer in the contour feature enhancement module to obtain the feature map C2. Among them, the Sobel module contains two 3×3 matrices, representing the horizontal and vertical directions respectively. The feature map C2 is convolved with the original images {X vis , X ir} in the plane, and the horizontal and vertical image gray values can be obtained respectively. Specifically as follows:
[0107] Let A represent the original image in the image set D, G x and G y represent the horizontal and vertical image gray values respectively. The formula is as follows:
[0108]
[0109] Then, at each point in the image, by combining the horizontal and vertical image gray values calculated above, the approximate gradient of each point can be obtained:
[0110]
[0111] The feature map C2 is obtained according to the approximate gradient.
[0112] C1 and C2 are subjected to element-wise addition. The feature map after element-wise addition is trained through the first to the fourth blocks of the VGG-16 network. Let ψ = {ψ1, ψ2, ψ3, ψ4} represent the four blocks of the VGG-16 network. Denote the input as the feature map obtained after passing through the i-th block of the VGG-16 network. Denote the feature map extracted after passing through the (i - 1)-th block of the VGG-16 network, where i ∈ [1, 4]. When i = 1 denotes the feature map initially input into the VGG-16 network, and it does not mean that the VGG-16 network has a 0-th block. The shape of x is C i ×H i ×W i , C i represents the total number of channels, H i represents the height, and W i represents the width. The feature map can be divided into and which respectively represent the original image feature map and the contour feature map obtained after passing through the i-th layer of the VGG-16 network, where m ∈ {vis, ir} represents the two modalities of visible light and infrared light.
[0113] Due to the existence of two modalities, the images of the two modalities are both processed to obtain the corresponding C1 and C2 after processing. Therefore, the feature maps after element-wise addition will be obtained respectively under the two modalities.
[0114] is the total perceptual contour loss finally obtained, which can be expressed by the formula:
[0115]
[0116] where, denotes the original image feature map obtained after passing through the i-th block of the VGG-16 network for the input , denotes the original image feature map extracted after passing through the (i - 1)-th block of the VGG-16 network for the feature map of the visible light modality, denotes the input as the contour feature map obtained after passing through the i-th block of the VGG-16 network, denotes the contour feature map extracted after passing through the (i - 1)-th block of the VGG-16 network for the feature map of the visible light modality, denotes the input as the original image feature map obtained after passing through the i-th block of the VGG-16 network, Denote the original image feature map obtained by extracting the feature map of the infrared modality through the (i - 1)-th block of the VGG-16 network. Denote the input The contour feature map obtained by passing through the i-th block of the VGG-16 network. Denote the contour feature map obtained by extracting the feature map of the infrared modality through the (i - 1)-th block of the VGG-16 network.
[0117] According to the original image feature map And the contour feature map Obtain the perceptual contour loss, which can be calculated and expressed by the following formula:
[0118]
[0119] Among them, Denote the perceptual contour loss, C i Denote the total number of channels, C represents calculating the loss for the current channel, m ∈ {vis, ir} represents the two modalities of visible light and infrared, Denote calculating and summing the losses for each channel in the features of each modality, L δ Denote the Huber loss, which can be expressed by the formula:
[0120]
[0121] Among them, δ represents the hyperparameter, and the loss will determine which loss to specifically adopt according to the value of δ.
[0122] The overall loss in the first stage Can be expressed by the following formula:
[0123]
[0124] Among them, Denote the id loss, calculate the loss for the features obtained by processing the input images {X vis 、X ir} through the multi-feature joint progressive learning network. The formula is as follows:
[0125]
[0126] Among them, P is the number of categories in the image set, y represents the category of the pedestrian, p a Is the predicted probability of category a, ε is a small constant, q a Represents the actual label distribution probability, p a Based on the original image feature map The processed feature vector is obtained through the softmax function. To correspond to p a Correspondingly, q aIt also needs to be further processed by the softmax function.
[0127] For the original image feature map The processed feature vectors are subjected to average pooling and batch normalization (BN) to obtain the final feature vectors. Finally, the final feature vectors are further processed by a softmax function to convert the original input into a class probability distribution.
[0128] Triplet loss between modalities For the input images {X vis , X ir}, the loss is calculated for the features obtained by processing through the multi-feature joint progressive learning network, which can be expressed by the formula:
[0129]
[0130] Among them, represents the triplet loss calculated for the positive sample pairs from the VIS and IR modalities and the negative sample pairs from the VIS modality, represents the triplet loss calculated for the positive sample pairs from the IR and VIS modalities and the negative sample pairs from the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality. and can be expressed by the formula:
[0131]
[0132]
[0133] Among them, M represents randomly selecting M pedestrians each iteration. D(V b , I c ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the c-th pedestrian in the IR image. D(V b , I k ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the k-th pedestrian in the IR image. D(I b , V c ) represents the Euclidean distance between the b-th pedestrian in the IR image and the c-th pedestrian in the VIS image. D(I b , V k ) represents the Euclidean distance between the b-th pedestrian in the IR image and the k-th pedestrian in the VIS image. b and c are the same, representing the same pedestrian; k is different from b and c, representing different pedestrians. y b represents the class label of the b-th pedestrian in the current modality, and y c represents the class label of the same pedestrian in a different modality from y b and yk denotes the class labels of different pedestrians, ξ is the interval parameter, and [z] b and y c = max(z, 0). (V + = max(z, 0). (V b , I c ), (I b , V c ) are positive sample pairs, and (V b , I k ), (I b , V k ) are negative sample pairs.
[0134] Step S4: Extract the grayscale features of the person and calculate the loss.
[0135] In this example, in the second training stage, X vis is processed through grayscale conversion to obtain the grayscale image set x gray . The grayscale image set x gray and the infrared image set X ir are input into the dual-stream Resnet-50 network to output the grayscale features of the person in the grayscale image and calculate the id loss hetero-center cluster loss and triplet loss
[0136] The overall loss in the second stage can be expressed by the following formula:
[0137]
[0138] For the features obtained by the input images {X gray , X ir} through the multi-feature joint progressive learning network, the loss is calculated as follows:
[0139]
[0140] where N is the total number of and in the training stage batch, f() is the output before the fully connected layer of the multi-feature joint progressive learning network, mean[A - B] is the mean operation of the difference between A and B, denotes the grayscale image generated by the visible light modality, denotes the input image of the infrared modality.
[0141] id loss For the input images {X gray , X irCalculate the loss for the features obtained through the multi-feature joint progressive learning network, and the formula is as follows:
[0142]
[0143] Among them, P is the number of categories in the image set, y represents the category label of the pedestrian, and p a is the predicted probability of category a, ε is a small constant, and q a represents the actual label distribution probability, and p a is obtained by processing the grayscale features of the person through the softmax function. To correspond to p a , q a also needs to be further processed through the softmax function.
[0144] After being processed by the two-stream Resnet-50 network, the feature vectors of the grayscale features of the person will be obtained. Average pooling and batch normalization (BN) are performed on these feature vectors to obtain the final feature vectors. Finally, the final feature vectors are processed through a softmax function to convert the original input into a category probability distribution.
[0145] Triplet loss between modalities Calculate the loss for the features obtained by processing the input images {X gray , X ir} through the multi-feature joint progressive learning network, which can be expressed by the formula:
[0146]
[0147] Among them, represents the triplet loss calculated for the positive sample pairs from the VIS and IR modalities and the negative sample pairs from the VIS modality, represents the triplet loss calculated for the positive sample pairs from the IR and VIS modalities and the negative sample pairs from the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality, and can be expressed by the formula:
[0148]
[0149]
[0150] Among them, M represents randomly selecting M pedestrians in each iteration, and D(V b , I c ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the c-th pedestrian in the IR image, and D(V b , I k) represents the Euclidean distance between the bth pedestrian in the VIS image and the kth pedestrian in the IR image, D(I b ,V c ) represents the Euclidean distance between the bth pedestrian in the IR image and the cth pedestrian in the VIS image, D(I b ,V k ) represents the Euclidean distance between the bth pedestrian in the IR image and the kth pedestrian in the VIS image. b and c are the same, indicating the same pedestrian; k is different from b and c, indicating different pedestrians. b Indicates the category label of the bth pedestrian in the current mode, y c Indicates that b Category labels of the same person in different modalities, y k Indicates that b and c The category labels of different pedestrians, ξ is the interval parameter, [z] + =max(z,0). (V b ,I c )、(I b ,V c ) is a positive sample pair, (V b ,I k )、(I b ,V k ) is a negative sample pair.
[0151] Step S5: Design a batch adaptation module to establish connections in the batch dimension.
[0152] In this example, in the third training stage, the original images of the image set D {X vis ,X ir} as the input of the two-stream Resnet-50 network, and obtain the feature sequence Y, where Y∈R E×F , E and F represent the length and dimension of the feature sequence respectively, and replace the dimension of the feature sequence Y:
[0153]
[0154] Will Input batch adaptive module to get The batch adaptation module can be expressed as:
[0155]
[0156]
[0157] Among them, MHSA represents multi-head attention, FF represents fully connected layer, LN1 and LN2 represent linear layers;
[0158] The feature sequence Restore the dimension:
[0159]
[0160] Calculate the loss for the feature sequence after the designed batch adaptive module to obtain the overall loss in the third stage It can be expressed by the following formula:
[0161]
[0162] Where, Represents the heterocentric cluster loss, Represents the id loss, Represents the triplet loss;
[0163] Heterocentric cluster loss Calculate the loss for the features obtained by the input images {X vis , X ir} through the multi-feature joint progressive learning network. The formula is as follows:
[0164]
[0165] Where, N is the total number of and in the training stage batch, f() is the output before the fully connected layer of the multi-feature joint progressive learning network, mean[A - B] is the mean operation of the difference between A and B, Represents the feature sequence Y' generated by the visible light modality, Represents the feature sequence Y' generated by the infrared modality.
[0166] Id loss Calculate the loss for the features obtained by the input images {X vis , X ir} through the multi-feature joint progressive learning network. The formula is as follows:
[0167]
[0168] Where, P is the number of categories in the image set, y represents the category label of the pedestrian, p a is the predicted probability of category a, ε is a small constant, q a represents the actual label distribution probability, p a is obtained by processing the feature sequence Y' through the softmax function. To correspond to p a , q a also needs to be further processed through the softmax function.
[0169] Average pooling and batch normalization (BN) are performed on the feature sequence Y' to obtain the final feature vector. Finally, the final feature vector is processed through a softmax function to convert the original input into a class probability distribution.
[0170] Triplet loss between modalities For the features obtained by processing the input images {X vis , X ir} through the network, the loss is calculated, which can be expressed by the formula:
[0171]
[0172] where represents the triplet loss calculated from positive sample pairs from the VIS and IR modalities and negative sample pairs from the VIS modality, represents the triplet loss calculated from positive sample pairs from the IR and VIS modalities and negative sample pairs from the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality. and can be expressed by the formula:
[0173]
[0174]
[0175] where M represents randomly selecting M pedestrians in each iteration, D(V b , I c ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the c-th pedestrian in the IR image, D(V b , I k ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the k-th pedestrian in the IR image, D(I b , V c ) represents the Euclidean distance between the b-th pedestrian in the IR image and the c-th pedestrian in the VIS image, D(I b , V k ) represents the Euclidean distance between the b-th pedestrian in the IR image and the k-th pedestrian in the VIS image. b and c are the same, representing the same pedestrian; k is different from b and c, representing different pedestrians. y b represents the class label of the b-th pedestrian in the current modality, y c represents the class label of the same pedestrian in a different modality from y b , y k represents the class label of a pedestrian different from y b and y c , and ξ is the margin parameter, [z] += max(z, 0). (V b , I c ), (I b , V c ) are positive sample pairs, (V b , I k ), (I b , V k ) are negative sample pairs.
[0176] Step S6, Backpropagation of the jointly optimized loss function.
[0177] The obtained loss is jointly optimized, and the model can perform progressive learning from coarse-grained to fine-grained, thus learning more diverse features. The overall loss is composed of the losses in three stages:
[0178]
[0179] Among them, is the overall loss of the first stage, is the overall loss of the second stage, is the overall loss of the third stage, and can be expressed by the formula:
[0180]
[0181]
[0182]
[0183] Each stage calculates and The second and third stages calculate However, since the data used to calculate the loss in each stage is different, the calculation results are different. For the convenience of distinction, the in the first stage can be denoted as The in the second stage is denoted as The in the third stage is denoted as The in the first stage is denoted as The in the second stage is denoted as The in the third stage is denoted as The in the second stage is denoted as The in the third stage is denoted as
[0184] Step S7: After the features of the query image and the images in the image library are extracted by the dual-stream Resnet-50 network respectively, they are matched.
[0185] In the inference stage, a query image in the visible light (infrared) modality and an infrared (visible light) image library are given. In this example, the query image and the images in the image library are used as the inputs of the multi-feature joint progressive learning network, and these images are projected into the same feature space, and the dual-stream Resnet-50 network outputs the human features. By calculating the feature similarity between the human features, the e images in the image library with the highest feature similarity to the query image are used as candidate images. These candidate images are the re-identification results obtained by the multi-feature joint progressive learning network.
[0186] The above-described embodiments only represent one or several implementation manners of the present invention. The description is relatively specific and detailed, but it cannot be understood as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.
Claims
1. A cross-modal pedestrian re-identification method based on multi-feature joint progressive learning, characterized in that The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning includes: Divide the given image set \(D = \{X vis , X ir \}\) into a visible light image set \(X vis \) and an infrared image set \(X ir \); Construct a multi-feature joint progressive learning network, which includes a two-stream Resnet-50 network, a contour feature enhancement module, and a batch adaptive module. The two-stream Resnet-50 network contains two branches and a softmax function, and the two branches are respectively used to extract the features of visible light images and infrared images; Input the image set D into the contour feature enhancement module and the two-stream Resnet-50 network respectively to obtain the original image feature map and the contour feature map, obtain the perceptual contour loss according to the original image feature map and the contour feature map, and calculate the overall loss in the first stage; The visible light images in the image set D are grayscale processed to obtain corresponding grayscale images, and the grayscale image set X gray and the infrared image set X ir are used as the inputs of the dual-stream Resnet-50 network to calculate the overall loss in the second stage; Take the image set D as the input of the two-stream Resnet-50 network, and input the obtained feature sequence into the batch adaptive module to calculate the overall loss in the third stage; Jointly optimize the overall losses of the three stages, and optimize the parameters of the multi-feature joint progressive learning network through backpropagation; Given a pedestrian query image and an image gallery as the input of the trained multi-feature joint progressive learning network, calculate the similarity between the query image and all images in the image gallery by using the output features to complete the re-identification process.
2. The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning according to claim 1, wherein Both branches of the dual-stream ResNet-50 network include a sub-network with independent parameters and a sub-network with shared parameters. The sub-networks with independent parameters of the two branches are denoted as E V and E IR , which are composed of the first layer of each branch of the dual-stream ResNet-50 network. The sub-network with shared parameters is denoted as E S , which is composed of the remaining layers of each branch of the dual-stream ResNet-50 network.
3. The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning according to claim 2, wherein The step of inputting the image set D into the contour feature enhancement module and the two-stream Resnet-50 network respectively to obtain the original image feature map and the contour feature map, obtaining the perceptual contour loss according to the original image feature map and the contour feature map, and calculating the overall loss in the first stage includes: Input the image set D into the sub-network E with independent parameters of the two-stream Resnet-50 network V 、E IR to obtain the feature map C1; Input the image set D into the Sobel convolutional layer of the contour feature enhancement module to obtain the feature map C2; Perform element-wise addition on C1 and C2, and the element-wise added feature map is trained through the first to fourth blocks of the VGG-16 network in the contour feature enhancement module. Let ψ = {ψ1, ψ2, ψ3, ψ4} represent the four blocks of VGG-16; The total perceptual contour loss finally obtained can be expressed by the formula as follows: Among them, represents the original image feature map obtained by passing the input through the i-th block of the VGG-16 network, where i ∈ [1, 4]. represents the original image feature map obtained by extracting the feature map of the visible light modality through the (i - 1)-th block of the VGG-16 network. represents the input and the contour feature map obtained by passing it through the i-th block of the VGG-16 network. represents the contour feature map obtained by extracting the feature map of the visible light modality through the (i - 1)-th block of the VGG-16 network. represents the input and the original image feature map obtained by passing it through the i-th block of the VGG-16 network. represents the original image feature map obtained by extracting the feature map of the infrared modality through the (i - 1)-th block of the VGG-16 network. represents the input and the contour feature map obtained by passing it through the i-th block of the VGG-16 network. represents the contour feature map obtained by extracting the feature map of the infrared modality through the (i - 1)-th block of the VGG-16 network; According to the original image feature map and the contour feature map the perceptual contour loss is obtained and can be calculated and expressed by the following formula: Among them, represents the perceptual contour loss, C represents calculating the loss for the current channel, and the original image feature map and the contour feature map are of the shape C i ×H i ×W i , C i represents the total number of channels, H i represents the height, W i represents the width, m ∈ {vis, ir} represents two modalities of visible light and infrared, represents calculating and summing the losses for each channel in the features of each modality, and L δ represents the Huber loss, which can be expressed by the formula: Among them, δ represents a hyperparameter; Total loss in the first stage It can be expressed by the following formula: Among them, represents the id loss, and calculates the loss of the features obtained by processing the input images {X vis , X ir} through the multi-feature joint progressive learning network. The formula is as follows: Among them, P is the number of categories in the image set, y represents the category label of the pedestrian, and p a is the predicted probability of category a, ε is a small constant, and q a represents the actual label distribution probability, and p a is obtained based on the original image feature map after the processed feature vector is processed by the softmax function; Triplet loss between modalities Calculate the loss of the features obtained by processing the input images {X vis , X ir} through the multi-feature joint progressive learning network. It can be expressed by the formula: Among them, represents the triplet loss calculated from the positive sample pairs of the VIS and IR modalities and the negative sample pairs of the VIS modality, represents the triplet loss calculated from the positive sample pairs of the IR and VIS modalities and the negative sample pairs of the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality. and can be expressed by the formula: Among them, M represents randomly selecting M pedestrians each time for iteration, D(V b , I c ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the c-th pedestrian in the IR image, D(V b , I k ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the k-th pedestrian in the IR image, D(I b , V c ) represents the Euclidean distance between the b-th pedestrian in the IR image and the c-th pedestrian in the VIS image, D(I b , V k ) represents the Euclidean distance between the b-th pedestrian in the IR image and the k-th pedestrian in the VIS image, b and c are the same, representing the same pedestrian; k is different from b and c, representing different pedestrians, y b represents the class label of the b-th pedestrian in the current modality, y c represents the class label of the same pedestrian in a different modality from y b , y k represents the class label of a pedestrian different from y b and y c , ξ is the interval parameter, [z] + = max(z, 0).
4. The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning according to claim 3, wherein The step of inputting the image set D into the Sobel convolutional layer of the contour feature enhancement module to obtain the feature map C2 includes: By inputting the image set D into the Sobel convolutional layer of the contour feature enhancement module, obtain the horizontal and vertical image gray values. The formula is as follows: Among them, A represents the original image in the image set D, and G x represents the horizontal image gray value, and G y represents the vertical image gray value; At each point of the image, combine the horizontal and vertical image gray values, and calculate the approximate gradient of each point through the following formula: Obtain the feature map C2 according to the approximate gradient.
5. The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning according to claim 1, characterized in that Performing grayscale processing on the visible light images in the image set D to obtain corresponding grayscale images, and using the grayscale image set X gray and the infrared image set X ir as the input of the dual-stream Resnet-50 network to calculate the overall loss in the second stage, including: The visible light images in the image set D are grayscaled to obtain grayscale images; Input the grayscale image set X gray and the infrared image set X ir into the two-stream Resnet-50 network, output the human grayscale features of the grayscale images, and calculate the overall loss in the second stage which can be expressed by the following formula: Among them, represents the off-center cluster loss, represents the id loss, represents the triplet loss; Off-center cluster loss Calculate the loss of the features obtained from the input images {X gray , X ir} through the multi-feature joint progressive learning network. The formula is as follows: where N is the total number in the training phase batches and f() is the output before the fully connected layer of the dual-stream Resnet-50 network, and mean[A - B] is the mean operation of the difference between A and B. denotes the grayscale image generated by the visible light modality, denotes the input image of the infrared modality; id loss Calculate the loss of the features obtained from the input images {X gray , X ir} through the multi-feature joint progressive learning network. The formula is as follows: where P is the number of categories in the image set, y represents the category label of the pedestrian, and p a is the predicted probability of category a, ε is a small constant, and q a represents the actual label distribution probability, and p a is obtained by processing the grayscale features of the person through the softmax function; Triplet loss between modalities Calculate the loss of the features obtained by processing the input images {X gray , X ir} through a multi-feature joint progressive learning network. It can be expressed by the formula: Among them, represents the triplet loss calculated from the positive sample pairs of the VIS and IR modalities and the negative sample pairs of the VIS modality, represents the triplet loss calculated from the positive sample pairs of the IR and VIS modalities and the negative sample pairs of the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality. and can be expressed by the formula: Among them, M represents randomly selecting M pedestrians each time for iteration, D(V b , I c ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the c-th pedestrian in the IR image, D(V b , I k ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the k-th pedestrian in the IR image, D(I b , V c ) represents the Euclidean distance between the b-th pedestrian in the IR image and the c-th pedestrian in the VIS image, D(I b , V k ) represents the Euclidean distance between the b-th pedestrian in the IR image and the k-th pedestrian in the VIS image, b and c are the same, representing the same pedestrian; k is different from b and c, representing different pedestrians, y b represents the class label of the b-th pedestrian in the current modality, y c represents the class label of the same pedestrian in a different modality from y b , y k represents the class label of a pedestrian different from y b and y c , ξ is the interval parameter, [z] + = max(z, 0).
6. The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning according to claim 1, characterized in that The step of taking the image set D as the input of the two-stream Resnet-50 network, inputting the obtained feature map into the batch adaptive module, and calculating the overall loss in the third stage includes: Input the image set D into the two-stream Resnet-50 network to obtain the feature sequence Y, where Y ∈ R E×F , E and F respectively represent the length and dimension of the feature sequence. Replace the dimension of the feature sequence Y: Input into the batch adaptive module to obtain The batch adaptive module can be expressed as: Among them, MHSA represents multi-head attention, FF represents a fully connected layer, and LN1, LN2 represent linear layers; Restore the dimension of the feature sequence : Calculate the loss of the feature sequence after passing through the batch adaptive module to obtain the overall loss in the third stage It can be expressed by the following formula: Among them, represents the off-center cluster loss, represents the id loss, represents the triplet loss; Off-center cluster loss Calculate the loss of the features obtained from the input images {X vis , X ir} through the multi-feature joint progressive learning network. The formula is as follows: where N is the total number in the batches during the training phase and , f() is the output before the fully connected layer of the multi-feature joint progressive learning network, mean[A - B] is the mean operation of the difference between A and B, denotes the feature sequence Y' generated by the visible light modality, denotes the feature sequence Y' generated by the infrared modality; id loss For the features obtained from the input images {X vis , X ir} through the multi-feature joint progressive learning network, the loss is calculated as follows: Among them, P is the number of categories in the image set, y represents the category label of the pedestrian, and p a is the predicted probability of category a, ε is a small constant, and q a represents the actual label distribution probability, and p a is obtained by processing the feature sequence Y' through the softmax function; Triplet loss between modalities Calculate the loss for the features obtained by processing the input images {X vis , X ir} through the network. It can be expressed by the formula: Among them, represents the triplet loss calculated from positive sample pairs from the VIS and IR modalities and negative sample pairs from the VIS modality, represents the triplet loss calculated from positive sample pairs from the IR and VIS modalities and negative sample pairs from the IR modality. The VIS modality represents the visible light modality, and the IR modality represents the infrared modality. and can be expressed by the formula: Among them, M represents randomly selecting M pedestrians each time, D(V b , I c ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the c-th pedestrian in the IR image, D(V b , I k ) represents the Euclidean distance between the b-th pedestrian in the VIS image and the k-th pedestrian in the IR image, D(I b , V c ) represents the Euclidean distance between the b-th pedestrian in the IR image and the c-th pedestrian in the VIS image, D(I b , V k ) represents the Euclidean distance between the b-th pedestrian in the IR image and the k-th pedestrian in the VIS image, b and c are the same, representing the same pedestrian; k is different from b and c, representing different pedestrians, y b represents the class label of the b-th pedestrian in the current modality, y c represents the class label of the same pedestrian in a different modality from y b , y k represents the class label of a pedestrian different from y b and y c , ξ is the interval parameter, [z] + = max(z, 0).
7. The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning according to claim 1, wherein The step of jointly optimizing the overall losses of the three stages includes: The overall losses of the three stages are obtained by adding the overall loss in the first stage, the overall loss in the second stage, and the overall loss in the third stage.
8. The cross-modal pedestrian re-identification method based on multi-feature joint progressive learning according to claim 1, characterized in that The re-identification process includes: Given a query image in the visible light modality and an infrared image gallery, or given a query image in the infrared modality and a visible light image gallery; The given query image and image library are input into the multi-feature joint progressive learning network, and the two-stream Resnet-50 network outputs the character features of the query image and the images in the image library; The feature similarity between the query image and all images in the image library is calculated, and the e images in the image library with the highest feature similarity to the query image are taken as candidate images, which are recorded as the re-identification results obtained by the multi-feature joint progressive learning network.