Cross-modal pedestrian re-identification method based on feature enhancement
By introducing feature enhancement methods and biological evolution simulations in the cross-modal pedestrian re-identification technology, optimizing feature fusion and adapting to cross-modal data differences, the problem of poor recognition effect in the existing technology under different lighting conditions is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510207764.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing cross-modal pedestrian re-identification technology has poor recognition effect under different lighting conditions, and due to insufficient training data, there is a possibility of overfitting.
A cross-modal pedestrian re-identification method based on feature enhancement is adopted, a Vision Transformer with shared parameters is used as the backbone network, and the intersection and mutation process of biological evolution is simulated in the patch embedding process, and a attention fusion feature module and a Gaussian variant feature module are introduced to optimize feature fusion and adapt to the differences in cross-modal data.
The recognition accuracy and robustness of the model under different imaging conditions are improved, and the adaptability to interference factors such as noise and occlusion is enhanced, and the accuracy of cross-modal pedestrian re-identification is achieved.
Smart Images

Figure CN120220043A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a cross-modal pedestrian re-identification method based on feature enhancement. Background Art
[0002] In modern society, intelligent monitoring systems play a crucial role in maintaining public safety and improving urban management efficiency. As the core of the monitoring system, pedestrian re-identification technology mainly aims to identify and track the same pedestrian in images captured by different cameras. Although visible light images provide rich visual information for pedestrian identification, the recognition effect will decline significantly at night or under poor lighting conditions. To overcome this limitation, cross-modal pedestrian re-identification technology has emerged, which combines visible light and infrared images to achieve stable recognition under various lighting conditions. However, the current scale of cross-modal pedestrian re-identification datasets is limited, and due to insufficient training data, the possibility of overfitting is very high. A feasible solution to this problem is to implement feature enhancement technology during the training process.
[0003] With the continuous progress of deep learning technology, especially the introduction of the Vision Transformer architecture, it provides a new perspective for processing such cross-modal data. The present invention will explore the use of feature enhancement algorithms to optimize the cross-modal pedestrian re-identification model, aiming to improve its recognition accuracy and robustness under different imaging conditions, so as to achieve more reliable pedestrian identification in all-weather monitoring scenarios. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a cross-modal pedestrian re-identification method based on feature enhancement, which uses a Vision Transformer with shared parameters as the backbone network, simulates the crossover and mutation processes of biological evolution in the patch embedding link, and proposes an attention fusion feature module and a Gaussian mutation feature module. This not only optimizes the effect of feature fusion, but also enables the model to better adapt to the differences between cross-modal data and improves the adaptability to environmental changes.
[0005] The technical solution of the present invention is as follows:
[0006] A cross-modal pedestrian re-identification method based on feature enhancement, comprising the following steps:
[0007] Step 1, obtain a dataset and perform preprocessing, and generate new modalities using a channel enhancement strategy;
[0008] Step 2, construct a cross-modal pedestrian re-identification network model based on feature enhancement;
[0009] Step 3, construct a loss function, and train and optimize the model based on the training dataset and the loss function;
[0010] Step 4. Perform cross-modal pedestrian re-identification based on the trained model.
[0011] Further, the specific process of the said Step 1 is as follows:
[0012] Step 1.1. Obtain the public datasets SYSU-MM01 and RegDB as the training datasets; collect all pedestrian images captured by each surveillance camera as the test datasets; the test datasets include two parts, namely the query set and the gallery set. The query set is the set of current pedestrian images to be queried, and the gallery set is the set of candidate pedestrian images to be matched with the query set; the data in the datasets includes visible light images and infrared images.
[0013] Step 1.2. Generate new enhanced images from the visible light images in the training datasets by using the method of channel enhancement. At this time, the training datasets contain images of three modalities, namely visible light images, infrared images, and enhanced images. One type of image corresponds to one modality.
[0014] Step 1.3. Perform image preprocessing on the images of the three modalities in the training datasets. Randomly horizontally flip and regularize the visible light images, perform Gaussian blur and brightness adjustment on the infrared images, and perform random erasing on the enhanced images.
[0015] Step 1.4. Resize the images of these three modalities to 256*128 pixels.
[0016] Further, in the said Step 2, the cross-modal pedestrian re-identification network model includes three parts, namely: the patch embedding part, the feature enhancement algorithm part, and the encoding part; the patch embedding part includes image segmentation and embedding operations; the feature enhancement algorithm part contains an attention fusion feature module and a Gaussian mutation feature module; the encoding part includes three weight-sharing encoders.
[0017] The working process of the cross-modal pedestrian re-identification network model is as follows:
[0018] Step 2.1. Perform image segmentation and embedding in the patch embedding part to obtain a series of patch sets p of the image set t t ;
[0019] Step 2.2. Obtain the fused patches based on the attention fusion feature module and update the patch set.
[0020] Step 2.3. The Gaussian mutation feature module mutates the patches by using the Gaussian mutation strategy and updates the patch set.
[0021] Step 2.4. A series of patch sets p of the original input image set t tIn the input feature enhancement algorithm part, the attention fusion feature module and the Gaussian mutation feature module in the feature enhancement algorithm part are executed in parallel according to the processes of step 2.2 and step 2.3, and finally a new series of patch sets p are optimized. t′ , and p t′ is input into three weight - shared encoders. One encoder processes images of one modality to obtain the predicted value features of each modality.
[0022] Further, the specific process of step 2.1 is as follows: The visible - light image set the infrared image set X r = the enhanced image set is input into the model, where is the i - th visible - light image, and N v is the number of visible - light images; is the - th infrared image, and N r is the number of infrared images; is the - th enhanced image, and N c is the number of enhanced images; For the visible - light, infrared, and enhanced image sets, corresponding identification label sets are defined, which are the visible - light image identification label set Y v = the infrared - image identification label set and the enhanced - image identification label set Define the image set X t to represent one of the visible - light image set X v , the infrared - image set X r or the enhanced - image set X c , X t ∈ {X v , X r , X c};
[0023] Before inputting into the model, it is necessary to perform a patch embedding operation on the input images: Each image in the image set is sliced into a series of patches, and the patches are mapped to a high - dimensional feature space to obtain the corresponding patch set. The formula is:
[0024] p t = PatchEmbedding(X t );
[0025] Among them, p t is a series of patch sets of the image set X t ; PatchEmbedding(·) is the patch embedding operation; The number of images in the image set corresponds to the number of patch sets.
[0026] Further, the specific process of step 2.2 is as follows:
[0027] Step 2.2.1: Preset the crossover rate, and randomly select two different patch sets with the same identity according to the crossover rate Perform attention fusion feature operation, which are the i-th and I-th patches in p a respectively; which are the j-th and J-th patches in p b respectively. Multiply and normalize these two patch sets to construct an attention matrix, and the formula is:
[0028] W = softmax(p a * p b );
[0029] where W is the attention matrix; softmax(·) is the softmax function;
[0030] Step 2.2.2: Randomly select a patch a from the patch set p , and use the attention matrix to find the patch b in the patch set p with the highest similarity to . Determine the index j of the current and the adjacent indices around j as valid indices; Extract the corresponding eigenvalues from p b according to the valid indices, and extract the attention scores corresponding to the indices from the attention matrix W. Then, renormalize the extracted attention scores and construct a new attention matrix, specifically:
[0031] W' = softmax(w j ), j ∈ vaild indices ;
[0032] where W' is the new attention matrix; w j is the attention score at index j in the attention matrix W; vaild indices is the valid index;
[0033] Step 2.2.3: Obtain the fused patch according to the new attention matrix and the corresponding eigenvalues, and the formula is:
[0034]
[0035] where p r is the fused patch; w j ' is the attention score at index j in the new attention matrix W', and the index numbers of the attention matrix and the new attention matrix correspond;
[0036] Step 2.2.4. Finally, replace p a in with the fused patch p r to obtain the updated patch set
[0037] Furthermore, the specific process of Step 2.3 is as follows: First, define a series of patches with the same identity in the input as the original patch set p s , calculate the statistics of p s , including the mean and variance; according to the calculated mean μ and variance σ 2 , define a Gaussian distribution N(μ, σ 2 ); select a series of patches of one identity from p s to form the original mutant patch set where s is a specific identity and m is a sample in a specific identity, is the m -th patch in p , is the m -th patch in p ; sample from the Gaussian distribution, traverse each patch in the original mutant patch set p m , preset the mutation rate, and replace the individual with a point sampled from the Gaussian distribution according to the mutation rate; the calculation formulas for the mean and variance are:
[0038]
[0039]
[0040] where is the s -th patch in the original patch set p s ; H is the number of patches in the original patch set p
[0041] The new feature sample p new sampled from the Gaussian model conforms to the Gaussian distribution, specifically:
[0042] p new ~N(μ, σ 2 );
[0043] Finally, replace the samples in the original mutant patch set with the new feature sample p new sampled to obtain the new mutant patch set
[0044] Furthermore, the formula in step 2.4 is as follows:
[0045]
[0046] where F t is the predicted value feature of the image set X t ; is an encoder with shared weights.
[0047] Furthermore, the specific process of step 3 is as follows:
[0048] Step 3.1: Input the predicted value features of each image set into the classification layer to obtain the predicted probability of the image for each identity, and calculate the identity loss L id :
[0049] S t = Classification(F t );
[0050]
[0051] where S t is the predicted probability of the image set X t ; Classification(·) is the classification layer; Y t represents the true label of the image set X t ; represents the total number of identity samples; represents the predicted probability of the image set X t for the th identity;
[0052] Step 3.2: Calculate the triplet loss L tri :
[0053]
[0054] where F is the predicted value feature; represents a randomly selected sample image; o represents the positive sample image with the same identity as the sample; o′ represents the negative sample image with a different identity from the sample; margin is the parameter boundary;
[0055] Step 3.3: Calculate the discriminative center loss L dcl :
[0056]
[0057] where is the center vector; is the z-th identity; K represents the number of predicted value features of the same modality in the same identity; is the z-th predicted value feature in the visible light image; is the k-th predicted value feature in the infrared image; is the average distance between the predicted value features of all identities except and the central vector of ; F represents the predicted value feature of the z'-th sample different from the -th identity; F is the predicted value feature of the k'-th sample with the same identity as the z, -th identity; Step 3.4. Construct the least squares error loss function L k′ Randomly select a predicted value feature of an infrared image or a visible light image, which are respectively represented as the predicted value feature of the -th infrared image
[0058] or the predicted value feature of the msel -th visible light image First, calculate the average distance between and other samples with the same identity in the intra-modal and cross-modal cases. The calculation formula is:
[0059]
[0060] intra where S cross is the average distance in the intra-modal case; S is the average distance in the cross-modal case; S(·) is the Euclidean distance; is the predicted value feature of the -th infrared image; is the predicted value feature of the
[0061] Calculate the average distance between and other samples with the same identity in the intra-modal and cross-modal cases in the same way as the above formula;
[0062] Calculate the difference L msel between the average distance in the intra-modal case and the average distance in the cross-modal case:
[0063]
[0064] where is the average distance in the intra-modal case corresponding to the -th predicted value feature in the same identity; is the average distance in the cross-modal corresponding to the th predicted value feature in the same identity;
[0065] Step 3.5. Finally, the overall loss function L of the training process is:
[0066] L = λL msel + λL dcl + L id + L tri ;
[0067] where λ is a hyperparameter.
[0068] Furthermore, the specific process of step 4 is as follows:
[0069] Step 4.1. Use the query set and gallery set of the test data set as the input of the cross-modal person re-identification model trained in step 3, and splice the image features of the three modalities output by the model together in the channel dimension to obtain the final person predicted value feature;
[0070] Step 4.2. Calculate the similarity between the person images in the query set and each person image in the gallery set; the similarity calculation formula is:
[0071]
[0072] where is the similarity between the person image in the query set and the person image in the gallery set; represents the predicted value feature of the person image in the query set; represents the predicted value feature of the person image in the gallery set; |·| represents the modulus length;
[0073] Step 4.3. Sort all the similarity values in descending order, and output the images corresponding to the top ten person predicted value features with the highest similarity values as the re-identification results.
[0074] The beneficial technical effects brought by the present invention: The method of the present invention proposes cross-modal person re-identification based on feature enhancement to solve the significant differences between visible light and infrared images. By simulating the crossover and mutation operations in the biological evolution process, it aims to optimize the feature fusion process. This method not only improves the adaptability of the model to changes between different modalities but also enhances its robustness to interference factors such as noise and occlusion. Through the feature enhancement strategy, the model can automatically explore and utilize the complementary information in multi-modal data, thus achieving higher accuracy in complex cross-modal recognition tasks. Description of the Drawings
[0075] Figure 1 This is the flowchart of the cross-modal person re-identification method based on feature enhancement of the present invention.
[0076] Figure 2 This is the structural schematic diagram of the cross-modal person re-identification model based on feature enhancement of the present invention.
[0077] Figure 3 is Figure 2 the schematic diagram of the attention fusion feature module in
[0078] Figure 4 is Figure 2 the schematic diagram of the Gaussian mutation feature module in Detailed implementation manners
[0079] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners:
[0080] First, the explanations of the following terms are given:
[0081] PatchEmbedding: PatchEmbedding is a commonly used technique in Vision Transformer. Its core function is to divide the input image into multiple small patches and convert each small patch into a vector with a fixed dimension so that the subsequent Transformer model can process it.
[0082] Softmax: Softmax is a commonly used activation function, mainly used in multi-classification problems to convert the input real values into a probability distribution. It is widely used in machine learning and deep learning, especially in the output layer of neural networks, to convert the output of the model into class probabilities.
[0083] CA: CA (Channel Attention) channel enhancement is a technique in deep learning, especially in the field of image processing, used to enhance the model's attention to image channels (features). The core idea of this technique is to enable the model to adaptively emphasize important channel features while suppressing less important channel features. CA channel enhancement is usually used in convolutional neural networks (CNNs) to improve the model's ability to represent image features.
[0084] Vision transformer: Vision Transformer (ViT) is a deep learning model based on the Transformer architecture. It was originally designed for natural language processing (NLP) tasks but has since been successfully applied to the field of computer vision, particularly in image classification tasks. The core idea of ViT is to divide an image into multiple small patches, and then treat these patches as words or tokens in a sequence. The Transformer model is used to process these sequences to achieve image understanding and classification.
[0085] SYSU-MM01 dataset: The SYSU-MM01 dataset was created by researchers at South China University of Technology to address the problem of identifying the same pedestrian under different imaging modalities. SYSU-MM01 was proposed in 2017. It comes from 6 cameras including 2 infrared cameras and 4 RGB cameras. It includes RGB images and IR images of 491 identities, resulting in a total of 287,628 RGB images and 15,792 IR images. The dataset has two modes, namely the full search mode and the indoor search mode.
[0086] RegDB dataset: The RegDB dataset was proposed in 2017. It is a small-scale dataset collected by a dual-camera system, including a visible light camera and a thermal camera. The dataset contains a total of 412 personal identities, with 10 visible light images and 10 infrared images for each identity. The database contains 4,120 visible light images and 4,120 corresponding infrared images. The training set and the test set each have 206 pedestrians.
[0087] As Figure 1 shown, the method of the present invention includes the following steps:
[0088] Step 1: Obtain the dataset and perform preprocessing, and generate new modalities using the channel enhancement strategy. The specific process is as follows:
[0089] Step 1.1: Obtain the publicly available datasets SYSU-MM01 and RegDB as the training datasets; collect all pedestrian images under each monitoring camera as the test dataset; the test dataset contains two parts, the query set and the gallery set. The query set is the set of current pedestrian images to be queried, and the gallery set is the set of candidate pedestrian images to be matched with the query set; the data in the dataset are visible light images and infrared images;
[0090] Step 1.2: Generate new enhanced images from the visible light images in the training dataset using the channel enhancement method. At this time, the training dataset contains three modalities of images: visible light images, infrared images, and enhanced images. One type of image corresponds to one modality;
[0091] Step 1.3: Perform image preprocessing on the images of the three modalities in the training dataset. Specifically, perform operations such as random horizontal flipping and regularization on the visible light images, Gaussian blurring and brightness adjustment on the infrared images, and random erasing on the enhanced images;
[0092] Step 1.4: Resize the images of these three modalities to 256*128 pixels.
[0093] Step 2: Construct a cross-modal person re-identification network model based on feature enhancement, which incorporates an attention fusion feature module and a Gaussian mutation feature module;
[0094] The cross-modal person re-identification network model based on feature enhancement consists of three parts, namely: the patch embedding part, the feature enhancement algorithm part, and the encoding part; the patch embedding part includes image segmentation and embedding operations; the feature enhancement algorithm part contains an attention fusion feature module and a Gaussian mutation feature module; the encoding part is a Vision transformer structure, which includes regularization, attention mechanism modules, etc.
[0095] As Figure 2 , Figure 3 , Figure 4 shown, the working process of the cross-modal person re-identification network model based on feature enhancement is as follows:
[0096] Step 2.1: Perform image segmentation and embedding in the patch embedding part; the specific process is as follows: Input the visible light image set (RGB images) infrared image set enhanced image set into the model, where is the i-th visible light image, and N v is the number of visible light images; is the -th infrared image, and N r is the number of infrared images; is the -th enhanced image, and N c is the number of enhanced images. For the visible light, infrared, and enhanced image sets, define the corresponding identification label sets, namely the visible light image identification label set infrared image identification label set and the enhanced image identification label set whose label candidate sets are shared. For simplicity, define the image set X t to represent one of the visible light image set X v , infrared image set X r , or enhanced image set X c , where X t ∈{X v,X r ,X c}。
[0097] Before inputting into the model, it is necessary to perform patch embedding operation on the input image: each image in the image set is sliced into a series of patches, and the patches are mapped to a high-dimensional feature space to obtain the corresponding patch set. The formula is:
[0098] p t = PatchEmbedding(X t );
[0099] Among them, p t is a series of patch sets of the image set X t ; PatchEmbedding(·) is the patch embedding operation; the number of images in the image set corresponds to the number of patch sets;
[0100] Step 2.2. Obtain the fused patches based on the attention fusion feature module and update the patch set; the specific process is as follows:
[0101] Step 2.2.1. Preset the crossover rate (set to 10% in the present invention), and randomly select two groups of patch sets with the same identity according to the crossover rate to perform the attention fusion feature operation, are the i-th and I-th patches in p a respectively; are the j-th and J-th patches in p b respectively. Multiply and normalize these two groups of patch sets to construct the attention matrix. The formula is:
[0102] W = sogmtax(p a *p b );
[0103] Among them, W is the attention matrix; softmax(·) is the softmax function; p a , p b are two different groups of patch sets;
[0104] Step 2.2.2. Randomly select a patch from the patch set p a Using the attention matrix, find the patch in the patch set p with the highest similarity to b Determine the index j of the current and the adjacent indices around j (such as j + 1, j - 1) as valid indices. According to the valid indices, select from p Determine the current index j and the adjacent indices around j (such as j + 1, j - 1) as valid indices. According to the valid indices, select from p bExtract the corresponding eigenvalue, extract the corresponding attention score from the attention matrix W, and then renormalize the extracted attention score and construct a new attention matrix. Specifically:
[0105] W′ = softmax(w j ), j ∈ valid indices ;
[0106] where W′ is the new attention matrix; w j is the attention score at index j in the attention matrix W; valid indices is the valid index;
[0107] Step 2.2.3, According to the new attention matrix and the corresponding eigenvalue, obtain the fused patch. The formula is:
[0108]
[0109] where p r is the fused patch; w j ′ is the attention score at index j in the new attention matrix W′, and the index numbers of the attention matrix and the new attention matrix correspond;
[0110] Step 2.2.4, Finally, replace the a in p with the fused patch p r , to obtain the updated patch set
[0111] Step 2.3, The Gaussian mutation feature module mutates the patches using the Gaussian mutation strategy and updates the patch set; the core of the mutation strategy is to select elements from the input tensor, perform splicing and statistical analysis, and perform random mutation based on the Gaussian distribution. First, define a series of patches with the same identity in the input as the original patch set p s , calculate the statistics of p s , including the mean and variance. These statistics are used to define a Gaussian distribution, which will be used to generate random mutation points. According to the calculated mean μ and variance σ 2 , a Gaussian distribution N(μ, σ 2 ) is defined. Select a series of patches of one of the identities from p s to form the original mutated patch set s is a specific identity, m is a sample in a specific identity, is the m th patch in p , is the m th patch in p A patch; sample from a Gaussian distribution, and traverse each patch in the original mutated patch set p m in the set, preset the mutation rate (set to 20% in the present invention), and replace the individual with a point sampled from the Gaussian distribution according to the mutation rate. The calculation formulas for the mean and variance are:
[0112]
[0113]
[0114] wherein, is the h-th patch in the original patch set p s ; H is the number of patches in the original patch set p s .
[0115] The new feature sample p new extracted from the Gaussian model conforms to the Gaussian distribution, specifically:
[0116] p new ~N(μ,σ 2 );
[0117] Finally, replace the samples in the original mutated patch set with the newly sampled feature sample p new to obtain a new mutated patch set
[0118] Step 2.4: Input a series of patch sets p t of the original input image set t into the feature enhancement algorithm part. The attention fusion feature module and the Gaussian mutation feature module in the feature enhancement algorithm part are executed in parallel according to the processes of Step 2.3 and Step 2.4, and finally a new series of patch sets p t′ is optimized, and p t′ is input into three weight-sharing encoders. One encoder processes the images of one modality to obtain the predicted value features of each modality, specifically:
[0119]
[0120] wherein, F t is the predicted value feature of the image set X t ; is the weight-sharing encoder.
[0121] Step 3: Construct a loss function, and train and optimize the model based on the training data set and the loss function. The specific process is as follows:
[0122] Step 3.1: Input the predicted value features of each image set into the classification layer to obtain the prediction probability of the image for each identity, and calculate the identity loss L based on the prediction probability id, the calculation formula is as follows:
[0123] S t = Classification(F t );
[0124]
[0125] Among them, S t is the predicted probability of the image set X t ; Classification(·) is the classification layer; Y t represents the true label of the image set X t ; represents the total number of identity samples; represents the predicted probability of the image set X t for the th identity.
[0126] Step 3.2, calculate the triplet loss L tri of the image set, and the calculation formula is as follows:
[0127]
[0128] Among them, F is the predicted value feature; represents a randomly selected sample image; o represents the positive sample image with the same identity as the sample; o' represents the negative sample image with a different identity from the sample; margin is the parameter boundary, set to 0.3.
[0129] Step 3.3, calculate the discriminative center loss of the image set, and the calculation formula of the discriminative center loss L dcl is as follows:
[0130]
[0131] Among them, is the center vector; is the th identity; K represents the number of predicted value features of the same modality in the same identity; is the zth predicted value feature of in the visible light image; is the kth predicted value feature of in the infrared image; is the average value of the distances between the predicted value features of all identities except and the center vector; F z′ represents the predicted value feature of the z'th sample different from the th identity; Fk′ For the predicted value feature of the k'-th sample that is the same as the th identity;
[0132] Step 3.4, construct the least squares error loss function L msel , randomly select the predicted value feature of an infrared image or a visible light image, which are respectively expressed as the predicted value feature of the th infrared image or the predicted value feature of the th visible light image First, calculate the average distance from other samples with the same identity in the intra-modal and cross-modal modes. The calculation formula is:
[0133]
[0134] where D intra is the average distance in the intra-modal mode; D cross is the average distance in the cross-modal mode; D(·) is the Euclidean distance; is the predicted value feature of the th infrared image; is the predicted value feature of the th visible light image;
[0135] Calculate the average distance from other samples with the same identity in the intra-modal and cross-modal modes in the same way as the above formula;
[0136] Calculate the difference L msel between the average distance in the intra-modal mode and the average distance in the cross-modal mode:
[0137]
[0138] where is the average distance in the intra-modal mode corresponding to the th predicted value feature in the same identity; is the average distance in the cross-modal mode corresponding to the th predicted value feature in the same identity;
[0139] Step 3.5, finally, the overall loss function L of the training process is defined as:
[0140] L = λL msel + λL dcl + L id + L tri ;
[0141] where λ is a hyperparameter, set to 0.5, which is used to balance the importance of the loss.
[0142] The cross-modal person re-identification model is constrained by an overall loss function to train and optimize to obtain a more effective and robust cross-modal person re-identification model.
[0143] Step 4: Perform cross-modal person re-identification based on the trained model. The specific process is as follows:
[0144] Step 4.1: Use the query set and gallery set of the test data set as the input of the cross-modal person re-identification model trained in Step 3, and splice the image features of the three modalities output by the model together in the channel dimension to obtain the final pedestrian prediction value features;
[0145] Step 4.2: Calculate the similarity between the pedestrian images in the query set and each pedestrian image in the gallery set;
[0146] The similarity calculation formula is:
[0147]
[0148] Where, is the pedestrian image in the query set and the pedestrian image in the gallery set is the similarity; represents the predicted value feature of the pedestrian image in the query set ; represents the predicted value feature of the pedestrian image in the gallery set ; |·| represents the modulus length;
[0149] Step 4.3: Sort all the similarity values in descending order, and output the images corresponding to the top ten pedestrian prediction value features with the highest similarity values as the re-identification results.
[0150] In the embodiment of the present invention, the dimension of the final feature vector for recognition is 768. The present invention is implemented under the PyTorch framework, uses the Adam algorithm to optimize the model, sets the learning rate to 3.5e-4, and the maximum number of iterations is 100.
[0151] To verify the feasibility and superiority of the present invention, the following comparative experiments were carried out. The experiments were all carried out on two cross-modal pedestrian data sets, SYSU-MM01 and RegDB.
[0152] Four methods, namely SPOT, FMCNet, PMT, and DART, are selected for cross-modal pedestrian re-identification, and the recognition results are compared with those of the present invention. The comparison results are shown in Table 1. The SPOT method proposes a structure-aware position transformer network that utilizes human body structure information to enhance the representation ability of cross-modal features, thereby improving the accuracy of pedestrian recognition under different lighting and environmental conditions. The FMCNet method proposes a feature-level modal compensation network that improves the performance of visible-light-infrared pedestrian re-identification by compensating for missing modality-specific information. The PMT method proposes a progressive learning strategy that uses grayscale images as an auxiliary modality and utilizes the Vision Transformer model to learn modality-independent features. The DART method proposes a new framework for the dual-noise label problem in visible-light-infrared pedestrian re-identification, which corrects the noise correspondence by estimating the confidence of clean annotations and divides the data into four groups to achieve robust learning. The present invention selects two evaluation metrics, namely the first hit rate Rank-1 and the mean average precision mAP, to evaluate the trained model. The higher the values of the first hit rate Rank-1 and the mean average precision mAP, the higher the accuracy of the model.
[0153] Table 1 Comparison results of the method of the present invention and other methods on the SYSU-MM01 and RegDB datasets
[0154]
[0155]
[0156] As can be seen from Table 1, using the method proposed by the present invention, the Rank-1 values of 73.02% and 90.49% and the mAP values of 68.88% and 87.45% can be achieved on the SYSU-MM01 and RegDB datasets for re-identifying pedestrians with changed clothes, respectively. The optimal results are obtained on the SYSU-MM01 dataset and also on the RegDB dataset, effectively improving the accuracy of cross-modal pedestrian re-identification.
[0157] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A cross-modal person re-identification method based on feature enhancement, characterized in that: The steps include: Step 1: Obtain the data set and preprocess it, and use the channel enhancement strategy to generate a new modality; Step 2: Construct a cross-modal person re-identification network model based on feature enhancement; Step 3: Construct a loss function and optimize the model based on the training data set and loss function; Step 4: Perform cross-modal person re-identification based on the trained model.
2. The cross-modal person re-identification method based on feature enhancement according to claim 1, characterized in that: The specific process of step 1 is as follows: Step 1.1, obtain the public datasets SYSU-MM01 and RegDB as training datasets; collect all pedestrian images under each surveillance camera as test datasets; the test dataset contains two parts: query set and gallery set. The query set is the set of pedestrian images to be queried, and the gallery set is the set of candidate pedestrian images matched with the query set; the data in the dataset includes visible light images and infrared images; Step 1.2: Generate a new enhanced image by using the channel enhancement method for the visible light image in the training data set. At this time, the training data set contains images of three modalities: visible light image, infrared image, and enhanced image, and one image corresponds to one modality. Step 1.3: Perform image preprocessing on the images of the three modalities in the training data set, perform random horizontal flipping and regularization on the visible light images, perform Gaussian blurring and brightness adjustment on the infrared images, and perform random erasing on the enhanced images; Step 1.4: Resize the images of these three modalities to 256*128 pixels.
3. The cross-modal person re-identification method based on feature enhancement according to claim 2, characterized in that: In the step 2, the cross-modal person re-identification network model includes three parts, namely: a patch embedding part, a feature enhancement algorithm part and an encoding part; the patch embedding part includes image segmentation and embedding operations; the feature enhancement algorithm part includes an attention fusion feature module and a Gaussian variation feature module; the encoding part includes three weight-sharing encoders; The working process of the cross-modal person re-identification network model is as follows: Step 2.1: Perform image segmentation and embedding in the patch embedding part to obtain a series of patch sets p of the image set t t ; Step 2.2, obtain the fused patches based on the attention fusion feature module and update the patch set; Step 2.3, the Gaussian mutation feature module uses the Gaussian mutation strategy to mutate the patch and update the patch set; Step 2.4: transform the series of patches p of the original input image set t into t Input feature enhancement algorithm part, the attention fusion feature module and Gaussian variation feature module in the feature enhancement algorithm part are executed in parallel according to the process of step 2.2 and step 2.3, and finally optimize to obtain a new series of patch sets p t′ , and p t′ The input is fed into three weight-sharing encoders, where one encoder processes images of one modality to obtain predicted value features for each modality.
4. The cross-modal person re-identification method based on feature enhancement according to claim 3 is characterized in that: The specific process of step 2.1 is: Infrared image collection Enhanced Image Set Input into the model, where It is visible light images, N v is the number of visible light images; It is Infrared images, N r is the number of infrared images; It is enhanced images, N c is the number of enhanced images; for the visible light, infrared and enhanced image sets, the corresponding identification label sets are defined, namely, the visible light image identification label set Infrared image identification label set and enhanced image identification label set Define image set X t Represents the visible light image set X v 、Infrared image set X r Or Enhanced Image Set X c One of them, X t ∈{X v ,X r ,X c }; Before inputting into the model, the input image needs to be patch-embedded: each image in the image set is divided into a series of patches, and the patches are mapped to a high-dimensional feature space to obtain the corresponding patch set. The formula is: p t =PatchEmbedding(X t ); Among them, p t For image set X t is a series of patch sets; PatchEmbedding(·) is a patch embedding operation; the number of images in the image set corresponds to the number of patch sets.
5. The cross-modal person re-identification method based on feature enhancement according to claim 4, characterized in that: The specific process of step 2.2 is as follows: Step 2.2.1: Preset the crossover rate and randomly select two different patch sets with the same identity based on the crossover rate. Perform attention fusion feature operation, They are p a The i-th and I-th patches in ; They are p b The jth and Jth patches in , multiply and normalize these two sets of patches to construct the attention matrix, the formula is: W=softmax(p a *p b ); Where W is the attention matrix; softmax(·) is the softmax function; Step 2.2.2, from patch set p a Randomly select a patch from Using the attention matrix, find the patch set p b Zhongyu The most similar patch Determine the current The index j and the adjacent indexes around j are valid indexes; according to the valid indexes, from p b Extract the corresponding eigenvalue from , and extract the attention score corresponding to the index from the attention matrix W, then renormalize the extracted attention score and construct a new attention matrix, specifically: W′=softmax(w j ),j∈vaild indices ; Among them, W′ is the new attention matrix; w j is the attention score at index j in the attention matrix W; vaild indices is a valid index; Step 2.2.3, according to the new attention matrix and the corresponding eigenvalue, the fused patch is obtained, the formula is: Among them, p r is the fused patch; w j ′ is the attention score at index j in the new attention matrix W′, and the attention matrix corresponds to the index number of the new attention matrix; Step 2.2.4, finally, p a In Replaced with the fused patch p r , get the updated patch set 6. The cross-modal person re-identification method based on feature enhancement according to claim 5, characterized in that: The specific process of step 2.3 is as follows: First, define a series of input patches with the same identity as the original patch set p s , calculate p s Statistics, including mean and variance; according to the calculated mean μ and variance σ 2 , defines a Gaussian distribution N(μ,σ 2 ); from p s A series of patches of one of the identities are selected to form the original mutation patch set m∈s, s is a specific identity, m is a sample in a specific identity, For p m Middle Patches, For p m Middle patches; sample from the Gaussian distribution and traverse the original mutation patch set p m For each patch in , the mutation rate is pre-set, and the individuals are replaced with points sampled from the Gaussian distribution according to the mutation rate; the calculation formulas for the mean and variance are: in, For the original patch set p s The hth patch in the set; H is the original patch set p s The number of patches in New feature samples p extracted from the Gaussian model new It conforms to the Gaussian distribution, specifically: p nee ~N(μ,σ 2 ); Finally, the samples in the original mutation patch set Replace with the new feature sample p obtained by sampling new , get the new mutation patch set 7. The cross-modal person re-identification method based on feature enhancement according to claim 6, characterized in that: The formula of step 2.4 is: Among them, F t For image set X t The predicted value characteristics of is a weight-sharing encoder.
8. The cross-modal person re-identification method based on feature enhancement according to claim 7, characterized in that: The specific process of step 3 is as follows: Step 3.1: Input the predicted value features of each image set into the classification layer to obtain the predicted probability of the image for each identity, and calculate the identity loss L based on the predicted probability. id : S t =Classification(F t ); Among them, S t For image set X t The predicted probability of Classification(·) is the classification layer; Y t Represents the image set X t The true label of represents the total number of identity samples; Represents the image set X t For The predicted probability of each identity; Step 3.2: Calculate the triplet loss L of the image set tri : Among them, F is the predicted value feature; represents a randomly selected sample image; o represents Positive sample images with the same sample identity; o′ represents Negative sample images with different sample identities; margin is the parameter boundary; Step 3.3: Calculate the discriminant center loss L of the image set dcl : in, for The center vector of For the identities; K represents the number of predictive value features of the same modality in the same identity; For visible light images The z-th predicted value feature; Infrared image The k-th predicted value feature; For The predicted value characteristics of all identities except The average distance of the center vector of z′ Indicates The predicted value features of the z′th sample with different identities; F k′ For the The predicted value features of the k′th sample with the same identity; Step 3.4: Construct the least squares error loss function L msel , randomly select a prediction value feature of an infrared image or a visible light image, denoted as Prediction value features of infrared images or Prediction value features of visible light images First calculate The average distance to other samples with the same identity in intramodality and cross-modality is calculated as: Among them, D intra is the average distance under the internal mode; D cross is the average distance under cross-modality; D(·) is the Euclidean distance; For the The predicted value features of infrared images; For the The predicted value features of visible light images; Calculate in the same way as above The average distance to other samples of the same identity in both intramodality and crossmodality; Calculate the difference L between the mean distance under intramodality and the mean distance under cross-modality msel : in, For the same identity The average distance under the internal mode corresponding to the predicted value features; For the same identity The average cross-modal distance corresponding to the predicted value features; Step 3.
5. Finally, the overall loss function L of the training process is: L=λL msel +λL dcl +L id +L tri ; Among them, λ is a hyperparameter.
9. The cross-modal person re-identification method based on feature enhancement according to claim 8, characterized in that: The specific process of step 4 is as follows: Step 4.1: Use the query set and gallery set of the test dataset as the input of the cross-modal person re-identification model trained in step 3, and concatenate the image features of the three modalities output by the model in the channel dimension to obtain the final pedestrian prediction value features; Step 4.2: Calculate the similarity between the pedestrian images in the query set and the pedestrian images in the gallery set. The similarity calculation formula is: in, Pedestrian images in the query set Pedestrian images with gallery set similarity; Pedestrian images representing the query set The predicted value characteristics of Pedestrian images representing the gallery set The predicted value features of ; |·| represents the modulus length; Step 4.3: Sort all similarity values in descending order, and output the images corresponding to the features of the top ten pedestrian prediction values with the highest similarity values as the re-identification results.
Citation Information
Patent Citations
Pedestrian rerecognition method based on distance distribution metric learning
CN109063591A
Lightweight cross-modal pedestrian re-identification method combined with data enhancement
CN115775394A
Cross-modal pedestrian re-identification method based on attitude feature alignment
CN117333908A
Unsupervised cross-modal pedestrian re-identification method and system based on hierarchical difference
CN117351518A
Semantic perception-based cross-modal pedestrian re-identification method
CN117542084A