Reloading pedestrian re-identification method based on gradient inversion and identity attention focusing
By using gradient inversion and identity attention-focusing methods in re-identification of dress-changing pedestrians, the clothing features of pedestrian images are extracted and decoupled, and the identification inaccuracy problem caused by changes in pedestrian clothing in the prior art is solved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202510332313.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
The existing pedestrian re-identification method for changing clothes is difficult to effectively constrain the model, resulting in inaccurate identification of pedestrian clothing changes, especially in the case of long-term spans and perspective changes.
Using a method based on gradient inversion and identity attention aggregation, the global and local features of pedestrian images are extracted through a pre-trained backbone network, combining the mask cross attention module and gradient inversion layer to decouple clothing features and aggregate identity features.
It effectively reduces the model's attention to clothing and background, improves attention to pedestrian identity areas, and improves the accuracy of pedestrian re-identification of dress-up pedestrians.
Smart Images

Figure CN120183000A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of digital image processing and pattern recognition, and in particular to a method for re-identifying a pedestrian in a changed outfit based on gradient inversion and identity attention aggregation. Background Art
[0002] Person re-identification is an important automated pedestrian retrieval technology in the field of computer vision. It supports pedestrian tracking across time and location under different perspectives and cameras. Person re-identification with changing clothes is a special scenario of pedestrian re-identification. In the case of a long span, pedestrians will cause significant changes in discriminable features due to changing clothes. Therefore, person re-identification with changing clothes is a very challenging task. Person re-identification with changing clothes has made some progress in recent years, but the existing methods lack effective constraints, making it difficult for the model to always focus on the important discriminative areas of pedestrians, and are still largely affected by clothing changes. This technology can be widely used in fields such as intelligent monitoring, criminal investigation and public security management, especially in the field of criminal investigation, it can be used to identify the suspect's clothing disguise. In order to simulate real scenes, the data samples used for person re-identification with changing clothes contain occlusion, clothing changes, image blur and perspective changes. In addition, the current person re-identification method with changing clothes based on feature decoupling cannot keep the focus on the pedestrian's identity area while decoupling clothing and identity features. Summary of the invention
[0003] In order to solve the problem of inaccurate recognition when pedestrians change clothes over a long period of time, when pedestrian images are blurred, and when the viewing angle changes, the present invention proposes a pedestrian re-identification method based on gradient inversion and identity attention aggregation. A pre-trained backbone network is used to extract global features from pedestrian images, and the similarity between the global features and the features in the retrieval library is calculated. The images corresponding to the features whose similarity is greater than a set threshold are pushed to the user. The training process of the backbone network specifically includes the following steps:
[0004] Obtain pedestrian images and use the backbone network to extract global features and patch-based local features from them;
[0005] The local features are input into the mask predictor to obtain the predicted mask, the pedestrian image is segmented through the human body parsing operation, and then the masks of each part are obtained through the mask encoder, and the cross entropy loss between the mask and the predicted mask is calculated;
[0006] The key vector and value vector of the mask cross attention module are constructed according to the global features. The query vector of the mask cross attention module is constructed according to the local features and the randomly sampled learnable embedding vector. The mask predictor obtains the predicted mask as the mask to obtain the fused feature vector.
[0007] Extract the attention aggregation loss by separating the identity feature and the local feature from the fused feature vector;
[0008] Construct a borderless triplet loss function based on the local feature, batch-normalize the local feature and then input it into the identity classifier, and calculate the cross-entropy loss according to the result output by the classifier;
[0009] Input the batch-normalized local feature into the gradient reversal layer, process it based on the clothing attention score, then input the feature obtained from the gradient reversal layer into the clothing classifier for classification, and calculate the cross-entropy loss according to the classification result;
[0010] Train the backbone network according to the triplet loss function of the local feature, the cross-entropy losses of the clothing classifier and the identity classifier, the cross-entropy loss between the mask and the predicted mask, and the attention aggregation loss.
[0011] Furthermore, the process of obtaining the fused feature of the i-th sample includes:
[0012] Concatenate the global feature of the i-th sample with the randomly sampled embedding vector as the input of the multi-head self-attention module for aggregation to obtain the aggregation result Q;
[0013] Use the aggregation result Q as the query vector in the masked cross-attention module, the global feature of the i-th sample as the key vector and value vector, and the predicted mask as the mask, and the masked cross-attention module outputs the fused feature.
[0014] Furthermore, the process of the masked cross-attention module processing the input feature includes:
[0015] F = FFN(LN(X))
[0016] X c = Softmax(H c + Q c K T )V
[0017]
[0018] where F is the fused feature output by the masked cross-attention module; X is the set of all category local features, X c represents the local feature of category c; FFN represents the feed forward network operation, LN represents the layer normalization operation; Q c represents the feature embedding of category c; H c is the feature obtained by the masking operation; H c (i) represents the masking information of whether the i-th image patch belongs to attribute c; M pre(c, i) represents the mask indicating that the i-th image patch belongs to attribute c.
[0019] Further, the gradient reversal layer multiplies the gradient backpropagated by the clothing classifier by -λ and passes it to the backbone network, and this process is expressed as:
[0020]
[0021] where GRL(·) represents the gradient reversal layer; represents taking the partial derivative; θ e represents the network parameters of the backbone network δ e of, θ c represents the network parameters of the clothing classifier δ c of; λ is a training parameter that gradually increases from 0.1 to 1 during the training process.
[0022] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0023] 1. The present invention designs a gradient reversal module based on clothing attention scores. In the TransReID model adopted in the present invention, there are 12 transformer blocks. The present invention uses the predicted clothing mask to extract the attention scores of the model for the clothing region from the last transformer block, normalizes it, and uses it as part of the reversal weight, enabling the model to stably decouple clothing features and non-clothing features only through simple gradient reversal;
[0024] 2. The present invention designs an attention aggregation module, which uses the predicted identity mask and the often-neglected local features to obtain pedestrian features related to the identity region. At the same time, an attention aggregation loss function is proposed. By narrowing the similarity distance between the global features and the identity features, it helps the model understand the identity semantic information of pedestrians. While reducing the model's attention to clothing and background, the present invention aggregates the attention to the pedestrian's identity-related regions, such as discriminative regions like the head and arms. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is the training flow chart of the backbone network in a clothing-changing pedestrian re-identification method based on gradient reversal and identity attention aggregation according to the present invention;
[0026] Figure 2 is the schematic structural diagram of the network model of a clothing-changing pedestrian re-identification method based on gradient reversal and identity attention aggregation according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] The present invention proposes a clothing-changing pedestrian re-identification method based on gradient reversal and identity attention aggregation. The global features are extracted from pedestrian images by using a pre-trained backbone network, the similarity between the global features and the features in the retrieval library is calculated, and the images corresponding to the features with similarity greater than the set threshold are pushed to the user. For example Figure 1 , the training process of the backbone network specifically includes the following steps:
[0029] Obtain pedestrian images, and use the backbone network to extract global features and local features based on patch blocks from them;
[0030] Input the local features into the mask predictor to obtain the predicted mask, segment the pedestrian image through human parsing operation, then obtain the masks of each part through the mask encoder, and calculate the cross-entropy loss between the mask and the predicted mask;
[0031] Construct the key vector and value vector of the mask cross-attention module according to the global features, construct the query vector of the mask cross-attention module by splicing the local features with the randomly sampled learnable embedding vector, and use the predicted mask obtained by the mask predictor as the mask to obtain the fused feature vector;
[0032] Separate the identity features from the fused feature vector and extract the attention aggregation loss from the local features;
[0033] Construct a borderless triplet loss function according to the local features, input the local features into the identity classifier after batch normalization, and calculate the cross-entropy loss according to the output result of the classifier;
[0034] Input the batch-normalized local features into the gradient reversal layer, process them based on the clothing attention score, then input the features obtained by the gradient reversal layer into the clothing classifier for classification, and calculate the cross-entropy loss according to the classification result;
[0035] Train the backbone network according to the triplet loss function of the local features, the cross-entropy losses of the clothing classifier and the identity classifier, the cross-entropy loss between the mask and the predicted mask, and the attention aggregation loss.
[0036] In this embodiment, the implementation manner of a clothing-changing pedestrian re-identification method based on gradient reversal and identity attention aggregation is divided into 7 stages for description respectively. For example Figure 2 , specifically including the following steps:
[0037] S1. Data preparation stage: Use the pre-trained human parsing model SCHP to parse the pedestrian image and obtain the segmented human semantic image; then perform data enhancement on the image and adjust the size of the pedestrian image and semantic image to 384×128×3, where 384×128×3 corresponds to the height, width and number of channels of the image respectively; finally, use the semantic image to generate a clothing mask to adjust the RGB channels of the clothing area in the pedestrian image, and use a random combination method to enrich the clothing color;
[0038] S2, feature extraction stage: the present invention uses TransReID as the backbone network, which extracts the global and local features of pedestrians. The global feature refers to the feature of all image blocks in the image, including the overall attributes of the pedestrian object, in which identity, clothing and background features are redundant together; the local feature extracts the features of each image block, corresponding to the local attributes in the image, and splices the features of the image blocks together as the local features of the entire image;
[0039] S3, mask prediction stage: the network model can predict the category corresponding to each image using local features and generate a category mask, where the category includes at least background, identity and clothing;
[0040] S4, feature decoupling stage: The feature decoupling module in the network model will pass the features extracted by the backbone network to the gradient reversal layer, and then to the clothing classifier to classify the clothing worn by pedestrians. In the back propagation stage of training, the clothing classifier parameters are forward optimized by calculating the clothing classification loss. Due to the existence of the gradient reversal layer, the backbone network tends to learn clothing-irrelevant features to achieve the effect of deceiving the clothing classifier, and finally achieve the decoupling of clothing features and non-clothing features; at the same time, the model's attention score for the clothing area is extracted using the predicted mask, and it is applied to the gradient reversal to help the model stabilize training;
[0041] S5, identity attention aggregation stage: using the predicted category mask and the local features extracted by the backbone network, the mask cross attention module can obtain three features related only to the background area, identity area and clothing area; in the feature decoupling stage, the network model can well decouple clothing features and non-clothing features, but its attention is in a divergent state. By using the attention aggregation loss to shorten the cosine similarity distance between the decoupled features of the backbone network and the features related only to the identity area, the decoupled features can be closer to the identity features in the feature space, achieving the purpose of focusing attention on the identity area;
[0042] S6. Model training stage: The global features extracted by the backbone network are input into the triplet loss function for metric learning, and the global features are classified for identity and clothing respectively and then input into the identity classification loss function and the clothing classification loss function for classification learning; for the mask prediction module, the predicted mask is input into the cross-entropy loss function for classification learning; for the attention aggregation module, the global features and the identity features are input into the attention aggregation loss function for metric learning.
[0043] S7. Model evaluation stage: The feature vectors of the target pedestrian and the pedestrians in the retrieval library are extracted using the backbone network, sorted by similarity, and finally the classification results of cross-dressing pedestrian re-identification are obtained.
[0044] As an alternative implementation, this embodiment gives the specific implementation of the data preparation stage, which specifically includes the following steps:
[0045] Use the human parsing model SCHP to obtain the human semantic image x of the pedestrian image x sc , and the parsing formula is:
[0046] x sc = SCHP(x)
[0047] where SCHP(·) represents the human parsing model.
[0048] Then, the pedestrian image and the human semantic image are uniformly resized to 384×128×3;
[0049] At the same time, use the semantic image to adjust the RGB channels of the clothing area to enrich the clothing colors;
[0050] To avoid destroying the semantic information in the semantic image, mode pooling is used to classify the semantics of image patches to obtain the class mask M. Mode pooling will select the value that appears most frequently within the pooling window.
[0051] As an alternative implementation, this embodiment gives the specific implementation of the feature extraction stage, which specifically includes: inputting the pedestrian image X into the backbone network TransReID of the network model for feature extraction to obtain the global feature F g and the local feature F based on patch blocks p , the global feature F g has a size of 1×768, and the local feature F based on patch blocks p has a size of 192×768, where 192 is the number of image patches.
[0052] As an alternative implementation, this embodiment gives the specific implementation of the mask prediction stage. In the mask prediction module, the local feature F based on patch blocks pObtain the predicted class mask M in the input mask predictor pre , the mask predictor is processed sequentially through a batch normalization layer (Batch Normalization), a linear layer (Linear), and an activation function layer σ (Softmax activation function layer), expressed as:
[0053] M pre = σ(Linear(BN(F p )))
[0054] Among them, the mask prediction result obtained by the mask predictor has a size of 192×3, where one 192-dimensional vector is an identity feature vector, one 192-dimensional vector is a clothing feature vector, and one 192-dimensional vector is a background feature vector.
[0055] In the present invention, for decoupling clothing and non-clothing features, gradient reversal is a very effective method, and a good decoupling effect can be achieved through a simple gradient reversal operation. During the forward propagation of the TransReID model, the gradient reversal layer does not change the input data and behaves as an identity mapping; but in the backward propagation stage of training, by minimizing the clothing classification loss L c perform forward optimization on the parameters of the clothing classifier, and the calculated gradient will point in the direction conducive to extracting clothing features; the gradient reversal layer will multiply the gradient backpropagated by the clothing classifier by -λ and pass it to the feature extraction network.
[0056] As an optional implementation manner, this embodiment gives the specific implementation manner of the mask prediction stage, which specifically includes the following steps:
[0057] Perform batch normalization on the global feature F g and then pass it into the gradient reversal layer, and then pass it into the clothing classifier L c , to obtain the clothing classification result f c , and the classification formula is as follows:
[0058] f c = L c (BatchNorm(F g ))
[0059] During the forward propagation process, the gradient reversal layer GRL does not change the input data and behaves as an identity mapping, as follows:
[0060] F g = GRL(F g )
[0061]
[0062] However, during backpropagation, the sign of the backpropagated gradient is reversed and then multiplied by α to scale the gradient. The gradient reversal formula is as follows:
[0063]
[0064] Since the clothing classifier is optimized in the forward direction, the calculated gradient will point in the direction that is conducive to extracting clothing features. However, the backbone network receives the reversed gradient. Therefore, the backbone network will be more inclined to extract non-clothing features, thus achieving the purpose of decoupling clothing features and non-clothing features.
[0065] In this embodiment, the value of λ determines the degree to which the gradient is amplified or reduced during backpropagation and is used to adjust the balance between the feature extraction network and the clothing classifier. In the domain adaptation task, the classifier only classifies the source domain and the target domain, and the complexity of the domain classification task is relatively low. Therefore, λ in DANN is set to gradually increase from 0.1 to 1 as the training process progresses. However, due to the higher complexity of the clothing classification task, which requires distinguishing many different clothing categories, gradually increasing λ as in the domain adaptation task is not sufficient to effectively reduce the influence of clothing information. Therefore, in the task of re-identifying pedestrians with changed clothes, an appropriate base value needs to be set for λ according to different classification complexities, so as to exert sufficient reverse optimization pressure on the feature extractor in the early stage of training to suppress the influence of clothing features and ensure that the model can better learn feature representations that are independent of identity. As the training progresses, the model's attention to clothing will gradually change. Therefore, as a preferred implementation method, the present invention believes that λ should be combined with the model's attention to the clothing region to form a dynamic weight λ d , and the dynamic weight is expressed as:
[0066]
[0067] λ d = λ base + α c
[0068] where B is the batch size, N is the number of image patches, is the attention weight between the cls token and all other patches in the last transformer block of the TransReID model for the j-th image patch of the i-th sample; represents the prediction mask that the j-th image patch of the i-th sample belongs to the category c. When the model's attention to the clothing region is too large, will also increase accordingly, and the gradient reversal layer will enhance the suppression intensity of clothing features. This can not only reduce the model's attention to clothing but also stabilize the training of the model.
[0069] As an alternative implementation, this embodiment provides a specific implementation of the identity attention aggregation stage, which specifically includes the following steps:
[0070] In the attention aggregation module, the local feature F extracted by the backbone network p can be used to further extract features related only to the identity region, clothing region, and background region respectively, that is, a learnable embedding vector q is obtained through random sampling embbed =[q id ,q clothe ,q bg , with a size of 3×768, where q id is the identity embedding, q clothe is the clothing embedding, and q bg is the background embedding; using the multi-head self-attention mechanism, q embbed can learn the global features in F g to help q embbed better aggregate image features;
[0071] After concatenating q embbed and F g and passing them into the multi-head self-attention module, the aggregated result Q is obtained, which is expressed as:
[0072] q = MHSA(concat(q embbed ,F g ))
[0073] where the obtained aggregated result Q has a size of 3×768; concat(·) represents the concatenation operation; MHSA(·) represents the multi-head self-attention layer;
[0074] Taking the class mask M pre obtained by the mask prediction module as the mask, with a size of 3×192; the aggregated result Q as the Query; the local feature F extracted by the backbone network p as the Key and Value, with a size of 192×768, and passing them into the Mask Cross Attention module to extract features, which is expressed as:
[0075] feat = MCA(Q,F p ,F p ,M pre )
[0076] where feat represents the output of the mask cross-attention module; the Mask Cross Attention module adjusts the calculated attention scores according to the M pre mask, using the classes of each patch block, through the following formula:
[0077] Xc = Softmax(H c + Q c K T )V
[0078] where X c represents the local feature of class c, Q c represents the feature embedding of class c, and H c is obtained by the following formula:
[0079]
[0080] F = FFN(LN(X))
[0081] where FFN represents the feed forward network operation; LN represents the layer normalization operation (LayerNormalization); H c (i) represents the mask information of the i-th patch in the sample with class c, and this value is the mask M pre (c, i) of the i-th patch with class c. For example, in the present invention, there are two classes. When c = 0, it represents the identity class, and when c = 1, it represents the clothing class. If the current image patch includes clothing information, then M pre (1, i) = 1, otherwise M pre (c, i) = 0; the size of the feat extracted by the MaskCross Attention module is 3×768, and the identity feature feat id = feat[0], with a size of 1×768.
[0082] As an alternative implementation, this embodiment gives the specific implementation in the model training stage, which specifically includes the following steps:
[0083] Input the feature F extracted by the backbone network g into Loss 1 for feature metric learning. The Loss 1 is calculated using an unbounded triplet loss function, and its formula is as follows:
[0084]
[0085] where B represents the Batch size of the data during training; f ori is the feature extracted from the original image, represents the feature extracted from the positive sample, represents the feature extracted from the negative sample;
[0086] Then, after batch normalization, F g is respectively input into the identity classifier L id and the clothing classifier L cGet the identity classification result f id And clothing classification results f c , and then pass it into loss 2 for classification learning. The loss 2 is calculated using the cross entropy loss function, and its formula is as follows:
[0087]
[0088] in, and Indicates that the i-th sample predicts the true label under the identity feature and clothing feature and probability;
[0089] Prediction mask M pre Loss 3 is used for constraint, where Loss 3 is a cross entropy loss function and uses the same calculation formula as Loss 2;
[0090] Finally, since the feature decoupling module cannot accurately focus on identity-related areas while decoupling features, the network model uses the calculated identity features to help the backbone network focus attention and adopts the attention aggregation loss function to bring F g and feat id The cosine similarity distance between them is as follows:
[0091]
[0092] Among them, L focused attntion represents the attention aggregation loss; B represents the batch size of the data during training; represents the local features of the i-th sample, Represents the eigenvector Model; Represents the identity feature separated from the fusion feature of the i-th sample.
[0093] A pedestrian re-identification method based on gradient inversion and identity attention aggregation is constructed. Figure 1The network model structure shown is tested on the PRCC, LTCC, and VC-Clothes datasets. The Cumulative Matching Characteristic (CMC) and mean Average Precision (mAP), which are commonly used evaluation metrics in the field of pedestrian re-identification research, are adopted to evaluate the effectiveness of the method on the dataset of re-identifying pedestrians with changed clothes. At the same time, in Table 1, Method 1 that decouples clothing and identity features by using the theory of causal intervention, Method 2 that enhances the robustness of the model by expanding the data of changed clothes in the feature space, Method 3 that uses a three-branch network to enhance head attention and reduce clothing attention through a human parsing model, and Method 4 of the present invention are compared.
[0094] The following table presents the results of the tests on the database. It can be seen that based on the network model of the present invention, excellent performance is achieved on each dataset in terms of the Rank1 and mAP metrics.
[0095] Table 1
[0096]
[0097] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for re-identifying pedestrians in costume based on gradient inversion and identity attention aggregation. The pre-trained TransReID model is used to identify the input pedestrian image. The feature extraction network of the TransReID model extracts global features from the pedestrian image, calculates the similarity between the global features and the features in the retrieval library, and pushes the images corresponding to the features whose similarity is greater than the set threshold to the user. The characteristics are: The training process of the TransReID model specifically includes the following steps: Obtain pedestrian images and use the feature extraction network of the TransReID model to extract global features and patch-based local features from them; The local features are input into the mask predictor to obtain the predicted mask, the pedestrian image is segmented through the human body parsing operation, and then the masks of each part are obtained through the mask encoder, and the cross entropy loss between the mask and the predicted mask is calculated; The key vector and value vector of the mask cross attention module are constructed according to the global features. The query vector of the mask cross attention module is constructed according to the local features and the randomly sampled learnable embedding vector. The mask predictor obtains the predicted mask as the mask to obtain the fused feature vector. The attention aggregation loss is obtained by separating the identity feature and the local feature from the fused feature vector; An unbounded triplet loss function is constructed based on local features. The local features are batch-normalized and then input into the identity classifier. The cross entropy loss is calculated based on the output of the classifier. The local features after batch normalization are input into the gradient reversal layer and processed based on the clothing attention score. Then, the features obtained by the gradient reversal layer are input into the clothing classifier for classification, and the cross entropy loss is calculated based on the classification results. The backbone network is trained according to the triplet loss function of local features, the cross entropy loss of clothing classifier and identity classifier, the cross entropy loss of mask and predicted mask, the diversity loss, and the attention aggregation loss.
2. According to claim 1, a method for pedestrian re-identification based on gradient inversion and identity attention aggregation is characterized in that: The global feature F of the pedestrian image p Input the predicted category mask M into the mask predictor pre , the process is expressed as: M pre =σ(Linear(BN(F p ))) Among them, BatchNorm(·) represents the batch normalization layer; Linear(·) represents the linear layer; σ(·) represents the Softmax activation function.
3. According to claim 1, the method for pedestrian re-identification based on gradient inversion and identity attention aggregation is characterized in that: The calculation formula of the unbounded triplet loss function is expressed as: Among them, L triplet represents the unbounded triplet loss function; B represents the batch size of the data during training; f ori (i) is the feature extracted from the i-th original image sample; represents the features extracted from the i-th positive sample, represents the features extracted from the i-th negative sample; Represents the distance between the features extracted from the i-th original image sample and the i-th positive sample; Represents the distance between the features extracted from the ith original image sample and the ith negative sample.
4. According to claim 1, the method for pedestrian re-identification based on gradient inversion and identity attention aggregation is characterized in that: The calculation formula of cross entropy loss is expressed as: Among them, L cls represents the cross entropy loss; B represents the batch size of the data during training; Indicates that the i-th sample has the identity feature Predict the true identity label Probability; Indicates that the i-th sample predicts the true clothing label under the clothing feature probability.
5. According to claim 1, the method for pedestrian re-identification based on gradient inversion and identity attention aggregation is characterized in that: The calculation formula of attention aggregation loss is expressed as: Among them, L focusedattntion represents the attention aggregation loss; B represents the batch size of the data during training; represents the local features of the i-th sample, Represents the eigenvector Model; Represents the identity feature separated from the fusion feature of the i-th sample.
6. The method for pedestrian re-identification based on gradient inversion and identity attention aggregation according to claim 5, characterized in that: The process of obtaining the fusion features of the i-th sample includes: The global feature of the i-th sample is concatenated with the embedding vector obtained by random sampling as the input of the multi-head self-attention module to obtain the aggregation result Q; The aggregation result Q is used as the query vector in the masked cross-attention module, the global features of the i-th sample are used as the key vector and value vector, and the predicted mask is used as the mask. The masked cross-attention module outputs the fused features.
7. The method for pedestrian re-identification based on gradient inversion and identity attention aggregation according to claim 6, characterized in that: The process of processing the input features by the masked cross-attention module includes: F=FFN(LN(X)) X c =Softmax(H c +Q c K T )V Among them, F is the fusion feature output by the mask cross attention module; X is the set of local features of all categories, X c represents the local feature of category c; FFN represents the feed forward network operation, LN represents the layer normalization operation; Q c represents the feature embedding of category c; H c is the feature obtained by mask operation; H c (i) indicates whether the i-th image block belongs to attribute c; M pre (c,i) represents the mask that the i-th image patch belongs to attribute c.
8. The method for pedestrian re-identification based on gradient inversion and identity attention aggregation according to claim 1, characterized in that: The gradient reversal layer multiplies the gradient returned by the clothing classifier by -λ and passes it to the backbone network. The process is expressed as: Where GRL(·) represents the gradient reversal layer; represents partial derivative; θ e Represents the backbone network δ e The network parameters, θ c represents the clothing classifier δ c The network parameters; λ is the training parameter, which gradually increases from 0.1 to 1 during the training process.
9. The method for pedestrian re-identification based on gradient inversion and identity attention aggregation according to claim 8, characterized in that: The training parameter λ is combined with the backbone network's attention to the clothing area to form a dynamic weight. The backbone network's attention to the clothing area is expressed as: Among them, B represents the batch size of the data during training; N is the number of image blocks; α c Indicates the backbone network's attention to the clothing area; is the attention weight of the cls token and all other patches in the last transformer block of the i-th sample in the TransReID model; Represents the predicted mask of the jth image block category c of the i-th sample.
10. The method for re-identifying pedestrians in costume based on gradient inversion and identity attention aggregation according to claim 1, characterized in that: The diversity loss is expressed as: Among them, L div represents diversity loss; N is the number of image blocks, Sim represents the cosine similarity calculation function; F p [i] represents the local features of the i-th image block; F p [j] represents the local features of the j-th image block.