Cross-modal Person Re-identification Method with Adaptive Mask Enhancement for Key Detail Attributes
By adopting the adaptive mask enhancement method of key detail attributes in the cross-modal pedestrian re-identification task, the problem of the model ignoring key detail attributes is solved and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202310368238.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-04-08
AI Technical Summary
The prior art is prone to ignore key details attributes in cross-modal pedestrian re-identification tasks, making it difficult to distinguish similar pedestrian images, and the model has poor ability to model key details attributes.
A method based on adaptive mask enhancement based on key detail attributes is proposed. The model is forced to focus on key detail attributes through single-modal and cross-modal significant attribute mask modules, and the modeling balance of significant attributes and key detail attributes is ensured through the attribute modeling balance module.
The model's attention to key detail attributes is improved, the problem of key detail attribute ignorance under the influence of significant attributes is alleviated, and the accuracy of cross-modal pedestrian re-identification is significantly improved.
Smart Images

Figure CN116503904B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision, information retrieval, and multimodal computing, and relates to a cross-modal pedestrian re-identification method with adaptive mask enhancement for key detail attributes, and particularly to a natural language pedestrian search method for adaptive mask enhancement of key detail attribute modeling. Background Art
[0002] Pedestrian re-identification based on natural language is an important and challenging computer vision task, which has extensive applications in the fields of security monitoring, intelligent video analysis, personnel search and rescue, etc. At present, there have been a large number of research progresses in extracting easily learnable significant attribute features and performing cross-modal alignment through significant attributes. However, due to the fact that the significant attributes of similar pedestrian images often have small differences, it is difficult to judge solely based on significant attributes, and prominent significant attributes are likely to cause the model to ignore other key detail attributes and other problems. Under the influence of significant attributes, the model has poor ability to model key detail attributes and is difficult to distinguish similar pedestrian images. In response to this, we designed a significant attribute mask method to mask easily learnable significant attributes and force the model to focus on key detail attributes. The problem brought by this forced masking operation is that the model ignores easily learnable significant attributes and only focuses on key detail attributes, which may cause unbalanced attribute modeling and affect the retrieval accuracy. In response to this, we designed a reasonable method to balance the modeling of easily learnable significant attributes and key detail attributes. Finally, the cross-modal pedestrian re-identification method based on adaptive mask enhancement of key detail attributes proposed by us can better focus on easily learnable significant attributes and key detail attributes, and the retrieval performance has been improved. Summary of the Invention
[0003] Technical Problems to be Solved
[0004] To avoid the deficiencies of the prior art, the present invention proposes a cross-modal pedestrian re-identification method with adaptive mask enhancement for key detail attributes. Aiming at the difficulty in distinguishing similar pedestrian images due to the neglect of key detail attributes in the cross-modal pedestrian re-identification task in the prior art, a method based on adaptive mask enhancement for key detail attributes is proposed for the first time, which is used to improve the model's attention to key detail attributes, and then alleviate the problem that the model is prone to ignore key detail attributes when facing easily learnable salient attributes, so as to obtain more accurate cross-modal re-identification results. First, the single-modal salient attribute mask module is used to clarify the importance of different attributes with reference to the global context semantics; then the cross-modal salient attribute mask module is proposed, and according to the fine-grained cross-modal relationship, the easily learnable salient attribute mask is found to force the model to improve the modeling ability of key detail attributes; finally, through the attribute modeling balance module, it is ensured that the modeling ability of easily learnable salient attributes and key detail attributes is balanced, and while not losing the modeling ability of easily learnable salient attributes, the attention to key detail attributes is improved, thereby enhancing the accuracy of cross-modal pedestrian re-identification.
[0005] Technical solution
[0006] A cross-modal pedestrian re-identification method based on adaptive mask enhancement for key detail attributes, characterized in that the steps are as follows:
[0007] Step 1: In the image single-modal mask branch and the text single-modal mask branch, calculate the visual feature map of the single-modal salient attribute mask and the text feature map of the single-modal salient attribute mask
[0008] On the image side:
[0009] Step a1: Normalize the training set images to a unified size, perform data augmentation on the training set images, and use a convolutional network to extract image features to obtain the initial visual feature map V;
[0010] Step a2: In the image single-modal mask branch, calculate the cosine similarity between the initial visual feature map V and the global visual feature v to obtain the single-modal visual similarity matrix S v , where the global visual feature v is extracted from the initial visual feature map V through a max pooling layer;
[0011] Step a3: Calculate the v largest values in the single-modal visual similarity matrix S:
[0012]
[0013] Where: h v , wv represent the height and width of the visual feature map V, respectively, r m represents the mask position ratio parameter;
[0014] Set all channel values corresponding to the selected maximum pixel positions of the initial visual feature map V to 0 to obtain the visual feature map of the single-modal significant attribute mask
[0015] At the text end:
[0016] Step b1: Unify the number of words in the sentences of the original training set, encode the words into word vectors using the existing word vector embedding method, and then obtain the initial text feature map T through 1×1 convolution, i.e., the text convolution layer;
[0017] Step b2: Calculate the cosine similarity between the initial text feature map T and the global text feature t to obtain the single-modal text similarity matrix S t , where the global text feature t is extracted from the initial text feature map T by the max pooling layer;
[0018] Step b3: Calculate the t maximum similarity values in the single-modal text similarity matrix S values:
[0019]
[0020] where h t , w t represent the height and width of the text feature map T, respectively, r m is the same mask position ratio parameter as that of the image single-modal mask branch;
[0021] Set all channel values corresponding to the selected maximum word positions of the initial text feature map T to 0 to obtain the text feature map of the single-modal significant attribute mask
[0022] Step 2: In the cross-modal mask branch, calculate the cosine similarity between the initial visual feature map V and the initial text feature map T, and obtain the cross-modal similarity matrix S c ;
[0023] Step 3: According to the cross-modal similarity matrix S c , respectively find the and maximum similarity values, which are the most significant visual and text attributes considered in the cross-modal search. By introducing the same mask position ratio parameter r m as that of the single-modal mask branch, obtain where, corresponding to the number of pixels in the image, corresponding to the number of words in the text;
[0024] Step 4: Mask the eigenvalues of the most significant regions, and find the position with the maximum similarity in S c which corresponds to the pixels in the image, and set the pixels in V to 0 in the entire channel to obtain the visual feature map of the cross-modal significant attribute mask
[0025] Step 5: The position with the maximum similarity in S c corresponds to the words in the text, and set the words in T to 0 in the entire channel to obtain the text feature map of the cross-modal significant attribute mask
[0026] Step 6: Adopt the attribute modeling balance module, randomly select a sample with a probability in a training batch for masking, and set the training batch random masking ratio parameter r b , and finally the number of feature maps masked in a training batch is n b :
[0027]
[0028] where: b represents the size of a training batch, represents rounding down;
[0029] Step 7: Input the feature maps trained in Step 6 into the residual network and max pooling layer in the attribute modeling balance module to obtain the masked feature vectors; there are four attribute modeling balance modules, where, and V pass through the single-modal image attribute modeling balance module to obtain the single-modal adaptive masked visual feature vector V u , and T pass through the single-modal text attribute modeling balance module to obtain the single-modal adaptive masked text feature vector T u , and V pass through the cross-modal image attribute modeling balance module to obtain the cross-modal adaptive masked visual feature vector V c , and T pass through the single-modal image attribute modeling balance module to obtain the cross-modal adaptive masked text feature vector T c ;
[0030] Step 8: Respectively for (V u , T u ), (V c , Tc ) Perform cross-modal matching and use the "Adam optimization algorithm" for training until convergence;
[0031] Step 9: During testing, use the trained network to extract features from the image and the sentence respectively without any masking operation. Obtain the image features through the image single-modal masking branch Obtain the image features through the cross-modal masking branch Obtain the text features through the text single-modal masking branch Obtain the text features through the cross-modal masking branch And the image features Are concatenated in the channel dimension to obtain the final visual feature V f , the text features Are concatenated in the channel dimension to obtain the text feature T f ;
[0032] Step 10: The visual feature V obtained in Step 9 f , the text feature T f , and then sort according to the similarity between different samples to obtain the final retrieval result sequence.
[0033] The data augmentation method in Step 1 includes but is not limited to image flipping.
[0034] The cross-modal similarity matrix S in Step 2 c The value represents the similarity between each pixel and each word.
[0035] In Step a1, the convolutional network is used to extract the image features, including but not limited to Resnet18 or Resnet50.
[0036] Beneficial effects
[0037] A cross-modal pedestrian re-identification method with adaptive mask enhancement for key detail attributes proposed by the present invention aims at the problem of natural language cross-modal pedestrian re-identification. It innovatively proposes a key detail attribute adaptive mask enhancement network. By adaptively masking the significant attributes that are easy to learn in the modal content, it improves the model's attention to key detail attributes, thereby alleviating the problem that the model is prone to ignore key detail attributes when facing significant attributes that are easy to learn. Finally, it significantly improves the accuracy of cross-modal pedestrian re-identification. The proposed method first uses a single-modal significant attribute masking module to clarify the importance of different attributes by referring to the global context semantics in the same modality. Then, a cross-modal significant attribute masking module is proposed to determine the importance of different attributes through fine-grained cross-modal relationships, that is, pixel-word correspondence. In addition, we propose an attribute modeling balance module to ensure the modeling balance between easily learned significant attributes and key detail attributes by randomly selecting image-text pairs of masked features for cross-modal alignment. The method proposed in this paper is the first to consider adaptively masking easily learned significant attributes, screening out easily learned significant attributes by referring to single-modal and cross-modal relationships, and driving the model to improve the modeling ability of key detail attributes through the masking mechanism, so as to be able to more accurately distinguish similar pedestrians, and has achieved significant improvement in retrieval accuracy in both natural language pedestrian search and image-text matching tasks.
[0038] In the present invention, aiming at the problem of cross-modal pedestrian re-identification based on natural language, a key detail attribute adaptive mask enhancement network is proposed, which attempts to forcefully mask the significant attributes that are easy to learn in the feature map, improve the model's attention to key detail attributes, and thereby alleviate the problem that the model is prone to ignore key detail attributes when facing significant attributes that are easy to learn.
[0039] In the present invention, a single-modal significant attribute masking module is proposed to mask the regions that the model focuses on within a single modality, forcing the model to complete accurate cross-modal identification through key detail attributes that are difficult to focus on.
[0040] In the present invention, a cross-modal significant attribute masking module is proposed to find and mask significant regions from the cross-modal fine-grained level of pixels and words, reducing the impact of easily learned significant attributes between the two modalities on the performance of pedestrian re-identification, and enhancing the model's modeling ability for key detail attributes.
[0041] The attribute modeling balance module proposed in the present invention is designed to ensure the modeling balance between easily learned significant attributes and key detail attributes. While the module focuses on the modeling of key detail attributes that are difficult to learn, it avoids the decline in the model's modeling ability for easily learned significant attributes, further improves the model's modeling ability for various attributes, and enhances the robustness of the model.
[0042] In the present invention, for the cross-modal person re-identification task based on natural language, the proposed key detail attribute adaptive mask enhancement method has achieved a significant improvement in retrieval accuracy compared to existing solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic framework diagram of the key detail attribute adaptive mask enhancement method;
[0044] Figure 2 is a flowchart of the algorithm; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The present invention will be further described in conjunction with embodiments and the accompanying drawings:
[0046] The key detail attribute adaptive mask enhancement method proposed by the present invention introduces: a significant attribute mask module and an attribute modeling balance module. The significant attribute mask module searches for easily learnable significant attributes from both single-modal and cross-modal directions. By masking the significant attributes, the model is forced to mine the semantic information contained in the key detail attributes. The attribute modeling balance module adaptively enhances the model's ability to model key detail attributes by selecting an appropriate mask ratio, and finally obtains a feature vector that contains both easily learnable significant attributes and key detail attribute information. Aiming at the difficulty of key detail attribute modeling in the existing technology for cross-modal person re-identification tasks based on natural language, the proposed improved method can improve the model's attention to key detail attributes, alleviate the problem of poor key detail attribute modeling ability to a certain extent, and thus improve the accuracy of cross-modal person re-identification. The accuracy has been significantly improved compared to existing methods.
[0047] Figure 2 is a flowchart of the algorithm for the key detail attribute adaptive mask enhancement method in the present invention. As shown in the figure, the present invention specifically includes the following steps:
[0048] Step S0, on the image side, normalize the training set images to a unified size, and perform data augmentation by means such as image flipping. Use a convolutional network (e.g., Resnet18, Resnet50, etc.) to extract image features and obtain an initial visual feature map V;
[0049] Step S1, since the significant regions concerned in single-modal and cross-modal are different, we introduce two mask branches: a single-modal mask and a cross-modal mask. In the single-modal mask branch of the image, to find the most easily learnable significant attributes (pixels) in the initial visual feature map V, we calculate the cosine similarity between the initial visual feature map V and the global visual feature v to obtain a single-modal visual similarity matrix S v . Among them, the global visual feature v is extracted from the initial visual feature map V by a max pooling layer;
[0050] Step S2. According to the unimodal visual similarity matrix S v , find the values with the largest similarity, which correspond to the most prominent visual pixels. To find a mask region of appropriate size, we introduce a mask position ratio parameter r m , and obtain the calculation method:
[0051]
[0052] where h v , w v respectively represent the height and width of the visual feature map V. After finding the visually salient regions that are easy to learn, mask the feature values at these positions. Set all channel values corresponding to the selected largest pixel positions of the initial visual feature map V to 0 to obtain the visual feature map of the unimodal salient attribute mask
[0053] Step S3. On the text side, unify the number of words in the sentences of the original training set, introduce an existing word vector embedding method (e.g., BERT, etc.) to encode the words into word vectors, and then obtain the initial text feature map T through a 1×1 convolutional (text convolution) layer;
[0054] Step S4. In the text unimodal mask branch, similar to the image unimodal mask branch, we calculate the cosine similarity between the initial text feature map T and the global text feature t to obtain the unimodal text similarity matrix S t . Among them, the global text feature t is extracted from the initial text feature map T through a max pooling layer;
[0055] Step S5. According to the unimodal text similarity matrix S t , find the values with the largest similarity, which correspond to the most prominent words. To find a mask region of appropriate size, we use the same mask position ratio parameter r as in the image unimodal mask branch m , and obtain the calculation method:
[0056]
[0057] where h t , w t respectively represent the height and width of the text feature map T. After finding the textually salient regions that are easy to learn, mask the feature values at these positions. Set all channel values corresponding to the selected largest word positions of the initial text feature map T to 0 to obtain the text feature map of the unimodal salient attribute mask
[0058] In step S6, in the cross-modal masking branch, we calculate the cosine similarity between the initial visual feature map V and the initial text feature map T, and obtain the cross-modal similarity matrix S c . The cross-modal similarity matrix S c represents the similarity between each pixel and each word;
[0059] In step S7, based on the cross-modal similarity matrix S c , we respectively find the and values with the greatest similarity, which are the most significant visual and text attributes considered in the cross-modal search. Similar to steps S2 and S5, by introducing the masking position ratio parameter r m , we calculate and Different from the single-modal mask, and are different from the number of positions in the similarity matrix S c . Among them, corresponds to the number of pixels in the image, corresponds to the number of words in the text;
[0060] In step S8, after finding the most significant regions, we mask the feature values at these positions. We find the positions with the greatest similarity in S c , and these positions correspond to pixels in the image. We set pixels in V to 0 in the entire channel, and obtain the visual feature map of the cross-modal significant attribute mask
[0061] In step S9, referring to step S8, the positions with the greatest similarity in S c correspond to words in the text. We set words in T to 0 in the entire channel, and can obtain the text feature map of the cross-modal significant attribute mask
[0062] In step S10, to prevent the model from possibly never seeing the easily learnable significant attributes during training, resulting in an imbalance in the modeling capabilities of the easily learnable significant attributes and the key detail attributes, we design an attribute modeling balance module. This module randomly selects samples with a certain probability for masking in a training batch, and sets the training batch random masking ratio parameter r b . Finally, the number of feature maps masked in a training batch is n b :
[0063]
[0064] where b represents the size of a training batch, denotes rounding down;
[0065] Step S11: Feed the feature map obtained in Step S10 into the residual network and the max pooling layer in the attribute modeling balance module to obtain the masked feature vector. Figure 1 The display model has a total of four different attribute modeling balance modules. Among them, and V pass through the single-modal image attribute modeling balance module to obtain the single-modal adaptive masked visual feature vector V u , and T pass through the single-modal text attribute modeling balance module to obtain the single-modal adaptive masked text feature vector T u , and V pass through the cross-modal image attribute modeling balance module to obtain the cross-modal adaptive masked visual feature vector V c , and T pass through the single-modal image attribute modeling balance module to obtain the cross-modal adaptive masked text feature vector T c ;
[0066] Step S12: Perform cross-modal matching on (V u , T u ), (V c , T c ) (for example: using a triplet loss function, etc.), and use the "Adam optimization algorithm" for training until convergence;
[0067] Step S13: During testing, use the trained network to extract features from the picture and the statement respectively without any masking operation. Obtain the image features through the image single-modal masking branch Obtain the image features through the cross-modal masking branch Obtain the text features through the text single-modal masking branch Obtain the text features through the cross-modal masking branch And splice the image features in the channel dimension to obtain the final visual feature V f , and splice the text features in the channel dimension to obtain the text feature T f ;
[0068] Step S14: Obtain the visual feature V f , the text feature T f , and then sort according to the similarity between different samples to obtain the final retrieval result sequence.
Claims
1. A cross-modal pedestrian re-identification method based on adaptive mask enhancement of key detail attributes, characterized in that The steps are as follows: Step 1: Calculate and obtain the visual feature map of the unimodal significant attribute mask and the text feature map of the unimodal significant attribute mask in the image unimodal mask branch and the text unimodal mask branch, respectively. And the text feature map of the unimodal significant attribute mask On the image side: Step a1: Normalize the training set images to a unified size, perform data augmentation on the training set images, extract image features, and obtain the initial visual feature map V; Step a2: In the image single-modal mask branch, calculate the cosine similarity between the initial visual feature map V and the global visual feature v to obtain the single-modal visual similarity matrix S v , where the global visual feature v is extracted from the initial visual feature map V through a max pooling layer; Step a3: Calculate the unimodal visual similarity matrix S v with the largest similarity values: Where: h v , w v respectively represent the height and width of the initial visual feature map V, and r m represents the proportionality parameter of the mask position; Set all channel values corresponding to the selected maximum pixel positions of the initial visual feature map V to 0 to obtain the visual feature map of the unimodal saliency attribute mask On the text side: Step b1: Unify the number of words in the original training set sentences, encode the words into word vectors using existing word vector embedding methods, and then obtain the initial text feature map T through 1×1 convolution, i.e., the text convolution layer; Step b2: Calculate the cosine similarity between the initial text feature map T and the global text feature t to obtain the unimodal text similarity matrix S t , where the global text feature t is extracted from the initial text feature map T through a max pooling layer; Step b3: Calculate the single-modal text similarity matrix S t with the largest similarity in values: where h t and w t represent the height and width of the initial text feature map T respectively, and r m is the same mask position ratio parameter as that of the image unimodal mask branch; Set all channel values corresponding to the selected maximum word positions of the initial text feature map T to 0 to obtain the text feature map of the unimodal salient attribute mask Step 2: In the cross-modal mask branch, calculate the cosine similarity between the initial visual feature map V and the initial text feature map T, and obtain the cross-modal similarity matrix S c ; Step 3: According to the cross-modal similarity matrix S c , respectively find the and values with the largest similarity, which are the most significant visual and text attributes considered in cross-modal search. By introducing the same mask position ratio parameter r m as that of the image unimodal mask branch, we get where corresponds to the number of pixels in the image, corresponds to the number of words in the text; Step 4: Mask the eigenvalues of the most significant regions and find the position with the maximum similarity in c . This position corresponds to c a pixel in the image. Set pixels in V to 0 in the entire channel to obtain the visual feature map of the cross-modal significant attribute mask Step 5: Correlate the position with the greatest similarity in S c to the words in the text, and set the words in T to 0 across the entire channel to obtain the text feature map of the cross-modal salient attribute mask Step 6: Use the attribute modeling balance module to randomly select a sample with a probability in a training batch for masking, and set the training batch random masking ratio parameter r b , and finally the number of feature maps masked in a training batch is n b : where: b represents the size of a training batch, denotes rounding down; Step 7: Input the feature maps trained in Step 6 into the residual network and the max pooling layer in the attribute modeling balance module to obtain the masked feature vectors; there are four attribute modeling balance modules, where and V pass through the single-modal image attribute modeling balance module to obtain the single-modal adaptive masked visual feature vector V u , and T pass through the single-modal text attribute modeling balance module to obtain the single-modal adaptive masked text feature vector T u , and V pass through the cross-modal image attribute modeling balance module to obtain the cross-modal adaptive masked visual feature vector V c , and T pass through the cross-modal text attribute modeling balance module to obtain the cross-modal adaptive masked text feature vector T c ; Step 8: Perform cross-modal matching on (V u , T u ), (V c , T c ), and use the "Adam optimization algorithm" for training until convergence; Step 9: During testing, the image and the statement are respectively subjected to feature extraction using the trained network without any masking operation, and the image features are obtained through the image single-modal masking branch The image features are obtained through the cross-modal masking branch The text features are obtained through the text single-modal masking branch The text features are obtained through the cross-modal masking branch And the image features are concatenated in the channel dimension to obtain the final visual feature V f , the text features are concatenated in the channel dimension to obtain the text feature T f ; Step 10: The visual feature V is obtained in Step 9 f , and the text feature T f . Then, they are sorted according to the similarity between different samples to obtain the final retrieval result sequence.
2. The cross-modal pedestrian re-identification method based on adaptive mask enhancement of key detail attributes according to claim 1, characterized in that: The data augmentation method in step 1 includes but is not limited to image flipping.
3. The cross-modal pedestrian re-identification method based on adaptive mask enhancement of key detail attributes according to claim 1, characterized in that: The cross-modal similarity matrix S in step 2 c represents the similarity between each pixel and each word.
4. The cross-modal pedestrian re-identification method based on adaptive mask enhancement of key detail attributes according to claim 1, characterized in that: In step a1, the image features are extracted using a convolutional network, including but not limited to Resnet18 or Resnet50.
5. The cross-modal pedestrian re-identification method based on adaptive mask enhancement according to claim 1, characterized in that: The word vector embedding method in step b1 includes but is not limited to BERT.
Citation Information
Patent Citations
Cross-modal image-text matching method and device and computer readable storage medium
CN112905827A
Text pedestrian re-recognition algorithm based on cross-modal correlation graph inference method
CN115116096A