A pose-guided attention-based person re-identification method with occlusion resistance
Through the anti-occlusion pedestrian re-identification method of pose-guided attention, the joint node heat map and global feature map are extracted, and feature matching is performed in combination with local and spatial attention feature maps, which solves the feature mismatch problem of pedestrian re-identification in the case of occlusion and improves the recognition accuracy.
Patent Information
- Application Number
- CN202210645693.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-08
AI Technical Summary
Existing pedestrian re-identification technology is difficult to effectively identify different human parts of pedestrians under occlusion, resulting in feature mismatch and reduced accuracy.
The anti-occlusion pedestrian re-identification method of pose-guided attention is adopted. The joint node heat map and global feature map are extracted through the pose estimation network, combined with the joint node local feature map and spatial attention feature map, and the attitude guide spatial attention module and channel attention module are used to generate spatial and channel attention feature maps, and feature matching is performed to reduce the impact of occlusion.
The accuracy of occluding pedestrian re-identification is improved, and the problem of singleness and feature mismatch of simply using global and local features is overcome.
Smart Images

Figure CN115527233B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of pedestrian re-identification, and in particular, relates to an anti-occlusion pedestrian re-identification method with posture-guided attention. Background Art
[0002] Person re-identification is a sub-problem of image retrieval. It is a technology that re-identifies a person from images captured by different cameras from different perspectives given an image of a specific person captured by a camera. In recent years, with the continuous development of artificial intelligence technology, existing video surveillance systems have also developed in the direction of intelligence, and intelligent systems have been widely used in people's work and life. Person re-identification is one of the important research directions that is closely related to the location and identity of pedestrians. It can help to quickly lock the target in huge video data. Among them, in the occluded pedestrian re-identification, the pedestrian to be retrieved will be occluded by other pedestrians or site occlusions. Most pedestrians only leave a half-body image. Compared with the full-body image, some body parts are occluded and the obtained pedestrian boundary box has noise information of the occlusion. Therefore, it is more practical than ordinary pedestrian re-identification.
[0003] Pedestrian re-identification is to distinguish and identify pedestrians based on their appearance features such as body shape and jewelry. In actual scenes, pedestrians may be affected by other pedestrians or other objects, which may make some body parts invisible and cause the network to learn the noise of occluders, making some pedestrian features inconsistent. In addition, in occluded pedestrian re-identification, the number of images in the pedestrian image dataset is small, and the network cannot learn sufficiently robust features. Therefore, how to extract a more robust feature still has strong theoretical and application value.
[0004] At present, with the continuous maturity and development of various neural network models and computing resources, deep learning has achieved relatively good performance in images. Convolutional neural networks have achieved good accuracy in various upstream and downstream tasks of computer vision. Therefore, convolutional neural networks are generally used for pedestrian re-identification. In pedestrian re-identification based on deep learning, the global features of pedestrians are mainly used. The pedestrian images are passed through the convolutional neural network to extract the global pedestrian feature vector, and each feature vector is sorted by comparing the similarity to finally obtain the sorting result. In the task of occluded pedestrian re-identification, since the queried pedestrian images are all occluded images, and there are unoccluded images and occluded images in the gallery, the queried pedestrian image feature vector does not match the pedestrian image feature vector in the gallery. For the pedestrian re-identification network, it is difficult to identify different human body parts of pedestrians for matching, which increases the difficulty of alignment. Therefore, it is necessary to learn the unoccluded part features of pedestrians in the feature learning stage and to match features in the similarity measurement stage to reduce the impact of occlusions on the final result. Both single local features and global features have certain limitations, which will affect the improvement of the final accuracy of pedestrian re-identification. Summary of the invention
[0005] The purpose of this application is to provide an anti-occlusion pedestrian re-identification method with posture-guided attention, which overcomes the singleness problem of simply using global features and local features, as well as the problem of mismatch between features.
[0006] In order to achieve the above purpose, the technical solution of this application is as follows:
[0007] A posture-guided attention-based anti-occlusion pedestrian re-identification method, comprising:
[0008] Obtain pedestrian images, extract joint point heat maps and joint point confidences through the posture estimation network, and extract global feature maps through the global feature extraction network;
[0009] The global feature map is element-wise multiplied with the joint point heat map to obtain the joint point local feature map. The joint point local feature map and the global feature map are then guided by the posture spatial attention module to obtain the spatial attention feature map.
[0010] The joint local feature map and the spatial attention feature map are passed through the posture-guided channel attention module to obtain the channel attention feature map;
[0011] When the joint point confidence is greater than or equal to the preset threshold, the joint point confidence is used as the joint point feature weight, otherwise the joint point feature weight is reset to zero;
[0012] The similarity distance between the pedestrian image to be retrieved and the library image is calculated based on the global feature map, joint point local feature map and channel attention feature map of the pedestrian image to be retrieved, and the most similar library image is selected as the final recognition result. The similarity distance calculation formula is as follows:
[0013]
[0014] Among them, d total Represents the similarity distance between the pedestrian image to be retrieved and the gallery image, represents the joint feature weights of the i-th joint in the pedestrian image to be retrieved and the gallery image, respectively, and b global Represents the distance between the global features of the pedestrian image to be retrieved and the gallery image, d pga represents the distance between the channel attention features of the pedestrian image to be retrieved and the gallery image, d i represents the distance between the i-th joint point between the pedestrian image to be retrieved and the gallery image, and N represents the number of joint points.
[0015] Furthermore, the joint point confidence is the maximum value in the corresponding joint point heat map.
[0016] Furthermore, the method of obtaining a spatial attention feature map by guiding the joint point local feature map and the global feature map through a posture-guided spatial attention module includes:
[0017] Perform the maximum pooling and ReLU activation function operations on the pixels in each channel of the joint local feature map to generate the joint spatial attention map;
[0018] After the global feature map is convolved, a global average pooling operation is performed on the channel dimension to generate a global spatial attention map. The joint point spatial attention map is concatenated with the global spatial attention map to obtain a spatial posture guided feature map.
[0019] The spliced spatial posture guided feature map is convolved and subjected to the Sigmoid function to obtain the spatial attention mask.
[0020] Multiplying the spatial attention mask with the global feature obtains the spatial attention feature map.
[0021] Furthermore, the method of obtaining a channel attention feature map by guiding the channel attention module through the posture-guided channel attention module includes:
[0022] Perform the maximum pooling and ReLU activation function operations on each pixel in the joint point local feature map to generate the joint point channel attention map;
[0023] After convolution operation on the spatial attention feature map, global average pooling operation is performed in the spatial dimension to generate a global channel attention map. The joint channel attention map is concatenated with the global channel attention map to obtain the channel posture guidance feature map.
[0024] The spliced channel posture guide feature map is convolved and subjected to the sigmoid function to obtain the channel attention mask.
[0025] Multiply the channel attention mask with the global feature to get the channel attention feature map.
[0026] Furthermore, the distance of global features between the pedestrian image to be retrieved and the gallery image, the distance of channel attention features between the pedestrian image to be retrieved and the gallery image, and the distance of the i-th joint point between the pedestrian image to be retrieved and the gallery image are all measured using cosine distance.
[0027] Furthermore, the posture estimation network adopts the high-resolution network HRNet-W32, and the global feature extraction network adopts ResNet-50.
[0028] The present application proposes a posture-guided attention-resistant pedestrian re-identification method, which uses joint point feature maps as attention to guide feature generation and performs feature matching during testing. It can overcome the singleness problem of simply using global features and local features, as well as the problem of mismatch between features, and improve the accuracy of occluded pedestrian re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Flowchart of the anti-occlusion pedestrian re-identification method with posture-guided attention for this application;
[0030] Figure 2 This is a high-resolution network model structure diagram of an embodiment of the present application;
[0031] Figure 3 This is a schematic diagram of a posture-guided spatial attention module according to an embodiment of the present application;
[0032] Figure 4 This is a schematic diagram of the posture-guided channel attention module of an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0034] In one embodiment, Figure 1 As shown in the figure, a posture-guided attention-based anti-occlusion pedestrian re-identification method is proposed, including:
[0035] Step S1, obtain a pedestrian image, extract joint point heat map and joint point confidence through a posture estimation network, and extract a global feature map through a global feature extraction network.
[0036] This application uses a trained global feature extraction network and posture estimation network to extract global feature maps and joint point heat maps.
[0037] During training, after obtaining the pedestrian images, the pedestrian images can also be preprocessed by horizontal image flipping and image erasing to enrich the training samples. When performing data preprocessing, each pedestrian image is first adjusted to an image with a height of 256 and a width of 128. A threshold of 0.5 is given to generate a horizontal flip probability for each pedestrian image. When the probability is greater than the threshold, the left half and the right half of the image are horizontally flipped with the vertical center axis of the pedestrian image as the center axis; a random erasing probability is generated for each pedestrian image. When the probability is greater than the threshold, a rectangular area is randomly selected from the pedestrian image. The height of the rectangular area is 118 and the width is 88, and the pixel value of each channel in the area is set to 0.
[0038] After training the global feature extraction network and the posture estimation network, for the pedestrian image to be identified, you only need to resize the pedestrian image to a size of 256 in height and 128 in width, and input it into the trained network to extract features.
[0039] After obtaining the pedestrian image, this embodiment inputs it into the global feature extraction network and the posture estimation network respectively, extracts the global feature map and the joint point heat map, and obtains the joint point confidence based on the heat map. In this embodiment, the posture estimation network uses the high-resolution network HRNet-W32, and the global feature extraction network uses the residual network ResNet-50.
[0040] In a specific embodiment, extracting a joint point heat map through a posture estimation network includes:
[0041] High-resolution network HRNet-W32 Figure 2 As shown in Figure 1, the output of the high-resolution network HRNet-W32 is convolved to obtain a joint point heat map, and the maximum value in the joint point heat map is used as the corresponding joint point confidence.
[0042] The global feature map is extracted through the global feature extraction network, including:
[0043] The step size in the last layer Layer4 in the residual network ResNet-50 is set to 1 to obtain the global feature map.
[0044] Step S2: perform element-wise multiplication of the global feature map and the joint point heat map to obtain the joint point local feature map, and then use the posture-guided spatial attention module to obtain the spatial attention feature map through the joint point local feature map and the global feature map.
[0045] like Figure 3 As shown in the figure, firstly, the global feature map is multiplied with the joint point heat map to obtain the local feature map of the joint point.
[0046] In this embodiment, the global feature map and the joint point heat map are multiplied to obtain N joint point local features {F1, F2, ..., F N Then, the spatial attention feature map F is obtained through the posture-guided spatial attention module. s .
[0047] Specific as Figure 3 As shown, including:
[0048] Step S2.1, perform maximum pooling and ReLU activation function operations on the pixels in each channel of the joint point local feature map to generate a joint point spatial attention map.
[0049] The local feature map of each joint point is regarded as a feature map set with C channels and a size of H×W×1, as shown in the following formula:
[0050] F i ={s i1 ,s i2 ,…,s iC}
[0051] Where i represents the i-th joint point, s is the feature map of H×W×1, and C is the number of channels. i The pixels in each channel are pooled at the maximum, and the ReLU activation function is used to set the nonlinear mapping, as shown in the following formula:
[0052] S i =ReLU(Pool c (s i1 ,s i2 ,…,s iC ))
[0053] Generate a spatial attention map for each local feature of the joint point, and get the set
[0054] Step S2.2: After performing a convolution operation on the global feature map, a global average pooling operation is performed on the channel dimension to generate a global spatial attention map. The joint point spatial attention map is concatenated with the global spatial attention map to obtain a spatial posture guided feature map.
[0055] This step uses global features to generate an attention map with original information and concatenates it into a spatial posture guided feature map Y s , as shown below:
[0056]
[0057] where ψ s The embedding function representing the global feature F is implemented by using a 1×1 convolution in space and a batch normalization (Batch Normalization, BN) layer and ReLU activation function. c represents the global average pooling in the channel dimension, Y s The size of the feature map Y is H×W×(N+1). s It contains rich semantic information.
[0058] Step S2.3: The spliced spatial posture guided feature map is subjected to convolution operation and Sigmoid function to obtain the spatial attention mask.
[0059] This step uses convolution operation and Sigmoid function to obtain the spatial attention mask a s As shown below:
[0060] a s =Sigmoid(Conv(ReLU(Y s )))
[0061] Among them, Conv is a 1×1 convolutional layer, which converts the channel dimension to 1 to obtain the final spatial attention mask.
[0062] Step S2.4, multiply the spatial attention mask with the global feature to obtain the spatial attention feature map.
[0063] Finally, the spatial attention mask a s Multiplying with the global feature F can achieve more attention to the joint information of pedestrians at the spatial level, distinguishing pedestrians from the background and occlusion to obtain the spatial attention feature map F s , as shown below:
[0064] F s =F·a s .
[0065] Step S3: The local feature map of the joint points and the spatial attention feature map are passed through the posture-guided channel attention module to obtain the channel attention feature map.
[0066] Specific as Figure 4 As shown, including:
[0067] Step S3.1, perform maximum pooling and ReLU activation function operations on each pixel point in the local feature map of the joint point to generate a joint point channel attention map.
[0068] Specifically, Figure 4 As shown, each joint point local feature map is regarded as a set of feature vectors with size H×W and C×1, as shown in the following formula:
[0069]
[0070] Where M = H × W, is the size of the corresponding feature map, c is the feature vector of size C × 1, by in the set {c1, c2, …, c M} performs maximum pooling on each pixel in the space, and uses the ReLU activation function to set the nonlinear mapping, as shown in the following formula:
[0071] C i =ReLU(Pool s (c i1 ,c i2 ,…,c iM ))
[0072] Among them, C i The size is 1×1×C, and a channel attention map is generated for each local feature of the joint point to obtain the set
[0073] Step S3.2: After performing a convolution operation on the spatial attention feature map, a global average pooling operation is performed in the spatial dimension to generate a global channel attention map. The joint channel attention map is concatenated with the global channel attention map to obtain a channel posture guidance feature map.
[0074] Similar to posture-guided spatial attention, we use the spatial attention feature map F s As the global feature F, we generate an attention map with the original channel information and concatenate it to obtain the channel posture guided feature map Y c , as shown below:
[0075]
[0076] where ψ c To use 1×1 convolution, BN layer and ReLU activation function on the channel, pool s Indicates global average pooling in the spatial dimension, Y c The size of is 1×C×(N+1).
[0077] Step S3.3: The spliced channel posture guide feature map is subjected to convolution operation and Sigmoid function to obtain the channel attention mask.
[0078] And for Y c Perform convolution operation and Sigmoid function to extract channel attention mask a c , as follows:
[0079] a c =Sigmoid(Conv(ReLU(Y c )))
[0080] Conv is a 1×1 convolutional layer. The spatial dimension is converted to 1 to obtain the final channel attention mask a c .
[0081] Step S3.4, multiply the channel attention mask by the global feature to obtain the channel attention feature map.
[0082] Finally, the channel attention mask a s and the spatial attention feature map F as the global feature map F s Multiplication achieves more attention to the joint information of pedestrians at the channel level, and finally obtains the channel attention feature map F c , as shown below:
[0083] F c =F·a c .
[0084] Step S4: When the joint point confidence is greater than or equal to a preset threshold, the joint point confidence is used as the joint point feature weight, otherwise the joint point feature weight is reset to zero.
[0085] Specifically, during testing, in order to prevent severely occluded joint point features from entering the distance similarity measurement, a threshold is set for filtering, as shown in the following formula:
[0086]
[0087] where h i represents the confidence score of the i-th joint point, γ is the set threshold, and v i is the joint feature weight during the test. When the i-th joint confidence score is too small, the joint is considered to be occluded, and its joint feature weight is reset to 0, which is equivalent to not performing similarity measurement between features. When the i-th joint confidence score exceeds the set threshold, the joint feature weight is the joint confidence score, and similarity measurement is performed.
[0088] Step S5: Calculate the similarity distance between the image of the pedestrian to be retrieved and the library image based on the global feature map, joint point local feature map and channel attention feature map of the image of the pedestrian to be retrieved, and select the most similar library image as the final recognition result.
[0089] In this embodiment, the global feature map of the pedestrian image to be retrieved (query image), the local feature maps of each joint point and the channel attention feature map are respectively compared with the feature vectors of the library image for similarity, and the distance between each feature vector is obtained, which is then calculated with the confidence of each joint point as the final similarity distance.
[0090] When measuring similarity, the constructed global features, channel attention features and joint point local features are used to calculate the distance, where the distance between the query image and the gallery image at the i-th joint point is:
[0091]
[0092] Where D(·) represents the distance measurement function. This application uses cosine distance for measurement. are the local feature vectors of the i-th joint point of the query image and gallery image respectively.
[0093] In addition, the distance between the channel attention features and the global features between the query image and the gallery image is also calculated:
[0094]
[0095]
[0096] Among them are the global feature vectors of the query image and the gallery image, respectively. are the channel attention feature vectors of the query image and gallery image respectively.
[0097] The similarity distance between the query image and the gallery image is calculated as follows:
[0098]
[0099] where d total represents the similarity distance between the query image and the gallery image, are the joint point feature weights of the i-th joint point in the query image and the gallery image respectively, and N+2 is the total number of feature vectors.
[0100] The final similarity distance is used for similarity comparison, the ranking result of the similarity distance is output, and the most similar gallery image is selected as the final recognition result.
[0101] In this application, the technical solution of this application is also verified through experiments. In the experiment, based on the Occluded-DukeMTMC dataset and the Partial-ReID dataset, the current mainstream pedestrian re-identification algorithm is compared, and the mean average precision (mAP) and the k-th hit rate (Rank-1) are used to measure the recognition performance. The experimental results are shown in Tables 1 and 2:
[0102]
[0103] Table 1
[0104]
[0105]
[0106] Table 2
[0107] Table 1 and Table 2 are the test results on the Occluded-DukeMTMC dataset and the Partial-reID dataset, respectively. By comparing the two tables, it can be seen that the method of this application is superior to other occluded pedestrian re-identification algorithms in 2019, 2020 and 2021 in occluded pedestrian re-identification. The mAP and Rank-1 indicators are relatively excellent. Therefore, the method of this application has certain advantages over other algorithms.
[0108] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A posture-guided attention-based anti-occlusion pedestrian re-identification method, characterized in that: The posture-guided attention anti-occlusion pedestrian re-identification method comprises: Obtain pedestrian images, extract joint point heat maps and joint point confidences through the posture estimation network, and extract global feature maps through the global feature extraction network; The global feature map is element-wise multiplied with the joint point heat map to obtain the joint point local feature map. The joint point local feature map and the global feature map are then guided by the posture spatial attention module to obtain the spatial attention feature map. The joint local feature map and the spatial attention feature map are passed through the posture-guided channel attention module to obtain the channel attention feature map; When the joint point confidence is greater than or equal to the preset threshold, the joint point confidence is used as the joint point feature weight, otherwise the joint point feature weight is reset to zero; The similarity distance between the pedestrian image to be retrieved and the library image is calculated based on the global feature map, joint point local feature map and channel attention feature map of the pedestrian image to be retrieved, and the most similar library image is selected as the final recognition result. The similarity distance calculation formula is as follows: Among them, d total Represents the similarity distance between the pedestrian image to be retrieved and the gallery image, denotes the joint feature weights of the i-th joint in the pedestrian image to be retrieved and the gallery image, respectively, and d global Represents the distance between the global features of the pedestrian image to be retrieved and the gallery image, d pga represents the distance between the channel attention features of the pedestrian image to be retrieved and the gallery image, d i represents the distance between the i-th joint point between the pedestrian image to be retrieved and the gallery image, and N represents the number of joint points.
2. The posture-guided attention anti-occlusion pedestrian re-identification method according to claim 1 is characterized in that: The joint point confidence is the maximum value in the corresponding joint point heat map.
3. The posture-guided attention anti-occlusion pedestrian re-identification method according to claim 1, characterized in that: The method of obtaining a spatial attention feature map by guiding the joint point local feature map and the global feature map through a posture-guided spatial attention module includes: Perform the maximum pooling and ReLU activation function operations on the pixels in each channel of the joint local feature map to generate the joint spatial attention map; After the global feature map is convolved, a global average pooling operation is performed on the channel dimension to generate a global spatial attention map. The joint point spatial attention map is concatenated with the global spatial attention map to obtain a spatial posture guided feature map. The spliced spatial posture guided feature map is convolved and subjected to the Sigmoid function to obtain the spatial attention mask. Multiplying the spatial attention mask with the global feature obtains the spatial attention feature map.
4. The posture-guided attention anti-occlusion pedestrian re-identification method according to claim 1, characterized in that: The method of obtaining a channel attention feature map by guiding the joint point local feature map and the spatial attention feature map through a posture-guided channel attention module comprises: Perform the maximum pooling and ReLU activation function operations on each pixel in the joint point local feature map to generate the joint point channel attention map; After convolution operation on the spatial attention feature map, global average pooling operation is performed in the spatial dimension to generate a global channel attention map. The joint channel attention map is concatenated with the global channel attention map to obtain the channel posture guidance feature map. The spliced channel posture guide feature map is convolved and subjected to the sigmoid function to obtain the channel attention mask. Multiply the channel attention mask with the global feature to get the channel attention feature map.
5. The posture-guided attention anti-occlusion pedestrian re-identification method according to claim 1, characterized in that: The distance of global features between the pedestrian image to be retrieved and the gallery image, the distance of channel attention features between the pedestrian image to be retrieved and the gallery image, and the distance of the i-th joint point between the pedestrian image to be retrieved and the gallery image are all measured using cosine distance.
6. The posture-guided attention anti-occlusion pedestrian re-identification method according to claim 1, characterized in that: The posture estimation network adopts the high-resolution network HRNet-W32, and the global feature extraction network adopts ResNet-50.