A Occluded Pedestrian Re-identification Method Based on Multi-Granularity Occlusion Perception
Through the multi-grained occlusion perception method, the multi-grained feature extractor and attention occlusion erasing technology are used to solve the problem of feature diversity and occlusion area erasing in occlusion pedestrian re-identification, and achieve higher accuracy pedestrian recognition.
Patent Information
- Application Number
- CN202111255008.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-10-27
AI Technical Summary
The prior art has problems with consistency caused by weak feature diversity and uneven distribution of occlusion information in the re-identification of occlusion pedestrians, and the occlusion area cannot be erased accurately, resulting in low recognition accuracy.
The multi-grained occlusion perception method is adopted to extract pedestrian features through a multi-grained feature extractor, divide feature maps in horizontal and vertical directions, calculate attention occlusion perception, realize the erasure of occlusion from coarse to thin, and improve feature representation through multi-level prediction.
It enhances the diversity of features, improves the accuracy of re-identification of occlusion pedestrians, and can erase occlusion areas more accurately, improving the accuracy of recognition.
Smart Images

Figure CN114091519B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and relates to a pedestrian re-identification method, in particular to an occluded pedestrian re-identification method based on multi-granularity occlusion perception. Background Art
[0002] Pedestrian re-identification (ReID), also known as pedestrian re-recognition, is a technology that uses computer vision technology to determine whether a specific pedestrian exists in an image or video sequence. It is widely regarded as a sub-problem of image retrieval. Given a monitored pedestrian image, retrieve the pedestrian image across devices. It aims to make up for the visual limitations of fixed cameras and can be combined with pedestrian detection and pedestrian tracking technologies, and can be widely applied to fields such as intelligent video surveillance and intelligent security.
[0003] Since people are easily occluded by some obstacles (such as luggage, counters, crowds, cars, trees), or due to part of the body walking out of the camera shooting area, occlusion occurs. Therefore, it is necessary to accurately match pedestrian pictures with only partial observability, which is the so-called occluded pedestrian re-identification problem. The goal of the pedestrian re-identification (ReID) task is to match images of the same person taken by different cameras. Although various methods for ReID have been proposed in recent years, one common method for ReID is based on multi-granularity perception. However, most of the existing methods for solving the pedestrian re-identification problem based on multi-granularity perception focus on the full-body images of people, ignoring the more challenging and frequently occurring pedestrian occlusion problem in practical applications. For the occluded pedestrian re-identification problem, there are three limiting factors in the existing technology. First, the features extracted by the existing technology are restricted by fixed segmentation and a single level, which results in weak feature diversity. Second, the uniform horizontal segmentation along the height dimension conflicts with the actual distribution of semantic information, that is, the occlusion information is not all horizontally distributed, which leads to a consistency problem. Finally, since the existing methods usually perform one-step prediction at a single level, the occluded area cannot be accurately removed. Summary of the Invention
[0004] To solve the above problems of the prior art, the present invention proposes an occluded pedestrian re-identification method based on multi-granularity occlusion perception. The method obtains the pedestrian image data to be re-identified in real time; inputs the data into a trained occluded pedestrian re-identification model based on multi-granularity occlusion perception, and outputs the occluded pedestrian re-identification result;
[0005] The process of training the occluded pedestrian re-identification model based on multi-granularity occlusion perception includes:
[0006] S1: Obtain the original pedestrian image data; preprocess the obtained data, and divide the preprocessed data to obtain a training set and a test set;
[0007] S2: Input the data of the training set into the multi-granularity feature extractor of the model to extract the features of pedestrians, obtain feature maps of different scales, and input the feature maps of different scales into different branches;
[0008] S3: Calculate the attention occlusion perception based on the feature maps to obtain the enhanced pedestrian features of the previous branch;
[0009] S4: Use multi-granularity information transfer to transfer the feature maps of the previous branch to the current branch;
[0010] S5: Calculate the attention occlusion perception based on the transferred feature maps and the feature maps of the current branch to obtain the enhanced pedestrian features of the current branch, and realize the occlusion erasure from coarse to fine;
[0011] S6: Concatenate the pedestrian features obtained after processing all branches together, and output the pedestrian features after erasing the occlusion of the occlusion pedestrian re-identification model based on multi-granularity occlusion perception;
[0012] S7: Construct the loss function of the model according to the obtained pedestrian features;
[0013] S8: Input the data in the test set into the model, continuously adjust the parameters of the model, and complete the training of the model when the value of the loss function is the smallest.
[0014] Further, the features of pedestrians extracted by the multi-granularity feature extractor include: using a single-frame extractor to extract the data at different stages to obtain feature maps of different scales, and inputting the feature maps of different scales into the convolutional layers in front of different branches. The convolutional layers in front of each branch adjust the feature maps of different scales to obtain the adjusted feature maps.
[0015] Further, the process of calculating the attention occlusion perception based on the feature maps includes:
[0016] Divide the feature maps of the branch along the horizontal and vertical directions respectively to obtain partitions;
[0017] Calculate the attention scores of each partition;
[0018] Obtain the attention mask according to the attention scores;
[0019] Obtain the enhanced pedestrian features according to the attention mask.
[0020] Further, the process of calculating the attention score for each partition is as follows: After the mapped features in the partition are processed by multiple 3×3 convolutional filters and 1×1 convolutional filters, and then through batch normalization, the relationship between the pixels in each partition and all the pixels in the entire image is calculated from the perspectives of channels and spatial positions, and the attention score for each partition is calculated based on the relationship between the pixels in each partition and all the pixels in the entire image.
[0021] Preferably, using multi-granularity information transfer, the process of transferring the feature map of the previous branch to the current branch includes:
[0022] Performing preliminary processing on the feature map of the previous branch and the feature map of the current branch to make the two feature maps have the same dimension;
[0023] Concatenating the two preliminarily processed feature maps;
[0024] Completing the aggregation of the concatenated feature maps;
[0025] Inputting the aggregated feature map into the current branch.
[0026] Further, the process of preliminary processing includes:
[0027] Using a 1×1 convolutional filter to adjust the number of channels C of the feature map of the previous branch so that the number of channels C of the feature map of the previous branch is the same as that of the feature map of the current branch;
[0028] Using bilinear interpolation to sample the feature map of the previous branch and the feature map of the current branch so that the size of the feature map of the previous branch is the same as that of the feature map of the current branch.
[0029] Further, the process of aggregating the concatenated feature maps includes: Passing the feature map through a 1×1 convolutional filter and batch normalization, and finally completing the aggregation after activation by ReLU.
[0030] Further, the multi-granularity feature extractor is a single-frame feature extractor ResNet-50 with shared weights.
[0031] Preferably, the expression of the loss function of the occluded pedestrian re-identification model based on multi-granularity occlusion awareness is:
[0032]
[0033]
[0034]
[0035] Among them, L is the total loss of the model, L tris the hard sampling triplet loss, L ce is the cross-entropy loss for smooth labels, f t a 、f t p 、f t n respectively represent the feature maps of the target, positive sample, and negative sample; L tr (f i ) represents the hard sampling triplet loss of the branch, L ce (v i ) represents the cross-entropy loss of the branch for smooth labels, N represents the number of samples, and m represents the margin distance; represents the feature embedding output by the model, W t 、b t 、b k represent learnable parameter matrices, y t is the corresponding true identity label, and C represents the number of label categories in the training set.
[0036] Advantages of the present invention:
[0037] The present invention improves the feature representation by locating and removing occlusions; first, pedestrian features are extracted using a multi-granularity feature extractor; second, attention-aware occlusion erasure is designed to learn local feature representations by dividing the feature map along the horizontal and vertical directions respectively; finally, considering the correlation between granularities, the relationship between granularities is established through information transmission to achieve the erasure of occlusions from coarse to fine; compared with fixed segmentation and single-level methods, the present invention enhances feature diversity; compared with the method that only uses horizontal segmentation along the height dimension, vertical division is added to handle arbitrarily distributed semantic information; compared with the one-step prediction method, multi-step prediction is performed at multiple levels to improve accuracy and achieve more accurate occluded pedestrian re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the multi-granularity occlusion perception schematic diagram of the present invention;
[0039] Figure 2 is the flowchart of attention-aware occlusion erasure of the present invention;
[0040] Figure 3 is an example display of the multi-granularity region of interest of a single-frame image of the present invention;
[0041] Figure 4 is an example display of the regions of interest of all input images of the present invention;
[0042] Figure 5 is the pseudo-code diagram of multi-granularity information transmission. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0044] The present invention proposes an occluded pedestrian re-identification method based on multi-granularity occlusion perception, as Figure 1 shown. This method obtains the pedestrian image data to be re-identified in real time; inputs this data into a trained occluded pedestrian re-identification model based on multi-granularity occlusion perception, and outputs the occluded pedestrian re-identification result;
[0045] The process of training the occluded pedestrian re-identification model based on multi-granularity occlusion perception includes:
[0046] S1: Obtain the original pedestrian image data; preprocess the obtained data, and divide the preprocessed data to obtain a training set and a test set;
[0047] S2: Input the data in the training set into the multi-granularity feature extractor of the model to extract the features of the pedestrian, obtain feature maps of different scales, and input the feature maps of different scales into different branches;
[0048] S3: Calculate the attention occlusion perception according to the feature maps to obtain the enhanced pedestrian features of the previous branch;
[0049] S4: Use multi-granularity information transfer to transfer the feature maps of the previous branch to the current branch;
[0050] S5: Calculate the attention occlusion perception according to the transferred feature maps and the feature maps of the current branch to obtain the enhanced pedestrian features of the current branch, and realize the occlusion erasure from coarse to fine;
[0051] S6: Concatenate the pedestrian features obtained after processing all branches together, and output the pedestrian features after erasing the occlusion of the occluded pedestrian re-identification model based on multi-granularity occlusion perception;
[0052] S7: Construct the loss function of the model according to the obtained pedestrian features;
[0053] S8: Input the data in the test set into the model, continuously adjust the parameters of the model, and complete the training of the model when the value of the loss function is the smallest.
[0054] Further, the process of obtaining the feature maps of different scales is as follows:
[0055] Randomly sample T frames from the input video, and use a multi-granularity feature extractor to extract the features of pedestrians. Using a multi-granularity feature extractor to extract the features of pedestrians includes: using a single-frame extractor to extract data at different stages to obtain feature maps of different scales, inputting the feature maps of different scales into the convolutional layers at the front of different branches, and each convolutional layer at the front of a branch adjusts the feature maps of different scales to obtain adjusted feature maps. Each scale of feature map corresponds to the input of a branch, and the corresponding feature map is input to each branch; the feature map is denoted as F i ={F t |t = 1, 2, ..., T}, where F t ∈R H×W×C is the single-frame feature map, and H, W, and C represent the width, height, and number of channels respectively.
[0056] Furthermore, the attention-aware occlusion erasure (AA-OE) process is as Figure 2 shown. The process of calculating the attention occlusion perception based on the feature map to obtain the enhanced pedestrian features is: dividing the feature map of the branch into a certain number of partitions along the horizontal and vertical directions respectively, and calculating the attention score of each partition; obtaining the attention mask according to the attention score; and obtaining the enhanced pedestrian features according to the attention mask.
[0057] The specific process is as follows: In different branches, the input feature map F t ∈R H×W×C is divided into a certain number of partitions P i along the horizontal and vertical directions respectively to comprehensively explore the pedestrian features. The size of the partition is the granularity, the large partition is the coarse granularity, and the small partition is the fine granularity. The number of partitions in each branch is fixed. For example, the partition of the first branch is divided into 2, the partition of the second branch is divided into 4, ...; due to the limitation of the image storage method in memory, F t is transposed and then divided to achieve vertical division. The division is expressed as follows:
[0058] x h =[x 1 ,x 2 ,…,x p
[0059] x v =[y 1 ,y 2 ,…,y p
[0060] where, x i and y i are F t and F t T and the final x h and x v both contain D = T × P partitions;
[0061] For the pixels in x h and x v i.e., the feature map, after being processed by multiple 3×3 convolutional filters and 1×1 convolutional filters, and then through batch normalization, calculate the relationship between the pixels in each partition and all the pixels in the entire image from the perspectives of channels and spatial positions, and calculate the attention score for each partition according to the relationship between the pixels in each partition and all the pixels in the entire image;
[0062] According to the attention scores, use the Sigmoid function to obtain the attention mask; the attention mask is an attention score matrix, denoted as At the same time, obtain the horizontal attention mask and the vertical attention mask After that, add the horizontal attention mask and the vertical attention mask and aggregate them into the initial feature map to obtain the enhanced pedestrian features, i.e., the complete feature map, and the output complete feature map is denoted as A i ∈R H×W×C with the same size as F i The specific formula is:
[0063]
[0064] Examples of multi - granularity regions of interest in a single - frame image are shown as Figure 3 shown, and examples of regions of interest for all input images are shown as Figure 4 shown.
[0065] Furthermore, the multi - granularity feature extractor is a single - frame feature extractor ResNet - 50 with shared weights.
[0066] Preferably, the process of transmitting the feature map of the previous branch to the current branch using multi - granularity information transfer includes:
[0067] Perform preliminary processing on the feature map of the previous branch and the feature map of the current branch to make the two feature maps have the same dimension;
[0068] Concatenate the two preliminarily processed feature maps;
[0069] Complete the aggregation of the concatenated feature maps;
[0070] Input the aggregated feature map into the current branch.
[0071] The pseudo-code for multi-granularity information transfer is as follows Figure 5 shown, indicating that this module aggregates two input images to achieve a coarse-to-fine transfer process.
[0072] Furthermore, the preliminary processing process includes: since the input is F of the previous branch i and F of the current branch j , and the number of channels C of F i is twice that of F i , the present invention uses a 1×1 convolutional filter to adjust the number of channels C of the feature map of the previous branch, so that the number of channels C of the feature map of the previous branch is the same as that of the feature map of the current branch; bilinear interpolation is used to sample the feature map of the previous branch and the feature map of the current branch, so that the size of the feature map of the previous branch is the same as that of the feature map of the current branch; preferably, the coefficient of bilinear interpolation is 2.
[0073] Furthermore, the aggregation process of the feature maps spliced together includes: passing the feature maps through a 1×1 convolutional filter and batch normalization, and finally completing the aggregation after activation by ReLU.
[0074] Furthermore, the loss function is used to supervise the pedestrian features output by the occlusion pedestrian re-identification model based on multi-granularity occlusion perception to confirm the pedestrian identity output by the model.
[0075] A total of two losses are used, namely the triplet loss L tr with hard sampling and the cross-entropy loss L ce with label smoothing, and the formulas are as follows
[0076]
[0077]
[0078] where L tr is the hard sampling triplet loss, and f t a , f t p , f t n represent the feature maps of the target, positive sample, and negative sample respectively; N represents the number of samples, m represents the margin distance, which is used to set the absolute distance between the positive and negative samples, and ||f t a -f t p ||2 represents the distance between the target and the positive sample, and ||f t a -f t n||2 represents the distance between the target and the negative samples. By calculating the hard sampling triplet loss to calculate the distances between the target and the positive and negative samples, the distance between the target and the positive sample is reduced, and the distance between the target and the negative sample is enlarged, so as to evaluate whether the features generated by the model belong to the positive sample.
[0079] L ce is the cross-entropy loss of the smoothed label. Among them, represents the feature embedding output by the model, and W t , b t , b k represent the learnable parameter matrices, and y t is the corresponding true identity label. C represents the number of label categories in the training set. By calculating the cross-entropy loss of the smoothed label, the gap between the target identity estimated by the model and the true identity label is calculated, so as to evaluate whether the target identity generated by the model is correct.
[0080] In order to supervise each branch, losses are added to each branch. The final loss function is:
[0081]
[0082] Among them, L is the total loss of the model, and L tr (f i ) represents the hard sampling triplet loss of the branch, and L ce (v i ) represents the cross-entropy loss of the smoothed label of the branch.
[0083] The present invention improves the feature representation by locating and removing occlusions. First, a multi-granularity feature extractor is used to extract pedestrian features. Secondly, an attention-aware occlusion erasure is designed to divide the feature map along the horizontal and vertical directions respectively to learn local feature representations. Finally, considering the correlation between the granularities, the relationship between the granularities is established through information transmission to realize the occlusion erasure from coarse to fine. The present invention enhances the feature diversity compared with the fixed segmentation and single-level methods; compared with the method that only uses horizontal segmentation along the height dimension, vertical division is added to cope with arbitrarily distributed semantic information; compared with the one-step prediction method, multi-step prediction is performed at multiple levels to improve the accuracy, realizing more accurate occluded pedestrian re-identification.
[0084] The above-mentioned embodiments further illustrate the purpose, technical solutions and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for occluded pedestrian re-identification based on multi-granularity occlusion perception, characterized in that Including: Real-time acquisition of pedestrian image data to be re-identified; Inputting this data into a trained occlusion pedestrian re-identification model based on multi-granularity occlusion perception to output the occlusion pedestrian re-identification result; The process of training the occlusion pedestrian re-identification model based on multi-granularity occlusion perception includes: S1: Acquire the original pedestrian image data; preprocess the acquired data, and divide the preprocessed data to obtain a training set and a test set; S2: Input the data of the training set into the multi-granularity feature extractor of the model to extract the features of the pedestrian, obtaining feature maps of different scales, and input the feature maps of different scales into different branches; S3: Calculate the attention occlusion perception based on the feature maps to obtain the enhanced pedestrian features of the previous branch; the process of calculating the attention occlusion perception based on the feature maps includes: Divide the feature maps of the branch along the horizontal and vertical directions respectively to obtain partitions; Calculate the attention score of each partition; Obtain the attention mask according to the attention score; Obtain the enhanced pedestrian features according to the attention mask; S4: Use multi-granularity information transfer to transfer the feature map of the previous branch to the current branch; the process of transferring the feature map of the previous branch to the current branch includes: Perform preliminary processing on the feature map of the previous branch and the feature map of the current branch to make the two feature maps have the same dimension; Concat and splice the two preliminarily processed feature maps; Complete the aggregation of the concatenated feature maps; Input the aggregated feature map into the current branch; S5: Calculate the attention occlusion perception based on the transferred feature map and the feature map of the current branch to obtain the enhanced pedestrian features of the current branch, realizing the occlusion erasure from coarse to fine; S6: Concat and splice the pedestrian features obtained after being processed by all branches together, and output the pedestrian features after erasing the occlusion of the occlusion pedestrian re-identification model based on multi-granularity occlusion perception; S7: Construct the loss function of the model according to the obtained pedestrian features; S8: Input the data in the test set into the model, continuously adjust the parameters of the model, and complete the training of the model when the value of the loss function is the smallest.
2. The occluded pedestrian re-identification method based on multi-granularity occlusion perception according to claim 1, wherein, The multi-granularity feature extractor extracts the features of the pedestrian, including: using a single-frame extractor to extract the data at different stages to obtain feature maps of different scales; inputting the feature maps of different scales into the convolutional layers in front of different branches, and each convolutional layer in front of the branch adjusts the feature maps of different scales to obtain the adjusted feature maps.
3. A method for occluded pedestrian re-identification based on multi-granularity occlusion perception according to claim 1, characterized in that The process of calculating the attention score of each partition is: after the mapped features in the partition are processed by 3×3 convolutional filters and 1×1 convolutional filters, and then batch normalization processing is performed, calculate the relationship between the pixels in each partition and all the pixels in the whole image from the perspectives of channels and spatial positions, and calculate the attention score of each partition according to the relationship between the pixels in each partition and all the pixels in the whole image.
4. A method for occluded pedestrian re-identification based on multi-granularity occlusion perception according to claim 1, characterized in that The process of preliminary processing includes: Use a 1×1 convolutional filter to adjust the number of channels C of the feature map of the previous branch so that the number of channels C of the feature map of the previous branch is the same as the number of channels C of the feature map of the current branch; Use bilinear interpolation to sample the previous branch feature map and the current branch feature map so that the size of the previous branch feature map is the same as that of the current branch feature map.
5. A method for occluded pedestrian re-identification based on multi-granularity occlusion perception according to claim 1, characterized in that, The aggregation process of the concatenated feature maps includes: passing the feature maps through a 1×1 convolutional filter and batch normalization, and finally completing the aggregation after activation with ReLU.
6. The occluded pedestrian re-identification method based on multi-granularity occlusion perception according to claim 1, wherein The multi-granularity feature extractor is the single-frame feature extractor ResNet-50 with shared weights.
7. A method for occluded pedestrian re-identification based on multi-granularity occlusion perception according to claim 1, characterized in that, The loss function expression of the occlusion pedestrian re-identification model based on multi-granularity occlusion awareness is: Among them, \(L\) is the total loss of the model, \(L\) tr is the hard sampling triplet loss, \(L\) ce is the cross-entropy loss of smooth labels, \(f\) t a 、\(f\) t p 、\(f\) t n respectively represent the feature maps of the target, positive sample, and negative sample; \(L\) tr \((f\) i ) represents the hard sampling triplet loss of the branch, \(L\) ce \((v\) i ) represents the cross-entropy loss of smooth labels of the branch, \(N\) represents the number of samples, and \(m\) represents the margin distance; represents the feature embedding output by the model, \(W\) t 、\(b\) t 、\(b\) k represent learnable parameter matrices, \(y\) t is corresponding true identity label, and \(C\) represents the number of label categories in the training set.
Citation Information
Cited By
Shielding pedestrian re-identification method based on attitude estimation
CN116935301A
A pose estimation based occluded pedestrian re-identification method
CN116935301B