Shielding pedestrian re-identification method based on random shielding enhancement and Transform
By introducing random occlusion enhancement and Transformer architectures into pedestrian re-identification technology, the problem of poor recognition performance in occlusion scenarios is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510216049.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing pedestrian re-identification technology shows great limitations in occlusion scenarios, mainly due to the lack of effective occlusion adaptability and feature extraction capabilities.
The occlusion pedestrian re-identification method based on random occlusion enhancement and Transformer is adopted. By adding random shape occlusion areas to the input image, complex occlusion situations in real scenes are simulated, and block characteristics of pedestrian images are extracted in combination with the Transformer architecture.
It significantly improves the recognition performance and robustness of the model in the case of occlusion, can extract pedestrian characteristics more effectively, and improves the accuracy and stability of pedestrian re-identification of occlusion.
Smart Images

Figure CN120220045A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pedestrian re-identification, and particularly relates to an occluded pedestrian re-identification method based on random occlusion enhancement and Transformer. Background Technique
[0002] In the fields of intelligent monitoring and public security, the pedestrian re-identification (Person Re-Identification, ReID) technology has become an important means to improve the intelligence level of security systems. Its goal is to achieve the matching and identification of pedestrians across camera scenarios, providing efficient and accurate pedestrian tracking for the monitoring system. However, pedestrian images in practical applications are often difficult to identify due to factors such as occlusion, pose changes, and perspective differences. Especially in occlusion scenarios, existing pedestrian re-identification technologies show great limitations. Miao et al. [1] proposed an occluded pedestrian re-identification (Occluded Person Re-identification, abbreviated as Occluded Re-ID) method and constructed a representative dataset Occluded-DukeMTMC to address the occlusion problem (as Figure 1 shown).
[0003] First, the adaptability to occlusion scenarios is insufficient. In the training data of current mainstream pedestrian re-identification models, there are usually few samples with diverse occlusions, resulting in a significant decline in the recognition performance of the model when encountering complex occlusions. Traditional data augmentation methods, such as random cropping, horizontal flipping, random erasing [2] etc., although they can increase the diversity of samples, are difficult to effectively simulate complex occlusion situations in real scenarios. Especially for random occlusion phenomena with various shapes and sizes, simple cropping and erasing cannot provide sufficient occlusion diversity, resulting in still much room for improvement in the performance of the model in complex occlusion scenarios. Wang et al. [3] proposed the NPO Augmentation Strategy to enhance the occlusion robustness of the model by simulating non-pedestrian occlusions (NPO). However, this method relies on accurate occlusion masks and has limited adaptability in different datasets or occlusion types, which may affect the performance of the model in new or unseen occlusion scenarios.
[0004] Secondly, there is a lack of feature extraction ability for occlusion. Common person re-identification models are mostly based on global feature extraction methods. However, in occlusion scenarios, global features are easily interfered by the occluded areas, resulting in a decrease in recognition accuracy. Some studies have tried to adopt local feature learning methods. However, due to the unpredictability of the occlusion positions, it is difficult to ensure the effectiveness of local feature learning. The reliability of local feature methods depends on the accurate positioning of the occluded areas. However, in practical applications, occlusion situations are complex and diverse, and such features cannot be stably provided.
[0005] In recent years, the Transformer architecture has demonstrated superior modeling capabilities and accuracy in the field of vision, especially having significant advantages in feature representation and global dependency modeling. In the 2021 paper "TransReID: Transformer-based Object Re-Identification" by He et al. [4] first applied a pure Transformer-based network architecture to the field of person re-identification, and its model diagram is as Figure 2 shown. However, the Transformer model applied to the field of person re-identification also faces challenges when dealing with occlusion problems. The Transformer is good at capturing global information. However, in occlusion scenarios, due to noise interference, the feature extraction process may cause the model to focus on invalid areas, thereby affecting the recognition effect. Therefore, how to effectively introduce the Transformer model while overcoming the noise interference brought by occlusion is an important research direction in the current field of person re-identification.
[0006] To solve this problem, Wang et al. [5] proposed the FCFormer model, which uses feature completion technology to alleviate the occlusion problem and enhances the training data through Occlusion Instance Augmentation. Although this method improves the occlusion robustness to a certain extent, it relies on complex occlusion enhancement strategies and may still face the situation of insufficient sample generation when dealing with the diversity of occlusions in real scenarios. Especially when the occlusion information is severely lost, the FCFormer is difficult to effectively recover features, affecting the recognition accuracy, and has poor adaptability to new types of occlusions. At the same time, it also requires more computing resources.
[0007] References
[0008] [1] Miao J, Wu Y, Liu P, et al. Pose-guided feature alignment for occluded person re-identification[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2019:542-551.
[0009] [2] Zhong Z, Zheng L, Kang G, et al. Random erasing data augmentation[C] / / Proceedings of the AAAI conference on artificial intelligence. 2020, 34(07):13001-13008.
[0010] [3] Wang Z, Zhu F, Tang S, et al. Feature erasing and diffusion network for occluded person re-identification[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022:4754-4763.
[0011] [4] He S, Luo H, Wang P, et al. Transreid: Transformer-based object re-identification[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021:15013-15022.
[0012] Wang T, Liu M, Liu H, et al. Feature completion transformer for occluded person re-identification[J]. IEEE Transactions on Multimedia, 2024. Summary of the Invention
[0013] In view of the problems existing in the prior art, the present invention provides an occluded pedestrian re-identification method based on random occlusion enhancement and Transformer. This method adds occluded regions with random shapes to the input image to simulate complex occlusion situations in real scenes, thereby enhancing the adaptability of the model to occluded scenarios. In addition, this method combines the powerful global dependency modeling ability of the Transformer architecture to effectively extract the block features of pedestrian images, and at the same time uses the occlusion enhancement strategy to enable the model to focus on non-occluded regions, thus significantly improving the recognition performance in the case of occlusion.
[0014] To solve the above technical problems, the present invention provides the following technical solutions: An occluded pedestrian re-identification method based on random occlusion enhancement and Transformer, comprising the following steps:
[0015] S1. Construct a random occlusion enhancement module to perform random occlusion enhancement on a number of color pictures;
[0016] S2. Construct an image block projection module to divide the occluded and enhanced image into blocks, and map them to a high-dimensional feature space through a linear projection layer, while adding position and camera information embeddings;
[0017] S3. Construct a self-attention feature extraction module, input the projected image blocks and embedding information into a multi-layer Transformer encoder, extract global features and local features with strong occlusion robustness through the self-attention mechanism, and fuse the global features and local features;
[0018] S4. Based on the random occlusion enhancement module, the image block projection module, the self-attention feature extraction module, and the prediction module, construct a random occlusion enhancement pedestrian re-identification network, train the random occlusion enhancement pedestrian re-identification network with a number of color pictures as input and the corresponding pedestrian recognition classification results as output, use a joint loss function based on the triplet loss function and the cross-entropy loss function with label smoothing to train the network during the training process, obtain a random occlusion enhancement pedestrian re-identification model, and finally use this model to realize the re-identification of occluded pedestrian pictures.
[0019] Further, the aforementioned random occlusion enhancement module is configured to perform the following steps:
[0020] S1.1. Determine the area of the occluded region;
[0021] S1.2. Calculate the aspect ratio of the occluded region;
[0022] S1.3. Randomly select the position of the occluded region;
[0023] S1.4. Generate occlusion masks with different shapes;
[0024] S1.5. Perform occlusion filling using a mask.
[0025] Furthermore, for the aforementioned occlusion pedestrian re-identification method based on random occlusion enhancement and Transformer, random occlusion enhancement performs data augmentation on the input feature map through the random occlusion enhancement method ROA. Specifically, given a predefined occlusion probability P, ROA randomly selects a rectangular region of varying size at any position in the image during the training process and generates a corresponding occlusion mask M. Using this mask, the pixel values within the occluded region are replaced with random values to simulate uncertain occlusion situations, making the training process closer to the complex occlusion environment in the actual scene.
[0026] Furthermore, in the aforementioned step S1.4, the shape S of the occluded region t is randomly selected from a rectangle, an ellipse, a triangle, and a rhombus. The mask M is generated according to the selected shape and is used to define the specific position of the occluded region. Specifically, as follows: If S t is a rectangle, all pixels within its mask M r are 1, and the remaining regions are 0;
[0027] If S t is an ellipse, the mask M e satisfies the ellipse equation where (x c , y c ) is the center, and r x and r y are the semi-axes of the ellipse;
[0028] If S t is a triangle, the mask M t satisfies where (x0, y0) is the vertex of the triangle;
[0029] If S t is a rhombus, the mask M d satisfies where (x c , y c ) is the center of the rhombus.
[0030] Furthermore, in the aforementioned step S1.5, according to the generated mask M, the pixel values of the occluded region are replaced with random filling values F v , to achieve a random occlusion effect. The occluded image is represented as:
[0031] I masked = I ⊙ (1 - M) + F v ⊙ M (1)
[0032] where: I maskedDenote the image after occlusion as, I is the original input image, and ⊙ represents the element-wise multiplication operation.
[0033] F v Fill the randomly generated occlusion region with values, of size (C, H o , W o ), where C is the number of channels.
[0034] Furthermore, in the aforementioned step S2, the image patch projection module is configured to perform the following actions:
[0035] S2.1. Divide the preprocessed image into N image patches of a fixed size: where the size of each image patch is P and the stride is S. The number of image patches N is shown in Equation 2: When the stride S is less than the image patch size P, there will be overlapping regions between the image patches, enabling the model to capture local details and subtle features;
[0036]
[0037] S2.2. Map each image patch to a D-dimensional feature space through linear projection to form an image patch embedding and add a learnable class token x cls as the global feature representation to retain global information;
[0038] S2.3. Further introduce a learnable position encoding P E , retain the spatial position information of the image patch, and add a camera information embedding C id , to reduce the influence of different camera angles on the feature representation. The final embedded data is shown in Equation 3:
[0039] E input = {x class ; E i} + P E + λ cm C id (3)
[0040] where P E represents the position embedding, C id represents the camera information embedding, λ cm is a hyperparameter used to control the weight of the camera embedding, and {x class ; E i} represents the combination of the learnable image class label and the corresponding feature embedding vector of the image patch.
[0041] Furthermore, in the aforementioned step S3, the self-attention feature extraction module is configured to perform the following steps:
[0042] S3.1. Input the embedded data into several Transformer encoding layers of the ViT backbone network, and use the pre-trained weights to accelerate model convergence and improve feature expression ability;
[0043] S3.2. After the embedded data is processed by the Transformer encoding layer, the encoder output features are obtained The output features include global features and local features representing the semantic information of the overall image and the detail information of the image patches respectively;
[0044] S3.3. Divide the local feature f part into K groups, each group with a size of (N / / K)×D, denoted as f part1 , f part2 , …, f partk ;
[0045] S3.4. Concatenate the global feature f gb with each group of local features respectively to form the local visual semantic feature f vp , denoted as: f vp = [[f gb , f part1 , [f gb , f part2 ,... [f gb , f partk ;
[0046] S3.5. Input the fused feature f vp and the global feature f gb together into the subsequent prediction module to achieve accurate classification and matching of pedestrian identities.
[0047] Furthermore, in the aforementioned step S4, the label-smoothing cross-entropy loss function is as follows:
[0048]
[0049] where B represents the training batch size, C represents the number of classes, y i,c is the true label of sample i, is the probability distribution predicted by the model, α is the label-smoothing coefficient.
[0050] Furthermore, in the aforementioned step S4, the triplet loss function is as follows:
[0051]
[0052] Among them, B represents the training batch size, p represents the positive samples, n represents the negative samples, d(·, ·) is the distance metric between samples, and β is a hyperparameter used to control the margin.
[0053] Furthermore, in the aforementioned step S4, the joint loss function is as follows:
[0054]
[0055] Among them, f gb is the global feature, is the local visual semantic feature, and K is the number of groups of local visual semantic features.
[0056] Compared with the prior art, the beneficial technical effects of the present invention adopting the above technical solutions are as follows:
[0057] The present invention proposes an occluded pedestrian re-identification method based on random occlusion augmentation and Transformer, aiming to improve the recognition performance of the pedestrian re-identification model in complex occlusion scenarios. Based on the existing pedestrian re-identification model, this method significantly improves the robustness and accuracy of the model in the case of occlusion by introducing random occlusion augmentation (ROA) and Transformer encoding structure.
[0058] The present invention designs a model with Vision Transformer as the backbone network and adds a random occlusion augmentation module to simulate diverse occlusions in the real scene. By generating occlusion regions with random sizes and shapes on the preprocessed images, the model can be exposed to a more diverse set of occluded samples during training, thus enhancing the generalization ability of the model. In addition, the present invention combines the cross-entropy loss with label smoothing regularization and the triplet loss. By making a refined distinction of the sample features, the samples of the same identity are more closely clustered in the feature space.
[0059] The random occlusion augmentation method proposed by the present invention effectively increases the diversity of training samples and significantly enhances the adaptability of the model in occlusion scenarios. This method not only improves the model's ability to handle local occlusions but also makes the recognition performance of the model more stable under different camera angles. Secondly, the Transformer-based encoder structure designed by the present invention captures the global dependencies of image features through the self-attention mechanism, enabling the model to focus on important unoccluded regions and suppressing the interference of occlusion noise, thereby improving the overall feature expression effect.
[0060] Experiments on the Occluded-DukeMTMC dataset show that the method of the present invention has excellent performance in terms of the correct recognition rate (Rank-1), correct recognition rate (Rank-5), correct recognition rate (Rank-10), and mean average precision (mAP), which are 68.30%, 82.30%, 86.70%, and 60.00% respectively. Compared with existing occluded pedestrian re-identification models, the method of the present invention significantly improves the recognition accuracy in occluded situations, can more effectively extract pedestrian features, and provides a more robust solution for the occluded pedestrian re-identification task. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 FIG. is a schematic diagram showing the occlusion situation of two different pedestrians in the Occluded-DukeMTMC dataset.
[0062] Figure 2 FIG. is an architecture diagram of a ReID model based on pure Transformer.
[0063] Figure 3 FIG. is a network framework diagram of the occluded pedestrian re-identification method of the present invention with random occlusion augmentation and Transformer.
[0064] Figure 4 FIG. is a flowchart of the training stage of the present invention.
[0065] Figure 5 FIG. is an effect diagram of random occlusion augmentation. DETAILED DESCRIPTION OF THE INVENTION
[0066] In order to better understand the technical content of the present invention, specific embodiments are given below in conjunction with the accompanying drawings for illustration.
[0067] In the present invention, various aspects of the present invention are described with reference to the accompanying drawings, in which many illustrative embodiments are shown. The embodiments of the present invention are not limited to those described in the drawings. It should be understood that the present invention can be implemented by any one of the various concepts and embodiments introduced above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in the present invention are not limited to any embodiment. In addition, some aspects disclosed in the present invention can be used alone, or in any appropriate combination with other aspects disclosed in the present invention.
[0068] The present invention provides an occluded pedestrian re-identification method based on random occlusion augmentation and Transformer to improve the performance of pedestrian re-identification in occluded situations. First, a random occlusion augmentation module, an image block projection module, and a self-attention feature extraction module are constructed through steps S1 to S3, and then a random occlusion augmentation pedestrian re-identification network is constructed and trained. The network structure is as Figure 3 shown, and the training process is asFigure 4 As shown below. The specific steps are as follows:
[0069] S1. Construct a random occlusion enhancement module to perform random occlusion enhancement on a number of color pictures;
[0070] S2. Construct an image block projection module to divide the occluded and enhanced image into blocks, and map them to a high-dimensional feature space through a linear projection layer, while adding position and camera information embeddings;
[0071] S3. Construct a self-attention feature extraction module, input the projected image blocks and embedding information into a multi-layer Transformer encoder, extract global features and local features with strong occlusion robustness through the self-attention mechanism, and fuse the global features and local features;
[0072] S4. Based on the random occlusion enhancement module, the image block projection module, the self-attention feature extraction module, and the prediction module, construct a random occlusion enhanced pedestrian re-identification network. Use a number of color pictures as input and the corresponding pedestrian recognition classification results as output to train the random occlusion enhanced pedestrian re-identification network. During the training process, use a joint loss function based on the triplet loss function and the cross-entropy loss function with label smoothing to train the network to obtain a random occlusion enhanced pedestrian re-identification model. Finally, use this model to achieve the re-identification of occluded pedestrian pictures.
[0073] As a preferred embodiment of the present invention, the color images in step S1 are all from the standard dataset for occluded pedestrian re-identification, the Occluded-DukeMTMC dataset. Each image is composed of three primary colors, red, green, and blue, and contains three channels, and each channel corresponds to the corresponding primary color. In the embodiment, there are n1 pedestrian images in total, and the image samples can be expressed as where, I i represents the i-th pedestrian image, and y i respectively represent the identities corresponding to I i . Consistent with the above description, the present invention takes the input of B samples {I i , y i} as an example to introduce the working principle of the present invention during the training process.
[0074] As a preferred embodiment of the present invention, in step S1, the present invention uses the Random Occlusion Augmentation (ROA) method to perform data augmentation on the input feature map, increasing the diversity of training samples and enhancing the generalization ability and robustness of the model. The basic idea of ROA is to given a predefined occlusion probability P. Different from ordinary random erasing, ROA selects a rectangular region of random size at any position in the image during the training process, generates a corresponding occlusion mask M, and combines this occlusion mask to replace the pixel values in this region with random values, thereby effectively simulating uncertain occlusion and making the training process closer to the complex occlusion situation in the actual scenario. As Figure 5 shown in the schematic diagram of random occlusion augmentation.
[0075] As a preference of the present invention, the implementation process of the random occlusion augmentation method in step S1 is as follows:
[0076] Given an input image This module enhances the image by randomly selecting occlusion regions of different shapes (including rectangles, ellipses, triangles, rhombuses, etc.). The occlusion probability is p, that is, each image is occluded with probability p, and the probability of not being occluded is 1 - p. The size, shape, and position of the occlusion region are all randomly generated to simulate the diverse occlusion situations in the real scenario. The specific steps are as follows:
[0077] S1.1. Determine the area of the occlusion region
[0078] The area S of the occlusion region o is a proportion of the total area S = H × W of the image, and this proportion is randomly selected within the range of the minimum value s l and the maximum value s h , that is:
[0079] S o = rand(s l , s h ) × S
[0080] S1.2. Calculate the aspect ratio of the occlusion region
[0081] The aspect ratio r of the occlusion region o is randomly generated within the interval [r1, r2]. According to the area S o and the aspect ratio r o , calculate the height H o and width W o of the occlusion region I o :
[0082]
[0083] S1.3. Randomly select the position of the occlusion region
[0084] Randomly initialize the upper-left position P = (x o , y o ), as the starting point of the occlusion area. If x o +W o ≤W and y o +H o ≤H, then select the area I o =(x o , y o , x o +W o , y o +H o ) as the occlusion area; otherwise, repeat this process until a suitable occlusion area I o is selected.
[0085] S1.4. Generate occlusion masks of different shapes
[0086] To simulate diverse occlusion scenarios, the shape S t of the occlusion area is randomly selected from a rectangle, an ellipse, a triangle, and a rhombus. The mask M is generated according to the selected shape and is used to define the specific position of the occlusion area:
[0087] If S t is a rectangle, all pixels within its mask M r are 1, and the rest are 0.
[0088] If S t is an ellipse, the mask M e satisfies the ellipse equation where (x c , y c ) is the center, and r x and r y are the semi-axes of the ellipse.
[0089] If S t is a triangle, the mask M t satisfies where (x0, y0) is the vertex of the triangle.
[0090] If S t is a rhombus, the mask M d satisfies where (x c , y c ) is the center of the rhombus.
[0091] S1.5. Use the mask for occlusion filling
[0092] According to the generated mask M, replace the pixel values of the occlusion area with the randomly filled value F v, to achieve a random occlusion effect. The occluded image is represented as:
[0093] I masked = I ⊙ (1 - M) + F v ⊙ M (1)
[0094] Where: I masked represents the image after occlusion, I is the original input image, ⊙ represents the element-wise multiplication operation, F v is the filling value for the randomly generated occlusion area, with a size of (C, H o , W o ), where C is the number of channels.
[0095] Through this formula, the non-occluded areas of the image remain unchanged, while the occluded areas are randomly filled with F v to finally achieve the random occlusion effect of the image.
[0096] As a preferred embodiment of the present invention, in step S2, the image block projection module is configured to perform the following actions: S2.1, divide the preprocessed image into N image blocks (patches) of a fixed size: where the size of each image block is P, the step size is S, and the number of image blocks N, as shown in formula (2): when the step size S is less than the image block size P, there will be overlapping areas between the image blocks, enabling the model to capture local details and subtle features;
[0097]
[0098] When the step size S is less than P, overlapping areas will appear between the generated image blocks, enabling the model to capture local details and subtle features.
[0099] S2.2, map each image block to a D-dimensional feature space through linear projection to form an image block embedding and add a learnable class token x cls as the global feature representation to retain global information.
[0100] S2.3,, further introduce a learnable position encoding P E , retain the spatial position information of the image blocks, and add a camera information embedding C id , to reduce the influence of different camera angles on the feature representation. Finally, the embedded data is as shown in formula 3:
[0101] E input = {x class ; E i}+ P E + λ cm C id (3)
[0102] Among them, P E represents the position embedding, C id represents the camera information embedding, λ cm is a hyperparameter used to control the weight of the camera embedding, {x class ; E i} represents the combination of the learnable image class label and the feature embedding vector corresponding to the image patch.
[0103] Through the above processing, the embedding vector of the image patch contains rich context information, providing a better expression basis for subsequent Transformer encoder feature extraction.
[0104] As a preferred embodiment of the present invention, the self-attention feature extraction module in step S3 is configured to perform the following steps: S3.1, input the embedding data into several Transformer encoding layers of the ViT backbone network, and use the pre-trained weights to accelerate the model convergence and improve the feature expression ability;
[0105] S3.2, after the embedding data is processed by the Transformer encoding layer, the encoder output features are obtained The output features include global features and local features representing the semantic information of the overall image and the detailed information of the image patch respectively;
[0106] S3.3, divide the local feature f part into K groups, each group with a size of (N / / K)×D, denoted as: f part1 , f part2 ,..., f partk ;
[0107] S3.4, concatenate the global feature f gb with each group of local features respectively to form local visual semantic features f vp , denoted as: f vp = [[f gb , f part1 , [f gb , f part2 ,... [f gb , f partk ; By the concatenation operation, the global and local features are fused together, so that each local feature group not only contains local detailed information but also has overall semantic understanding, thereby enhancing the robustness and recognition performance of the model in occlusion scenarios.
[0108] S3.5, the fused feature f vp and the global feature f gbThey are jointly input into the subsequent prediction module to achieve accurate classification and matching of pedestrian identities. This design enables the model to have strong feature expression capabilities at both the global and local levels, especially showing higher recognition accuracy in complex scenes with occlusions.
[0109] In step S4 of the present invention, the backbone network Vision Transformer is used to extract features. Using the pre-trained weights of ViT Base, which is pre-trained on ImageNet21K and then fine-tuned on ImageNet-1K, can accelerate the convergence speed of network training and the final prediction accuracy, and improve the feature expression ability. In the training process of the present invention, the network is jointly trained using the triplet loss function and the cross-entropy loss function with label smoothing regularization. These two loss functions are commonly used loss functions in the field of pedestrian re-identification.
[0110] In step S4, the cross-entropy loss function with label smoothing is as follows:
[0111]
[0112] where B represents the training batch size, C represents the number of classes, y i,c is the true label of sample i, is the probability distribution predicted by the model, α is the label smoothing coefficient.
[0113] In step S4, the triplet loss function is as follows:
[0114]
[0115] where B represents the training batch size, p represents the positive sample, n represents the negative sample, d(·,·) is the distance metric between samples, and β is a hyperparameter used to control the margin.
[0116] Finally, the joint loss function is as follows:
[0117]
[0118] where f gb is the global feature, is the local visual semantic feature, and K is the number of groups of local visual semantic features.
[0119] The test process of the present invention is as follows:
[0120] Step 1: Input the query set and the gallery set, and enter step 2;
[0121] Step 2: Use the model obtained in the training process to extract features from all pedestrian images in the query set and the gallery set input in Step 1, and proceed to Step 3;
[0122] Step 3: Calculate the similarity between the query set features and the gallery set features, and proceed to Step 4;
[0123] Step 4: Based on the similarity level, obtain the matching results corresponding to each pedestrian image in the query set, and proceed to Step 5;
[0124] Step 5: End.
[0125] Preferably, in Step 1 of the test process, the query set represents the set of pedestrian images to be recognized, while the gallery set represents the set of target pedestrian images that the query set matches.
[0126] Preferably, in Step 3 of the test process, the similarity calculation method is cosine similarity, which is used to evaluate the similarity of the image features between the query set and the gallery set.
[0127] Preferably, in Step 4 of the test process, each query set image corresponds to several matching images in the gallery set, and the test results use the Cumulative Matching Characteristic (CMC) and the mean Average Precision (mAP) as evaluation indicators. Among them, the Rank-k accuracy of CMC measures the probability of a correct match appearing in the top k retrieval results, and mAP reflects the average retrieval performance of the method over the entire query set.
[0128] Although the present invention has been described above with preferred embodiments, it is not intended to limit the present invention. Those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the scope defined in the claims.
[0129] In this embodiment, the VIT network is used as the backbone of the visual feature extraction network, and the sizes of the training and test images are both adjusted to 256×128. The training images are enhanced by random horizontal flipping, padding, random cropping, and random occlusion. The initial weights of the VIT network are pre-trained on ImageNet-21K and then fine-tuned on ImageNet1K. In this paper, the number of split groups K is set to 17. The batch size is set to 64, with 4 images per ID. The initial learning rate is 0.008, and the learning rate decay is implemented using cosine annealing. The sizes of all images are set to 256×128, and the sliding stride is set to [11,11]. The experiment is implemented using the PyTorch 1.8 deep learning module and accelerated on an NVIDIA 4090 GPU.
[0130] This embodiment uses the Occluded-Duke dataset to complete the occluded pedestrian re-identification task to verify the recognition performance of the method of the present invention in occluded scenarios.
[0131] This experiment was evaluated on the Occluded-Duke dataset, and the performance of the method of the present invention and multiple existing mainstream pedestrian re-identification methods was tested. The comparison results are shown in Table 1, the performance comparison table of the method of the present invention and other methods on the Occluded-Duke dataset.
[0132] Table 1
[0133]
[0134]
[0135] In this experiment, a variety of existing pedestrian re-identification methods were selected for comparison to verify the effectiveness of the present invention. The compared models cover several representative methods, including: global feature-based methods, such as PCB (ECCV2018), which extract pedestrian global features through partial pooling. Pose-guided partial matching methods, such as PVPM (CVPR2020), which guide the model to match visible parts in occluded scenarios. High-order information-based methods, such as HOReID (CVPR2020), which improve the model's adaptability to occluded scenarios through relationship and topology learning. Transformer-based models, such as PAT (CVPR 2021) and TransReID (ICCV 2021), which use the Transformer framework to extract pedestrian features. Feature erasure and diffusion-based networks, such as FED (CVPR 2022), which enhance the occlusion robustness of the model through feature erasure.
[0136] All experiments adopted the single-frame query mode. For the performance metrics of the model, the Rank-1 accuracy and mean average precision (mAP) were mainly used as the evaluation criteria.
[0137] As can be seen from Table 1, the Rank-1 and mAP of the method of the present invention on the Occluded-Duke dataset reached 68.3% and 60.0% respectively, showing significant advantages in the occluded pedestrian re-identification scenario. Compared with PCB, the Rank-1 and mAP increased by 25.7% and 26.3% respectively, and compared with PVPM, they increased by 21.3% and 22.3% respectively. Compared with HOReID based on high-order information, the present invention improved by 13.2% and 16.2% respectively in Rank-1 and mAP.
[0138] In addition, compared with other Transformer-based methods, such as PAT and TransReID, the present invention achieved an improvement of more than 3.8% in Rank-1 and mAP on the Occluded-Duke dataset. This shows that in the occluded scenario, the model structure combining random occlusion augmentation (ROA) and Transformer can effectively enhance the feature extraction ability of the model in the occluded situation, enabling the model to still obtain excellent recognition performance in relatively complex occluded scenarios.
[0139] In summary, the occluded pedestrian re-identification method based on random occlusion augmentation and Transformer of the present invention achieved better performance than the existing methods on the Occluded-Duke dataset, proving the effectiveness and robustness of the method of the present invention in the occluded scenario, and providing a new solution idea for the occluded pedestrian re-identification task.
[0140] Example 2:
[0141] This example will introduce an applicable scenario of the present invention.
[0142] In environments such as large transportation hubs, commercial centers, and public places with dense monitoring coverage areas, pedestrian re-identification systems often face complex occlusion problems. For example, in stations, airports, shopping malls, etc., pedestrians are dense and occlusion phenomena occur frequently, resulting in a significant decline in the recognition effect of existing pedestrian re-identification models in such scenarios.
[0143] First, apply the occluded pedestrian re-identification method based on random occlusion augmentation and Transformer of the present invention to the pedestrian re-identification system to enhance its adaptability to occlusion situations, thereby improving the recognition accuracy and stability in the occluded scenario.
[0144] Next, by combining the method of the present invention with the existing surveillance camera system, various occlusion scenarios are simulated during the training stage, enabling the system to accurately identify pedestrian identities under different camera angles and occlusion conditions, ensuring the effectiveness of the surveillance system.
[0145] Finally, to adapt to the ever-changing occlusion scenarios, the pedestrian re-identification system of the present invention can regularly update the occlusion enhancement model, enabling it to maintain a high recognition accuracy when facing different occlusion degrees and scenarios, providing more robust technical support for public safety.
Claims
1. A method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer, characterized in that: The following steps are involved: S1. Construct a random occlusion enhancement module to perform random occlusion enhancement on several color images; S2, build an image block projection module, divide the image after occlusion enhancement into blocks, and map it to a high-dimensional feature space through a linear projection layer, while embedding the position and camera information; S3. Construct a self-attention feature extraction module, input the projected image blocks and embedded information into the multi-layer Transformer encoder, extract global features and local features with strong occlusion robustness through the self-attention mechanism, and fuse the global features and local features; S4. Based on the random occlusion enhancement module, the image block projection module, the self-attention feature extraction module, and the prediction module, a random occlusion enhanced pedestrian re-identification network is constructed. The random occlusion enhanced pedestrian re-identification network is trained with several color pictures as input and the corresponding pedestrian recognition classification results as output. During the training process, the network is trained using a joint loss function based on the triplet loss function and the label smoothed cross entropy loss function to obtain a random occlusion enhanced pedestrian re-identification model. Finally, the model is used to realize the re-identification of occluded pedestrian pictures.
2. According to claim 1, the method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer is characterized in that: The random occlusion enhancement module is configured to perform the following steps: S1.
1. Determine the area of the blocked area; S1.2, calculate the aspect ratio of the occluded area; S1.3, randomly select the location of the occluded area; S1.4, generating occlusion masks of different shapes; S1.
5. Use mask to perform occlusion and filling.
3. The method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer according to claim 2, characterized in that: Random occlusion enhancement uses the random occlusion enhancement method ROA to perform data expansion on the input feature map. Specifically, given a predefined occlusion probability P, ROA will randomly select a rectangular area of different sizes at any position in the image during training and generate a corresponding occlusion mask M. Using this mask, the pixel values in the occluded area are replaced with random values to simulate uncertain occlusion situations, making the training process closer to the complex occlusion environment in the actual scene.
4. The method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer according to claim 3, characterized in that: In step S1.4, the shape S of the occluded area t A random selection is made among rectangles, ellipses, triangles and diamonds. The mask M is generated according to the selected shape to define the specific location of the occluded area, as follows: If S t is a rectangle, whose mask M r All pixels in the region are 1, and the rest of the region is 0; If S t is an ellipse, mask M e Satisfies the ellipse equation Where (x c ,y c ) is the center, r x and r y is the semi-axis of the ellipse; If S t is a triangle, mask M t satisfy Where (x0, y0) is the vertex of the triangle; If S t is a diamond shape, mask M d satisfy Where (x c ,y c ) is the center of the rhombus.
5. According to claim 3, the method for re-identifying pedestrians under occlusion based on random occlusion enhancement and Transformer is characterized in that: Step S1.5: Replace the pixel values of the occluded area with random fill values F according to the generated mask M. v , to achieve the random occlusion effect, the occluded image is expressed as: I masked =I(1-M)+F v ⊙M(1) Where: I masked represents the image after occlusion, I is the original input image, represents the element-by-element multiplication operation, and F v Fill the randomly generated occlusion area with a size of (C,H o ,W o ), where C is the number of channels.
6. The method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer according to claim 3, characterized in that: In step S2, the image block projection module is configured to perform the following actions: S2.
1. Divide the preprocessed image into N image blocks of fixed size: The size of each image block is P, the step length is S, and the number of image blocks N is as shown in Formula 2: When the step length S is smaller than the image block size P, overlapping areas will appear between the image blocks, allowing the model to capture local details and subtle features; S2.
2. Map each image block to the D-dimensional feature space through linear projection to form an image block embedding And add a learnable class label x cls As a global feature representation to preserve global information; S2.3, further introduce the learnable position encoding P E , retain the spatial position information of the image block and add the camera information to embed C id , in order to reduce the impact of different camera angles on feature representation, the final embedded data is shown in Formula 3: AND input {x class ;AND i }+P E +λ cm C id (3) Among them, P E represents the position embedding, C id represents the camera information embedding, λ cm is a hyperparameter used to control the camera embedding weights, {x class ; E i } represents the combination of the learnable image category label and the feature embedding vector corresponding to the image patch.
7. The method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer according to claim 3, characterized in that: The self-attention feature extraction module in step S3 is configured to perform the following steps: S3.
1. Input the embedded data into several Transformer encoding layers of the ViT backbone network, use the pre-trained weights to accelerate model convergence and improve feature expression capabilities; S3.
2. After the embedded data is processed by the Transformer encoding layer, the encoder output features are obtained The output features include global features and local features Respectively represent the semantic information of the overall image and the detailed information of the image block; S3.3, local feature f part Divide into K groups, each group size is (N / / K)×D, denoted as f part1 ,f part2 ,…,f partk ; S3.4, the global feature f gb It is concatenated with each group of local features to form a local visual semantic feature f vp , expressed as: f vp =[[f gb ,f part1 ],[f gb ,f part2 ],...[f gb ,f partk ]]; S3.5, the fused feature f vp and the global feature f gb They are then input into the subsequent prediction module to achieve accurate classification and matching of pedestrian identities.
8. The method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer according to claim 1, characterized in that: In step S4, the cross entropy loss function of label smoothing is as follows: Among them, B represents the training batch size, C represents the number of categories, and y i,c is the true label of sample i, is the probability distribution predicted by the model, α is the label smoothing coefficient.
9. The method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer according to claim 8, characterized in that: In step S4, the triple loss function is as follows: Where B represents the training batch size, p represents positive samples, n represents negative samples, d(·,·) is the distance metric between samples, and β is a hyperparameter used to control the margin.
10. The method for occluded pedestrian re-identification based on random occlusion enhancement and Transformer according to claim 9, characterized in that: In step S4, the joint loss function is as follows: Among them, f gb is a global feature, is the local visual semantic feature, and K is the number of groups of the local visual semantic feature.
Citation Information
Cited By
Pedestrian re-identification method, apparatus and device, and medium
CN121281095A