Shielding pedestrian re-identification method based on attitude guidance and feature fusion
By adopting the method of pose guidance and feature fusion in the occlusion pedestrian re-identification technology, the problem of unsolidity of feature extraction in complex occlusion scenarios is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510216046.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing occlusion pedestrian re-identification technology is not robust in complex occlusion scenarios, and the modal information fusion effect is limited, resulting in a decrease in recognition accuracy.
Using a method based on pose guidance and feature fusion, a pose heat map is generated through the pose estimation network, and a multi-head cross-attention mechanism is deeply integrated with visual features to generate robust local visual semantic features.
The adaptability of the pedestrian re-identification model in local occlusion and multi-camera scenarios has been significantly improved, and the recognition accuracy and robustness have been improved.
Smart Images

Figure CN120220044A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pedestrian re-identification, and particularly relates to an occluded pedestrian re-identification method based on pose guidance and feature fusion. Background Art
[0002] Re-ID (Person Re-identification) is an important research direction in the field of computer vision, and its goal is to identify specific pedestrians in images or video sequences captured by different cameras. In recent years, with the growth of the needs of intelligent security, urban management, and public safety, pedestrian re-identification technology has become increasingly important in practical applications. However, in real-world scenarios, pedestrians are often partially occluded by other objects or pedestrians, which poses a great challenge to pedestrian re-identification.
[0003] Occluded Re-ID (Occluded Person Re-identification), as an important branch of pedestrian re-identification research, aims to solve the problem of pedestrian identification in occluded situations and was proposed by Zhou et al. [1] . Miao et al. [2] constructed the representative dataset Occluded-DukeMTMC to address occlusion problems, as Figure 1 shown. The main challenge is that occlusion may lead to the loss, deformation, or interference of the appearance features of pedestrians. The occluders can be other pedestrians, static obstacles (such as trees, walls, vehicles), etc., which significantly reduces the performance of traditional Re-ID methods. The occlusion problem disrupts the global feature extraction mechanism relied on by traditional Re-ID methods, resulting in a significant decrease in the recognition accuracy. Therefore, how to extract robust pedestrian features in an occluded environment has become a hot topic and a difficult point in current research.
[0004] Existing Re-ID methods usually assume that the whole body of the pedestrian is visible, so their performance is poor in occluded scenarios. To solve this problem, in recent years, researchers have tried to introduce external information to assist the backbone visual feature extraction network. However, the pose information heatmap output by the pose estimation network and the features output by the visual feature extraction network belong to two different modalities. How to better fuse the information of the two different modalities and make the pose information assist the backbone visual feature extraction network is the key issue we need to focus on. Miao et al. used the pose information to mark the unoccluded body parts on the spatial feature map and divided the global features into local features. Although this method is intuitive and effective, it requires strict spatial feature alignment. Gao et al. [3] adopted a graph-based model to model the topological structure by learning the correspondence of nodes or edges to further explore the visible parts, and its model graph is as Figure 2As shown, Wang et al. [4] regard the fusion of local visual features and pose features as a set similarity calculation problem and achieve it by calculating the cosine similarity between two sets. However, these methods usually adopt one-to-one matching between local visual features and key-point features, which is not accurate in practical applications because a local visual area in an image may be affected by multiple key points simultaneously. For example, the local visual features near the arm may be guided by multiple key points such as the shoulder, elbow, and wrist.
[0005] In recent years, occlusion person re-identification methods have gradually introduced more refined occlusion enhancement and feature decoupling techniques. Liu et al. [5] proposed the Fine-grained Occlusion Disentanglement Network (FODN) model, which generates diverse occlusion data through a fine-grained occlusion enhancement scheme and obtains fine-grained occlusion labels through bilinear interpolation and downsampling strategies. This model introduces an occlusion feature decoupling module that can decouple the norm and angle related to occlusion from the features and balances the importance of different human body parts through a dynamic local weight controller. Although FODN has achieved good performance on multiple benchmark datasets, this method relies on a large amount of training data and computing resources, and its adaptability in complex occlusion patterns is still limited. Especially in the case where occlusion information is severely lost, the model may have difficulty recovering the lost features.
[0006] In this context, how to design a robust, efficient, and occlusion-scene-adaptive person re-identification method has become a key issue in current research. Existing occlusion person re-identification methods still have significant deficiencies in feature extraction, information fusion, and model robustness. Specifically, the fusion of pose guidance and visual features often fails to fully exploit the complementarity between different modal information; there are technical bottlenecks in the spatial alignment and consistency modeling of local features; at the same time, the diverse occlusion patterns in complex scenes make the actual performance of existing methods difficult to meet the application requirements.
[0007] References
[0008] [1] Zhuo J, Chen Z, Lai J, et al. Occluded person re-identification [C] / / 2018 IEEE international conference on multimedia and expo (ICME). IEEE, 2018: 1-6.
[0009] [2]Miao J,Wu Y,Liu P,et al.Pose-guided feature alignment for occludedpersonre-identification[C] / / Proceedings of the IEEE / CVF internationalconference on computer vision.2019:542-551.
[0011] [3]Gao S,Wang J,Lu H,et al.Pose-guided visible part matching foroccluded personreid[C] / / Proceedings of the IEEE / CVF conference on computervision and pattern recognition.2020:11744-11752.
[0013] [4]Wang T,Liu H,Song P,et al.Pose-guided feature disentangling foroccluded person re-identificationbased on transformer[C] / / Proceedings of theAAAI conference on artificial intelligence.2022,36(3):2540-2549.
[0015] Liu W,Wang X,Tan L,et al.Learning Occlusion Disentanglement withFine-grained Localization for Occluded Person Re-identification[C] / / Proceedings of the 31st ACM International Conference on Multimedia.2023:6462-6471. Summary of the Invention
[0016] In view of the problems in the prior art that the feature extraction of occluded person re-identification technology is not robust in complex occlusion scenarios and the effect of modality information fusion is limited, the present invention provides an occluded person re-identification method based on pose guidance and feature fusion. By making full use of the complementarity between pose estimation and visual features, as well as an innovative feature fusion strategy, the problems of feature loss and interference in occlusion environments are solved, thereby significantly improving the performance and adaptability of person re-identification.
[0017] To solve the above technical problems, the present invention provides the following technical solutions: An occluded person re-identification method based on pose guidance and feature fusion, comprising the following steps:
[0018] S1. Construct a preprocessing module to normalize and adjust the size of a number of pedestrian color images;
[0019] S2. Construct a pose estimation module to extract key point information from a number of pedestrian color images, obtain corresponding pose heatmaps, and generate pose features and heatmap features;
[0020] S3. Construct a feature extraction module to divide the images output by the preprocessing module into blocks, and map them to a high-dimensional feature space through a linear projection layer. At the same time, add position embedding and camera information embedding. After obtaining the fused features, input them into a multi-layer Transformer encoder, and use the self-attention mechanism to extract occlusion-robust visual features, including global features and local features. Then, perform feature fusion on the global features and local features to obtain local visual semantic features;
[0021] S4. Construct a feature fusion module to fuse the heatmap features and visual features to obtain a set of local semantic features, fuse the local visual semantic features and pose features to obtain a set of pose information fusion, and then obtain a set of local features based on the set of local semantic features and the set of pose information fusion;
[0022] S5. Construct a prediction to predict the pedestrian identity based on the global features, local visual semantic features, and the set of local features; S6. Construct an occluded person re-identification network based on the preprocessing module, pose estimation module, feature extraction module, feature fusion module, and prediction module. Use a number of pedestrian color images as input and the corresponding pedestrian identity recognition results as output to train the occluded person re-identification network to obtain an occluded person re-identification model. During the training process, use a joint loss function based on a classification loss function, a metric loss function, and a pose constraint loss function to optimize the learning of the global features, local features, and pose features;
[0023] S7. Use the occluded person re-identification model to identify pedestrian color images to obtain the recognition results of the pedestrian identity.
[0024] Further, the aforementioned step S1 is specifically as follows: For the pedestrian image perform data augmentation, including random horizontal flipping, padding, random cropping, and random erasing, and then adjust it to a unified size; where, I i represents the i-th pedestrian image, and y i respectively represent the identities corresponding to I i , and there are n1 pedestrian images in total.
[0025] Further, the aforementioned step S2 for constructing the pose estimation module is configured to perform the following steps:
[0026] S2.1. Heatmap generation and downsampling: Extract N key points of the input image to generate the corresponding pose heatmap H = [h1, h2,... h N , and downsample the heatmap to the size (H / 4) × (W / 4)(H / 4), where H and W are the height and width of the input image, and h i represents the spatial distribution and confidence score of the key point i;
[0027] S2.2. Confidence filtering and label assignment: Assign a binary label to each key point i, and the label assignment rule is as follows:
[0028]
[0029] Among them, the maximum response value of each heatmap h i represents the confidence score C ij of the corresponding key point, and the threshold γ is used to distinguish high-confidence and low-confidence key points;
[0030] S2.3. Extract pose features and heatmap features: Input the pose heatmap H = [h1, h2,... h N into the fully connected layer to generate the pose feature f pose , as shown in the following formula:
[0031] f pose = Φ fc (H) (2)
[0032] Among them, Φ fc represents the fully connected layer mapping function, f pose is the pose feature, and the pose feature is used to describe the global structure information of capturing human key points;
[0033] Input the pose heatmap H = [h1, h2,... h N into the average pooling layer to generate the heatmap feature f heat , as shown in the following formula:
[0034]
[0035] Among them, H and W are the height and width of the heat map, respectively, and H ij represents the pixel value at position (i, j) in the heat map.
[0036] Furthermore, in the aforementioned step S3, a feature extraction module is constructed based on the ViT backbone network, and the feature extraction module is configured to perform the following actions:
[0037] S3.1. Perform a chunking process on the pedestrian image, dividing it into N image chunks of a fixed size The size of each image chunk is P, and the stride is S. When the stride S is less than P, there will be overlapping regions between the generated image chunks, enabling the model to capture local details and subtle features; among them, the number of N is expressed as:
[0038]
[0039] Map each image chunk to a D-dimensional feature space through linear projection to form patch embeddings To retain global information, add a learnable class token x cls As the global feature representation, at the same time, in order to retain the spatial position information of the image chunks, introduce a learnable position encoding P E , and add a camera information embedding C id , to reduce the influence of different camera angles, and the final input sequence is expressed as:
[0040] E input ={x class ; E i}+P E +λ cm C id (5)
[0041] Among them, P E is the position embedding, C id is the camera information embedding, and λ cm is a hyperparameter for controlling the weight of the camera embedding;
[0042] S3.2. Multiple Transformer encoding layers of the ViT backbone network extract features from the input embedding data, use the pre-trained weights to accelerate the model convergence, and output visual features containing global features and local features representing the semantic information of the overall image and the detail information of the image chunks respectively;
[0043] S3.3. Divide the local feature f part into K groups, with each group having a size of (N / / K)×D, and is expressed as:
[0044] f part1 , f part2 ,..., f partk , concatenate the global feature f gb with each group of local features respectively to form the local visual semantic feature f vp , which is expressed as:
[0045] f vp = [[f gb , f part1 , [f gb , f part2 ,... [f gb , f partk (6)
[0046] Furthermore, the aforementioned feature fusion module in step S4 is configured to execute the following steps:
[0047] S4.1. Preliminary fusion of visual features and heatmap features: Multiply the visual feature f en extracted by the ViT backbone network and the heatmap feature f heat element - by - element to obtain the input of the Transformer decoder and use it as the key and value in the input of the Transformer decoder. Define a learnable visual semantic feature to learn the differences of different body parts and use it as the query in the input of the Transformer decoder to obtain k sets of local semantic feature sets PS = {PS i |i = 1, 2,... k}. The calculation formulas for the query, key, and value are as follows:
[0048]
[0049] where i = 1, 2,..., N, j = 1, 2,..., D, and the projection parameter matrix S4.2. Fusion of local visual semantic feature f vp and pose feature f pose : Input the local visual semantic feature f vp and the pose feature f pose into the pose information integration module PIIM, and perform deep fusion through the multi - head cross - attention mechanism. The calculation process is as follows:
[0050]
[0051] where i = 1, 2,..., k, j = 1, 2,..., D, and the projection parameter matrix d k is the dimension of the key vector;
[0052] Under the multi-head attention mechanism, each attention head is calculated independently, and the results are finally concatenated together, as shown in the following formula:
[0053] MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O (10)
[0054] head i = CrossAttention(QW i Q , KW i K , VW i V ) (11)
[0055] Among them, W i Q , W i K , W i V is the weight matrix for linear projection, and W O is the linear projection matrix of the output;
[0056] S4.3. The pose information fusion module PIIM outputs the fusion feature of the local visual semantic feature f vp and the pose feature f pose as follows:
[0057] PI = LayerNorm(Q + Dropout(MultiHead(Q, K, V))) (12)
[0058] Among them, PI = {PI i | i = 1, 2,..., k} represents the pose information fusion set after being fused by PIIM. The LayerNorm and Dropout operations are used to enhance the stability of the model and further normalize on the basis of the residual connection;
[0059] S4.4. Matching and fusion of local semantic features and pose information: Match the local semantic feature set PS = {PS i | i = 1, 2,... k} with the pose information fusion set PI = {PI i | i = 1, 2,..., k}, calculate the cosine similarity between each PS i and the corresponding PI, and the most similar features will be fused to obtain the final fused local feature set Specifically, as shown in the following formula:
[0060]
[0061] Among them, F v represents the finally fused local feature set;
[0062] S4.5. Classification of high-confidence and low-confidence features: Based on the heatmap labels Classify the features with heatmap label value of 1 and the features with label value of 0 to obtain the high-confidence local feature set and the low-confidence local feature set F l ={f l i |i = 1, 2,..., k - n}, where n represents the number of features with heatmap label of 1. The features with label value of 1 are high-confidence features, and the features with label value of 0 are low-confidence features.
[0063] Furthermore, in the aforementioned step S6, the classification loss function adopts the form of cross-entropy loss function, and the expression is as follows:
[0064]
[0065] Among them, B represents the training batch size, C represents the number of categories, y i,c is the true label of sample i, is the probability distribution predicted by the model.
[0066] Furthermore, in the aforementioned step S6, the metric loss function adopts the Triplet Loss loss function, and the calculation formula is:
[0067]
[0068] Among them, B represents the training batch size, p represents the positive sample, n represents the negative sample, d(·,·) is the distance metric between samples, and β is a hyperparameter used to control the margin.
[0069] Furthermore, in the aforementioned step S6, the pose constraint loss function is as follows:
[0070]
[0071] Among them, B is the size of the training batch, AP represents the average pooling operation, f h is the high-confidence local feature set, f l is the low-confidence local feature set.
[0072] Furthermore, in the aforementioned step S6, the joint loss function is as follows:
[0073]
[0074] Among them, f gb is the global feature, is a local visual semantic feature, and K is the number of groups of local visual semantic features.
[0075] Compared with the prior art, the beneficial technical effects of the present invention adopting the above technical solutions are as follows: The present invention proposes an occluded pedestrian re-identification method based on pose guidance and feature fusion, aiming to improve the robustness and recognition accuracy of the pedestrian re-identification model in complex occlusion scenarios. By combining pose guidance information and a multi-modal feature fusion strategy, the present invention significantly improves the adaptability of the model in local occlusion and multi-camera scenarios.
[0076] The present invention designs a model with Vision Transformer as the backbone network, generates a pose heatmap through a pose estimation network, and uses key point information to mark the unoccluded areas of pedestrians. The feature fusion module deeply combines visual features and pose features through a multi-head cross-attention mechanism, enabling the model to capture information of key parts at the local level, while focusing on unoccluded areas at the global level and suppressing the interference of occlusion noise. Through this feature fusion method, the generated fusion features have stronger expression ability and higher robustness.
[0077] In addition, the present invention adopts a camera information embedding strategy to incorporate the angle information of different cameras into the feature extraction process, effectively reducing the interference of cross-camera scenarios on the consistency of pedestrian feature representation. At the same time, by combining classification loss and triplet loss, the present invention further optimizes the distribution of sample features, making samples of the same identity cluster more closely in the feature space, the intervals between samples of different identities are clearer, and a pose-guided separation loss is introduced to enhance the role of pose information in the model.
[0078] Experimental results show that the Rank-1, Rank-5, Rank-10, and mean average precision (mAP) of the present invention on the Occluded-DukeMTMC dataset reach 68.9%, 82.6%, 86.9%, and 61.6% respectively, significantly superior to existing occluded pedestrian re-identification methods. Especially in occlusion scenarios, the method of the present invention demonstrates stronger feature extraction ability and higher recognition accuracy. Through the innovative design of pose guidance and feature fusion, the present invention provides a robust and efficient solution for the occluded pedestrian re-identification task and has broad application prospects in practical scenarios. Brief Description of the Drawings
[0079] Figure 1 is a schematic diagram of the occlusion situation of 2 different pedestrians in the Occluded-DukeMTMC dataset.
[0080] Figure 2It is a framework diagram of a graph matching-based ReID model proposed in "Pose-guided Visible Part Matching for Occluded Person ReID".
[0081] Figure 3 It is a network framework diagram of the occluded person re-identification method based on pose guidance and feature fusion of the present invention.
[0082] Figure 4 It is a flowchart of the training stage of the present invention. Specific implementation manners
[0083] To better understand the technical content of the present invention, specific embodiments are hereby given and described in conjunction with the accompanying drawings as follows.
[0084] In the present invention, various aspects of the present invention are described with reference to the accompanying drawings, in which many illustrative embodiments are shown. The embodiments of the present invention are not limited to those described in the drawings. It should be understood that the present invention can be implemented by any one of the various concepts and embodiments introduced above, as well as the concepts and implementation manners described in detail below, because the concepts and embodiments disclosed in the present invention are not limited to any implementation manner. In addition, some aspects disclosed in the present invention can be used alone, or in any suitable combination with other aspects disclosed in the present invention.
[0085] Refer to Figure 3 , the present invention provides an occluded person re-identification method based on pose guidance and feature fusion, including the following steps:
[0086] S1. Construct a preprocessing module to normalize and adjust the size of several pedestrian color images;
[0087] S2. Construct a pose estimation module to extract key point information from several pedestrian color images, obtain corresponding pose heatmaps, and generate pose features and heatmap features;
[0088] S3. Construct a feature extraction module to divide the images output by the preprocessing module into blocks, and map them to a high-dimensional feature space through a linear projection layer. At the same time, position embedding and camera information embedding are added. After obtaining the fused features, input them into a multi-layer Transformer encoder, and use the self-attention mechanism to extract occlusion-robust visual features, including global features and local features. Then, perform feature fusion on the global features and local features to obtain local visual semantic features;
[0089] S4. Construct a feature fusion module to fuse the heatmap features and visual features to obtain a set of local semantic features, fuse the local visual semantic features and pose features to obtain a set of pose information fusion, and then obtain a set of local features based on the set of local semantic features and the set of pose information fusion;
[0090] S5. Build a prediction. Based on the global features, local visual semantic features, and local feature set, perform pedestrian identity prediction; S6. Based on the preprocessing module, pose estimation module, feature extraction module, feature fusion module, and prediction module, build an occluded pedestrian re-identification network with pose guidance and feature fusion. Use several pedestrian color images as input and the corresponding pedestrian identity recognition results as output to train the occluded pedestrian re-identification network, obtain the occluded pedestrian re-identification model, and use the joint loss function based on the classification loss function, metric loss function, and pose constraint loss function to optimize the learning of global features, local features, and pose features during the training process;
[0091] S7. Use the occluded pedestrian re-identification model to identify the pedestrian color images and obtain the recognition results of the pedestrian identities.
[0092] In this embodiment, Vision Transformer (ViT) is used as the backbone network to build a pose estimation module for generating pose features. The sizes of the training and test images are both adjusted to 256×128. The training images are processed by data augmentation methods such as random horizontal flipping, padding, and random cropping. At the same time, camera information embedding is introduced to improve the recognition performance in cross-camera scenarios. The initial weights of the ViT network are pre-trained on the ImageNet-21K dataset and then fine-tuned on ImageNet1K to accelerate the convergence speed and improve the recognition accuracy. In the feature fusion module, a multi-head cross-attention mechanism is used to deeply fuse visual features and pose features.
[0093] The training process of this embodiment adopts the following settings: the batch size is 64, and each ID contains 4 images; the initial learning rate is set to 0.008, and the cosine annealing learning rate decay strategy is adopted; the number of split groups K is set to 17 for the segmentation of local visual semantic features; the sliding stride of the image patches is set to [11,11]. All experiments are implemented under the PyTorch 1.8 deep learning framework and accelerated on the NVIDIA 4090 GPU.
[0094] This embodiment uses the Occluded-Duke dataset to verify the occluded pedestrian re-identification task.
[0095] The Occluded-Duke dataset is a pedestrian re-identification dataset with rich occlusion scenarios, which can truly reflect the performance of the model under complex occlusion conditions. The experiment adopts the single-frame query mode, and the main performance evaluation indicators include the Rank-1 accuracy and the mean average precision (mAP). Each image in the Occluded-DukeMTMC dataset consists of three primary colors: red, green, and blue, with three channels, and each channel matches the corresponding primary color.
[0096] As a preferred embodiment of the present invention, in step S1, there are n1 pedestrian images in total, and the image samples are represented as where I i represents the i-th pedestrian image, and y i respectively represent the identities corresponding to I i . To adapt to the network requirements, the training and test images are uniformly adjusted to a size of 256×128; data augmentation is performed on the training images, including random horizontal flipping, padding, random cropping, and random erasing, to enrich data diversity and improve the generalization ability of the model. Consistent with the above description, the present invention takes the input of B samples {I i , y i} as an example to introduce the working principle of the present invention during the training process.
[0097] The present invention uses the high-resolution network HRNet to construct a pose estimation module. HRNet has the ability to fuse multi-resolution features, can generate high-precision pose heatmaps, and shows high robustness especially in occlusion scenarios. To improve the model initialization performance, HRNet uses the weights pre-trained on the COCO dataset as the initialization parameters of the present invention.
[0098] As a preferred embodiment of the present invention, step S2 for constructing the pose estimation module is configured to perform the following steps:
[0099] S2.1. Heatmap generation and downsampling: Extract N key points of the input image to generate the corresponding pose heatmap H = [h1, h2,... h N , and downsample the heatmap to a size of (H / 4)×(W / 4)(H / 4), where H and W are the height and width of the input image, and h i represents the spatial distribution and confidence score of key point i. HRNet generates pose heatmaps through multi-resolution feature fusion, where the high-resolution branch captures local details and the low-resolution branch provides global context information, and finally outputs high-precision key point heatmaps. Downsampling retains the main features of the key point confidence distribution while reducing the computational cost.
[0100] S2.2. Confidence filtering and label assignment: Different from the existing methods, the present invention does not directly set the low-confidence heatmaps to zero, but assigns a binary label to each key point i. The label assignment rule is as follows:
[0101]
[0102] where, the maximum response value of each heatmap h i represents the confidence score C ij, the threshold γ is used to distinguish high-confidence and low-confidence key points;
[0103] S2.3. Extract pose features and heatmap features: Input the pose heatmap H = [h1, h2,... h N into the fully connected layer to generate pose features f pose , as follows:
[0104] f pose = Φ fc (H) (2)
[0105] where Φ fc represents the fully connected layer mapping function, and f pose is the pose feature, which is used to describe the global structural information of the captured human key points;
[0106] Input the pose heatmap H = [h1, h2,... h N into the average pooling layer to generate heatmap features f heat , and the heatmap features are used to describe the pixel-level distribution information of local key points, as follows:
[0107]
[0108] where H and W are the height and width of the heatmap respectively, and H ij represents the pixel value at position (i, j) in the heatmap.
[0109] Reference Figure 4 , as a preferred embodiment of the present invention, in step S3, the present invention uses the backbone network VisionTransformer to extract features. Using the ViT Base pre-trained weights, which are pre-trained on ImageNet21K and then fine-tuned on ImageNet-1K, can accelerate the convergence speed of network training and the final prediction accuracy, and improve the feature expression ability.
[0110] The feature extraction module is configured to perform the following actions:
[0111] S3.1. Divide the pedestrian image into blocks, divided into N fixed-size image patches (patches)
[0112] The size of each image patch is P, and the stride is S. When the stride S is less than P, there will be overlapping regions between the generated image patches, enabling the model to capture local details and subtle features; among them, the number of N is expressed as:
[0113]
[0114] Map each image patch to the D-dimensional feature space through linear projection to form patch embeddings To retain global information, a learnable class token x is added cls as the global feature representation. Meanwhile, to retain the spatial position information of the image patches, a learnable position encoding P is introduced E , and a camera information embedding C is added id to reduce the influence of different camera angles. Finally, the input sequence is represented as:
[0115] E mput ={x class ; E i}+P E +λ cm C id (5)
[0116] where P E is the position embedding, C id is the camera information embedding, and λ cm is a hyperparameter that controls the weight of the camera embedding.
[0117] Through the above processing, the embedding vectors of the image patches contain rich context information, providing a better expression basis for subsequent Transformer encoder feature extraction.
[0118] S3.2. The m Transformer encoding layers of the ViT backbone network extract features from the input embedding data, and use pre-trained weights to accelerate model convergence, outputting visual features containing global features and local features representing the semantic information of the overall image and the detailed information of the image patches respectively;
[0119] S3.3. To further improve the model's ability to extract local information, the local feature f part is divided into K groups, each group with a size of (N / / K)×D, denoted as: f part1 , f part2 , …, f partk ,
[0120] The global feature f gb is concatenated with each group of local features respectively to form local visual semantic features f vp , denoted as:
[0121] f vp =[[f gb , f part1 , [f gb , f part2 ,... [f gb , f partk (6)
[0122] The global and local features are fused through a splicing operation, so that each local feature group not only contains local detailed information but also has overall semantic understanding, thereby enhancing the robustness and recognition performance of the model in occluded scenarios.
[0123] As a preferred embodiment of the present invention, the feature fusion module in step S4 is configured to perform the following steps:
[0124] S4.1. Preliminary fusion of visual features and heatmap features: Multiply the visual feature f en extracted by the ViT backbone network and the heatmap feature f heat element by element to obtain the input of the Transformer decoder and use it as the key and value in the input of the Transformer decoder. Define a learnable visual semantic feature to learn the differences between different body parts, use it as the query in the input of the Transformer decoder, and obtain k sets of local semantic features PS = {PS i |i = 1, 2,... k}. The calculation formulas for the query, key, and value are as follows:
[0125]
[0126] where i = 1, 2,..., N, j = 1, 2,..., D, and the projection parameter matrix S4.2. Fusion of local visual semantic features f vp and pose features f pose : Input the local visual semantic feature f vp and the pose feature f pose into the Pose-Image Information Module (PIIM). Perform deep fusion through the multi-head cross-attention mechanism, and the calculation process is as follows:
[0127]
[0128]
[0129] where i = 1, 2,..., k, j = 1, 2,..., D, and the projection parameter matrix d k is the dimension of the key vector;
[0130] Under the multi-head attention mechanism, each attention head calculates independently and finally splices the results together, as follows:
[0131] MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O (10)
[0132] head i = CrossAttention(QW i Q , KW i K , VW i V ) (11)
[0133] where W i Q , W i K , W i V is the weight matrix for linear projection, and W O is the linear projection matrix of the output;
[0134] S4.3. The Pose Information Integration Module PIIM outputs the integrated feature of the local visual semantic feature f vp and the pose feature f pose , as follows:
[0135] PI = LayerNorm(Q + Dropout(MultiHead(Q, K, V))) (12)
[0136] where PI = {PI i | i = 1, 2,..., k} represents the pose information integration set after being integrated by PIIM. The LayerNorm and Dropout operations are used to enhance the stability of the model and further normalize it on the basis of the residual connection;
[0137] S4.4. Matching and integration of local semantic features and pose information: Since each local feature is related to the information of a certain key point of the human body, the local semantic feature set PS = {PS i | i = 1, 2,... k} is matched with the pose information integration set PI = {PI i | i = 1, 2,..., k}, and the cosine similarity between each PS i and the corresponding PI is calculated. The most similar features are integrated to obtain the final integrated local feature set Specifically, the following formula:
[0138]
[0139] where F vRepresents the finally fused local feature set;
[0140] S4.5. Classification of high-confidence and low-confidence features: Based on the heatmap labels Classify the features with heatmap label value of 1 and the features with label value of 0 to obtain the high-confidence local feature set and the low-confidence local feature set F l ={f l i |i = 1, 2,..., k - n}, where n represents the number of features with heatmap label of 1. The features with label value of 1 are high-confidence features, and the features with label value of 0 are low-confidence features.
[0141] The design of the loss function is crucial, directly affecting the learning effect and final performance of the model. The present invention adopts a joint loss function to calculate the classification loss and the metric loss respectively to ensure that the model has high accuracy and robustness in the occluded pedestrian re-identification task. Therefore, in the training process of the present invention, the joint loss function based on the classification loss function, the metric loss function and the pose constraint loss function is used to optimize the learning of the global features, local features and pose features;
[0142] Classification loss (ID Loss): Used to measure the performance of the encoder in the pedestrian identity recognition task. The cross-entropy loss function is adopted, and its form is:
[0143]
[0144] where B represents the training batch size, C represents the number of classes, y i,c is the true label of sample i, is the probability distribution predicted by the model.
[0145] Metric loss: Used to optimize the similarity and difference of features to ensure that the distance between similar samples is less than the distance between dissimilar samples. The Triplet Loss loss function is adopted, and the calculation formula is:
[0146]
[0147] where B represents the training batch size, p represents the positive sample, n represents the negative sample, d(·,·) is the distance metric between samples, and β is a hyperparameter used to control the margin.
[0148] The Pose Guided Push Loss aims to enhance the role of pose information in the model by minimizing f h and f lThe cosine similarity between them is used to minimize the similarity between the features of human body parts and non - human body parts, so as to achieve the effect of making the model pay more attention to the unoccluded body features. The specific form of this loss function is as follows:
[0149]
[0150] where B is the size of the training batch, AP represents the average pooling operation, f h is the high - confidence local feature set, and f l is the low - confidence local feature set.
[0151] The final combined loss function is specifically defined as follows:
[0152]
[0153] where f gb is the global feature, is the local visual semantic feature, and K is the number of groups of local visual semantic features.
[0154] The test process of the present invention is as follows:
[0155] Step 1: Input the query set and the gallery set, and enter Step 2;
[0156] Step 2: Use the model obtained in the training process to extract features from all pedestrian images in the query set and the gallery set input in Step 1, and enter Step 3;
[0157] Step 3: Calculate the similarity between the query set features and the gallery set features, and enter Step 4;
[0158] Step 4: According to the level of similarity, obtain the matching result corresponding to each pedestrian image in the query set, and enter Step 5;
[0159] Step 5: End.
[0160] Preferably, the query set in Step 1 of the test process represents the set of pedestrian images to be recognized, and the gallery set represents the set of target pedestrian images matched by the query set.
[0161] Preferably, the similarity calculation method in Step 3 of the test process is the cosine similarity, which is used to evaluate the similarity between the query set and the gallery set image features.
[0162] Preferably, in step 4 of the test process, each query set image corresponds to several matching images in the gallery set, and the test results are evaluated using the Cumulative Matching Characteristic (CMC) and the mean Average Precision (mAP). Among them, the Rank-k accuracy of CMC measures the probability of a correct match among the top k retrieval results, while mAP reflects the average retrieval performance of the method over the entire query set.
[0163] This experiment was evaluated on the Occluded-Duke dataset, and the performance of the method of the present invention and multiple existing mainstream person re-identification methods was tested, including: global feature-based methods, such as PCB (ECCV 2018), which extract pedestrian global features through partial pooling. Pose-guided partial matching methods, such as PVPM (CVPR 2020), which guide the model to match visible parts in occlusion scenarios. Higher-order information-based methods, such as HOReID (CVPR 2020), which improve the model's adaptability to occlusion scenarios through relationship and topological structure learning. Transformer-based models, such as PAT (CVPR 2021) and TransReID (ICCV 2021), which use the Transformer framework to extract pedestrian features. Feature erasure and diffusion-based networks, such as FED (CVPR 2022), which enhance the occlusion robustness of the model through feature erasure. The comparison results are shown in Table 1, the performance comparison table of the method of the present invention and other methods on the Occluded-Duke dataset.
[0164] Table 1
[0165]
[0166] As can be seen from Table 1, the Rank-1 and mAP of the present invention on the Occluded-Duke dataset reached 68.9% and 61.6% respectively, showing significant advantages in the occluded person re-identification task. Compared with PCB, Rank-1 and mAP increased by 26.3% and 27.9% respectively; compared with PVPM, they increased by 21.9% and 23.9% respectively; compared with HOReID, they increased by 13.8% and 17.8% respectively.
[0167] In addition, compared with Transformer-based methods (such as PAT and TransReID), the method of the present invention has achieved improvements of 4.4% and 5.9% in Rank-1 and mAP respectively. This shows that the pose-guided feature fusion module combined with the multi-head cross-attention mechanism can effectively combine pose information and visual features, improve the model's ability to capture detailed information in occluded scenarios, and enhance the suppression effect on occlusion noise.
[0168] In summary, the occluded pedestrian re-identification method based on pose guidance and feature fusion proposed by the present invention fully combines the advantages of visual features and pose features, optimizes the feature distribution using the feature fusion module and the pose-guided separation loss, and significantly improves the recognition performance of the model in complex occluded scenarios. The experimental results on the Occluded-Duke dataset show that the present invention has significant advantages in the occluded pedestrian re-identification task, providing a robust and efficient solution for related fields.
[0169] This embodiment will introduce an applicable scenario of the present invention.
[0170] In scenarios such as urban blocks, subway stations, and public safety monitoring areas, the occlusion problem is a major challenge faced by pedestrian re-identification systems. For example, in high-traffic areas such as subway station entrances, shopping mall entrances, or city streets, pedestrians are frequently occluded by each other or by other objects (such as luggage, building obstacles), and the performance of traditional pedestrian re-identification models in these scenarios is often unsatisfactory.
[0171] First, apply the occluded pedestrian re-identification method based on pose guidance and feature fusion of the present invention to the pedestrian re-identification system. The present invention effectively alleviates the interference of occlusion on feature extraction by fusing visual features and pose features, and enhances the adaptability of the system in occluded scenarios. This method can make full use of pose information to guide the model to focus on unoccluded areas, thereby significantly improving the recognition accuracy.
[0172] Next, combine the method of the present invention with existing intelligent monitoring systems. By introducing occluded data from real scenarios (such as the Occluded-Duke dataset) and sample features from different camera perspectives during the training phase, the system can accurately identify pedestrian identities in various complex scenarios. Especially in the case of cross-cameras, the present invention improves the robustness of the model to the consistency of feature expressions under different perspectives by adding camera information embedding.
[0173] Finally, to cope with the complex and changeable occlusion environment in public places, the pedestrian re-identification system of the present invention supports dynamic model updating. The system can regularly retrain the model with newly collected occluded data to maintain the model's efficient adaptability to different occluded scenarios, thereby ensuring the stability and efficiency of the system in long-term applications.
[0174] The pedestrian re-identification method of this embodiment not only significantly improves the recognition effect in occlusion scenarios, but also demonstrates extensive application potential in the fields of actual public safety, intelligent transportation, etc., providing highly robust technical support for monitoring systems in complex scenarios.
[0175] Although the present invention has been described above with reference to preferred embodiments, it is not intended to limit the present invention. Those of ordinary skill in the technical field to which the present invention pertains can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the scope defined in the claims.
Claims
1. A method for re-identifying occluded pedestrians based on posture guidance and feature fusion, characterized in that: The following steps are involved: S1. Build a preprocessing module to normalize and resize several pedestrian color images; S2. Construct a posture estimation module to extract key point information from several pedestrian color images, obtain corresponding posture heat maps, and generate posture features and heat map features; S3. Construct a feature extraction module, divide the image output by the preprocessing module into blocks, and map it to a high-dimensional feature space through a linear projection layer. At the same time, add position embedding and camera information embedding, obtain fused features, and input them into a multi-layer Transformer encoder. Use the self-attention mechanism to extract occlusion-robust visual features, including global features and local features, and then perform feature fusion on the global features and local features to obtain local visual semantic features. S4, constructing a feature fusion module, fusing the heat map features and the visual features to obtain a local semantic feature set, fusing the local visual semantic features and the posture features to obtain a posture information fusion set, and then obtaining a local feature set based on the local semantic feature set and the posture information fusion set; S5, construct prediction, and perform pedestrian identity prediction based on global features, local visual semantic features, and local feature sets; S6. Based on the preprocessing module, the posture estimation module, the feature extraction module, the feature fusion module, and the prediction module, an occluded pedestrian re-identification network based on posture guidance and feature fusion is constructed. The occluded pedestrian re-identification network is trained with a number of pedestrian color images as input and the corresponding pedestrian identity recognition results as output to obtain an occluded pedestrian re-identification model. During the training process, a joint loss function based on a classification loss function, a metric loss function, and a posture constraint loss function is used to optimize the learning of global features, local features, and posture features. S7. Use the occluded pedestrian re-identification model to identify the pedestrian color image and obtain the pedestrian identity recognition result.
2. The method for re-identifying an occluded pedestrian based on posture guidance and feature fusion according to claim 1, characterized in that: Step S1 specifically includes: for pedestrian images Perform data enhancement, including random horizontal flipping, padding, random cropping, and random erasing, and then adjust to a uniform size; I i represents the i-th pedestrian image, y i I i Corresponding to the identity, there are n1 pedestrian images in total.
3. The method for re-identifying occluded pedestrians based on posture guidance and feature fusion according to claim 1, characterized in that: Step S2 constructs a posture estimation module configured to perform the following steps: S2.
1. Heatmap Generation and Downsampling: Extracting Input Images N key points of the image, generate the corresponding posture heat map H = [h1,h2,...h N ], downsample the heatmap to size (H / 4)×(W / 4)(H / 4), where H and W are the height and width of the input image, and h i Represents the spatial distribution and confidence score of key point i; S2.2, Confidence filtering and label assignment: Assign a binary label to each key point i The label assignment rules are as follows: Among them, each heat map h i The maximum response value represents the confidence score C of the corresponding key point ij , the threshold γ is used to distinguish high-confidence and low-confidence key points; S2.3, extract posture features and heat map features: extract posture heat map H = [h1,h2,...h N ] is input to the fully connected layer to generate posture features f pose , as follows: f pose =Φ fc (H) (2) Among them, Φ fc represents the fully connected layer mapping function, f pose Posture features are used to describe and capture the global structural information of key points of the human body; The posture heat map H = [h1,h2,...h N ] Input to the average pooling layer to generate the heat map feature f heat , as follows: Among them, H and W are the height and width of the heat map respectively. ij Represents the pixel value of the location (i, j) in the heat map.
4. The method for re-identifying an occluded pedestrian based on posture guidance and feature fusion according to claim 1, characterized in that: In step S3, a feature extraction module is constructed based on the ViT backbone network, and the feature extraction module is configured to perform the following actions: S3.1, the pedestrian image is divided into N fixed-size image blocks The size of each image block is P, and the step length is S. When the step length S is smaller than P, there will be overlapping areas between the generated image blocks, so that the model can capture local details and subtle features; where the number of N is expressed as: Map each image block to the D-dimensional feature space through linear projection to form a patch embedding To preserve global information, add a learnable class label x cls As a global feature representation, in order to retain the spatial position information of the image block, a learnable position encoding P is introduced E , and add the camera information to embed C id , to reduce the impact of different camera angles, the final input sequence is expressed as: AND input {x class ;AND i }+P E +λ cm C id (5) Among them, P E is the position embedding, C id is the camera information embedding, λ cm Hyperparameters for controlling camera embedding weights; S3.2, the multiple Transformer encoding layers of the ViT backbone network extract features from the input embedded data, use pre-trained weights to accelerate model convergence, and output visual features Include global features and local features Respectively represent the semantic information of the overall image and the detailed information of the image block; S3.3, local feature f part Divide into K groups, each group size is (N / / K)×D, expressed as: f part1 ,f part2 ,…,f partk , the global feature f gb It is concatenated with each group of local features to form a local visual semantic feature f vp , expressed as: f vp =[[f gb ,f part1 ],[f gb ,f part2 ],...[f gb ,f partk ]] (6) 5. The method for re-identifying an occluded pedestrian based on posture guidance and feature fusion according to claim 1, characterized in that: Step S4: The feature fusion module is configured to perform the following steps: S4.
1. Preliminary fusion of visual features and heat map features: The visual features extracted by the ViT backbone network are en and heatmap feature f heat Multiply element by element to get the Transformer decoder input And as the key and value in the Transformer decoder input, define a learnable visual semantic feature To learn the differences between different body parts, we use them as queries in the Transformer decoder input and obtain a set of k local semantic features PS = {PS i |i=1,2,...k}, the calculation formula for query, key, and value is as follows: Where i = 1, 2, ..., N, j = 1, 2, ..., D, the projection parameter matrix S4.
2. Local visual semantic features f vp With the posture feature f pose Fusion: local visual semantic features f vp With the posture feature f pose The input is sent to the posture information fusion module PIIM, and deep fusion is performed through the multi-head cross attention mechanism. The calculation process is as follows: where i = 1, 2, ..., k, j = 1, 2, ..., D, the projection parameter matrix d k is the dimension of the key vector; In the multi-head attention mechanism, each attention head is calculated independently, and the results are concatenated together in the end, as shown in the following formula: MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O (10) head i =CrossAttention(QW i Q ,KW i K ,VW i V ) (11) Among them, W i Q ,W i K ,W i V is the weight matrix for linear projection, W O is the output linear projection matrix; S4.3, the posture information fusion module PIIM outputs local visual semantic features f vp With the posture feature f pose The fusion features are as follows: PI=LayerNorm(Q+Dropout(MultiHead(Q,K,V))) (12) Where PI = {PI i |i=1,2,...,k} represents the fusion set of posture information after PIIM fusion. LayerNorm and Dropout operations are used to enhance the stability of the model and further normalize it based on the residual connection; S4.4, matching and fusion of local semantic features and posture information: The local semantic feature set PS = {PS i |i=1,2,...k} and the attitude information fusion set PI={PI i |i=1,2,...,k} to match and calculate each PS i The cosine similarity between the corresponding PIs, the most similar features will be fused to obtain the final fused local feature set The specific formula is as follows: Among them, F v Represents the final fused local feature set; S4.
5. Classification of high-confidence and low-confidence features: based on heatmap labels By classifying the features with heat map label value 1 and the features with label value 0, we can obtain high confidence local feature sets. and low confidence local feature set F l ={f l i |i=1,2,...,kn}, where n represents the number of features with the heat map label of 1. Features with a label value of 1 are high-confidence features, and features with a label value of 0 are low-confidence features.
6. The method for re-identifying occluded pedestrians based on posture guidance and feature fusion according to claim 1, characterized in that: In step S6, the classification loss function takes the form of a cross entropy loss function, which is expressed as follows: Among them, B represents the training batch size, C represents the number of categories, and y i,c is the true label of sample i, is the probability distribution predicted by the model.
7. The method for re-identifying occluded pedestrians based on posture guidance and feature fusion according to claim 6, characterized in that: In step S6, the metric loss function adopts the Triplet Loss loss function, and the calculation formula is: Where B represents the training batch size, p represents positive samples, n represents negative samples, d(·,·) is the distance metric between samples, and β is a hyperparameter used to control the margin.
8. The method for re-identifying occluded pedestrians based on posture guidance and feature fusion according to claim 7, characterized in that: In step S6, the posture constraint loss function is as follows: Where B is the size of the training batch, AP represents the average pooling operation, and f h is a high confidence local feature set, f l is a low confidence local feature set.
9. The method for re-identifying occluded pedestrians based on posture guidance and feature fusion according to claim 7, characterized in that: In step S6, the joint loss function is as follows: where f gb is a global feature, is the local visual semantic feature, and K is the number of groups of the local visual semantic feature.
Citation Information
Cited By
Cornea image death time real-time analysis system based on Transform-diffusion model
CN121053683A
Fine-grained feature decoupling pedestrian reloading re-identification method and system based on text semantic guidance
CN121170709A
Aesthetic perception posture generation method based on vision and text prompt
CN122369124A