A method and system for occluded pedestrian re-identification based on node and trajectory transformation network

CN118537888BActive Publication Date: 2026-09-25ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410443573.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-13
Publication Date
2026-09-25
Estimated Expiration
2044-04-13

AI Technical Summary

Technical Problem

虽然外部辅助模型在一定程度上提高了特征提取精度,但是所依赖的外部辅助模型可能会受域差异问题导致特征预测错误,图像中的行人经常被前景物体或其他拥挤的人群遮挡,会导致外部模型的匹配效率降低,尤其是在拥挤人群遮挡时,甚至会造成关节点的预测错误

Benefits of technology

[0045]1.本发明设计了一个骨架图构建网络(SGM),利用图卷积神经网络,通过学习行人局部特征之间的拓扑关系,自适应地聚合节点的结构和轨迹,构造行人的骨架特征图。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118537888B_ABST
    Figure CN118537888B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on node and trajectory conversion network's occluded pedestrian re-identification method and system, belong to computer vision technical field.The model is constructed using node trajectory to carry out occluded pedestrian re-identification, first pose estimation feature extraction is used, the semantic feature of pedestrian is obtained, then the powerful patch feature extraction characteristics of visual converter are used to extract the patch feature of picture, in pose estimation global feature matching, the patch feature of picture is used, the semantic feature of the pose of the pedestrian in image is efficiently matched, and the key pedestrian semantic feature is enhanced by fusion.Then a directed graph pedestrian skeleton is constructed, and the edges of each joint are adaptively adjusted.Then use the pose guided skeleton feature converter model constructed by node trajectory, according to the confidence of joint, the joint components of the constructed skeleton graph are reasonably used to make up the occluded object or noise suppressed pedestrian part, so as to realize the accurate matching of pedestrian.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and more specifically to a method and system for re-identifying occluded pedestrians based on node and trajectory transformation networks. Background Technology

[0002] Occluded pedestrian re-identification is a pedestrian retrieval task that matches occluded pedestrian images with complete pedestrian images. Most recognition methods utilize semantic cues provided by external models to align the visible parts of the pedestrian in the feature space to match the pedestrian. This typically leads to two problems: firstly, by discarding features of the occluded parts to achieve matching of the visible body, some semantic features of the occluded areas are lost. Secondly, in heavily occluded areas, such as crowded places, incorrect features are provided, contaminating the overall pedestrian features.

[0003] Currently, mainstream algorithms for solving the problem of occluded pedestrian re-identification can be broadly divided into two categories: one is image feature map stripe segmentation algorithms, which manually cut the feature map and typically employ complex computational modules to extract useful key information about pedestrians; the other is algorithms that use external cues, introducing external auxiliary models and utilizing algorithms such as pose estimation, probability, and human body parsing. Accurate pose estimation provides pedestrian pose information and is widely used for alignment guidance in occluded pedestrian re-identification; human body parsing provides pixel-level partial predictions with semantic correspondence, which can be used to match specific parts of pedestrians without being affected by the overall appearance. Although external auxiliary models improve feature extraction accuracy to some extent, the external auxiliary models may suffer from domain differences, leading to feature prediction errors. Pedestrians in images are often occluded by foreground objects or other crowded people, which reduces the matching efficiency of external models, especially when occluded by crowded people, and may even cause errors in keypoint prediction.

[0004] Therefore, proposing a method and system for re-identifying occluded pedestrians based on node and trajectory transformation networks to effectively eliminate obstacle interference and achieve accurate pedestrian matching is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a method and system for re-identifying occluded pedestrians based on node and trajectory transformation networks. The method applies the node trajectory to construct a transformer model for re-identifying occluded pedestrians, which effectively solves the problem of mismatched pedestrian poses.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] On one hand, this invention discloses an occluded pedestrian re-identification method based on node and trajectory transformation networks, comprising the following steps:

[0008] Acquire pedestrian images;

[0009] A node trajectory construction model is constructed and trained, comprising a Partial Feature Extraction Network (PFE), a Skeleton Graph Construction Network (SGM), and a Node Trajectory Construction Network (NTC) connected sequentially. The NTC model uses a Vision Transformer to extract pedestrian image features and a pose estimation model to extract semantic features, matching and fusing the two to activate key local pedestrian features. A graph convolutional neural network is then used to construct a skeleton graph, using the local pedestrian features as graph nodes. The structure and trajectory of the skeleton graph nodes, along with the semantic features of the pedestrian image, are used to aggregate pedestrian features and generate a new skeleton graph. The features of the occluded pedestrian portion are then corrected based on the node confidence scores of the visible and occluded portions.

[0010] The pedestrian image is input into the trained node trajectory to construct a model for pedestrian re-identification.

[0011] Preferably, the partial feature extraction network (PFE) includes a feature extraction network and a pose estimation network;

[0012] The feature extraction network extracts features from the pedestrian image based on Vit, obtaining global features f. global and some features f part ;

[0013] The pose estimation network introduces a key point prediction model, and uses the key point prediction model to obtain the pedestrian's key point heat map H.

[0014] The aforementioned partial feature f part The enhanced local feature f is obtained by matching the joint heatmap H with the joint heatmap. local .

[0015] Preferably, the loss function of the partial feature extraction network (PFE) for:

[0016]

[0017]

[0018]

[0019] in, It is a probability function. These are the identity loss function and the triplet loss function, respectively.

[0020] Preferably, the skeleton graph construction network (SGM) constructs a skeleton graph based on graph convolution, including:

[0021] Using the global feature f global and the local feature f local The edge weights between nodes are dynamically updated based on their differences, as shown in the following formula:

[0022]

[0023]

[0024] Where bn(·) is the batch normalization operation, abs(·) is the absolute value function, repeat(·) is the dimension copy operation, and fc(·) represents a fully connected layer. is the edge information propagated from node i to node j, and A is a user-defined initial adjacency matrix. V n It is the output of the convolutional layer, i.e., the skeleton node features.

[0025] Preferably, the pose estimation network also generates joint confidence scores, and the local feature set is calculated based on the joint confidence scores. Divided into high-confidence feature sets and low confidence feature set

[0026] Preferably, the loss function of the skeleton graph construction network (SGM) for:

[0027]

[0028] Where k represents the number of key points for the pedestrian.

[0029] Preferably, the Node Trajectory Construction Network (NTC) is implemented based on the Transformer model and includes:

[0030] The skeleton node features are used as input to the Node Trajectory Construction Network (NTC), and a new skeleton map is obtained using multiple attention heads.

[0031] Generate a skeleton feature map f based on the new skeleton map. s The formula is as follows:

[0032]

[0033] in is the feature of the i-th node in the t-th new skeleton image, k represents the number of joints, and B is the number of similar pedestrian images;

[0034] Using the skeleton feature map f sThe node features in the low-confidence feature set are corrected to obtain the corrected skeleton feature map S.

[0035] Preferably, the loss function of the Node Trajectory Construction Network (NTC) for:

[0036]

[0037] Preferably, the loss function of the node trajectory construction model is:

[0038]

[0039] Where λ is a hyperparameter.

[0040] On the other hand, the present invention also discloses an occluded pedestrian re-identification system based on node and trajectory transformation networks, comprising:

[0041] The image acquisition module is used to acquire pedestrian images;

[0042] The model building module is used to build and train a node trajectory construction model, which includes a Partial Feature Extraction Network (PFE), a Skeleton Graph Construction Network (SGM), and a Node Trajectory Construction Network (NTC) connected in sequence.

[0043] The recognition module is used to input the pedestrian image into a pre-trained node trajectory to build a model and perform pedestrian re-recognition.

[0044] As can be seen from the above technical solutions, the present invention discloses a method and system for re-identifying occluded pedestrians based on node and trajectory transformation networks, which has the following advantages compared with the prior art:

[0045] 1. This invention designs a skeleton graph construction network (SGM), which uses graph convolutional neural networks to learn the topological relationships between local features of pedestrians, adaptively aggregates the structure and trajectory of nodes, and constructs the skeleton feature map of pedestrians.

[0046] 2. This invention utilizes a Node Trajectory Network (NTC) to diffuse feature information from the visible portion of a similar pedestrian skeleton map to the occluded area, intentionally weakening the features of the occluded area and reconstructing the local pedestrian features in the occluded area to obtain a complete pedestrian skeleton feature map. Specifically, the visible portion and the occluded area are separated based on the joint confidence scores of the visible and occluded portions of the pedestrian, and the occluded area or contaminated semantic features are reconstructed using the corresponding node features in the skeleton map to achieve accurate pedestrian matching. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0048] Figure 1 This is a schematic diagram of the node trajectory construction model structure of the present invention;

[0049] Figure 2 This is a schematic diagram of the partial feature extraction network (PFE) structure of the present invention;

[0050] Figure 3 This is a schematic diagram of the skeleton graph construction network (SGM) structure of the present invention;

[0051] Figure 4 This is a schematic diagram of the node trajectory construction network (NTC) structure of the present invention;

[0052] Figure 5(a) shows the effect of different λ values ​​on model performance, and Figure 5(b) shows the effect of different γ values ​​on model performance.

[0053] Figure 6 A comparison chart showing the results of using the TransReID algorithm and the method of this invention;

[0054] Figure 7(a) is a heat map obtained by processing a whole image of a person (unobstructed image) and a pedestrian image obscured by objects using the method proposed in this invention; Figure 7(b) is a heat map obtained by processing a whole image of a person (unobstructed image) and a pedestrian image obscured by other pedestrians using the method proposed in this invention. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] On one hand, embodiments of the present invention disclose a method for re-identifying occluded pedestrians based on node and trajectory transformation networks, comprising the following steps:

[0057] S1. Obtain pedestrian images.

[0058] S2. Construct and train the node trajectory construction model, the overall structure of which is as follows: Figure 1As shown, the node trajectory construction model consists of a Partial Feature Extraction Network (PFE), a Skeleton Graph Construction Network (SGM), and a Node Trajectory Construction Network (NTC) connected sequentially. The NTC model uses a Vision Transformer to extract pedestrian image features and a pose estimation model to extract semantic features, matching and fusing the two to activate key local pedestrian features. It then uses a graph convolutional neural network to construct a skeleton graph, using the local pedestrian features as graph nodes. The model aggregates pedestrian features using the structure and trajectory of the skeleton graph nodes and the semantic features of the pedestrian image, generating a new skeleton graph. Finally, it corrects the features of the occluded parts of the pedestrian based on the node confidence scores of the visible and occluded parts.

[0059] S21. Partial Feature Extraction Network (PFE) structure as follows Figure 2 As shown, it includes a feature extraction network and a pose estimation network. Figure 2 The output [class] marked with an asterisk (*) is used as the global feature f. global The features output from other image patches are used as partial features f. part ).

[0060] S211. The feature extraction network extracts features from pedestrian images based on Vit, obtaining global features f. global and some features f part .

[0061] Specifically, the input pedestrian image is transformed from batchsize×H×W×C to batchsize×N×D, i.e., the RGB image is converted into a two-dimensional patch sequence. Here, H and W are the length and width of the image, and C is the channel dimension. The pedestrian image is then processed by a separation module to obtain N fixed-size blocks. N can be calculated using the following formula:

[0062]

[0063] Where S represents the step size for dividing the image region, and P represents the length of each divided image block. It is the floor function.

[0064] The segmented images are used as sequence input to the encoder to obtain the embedding patch. Then mark a learnable [class] as X class To preserve positional information during the pre-patch embedding process, this embodiment applies learnable positional encoding, ultimately yielding... The formula is as follows:

[0065]

[0066] Where ∪ is the union operation, P EIt is a positional encoding.

[0067] Next, the output E input It will go through m layers of Transformers to finally get the output. The obtained output f en Divided into two parts: a global feature and some features

[0068] S212. A joint prediction model is introduced into the pose estimation network, and the joint heat map H of the pedestrian is obtained by using the joint prediction model.

[0069] Using a key prediction model, a key heatmap of pedestrians after global average pooling is obtained. Where k is the number of nodes, and D is the dimension of the heatmap after the avgpooling (·) average pooling operation.

[0070] S213. Partial feature f part Matching the joint heatmap H with the local feature f yields the enhanced local feature f. local .

[0071] After obtaining the heatmap H of the joints, the pedestrian components corresponding to some features are unknown. Based on this, in this embodiment, the matching function for any two vectors a and b is defined as follows:

[0072]

[0073] Where <·> is the inner product operation, ||·|| is the L2 norm, and f(·) is the function for calculating similarity, for any A matching function can be used to find the most similar vector in H.

[0074] definition as follows:

[0075]

[0076] Where k and N are equal, Indicates in and The most similar vectors in two vector groups.

[0077] Partial regional features f part After matching with the keypoint heatmap H, two vectors with high similarity are added together to enhance key features in the pedestrian visible area. Finally, a local feature set is obtained. Then, the global and local features of the pedestrian are merged to obtain the feature set X = [f global[,F] and [·,·] are merge operations.

[0078] This embodiment also sets a binary score for the confidence of keypoints generated by the pose estimator. The score is set to 0 when the confidence of the obtained keypoints is less than a threshold γ, and set to 1 when the confidence of the obtained keypoints is greater than the threshold γ.

[0079]

[0080] Among them l i ∈{0,1}, i=1,2,...,k, corresponding to the binary score of the confidence level for each keypoint, c i This indicates the confidence score.

[0081] Using the set confidence level, the binary score l i The local feature set F is divided into two parts: the binary score of the confidence level of the corresponding key point in the original pedestrian feature map and the other part. i Features with a value of 0 form a low-confidence key feature set. Furthermore, the remaining features form a high-confidence key feature set. Where L represents the number of features with a binary confidence score of 0 at key points. and This represents the key feature corresponding to the key confidence score.

[0082] S214. In this embodiment, the cross-entropy loss function is used as the identity loss. The identity loss and triplet loss are used to calculate the difference in output features X after passing through the Partial Feature Extraction Network (PFE). It is expressed as follows:

[0083]

[0084]

[0085]

[0086] in, It is a probability function. These are the identity loss function and the triplet loss function, respectively.

[0087] S22. Skeleton Graph Construction Network (SGM) constructs a skeleton graph based on graph convolution. The network structure is as follows: Figure 3 As shown.

[0088] Graph convolution is used to calculate the correlation between local features in pedestrian images and to construct a skeleton map of the pedestrians. Specifically, the graph convolutional layer is input with two values: a custom initial adjacency matrix A and pedestrian image features X. In pedestrian images, the local features of the visible regions are enhanced. Therefore, the local features of the visible regions are closer to the global features than the invisible regions obscured by occlusions.

[0089] Using global features f global and local features f local The edge weights between nodes are dynamically updated based on their differences, as shown in the following formula:

[0090]

[0091]

[0092] Where bn(·) is the batch normalization operation, abs(·) is the absolute value function, repeat(·) is the dimension copy operation, and fc(·) represents a fully connected layer. It is the edge information propagated from node i to node j. V n The output of the convolutional layer, i.e. the skeleton node features, can be used as the input of the local features of the next graph convolutional layer to perform a recursive operation to accumulate the number of graph convolutional layers. In this embodiment, the number of graph convolutional layers J is set to 2.

[0093] S23. Node Trajectory Construction Network (NTC) is implemented based on the Transformer model, with the following structure: Figure 4 As shown, it includes:

[0094] S231. Using the skeleton node features as input to the Node Trajectory Construction Network (NTC), a new skeleton map is obtained using multiple attention heads.

[0095] Specifically, the skeleton graph construction yielded a skeleton graph with relevant node and edge relationships. The obtained skeleton node features V n As input to the Transformer, in spatial features, nodes that are close together will have similar positional features, while nodes that are farther apart will have richer positional features, thus obtaining a richer representation of node features. Queries, key-value pairs, and values ​​can be expressed using the following formula:

[0096]

[0097] Among them W Q W K , It is a linear projection, Q r,l ,K r,l V r,lThis is the parameter matrix of queries, keys, and values ​​used in the r-th head of the transformer layer for constructing the trajectory of the l-th human node. The weights of each attention head are calculated based on the dot product similarity between each query and key as follows:

[0098]

[0099] in It represents the normalized relation weight between the i-th node and the j-th node in the r-th head of the l-th layer. It is a scaling factor for the scaled dot product similarity. Each parameter of the attention matrix can indirectly represent Q. r,l ,K r,l The connection between them, and and V r,l The cross product extends the attention mechanism to multifaceted relationship learning of graph nodes. Because the input pedestrian skeleton graphs are similar in feature space, it can simultaneously capture the structural and behavioral relationships between adjacent and non-adjacent body component nodes from similar images. By leveraging multiple heads to collaboratively focus on node relationships in different feature subspaces, it captures the diversity of pedestrian poses and integrates more key node features into the final node representation.

[0100]

[0101] Then, the attention matrices of different heads are merged to obtain the trajectory of the l-th human node, which is used to construct the output transformation matrix of the transformer layer. For clarity, this embodiment uses This represents the feature of the i-th node learned in different heads in layer l. The pedestrian node features, aggregating different node relationships in the feature subspace, are output by the following feedforward network (FFN) with residual connections and batch normalization:

[0102]

[0103]

[0104] Where bn(·) represents the batch normalization operation. σ(·) represents the learnable parameters of the feedforward network, and σ(·) represents the ReLU activation function. These represent the intermediate layer output and the next layer output of the l-th layer of the human node trajectory construction transformer, respectively. The resulting node representations are then merged to obtain a skeleton graph obtained by the Transformer.

[0105] S232. Generate skeleton feature map f based on the new skeleton map. s .

[0106] The node features in each skeleton graph are averaged to obtain the corresponding graph representation. Then, the consecutive graph representations are integrated into the final feature-level graph, resulting in a new skeleton feature map f. s The formula is as follows:

[0107]

[0108] in This represents the skeleton feature map after the human node trajectory is transformed by the transformer. The feature of the i-th node in the t-th skeleton graph is used to average the features of k nodes and their components. After assembling these k nodes into a complete skeleton graph, averaging the features across the set yields a final, information-rich skeleton feature map f. s B is the number of similar pedestrian images.

[0109] S233. Utilizing skeleton feature map f s The node features in the low-confidence feature set are corrected to obtain the corrected skeleton feature map S.

[0110] The binary confidence score is used to determine the occluded or noise-suppressed parts of the pedestrian, and then the newly generated skeleton feature map f is selected. s The corresponding joints are used to correct the occluded pedestrian components.

[0111] Because noise inevitably propagates during information transmission, we retain the original visible features of the pedestrian and use the newly generated skeleton feature map f. s This is used to correct pedestrian keypoint features with a 2D confidence score of 0. The corrected skeleton feature map is as follows. Each of these elements can be described by the following formula.

[0112]

[0113] S234. In order to better selectively separate and block noise from non-human parts, this embodiment designs a skeleton construction loss function. The formula is as follows:

[0114]

[0115] Where k represents the number of key points of a pedestrian, the motivation for this loss is obvious: visible and invisible parts of the human body should not have strong similarity, the purpose of which is to separate the visible and invisible parts of the human body.

[0116] Then, identity loss and triplet loss are used to guide pedestrians in using the newly generated skeleton feature map f. s The formula for learning the missing semantic features corresponding to the pedestrian occlusion region is as follows:

[0117]

[0118] in It is a probability function. These are the identity loss function and the triplet loss function, respectively.

[0119] S24. Model training and inference.

[0120] During the training phase, pose estimation uses a pre-trained model, and the remaining components are trained together with the overall objective loss. The overall loss function of the model is as follows:

[0121]

[0122] The hyperparameter λ is set to 0.5.

[0123] During the inference phase, given the query image and gallery image, the pedestrian component features are first obtained through the component feature extraction module. and Then, new skeleton features are obtained through skeleton construction and human node trajectory construction transformers, respectively. and The cosine distance between the query and gallery is calculated using their skeleton graph features:

[0124]

[0125] Through D(S) q ,S g The performance of re-identification is improved. By reconstructing occluded pedestrian images, this method can significantly improve image matching accuracy.

[0126] S3. Input the pedestrian image into the trained node trajectory to build a model and perform pedestrian re-identification.

[0127] The performance of this invention was also verified through experiments, as detailed below:

[0128] ① To verify the effectiveness of the method proposed in this invention, experiments were conducted on four datasets: Occluded-Duke, Occluded-reid, Market-1501, and DukeMTMC-reid. Comparisons were also made with some state-of-the-art methods, including global person re-identification methods, pose-aligned person re-identification methods, Transformer-based person re-identification methods, and occluded person re-identification methods.

[0129] 1) Dataset and Evaluation Metrics

[0130] Occluded-Duke: Consists of 15,618 training images, 2,210 occluded query images, and 17,661 library images. It is a subset of DukeMTMC-reID, with occluded images classified and some overlapping images removed.

[0131] Occluded-reid consists of 2000 images of 200 occluded individuals. Each identity has 5 full-body images and 5 images of heavily occluded individuals of different types.

[0132] Market-1501: Contains 12,936 training images for 1,501 identities and 751 identities observed from 6 camera viewpoints, 19,732 library images, and 2,228 queries.

[0133] DukeMTMC-reID contains 36,411 images representing 1,404 identities captured from 8 camera viewpoints. It includes 16,522 training images, 17,661 library images, and 2,228 queries.

[0134] Evaluation metrics: The cumulative matching feature (CMC) curve was used to evaluate the quality of different person re-identification models using Rank-1 and mean accuracy (mAP). mAP represents the average accuracy of all queries.

[0135] 2) Implementation details

[0136] Unless otherwise specified, all training and testing images were resized to 256×128. Training images were augmented using random horizontal flipping, padding, random cropping, and random erasing. The initial weights of Vit were pre-trained on ImageNet-21K and then fine-tuned on ImageNet-1K. In this embodiment, the number of segmented pedestrian images and human joints were both set to 13, and an adjacency matrix of relationships between k sets of edges was set. The hidden size D was set to 768. The batch size was set to 64, with 16 IDs and 4 images per ID. The learning rate was initialized to 0.008 with cosine decay. HRNet was pre-trained on the COCO dataset to activate pedestrian pose information.

[0137] 3) Comparison with the latest methods

[0138] 31) On the occluded pedestrian datasets Occluded-DukeMTMC and Occluded-reid, the PNTCT algorithm proposed in this embodiment was compared with the stripe segmentation algorithm, the joint alignment algorithm, the occlusion-based algorithm, and the Vit-based algorithm. The experimental results are shown in Table 1.

[0139] Table 1 shows the comparison results based on the datasets Included-DukeMTMC and Included-reid.

[0140]

[0141] When comparing with Vit-based algorithms, the image size was set to 256×128 for a fair comparison. As shown in Table 1, the PNTCT method proposed in this invention achieves the best results, with a Rank-1 accuracy of 69.0% and mAP of 61.0% on the Occluded-DukeMTMC dataset, which are 2.6 and 0.8 percentage points higher than the second-best algorithm, TransReID, respectively. On the Occluded-ReID dataset, the Rank-1 accuracy is 81.0% and the mAP is 75.7%, with the mAP being 2.4 percentage points higher than the second-best algorithm's PAT, and the Rank-1 algorithm's PAT being 0.6 percentage points lower.

[0142] 32) In order to verify the effect on the overall dataset, this embodiment conducted comparative experiments on the Market-1501 and DukeMTMC datasets, comparing the method proposed in this invention with the stripe segmentation algorithm, the key point alignment algorithm, the occlusion algorithm and the Vit algorithm. The experimental results are shown in Table 2.

[0143] Table 2 shows the comparison results based on the Market-1501 and DukeMTMC datasets.

[0144]

[0145] As can be seen, on the Market-1501 dataset, the method proposed in this invention achieves a Rank-1 accuracy of 95.3% and an mAP of 89.3%, which are 0.1 and 1.6 percentage points higher than the second-ranked algorithm TransReID, respectively; and on the DukeMTMC dataset, it achieves the best results with a Rank-1 accuracy of 91.0% and an mAP of 82.2%.

[0146] ② The effectiveness of the method proposed in this invention was verified through ablation experiments.

[0147] 1) Ablation experiments were conducted on the Occluded-Duke dataset to verify the effectiveness of the proposed algorithm by progressively adding modules of the algorithm to the baseline.

[0148] Table 3 presents the experimental results. Index 0 represents the baseline, and indices 1-4 represent additions to the baseline: PFE, PFE+NTC, and PFE+SGM+NTC, respectively. As can be seen from the table, index 1, which only adds the semantic feature extraction module, outperforms the baseline, showing a 1.8% improvement in Rank-1 accuracy and a 3.3% improvement in mAP. Index 2, PFE+SGM, significantly improves performance by +1.1% in Rank-1 accuracy and +3.1% in mAP. This indicates that the higher-order semantic relationships introduced by skeleton graph construction can bring about significant performance improvements. Index 3 shows that the transformer for node trajectory construction proposed in this invention is also effective. Comparing indices 2 and 3 reveals that the combination of SGM and NTC improves performance by +5.8% in Rank-1 accuracy and +2.7% in mAP, indicating that SGM and NTC are crucial. This achieves the optimal overall performance of the node trajectory construction model proposed in this invention.

[0149] Table 3 Ablation Experiment Table

[0150]

[0151] 2) Analysis of the number of key points and patch feature regions.

[0152] This section evaluates the effectiveness of the number of local features on the Occluded-DukeMTMC dataset. The local features are fused by weighted averaging, and the number of local features corresponds to the number of keypoints. Using the HRnet model, a total of 17 keypoints were predicted (i.e., 1 represents the nose, 2 and 3 represent the left and right eyes, 4 and 5 represent the left and right ears, 6 and 7 represent the left and right shoulders, 8 and 9 represent the left and right arms, 10 and 11 represent the left and right wrists, 12 and 13 represent the left and right hips, 14 and 15 represent the left and right knees, and 16 and 17 represent the left and right ankles). The head is represented by Part 1 = {1, 2, 3, 4, 5}, the left limb by Part 2 = {6, 8, 10}, the right limb by Part 3 = {7, 9, 11}, the left leg by Part 4 = {14, 16}, the right leg by Part 5 = {15, 17}, the torso by Part 6 = {6, 7, 8, 9, 10, 11}, the upper body by Part 7 = {12, 13}, the lower body by Part 8 = {12, 13, 14, 15, 16, 17}, upper body 1 by Part 9 = {6, 7, 8, 9}, and upper body 2 by Part 10={10,11,12,13}, these are all weighted average set of key points. In Table 1, when index k is 3, it means taking three fusion key points of the pedestrian, which are {Part1,Part6,Part8}. When index k is 6, it means taking six key points of the pedestrian, which are {Part1,Part2,Part3,Part4,Part5,Part7}. When index k is 7, it means taking seven key points of the pedestrian, which are {Part1,Part2,Part3,Part4,Part5,Part9,Part...} 10}. 13 indicates that only the head joints {Part1,6,7,8,9,10,11,12,13,14,15,16,17} are fused. 17 indicates that no discrete joint features are fused. As shown in Table 4, when there are only three joints and corresponding patch features, the semantic features are relatively concentrated, some detailed information is ignored, and the performance is not very good. As the number increases, the model in this embodiment notices more semantic feature information. The method of this invention gradually improves, but the performance reaches its best when the number is 13. When the number is 17, which corresponds to the number of joints, the head features are not fused, and the performance decreases. Therefore, it can be concluded that fusing the head region gives better results than separating it into 5 parts, because the head is usually fully visible, and separating the head information may destroy the semantic features of the head. This also proves that the integrity of the semantic features of the head region is very important in person re-identification.

[0153] Table 4 Analysis of the number of key points and patch feature regions

[0154]

[0155] 3) The effect of threshold γ.

[0156] In the pose estimation network, the confidence score of keypoints above γ is set to 1, and the confidence score of keypoints below γ is set to 0. This threshold clearly distinguishes the visible parts of pedestrians. As shown in Figure 5(b), when the value of γ is small, occlusions may be treated as visible areas of the human body, introducing noise and increasing the difficulty of recognition. When γ is large, it may lead to the loss of the pedestrian's true features. The proposed method performs best when γ = 0.5.

[0157] 4) The effect of different λ values.

[0158] In the overall loss function loss function The sum of the parameters is limited to 1. As can be seen from Figure 5(a), this embodiment tests the performance by traversing λ from 0 to 1 with a step size of 0.1. When λ is set to 0.5, the proposed method exhibits the best performance. This indicates that this loss makes a competitive contribution to the performance of the proposed method.

[0159] ③ Visual comparison.

[0160] 1) Visualize the retrieval results of the TransReID algorithm and the method (PNTCT) proposed in this invention.

[0161] Since the method proposed in this invention is based on ViT feature extraction, and the TransReID algorithm is the first pure Transformer architecture, it is used as a comparison sample for visualization results. The comparison results of the two algorithms are as follows: Figure 6 As shown, the left column contains occluded images of the query person; the first two images are obscured by objects, and the last two are obscured by crowds. The two columns on the right, separated by the dotted line, represent the top 10 images with the highest matching scores generated using TransReID and the PNTCT method proposed in this invention, respectively. Figure 7 shows that the PNTCT method can overcome occlusion and correctly identify images of people in the same row (green boxes). In contrast, TransReID is very sensitive to occlusion and returns a large number of incorrectly matched images (red boxes).

[0162] 2) As shown in Figure 7(a), from left to right, the figures show the overall image and heatmap of a complete pedestrian, the image of an occluded pedestrian, and the heatmaps generated by the PFE+SGM and PFE+SGM+NTC modules when the person is occluded by objects; Figure 7(b), from left to right, shows the overall image and heatmap of a complete pedestrian, the image of an occluded pedestrian, and the heatmaps generated by the PFE+SGM and PFE+SGM+NTC modules when the person is occluded by other pedestrians. We can observe the effectiveness of the PNTCT method on complete person images. In the case of occlusion, adding the PFE+SGM module tends to focus on the visible parts of the person, and subsequently adding the NTC module can enhance the features of the visible areas of the person.

[0163] On the other hand, embodiments of the present invention also disclose an occluded pedestrian re-identification system based on node and trajectory transformation networks, used to implement the above-mentioned occluded pedestrian re-identification method, the system comprising:

[0164] The image acquisition module is used to acquire pedestrian images;

[0165] The model building module is used to build and train the node trajectory construction model, which includes a partially feature extraction network (PFE), a skeleton graph construction network (SGM), and a node trajectory construction network (NTC) connected in sequence.

[0166] The recognition module is used to input pedestrian images into a pre-trained node trajectory model for pedestrian re-recognition.

[0167] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0168] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for re-identifying occluded pedestrians based on node and trajectory transformation networks, characterized in that, Includes the following steps: Acquire pedestrian images; A node trajectory construction model is constructed and trained, comprising a partial feature extraction network, a skeleton graph construction network, and a node trajectory construction network connected sequentially. The node trajectory construction model uses VisionTransformer to extract pedestrian image features and a pose estimation model to extract semantic features, matching and fusing the two to activate key local pedestrian features. A graph convolutional neural network is then used to construct a skeleton graph, using the local features of the pedestrian as graph nodes. The structure and trajectory of the skeleton graph nodes, along with the semantic features of the pedestrian image, are used to aggregate pedestrian features and generate a new skeleton graph. The features of the occluded portion of the pedestrian are then corrected based on the node confidence scores of the visible and occluded portions. The skeleton graph construction network constructs a skeleton graph based on graph convolution, including: Utilizing global features f global and local features f local The edge weights between nodes are dynamically updated based on their differences, as shown in the following formula: ; ; in, It is a batch normalization operation. It is an absolute value function. It is a dimension copy operation. Indicates a fully connected layer. It is a node i propagation to nodes j The edge information, where A is a custom initial adjacency matrix. For a set of local features, , It is the output of the convolutional layer, i.e., the skeleton node features; The node trajectory construction network is implemented based on the Transformer model and includes: The skeleton node features are used as input to the node trajectory construction network, and a new skeleton map is obtained using multiple attention heads. ; Generate a skeleton feature map based on the new skeleton map. The formula is as follows: ; in It is the feature of the i-th node in the t-th new skeleton graph. k B represents the number of key points, and B is the number of similar pedestrian images. Using the skeleton feature map The node features in the low-confidence feature set are corrected to obtain the corrected skeleton feature map; The pedestrian image is input into the trained node trajectory to construct a model for pedestrian re-identification.

2. The occlusion pedestrian re-identification method based on node and trajectory transformation network according to claim 1, characterized in that, The feature extraction network includes a feature extraction network and a pose estimation network; The feature extraction network extracts features from the pedestrian image based on Vit to obtain global features. f global and some features f part ; The pose estimation network introduces a key point prediction model, and uses the key point prediction model to obtain the pedestrian's key point heat map H. The aforementioned features f part The enhanced local features are obtained by matching them with the joint heatmap H. f local .

3. The occlusion pedestrian re-identification method based on node and trajectory transformation network according to claim 2, characterized in that, The loss function of the feature extraction network for: ; ; ; in, It is a probability function. These are the identity loss function and the triplet loss function, respectively.

4. The occlusion pedestrian re-identification method based on node and trajectory transformation network according to claim 3, characterized in that, The pose estimation network also generates joint confidence scores, and the local feature set is then analyzed based on these joint confidence scores. Divided into high-confidence feature sets and low confidence feature set .

5. The occlusion pedestrian re-identification method based on node and trajectory transformation network according to claim 4, characterized in that, The loss function of the skeleton graph-constructed network for: ; in k This indicates the number of key points for pedestrians.

6. The occlusion pedestrian re-identification method based on node and trajectory transformation network according to claim 5, characterized in that, The loss function of the network constructed from the node trajectories for: 。 7. The occlusion pedestrian re-identification method based on node and trajectory transformation network according to claim 6, characterized in that, The loss function for the node trajectory construction model is: ; in, This is a hyperparameter.

8. An occluded pedestrian re-identification system based on node and trajectory transformation network, characterized in that, include: The image acquisition module is used to acquire pedestrian images; The model building module is used to build and train a node trajectory construction model, which includes a partial feature extraction network, a skeleton graph construction network and a node trajectory construction network connected in sequence. The skeleton graph construction network constructs a skeleton graph based on graph convolution, including: Utilizing global features f global and local features f local The edge weights between nodes are dynamically updated based on their differences, as shown in the following formula: ; ; in, It is a batch normalization operation. It is an absolute value function. It is a dimension copy operation. Indicates a fully connected layer. It is a node i propagation to nodes j The edge information, where A is a custom initial adjacency matrix. For a set of local features, , It is the output of the convolutional layer, i.e., the skeleton node features; The node trajectory construction network is implemented based on the Transformer model and includes: The skeleton node features are used as input to the node trajectory construction network, and a new skeleton map is obtained using multiple attention heads. ; Generate a skeleton feature map based on the new skeleton map. The formula is as follows: ; in It is the feature of the i-th node in the t-th new skeleton graph. k B represents the number of key points, and B is the number of similar pedestrian images. Using the skeleton feature map The node features in the low-confidence feature set are corrected to obtain the corrected skeleton feature map; The recognition module is used to input the pedestrian image into a pre-trained node trajectory to build a model and perform pedestrian re-recognition.