A Pedestrian Re-identification Method and Device Based on Pose Estimation and Feature Association

Through the method based on pose estimation and feature association, local key point features and global features of pedestrian images are extracted and fused, which solves the shortcomings of local feature irrelevance and the fusion of global features and local features, and improves the accuracy of pedestrian re-identification.

CN116844187BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310713891.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2025-06-13
Estimated Expiration
2043-06-15

AI Technical Summary

Technical Problem

The existing pedestrian re-identification methods have shortcomings in local feature independence and fusion of global features and local features, which leads to insensitive to large-scale action changes and scaling, and it is difficult to effectively integrate coarse-grained global features and fine-grained local features.

Method used

Using a method based on pose estimation and feature association, local key point heat maps and global feature maps are extracted through pose estimation network and global feature extraction network, combined with foreground pose non-local network and fusion module, non-linear fusion of local key point feature sequences and global features is achieved to generate fusion features.

Benefits of technology

Effectively extract the correlation between local features of each key point, realize the effective fusion of global features and local features, improve the accuracy of pedestrian re-identification, and overcome the shortcomings of local feature independence and the fusion of global features and local features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844187B_ABST
    Figure CN116844187B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian re-identification method and device based on pose estimation and feature association. After obtaining the query pedestrian image, a local key point heat map and a global feature map are respectively extracted through a pose estimation network and a global feature extraction network. Then, the dimension-reduced global feature map is multiplied element-wise with the local key point heat map to obtain a local key point feature sequence. Next, the local key point feature sequence and the global feature map are passed through a foreground pose non-local network to obtain a local pose joint feature. An adaptive average pooling operation is performed on the global feature map, and then non-linearly fused with the local pose joint feature to obtain a fused feature. Finally, the similarity between the query pedestrian image and the images in the database is calculated based on the fused feature, and the query result is output. The present invention overcomes the problem of local feature irrelevance and the problem of global feature and local feature fusion, and improves the accuracy of pedestrian re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of pedestrian re-identification, and particularly relates to a pedestrian re-identification method and device based on pose estimation and feature association. Background Technique

[0002] Pedestrian re-identification (Person Re-ID) is an important field and research hotspot in computer vision. The goal of pedestrian re-identification is to automatically find all pedestrian images of a query object among multiple non-overlapping cameras given a query pedestrian image. In practical applications, by combining technologies such as pedestrian re-identification, pedestrian detection, and pedestrian tracking as a pedestrian re-identification system, it can be applied to fields such as video surveillance and unmanned supermarkets.

[0003] In recent years, due to the development of deep learning and the improvement of computer hardware, pedestrian re-identification has made great progress. Especially in feature extraction, deep neural networks can often extract distinguishable features. For the features extracted by the neural network, according to the size of the region, they can be divided into global features and local features. Global features can represent the overall attributes of pedestrians and have good invariance, but are easily interfered by the background. For similar appearances, it is also difficult for global features to distinguish small differences, so it is difficult to be used alone in the pedestrian re-identification task. Local features are to extract features for a certain region in the image, which contain rich fine-grained information. Many pedestrian re-identification methods are based on local features to judge pedestrian identities. However, at present, although many methods extract local features of pedestrians, due to the small correlation between these local features, they are not sensitive to large-scale motion changes and scaling. Therefore, a method of feature association is needed to associate local features and form an associated feature to judge pedestrian identities.

[0004] Although local features can represent fine-grained information, a large number of disordered and unassociated local features lack unified global guidance. Therefore, many current studies usually combine local features with global features. Although the combination of global features and local features has achieved good results, there is still a lack of an effective mechanism to effectively fuse coarse-grained global features and fine-grained local features, so that the fused features have both sufficient detailed information of local features and global guidance of global features. Summary of the Invention

[0005] The purpose of this application is to provide a pedestrian re-identification method and device based on pose estimation and feature association, which overcomes the problem of local feature irrelevance and the problem of the fusion of global features and local features.

[0006] To achieve the above purpose, the technical solution of this application is as follows:

[0007] A pedestrian re-identification method based on pose estimation and feature association, comprising:

[0008] Obtain the pedestrian image to be queried, and extract the local key-point heat map and the global feature map through a pose estimation network and a global feature extraction network respectively;

[0009] Multiply the dimension-reduced global feature map element-wise with the local key-point heat map to obtain the local key-point feature sequence, and then pass the local key-point feature sequence and the global feature map through a foreground pose non-local network to obtain the local pose joint feature;

[0010] Perform an adaptive average pooling operation on the global feature map, and then input it together with the local pose joint feature into a fusion module for non-linear fusion to obtain the fusion feature;

[0011] Calculate the similarity between the pedestrian image to be queried and the images in the database according to the fusion feature, sort them according to the similarity, and then output the query result.

[0012] Further, the step of obtaining the local pose joint feature by passing the local key-point feature sequence and the global feature map through a foreground pose non-local network includes:

[0013] Add each feature in the local key-point feature sequence to the global feature map and then perform an adaptive average pooling operation to obtain the local key-point feature sequence embedded with the global feature;

[0014] Perform a normalization operation on the local key-point feature sequence embedded with the global feature, then further mine the fine-grained information through a feed-forward neural network, and then perform a residual connection with the feature sequence embedded with the global feature to obtain the fine-grained local feature sequence;

[0015] Perform a normalization operation and a multi-head self-attention operation on the fine-grained local feature sequence, and then perform a residual connection with the fine-grained local feature sequence to obtain the local feature association sequence;

[0016] Concatenate the local feature association sequences, and then pass through a linear layer to obtain the local pose joint feature.

[0017] Further, the step of performing an adaptive average pooling operation on the global feature map, then inputting it together with the local pose joint feature into a fusion module for non-linear fusion to obtain the fusion feature includes:

[0018] Perform an adaptive average pooling operation on the global feature map to obtain the coarse-grained feature;

[0019] Input the coarse-grained feature and the local pose joint feature into their respective feed-forward neural networks, and then perform a residual connection to obtain the first feature and the second feature;

[0020] The coarse-grained features and local pose joint features are evenly cut into a preset number of blocks according to the dimension of the features, and then recombined into matrices K and Q. After that, the matrix Q and K are multiplied to obtain the feature supervision similarity, and the feature supervision similarity is used to calculate the weights of matrix K and matrix S;

[0021] The weights and matrix S are horizontally cut into a preset number of blocks, and then spliced in order and recombined into associated features;

[0022] The first feature, the second feature and the associated features are accumulated to obtain the final fused feature.

[0023] Further, for the step of accumulating the first feature, the second feature and the associated features to obtain the final fused feature, the following formula is used for accumulation:

[0024] f = αf′g + βf′ p +(1 - α - β)f s ;

[0025] where, f′ g represents the first feature, f′ p represents the second feature, f s represents the associated feature, and α, β are weight parameters.

[0026] This application also proposes a pedestrian re-identification device based on pose estimation and feature association, including:

[0027] A feature extraction module, configured to obtain an image of a pedestrian to be queried, and extract a local key point heat map and a global feature map through a pose estimation network and a global feature extraction network respectively;

[0028] A pose joint module, configured to perform element-wise multiplication on the downsampled global feature map and the local key point heat map to obtain a local key point feature sequence, and then obtain local pose joint features by passing the local key point feature sequence and the global feature map through a foreground pose non-local network;

[0029] A fusion module, configured to perform an adaptive average pooling operation on the global feature map, and then input it together with the local pose joint features into the fusion module for non-linear fusion to obtain a fused feature;

[0030] A query module, configured to calculate the similarity between the image of the pedestrian to be queried and the images in the database according to the fused feature, sort them according to the size of the similarity, and then output the query result.

[0031] Further, the step of obtaining local pose joint features by passing the local key point feature sequence and the global feature map through a foreground pose non-local network includes:

[0032] After adding each feature in the local key-point feature sequence to the global feature map and performing an adaptive average pooling operation, a local key-point feature sequence embedded with the global feature is obtained;

[0033] After performing a normalization operation on the local key-point feature sequence embedded with the global feature and further mining fine-grained information through a feed-forward neural network, and then performing a residual connection with the feature sequence embedded with the global feature, a fine-grained local feature sequence is obtained;

[0034] After performing a normalization operation and a multi-head self-attention operation on the fine-grained local feature sequence, and then performing a residual connection with the fine-grained local feature sequence, a local feature correlation sequence is obtained;

[0035] The local feature correlation sequences are concatenated and then passed through a linear layer to obtain a local pose joint feature.

[0036] Further, the adaptive average pooling operation is performed on the global feature map, and then it is input into the fusion module together with the local pose joint feature for non-linear fusion to obtain a fusion feature, including:

[0037] The adaptive average pooling operation is performed on the global feature map to obtain a coarse-grained feature;

[0038] The coarse-grained feature and the local pose joint feature are respectively input into their own feed-forward neural networks, and then a residual connection is performed to obtain a first feature and a second feature;

[0039] The coarse-grained feature and the local pose joint feature are evenly cut into a preset number of blocks according to the dimension of the feature, and then recombined into matrices K and Q. After that, the matrix Q and K are multiplied to obtain a feature supervision similarity, and the weight of matrix K and matrix S are calculated using the feature supervision similarity;

[0040] The weight and matrix S are horizontally cut into a preset number of blocks and then concatenated in order to recombine into a correlation feature;

[0041] The first feature, the second feature, and the correlation feature are added up to obtain the final fusion feature.

[0042] Further, for the adding up of the first feature, the second feature, and the correlation feature to obtain the final fusion feature, the following formula is used for adding up:

[0043] f = αf′ g + βf′ p +(1 - α - β)f s ;

[0044] where f′ g represents the first feature, f′ p represents the second feature, and f sIndicates associated features, where α and β are weight parameters.

[0045] A person re-identification method and device based on pose estimation and feature association proposed in this application can extract the associations between local features of each key point, realize the effective fusion of coarse-grained global features and fine-grained local features, overcome the problem of local feature irrelevance, as well as the problem of global feature and local feature fusion, and improve the accuracy of person re-identification. Brief Description of the Drawings

[0046] Figure 1 It is a flow chart of the person re-identification method based on pose estimation and feature association in this application.

[0047] Figure 2 It is a schematic diagram of the overall network structure of an embodiment of this application.

[0048] Figure 3 It is a schematic diagram of the foreground pose non-local module of an embodiment of this application.

[0049] Figure 4 It is a schematic diagram of the fusion module of an embodiment of this application. Detailed Embodiments

[0050] In order to make the purpose, technical solutions and advantages of this application clearer, the following further details this application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.

[0051] In one embodiment, as Figure 1 shown, a person re-identification method based on pose estimation and feature association is proposed, including:

[0052] Step S1, obtain the image of the pedestrian to be queried, and extract the local key point heat map and the global feature map through the pose estimation network and the global feature extraction network respectively.

[0053] This application uses a trained global feature extraction network and a pose estimation network to extract the global feature map and the local key point heat map. In a specific embodiment, both the global feature extraction network and the pose estimation network use the high-resolution network HRNet-W32.

[0054] When training the global feature extraction network and the pose estimation network, for the obtained pedestrian image samples, data preprocessing can also be performed on the pedestrian images by horizontal flipping and random cropping of the images to enrich the training samples. When performing data preprocessing, first adjust each pedestrian image to an image with a height of 256 and a width of 128. Given a threshold of 0.5, generate a horizontal flipping probability for each pedestrian image. When the probability is greater than the threshold, flip the left half and the right half of the image horizontally with the vertical central axis of the pedestrian image as the central axis; a random cropping probability can also be generated for each pedestrian image. First, fill the image, and then randomly crop it to the required size.

[0055] After training the global feature extraction network and the pose estimation network, for the pedestrian image to be queried, it is only necessary to adjust the pedestrian image to an image with a height of 256 and a width of 128, and input it into the trained global feature extraction network and pose estimation network.

[0056] As Figure 2 shown, in a specific embodiment, the pose estimation network extracts local key point heatmaps. Specifically, the feature maps output by the high-resolution network HRNet-W32 are subjected to a convolution operation to reduce the number of channels to N, obtaining N local key point heatmaps, corresponding to N key points of the human body.

[0057] In a specific embodiment, the global feature extraction network extracts the global feature map. Specifically, the feature maps of four layers in the high-resolution network HRNet-W32 are enlarged to the corresponding size by bilinear interpolation and these four feature maps are spliced together to obtain the global feature map F g .

[0058] Step S2: Multiply the global feature map after dimensionality reduction and the local key point heatmaps element by element to obtain the local key point feature sequence, and then obtain the local pose joint feature by passing the local key point feature sequence and the global feature map through the foreground pose non-local network.

[0059] In this embodiment, the global feature map is obtained through HRNet-W32 where H, W, and C represent the height, width, and number of channels of the feature map. To reduce the computational complexity, the global feature map F with the number of channels C output by HRNet is reduced to the global feature map with the number of channels D through a 1×1 convolutional kernel g where C is defaulted to 1920 and D is set to 256. Then, multiply this global feature map

[0060] with the N local key point heatmaps element by element to obtain N corresponding local key point feature sequences

[0061] ​Finally, the global feature map after dimensionality reduction and the local key-point feature sequence Z 0 are processed through a foreground pose non-local network, as Figure 3 shown, and the following operations are performed:

[0062] Step 2.1: Add each feature in the local key-point feature sequence Z 0 to the global feature map , and then perform an adaptive average pooling operation to obtain the local key-point feature sequence Z 1 embedded with global features.

[0063] In this embodiment, the global spatial information is embedded into the local features of each key point. This process can be expressed as:

[0064]

[0065] where GeM refers to the Generalized Mean Pooling method of adaptive average pooling, and N represents the number of local features), represents the local key-point feature sequence embedded with global features.

[0066] Step 2.2: Perform a normalization operation on the local key-point feature sequence Z 1 embedded with global features, then further mine fine-grained information through a feed-forward neural network, and then perform a residual connection with the feature sequence Z 1 embedded with global features to obtain the fine-grained local feature sequence Z 2 .

[0067] In this embodiment, in order to reduce the differences between different samples while retaining the differences between different features, the LayerNorm layer is used to perform a normalization operation on the feature sequence Z 1 . Then, a feed-forward neural network is used to further mine fine-grained information. At the same time, a residual connection of the feature sequence Z 1 Z 1 is also added to better retain the originally useful information. As shown in the following formula:

[0068] Z 2 = FFN(LN(Z 1 )) + Z 1

[0069] where FFN() represents the feed-forward neural network, and LN() represents the normalization operation of the LayerNorm layer. Z 2 is the fine-grained local feature sequence mined by the feed-forward neural network (FFN).

[0070] Step 2.3: For the fine-grained local feature sequence Z2 Perform normalization operation and multi-head self-attention operation, and then combine with the fine-grained local feature sequence Z 2 Residual connection to obtain the local feature correlation sequence Z 3 .

[0071] In this embodiment, the fine-grained local feature sequence Z 2 is subjected to normalization operation and multi-head self-attention operation to explore the correlation between local features, and the process is expressed as:

[0072] Z 3 = MSA(LN(Z 2 )) + Z 2

[0073] where MSA() represents the multi-head self-attention module, and Z 3 is the local feature correlation sequence mined through the multi-head self-attention operation (MSA).

[0074] Step S2.4: Concatenate the local feature correlation sequence Z 3 , and then pass through a linear layer to obtain the local pose joint feature f p .

[0075] In this step, the local feature correlation sequence Z 3 is concatenated into a feature with a dimension of , and then passes through a linear layer to obtain the local pose joint feature f p , and the process is as follows:

[0076] f p = Linear(concate(Z 3 ))

[0077] where Linear() is the linear layer, which converts to

[0078] Step S3: Perform adaptive average pooling operation on the global feature map, and then input it together with the local pose joint feature into the fusion module for non-linear fusion to obtain the fusion feature.

[0079] As Figure 4 shown, the fusion module includes two input features: the global feature map F g The coarse-grained feature obtained through GeM pooling and the local pose joint feature where C is the number of channels of the feature map, which is 1920 in this embodiment.

[0080] The global feature f g is a coarse-grained feature, while the local pose joint feature map fp It is a fine-grained local feature that contains the semantic features of each key node. Combining the two can obtain a more robust feature representation. However, common fusion methods, such as addition or concatenation, cannot deeply fuse these two features. Therefore, this embodiment proposes a fusion module for nonlinear fusion, reduces feature redundancy, and integrates global features into fine-grained local features to generate fused features.

[0081] In a specific embodiment, an adaptive average pooling operation is performed on the global feature map, and then the global feature map is input into a fusion module together with the local posture joint feature for nonlinear fusion to obtain a fused feature, including:

[0082] Step 3.1: Perform an adaptive average pooling operation on the global feature map to obtain the coarse-grained feature f g .

[0083] This step is to transform the global feature map F g Perform adaptive average pooling GeM operation to obtain coarse-grained features

[0084] Step 3.2: The coarse-grained feature f g and local pose joint feature f p Input them into their respective feedforward neural networks, and then perform residual connection to obtain the first feature f′ g and the second feature f′p.

[0085] Specifically, in the fusion module, the coarse-grained feature f g and local pose joint feature f p They will enter their respective feedforward neural networks, each of which consists of two linear layers, a ReLU activation layer and two Dropout layers. The feedforward neural network here has the same structure as the feedforward neural network FFN in the foreground pose non-local network. The purpose of using this structure is to retain and further mine the unique information of each feature. The process is shown in the following formula:

[0086] f′ g =FFN(f g )+f g ;

[0087] f′ p =FFN(f p )+f p .

[0088] Step 3.3: The coarse-grained feature f g and local pose joint feature f pCut evenly into a preset number of blocks according to the dimension of the feature, then recombine them into matrices K and Q. After that, multiply matrices Q and K to obtain the feature supervision similarity, and use the feature supervision similarity to calculate the weights of matrix K and matrix S.

[0089] In this embodiment f g and f p will also enter a shared network to fuse them and extract the common information between the two. To this end, first f g and f p are evenly cut into M blocks according to the dimension of the feature, then recombined into matrices K and Q, and Specifically, the feature with dimension C = 1920 is evenly cut into M = 8 blocks, and then recombined into an 8 * 240 feature matrix, f g After cutting and recombination, matrix K is obtained, f p After cutting and recombination, matrix Q is obtained. Then multiply Q and K to obtain the feature supervision similarity, and use this similarity to calculate the weights of K and S. The calculation process is as follows:

[0090]

[0091] where softmax() is the softmax activation function, and C is the dimension of the feature.

[0092] Step 3.4: Horizontally cut the weights and matrix S into a preset number of blocks, then splice them in order and recombine them into the associated feature f s .

[0093] In this step, the weights and matrix are horizontally cut into M blocks, spliced in order, and recombined into the associated feature Specifically, reconnect the 8 * 240 feature matrices to reconstruct the feature with a dimension of C = 1920 again.

[0094] Step 3.5: Accumulate the first feature f′ g , the second feature f′ p and the associated feature f s to obtain the final fused feature f.

[0095] Since f s inevitably loses the unique information of f g and f p during the calculation process, in this step, use the weights to accumulate f′ g , f′ p and f s to obtain the final fused feature f, and the formula is as follows:

[0096] f = αf′ g+βf′ p +(1 - α - β)f s 。

[0097] Where α and β are the weights of each feature, and α, β, α + β ∈ (0, 1).

[0098] It should be noted that during the training of this embodiment, the following processing is also performed on the fused features:

[0099] f′ = cls(BN(f));

[0100] Thus, the triplet loss is calculated using f and the cross - entropy loss is calculated using f′ during the training stage to complete the training of the network model.

[0101] Step S4: Calculate the similarity between the pedestrian image to be queried and the images in the database based on the fused features, sort them according to the similarity size, and then output the query result.

[0102] The task of pedestrian re - identification is to find the images of the same pedestrian under other cameras (referred to as gallery) based on an image of a pedestrian under a certain camera (referred to as query). Among them, query is the input to be queried, also called the query image, and Gallery is the pedestrian image library. Specifically, the final fused features are calculated for both the query image and the images in the gallery using the method of this application. Then, the similarity is calculated for each fused feature between the query image and the images in the gallery, and they are sorted from largest to smallest according to the similarity size. The pedestrian corresponding to the most similar image in the gallery is the recognized result.

[0103] It should be noted that calculating similarity is a relatively mature technology in this field and will not be elaborated here.

[0104] This application also experimentally verified the technical solution of this application. During the experiment, based on the Market - 1501 dataset and the DukeMTMC - Re - ID dataset, it was compared with the current mainstream pedestrian re - identification algorithms, and the mean average precision (mAP) and the first - rank hit rate (Rank - 1) were used to measure the recognition performance. The experimental results are shown in Table 1:

[0105] Table 1

[0106]

[0107]

[0108] Table 1 shows the test results on the Market-1501 dataset and the DukeMTMC-Re-ID dataset. By comparison, it can be seen that in the Market-1501 dataset, the mAP accuracy of this application reaches 90.0%, which is the best result at present. Compared with the PFD method based on pose estimation, the method of this application has improved by 0.4% in both Rank-1 and mAP on Market-1501. For the Rank-1 metric, the accuracy of this application is only 0.2% lower than that of the SAN method, which is the second-best result at present. On the DukeMTMC-Re-ID dataset, the method of this application achieves the highest performance in the Rank-1 metric, with an accuracy of 90.9%. However, the mAP is still slightly lower than that of the PFD method by about 0.5%, and it is currently the second-best, with an accuracy of 81.7%. In summary, the method of this application has certain superiority compared with other algorithms.

[0109] In one embodiment, this application also provides a person re-identification device based on pose estimation and feature association, including:

[0110] A feature extraction module, configured to obtain an image of a pedestrian to be queried, and extract a local key point heat map and a global feature map respectively through a pose estimation network and a global feature extraction network;

[0111] A pose joint module, configured to perform element-wise multiplication on the reduced-dimensional global feature map and the local key point heat map to obtain a local key point feature sequence, and then obtain a local pose joint feature by passing the local key point feature sequence and the global feature map through a foreground pose non-local network;

[0112] A fusion module, configured to perform an adaptive average pooling operation on the global feature map, and then input it together with the local pose joint feature into the fusion module for non-linear fusion to obtain a fusion feature;

[0113] A query module, configured to calculate the similarity between the image of the pedestrian to be queried and the images in the database according to the fusion feature, sort them according to the size of the similarity, and then output the query result.

[0114] For the specific limitations of the person re-identification device based on pose estimation and feature association, reference can be made to the limitations of the person re-identification method based on pose estimation and feature association in the above text, which will not be elaborated here. Each module in the above person re-identification device based on pose estimation and feature association can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0115] The memory and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory stores a computer program that can run on the processor. The processor realizes the network topology layout method in the embodiments of the present invention by running the computer program stored in the memory.

[0116] Among them, the memory can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store the program, and the processor executes the program after receiving the execution instruction.

[0117] The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0118] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A pedestrian re-identification method based on pose estimation and feature association, characterized in that, the pedestrian re-identification method based on pose estimation and feature association includes: Obtain the query pedestrian image, and extract the local key point heat map and the global feature map through the pose estimation network and the global feature extraction network respectively; Multiply the dimension-reduced global feature map and the local key point heat map element by element to obtain the local key point feature sequence, and then pass the local key point feature sequence and the global feature map through the foreground pose non-local network to obtain the local pose joint feature; Perform an adaptive average pooling operation on the global feature map, and then input it together with the local pose joint feature into the fusion module for non-linear fusion to obtain the fusion feature; Calculate the similarity between the query pedestrian image and the images in the database according to the fusion feature, sort them according to the similarity size, and then output the query result; Among them, the step of performing an adaptive average pooling operation on the global feature map, then inputting it together with the local pose joint feature into the fusion module for non-linear fusion to obtain the fusion feature includes: Perform an adaptive average pooling operation on the global feature map to obtain the coarse-grained feature; Input the coarse-grained feature and the local pose joint feature into their respective feed-forward neural networks respectively, and then perform residual connection to obtain the first feature and the second feature; Uniformly cut the coarse-grained feature and the local pose joint feature into a preset number of blocks according to the dimension of the feature, and then recombine them into matrices K and Q. After that, multiply matrix Q and K to obtain the feature supervision similarity, and use the feature supervision similarity to calculate the weight of matrix K and matrix S; Horizontally cut the weight and matrix S into a preset number of blocks, and then splice them in order to recombine them into the associated feature; Accumulate the first feature, the second feature and the associated feature to obtain the final fusion feature; The step of accumulating the first feature, the second feature and the associated feature to obtain the final fusion feature is accumulated using the following formula: f = αf' g + βf' p +(1 - α - β)f s ; Among them, f′ g represents the first feature, f′ p represents the second feature, f s represents the associated feature, and α, β are weight parameters.

2. The pedestrian re-identification method based on pose estimation and feature association according to claim 1, characterized in that, the step of obtaining the local pose joint feature by passing the local key point feature sequence and the global feature map through the foreground pose non-local network includes: Add each feature in the local key point feature sequence to the global feature map and then perform an adaptive average pooling operation to obtain the local key point feature sequence after embedding the global feature; Perform a normalization operation on the local key point feature sequence after embedding the global feature, and then further mine the fine-grained information through a feed-forward neural network, and then perform residual connection with the feature sequence after embedding the global feature to obtain the fine-grained local feature sequence; Perform a normalization operation and a multi-head self-attention operation on the fine-grained local feature sequence, and then perform residual connection with the fine-grained local feature sequence to obtain the local feature association sequence; Splice the local feature association sequence, and then pass it through a linear layer to obtain the local pose joint feature.

3. A pedestrian re-identification device based on pose estimation and feature association, characterized in that, the pedestrian re-identification device based on pose estimation and feature association includes: A feature extraction module, which is used to obtain the pedestrian image to be queried, and extract the local key-point heat map and the global feature map through the pose estimation network and the global feature extraction network respectively; A pose joint module, which is used to perform element-wise multiplication on the downsampled global feature map and the local key-point heat map to obtain the local key-point feature sequence, and then obtain the local pose joint feature through the foreground pose non-local network for the local key-point feature sequence and the global feature map; A fusion module, which is used to perform adaptive average pooling operation on the global feature map, and then input it together with the local pose joint feature into the fusion module for non-linear fusion to obtain the fusion feature; A query module, which is used to calculate the similarity between the pedestrian image to be queried and the images in the database according to the fusion feature, sort them according to the similarity size, and then output the query result; Among them, the operation of performing adaptive average pooling on the global feature map, and then inputting it together with the local pose joint feature into the fusion module for non-linear fusion to obtain the fusion feature includes: Performing adaptive average pooling operation on the global feature map to obtain the coarse-grained feature; Inputting the coarse-grained feature and the local pose joint feature into their respective feed-forward neural networks respectively, and then performing residual connection to obtain the first feature and the second feature; Uniformly cutting the coarse-grained feature and the local pose joint feature into a preset number of blocks according to the dimension of the feature, and then recombining them into matrices K and Q. After that, multiplying matrix Q and matrix K to obtain the feature supervised similarity, and using the feature supervised similarity to calculate the weight of matrix K and matrix S; Horizontally cutting the weight and matrix S into a preset number of blocks, and then splicing them in order to recombine them into the associated feature; Accumulating the first feature, the second feature and the associated feature to obtain the final fusion feature; The operation of accumulating the first feature, the second feature and the associated feature to obtain the final fusion feature is performed by using the following formula for accumulation: f = αf' g + βf' p +(1 - α - β)f s ; where, f′ g represents the first feature, f′ p represents the second feature, f s represents the associated feature, and α, β are weight parameters.

4. The pedestrian re-identification device based on pose estimation and feature association according to claim 3, wherein, the operation of obtaining the local pose joint feature through the foreground pose non-local network for the local key-point feature sequence and the global feature map includes: Adding each feature in the local key-point feature sequence to the global feature map and then performing adaptive average pooling operation to obtain the local key-point feature sequence embedded with the global feature; Performing normalization operation on the local key-point feature sequence embedded with the global feature and then further mining the fine-grained information through the feed-forward neural network, and then performing residual connection with the feature sequence embedded with the global feature to obtain the fine-grained local feature sequence; Performing normalization operation and multi-head self-attention operation on the fine-grained local feature sequence, and then performing residual connection with the fine-grained local feature sequence to obtain the local feature association sequence; Splicing the local feature association sequence, and then passing through a linear layer to obtain the local pose joint feature.