A Person Re-identification Method Incorporating Location-Aware Attention

By introducing a position-aware attention module into the ResNet50 network, using position coding and non-local attention mechanisms, the problem of slow speed and confusing relationship between features during feature extraction is solved, and a more efficient pedestrian re-identification effect is achieved.

CN114663974BActive Publication Date: 2025-06-20NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210247905.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-06-20
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

In the ResNet50 network, the model believes that the importance of each sub-feature in the feature map is the same, resulting in slow training speed and inability to efficiently extract key features that are helpful to the task. At the same time, the attention module lacks the concept of positional relationship between features, which may lead to confusion between features.

Method used

The position-aware attention module is introduced, and the output feature maps obtained by the original input through the first two layers of the ResNet50 network are input into the module for processing, and they are integrated into the ResNet50 network for training and testing. Through position coding and non-local attention mechanism, this module captures the relationship between different position features in the feature map and embeds the attention module to enhance feature expression capabilities.

Benefits of technology

It effectively improves the ability to express features, can more accurately extract pedestrians' distinguishable features, suppress features with low correlation with pedestrian recognition tasks, and achieve better recognition results on multiple pedestrian re-identification standard data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663974B_ABST
    Figure CN114663974B_ABST
Patent Text Reader

Abstract

The present invention provides a pedestrian re-identification method incorporating position-aware attention: A position-aware attention module is introduced into the ResNet50 network. This module is an effective improvement of the non-local attention module. By embedding position information into the non-local attention module that captures long-range feature dependencies, the expressive power of the extracted features is effectively enhanced. The position-aware attention module proposed by the present invention belongs to a lightweight structure. Incorporating this module into the ResNet50 network can effectively extract distinguishable features of pedestrians while suppressing features with low relevance to the pedestrian recognition task, achieving better recognition results than traditional network models and other related methods on multiple popular pedestrian re-identification standard datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a pedestrian re-identification method incorporating location-aware attention. Background Art

[0002] Person Re-identification refers to retrieving pedestrian images with the same identity as a given query image from a pedestrian image database in a scenario with multiple non-overlapping cameras. Person re-identification can be widely applied in fields such as intelligent security and video surveillance.

[0003] Person re-identification can be regarded as a feature-embedding problem. Ideally, the intra-class distance (different images of the same person) should be less than the inter-class distance (images of different people). Unfortunately, most existing feature-embedding solutions require samples to be grouped in a pairwise manner, which is usually computationally intensive. In practice, due to the obvious advantage of the classification task in terms of the implementation complexity of training, classification methods are often used as feature-embedding solutions. Nowadays, most of the latest methods for pedestrian re-identification have evolved from a single metric learning problem or a single classification problem to a multi-task problem, that is, simultaneously adopting a classification loss and a triplet loss. Since each sample image is only labeled with a person ID, end-to-end training methods are usually difficult to learn diverse and rich features without carefully designing the underlying neural network and further using some regularization techniques.

[0004] In recent years, many algorithms based on the attention mechanism and position encoding have been applied to computer vision. Wang et al. (Wang, Xiaolong, et al. Non-local neural networks. / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.) proposed inserting a non-local attention module into the network model, enabling the model to focus on task-related features through the attention mechanism and ignoring a large amount of useless information; the algorithm proposed by Dosovitskiy et al. (Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16×16 words: Transformers for image recognition at sacle [J]. arXiv:2010.11929, 2020.) (Vision Transformer, ViT), by adding position encoding, makes full use of the positions where features appear as prior knowledge to enhance the representativeness of features and can efficiently complete image classification tasks; as a typical method of applying position encoding, the ViT algorithm has been proven to have significant effects on computer vision tasks. However, the position encoding in ViT is directly added to the input image, resulting in too many parameters, and the network may encounter difficulties in learning the corresponding features. One way to reduce the number of parameters is to add position encoding when the image size is small. At the same time, in order to make full use of the ability of attention to extract key features, this method proposes a position-aware attention module.

[0005] This method obtains long-range feature dependencies through the non-local attention module, effectively improving the pedestrian recognition accuracy of the ResNet50 network. To solve the problem of the lack of position relationship between image features, the present invention proposes a position-aware attention module, integrates it into the ResNet50 network for training and testing, obtains a similarity ranking through distance measurement, and obtains a more accurate pedestrian re-identification result. Summary of the Invention

[0006] An embodiment of the present invention provides a pedestrian re-identification method incorporating position-aware attention to solve the following problems in the prior art:

[0007] In the pedestrian re-identification method of the ResNet50 network, the model considers that the importance of each sub-feature in the feature map is the same and needs to consider all features, resulting in slow training speed and inability to efficiently extract key features helpful for the task;

[0008] During the training process, the attention module can only help the model extract key features related to the task. Without the concept of the positional relationship between features, it may cause problems with the disordered relationship between features.

[0009] To solve the above problems, the present invention adopts the following technical solutions:

[0010] A person re-identification method incorporating position-aware attention includes inputting the output feature map obtained by passing the original input through the first two layers of the ResNet50 network into the position-aware attention module for processing and integrating the position-aware attention module into the ResNet50 network for training and testing;

[0011] The process of inputting the output feature map obtained by passing the original image through the first two layers of the ResNet50 network into the position-aware attention module for processing includes:

[0012] S1: Obtain the input feature map, extract three different feature maps through a convolutional filter, perform pooling operations on two of the feature maps to obtain feature map φ and feature map g, and keep feature map θ unchanged; then flatten and straighten the above three-dimensional feature maps θ, φ, and g into two-dimensional feature matrices along the channel dimension, transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain two-dimensional feature matrices θ and g, and keep the two-dimensional feature matrix φ unchanged;

[0013] S2: Based on the features at different positions in the position-aware encoded feature map, construct a two-dimensional position encoding matrix PE; multiply the two-dimensional feature matrix θ by the two-dimensional feature matrix φ to obtain the relationship matrix R between features and features; multiply the two-dimensional position encoding matrix PE by the two-dimensional feature matrix θ to obtain the relationship matrix R between features and positions; θ,φ ; θ,PE ;

[0014] S3: Add the two relationship matrices R θ,φ and R θ,PE in S2 to achieve position information embedding, and use the normalized exponential function (Softmax function) to obtain the normalized self-correlation weight coefficient matrix f containing position information c = Softmax(R θ,PE + R θ,φ );

[0015] S4: Multiply the normalized self-correlation weight coefficient matrix f containing position information c by the two-dimensional feature matrix g representing the feature map to obtain the two-dimensional spatial position key information matrix, then restore it to the three-dimensional spatial position key information feature map along the channel, use a convolutional filter to increase the dimension, and finally use a structure similar to the residual structure to add the input and the three-dimensional spatial position key information feature map after dimension increase to obtain the output of the position-aware attention module;

[0016] Incorporating the position-aware attention module into the ResNet50 network for training and testing includes:

[0017] S5: Insert the position-aware attention module into the output position of the second layer of the ResNet50 network, and use the weighted form of cross-entropy and triplet loss functions as the total loss function to train along with the network, and input test images to obtain pedestrian matching recognition results.

[0018] Preferably, step S1 specifically includes:

[0019] S1.1 Pass the input feature map X ∈ R b×c×h×w through three 1×1 convolutional filters with different weight coefficients and the number of output channels being the number of input channels to obtain three different feature maps, denoted as θ, φ, and g respectively, where b, c, h, w, and r are the number of images per batch, the number of channels, height, width, and the channel dimension reduction factor respectively;

[0020] S1.2 Select the feature maps φ and g from the three different feature maps for pooling operations to obtain the feature maps and The feature map without pooling operation is denoted as

[0021] S1.3 Flatten and straighten the above three feature maps into two-dimensional feature matrices according to the channel dimension, and transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain the two-dimensional feature matrices and the two-dimensional feature matrix The two-dimensional feature matrix remains unchanged;

[0022] Preferably, step S2 includes:

[0023] S2.1 Randomly initialize different position embedding vectors at different positions. The initial value of each position embedding vector is randomly taken from a normal distribution with a mean of 0 and a variance of 1, and all the position embedding vectors are arranged row by row to form a two-dimensional position encoding matrix All the parameters in PE are updated during the training process;

[0024] S2.2 Multiply the two-dimensional feature matrices representing two different feature maps with to obtain the relationship matrix R θ,φ = θ × φ, where

[0025] S2.3 The two-dimensional feature matrix described above Multiply with the matrix representing the relationship of characteristic positions to obtain the relationship matrix R between the characteristics and positions θ,PE = θ × PE, where

[0026] Preferably, step S3 specifically includes:

[0027] S3.1 Add the relationship matrix R between the characteristics θ,φ and the relationship matrix R between the characteristics and positions θ,PE to embed the position information and obtain the self - correlation weight coefficient matrix with position information At this time it contains the position relationship between sub - characteristics in the feature map;

[0028] S3.2 Pass the self - correlation weight coefficient matrix f with position information through the softmax function to obtain the normalized self - correlation weight coefficient matrix f with position information c = Softmax(R θ,PE + R θ,φ ), where

[0029] Preferably, step S4 specifically includes:

[0030] S4.1 Multiply the normalized self - correlation weight coefficient matrix f with position information c with the two - dimensional feature matrix representing the feature map to obtain the two - dimensional spatial position key information matrix g f = f c × g, where

[0031] S4.2 Transpose the two - dimensional spatial position key information matrix and restore it to a three - dimensional spatial position key information feature map according to the channels Use a 1×1 convolution filter to increase the dimension to make it the same as the channel number dimension of the input feature map, and the output is denoted as g fc ∈ R b×c×h×w ;

[0032] S4.3 Add the input feature map X ∈ R b×c×h×w and the three - dimensional spatial position key information feature map g after dimension increase fc ∈ R b ×c×h×w to obtain the output Y = X + g of the position - aware attention module fc , where Y ∈ R b×c×h×w .

[0033] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0034] 1. A person re-identification method incorporating position-aware attention provided by the present invention introduces a position-aware attention module into the ResNet50 network. This module is an effective improvement of the non-local attention module. By embedding position information into the non-local attention module that captures long-range feature dependencies, the expressive ability of the extracted features is effectively enhanced.

[0035] 2. The position-aware attention module proposed by the present invention belongs to a lightweight structure. Incorporating this module into the ResNet50 network can effectively extract distinguishable features of pedestrians, while suppressing features with little relevance to the pedestrian recognition task, achieving better recognition effects than traditional network models and other related methods on multiple popular pedestrian re-identification standard datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a processing flow chart of a person re-identification method incorporating position-aware attention provided by the present invention;

[0037] Figure 2 is a basic architecture diagram of the non-local attention module;

[0038] Figure 3 is a basic architecture diagram of the position-aware attention module proposed in a person re-identification method incorporating position-aware attention provided by the present invention;

[0039] Figure 4 is the overall architecture diagram of the ResNet50 network in a person re-identification method incorporating position-aware attention provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following embodiments will further illustrate the present invention in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the present invention. On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made within the spirit and scope defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, in the following detailed description of the present invention, some specific details are described in detail. Those skilled in the art can fully understand the present invention without the description of these details.

[0041] Embodiment 1

[0042] See Figure 1, a person re-identification method incorporating position-aware attention provided by the present invention mainly includes two processes: inputting the output feature map obtained by passing the original image through the first two layers of the ResNet50 network into the position-aware attention module for processing, and integrating the position-aware attention module into the ResNet50 network for training and testing.

[0043] Among them, inputting the output feature map obtained by passing the original image through the first two layers of the ResNet50 network into the position-aware attention module for processing includes:

[0044] S1: Obtain the input feature map, extract three different feature maps through a convolutional filter, perform pooling operations on two of the feature maps to obtain feature maps φ and g, and keep the feature map θ unchanged; then flatten and straighten the above three-dimensional feature maps θ, φ, and g into two-dimensional feature matrices along the channel dimension, transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain two-dimensional feature matrices θ and two-dimensional feature matrix g, and keep the two-dimensional feature matrix φ unchanged;

[0045] S2: Based on the features at different positions in the position-aware encoded feature map, construct a two-dimensional position encoding matrix PE, multiply it with the two-dimensional feature matrix θ to obtain a relationship matrix R between features and positions θ,PE ; multiply the two-dimensional feature matrix θ with the two-dimensional feature matrix φ to obtain a relationship matrix R between features and features θ,φ ;

[0046] S3: Add the two relationship matrices in S2 to achieve position information embedding, and after Softmax, obtain a normalized self-correlation weight coefficient matrix f containing position information c = Softmax(R θ,PE + R θ,φ );

[0047] S4: Multiply the normalized self-correlation weight coefficient matrix f containing position information c with the two-dimensional feature matrix g representing the feature map to obtain a two-dimensional spatial position key information matrix, then restore it to a three-dimensional spatial position key information feature map along the channel, and use a convolutional filter to increase the dimension. Finally, use a structure similar to the residual structure to add the input and the three-dimensional spatial position key information feature map after dimension increase to obtain the output of the position-aware attention module.

[0048] In the embodiment provided by the present invention, a position-aware attention module is adopted. The position-aware attention module is mainly composed of the fusion of a non-local attention module and a position encoding mechanism. The basic architecture of the non-local attention module is as Figure 2 shown, and the basic architecture of the position-aware attention module is as Figure 3As shown. The position encoding encodes the position information of different features. Based on this, by using the attention module, not only can it learn which regions in the feature map are key features, but also the positional relationship between key features can be learned, enhancing the acquisition of discriminative features of the image and refining the features adaptively.

[0049] The sub-features in the deep feature map of the convolutional neural network can be regarded as responses to different semantic features and are interrelated. Non-local attention can mine the dependence relationship between each sub-feature in the feature map. In fact, the importance of each sub-feature in the special map is different. By assigning weights, the importance of each sub-feature to key information is extracted, and the information with large weight values is selectively focused on, enhancing the feature representation of discriminative semantics and improving the feature classification performance.

[0050] Embodiment 2

[0051] The inventor found that in the pedestrian re-identification method of the ResNet50 network, the model considers that the importance of each sub-feature in the feature map is the same and all features need to be considered, resulting in slow training speed and inability to efficiently extract key features helpful for the task. To solve the above problems, in the preferred embodiment of the present invention, a non-local attention module is provided, and its basic architecture is as Figure 2 shown, and the specific steps are as follows:

[0052] S1.1 Pass the input feature map X∈R b×c×h×w through three 1×1 convolutional filters with different weight coefficients and the number of output channels being the number of input channels respectively to obtain three different feature maps, denoted as θ, φ, and g respectively, where b, c, h, w, and r are the number of images per batch, the number of channels, height, width, and the channel dimension reduction factor respectively;

[0053] S1.2 Select the feature maps φ and g from the three different feature maps for pooling operations to obtain the feature maps and The feature map without pooling operation is denoted as

[0054] S1.3 Flatten and straighten the above three feature maps into two-dimensional feature matrices according to the channel dimension, and transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain the two-dimensional feature matrix and the two-dimensional feature matrix The two-dimensional feature matrix remains unchanged.

[0055] S2 Multiply the two-dimensional feature matrices representing two different feature maps with to obtain the relationship matrix R of features θ,φ= θ × φ, where

[0056] S3 transforms the relationship matrix R between features θ,φ through Softmax to obtain the normalized self - correlation weight coefficient matrix R' θ,φ = Softmax(R θ,φ ), where

[0057] S4.1 multiplies the normalized self - correlation weight coefficient matrix R' θ,φ and the two - dimensional feature matrix g representing the feature map to obtain the two - dimensional space key information matrix g R = R' θ,φ × g, where

[0058] S4.2 transposes the two - dimensional space key information matrix and restores it to a three - dimensional space key information feature map by channels using a 1×1 convolutional filter to increase the dimension to make it the same as the channel number dimension of the input feature map, and the output is denoted as g Rc ∈ R b×c×h×w ;

[0059] S4.3 adds the input feature map X ∈ R b×c×h×w and the three - dimensional space key information feature map g after dimension increase Rc ∈ R b×c×h×w to obtain the output Y of the non - local attention module, Y = X + g Rc , where Y ∈ R b×c×h×w .

[0060] In the embodiment provided by the present invention, the basic architecture of the non - local attention adopted is as Figure 2 shown. During the training process, the attention module can only help the model extract the key features related to the task, without the concept of the positional relationship between features, which may cause the problem of disordered feature relationships. To address this drawback, the present invention incorporates the position encoding mechanism into the non - local attention module. In the above S2, the step is added: multiplying the two - dimensional feature matrix representing the feature map with the two - dimensional position encoding matrix to obtain the relationship matrix between features and positions Then in the above S3, the step is added: adding the relationship matrix between features and the relationship matrix between features and positions to achieve position information embedding, thus solving the problem that the model lacks the concept of the positional relationship between features.

[0061] The specific implementation steps of the present invention are as follows:

[0062] S1: Obtain the input feature map, extract three different feature maps through convolutional filters, perform pooling operations on two of the feature maps to obtain feature maps φ and g, and keep the feature map θ unchanged; then flatten and straighten the above three-dimensional feature maps θ, φ, and g into two-dimensional feature matrices along the channel dimension, transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain two-dimensional feature matrices θ and two-dimensional feature matrix g, and keep the two-dimensional feature matrix φ unchanged;

[0063] S1.1 Let the input feature map X ∈ R b×c×h×w Pass through three 1×1 convolutional filters with different weight coefficients and the number of output channels equal to the number of input channels respectively to obtain three different feature maps, denoted as θ, φ, and g respectively, where b, c, h, w, and r are the number of images per batch, the number of channels, the height, the width, and the channel dimension reduction factor respectively;

[0064] S1.2 Select feature maps φ and g from the three different feature maps for pooling operations to obtain feature map and feature map The feature map without pooling operation is denoted as

[0065] S1.3 Flatten and straighten the above three feature maps into two-dimensional feature matrices along the channel dimension, and transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain two-dimensional feature matrix and two-dimensional feature matrix The two-dimensional feature matrix remains unchanged.

[0066] S2: Based on the features at different positions in the position-aware encoded feature map, construct a two-dimensional position encoding matrix PE; multiply the two-dimensional feature matrix θ by the two-dimensional feature matrix φ to obtain the relationship matrix R θ,φ between features; multiply the two-dimensional position encoding matrix PE by the two-dimensional feature matrix θ to obtain the relationship matrix R θ,PE between features and positions;

[0067] S2.1 Randomly initialize different position embedding vectors for different positions. The initial value of each position embedding vector is randomly taken from a normal distribution with a mean of 0 and a variance of 1. All the position embedding vectors are arranged row by row to form a two-dimensional position encoding matrix All the parameters in PE are updated during the training process;

[0068] S2.2 Multiply the two-dimensional feature matrices representing two different feature maps by to obtain the relationship matrix R between featuresθ,φ = θ × φ, where

[0069] S2.3 Multiply the two-dimensional feature matrix with the two-dimensional position encoding matrix to obtain the relationship matrix R between features and positions θ,PE = θ × PE, where

[0070] S3: Add the two relationship matrices R θ,PE and R θ,φ in S2 to achieve position information embedding, and after Softmax, obtain the normalized self-correlation weight coefficient matrix f with position information c = Softmax(R θ,PE + R θ,φ );

[0071] S3.1 Add the relationship matrix R between features θ,φ and the relationship matrix R between features and positions θ,PE to achieve position information embedding, and obtain the self-correlation weight coefficient matrix with position information At this time contains the position relationship between sub-features in the feature map;

[0072] S3.2 Apply Softmax to the self-correlation weight coefficient matrix f with position information to obtain the normalized self-correlation weight coefficient matrix f with position information c = Softmax(R θ,PE + R θ,φ ), where

[0073] S4: Multiply the normalized self-correlation weight coefficient matrix f with position information c by the two-dimensional feature matrix g representing the feature map to obtain the two-dimensional spatial position key information matrix, then restore it to the three-dimensional spatial position key information feature map by channels, and use a convolutional filter to increase the dimension. Finally, use a structure similar to the residual structure to add the input and the three-dimensional spatial position key information feature map after dimension increase to obtain the output of the position-aware attention module;

[0074] S4.1 Multiply the normalized self-correlation weight coefficient matrix f with position information c by the two-dimensional feature matrix representing the feature map to obtain the two-dimensional spatial position key information matrix g f = f c × g, where

[0075] S4.2 Transpose the two-dimensional spatial position key information matrix and restore it to a three-dimensional spatial position key information feature map by channels. Use a 1×1 convolutional filter to increase the dimension so that it has the same number of channels as the input feature map, and denote the output as g. fc ∈R b×c×h×w ;

[0076] S4.3 For the input feature map X∈R b×c×h×w and the three-dimensional spatial position key information feature map g after dimension increase fc ∈R b ×c×h×w add them together to obtain the output Y = X + g of the position-aware attention module fc , where Y∈R b×c×h×w .

[0077] S5 Insert the position-aware attention module into the output position of the second layer of the ResNet50 network, and use the weighted form of cross-entropy and triplet loss functions as the total loss function to train with the network, and input test pictures to obtain the pedestrian matching recognition result.

[0078] Example 3

[0079] The present invention also provides an example for showing a specific experimental process of the method provided by the present invention.

[0080] In this example, three datasets, namely Market1501, DukeMTMC-ReID, and CUHK03, are used for training and testing. Market1501 was collected on the campus of Tsinghua University in the summer of 2015, containing 1501 pedestrian IDs, and a total of 32,668 pictures were collected through 6 cameras. Among them, the training set contains 751 pedestrian IDs with a total of 12,936 pictures, and the test set contains the remaining 750 IDs, 3,368 retrieval pictures, and 15,913 pictures to be inspected; DukeMTMC-reID was collected on the campus of Duke University in the winter of 2015, containing 1,812 pedestrian IDs, and there are a total of 36,411 pictures. Among them, the training set contains 702 pedestrian IDs with a total of 16,522 pictures, and the test set contains the remaining 702 pedestrian ID pictures. The CUHK03 dataset contains 14,096 manually labeled images and 14,097 detection-labeled images. These images are captured by two camera views, with a total of 1,467 IDs. Among them, the pictures of 767 IDs are used for training, and the rest are used for testing.

[0081] In the training stage, the method of data augmentation is adopted to cut the pictures into pedestrian images of 384×128 size, and the pictures are randomly mirrored and regularized, and then sent to the network model for training. In the testing stage, the global branch features and local branch features are concatenated, and the similarity ranking results are obtained through distance measurement.

[0082] In terms of training parameter settings, according to the GPU video memory, the batch size in the training process is set to 64 (including 16 pedestrian IDs, 4 pictures for each ID), the training cycle is set to 160, the Adam optimizer is selected, and the initial learning rate is 3.5×10 -5 , and the WarmUp strategy is adopted to increase the learning rate to 3.5×10 after 10 epochs -4 , and the learning rate is reduced to 3.5×10 at 30 epochs and 60 epochs respectively -5 and 3.5×10 -6 . After each Epoch in the training process, the model will be evaluated and saved through the test set. After all rounds of training are completed, the weights with the best recognition effect are saved as the final model file. The recognition effect of each batch of pedestrian pictures is tested through the saved model, and finally the experimental data are observed and recorded.

[0083] In summary, the present invention provides a pedestrian re-identification method incorporating position-aware attention: a position-aware attention module is introduced into the ResNet50 network. This module is an effective improvement of the non-local attention module. By embedding position information into the non-local attention module that captures long-range feature dependencies, the expression ability of the extracted features is effectively improved. The position-aware attention module proposed by the present invention belongs to a lightweight structure. Incorporating this module into the ResNet50 network can effectively extract distinguishable features of pedestrians, and at the same time suppress features with little relevance to the pedestrian recognition task, achieving better recognition effects than traditional network models and other related methods on multiple popular pedestrian re-identification standard datasets.

[0084] Those of ordinary skill in the art can understand that: the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0085] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0086] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A pedestrian re-identification method incorporating location-aware attention, characterized in that, Including inputting the output feature map obtained by passing the original image through the first two layers of the ResNet50 network into the position-aware attention module for processing, and integrating the position-aware attention module into the ResNet50 network for training and testing; The process of inputting the output feature map obtained by passing the original image through the first two layers of the ResNet50 network into the position-aware attention module for processing includes: S1: Obtain the input feature map, extract three different feature maps through a convolutional filter, perform pooling operations on two of the feature maps to obtain feature map φ and feature map g, and keep feature map θ unchanged; then flatten and straighten the above three-dimensional feature maps θ, φ, and g into two-dimensional feature matrices along the channel dimension, transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain two-dimensional feature matrices θ and g, and keep the two-dimensional feature matrix φ unchanged; S2: Based on the features at different positions in the position-aware coding feature map, a two-dimensional position encoding matrix PE is constructed; the two-dimensional feature matrix θ is multiplied by the two-dimensional feature matrix φ to obtain a relationship matrix R between features and features θ,φ ; Multiply the two-dimensional position encoding matrix PE by the two-dimensional feature matrix θ to obtain the relationship matrix R between features and positions θ,PE ; S3: Add the two relational matrices R θ,φ and R θ,PE to achieve position information embedding. After passing through the softmax function, obtain the normalized self-correlation weight coefficient matrix f c = Softmax(R θ,PE + R θ,φ ); S4: Multiply the normalized autocorrelation weight coefficient matrix f containing position information c with the two-dimensional feature matrix g representing the feature map to obtain a two-dimensional spatial position key information matrix, then restore it to a three-dimensional spatial position key information feature map by channel, and use a convolutional filter to increase the dimension. Finally, use a structure similar to the residual structure to add the input and the three-dimensional spatial position key information feature map after dimension increase to obtain the output of the position-aware attention module; The process of integrating the position-aware attention module into the ResNet50 network for training and testing includes: S5: Insert the position-aware attention module into the output position of the second layer of the ResNet50 network, use the weighted form of cross-entropy and triplet loss functions as the total loss function to train along with the network, and input test images to obtain pedestrian matching recognition results.

2. The method according to claim 1, characterized in that, Step S1 specifically includes: S1.1 Input the feature map X ∈ R b×c×h×w through three 1×1 convolutional filters with different weight coefficients and the number of output channels equal to the number of input channels to obtain three different feature maps, denoted as θ, φ, and g respectively, where b, c, h, w, and r are the number of images per batch, the number of channels, height, width, and the channel dimensionality reduction factor respectively; S1.2 Select feature maps φ and g from three different feature maps for pooling operations to obtain a feature map and a feature map The feature map without pooling operation is denoted as S1.3 Flatten the above three feature maps into two-dimensional feature matrices along the channel dimension, and transpose the two-dimensional feature matrices corresponding to the three-dimensional feature maps θ and g to obtain two-dimensional feature matrices and two-dimensional feature matrix The two-dimensional feature matrix Remains unchanged.

3. The method according to claim 1, characterized in that, Step S2 includes: S2.1 Randomly initialize different position embedding vectors for different positions The initial value of each position embedding vector is randomly sampled from a normal distribution with a mean of 0 and a variance of 1. All position embedding vectors are arranged row by row to form a two-dimensional position encoding matrix All parameters in PE are updated during the training process; S2.2 Multiply the two-dimensional feature matrices representing two different feature maps with to obtain the relationship matrix R of features θ,φ = θ × φ, where S2.3 Multiply the aforementioned two-dimensional feature matrix by the two-dimensional position encoding matrix to obtain the relationship matrix R θ,PE between features and positions, where 4. The method according to claim 1, characterized in that, Step S3 includes: S3.1 Add the relationship matrix R θ,φ between features and the relationship matrix R θ,PE between features and positions to embed the position information and obtain the self - correlation weight coefficient matrix containing position information At this time contains the position relationship between sub - features in the feature map; S3.2 Pass the self - correlation weight coefficient matrix f containing position information through the sigmoid function to obtain the normalized self - correlation weight coefficient matrix f containing position information c = Softmax(R θ,PE + R θ,φ ), where 5. The method according to claim 1, characterized in that, Step S4 specifically includes: S4.1 Multiply the normalized autocorrelation weight coefficient matrix f with the position information c by the two-dimensional feature matrix representing the feature map to obtain the two-dimensional spatial position key information matrix g f = f c × g, where S4.2 Transpose the two-dimensional spatial position key information matrix and restore it to a three-dimensional spatial position key information feature map by channels Use a 1×1 convolutional filter to increase the dimension to make it the same as the channel number dimension of the input feature map, and the output is denoted as g fc ∈R b×c×h×w ; S4.3 Add the input feature map X ∈ R b×c×h×w and the key information feature map g of the three-dimensional spatial position after dimensionality increase fc ∈ R b×c×h×w to obtain the output Y = X + g of the position-aware attention module fc , where Y ∈ R b×c×h×w .

Citation Information

Patent Citations

  • Pedestrian re-identification method based on multi-component self-attention mechanism

    CN111368815A

  • Residual network expression recognition method integrated with attention

    CN112541409A