Pedestrian re-identification method and system based on double-path attention mechanism
By introducing a dual-channel attention mechanism and ResNetRFA network into the pedestrian re-identification model, the pedestrian local and global features are extracted, and the problem that existing models are difficult to extract key features and utilize contextual features in an interference environment is solved, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510011144.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
Under the interference factors such as background noise, lighting changes, posture changes and severe occlusion, it is difficult for existing pedestrian re-identification models to extract key features of pedestrians and make full use of context feature areas for identification.
The pedestrian re-identification method based on the dual-channel attention mechanism is adopted. By constructing the ResNetRFA network and the dual-channel attention mechanism module, the pedestrian target key features and receptive field feature maps are extracted, and the local-global relationship feature maps are obtained through the global context attention mechanism and the local relation attention mechanism module, and finally the identity matching is performed through the calculation method of optimizing the Euclidean distance between vectors.
In an interfering environment, the accuracy and robustness of pedestrian re-identification are improved, the ability to combine pedestrian global and local context information is enhanced, and the computing complexity of identity matching is reduced.
Smart Images

Figure CN119942592A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and in particular relates to a pedestrian re-identification method and system based on a dual-path attention mechanism. Background Art
[0002] Person re-identification (Re-ID) focuses on identifying the target person from images taken by multiple cameras. Re-ID can be used to identify pedestrians in surveillance video streams in the field of intelligent security. With the continuous development of deep learning technology, Re-ID based on deep learning has become a hot research direction and has been applied to various fields. In an open community environment, due to the camera installation location and deployment area, the collected community pedestrian images have interference factors such as background noise, lighting changes, posture changes, and severe occlusion. The existing pedestrian re-identification model in the community cannot fully explore the local key features of pedestrians under interference factors, nor can it fully use the context feature area for identification. Therefore, it is necessary to propose a pedestrian re-identification method that can accurately extract strong discriminative and high-discriminative features. Summary of the invention
[0003] In order to solve the above technical problems, the present invention provides a pedestrian re-identification method based on a dual-path attention mechanism, comprising the following steps:
[0004] Step S1: construct a community pedestrian dataset and preprocess the community pedestrian dataset;
[0005] Step S2: Construct a ResNetRFA network based on channel-by-channel convolution, receptive field attention mechanism and residual network, input the community pedestrian dataset into the ResNetRFA network, extract the key features of pedestrian targets and generate a receptive field feature map ;
[0006] Step S3: After being divided into several local feature maps The two-way attention mechanism module is input together, and the local-global relationship feature map is obtained through the global context attention mechanism module and the local relationship attention mechanism module;
[0007] Step S4: converting the local-global relationship feature map into a feature vector of a predetermined dimension for representing pedestrians, optimizing the calculation method of the Euclidean distance by the angle difference between the vectors, and determining whether two pedestrians are the same person.
[0008] Beneficial effects:
[0009] 1. In order to solve the problem that the algorithm does not pay enough attention to the key features of community pedestrians under interference environments such as background noise and lighting changes, the present invention constructs a ResNetRFA network for feature extraction, emphasizes the important features of pedestrians, and strengthens the attention to the key features of pedestrians.
[0010] 2. In order to solve the problem of missing global and local context information features of pedestrians in interference environments such as background noise and lighting changes under cross-camera conditions in the community, a pedestrian re-identification model based on local relationship attention mechanism and global context attention mechanism is introduced to enhance the ability to combine global and local context information in interference environments and improve the accuracy of pedestrian re-identification.
[0011] 3. In order to reduce the computational complexity of community pedestrian identity matching and improve the accuracy of identity matching, the straight-line distance between local-global contrast feature vectors is converted into the angular difference between vectors, which improves the accuracy of identity matching while reducing the amount of computation. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A schematic diagram of a pedestrian re-identification method based on a dual-path attention mechanism of the present invention;
[0013] Figure 2 This is a schematic diagram of the architecture of the ResNetRFA network;
[0014] Figure 3 The schematic diagram shows the structure of the dual-path attention mechanism module;
[0015] Figure 4 The overall framework diagram of the pedestrian re-identification method based on the dual-path attention mechanism;
[0016] Figure 5 This is a structural block diagram of a pedestrian re-identification system based on a dual-path attention mechanism of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0018] Embodiment 1
[0019] like Figure 1 As shown, a pedestrian re-identification method based on a dual-path attention mechanism provided by an embodiment of the present invention includes the following steps:
[0020] Step S1: construct a community pedestrian dataset and preprocess the community pedestrian dataset;
[0021] Step S2: Construct a ResNetRFA network based on channel-by-channel convolution, receptive field attention mechanism and residual network, input the community pedestrian dataset into the ResNetRFA network, extract the key features of pedestrian targets and generate a receptive field feature map ;
[0022] Step S3: After being divided into several local feature maps The two-way attention mechanism module is input together, and the local-global relationship feature map is obtained through the global context attention mechanism module and the local relationship attention mechanism module;
[0023] Step S4: Convert the local-global relationship feature map into a feature vector of a predetermined dimension to represent pedestrians, and optimize the calculation method of the Euclidean distance by the angle difference between the vectors to determine whether two pedestrians are the same person.
[0024] In one embodiment, the above step S1: constructing a community pedestrian dataset and preprocessing the community pedestrian dataset specifically includes:
[0025] There are several key steps, including data collection, preprocessing, and data enhancement. Based on the video data obtained from community surveillance, the video data is diverse and authentic, covering pedestrian images in different scenes, different cameras, and different lighting conditions. The video data obtained from community surveillance is screened and classified, and image samples containing real pedestrians are selected, and damaged image samples are excluded. The pedestrian image sequence is annotated using annotation software. When performing target identity annotation, the camera category, pedestrian identity information, and time information are clearly defined. The annotated community pedestrian data is used as the community pedestrian dataset. After loading the community pedestrian dataset, the image size is resized to 128 by scaling and cropping. 384. Finally, the diversity of training data is increased by rotation and flipping to improve the generalization ability of the model.
[0026] In one embodiment, the above step S2: constructs a ResNetRFA network based on channel-by-channel convolution, receptive field attention mechanism and residual network, inputs the community pedestrian dataset into the ResNetRFA network, extracts the key features of pedestrian targets and generates a receptive field feature map , including:
[0027] The ResNetRFA network is constructed based on the residual network. The improvements are:
[0028] 1) At Layer 0, replace the original convolution kernel with channel-by-channel convolution to perform convolution on the input image, thus obtaining a convolution operation of size 192 64 feature map with 64 channels;
[0029] 2) The receptive field attention mechanism RFA is added to the Bottleneck of Layer 4, where RFA uses a sliding window with shared parameters to extract feature information based on the convolution operation;
[0030] In the residual network architecture, the present invention adds the receptive field attention mechanism to layer4, and gradually extracts features through the Bottleneck of ResNetRFA. In the initial feature extraction stage, due to the excessive size of the convolution kernel, some features are lost and the feature extraction of key positions is incomplete. Therefore, the original model convolution kernel is replaced with channel-by-channel convolution to increase the overall feature extraction capability of the algorithm. In order to adapt to the corresponding number of channels, the step size is set to 1 and the number of channels is set to 64. Compared with ordinary convolution, RFA has a sliding window with shared parameters to extract feature information, which overcomes the problems of many parameters and high computational overhead inherent in the construction of neural networks with fully connected layers.
[0031] Assume the input feature map is ,in Respectively represent the number of channels, height, and width of the feature map; , then the convolution operation to extract feature information from each receptive field slider can be expressed as:
[0032] (1)
[0033] in, Represents the value obtained by each convolution slider after calculation, Indicates the pixel value of the corresponding position in each slider. represents the convolution kernel, represents the number of parameters in the convolution kernel, Represents the total number of receiving domain sliders;
[0034] It can be seen from formula (1) that the features at the same position in each convolution slider share the same parameters ;
[0035] For RFA, learning the attention map by interacting with the feature information of the receiving domain can improve the network performance. Since interacting with each receiving domain feature will cause additional computational overhead, in order to minimize the computational overhead and the number of parameters, AvgPool is used to aggregate the global information of each receiving domain feature, convolution operation is used for information interaction, and Softmax is used to emphasize the importance of each feature in the receiving domain feature. Therefore, the calculation of RFA can be expressed as formula (2):
[0036] (2)
[0037] in, Represents a size of Grouped convolution; represents the size of the convolution kernel, represents normalization, represents the input feature map, Represents the receptive field feature map.
[0038] RFA is able to generate an attention map for each receptive field feature. The performance of convolutional neural networks is limited by standard convolution operations, which rely on shared parameters and are insensitive to information differences caused by position changes. RFA can solve this problem by emphasizing the importance of different features in the receptive field slider and prioritizing the receptive field spatial features. The feature map obtained by RFA is able to receptive field spatial features and has no overlap after adjusting the tensor. ResNetRFA can more effectively handle changes in background, lighting, occlusion, etc. in pedestrian re-identification tasks. By emphasizing important features and reducing attention to irrelevant information, the recognition accuracy and robustness of the model are improved.
[0039] Traditional deep learning models solve the gradient vanishing problem in deep network training by introducing residual learning, which allows the network to be designed deeper without affecting the training effect, thereby learning richer feature representations. However, traditional deep learning models such as ResNet50 may not capture key features in the residual connections between layers, resulting in low pedestrian recognition accuracy.
[0040] Although ResNet50 uses residual connections to enhance feature propagation, it cannot fully capture key features. In order to solve the problem of insufficient extraction of pedestrian key features, the present invention designs a ResNetRFA network architecture, adds receptive-field attention (RFA), flexibly adjusts the receptive field to extract pedestrian key features, and effectively improves the detection efficiency of pedestrian re-identification. The architecture of the ResNetRFA network is shown in the figure. Figure 2 shown.
[0041] In one embodiment, the above step S3: After being divided into several local feature maps The two-way attention mechanism module is input together, and the local-global relationship feature map is obtained through the global context attention mechanism module and the local relationship attention mechanism module, which specifically includes:
[0042] Step S31: Divide the height into six equal parts, leaving the width unchanged, and get six parts A local feature map and a The local feature map and the global feature map are input into the dual-path attention mechanism module respectively to extract the local-global relationship feature map; wherein the dual-path attention mechanism module includes two branches: the global context attention mechanism module GCA and the local relationship attention mechanism module PRA;
[0043] Step S32: First, the local feature map is input into the global context attention mechanism module GCA:
[0044] (3)
[0045] in, Represents the global context features of the local feature map after GCA, and the global attention pooling is performed through the weighted The weighted average groups the features at all positions together to obtain the global context feature, represents the feature transformation that captures channel dependencies, Represents the features of the location to be queried in the local feature map, express The characteristics of all possible positions except is the total number of all positions in the feature map, represents the fusion function that aggregates the global context features into the features at each position;
[0046] in, After expansion, it is as follows:
[0047] (4)
[0048] (5)
[0049] in, is the weight of the global attention pool; is an element in the input feature map; represents a specific position in the feature map, Representation and The corresponding feature weights, represents the weight of the first convolutional layer, represents the weight of the second convolutional layer; Representation layer normalization, represents the activation function;
[0050] GCA can better capture long-distance dependencies, and global context can benefit the fine-grained resolution task of person re-identification. The flexibility of GCA enables it to be inserted into the network architecture of person re-identification. Although it can enhance the accuracy of person re-identification, it ignores the importance of local features. When calculating the relationship between global and local features, it places too much emphasis on global information, resulting in the loss of local information. Therefore, the local relationship attention mechanism is introduced to strengthen local feature information.
[0051] Step S33: Segment the local feature map to obtain X 1 and X 2 , X 1 After the convolution operation, 2 After superposition, the PRA formula of the local relation attention mechanism module is as follows:
[0052] (6)
[0053] in, is the input mask distribution, is the convolution kernel, is the bias of the convolution operation;
[0054] In the embodiment of the present invention, X 1 The structure is , X 2 The structure is . The core advantage of the Partial Relation Attention (PRA) mechanism of the present invention is that it can flexibly cope with the situation of missing data. When processing images, missing data is a common phenomenon. Traditional convolution methods usually assume that the data is complete and do not work well when faced with missing data. The Partial Relation Attention mechanism intelligently identifies the validity of the data and only applies the convolution kernel to the existing data points, while ignoring the missing data points. When performing convolution, the Partial Relation Attention mechanism checks the data points within each window. For existing data points, it performs regular convolution; for missing data points, no calculation is performed. This mechanism enables the algorithm to change according to the actual situation of the data in each actual area of action, thereby enhancing the processing ability of local features and making more effective use of local feature information.
[0055] Step S34: and Superposition, get the local context feature map:
[0056] (7)
[0057] in, The local context feature map is the output of the local feature map through the dual-path attention mechanism module;
[0058] Step S35: Repeat steps S32 to S34 for the global feature map to obtain a global context feature map: perform a global relationship pooling operation with the local context feature map to finally obtain a local-global relationship feature map.
[0059] The local-global relationship feature map enhances the network's ability to capture global context features through the global context mechanism, which helps to understand the overall structure and context information of the image. At the same time, PRA can handle local missing or damaged areas in the image and effectively use the remaining complete information for feature extraction. The dual-path attention mechanism formed by the combination of GCA and PRA can better focus on the context information features of global and local features, solve the problem of insufficient context information feature extraction, and add context information features to local-global features, improving the accuracy of the model.
[0060] Figure 3 The structural diagram of the dual-path attention mechanism module is shown.
[0061] In one embodiment, the above step S4: converting the local-global relationship feature map into a feature vector of a predetermined dimension for representing pedestrians, optimizing the calculation method of the Euclidean distance by the angle difference between the vectors, and determining whether two pedestrians are the same person, specifically includes:
[0062] Step S41: converting the local-global relationship feature map into a feature vector of a predetermined dimension, which is used to represent the unique feature of the pedestrian;
[0063] The predetermined dimension in the embodiment of the present invention is set to 256.
[0064] Existing pedestrian re-identification methods use Euclidean distance to calculate two pedestrian feature vectors and The similarity between them is calculated as follows:
[0065] ;
[0066] in, yes and The Euclidean distance between yes and In between The values in the dimensions, is the dimension of the feature vector. In person re-identification, a smaller Euclidean distance usually indicates a higher similarity between the feature vectors of two pedestrians. By calculating the Euclidean distance between the feature vectors of the pedestrian to be identified and each pedestrian in the database, the pedestrian with the smallest Euclidean distance is the most similar pedestrian matched. Since the Euclidean distance is very sensitive to the length of the feature vector, the feature vector is processed as follows.
[0067] Step S42: Convert the distance difference between the feature vectors into the angle difference between the feature vectors:
[0068] (8)
[0069] in, is the eigenvector and The angle between
[0070] Step S43: If If it is greater than a preset threshold, it is considered to be the same pedestrian, otherwise, they are different pedestrians.
[0071] Figure 4 The overall framework diagram of the pedestrian re-identification method based on the dual-path attention mechanism is shown.
[0072] Embodiment 2
[0073] like Figure 5 As shown, the embodiment of the present invention provides a pedestrian re-identification system based on a dual-path attention mechanism, comprising the following modules:
[0074] The data collection and preprocessing module 51 is used to construct a community pedestrian dataset and preprocess the community pedestrian dataset;
[0075] Feature extraction module 52: used to build a ResNetRFA network based on channel-by-channel convolution, receptive field attention mechanism and residual network, input the community pedestrian dataset into the ResNetRFA network, extract the key features of pedestrian targets and generate a receptive field feature map ;
[0076] The feature fusion module 53 is used to After being divided into several local feature maps The two-way attention mechanism module is input together, and the local-global relationship feature map is obtained through the global context attention mechanism module and the local relationship attention mechanism module;
[0077] Pedestrian matching module 54: used to convert the local-global relationship feature map into a feature vector of a predetermined dimension for representing pedestrians, and optimize the calculation method of the Euclidean distance by the angle difference between the vectors to determine whether two pedestrians are the same person.
Claims
1. A pedestrian re-identification method based on a dual-path attention mechanism, characterized in that: include: Step S1: construct a community pedestrian dataset and preprocess the community pedestrian dataset; Step S2: Construct a ResNetRFA network based on channel-by-channel convolution, receptive field attention mechanism and residual network, input the community pedestrian dataset into the ResNetRFA network, extract the key features of pedestrian targets and generate a receptive field feature map ; Step S3: After being divided into several local feature maps The two-way attention mechanism module is input together, and the local-global relationship feature map is obtained through the global context attention mechanism module and the local relationship attention mechanism module; Step S4: converting the local-global relationship feature map into a feature vector of a predetermined dimension for representing pedestrians, optimizing the calculation method of the Euclidean distance by the angle difference between the vectors, and determining whether two pedestrians are the same person.
2. The pedestrian re-identification method based on the dual-path attention mechanism according to claim 1, characterized in that: The step S2: constructing a ResNetRFA network based on channel-by-channel convolution, receptive field attention mechanism and residual network, inputting the community pedestrian dataset into the ResNetRFA network, extracting key features of pedestrian targets and generating a receptive field feature map , specifically including: The ResNetRFA network is constructed based on the residual network. The improvements are: 1) At Layer 0, the original convolution kernel is replaced with channel-by-channel convolution to perform convolution operation on the input image; 2) The receptive field attention mechanism RFA is added to the Bottleneck of Layer 4, where RFA uses a sliding window with shared parameters to extract feature information based on the convolution operation; Assume the input feature map is ,in Respectively represent the number of channels, height, and width of the feature map; , then the convolution operation to extract feature information from each receptive field slider can be expressed as: (1) in, Represents the value obtained by each convolution slider after calculation, Indicates the pixel value of the corresponding position in each slider. represents the convolution kernel, represents the number of parameters in the convolution kernel, Represents the total number of receiving domain sliders; It can be seen from formula (1) that the features at the same position in each convolution slider share the same parameters ; Since interacting with each receptive domain feature will result in additional computational overhead, in order to minimize the computational overhead and the number of parameters, AvgPool is used to aggregate the global information of each receptive domain feature, convolution operation is used for information interaction, and Softmax is used to emphasize the importance of each feature in the receptive domain feature. Therefore, the calculation of RFA can be expressed as formula (2): (2) in, Represents a size of Grouped convolution; represents the size of the convolution kernel, represents normalization, represents the input feature map, Represents the receptive field feature map.
3. The pedestrian re-identification method based on the dual-path attention mechanism according to claim 2 is characterized in that: Step S3: After being divided into several local feature maps The two-way attention mechanism module is input together, and the local-global relationship feature map is obtained through the global context attention mechanism module and the local relationship attention mechanism module, which specifically includes: Step S31: The image is evenly divided into six parts according to the height, and the width remains unchanged, so as to obtain six local feature maps and one global feature map; the local feature map and the global feature map are respectively input into the dual-path attention mechanism module to extract the local-global relationship feature map; wherein the dual-path attention mechanism module includes two branches: a global context attention mechanism module GCA and a local relationship attention mechanism module PRA; Step S32: First, the local feature map is input into the global context attention mechanism module GCA: (3) in, Represents the global context features of the local feature map after GCA, and the global attention pooling is performed through the weighted The weighted average groups the features at all positions together to obtain the global context feature, represents the feature transformation that captures channel dependencies, represents the feature of the position to be queried in the local feature map, express The characteristics of all possible positions except is the total number of all positions in the feature map, represents the fusion function that aggregates the global context features into the features at each position; in, After expansion, it is as follows: (4) (5) in, is the weight of the global attention pool; is an element in the input feature map; represents a specific position in the feature map, Representation and The corresponding feature weights, represents the weight of the first convolutional layer, represents the weight of the second convolutional layer; Representation layer normalization, represents the activation function; Step S33: Segment the local feature map to obtain X1 and X2. X1 is convolved and then superimposed with X2. The formula of the local relation attention mechanism module PRA is as follows: (6) in, is the input mask distribution, is the convolution kernel, is the bias of the convolution operation; Step S34: and Superposition, get the local context feature map: (7) in, A local context feature map which is the output of the local feature map through the dual-path attention mechanism module; Step S35: Repeat steps S32 to S34 for the global feature map to obtain a global context feature map: perform a global relationship pooling operation with the local context feature map to finally obtain a local-global relationship feature map.
4. The pedestrian re-identification method based on the dual-path attention mechanism according to claim 3 is characterized in that: The step S4: converting the local-global relationship feature map into a feature vector of a predetermined dimension for representing pedestrians, optimizing the calculation method of the Euclidean distance by the angle difference between the vectors, and determining whether two pedestrians are the same person, specifically includes: Step S41: converting the local-global relationship feature map into a feature vector of a predetermined dimension for representing the appearance of the pedestrian; Step S42: Optimize the calculation method of the Euclidean distance by the angle difference between the vectors to determine whether the two pedestrians are the same person: (8) in, is the eigenvector and The angle between Step S43: If If it is greater than a preset threshold, it is considered to be the same pedestrian, otherwise, they are different pedestrians.
5. A pedestrian re-identification system based on a dual-path attention mechanism, characterized in that: Includes the following modules: A data collection and preprocessing module, used to construct a community pedestrian dataset and preprocess the community pedestrian dataset; Feature extraction module: used to build a ResNetRFA network based on channel-by-channel convolution, receptive field attention mechanism and residual network, input the community pedestrian dataset into the ResNetRFA network, extract the key features of pedestrian targets and generate a receptive field feature map ; Feature fusion module is used to After being divided into several local feature maps The two-way attention mechanism module is input together, and the local-global relationship feature map is obtained through the global context attention mechanism module and the local relationship attention mechanism module; Pedestrian matching module: used to convert the local-global relationship feature map into a feature vector of a predetermined dimension to represent pedestrians, optimize the calculation method of the Euclidean distance by the angle difference between the vectors, and determine whether two pedestrians are the same person.