A Person Re-identification Method Based on Enhancing Target Structure Relationships

By introducing a structural enhancement stackable attention module in ResNet50, the feature learning ability of pedestrian re-identification network is enhanced, and the challenge of pedestrian re-identification in cross-camera scenarios is solved, and higher accuracy and performance are achieved.

CN114529942BActive Publication Date: 2025-06-27CHONGQING UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210071827.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-06-27
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

The existing pedestrian re-identification algorithm faces challenges such as feature differences, background differences, occlusion conditions and diverse postures in cross-camera scenarios, resulting in unsatisfactory recognition results.

Method used

ResNet50 is used as the backbone network, and a structurally enhanced stackable attention module is added to its residual stacking module. Through the structurally enhanced vector learning module and structurally separated convolution module, the characteristics at each level are strengthened, and the ability of networks to learn distinguishable features is improved.

Benefits of technology

Through structural enhancement of stackable attention modules, more distinctive target structure features are extracted, which effectively suppresses unstructured noise, improves the accuracy of pedestrian re-identification, and achieves competitive performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529942B_ABST
    Figure CN114529942B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of pedestrian recognition, and in particular to a pedestrian re-identification method based on enhanced target structure relationship. It includes selecting ResNet50 as the backbone network; adding a structure-enhanced stackable attention module in the residual stacking module of ResNet50 to strengthen the features at each level, so as to improve the network's ability to learn distinguishable features through the structure enhancement factor; using label-smoothing cross-entropy loss combined with triplet loss to train the model. The structure-enhanced stackable attention module of the present invention can help the neural network establish the connection between target structure features by perceiving global structure information through local information, and strengthen the structure information, so as to refine more distinguishable target structure features; the modeling method establishes the interaction information between structures, making the structure information no longer independent, strengthening the interaction between the structure information and its own representation vector, and the final enhancement factor is more delicate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian recognition, and particularly to a pedestrian re-identification method based on enhanced target structure relationship. Background Art

[0002] Pedestrian re-identification is a fundamental but important image recognition task, which realizes target re-identification across scenarios such as cross-cameras by finding the target closest to the query target image among a large number of candidate pedestrian images. The accurate recognition of this task makes it possible to specifically implement the pedestrian re-identification task in a large number of video images. Similar to the face recognition task, pedestrian re-identification aims to study the extraction of the discriminative feature embedding expressions of pedestrian targets under different perspectives, different times, and different backgrounds, and uses a metric function to measure the similarity of feature embeddings between the query image and the candidate target image. However, due to the various differences presented by pedestrian target images in cross-camera scenarios, the pedestrian re-identification task faces greater challenges, such as Figure 1 in, the feature differences, background differences, target occlusion situations, and diverse pedestrian postures of the target caused by cross-cameras are all important factors that need to be considered in pedestrian re-identification algorithms.

[0003] Traditional pedestrian re-identification methods mainly use manually designed understandable features such as color or gradient to represent the target image. Existing literature A (Ojala T, Pietikainen M, Maenpaa T. Multiresolution grayscale and rotation invariant texture classification with local binary patterns[J]. IEEE Trans on Pattern Analysis and Machine Intelligence, 2002, 24(7): 971-987.) uses a specific local binary representation (LBP) to extract features and perform histogram statistics, and this histogram has been proven to be a very powerful texture representation. Existing literature B (LOWE D G. Distinctive image features from scale invariant keypoints[J]. International Journal of Computer Vision, 2004, 60(2): 91-110) proposed a scale invariant feature (SIFT) and it has been widely applied in the field of image recognition, such as tasks like image retrieval and image stitching. Literature C (Dalal N, Triggs B. Histograms of oriented gradients for human detection[C] / / Proc of IEEE Conference on Computer Vision and Pattern Recognition. San Diego, USA: IEEE Press, 2005: 886-893.) proposed a feature for representing local gradient (HOG) statistical representation. However, a single statistical feature cannot be applied to complex image tasks such as pedestrian re-identification. This task often involves difficulties such as environmental understanding, scale transformation, and inconsistent image data quality. Researchers often need to fuse multiple low-level local features and global features to represent the target image, and use various powerful classifiers to learn the optimal classification weights.Literature D (Layne R, Hospedales T M, Gong S, et al. Person reidentification by attributes[C] / / Proc of British Machine Vision Conference. Britain: BMVA Press, 2012: 24.1-24.11) believes that due to reasons such as occlusion, bottom-up features are not robust enough, and proposes a method of fusing intermediate semantic features and low-level features and using a support vector machine for classification.

[0004] Although researchers have studied and designed effective pedestrian target feature expressions, such algorithms are still limited by challenges such as complex display scenes, lens switching, pose transformation, and even changes in clothing, making the recognition algorithms designed with handcrafted features ineffective.

[0005] In recent years, deep learning was first applied to image classification. The use of deep learning technology to recognize handwritten digits was demonstrated in Document E (LeCun Y, Boser B, Denker JS, et al. Backpro-pagation app-lied to handwritten zip code recog-nition[J]. Neuralcomputation, 1989, 1(4):541-551). The entire neural network was used to extract the global information of the whole image and generate feature vectors for classification. For simple datasets such as handwritten digits, an accuracy rate of 95% was achieved. In order to generate feature vectors for distinguishing pedestrians, some researchers only used ResNet as the benchmark network for training pedestrian re-identification, and used loss functions such as softmax and triplet to train on datasets such as Market1501, and obtained very good accuracy. Document F (SUN Yifan, ZHENG Liang, DENG Weijian, et al. SVDNet for pedestrian retrieval[C] / / Proc of IEEE International Conferenceon Computer Vision. Washington D.C., USA:IEEE Press, 2017:3800-3808) based on the correlation hypothesis of convolutional layer weights, believes that the uncertainty of data distribution will cause redundancy of discriminative features and weaken the discriminative features. It is proposed to impose an orthogonal constraint on the neural network weights and perform decorrelation iterative training on the network weights by means of singular value decomposition. In this way, orthogonalized weight learning is used to improve the discriminability of features. Document G (ZHENG Zhedong, YANG Xiaodong, YU Zhiding, et al. Joint Discriminative and Generative Learningfor Person ReIdent-ification[C] / / Proc of the IEEE Conference on ComputerVision and Pattern Recognition. Long Beach, CA, USA:IEEE Press, 2019:2138-2147) innovatively combines the idea of generative adversarial network (GAN) into the pedestrian re-identification network, and uses a method combining generative and discriminative methods to enhance the network's learning of data. To a certain extent, it solves the recognition difficulties caused by problems such as cross-domain data and changes in pedestrian structures.Literature H (Dai, Zuozhuo, Mingqiang Chen, Xiaodong Gu, et al. Batch Feature Erasing for Person Re-identification and Beyond [C] / / Proc of the IEEE / CVF International Conference on Computer Vision. Seoul, Korea: IEEE Press, 2019: 3691-3701) believes that the occlusion and pose variation problems in person re-identification suppress the learning of some key information. A batch block discard module is proposed, which randomly discards sub-blocks at a certain position of the feature map to remove part of the information. The network is trained by splicing the feature maps of the base branch and the batch discard branch to represent the image features, strengthening the learning of key features. Literature I (SUN Yifan, ZHENGLiang, YANG Yi, et al. Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline) [C] / / Proc of the European Conference on Computer Vision. Berlin, Germany: Springer, 2018: 480-496) performs the extraction of image features in blocks on the strong baseline network ResNet, hoping to improve the person re-identification task from the perspective of local features. The authors propose a block-based convolutional neural network for the person re-identification task, strengthening the focused learning of image feature blocks. And a refined part pooling is proposed to adjust for different images to solve the problem of inconsistent boundaries between feature blocks and semantic blocks, further improving the performance of the person re-identification network. Global features, as representation vectors, can completely obtain the global information of the entire image, but global features are prone to carry non-important information such as the environment and noise. The idea of using local information as the target representation vector in this literature is more in line with cognition, and the cognitive idea from point to whole can filter out some non-important information.

[0006] Existing person re-identification works have focused on the extraction of discriminative features through data augmentation and ordinary location and channel attention, but have ignored the potential of the relationship information between channel structures to enhance the learning of structural features. Summary of the Invention

[0007] The purpose of the present invention is to provide a person re-identification method based on enhancing target structure relationships.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A pedestrian re-identification method based on enhanced target structure relationship, comprising the following steps:

[0010] Select ResNet50 as the backbone network;

[0011] Add a structure-enhanced stackable attention module in the residual stack module of ResNet50 to strengthen the features at each level, so as to improve the network's ability to learn distinguishable features through the structure enhancement factor;

[0012] Use label smoothing cross-entropy loss combined with triplet loss to train the model.

[0013] Further preferably, the structure-enhanced stackable attention module includes:

[0014] A structure-enhanced vector learning module for learning an embedding vector containing structural attention;

[0015] A structure separation convolution module for learning a specific mapping of the structure attention embedding vectors obtained from different structures.

[0016] Further preferably, the input of the structure-enhanced vector learning module is the feature map of a certain level of ResNet

[0017] X is obtained through an element rearrangement operation to get the feature map Xinput;

[0018] Xinput passes through three 1×1 convolutions with an input dimension of H×W and output dimensions of C1, C1, and C2 respectively, batch normalization, and ReLU activation function, and further obtains a query vector a response vector and an embedding vector representing its own information

[0019] The query vector is calculated as follows:

[0020] Q = ReLU(BN(W conv X input ));

[0021] The response vector is calculated as follows:

[0022] R = ReLU(BN(W conv X input ));

[0023] The embedding vector The calculation of

[0024] E = ReLU(BN(W conv X input ));

[0025] The request vector q i (q i ∈Q) is multiplied element - by - element with all response vectors r j (r j ∈R), and then a 1×1 convolution is performed on it to obtain the relationship response vector of the i - th channel;

[0026] Considering the bidirectional relationship between a specific channel representing structural information and other channels, that is, stacking the active response vector and the passive response vector of a certain channel as the structural relationship representation vector S i of channel i, the structural relationship vector S can be obtained by the following formula:

[0027]

[0028] where Φ(q i , r j ) = Conv(q i ×r j ), q i ×r j is the element - by - element multiplication of q i , r j , Conv is a convolution calculation with a 1×1 convolution kernel. At this time, S is a tensor of dimension C×(2C)×1. S then passes through a 1×1 convolution with an input channel of 2×C and an output channel of C2, batch normalization, and the ReLU activation function to obtain the structural vector. The structural vector is multiplied element - by - element with the embedding vector E to obtain the enhanced embedding vector E′. The calculation is as follows:

[0029] E′=(ReLU(BN(W conv S)))×E.

[0030] Further preferably, the structure - separation convolution module performs a separation mapping on the enhanced embedding vector E′ representing the specific structural relationship information of each channel;

[0031] For a feature map of dimension C, C convolution kernels W i , i∈C are used to separately learn the mapping relationship from the enhanced embedding vector E′ of each channel to the attention - enhancement factor. This mapping is a dedicated mapping for the enhanced embedding vector of the structural relationship information in a certain channel, rather than a unified measurement of the enhanced embedding vector;

[0032] The output formula of the structure - separation convolution module is:

[0033]

[0034] Further preferably, a structure-enhanced stackable attention module is added to the residual stack module of ResNet50 to strengthen the features at each level, so as to improve the network's ability to learn distinguishable features through the structure enhancement factor.

[0035] Further preferably, the model is trained using label-smoothing cross-entropy loss combined with triplet loss, and the combined loss function of the combined model is:

[0036] L loss = L triplet + L LSCE

[0037] L triplet = max(d(a, p) - d(a, n) + margin, 0)

[0038]

[0039]

[0040] where the function d is a distance metric function, and the Euclidean distance metric is adopted; a, p, and n respectively represent the query, matching, and non-matching image feature vectors; batch is the batch size of one training, class is the number of people in the dataset, y_pred is the two-dimensional vector predicted by the network; y_pred ij represents the probability that the i-th sample belongs to the j-th class; q j is the smoothing factor, λ is the smoothness, taking 0.1; y_label i is the class to which the i-th sample belongs.

[0041] The present invention has at least the following beneficial effects:

[0042] The present invention proposes a structure-enhanced stackable attention module, which can help the neural network establish the connection between target structure features by perceiving global structure information through local information and strengthen the structure information, so as to refine more distinguishable target structure features; the SES module mines the relationship structure information through the interaction modeling of the request vector and the response vector of the channel structure information, and then interacts and strengthens the calculation with its own representative vector to strengthen the structure representation features; finally, structure separation convolution is performed on the structure representation features to obtain a structure enhancement factor to strengthen the original features; this modeling method establishes the interaction information between structures, makes the structure information no longer independent, strengthens the interaction between the structure information and its own representation vector, and the final enhancement factor is more delicate; the structure features have a larger influence domain, can suppress unstructured noise, and thus learn distinguishable features with structure-enhanced information.

[0043] Through a large number of comparative experiments and neural network visualization techniques, the present invention verifies the enhanced feature learning ability of the structure-enhanced stackable attention module on the general pedestrian re-identification datasets CUHK03L and Market1501, achieves competitive performance on the CUHK03L and Market1501 datasets, and finally reaches an average precision mean of 78.0% and a rank-1 accuracy of 81.3% on the CUHK03L dataset, and realizes an average precision mean of 88.2% and a rank-1 accuracy of 96.2% on the Market1501 dataset. Brief Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 It is an example diagram of the challenges of the current pedestrian re-identification task;

[0046] Figure 2 It is a schematic diagram of the overall network structure of SESNet;

[0047] Figure 3 It is a schematic diagram of the structure-enhanced stackable attention module;

[0048] Figure 4 It is a schematic diagram of convolutional feature extraction;

[0049] Figure 5 It is a schematic diagram of structure-separated convolution;

[0050] Figure 6 It is a schematic diagram of sample data;

[0051] Figure 7 It is a diagram of the SESNet Top-5 query results;

[0052] Figure 8 It is a Grad-CAM visualization comparison heat map. Detailed Embodiments

[0053] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following further details the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] 1. Attention-Enhanced Network Based on Structural Features

[0055] In order to extract more discriminative features for characterizing pedestrians, the present invention designs a structure-enhanced stackable attention module based on the target structure relationship to learn global structure information in feature maps at different levels for strengthening the pedestrian representation vector. This module can be easily stacked after the feature maps of any network to enhance the feature expression of the feature maps. The structure-enhanced stackable attention module mainly consists of a structure-enhanced vector learning module and a structure-separated convolution module. Among them, the structure-enhanced vector learning module is used to learn the embedding vector containing structural attention, while the structure-separated convolution module is used to learn the specific mapping of the structural attention embedding vectors obtained from different structures.

[0056] 1.1. Overall network structure

[0057] The strong baseline network ResNet50 has good feature extraction ability and performs excellently in the pedestrian re-identification task. The present invention conducts research on the structure-enhanced stackable attention module based on ResNet50. ResNet50 can be divided into 5 sub-modules, consisting of the first low-level feature learning module and the latter four residual stacking modules. The latter four residual stacking modules perform convolutional calculations on the features step by step to extract semantic features at different levels for specific image processing tasks. The structure-enhanced stackable attention (Structure-Enhanced Stackable attention) module is added to the ResNet residual stacking modules respectively to strengthen the features at each level, so as to improve the network's ability to learn discriminative features through the structure-enhanced factor. The overall network structure (SESNet) is as Figure 2 shown. The pedestrian image passes through the ResNet backbone network enhanced by the SES module to obtain a feature map of 1024×16×8, and then through average pooling to obtain a 1024-dimensional feature vector for characterizing pedestrians. Finally, the metric distance is calculated by this vector and other pedestrian representation vectors.

[0058] 1.2. Structure-enhanced vector learning module

[0059] Convolutional neural networks can learn structural features at different levels to form feature map channels through the convolution calculation between the convolution kernel and the feature map. Different channels represent different structural feature information. The literature (Zeiler M D, Fergus R. Visualizing and understanding convolutional networks[C] / / Proc of European conference on computer. vision Berlin, Germany: Springer, 2014: 818-833.) observes the influence degree of original image pixels on features at different levels through the visualization technology of neural network feature maps. The low-level module undergoes less convolution calculation and is responsible for learning low-level semantic structure information, such as color, texture, and simple shape information, and cannot remove low-level image noise. The high-level module stacks convolution calculations to extract higher levels from low-level features and can obtain richer semantic structure information, such as hands and legs in different poses. Moreover, high-level information can suppress low-level environmental information and noise. In the past, researchers in person re-identification studied the high-level information learning ability brought by different network structures through experiments to learn powerful person representation vectors. However, they rarely paid attention to the impact of the interaction relationship between structural features at different levels on the learning of the final representation vector. The present invention proposes a structure-enhanced stackable module to enhance the learning ability of ResNet50 for structural features at different levels. Figure 3 The structure-enhanced stackable attention module proposed by the present invention is shown. The module consists of a structure-enhanced vector learning module and a structure-separated convolution module.

[0060] The input of the structure-enhanced vector learning module is the feature map of a certain level of ResNet Each channel of this feature map has learned specific image structure information. By establishing the interaction relationship between different channels, the network's mining of the structural relationship of the person image is enhanced to strengthen the network's learning ability for distinguishable features. X undergoes an element rearrangement (reshape) operation to obtain the feature map Xinput. Xinput passes through three 1×1 convolutions with an input dimension of H×W and output dimensions of C1, C1, and C2 respectively, batch normalization (Batch Normalization), and the ReLU activation function to further obtain the query vector the response vector and the embedding vector representing its own information That is, the query vector The calculation is as follows:

[0061] Q = ReLU(BN(Wconv X input )) (1)

[0062] The response vector R and the embedding vector E can be calculated from formula (1). Then, the request vector q i (q i ∈Q) is multiplied element-wise with all response r j (r j ∈R) vectors, and then a 1×1 convolution is performed on them to obtain the relationship response vector for the i-th channel. To maximize the potential of mining relationship information, the present invention will consider the bidirectional relationship between a specific channel representing structural information and other channels, that is, stacking the active response vector and the passive response vector of a certain channel as the structural relationship representation vector S i of channel i. The structural relationship vector S can be obtained from Equation 2.

[0063]

[0064] Where Φ(q i , r j ) = Conv(q i ×r j ), q i ×r j is the element-wise multiplication of q i , r j , and Conv is the convolution calculation with a 1×1 convolution kernel. At this time, S is a tensor of dimension C×(2C)×1. S then passes through a 1×1 convolution with an input channel of 2×C and an output channel of C2, batch normalization, and the ReLU activation function to obtain the structure vector. The structure vector is multiplied element-wise with the embedding vector E to obtain the enhanced embedding vector E':

[0065] E' = (ReLU(BN(W conv S)))×E (3)

[0066] 1.3. Structural Separation Convolution Module

[0067] The strength of a convolutional neural network lies in its ability to extract a certain semantic feature of corresponding pixels through the multiplication and addition operations of the convolution kernel and the feature image pixels, and this feature is automatically learned by the neural network to adaptively adjust to the most suitable abstract feature extractor. As Figure 4As shown in the figure, convolution calculation extracts the features of the original image by gradually moving the convolution kernel to the right and down in sequence to obtain a new feature map. This feature map represents the intensity of a similar structure to the convolution kernel. When the convolution kernel extracts the features in the upper left corner of the original image, the features at this position are exactly the same as the shape of the convolution kernel, and the maximum response value of 3 is obtained. When the convolution kernel extracts the features in the middle of the image, due to the existence of a similar structure, the response value of 2 is obtained. When the convolution kernel extracts the features in the lower right corner of the image, only a very small part of the features are similar, and the response value of 1 is obtained. There is no similar structure at the remaining positions, and the response value of 0 is obtained for all of them. It can be seen that convolution calculation is used to extract certain common features in the image to generate a feature map representing higher-level information. In practical applications, the neural network will perform backpropagation based on the difference metric between the sample inference result and the label, and then adjust the convolution kernel so that the convolution kernel adaptively learns a specific structural information response map to complete the extraction of image features. After the learning and inference of the network, each channel of the convolutional neural network stores a certain semantic structure information of the image.

[0068] Since the input images for person re-identification are all pedestrians, the image semantics all have similar structural information. For tasks such as object detection, the number of objects, the shape of the objects, and the background information in each image are all different. The structural information learned by each layer of channels in the object detection network has structural inconsistency, which is the response information for different structures of specific images. Regarding the structural consistency of the input for the person re-identification task, the responses of each layer of the feature map are similar. The present invention believes that there are structural relationship stability and relationship mapping difference in the connection between the channel structure information.

[0069] Structural relationship stability means that the relationship between each channel and other channels also has structural relationship stability in different inputs. Relationship mapping difference means that there is a mapping difference when the relationship embedding vectors between each channel and other channels are measured by the unifying measure as the strengthening factor. Convolution calculation is used to extract the common features of the feature map. However, the present invention believes that there is a difference in the measurement of structural relationship information. Based on this, the present invention designs a structure separation convolution module to perform a separation mapping on the strengthened embedding vector E' representing the specific structural relationship information of each channel. Figure 5 The structure of the structure separation convolution module is shown. For the feature map of dimension C, the present invention uses C convolution kernels W i,i ∈ C to separately learn the mapping relationship from the strengthened embedding vector E' of each channel to the attention strengthening factor. This mapping is a dedicated mapping for the strengthened embedding vector of the structural relationship information in a certain channel, rather than a unified measurement of the strengthened embedding vector.

[0070]

[0071] 1.4. Loss Function

[0072] Although the number of pedestrian categories in the pedestrian re-identification dataset is determined, the label smoothed cross-entropy loss can be used for the dataset to meet the training requirements. However, the pedestrian re-identification task often targets the pedestrian detection task in the open world. Therefore, the present invention uses the label smoothed cross-entropy loss (Label Smoothed Cross Entropy) combined with the triplet loss to train the model. The triplet loss makes the distance between the vectors representing the same target closer and the distance between the vectors representing different targets farther through the maximum margin factor Margin, so as to strengthen the learning ability of the neural network for discriminative feature vectors. The combined loss function can be expressed by Equation 5:

[0073] L loss = L triplet + L LSCE (5)

[0074] L triplet = max(d(a, p) - d(a, n) + margin, 0) (6)

[0075]

[0076]

[0077] where the function d is a distance metric function, and the Euclidean distance metric adopted in the present invention is used. a, p, and n represent the query, matching, and non-matching image feature vectors respectively. batch is the batch size of one training, class is the number of pedestrians included in the dataset, and y_pred is the two-dimensional vector predicted by the network. y_pred ij represents the probability that the i-th sample belongs to the j-th class; q j is the smoothing factor, λ is the smoothness, and λ takes 0.1. y_label i is the class to which the i-th sample belongs.

[0078] 2. Experimental Results and Analysis

[0079] In order for the network to find a suitable search space at the initial stage of training to ensure the stability of the deep layer of the model, the warmup training method is adopted at the initial stage of training. As the number of batches increases, the learning rate gradually decays exponentially, and the decay rate is reduced to half of the previous value for every 50 complete iterations of the dataset training. The experiment sets the random number seed to ensure that the initialization parameters are consistent. The experimental results are the average values of five experiments to exclude random results as much as possible. 2.1.

[0081] Experimental Dataset

[0082] To verify the robustness of the re-identification model, the present invention verifies the discriminative feature learning ability of the proposed network structure on the classical pedestrian re-identification datasets Market1501 and CUHK03. Both of the above-mentioned classical datasets have samples with different image qualities, and the pedestrians in the samples have diverse postures, diverse backgrounds, and different sizes, which well reflect the diversity of pedestrian image samples in the real world. The Matket1501 dataset was collected by Tsinghua University in summer. This dataset captured 1501 pedestrians, each of whom was photographed by different cameras, and a total of 32,668 detection frames were used to frame and identify the pedestrians. The training set of this dataset contains 751 pedestrians with a total of 12,936 multi-camera images, and the test set contains 750 pedestrians with a total of 19,732 multi-camera images. The CUHK03 dataset was collected on the campus of the Chinese University of Hong Kong, China.

[0083] This dataset is divided into detected, labeled, and testsets datasets. Among them, the pedestrian bounding boxes in the detected dataset are detected by a detector, and the pedestrian bounding boxes in the labeled dataset are manually annotated. The experiment was conducted on a total of 14,096 pedestrian images in the labeled dataset (CUHK03L). Each pedestrian was photographed by different cameras. The training set contains 767 pedestrians with a total of 7368 multi-camera images, and the test set contains 700 pedestrians with a total of 6728 multi-camera images. Figure 6 The dataset example diagram is shown. Table 1 shows the sample distributions of the Market1501 and CUHK03L datasets.

[0084] Table 1 Market1501 and CUHK03L Datasets

[0085]

[0086] 2.2. Experimental Preparation

[0087] The algorithm was implemented under the PyTorch (V 1.7.0) deep learning framework, and the operating system was ubuntu16.04. The hardware configuration is as follows: the CPU is an Intel Core i7-7700@3.6GHz ×8, the GPU is an NVIDIA GTX1080Ti×2, and the memory is 32GB. The inference batch size is 64, and the number of iterations is 500. The Stochastic Gradient Descent (SGD) optimization algorithm was used for model training, with a base learning rate of 0.0008, which was changed according to the warmup or decay strategy for specific batches.

[0088] 2.3. Experimental Evaluation Criteria

[0089] The experimental results were evaluated using the mean Average Precision (mAP) metric and the rank-1 level of the Cumulative Matching Characteristics (CMC) metric. During testing, the query images and candidate images within the test set were specified, and feature extraction was performed on all samples in the test set. The features of the query images were compared with those of all candidate images for similarity measurement. The Cumulative Matching Characteristics rank-n refers to the accuracy rate of having correct samples among the top n images ranked by similarity to the query image. The mean Average Precision refers to the mean of the areas (Average-Precision) representing the mean class precision under the precision-recall curve calculated for all samples. The precision (Pre) and recall (Rec) were calculated using equations (8) and (9).

[0090]

[0091]

[0092] Among them, TP represents True Positive, FP represents False Positive, and FN represents False Negative.

[0093] 2.4. Analysis of Experimental Results

[0094] The method of the present invention was verified on the Market1501 and CUHK03L datasets. The experimental results Figure 7 show the rank-5 query results of random samples. The blue block diagram represents the query image, the red solid block diagram represents the correct query result, and the yellow dashed block diagram represents the incorrect query result. It can be seen from the figure that SESNet can accurately find the pedestrian sample corresponding to the query image. Although there are mis-matched samples in the legend, the appearance and pose of this sample are very similar to the query image, which also reflects the accuracy of SESNet in finding features and the great challenges existing in this task from the side.

[0095] The present invention also conducted ablation experiments on the two sub-modules of the SES module respectively to verify the effectiveness of the SES module, and compared the performance of the feature learning network constructed by the strong baseline (Baseline) network ResNet. The experimental results are shown in Table 2, where SES- represents the SES module without structured separable convolution.

[0096] Table 2 Ablation Experiment of SES Module

[0097]

[0098] As can be seen from Table 2, when the SES module did not adopt structure-separated convolution, SES- still successfully established the relationship between target structures through the structure-enhanced vector module and ordinary 1×1 convolution, enabling the network to learn the attention enhancement factor and greatly improving the accuracy. Since SES- did not perform separate mapping on the enhancement vectors of different structures and could not fully exploit the relationship between structures, its accuracy was slightly lower than that of SES. By adopting structure-separated convolution for the SES- module, that is, the SES module, the accuracy was improved again after the enhancement of the pedestrian feature structure by the SES module. Finally, Baseline achieved a 3.5% mAP improvement and a 4.0% rank-1 accuracy improvement on the CUHK03L dataset. On the Market1501 dataset, it obtained a 4.5% mAP improvement and a 2.0% rank-1 accuracy improvement.

[0099] Table 3 Metrics of SESNet and Previous Methods on Market1501 and CUHK03L

[0100]

[0101] Table 3 compares the performance of different pedestrian re-identification methods on the Market1501 dataset and the CUHK03L dataset. Compared with the fine-grained feature method based on the fusion of global features and local features, SESNet strengthens the important features for the pedestrian re-identification task through the attention enhancement factor, thus no longer generating local feature representations through other branch networks and avoiding the problem of feature misalignment. On the other hand, compared with other attention-based methods, SESNet has a more fine-grained structural relationship and specific relationship semantics between different structural features. Therefore, it learns a more discriminative enhancement factor to enhance the structural feature representations at different levels.

[0102] To further verify the feature extraction effect of SESNet, the present invention uses the Grad-CAM visualization method to enhance the interpretability of the neural network and performs feature visualization on the high-level feature maps of the baseline network ResNet and the structure-enhanced attention network SESNet. Figure 8 Shows the regions of interest of the sample features for ResNet and SESNet. As can be seen from the heatmap, the regions of interest of ResNet change with the pose and are difficult to understand. However, even when the person presents different poses, SESNet can still stably notice the discriminative parts and suppress the influence of irrelevant regions.

[0103] Based on the above, it can be concluded that:

[0104] In order to enable the neural network to learn more discriminative representation vectors, the present invention proposes a structure-enhanced stackable attention module to strengthen the feature learning ability of the strong baseline network ResNet50, and uses a triplet loss function that strengthens the inter-class discriminative feature learning ability and jointly weakens the label smoothing cross-entropy loss that affects incorrect data to enhance the network learning ability. The structure-enhanced stackable attention module strengthens the structural representation vector by learning a more fine-grained attention enhancement vector, and uses structural separable convolution to separately map different enhanced structure vectors to obtain structure-specific enhancement factors to enhance different levels of features learned by the ResNet50 network.

[0105] The present invention verifies the enhanced feature learning ability of the structure-enhanced stackable attention module through a large number of comparative experiments and neural network visualization techniques on the pedestrian re-identification general datasets CUHK03L and Market1501, and achieves competitive performance on the CUHK03L and Market1501 datasets. Finally, on the CUHK03L dataset, it reaches an average precision mean of 78.0% and a rank-1 accuracy of 81.3%, and on the Market1501 dataset, it achieves an average precision mean of 88.2% and a rank-1 accuracy of 96.2%.

[0106] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A pedestrian re-identification method based on enhancing target structure relationship, characterized in that, It includes the following steps: Select ResNet50 as the backbone network; Add a structure-enhanced stackable attention module in the residual stack module of ResNet50 to strengthen the features at each level, so as to improve the network's ability to learn distinguishable features through the structure enhancement factor; Use label smoothing cross-entropy loss and triplet loss to train the model; The structure-enhanced stackable attention module includes: A structure-enhanced vector learning module, which is used to learn an embedding vector containing structural attention; A structure-separated convolution module, which is used to learn a specific mapping of the structure attention embedding vectors obtained from different structures; The input of the structure enhancement vector learning module is the feature map of a certain layer of ResNet X is obtained through the element rearrangement operation to get the feature map Xinput; Xinput passes through three 1×1 convolutions with input dimension H×W and output dimensions C1, C1, and C2 respectively, batch normalization, and ReLU activation function, and further obtains a request vector response vector and an embedding vector representing its own information Request vector is calculated as follows: Q = ReLU(BN(W conv X input )); Response vector is calculated as follows: R = ReLU(BN(W conv X input )); Embedding vector is calculated as follows: E = ReLU(BN(W conv X input )); The request vector q representing a certain structural information i (q i ∈Q) is multiplied element-wise with all response vectors r j (r j ∈R), and then 1×1 convolution is performed on the result to obtain the relationship response vector for the i-th channel; Consider the bidirectional relationship between a specific channel representing structural information and other channels, that is, stack the active response vector and the passive response vector of a certain channel as the structural relationship representation vector S of channel i i , and the structural relationship vector S can be obtained by the following formula: where Φ(q i , r j ) = Conv(q i × r j ), q i × r j is the element-wise multiplication of q i , r j . Conv is the convolution calculation with a convolution kernel of 1×1. At this time, S is a tensor of dimension C×(2C)×1. S then passes through a 1×1 convolution with an input channel of 2×C and an output channel of C2, batch normalization, and the ReLU activation function to obtain a structure vector. The structure vector and the embedding vector E are element-wise multiplied to obtain the enhanced embedding vector E′. The calculation is as follows: E′ = (ReLU(BN(W conv S))) × E。 2. The pedestrian re-identification method based on enhanced target structure relationship according to claim 1, characterized in that The structure-separated convolution module performs a separated mapping on the enhanced embedding vector E' representing the specific structural relationship information of each channel; For the feature map in the C dimension, C convolutional kernels W i , i ∈ C are used to separately learn the mapping relationship from the enhanced embedding vector E' of each channel to the attention enhancement factor. This mapping is specific to the enhanced embedding vector of the structural relationship information in a certain channel, rather than a unified measurement of the enhanced embedding vector; The output formula of the structure-separated convolution module is:

3. A pedestrian re-identification method based on enhanced target structure relationship according to claim 1, characterized in that, Where, Add a structure-enhanced stackable attention module in the residual stack module of ResNet50 to strengthen the features at each level, so as to improve the network's ability to learn distinguishable features through the structure enhancement factor.

4. The pedestrian re-identification method based on enhanced target structure relationship according to claim 1, characterized in that Where, Use label smoothing cross-entropy loss and triplet loss to train the model, and the combined loss function of this model is: L loss = L triplet + L LSCE L triplet = max(d(a, p) - d(a, n) + margin, 0) Among them, the function d is a distance metric function, and the Euclidean distance metric is adopted; a, p, and n respectively represent the query, matching, and non-matching image feature vectors; batch is the batch size for one training, class is the number of people in the dataset, and y_pred is the two-dimensional vector predicted by the network; y_pred ij represents the probability that the i-th sample belongs to the j-th class; q j is the smoothing factor, λ is the smoothness, taking 0.1; y_label i is the class to which the i-th sample belongs.

Citation Information

Patent Citations

  • Pedestrian re-identification method based on expectation maximization

    CN111738143A