A Pedestrian Re-identification Method and System Based on Local Inhibitory Self-Attention
By introducing technology that partially suppresses self-attention in the pedestrian recognition model, the problems of low accuracy of pedestrian recognition and high model complexity in the existing technology are solved, and higher recognition accuracy and robustness are achieved, which is suitable for pedestrian recognition in complex scenarios.
Patent Information
- Application Number
- CN202210102559.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-01-27
AI Technical Summary
The existing pedestrian re-identification technology has low accuracy when processing multi-camera data, anti-interference and occlusion, and the model parameters are large and the calculation is heavy, making it difficult to effectively identify in real scenarios.
The pedestrian re-identification method based on local suppression of self-attention is adopted. N self-attention branches are introduced into the convolutional backbone network optimized by the residual network, and the local semantic features of different limbs of pedestrians are extracted, and the key areas of different branches are suppressed by the gradient information and category activation heat map in backpropagation.
The recognition ability and robustness of the pedestrian re-identification model are improved, and pedestrians can be accurately identified under complex backgrounds and occlusions, and the adaptability to perspective differences and background changes is improved without significantly increasing the number of model parameters and calculations.
Smart Images

Figure CN114495170B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and particularly to a pedestrian re-identification method and system based on local suppression self-attention. Background Art
[0002] Pedestrian re-identification, also known as cross-camera tracking technology, aims to identify and retrieve pedestrian identities across different scenes and cameras. With the construction and implementation of smart cities and smart communities, more and more cameras are installed in various corners of communities, shopping malls, and streets, and a large amount of pedestrian video or image data can be obtained every day. However, the utilization of such data is far from sufficient at present. For example, in the intelligent security scenario, currently, the action trajectory of the target person is mainly determined by manually screening and integrating a large number of surveillance videos, which is not only time-consuming and laborious but also prone to misjudgment; in scenarios such as intelligent missing person search, when a child gets separated, it is often only possible to remind and notify by the staff's broadcast. In a noisy environment and due to the immature psychology of children, the effect is very limited.
[0003] Therefore, using a computer to intelligently integrate and analyze pedestrian data from multiple cameras is the preferred method for cross-camera tracking currently and in the future. A typical complete cross-camera tracking system is divided into three stages: pedestrian data collection and upload from multiple cameras at different times, pedestrian position detection based on video frames, and pedestrian identity recognition based on handcrafted features extraction or deep neural networks. Among them, due to the heterogeneity of data often brought by different manufacturers and models of cameras, and the accuracy of pedestrian position detection also depends on the performance of the detection model. In addition, due to factors such as weather changes, lighting conditions, pose differences, obstacle occlusion, and complex and variable backgrounds, it is extremely challenging to accurately retrieve pedestrians with the same identity from a large number of pedestrian identity databases.
[0004] Using a pedestrian re-identification model for identity recognition includes two steps: pedestrian image feature extraction and feature similarity calculation. Among them, the feature similarity calculation part generally calculates the cosine distance or Euclidean distance between the features of the image to be detected and the features of the images in the image library. The smaller the distance, the greater the similarity, and the sorting result of the detection is obtained according to the similarity score. The accuracy of this step depends on the performance of the model in the feature extraction stage. In a large-scale pedestrian image library, there are many different identity pedestrian samples with similar poses, clothing, and viewpoints. The discovery of some local fine-grained features is the key to distinguishing them. Therefore, designing a local feature extraction model that can resist background interference and occlusion is the core to improve the robustness of the pedestrian re-identification model.
[0005] Early work utilized manually designed feature extraction operators to extract local pedestrian features from the original image. For example, Karanam et al. divided the original image into six horizontal parts and separately calculated the grayscale histograms in different color spaces for each part. Matsukawa et al. proposed a two-level Gaussian modeling model. First, the image was divided into multiple local blocks, and Gaussian distributions were established among multiple local blocks at the same horizontal position. Then, a second-level Gaussian distribution was established between different horizontal positions, improving the ability to depict image textures. However, the above methods are greatly affected by texture features, lack the extraction of the overall semantic information of pedestrians, are prone to overfitting the training set, and have limited accuracy.
[0006] With the excellent performance of deep neural networks in the field of large-scale image classification, many researchers have also applied them to the problem of person re-identification. Sun et al. performed horizontal slicing on the feature maps output by the convolutional network, which respectively represented the head, upper body, thighs, calves, etc. Each slice was classified independently, effectively improving the accuracy of person re-identification. Rahul et al. directly performed horizontal slicing on the original image and then sent each slice into a long short-term memory network to obtain the fused features. Zhang et al. added a slice matching algorithm based on the shortest path on the basis of horizontal slicing, alleviating the misalignment problem of the hard slicing method to a certain extent. Such methods have good effects when the pedestrian postures are relatively standard and unified. However, in real scenarios, pedestrian postures vary greatly, and there are non-erect states such as cycling, partial occlusion states such as umbrellas, and occlusion by other pedestrians. Therefore, the hard slicing method will lead to a decrease in retrieval accuracy.
[0007] Zhao et al. used a human pose estimation model to extract multiple key points of the pedestrian skeleton. After obtaining the corresponding pixel regions according to the key points, they were trained together with the original image to achieve regional alignment. Zheng et al. used affine transformation to achieve pixel-level pose alignment based on the skeleton key points and then performed feature extraction. The local alignment effects of these methods rely on additional pose estimation models and inevitably increase the number of model parameters, which is not conducive to engineering deployment. Summary of the Invention
[0008] In order to overcome the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide a person re-identification method and system based on local suppression self-attention.
[0009] To achieve the above object of the present invention, the present invention provides a person re-identification method based on local suppression self-attention, including the following steps:
[0010] Collect original pedestrian picture samples and preprocess the samples;
[0011] Construct a network model, which includes a convolutional backbone network optimized by a residual network. The output of the convolutional backbone network is connected to N self-attention branches for local feature extraction. The outputs of the N local feature extraction branches are residually connected to the feature map output by the convolutional backbone network, where N is a positive integer;
[0012] Use the preprocessed samples to perform backpropagation training on the network model;
[0013] Perform person re-identification on the target image in the trained network model.
[0014] The person re-identification method based on local suppression self-attention
[0015] The person re-identification method based on local suppression self-attention introduces N self-attention branches and a global branch for residual connection on the basis of a convolutional backbone network optimized by a residual network to extract local semantic features of different body parts of pedestrians. Among them, due to the skip connection operation of the residual network, the model can have sufficient depth while avoiding the phenomenon of gradient disappearance or gradient explosion. The introduction of N self-attention branches improves the extraction accuracy of pedestrian local features, making the recognition ability of the person re-identification method based on local suppression self-attention higher.
[0016] The preferred scheme of the person re-identification method based on local suppression self-attention: The method for training in the network model is as follows:
[0017] Send the preprocessed image into the convolutional backbone network for global feature extraction to obtain a multi-channel feature map;
[0018] Send the obtained multi-channel feature map into the N self-attention branches respectively for local feature extraction;
[0019] Perform residual connection between the output of the self-attention branch and the feature map output by the convolutional backbone network, calculate the loss function through the pooling and classifier of the convolutional backbone network, perform backpropagation and update the network parameters, and save the model data until the iteration is completed.
[0020] The preferred scheme of the person re-identification method based on local suppression self-attention: Each self-attention branch includes a layer normalization layer and a self-attention block. Among them, the layer normalization layer normalizes the data of different channels of the same sample, and the self-attention block uses the structure of each head in the multi-head attention in the vision transformer to obtain a global receptive field by establishing the relationship between each feature and all other features.
[0021] The layer normalization layer avoids the distribution drift phenomenon by normalizing the data of different channels of the same sample, and solves the defect of the batch normalization layer being affected by the sample batch size, which can accelerate the convergence of the model. The self-attention block effectively avoids the problem that the local receptive field of the convolutional network is not enough to focus on global information. Since the background accounts for a large proportion of pedestrian images, and the backgrounds captured by different cameras vary greatly, the judgment of the identity of pedestrians is undoubtedly a huge interference. However, establishing a long-range dependency through the global receptive field can effectively shield the differences between different backgrounds, allowing the model to focus on the pedestrian's limb area. In addition, when the sample size is large, there will often be different pedestrians with similar appearances in certain parts, such as wearing the same shoes or the same sunglasses. The global receptive field can alleviate the identity misjudgment caused by such similar appearance to a certain extent.
[0022] The preferred solution of the pedestrian re-identification method based on local suppressed self-attention is as follows: after back propagation, the category activation heat maps corresponding to the output features of N self-attention branches are calculated respectively, and the input of the remaining branches is suppressed according to the heat map of each self-attention branch.
[0023] This can force the network to mine different local semantic features, which is more conducive to extracting multiple sub-salient local features of pedestrians. At the same time, it avoids the redundancy of N self-attention branches and avoids the problem of over-focusing on the most salient area and ignoring other equally important sub-salient areas, resulting in information loss.
[0024] Preferred solution: The method of suppressing the input of the remaining branches according to the heat map of each self-attention branch is: using the category activation heat map of each self-attention branch to obtain the salient area mask, and then superimposing it on the input of the remaining branches. This can achieve significant area suppression, where the category activation heat map can reflect the feature map space area that is strongly related to the identity of the pedestrian. The larger the value in the heat map, the more critical the area is to the judgment of the pedestrian's identity. The use of mask fusion can shield the key area of the current branch from other areas. Although the N branches have the same structure, they have different initialization parameters. Therefore, the parameter values after multiple iterations will also be different, allowing different branches to focus on different limb areas.
[0025] The preferred solution of the pedestrian re-identification method based on local suppressed self-attention is: the convolutional backbone network adds a non-local attention block after the first two residual blocks of the ResNet50 network, the pooling layer is a generalized average pooling layer with parameters, and the loss function is a weighted regularized triple loss function. The convolutional backbone network obtained in this way can improve the accuracy of the pedestrian re-identification network model.
[0026] Preferred solution of the person re-identification method based on local suppression self-attention: The N self-attention branch structures are the same, but the initialization parameters are different. Therefore, the parameter values after multiple iterations will also be different, enabling different self-attention branches to focus on different limb regions.
[0027] Preferred solution of the person re-identification method based on local suppression self-attention: The target image passes through the network model to obtain the global feature and the local features of N self-attention branches. After fusion, pooling is performed to obtain the feature vector. The similarity between this feature vector and the feature vectors of all images in the image library to be queried is calculated to obtain the recognition result.
[0028] This application also proposes a person re-identification system, including a processor and a memory communicatively connected to the processor. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned person re-identification method based on local suppression self-attention. This person re-identification system has the advantages of the above-mentioned person re-identification method based on local suppression self-attention.
[0029] The beneficial effects of the present invention are as follows: The present invention introduces multiple self-attention branches into the convolutional backbone network optimized by the residual network, which can adaptively focus on the local regions of the pedestrian's limbs, extract the local semantic features of different limb parts of the pedestrian, and reduce the interference of the background; using the gradient information in backpropagation and the class activation heatmap to achieve mutual suppression of the key regions of different branches, enabling different branches of the model to focus on different local body parts of the pedestrian, thereby obtaining rich fine-grained detail information. Without significantly increasing the number of parameters and the amount of computation of the network model, it can adaptively mine the body regions of pedestrians in the image, improving the robustness to perspective differences, background changes, etc.
[0030] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0031] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0032] Figure 1 is a schematic diagram of the network model structure;
[0033] Figure 2 is a visualization diagram of the class activation heatmap. Detailed Embodiments
[0034] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0035] In the description of the present invention, unless otherwise specified and defined, it should be noted that the terms "mounted", "connected", and "connected" should be understood in a broad sense. For example, it may be a mechanical connection or an electrical connection, or it may be the communication inside two elements. It may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.
[0036] The present invention provides a person re-identification method based on local suppression self-attention, including the following steps:
[0037] Collect original person image samples and preprocess the samples.
[0038] Set a data sampling strategy before collecting the original person image samples. In this embodiment, the data sampling strategy is to randomly select a fixed number of persons, and the same number of images of each person are combined into a batch. The specific number is determined by the video memory size of the used graphics card, and generally, the larger the better.
[0039] The preprocessing of the original person image samples includes common data augmentation operations such as resizing to a unified size, standardizing the three RGB channels, randomly flipping left and right, pixel value filling, randomly cropping the size, and random erasing. Through preprocessing, the diversity of the samples can be increased to a certain extent, avoiding overfitting of the model to a specific data set and improving the generalization performance.
[0040] Construct a network model. In this embodiment, a network model is established using a mainstream deep learning framework such as Pytorch or TensorFlow. The specific structure of this network model is as Figure 1 shown, including a convolutional backbone network and N local feature extraction branches.
[0041] In this embodiment, the convolutional backbone network partially loads the parameters pre-trained on the ImageNet dataset, and the remaining parameters are initialized by default. The convolutional backbone network uses the AGW network as the benchmark model, which is improved based on the ResNet50 residual network as follows: a non-local attention block is added after the first two residual blocks of the ResNet50 residual network respectively. The non-local attention belongs to a type of spatial attention structure, which is generally used to establish long-range dependencies in the shallow layer of the neural network; the original global average pooling layer is changed to a generalized average pooling with parameters; the original triplet loss function with a hyperparameter threshold is changed to a weighted regularized triplet loss function. The above three optimizations all improve the accuracy of the person re-identification model to varying degrees, and the performance on the current public dataset is better than that of ResNet50.
[0042] In this embodiment, the local feature extraction branch is a self-attention branch. Generally, N is not greater than 3. In this embodiment, considering the detection accuracy and computational complexity, it is preferred that there are three self-attention branches, such as Figure 1 Branch-1, Branch-2, and Branch-3 in. Each self-attention branch includes a layer normalization layer and a self-attention block. The layer normalization layer normalizes the data of different channels of the same sample, and the self-attention block uses the structure of each head of the multi-head attention in the vision transformer to obtain a global receptive field by establishing the relationship between each feature and all other features. The structures of the three self-attention branches are the same, but the initial parameters are different.
[0043] Specifically, the output of the convolutional backbone network is connected to three self-attention branches, and the outputs of the three local feature extraction branches are connected to the feature map output by the convolutional backbone network in a residual connection manner.
[0044] The preprocessed samples are used to train the network model. During the training process, one iteration includes two parts: calculating the gradient value of the parameters using backpropagation and updating the network parameters using the gradient value.
[0045] In this embodiment, it is preferred but not limited to using the Adam optimizer to train for 120 rounds. The initial learning rate is 0.00035. The learning rate doubles every time in the first 10 rounds, and the learning rate is reduced by 10 times at the 40th and 70th rounds. The model parameters are saved after the training ends.
[0046] Specifically, during training, the preprocessed pictures are sent into the convolutional backbone network (in this embodiment, the convolutional backbone network optimized by the residual network as described above) for global feature extraction to obtain a multi-channel feature map. Due to the skip connection operation of the residual network, the model can have sufficient depth while avoiding the phenomenon of gradient disappearance or gradient explosion.
[0047] The obtained multi-channel feature maps are respectively sent to the three self-attention branches for local feature extraction. In this embodiment, the feature maps output by the fourth residual block of the convolutional layer of the ResNet50 residual network are preferably sent to the three self-attention branches respectively.
[0048] The outputs of the three self-attention branches are residually connected to the feature maps output by the convolutional backbone network, and then pooling and loss function calculation are performed in the same way as the backbone network. The loss function includes cross entropy loss and weighted regularized triple loss. Backpropagation and network parameter update are performed until the model data is saved after the iteration is completed.
[0049] During the training process, the category activation heat maps corresponding to the output features of the three self-attention branches are calculated after back propagation, and the input of the remaining two branches is suppressed according to the heat map of each self-attention branch. That is, the category activation heat map of each self-attention branch is used to obtain the salient area mask, which is then superimposed on the input of the remaining branches.
[0050] The following takes branch-1 as an example: Figure 1 The heatmap-1 in the figure is the heatmap corresponding to the first self-attention branch. The category activation heatmap reflects the influence of different areas of the pedestrian image on its identity recognition. The key areas often have larger eigenvalues, which are reflected in the darker parts on the heatmap, such as Figure 1 The heat map of branch-1 indicates that branch-1 ignores the features of the person's shoes and calves. Therefore, this embodiment uses three parameter-unshared self-attention branches to extract multiple sub-significant local features of pedestrians. At the same time, in order to avoid redundancy in the three self-attention branches, the input of the remaining two self-attention branches is suppressed according to the heat map of each self-attention branch, forcing the network to mine different local semantic features. Specifically, taking branch-1 as an example, mask-1 is calculated according to its heat map heatmap-1 according to the following formula:
[0051]
[0052] Where m i,j It represents the value of the i-th row and j-th column in the mask image, h represents the heat map, α represents the suppression coefficient, and the value of the present invention is 0.1, and β represents the suppression factor, and the value of the present invention is 0.75. The mask-1 mask image calculates the Hadamard product with the feature map at the input of branch-2 and branch-3. This step is similar for branch-2 and branch-3, so as to achieve the purpose of different branches focusing on different human body areas.
[0053] The network parameters are updated after each iteration, and the network model data is saved after the iterative training is completed. After the training is over, the heat maps and mask maps are no longer needed. The training images pass through the network model to obtain global features and three branches of local features. After fusion, a feature vector is obtained through a generalized mean pooling layer and participates in the calculation of the similarity between samples. When the similarity reaches the preset requirement, it is considered that the network model obtained by this training meets the requirements, and the training of this network model is terminated.
[0054] When pedestrian re-identification is required for target images, the target images are used for pedestrian re-identification in the trained network model. Specifically, the target images pass through the network model to obtain global features and three self-attention branch local features. After fusion, pooling is performed to obtain a feature vector. The similarity (such as calculating the Euclidean distance) between this feature vector and the feature vectors of all images in the image library to be queried is calculated, and the specific recognition results can be obtained by sorting from small to large.
[0055] The following takes a specific example for a detailed introduction.
[0056] Experiments are carried out on 3 public datasets, including Market-1501, DukeMTMC-reID, and MSMT17. Among them, Market-1501 was made on the campus of Tsinghua University in 2015. The training set and the test set together include 1501 different pedestrians, with about 20 or more pictures per person on average. The test modes include the indoor retrieval mode and the full-scene retrieval mode. In the indoor retrieval mode, only indoor scenes are used for testing, and the background change is relatively small. The DukeMTMC-reID dataset is obtained by manual annotation in the multi-object pedestrian tracking dataset, with more than 30,000 pictures in total. The MSMT17 dataset is obtained by monitoring weather changes and different time periods, and its characteristics are that it includes a large number of pedestrian identities (more than 4000) and pedestrian images (more than 1.2 million).
[0057] The test metrics use two commonly used metrics in pedestrian re-identification problems, namely the cumulative match characteristic (CMC)
[12] and the mean average precision (mAP)
[13] . Among them, the cumulative match characteristic is in the form of rank-k. For example, rank-5 represents the ratio of the correct identity pedestrians included in the first 5 items of the picture sequence sorted by similarity. Generally, rank-1 needs to be concerned. The mean average precision represents the average value of the average precision of all images to be queried. The average precision of a single image is obtained by calculating the mean of the precision of each correct match in the picture sequence.
[0058] Table 1 Experimental results under different datasets
[0059]
[0060] Training was carried out on an Ubuntu system with an NVIDIA - 1080 GPU and 16G of memory. In the pre - processing stage, the image size was scaled down to 256x128, the weight decay coefficient was set to 0.0005, and other settings remained the same as those described in the above - mentioned embodiments. The results on three datasets are shown in Table 1. Among them, "-" indicates that the metric was not listed in the original paper. In addition, datasets not used in the original paper were not listed either. Since the method based on artificial features [1,2] has significantly lower accuracy than the method based on deep learning, it was not listed in the table either.
[0061] As can be seen from Table 1, the network model designed by the present invention has improved accuracy to varying degrees compared with the baseline model on three datasets. Especially for the MSMT17 dataset, the number of pedestrian identities and samples has increased significantly compared with other datasets. The weather involved is complex and diverse, and different lighting conditions and shadow situations at different times have significantly increased the complexity of this dataset. As can be seen from Table 1, the rank - 1 metric of the model of the present invention has increased by 2.7%, and the mAP metric has increased by 4%, both of which are significantly higher than the improvement amplitudes of the metrics on other datasets. This shows that the model of the present invention has good robustness and generalization performance and is superior to the existing methods.
[0062] To further illustrate the performance of the present invention, the class - activation heatmap tool grad - CAM commonly used in computer vision was used to visualize the features extracted by the model of the present invention and the baseline model. The comparison results are as Figure 2 shown.
[0063] From Figure 2 it can be seen that: the baseline model often only focuses on some parts of the pedestrian's body, such as the visualization result of pedestrian b, and is easily affected by the background and occlusion. For example, in pedestrian a, it is greatly interfered by the background, and in pedestrian c, the wheel occlusion leads to misjudgment; while the model of the present invention can better focus on all parts of the pedestrian's body and at the same time shield the interference of the background and occlusion areas.
[0064] To measure the changes in the number of parameters and computational amount brought by the model of the present invention, the open - source PyTorch - OpCounter tool was used for statistics. The results are shown in Table 2. It can be seen that compared with the baseline model, the number of parameters of the model of the present invention has only increased by 0.18%, and the computational amount has only increased by 0.54%, which can be ignored compared with the improvement amplitude of the recognition accuracy.
[0065] Table 2 Comparison of the number of model parameters and computational amount
[0066] Model Parameters FLOPs AGW 23.541M 4.076G ours 23.584M 4.098G
[0067] Thus, it can be seen that the present invention effectively improves the recognition accuracy and model robustness of the pedestrian re - identification model in scenarios with complex backgrounds, occlusions, and variable perspectives.
[0068] The present invention also provides an embodiment of a pedestrian re-identification system, which includes a processor and a memory communicatively connected to the processor. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned pedestrian re-identification method based on local suppression self-attention.
[0069] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0070] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A pedestrian re-identification method based on local suppression self-attention, characterized in that, it includes the following steps: Collect original pedestrian picture samples and preprocess the samples; Construct a network model, which includes a convolutional backbone network optimized by a residual network. A non-local attention block is added after the first two residual blocks of the ResNet50 network respectively. The pooling layer is a generalized mean pooling layer with parameters, and the loss function is a weighted regularized triplet loss function; The output of this convolutional backbone network is connected to N self-attention branches for local feature extraction. The outputs of the N local feature extraction branches are connected to the feature map output by the convolutional backbone network in a residual connection, where N is a positive integer; Use the preprocessed samples to perform backpropagation training on this network model; after backpropagation, calculate the class activation heatmaps corresponding to the output features of the N self-attention branches respectively, and perform input suppression on the remaining branches according to the heatmaps of each self-attention branch. Specifically, the method for performing input suppression on the remaining branches according to the heatmaps of each self-attention branch is: use the class activation heatmap of each self-attention branch to obtain a saliency region mask, and then superimpose it on the input of the remaining branches; Perform pedestrian re-identification on the target picture in the trained network model.
2. The pedestrian re-identification method based on local suppression self-attention according to claim 1, characterized in that, the method for training in the network model is: Send the preprocessed pictures into the convolutional backbone network for global feature extraction to obtain a multi-channel feature map; Send the obtained multi-channel feature maps into the N self-attention branches respectively for local feature extraction; Connect the output of the self-attention branch to the feature map output by the convolutional backbone network in a residual connection, calculate the loss function through the pooling and classifier of the convolutional backbone network, perform backpropagation and update the network parameters until the model data is saved after the iteration is completed.
3. The pedestrian re-identification method based on local suppression self-attention according to claim 1, characterized in that, each self-attention branch includes a layer normalization layer and a self-attention block. The layer normalization layer normalizes the data of different channels of the same sample, and the self-attention block uses the structure of each head in the multi-head attention in the vision transformer to obtain a global receptive field by establishing the relationship between each feature and all other features.
4. The pedestrian re-identification method based on local suppression self-attention according to claim 1, characterized in that, the N self-attention branches have the same structure and different initial parameters.
5. The pedestrian re-identification method based on local suppression self-attention according to claim 1, characterized in that, The target picture passes through the network model to obtain a feature vector after pooling after fusing the global feature and the local features of the N self-attention branches. Calculate the similarity between this feature vector and the feature vectors of all images in the image library to be queried to obtain the recognition result.
6. The pedestrian re-identification method based on local suppression self-attention according to claim 1, characterized in that, the N is not greater than 3.
7. A pedestrian re-identification system, It is characterized in that it includes a processor and a memory communicatively connected to the processor, the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the pedestrian re-identification method based on local suppression self-attention according to any one of claims 1-6.