Pedestrian re-identification method based on feature attention mechanism and GAN
By introducing feature attention mechanism to generate adversarial network (AmGAN) in pedestrian recognition, pedestrian images under multiple attention are generated, which solves the problem of insufficient feature vectors under single attention and improves the performance of pedestrian recognition model.
Patent Information
- Application Number
- CN202410217119.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-05-09
AI Technical Summary
The existing pedestrian recognition method based on deep learning is prone to ignore insignificant details of pedestrians during feature extraction, resulting in limited model performance and insufficient feature vectors under single attention.
A pedestrian recognition method based on feature attention mechanism generation adversarial network (AmGAN) is proposed. Pedestrian images under different attention are generated through AmGAN, the training set is expanded and the feature vectors of the query image are improved.
By generating pedestrian images with multiple attention, the training set is expanded, and the performance of pedestrian re-identification model is improved, especially in dealing with image aligning and feature detail extraction.
Smart Images

Figure CN119964197A_ABST
Abstract
Description
Technical field:
[0001] The present invention relates to the field of image generation, and in particular to a pedestrian re-identification method based on a feature attention mechanism and a generative adversarial network. Background technology:
[0002] In the early days, the research idea of person re-identification was usually to first extract manual features from pedestrian images, such as color histogram, HOG, etc., and then use similarity measurement methods to learn the metric matrix, such as LMNN, XQDA, etc. With the rise of deep learning, deep learning technology has been widely used to complete pedestrian re-identification tasks, and its performance far exceeds that of traditional methods.
[0003] At present, the pedestrian re-identification methods based on deep learning can be roughly divided into three categories. The first category is the pedestrian re-identification method based on global features. The main idea of this method is to treat pedestrian re-identification as an identity classification task to learn pedestrian features when training the network, that is, first extract the pedestrian features in the image through a convolutional neural network, and then judge whether it belongs to the same pedestrian based on the obtained features. The above are all based on global feature extraction, that is, a feature vector is obtained from the entire image. Later, researchers found that this type of global feature often ignores some insignificant details of pedestrians, causing bottlenecks in the performance of the model. Therefore, some researchers proposed the second type of method, which completes the pedestrian re-identification task based on the local features of pedestrians. In the early stage of research, local features were often extracted by cutting pictures. This method has high requirements for the degree of image alignment. If the two images are not aligned, there may be a phenomenon of comparison of different parts, which affects the performance of the model. In order to solve the problem of image misalignment, some researchers use prior knowledge to pre-align pedestrians, such as using human posture estimation and human skeleton key point extraction and MGN methods. Experiments have shown that by introducing an additional alignment model, although the system overhead is increased, more detailed information can be extracted, thereby improving the model performance. The third category is the pedestrian re-identification method based on metric learning. The main idea of this method is to make the distance between pedestrian images with the same ID small, while the distance between pedestrian images with different IDs large. Similar methods include Triplet loss, Quadruplet loss, and Group similarity learning. A visual analysis of the images identified by the existing pedestrian re-identification technology based on feature distance shows that images that belong to the same or similar attention as the query image often have a small feature distance with the query image, that is, they are easy to be identified; while images with a large attention deviation from the query image have a small feature distance with the query image and are difficult to be identified, thereby limiting the performance of the model.
[0004] The AmGAN-based person re-identification method proposed in the present invention can generate images of the pedestrian under three fixed attentions based on a given pedestrian image, which can not only expand the existing data set and improve network performance in the training stage; but also improve the semantic features of the query image under different attentions in the test stage, thereby further improving network performance. The AmGAN proposed in the present invention is highly flexible and can be directly used in combination with existing pedestrian re-identification methods, thereby making full use of the performance foundation of existing methods. Summary of the invention:
[0005] The purpose of the present invention is to overcome the shortcomings of the existing methods and propose a pedestrian re-identification method based on a feature attention mechanism to generate an adversarial network, especially to generate a pedestrian image feature attention mechanism based on AmGAN to solve the problem of insufficient feature vectors in pedestrian images under single attention.
[0006] A pedestrian re-identification method based on a feature attention mechanism to generate an adversarial network, characterized by comprising the following steps:
[0007] Step 1: Select images from the original person re-ID training set as the training set of the proposed AmGAN (Attention Mechanism Generation Adversarial Networks), and train the proposed AmGAN based on the selected training set;
[0008] Step 2: Use the trained AmGAN to generate images under other attention for the image under a given attention, and give the generated image the same ID label as the original image. Finally, add the generated image with the label to the original training set to obtain the expanded training set, and train the pedestrian re-identification network based on the expanded training set;
[0009] Step 3: Use AmGAN to generate a feature attention mechanism for the query image and improve the feature vector of the query image;
[0010] Step 4: Use the refined feature vector as the feature of the query image for the person re-identification task, and finally arrange the images by similarity to complete the person re-identification task.
[0011] The implementation of step 1 includes:
[0012] Step 1.1: Select three attention images of the front, side and back and 1 to 5 other attention images from the original pedestrian re-identification training set as the training set of the proposed AmGAN;
[0013] Step 1.2: Group the images according to the original pedestrian ID so that each group of images contains the front, side, back and several other attention images;
[0014] Step 1.3: The proposed AmGAN includes three generators for generating front, side and back images respectively and a multi-category discriminator. The feature attention mechanism generator takes a given image as input and outputs a pedestrian image with a certain attention. The feature attention mechanism generators G1, G2 and G3 have the same structure but do not share parameters. Each generator includes two small generators and the Monte Carlo search and attention mechanism therein. The discriminator takes the generated image or the real image and the corresponding attention as input and outputs the probability of the real image. The generator and the discriminator are trained alternately until the Nash equilibrium is reached.
[0015] The implementation of step 2 includes:
[0016] Step 2.1: Use the trained AmGAN generator to generate other images under a given attention for the image under a given attention, and assign the generated image the same ID label as the original image. Finally, add the generated image with the label to the original training set to obtain the expanded training set.
[0017] Step 2.2: Train a person re-identification network based on the expanded training set, where the person re-identification network can be an existing method or the proposed feature vector-based method. The finally trained person re-identification network has the ability to extract image features.
[0018] The implementation of step 3 includes:
[0019] Step 3.1: For a given query image, use AmGAN to generate a feature attention mechanism to obtain pedestrian images under three attentions: front, side, and back.
[0020] Step 3.2: Input the three generated images and the original given image into the pedestrian re-identification network for feature extraction, and fuse the four extracted feature vectors according to the maximum principle to obtain a complete feature vector.
[0021] The implementation of step 4 includes:
[0022] Step 4.1: Use the person re-identification network to extract features from all images in the test set and obtain feature vectors of all images in the test set;
[0023] Step 4.2: Take the perfected feature vector as the feature of the query image and measure the similarity with all the feature vectors in the test set. The similarity measurement can use Euclidean distance, but is not limited to this. Finally, arrange the images in the test set from large to small similarity, that is, from small to large Euclidean distance, to complete the pedestrian re-identification task. Description of the drawings:
[0024] Figure 1 The figure is a flowchart of a pedestrian re-identification method based on a feature attention mechanism to generate an adversarial network.
[0025] Figure 2 This is a model framework diagram of a pedestrian re-identification method based on a feature attention mechanism to generate an adversarial network.
[0026] Figure 3 This is the structure diagram of generator G1.
[0027] Figure 4 It is a multi-category discriminator structure diagram.
[0028] Figure 5 is the generated pedestrian image. Specific implementation method:
[0029] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0030] Figure 1 A specific schematic diagram of the process for implementing the present invention is shown. Figure 2 The overall framework diagram of the present invention is shown. The pedestrian re-identification method based on the feature attention mechanism to generate an adversarial network includes the following steps:
[0031] Step 1: First, three attention images of the front, side and back are selected from the original pedestrian re-identification training set as the training set of AmGAN, and grouped according to the original pedestrian ID, so that each group of images contains the front, side, back and several other attention images. The AmGAN proposed in this invention is trained based on the assembled image groups.
[0032] AmGAN consists of three generators and a multi-category discriminator. The three generators G1, G2 and G3 are used to generate pedestrian images under three different attentions, with the same structure but no shared parameters.
[0033] Figure 3The structure diagram of generator G1 is shown. The generation process includes three stages. In the first stage, the original image under a given attention is input into generator E1, and a coarse-grained image is generated using generator E1. In the second stage, the coarse-grained image is sampled six times using Monte Carlo search to obtain a larger semantic generation space J1~J6, and then the attention mechanism is used to extract features from the intermediate results of the six samples (where the attention mechanism is implemented using a convolutional network). In the third stage, the result extracted by the attention mechanism is input into generator F1 together with the original image to generate a fine-grained image.
[0034] Figure 4 The structure diagram of the multi-class discriminator is shown. The function of the multi-class discriminator is to distinguish between real images and generated images under different attention. The multi-class discriminator designed in the present invention is based on the discriminator of CGAN. The input includes the real image or the generated image and the corresponding attention label I i (I i The front, side and back images can be taken), and finally the probability of the real image is output.
[0035] The adversarial loss of AmGAN uses the objective function of WGAN, and based on the physical meaning of the matrix spectral norm, the multi-class discriminator satisfies the Lipschitz constraint globally. The physical meaning of the matrix spectral norm is that the length of any vector after matrix transformation is less than or equal to the length of the product of this vector and the matrix spectral norm, that is:
[0036]
[0037] Where σ(W) represents the spectral norm of the weight matrix, x represents the input vector of the layer, and δ represents the change in x.
[0038] In order to ensure the quality of generation, retain the original image features and improve visual satisfaction, pixel-wise mean squared error (pMSE) and perception loss are introduced based on the original objective function. pMSE is defined as:
[0039]
[0040] Where: I x,y and I′ x,y They represent the pixel values of the (x, y) pixel in the target attention image and the given attention image respectively, W and H represent the height and width of the image respectively, and θ is the generator parameter.
[0041] Since pMSE is a pixel-by-pixel loss calculation, it will inevitably lead to an overly smooth texture structure and poor visual perception. Therefore, the present invention introduces perception loss to improve visual satisfaction. The visual perception loss is defined as follows:
[0042]
[0043] Where: φ i,j W represents the feature map before the i-th maximum pooling layer and after the j-th convolutional layer in the pre-trained VGG19 network. I and I′ represent the target attention image and the given attention image, respectively. i,j and H i,j Represents the dimension of each feature map in the VGG network.
[0044] The overall objective function is:
[0045] L total =L WGAN +αL pMSE +βL pl
[0046] Where: L WGAN is the adversarial loss function of WGAN, L pMSE is the pixel-wise mean squre error, L pl is the perception loss, α and β are the hyperparameters that control the ratio.
[0047] The generator and the discriminator are trained alternately until the Nash equilibrium is reached.
[0048] Step 2: Use the trained AmGAN generator to generate images under other attention for the image under given attention, and give the generated image the same ID label as the original image. Finally, add the generated image with the label to the original training set to obtain the expanded training set.
[0049] A person re-identification network is trained based on the expanded training set, wherein the person re-identification network can be an existing method or a proposed feature vector-based method, and the finally trained person re-identification network has the ability to extract image features.
[0050] Step 3: For the given query image, use AmGAN to generate the feature attention mechanism to obtain the pedestrian image under the three attentions of the front, side and back of the given pedestrian image. Figure 5 The generated pedestrian image is shown.
[0051] The three generated images and the original given image are respectively input into the pedestrian re-identification network for feature extraction, and the four extracted feature vectors are fused according to the maximum principle to obtain a perfect feature vector.
[0052] Step 4: Use the person re-identification network to extract features from all images in the test set and obtain feature vectors of all images in the test set;
[0053] The perfected feature vector is used as the feature of the query image, and the similarity is measured with all the feature vectors in the test set. The similarity measurement can use Euclidean distance, but is not limited to this. Finally, the images in the test set are arranged from large to small in similarity, that is, from small to large in Euclidean distance, to complete the pedestrian re-identification task.
[0054] It should be understood that parts not elaborated in detail in this specification belong to the prior art.
[0055] The above description in combination with the accompanying drawings is only a specific implementation method and process of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art should understand that this is only an example, and various changes and substitutions can be made to this implementation method without departing from the essence of the present invention. The scope of the present invention is limited only by the attached claims.
[0056] The embodiments of the present invention described with reference to the accompanying drawings are exemplary and are only used to explain the present invention. They should not be understood as limiting the present invention. The specific scope of the embodiments of the present invention is not limited thereto. On the contrary, all embodiments of the present invention include all changes and modifications that fall within the spirit and connotation of the appended claims.
Claims
1. A pedestrian re-identification method based on feature attention mechanism and GAN, characterized in that: The steps include: Step 1: Select images from the original person re-ID training set as the training set for the proposed AmGAN (Attention Mechanism Generation Adversarial Networks); train the proposed AmGAN based on the selected training set; Step 2: Use the trained AmGAN to generate images under other attention for the image under a given attention, and give the generated image the same ID label as the original image. Finally, add the generated image with the label to the original training set to obtain the expanded training set, and train the pedestrian re-identification network based on the expanded training set; Step 3: Use AmGAN to perform multi-attention generation on the query image and improve the feature vector of the query image; Step 4: Use the improved feature vector as the feature of the query image for the person re-identification task, and finally arrange the images according to the similarity to complete the person re-identification task; The step 1 comprises the following steps: Step 1.1: Select three attention images of the front, side and back and 1 to 5 other attention images from the original pedestrian re-identification training set as the training set of the proposed AmGAN; Step 1.2: Group the images according to the original pedestrian ID so that each group of images contains the front, side, back and several other attention images; Step 1.3: The proposed AmGAN includes three generators for generating front, side and back images respectively and a multi-category discriminator; the generator takes a given image as input and outputs a pedestrian image under a certain attention; the discriminator takes the generated image or the real image and the corresponding attention as input and outputs the probability of the real image; the generator and the discriminator are trained alternately until the Nash equilibrium is reached; The step 2 comprises the following steps: Step 2.1: Use the trained AmGAN generator to generate other images under a given attention for the image under a given attention, and assign the generated image the same ID label as the original image. Finally, add the generated image with the label to the original training set to obtain the expanded training set. Step 2.2: Train a person re-identification network based on the expanded training set, where the person re-identification network can be an existing method or the proposed feature vector-based method, and the finally trained person re-identification network has the ability to extract image features; The step 3 comprises the following steps: Step 3.1: For a given query image, use AmGAN to perform multi-attention generation to obtain pedestrian images under three attentions: front, side, and back of the given pedestrian image; Step 3.2: Input the three generated images and the original given image into the pedestrian re-identification network for feature extraction, and fuse the four extracted feature vectors according to the maximum principle to obtain a complete feature vector.
2. The pedestrian re-identification method based on multi-attention generative adversarial network according to claim 1 is characterized in that: The step 4 comprises the following steps: Step 4.1: Use the person re-identification network to extract features from all images in the test set and obtain feature vectors of all images in the test set; Step 4.2: Use the perfected feature vector as the feature of the query image, and perform similarity measurement with all feature vectors in the obtained test set, where the similarity measurement can use Euclidean distance, but is not limited to this; finally, arrange the images in the test set from large to small similarity, that is, from small to large Euclidean distance, to complete the pedestrian re-identification task.