A Cross-modal Person Re-identification Method Based on Generative Adversarial Network

By generating adversarial networks and Resnet-50 feature extraction, combined with attention mechanisms and modal mitigation modules, the feature difference problem of mid-infrared and visible light images across modal pedestrian re-identification is solved, achieving higher recognition accuracy and robustness.

CN114743162BActive Publication Date: 2025-07-29ZHEJIANG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210364290.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-07-29
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

The existing cross-modal pedestrian re-identification technology has a large difference in characteristics under cross-channel conditions, especially between infrared and visible images, resulting in low recognition accuracy, and is difficult to match effectively in scenarios with insufficient lighting.

Method used

Generative adversarial networks are used for pixel alignment, natural light images are generated to generate cross-modal infrared images, and feature extraction is performed through Resnet-50, attention mechanism and modal mitigation module are added, and network model is optimized using joint loss function.

Benefits of technology

It improves the pedestrian re-recognition accuracy under different modes and poses, can effectively match cross-modal images, and improves the accuracy and robustness of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743162B_ABST
    Figure CN114743162B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-modal pedestrian re-identification method based on a generative adversarial network. Cross-modal images are generated through the generative adversarial network for pixel alignment, and then the real images and the generated cross-modal images under the same ID are input into the backbone network Resnet-50 for feature extraction and feature alignment. A joint loss function is created to screen out the features with identity distinctiveness in the modality-common features, so as to optimize the network model. The present invention utilizes the generative adversarial network and modifies the traditional Resnet-50, achieving good results in the cross-modal pedestrian re-identification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of computer vision, and particularly relates to a cross-modal pedestrian re-identification method based on a generative adversarial network. Background Art

[0002] ReID is a basic problem in image retrieval, which aims to match the target image in the query set query to the image in the database set gallery captured by different cameras. This is a challenge due to varying shooting perspectives, target morphologies, illuminations, and backgrounds. Most existing methods currently focus on the target ReID problem captured by visible light cameras, that is, the single-modal ReID problem. However, in some scenes with insufficient illumination (such as at night, in a dimly lit indoor environment), it is necessary to capture pedestrian images with an infrared camera. Therefore, under such cross-channel conditions, the ReID problem becomes extremely challenging, which is essentially a cross-channel retrieval problem.

[0003] For cross-modal pedestrian re-identification, the mainstream technical solutions include feature learning methods that bridge the gap between RGB and IR images through feature alignment and methods that eliminate modal differences or feature disentanglement through generative adversarial networks. The mainstream algorithms for feature learning, such as the Two-stream series, directly learn features by adding some operations to the two-stream network. The algorithms have high accuracy and speed, but when the appearance of pedestrians changes greatly, the ability to capture details is not strong. The method of generative adversarial networks aims to use the network to learn to generate images of another modality or disentangle modality-independent features. However, due to the existence of a large number of modality-related features, the quality of image generation is not ideal. Summary of the Invention

[0004] The purpose of this application is to provide a cross-modal pedestrian re-identification method based on a generative adversarial network. A generative adversarial network is introduced in the existing technical solution for pixel alignment, generating cross-modal infrared images from natural images, using Resnet-50 for feature extraction and adding an attention mechanism and a modality mitigation module, overcoming the cross-modal retrieval problem of images in different modalities and different postures.

[0005] To achieve the above purpose, the technical solution of this application is as follows:

[0006] A cross-modal pedestrian re-identification method based on a generative adversarial network, comprising:

[0007] Obtain a training data set, where each training sample in the training data set is a first image and a second image with identity annotations, and the first image and the second image are respectively one of a natural light image and an infrared image. Input the training samples into the generative adversarial network to train the generator;

[0008] The first image in the training samples is input into the generator to generate a pseudo-second image. The generated pseudo-second image and the real second image in the training samples are input into the constructed feature alignment network to extract the pseudo-second image features and the real second image features;

[0009] The pseudo-second image and the pseudo-second image features are combined into a pseudo-image feature pair, and the real second image and the real second image features in the training samples are combined into a real image feature pair, which are then input into the joint discriminator for discrimination;

[0010] Calculate the joint loss of the generative adversarial network, the feature alignment network, and the joint discriminator to complete the network training;

[0011] The images in the database are input into the generator in the trained generative adversarial network. The pseudo-images output by the generator and the images to be recognized are input into the feature alignment network to extract the corresponding image features respectively. Through the comparison of the image features, the recognition of the images to be recognized is completed.

[0012] Furthermore, the backbone network of the feature alignment network uses Resnet-50, including a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer. An NAM attention mechanism module is set after each convolutional layer, and an MAM modality mitigation module is also set after the NAM attention mechanism modules of the third convolutional layer and the fourth convolutional layer.

[0013] Furthermore, the pooling layers are removed from the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer.

[0014] Furthermore, the NAM attention mechanism module is located after the batch normalization layer of each convolutional layer.

[0015] Furthermore, the joint loss is expressed as follows:

[0016]

[0017]

[0018]

[0019] Among them, L pix represents the generative adversarial network loss, L feat represents the feature alignment network loss, L D represents the joint discriminator loss, represents the adversarial loss of the generative adversarial network, represents the adversarial loss of the feature alignment network, L cyc represents the cycle consistency loss of the generative adversarial network, λ cyc 、 represent the weights of the corresponding loss functions, Denotes the classification loss of the generated images in the generative adversarial network, Denotes the triplet loss calculated by the generative adversarial network for the generated images, Denotes the classification loss calculated for the features in the feature alignment stage, Denotes the triplet loss calculated for the features in the feature alignment stage, Denotes the weights of the generative adversarial network, Denotes the loss when the joint discriminator discriminates that the image-feature pair is true, Denotes the loss when the joint discriminator discriminates that the image-feature pair is false;

[0020]

[0021]

[0022] Among them, (x, m) represents the image-feature pair input to the joint discriminator, X′ ir Denotes the generated pseudo-second image, X ir Denotes the real second image, M ir Denotes the feature map extracted from the real second image through the feature alignment network, M′ ir Denotes the feature map extracted from the pseudo-second image through the feature alignment network, D j (x, m) represents the output of the joint discriminator;

[0023] Among them, the calculation formula for the joint discriminator loss is as follows:

[0024]

[0025]

[0026] Among them, Denotes that the joint discriminator discriminates that the image-feature pair is true, Denotes that the joint discriminator discriminates that the image-feature pair is false, D j (x, m) is the output of the joint discriminator, which outputs 1 when the discrimination is true and 0 when the discrimination is false. E represents the mathematical expectation. (X ir , M ir ) represents the real image-feature pair under the same identity, is the pseudo-image feature pair with the same identity as (X ir , M ir ), is the real image-feature pair with a different identity from (X ir , M ir );

[0027]

[0028]

[0029] Among them, represents calculating the classification loss for the features X ir and X' ir extracted from the feature alignment network, and p() represents the predicted probability of correctly classifying the input image to its true identity. represents calculating the triplet loss for the generated image;

[0030] L cyc = ||G p′ (G p (X rgb )) - X rgb ||1 + ||G p (G p′ (X ir )) - X ir ||1;

[0031]

[0032]

[0033] Among them, G p represents the generator that generates a pseudo-second image from the first image, and G p′ is also the generator that generates the first image back from the pseudo-second image. represents calculating the classification loss for the generated image X'. ir represents calculating the triplet loss for the generated image X' ir and the real infrared image X ir . L cyc represents the cycle loss function, and L tri represents the triplet loss function.

[0034] A cross-modal person re-identification method based on a generative adversarial network proposed in this application introduces a generative adversarial network for pixel alignment, generates a cross-modal infrared image from a natural image, uses Resnet-50 for feature extraction and adds an attention mechanism and a modality mitigation module, so as to achieve the purpose of pixel alignment and feature alignment, and overcome the cross-modal retrieval problem of images in different modalities and different poses. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flowchart of the cross-modal person re-identification method based on a generative adversarial network in this application;

[0036] Figure 2 is a schematic diagram of the network in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0038] In one embodiment, as Figure 1 shown, a cross-modal pedestrian re-identification method based on a generative adversarial network is proposed, including:

[0039] Step S1, obtain a training data set, where each training sample in the training data set is a first image and a second image with identity annotations, and the first image and the second image are respectively one of a natural light image and an infrared image. Input the training samples into the generative adversarial network to train the generator.

[0040] This application uses the SYSU-MM01 data set as the training data set. The training data set is a data set of infrared images and natural light images with identity annotations. The infrared image and the natural light image with the same identity ID are used as a training sample.

[0041] The training samples are sent into the generative adversarial network for pixel alignment. In a specific embodiment, as Figure 2 shown, G p represents the generator. The goal of the generator is to generate a cross-modal pseudo-infrared image from the natural light image. In the generative adversarial network, the generator learns a mapping from the natural light image to the infrared image. Let the input natural light image be X rgb , and the input infrared image be X ir , X rgb is generated into a pseudo-infrared image X' p by G ir . In addition, the generative adversarial network also includes a generative adversarial network discriminator D p (the generative adversarial network is a relatively mature technology, Figure 2 the complete generative adversarial network is not shown in p , only the generator G p′ is shown, not including another generator G p and the generative adversarial network discriminator D ir ), and its input is X' ir and X

[0042] to distinguish whether the generated image is consistent with the real infrared image. The generator and the discriminator are trained adversarially to reach a balance, so as to achieve the generation of cross-modal images. p′(Generative adversarial networks are relatively mature technologies. Figure 2 (not shown in Figure 2 ), the pseudo-IR image is regenerated back into an RGB image, and the L1 loss is calculated with the real RGB image to train the generator. It should be noted that the same operation is also performed on the IR image. In this embodiment, the first image and the second image are respectively one of a natural light image and an infrared image. When the first image is a natural light image, the second image is an infrared image; when the first image is an infrared image, the second image is a natural light image.

[0043] Step S2: The first image in the training samples passes through the generator to generate a pseudo-second image. The generated pseudo-second image and the real second image in the training samples are input into the constructed feature alignment network to extract the pseudo-second image features and the real second image features.

[0044] As Figure 2 shown in the embodiment, the RGB image passes through the generator to generate a pseudo-IR image, also called a cross-modal image, and then is input into the feature alignment network together with the real IR image in the training samples to extract image features.

[0045] In a specific embodiment, the backbone network of the feature alignment network uses Resnet-50, which includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer. A NAM attention mechanism module is set after each convolutional layer, and a MAM modality mitigation module is also set after the NAM attention mechanism modules of the third convolutional layer and the fourth convolutional layer.

[0046] In this embodiment, the generated cross-modal image X′ ir and the real infrared image X ir are linearly interpolated into a size of 384*192 and input into the backbone network Resnet-50. Resnet-50 includes a first convolutional layer Conv layer1, a second convolutional layer Convlayer2, a third convolutional layer Conv layer3, and a fourth convolutional layer Conv layer4. Although Resnet-50 can reduce the inter-modality differences, there are still large intra-modality differences here, which are mainly caused by factors such as pose, perspective, and illumination.

[0047] To solve this problem, in the feature alignment network constructed in this embodiment, the pooling layers are removed from layer1, layer2, layer3, and layer4 in the Resnet-50 network. The pooling layer reduces information and has a negative impact. In this embodiment, the pooling layer is removed to further retain the feature maps.

[0048] In this embodiment, an attention mechanism is added to the backbone network ResNet-50, and the network is weighted to focus on more discriminative features. Specifically, the NAM attention mechanism module is added after each batch normalization layer in layer 1, layer 2, layer 3, and layer 4.

[0049] In addition, in order to alleviate the feature differences in different modalities, in this embodiment, a MAM modality alleviation module is added after layer 3 and layer 4, enabling the network to learn the common feature representation of images.

[0050] Step S3: Combine the pseudo-second image and the pseudo-second image features to form a pseudo-image feature pair, and combine the real second image and the real second image features in the training samples to form a real image feature pair, and send them to the joint discriminator for discrimination.

[0051] As Figure 2 shown, in this embodiment, the pseudo-second image and the pseudo-second image features are also combined to form a pseudo-image feature pair, and the real second image and the real second image features in the training samples are combined to form a real image feature pair, and then sent to the joint discriminator for discrimination.

[0052] To better maintain identity consistency, this embodiment proposes a joint discrimination module to learn the joint data distribution of image feature pairs. Specifically, its input is an image-feature pair, and only real images and features from the same identity ID will be discriminated as true, otherwise as false.

[0053] Step S4: Calculate the joint losses of the generative adversarial network, the feature alignment network, and the joint discriminator to complete network training.

[0054] In this step, the joint losses of the generative adversarial network, the feature alignment network, and the joint discriminator are calculated, and the joint loss is expressed as follows:

[0055]

[0056]

[0057]

[0058] Among them, L pix represents the generative adversarial network loss, L feat represents the feature alignment network loss, L D represents the joint discriminator loss, represents the adversarial loss of the generative adversarial network, represents the adversarial loss of the feature alignment network, L cyc represents the cycle consistency loss of the generative adversarial network, λ cyc 、 Represents the weight of the corresponding loss function, Represents the classification loss of the generated images in the generative adversarial network, Represents the triplet loss calculated by the generative adversarial network for the generated images, Represents the classification loss calculated for the features in the feature alignment stage, Represents the triplet loss calculated for the features in the feature alignment stage, Represents the weight of the generative adversarial network, Represents the loss when the joint discriminator discriminates that the image feature pair is true, Represents the loss when the joint discriminator discriminates that the image feature pair is false.

[0059]

[0060]

[0061] In the above formulas: Represents the adversarial loss of the generative adversarial network, (x, m) represents the image feature pair input to the joint discriminator, X′ ir Represents the generated pseudo-second image, X ir Represents the real second image, M ir Represents the feature map extracted from the real second image by the feature alignment network, M′ ir Represents the feature map extracted from the pseudo-second image by the feature alignment network, D j (x, m) represents the output of the joint discriminator. Represents the generative adversarial loss of the feature alignment network, and the meanings of other letters are the same as above.

[0062] Among them, the calculation formula for the joint discriminator loss is as follows:

[0063]

[0064]

[0065] Among them, Represents the loss function when the joint discriminator discriminates that the image feature pair is true, Represents the loss function when the joint discriminator discriminates that the image feature pair is false, D j (x, m) is the output of the joint discriminator, which outputs 1 when the discrimination is true and 0 when the discrimination is false. E represents the mathematical expectation, (X ir , M ir ) represents the real image feature pair under the same identity, is the pseudo-image feature pair with the same identity as (X ir , M ir ), is the pair of real image features under different identities of (X ir , M ir ).

[0066]

[0067]

[0068] Among them represents calculating the classification loss (cross-entropy loss) for the features of X ir and X' ir extracted from the feature alignment network, and p() is the predicted probability of correctly classifying the input image to its true identity, represents calculating the triplet loss for the generated image.

[0069] The loss function of the generative adversarial network includes a cycle-consistency loss and an ID loss (classification loss + triplet loss). Among them, the cycle-consistency loss enables the generated picture to maintain the original structure and content (such as pose, angle, etc.), and the ID loss enables the synthesized picture to maintain the same identity information as the original picture as much as possible. These loss functions are as follows respectively:

[0070] L cyc = ||G p′ (G p (X rgb )) - X rgb ||1 + ||G p (G p′ (X ir )) - X ir ||1;

[0071]

[0072]

[0073] Among them, G p represents the generator that generates a pseudo-IR image from RGB, and G p′ is also the generator that generates the RGB image back from the pseudo-IR image, represents calculating the classification loss for the generated image X' ir , represents calculating the triplet loss for the generated image X' ir and the real infrared image X ir , L cyc represents the cycle loss function, and L tri represents the triplet loss function.

[0074] Step S5: Input the images in the database into the generator of the trained generative adversarial network. The generator outputs a pseudo-image and the image to be recognized are input into the feature alignment network to extract corresponding image features respectively. Through the comparison of the image features, the recognition of the image to be recognized is completed.

[0075] The specific implementation method is as follows: Input the images in the database (that is, the known images with pedestrian identities marked and saved in the database, usually a dataset) into the generator of the trained generative adversarial network, and the generator outputs a pseudo-image. Input the pseudo-image and the image to be recognized into the feature alignment network to extract corresponding image features respectively, save the extracted features, and calculate the cosine similarity between the features saved for the pseudo-image and the image to be recognized for matching. According to the cosine similarity, sort from largest to smallest to obtain the re-identification result.

[0076] The cosine similarity calculation formula is as follows:

[0077]

[0078] Where A and B are the real IR image features and pseudo-image features respectively, represented as n-dimensional vectors, · represents the vector inner product, and |||| represents taking the modulus of the vector. The cosine similarity measures the similarity between two vectors. The larger the cosine similarity, the more matching the features are.

[0079] It should be noted that during training, the generative adversarial network is a complete network, and a joint discriminator is added after the feature alignment network to train the generator and the feature alignment network well. When performing pedestrian re-identification after training, only the generator and the feature alignment network are needed. When performing pedestrian re-identification, input the RGB images in the database into the generator to generate pseudo-IR images, and then extract the pseudo-IR image features through the feature alignment network. Input the real IR image to be recognized into the feature alignment network to obtain the image features of the IR image to be recognized. Then make a comparison to find the RGB images of the same identity, so as to achieve the result of pedestrian re-identification.

[0080] This application will map the real infrared image and the generated infrared image to the same feature space, and use identity label-based classification and triplet loss to supervise the features. After the network extracts the features, calculate the loss with the real natural image to optimize the network parameters. When the pedestrian posture changes, the network can still extract similar features well.

[0081] In this application, the generated image and the real image are fed into the discriminator of the generative adversarial network. The parameters of the generative adversarial network are updated using the cycle consistency loss. The generated image and the real image are input into Resnet-50 for feature extraction. In order to make the network pay more attention to discriminative features, an attention mechanism is added to each layer. At the same time, a modality mitigation module is added to layers 3 and 4 in the deep layer. The combination of ID Loss and TripletLoss is used to train the global features, and cycle-Loss is used to train the generator and the discriminator. The parameters of the generative adversarial network and the backbone network Resnet-50 are optimized through the backpropagation of the loss, so as to achieve the purpose of pixel alignment and feature alignment. Jointly inputting the image and the features into the joint discriminator can improve the discrimination ability of the discriminator and the quality of image generation.

[0082] In a specific embodiment, the NAM attention mechanism module is expressed by the following formula:

[0083]

[0084] M c = sigmoid(W r (BN(F1)))

[0085] M s = sigmoid(W λ (BN s (F2)))

[0086]

[0087] The NAM attention mechanism is a mature attention mechanism improved on the basis of the CBAM mechanism. It includes two modules: channel attention and spatial attention, which can make the network pay more attention to the discriminative features of the image and has very few parameters, making it easy for network training.

[0088] Where μ β and σ β are the mean and standard deviation of the mini-batch B, γ and β are trainable affine transformation parameters, where Mc represents the output feature. γ is the example factor for each channel, and the weight is obtained by W γ = γ i / ∑ j=0 γ j x represents the input, y is the output, W represents the network weight, L() is the loss function, g() is the L1 loss function, and p is the threshold for balancing g(γ) and g(λ).

[0089] The goal of the NAM attention mechanism is to design a mechanism that reduces information and amplifies the interaction features of the global dimension. Adopting the sequence of the CBAM attention mechanism, the channel and spatial attention mechanisms, and redesigning the sub-module. Given the input feature map, The intermediate state F2 and the output F3 are defined as:

[0090]

[0091]

[0092] where Mc and Ms are the channel and spatial attention maps, denotes element-wise multiplication. The channel attention sub-module uses 3D permutations to preserve information across three dimensions. Then it uses two layers of MLP to amplify the cross-dimensional channel-spatial dependencies. In the spatial attention sub-module, to focus on spatial information, two convolutional layers are used for spatial information fusion.

[0093] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it cannot be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A cross-modal person re-identification method based on generative adversarial networks, characterized by: The cross-modal person re-identification method based on a generative adversarial network includes: Obtain a training data set. Each training sample in the training data set is a first image and a second image with identity annotations. The first image and the second image are respectively one of a natural light image and an infrared image. Input the training samples into the generator of the generative adversarial network for training; The first image in the training sample passes through the generator to generate a pseudo-second image. Input the generated pseudo-second image and the real second image in the training sample into the constructed feature alignment network to extract the pseudo-second image feature and the real second image feature; Form a pseudo-image feature pair with the pseudo-second image and the pseudo-second image feature, and form a real image feature pair with the real second image and the real second image feature in the training sample, and send them to the joint discriminator for discrimination; Calculate the joint loss of the generative adversarial network, the feature alignment network and the joint discriminator to complete the network training; Input the images in the database into the generator in the trained generative adversarial network. The generator outputs a pseudo-image and the image to be recognized and inputs them into the feature alignment network to extract the corresponding image features respectively. Through the comparison of the image features, the recognition of the image to be recognized is completed; Among them, the joint loss is expressed as follows: Among them, L pix represents the loss of the generative adversarial network, L feat represents the loss of the feature alignment network, L D represents the loss of the joint discriminator, represents the adversarial loss of the generative adversarial network, represents the adversarial loss of the feature alignment network, L cyc represents the cycle consistency loss of the generative adversarial network, λ cyc 、 represents the weight of the corresponding loss function, represents the classification loss of the generated images in the generative adversarial network, represents the triplet loss calculated by the generative adversarial network for the generated images, represents the classification loss calculated for the features in the feature alignment stage, represents the triplet loss calculated for the features in the feature alignment stage, represents the weight of the generative adversarial network, represents the loss when the joint discriminator discriminates that the image-feature pair is true, represents the loss when the joint discriminator discriminates that the image-feature pair is false; Among them, (x, m) represents the image feature pair input to the joint discriminator, and X′ ir represents the generated fake second image, X ir represents the real second image, M ir represents the feature map extracted from the real second image by the feature alignment network, M′ ir represents the feature map extracted from the fake second image by the feature alignment network, D j (x, m) represents the output of the joint discriminator; Among them, the calculation formula of the joint discriminator loss is as follows: in, Indicates that the joint discriminator identifies the image feature pair as true, Denotes that the joint discriminator identifies the image feature pair as false, D j (x,m) is the output of the joint discriminator. When the identification is true, the output is 1, and when the identification is false, the output is 0. E is the mathematical expectation. (X ir ,M ir ) represents the real image feature pair under the same identity, is the same as (X ir ,M ir )’s pseudo image feature pairs with the same identity, is the same as (X ir ,M ir ) Real image feature pairs under different identities; in, Represents the X extracted from the feature alignment network ir and X′ ir Features calculate the classification loss, p() is the predicted probability that the input image is correctly classified to its true identity, Indicates the calculation of triplet loss for the generated image; L cyc =‖G p' (G p (X rgb ))-X rgb ‖1+‖G p (G p′ (X ir ))-X ir ‖1; Among them, G p Represents the generator, which generates a pseudo second image from the first image, G p′ is also a generator that generates the pseudo second image back to the first image, Represents the generated image X′ ir Calculate the classification loss. Represents the generated image X′ ir and the real infrared image X ir Calculate triplet loss, L cyc represents the cycle loss function, L tri represents the triplet loss function; The backbone network of the feature alignment network adopts Resnet-50, including a first convolutional layer, a second convolutional layer, a third convolutional layer and a fourth convolutional layer. An NAM attention mechanism module is set after each convolutional layer, and an MAM modality mitigation module is also set after the NAM attention mechanism modules of the third convolutional layer and the fourth convolutional layer; The pooling layers are removed from the first convolutional layer, the second convolutional layer, the third convolutional layer and the fourth convolutional layer.

2. The cross-modal pedestrian re-identification method based on a generative adversarial network according to claim 1, characterized in that The NAM attention mechanism module is located after the batch normalization layer of each convolutional layer.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on difficult quintuple

    CN111597876A