Pedestrian re-identification method and system based on double attention network

By introducing a dual attention network into the pedestrian re-identification system, the problem that a single enhancement method cannot eliminate the impact of noise is solved, the recognition accuracy and stability are improved, and the model performance is improved.

CN119942636AActive Publication Date: 2025-05-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411946178.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

When the existing pedestrian re-identification system processes noise information in pictures, a single enhancement method cannot effectively eliminate the impact of noise, resulting in insufficient recognition accuracy and stability.

Method used

Using a dual attention network method, a dual attention network is formed by constructing the first and second attention modules and inserting them into the basic convolutional network to capture key information in the image more comprehensively.

Benefits of technology

It improves the accuracy and stability of pedestrian re-identification, comprehensively improves model performance, and can handle noise information more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942636A_ABST
    Figure CN119942636A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and particularly provides a pedestrian re-identification method and system based on a double attention network. The method comprises the following steps: extracting image data from a pedestrian re-identification data set, and carrying out standardization processing on the extracted image data to obtain a processed training set and a test set; constructing a basic convolutional network for pedestrian re-identification; constructing a dual attention module, wherein the dual attention module comprises a first attention module and a second attention module; inserting the double attention module into the basic convolutional network to obtain a double attention network; according to the method, the double attention network is trained according to the processed training set, and the trained double attention network is verified through the processed test set, so that pedestrian re-identification is realized, the double attention is introduced to be inserted into the basic convolutional network, and key information in a picture is enhanced for multiple times, so that pedestrian re-identification is realized. The accuracy and stability of re-identification are improved, and the model performance is comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a pedestrian re-identification method and system based on a dual attention network. Background Art

[0002] With the rapid development of science and technology, processing large amounts of video and image data is particularly critical. Large-scale processing cannot be achieved manually, so the task of using computers for pedestrian re-identification has emerged.

[0003] By giving a person image to be queried, the person re-identification system can be used to retrieve and visualize the same person images in the dataset. Traditional person re-identification systems mainly rely on supervised learning, which requires a lot of manual labeling and cannot be deployed on a large scale. Therefore, unsupervised learning methods are very popular. Unsupervised learning methods can assign pseudo labels through clustering and other methods without labeled data, and then use neural networks for training and recognition.

[0004] In most existing unsupervised training methods, a single attention mechanism or feature learning method is usually used to capture important information in the image for noise information such as the inherent background in the image. However, a single enhancement method cannot more comprehensively eliminate the impact of noise information. Summary of the invention

[0005] In view of this, the present invention provides a pedestrian re-identification method and system based on a dual attention network, so as to improve the accuracy and stability of re-identification and improve the model performance.

[0006] In a first aspect, the present invention provides a pedestrian re-identification method based on a dual attention network, the method comprising: Step 1: extract image data from the person re-identification dataset, and perform standardization on the extracted image data to obtain the processed training set and test set; Step 2: Build a basic convolutional network for person re-identification; Step 3: construct a dual attention module, which includes a first attention module and a second attention module; Step 4: Insert the dual attention module into the basic convolutional network to obtain a dual attention network; Step 5: Train the dual attention network according to the processed training set, and verify the trained dual attention network through the processed test set to achieve pedestrian re-identification.

[0007] Optionally, the step 2 includes: The basic convolutional network is a deep residual network, which contains multiple residual blocks, each residual block contains multiple residual units, each residual unit includes two 3×3 convolutional layers, and is activated by ReLU activation function; the main structure of the basic convolutional network is input layer, initial convolution layer, maximum pooling layer, first residual block, second residual block, third residual block, fourth residual block, global average pooling layer, fully connected layer, output layer; the first residual block contains 3 residual units, each with an output channel of 256, the second residual block contains 4 residual units, each with an output channel of 512, the third residual block contains 6 residual units, each with an output channel of 1024, and the fourth residual block contains 3 residual units, each with an output channel of 2048.

[0008] Optionally, the first attention module in step 3 includes: First, input the image feature map x, and obtain the feature y through adaptive average pooling and 1×1 convolution operation, which is expressed as: ; Then the feature y is subjected to 1×1 convolution and activation function Softmax to obtain the channel weight coefficient A1, which is expressed as: ; Then, the unit matrix A0 and the parameter matrix A2 are constructed to adjust the weighting coefficients of each channel; the final weighting coefficient matrix is: ; Multiply the feature y with the weighting coefficient matrix, and then go through 1×1 convolution, activation function ReLU, 1×1 convolution and activation function Sigmoid to get the final weighting matrix , whose expression is: ; Finally, the feature map x is multiplied by the weighting matrix y' to obtain the final weighted feature x'.

[0009] Optionally, the second attention module in step 3 includes channel attention and spatial attention; For spatial attention, the number of input channels d, the number of groups g, and the image After that, the number of channels of the image is first grouped according to the preset number of groups through the grouping operation. The number of channels processed in each group is the number of input channels d divided by the number of groups g, and the grouped features are obtained. ; Then, the images of each group are globally averaged pooled in the height and width directions respectively. The specific formula is: ,in represents the input feature of the cth channel, H and W represent the height and width of the feature respectively; the two pooling results are concatenated together, and two channel weighting coefficients are generated through 1×1 convolution, activation function Sigmoid and group normalization operations, which are and , which is used to indicate the importance of the image in the height and width directions; finally, the weighted coefficient is multiplied by the original feature map to obtain a new weighted feature map , the specific formula is: ; For channel attention, the grouped features Processed by 3×3 convolution to extract the feature map of the image Next, the feature maps and feature map Perform global average pooling and Softmax operations to obtain feature maps and feature map , where the specific formula of Softmax is: ,in is the pooling result of the cth channel, is the pooling result of the i-th channel; then, the feature map and feature map Perform reshape operation to get feature map and feature map ; and combine the coefficients by matrix multiplication to generate the weighting matrix , the specific formula is: ; Finally, the obtained weighted matrix Features after activation function Sigmoid and grouping Multiply and reshape to get the final weighted feature map .

[0010] Optionally, step 4 includes: The number of input channels of the first attention module is set to 256 and 2048, respectively, and they are inserted into the first residual block and the fourth residual block of the basic convolutional network respectively; the input channels of the second attention module are set to 512 and 1024, respectively, and they are inserted into the second residual block and the third residual block of the basic convolutional network respectively; the basic convolutional network after inserting the first attention module and the second attention module is the final dual attention network, and pedestrian re-identification is achieved through the dual attention network.

[0011] Optionally, step 5 includes: The processed training set is input into the dual attention network for training. Through the combination of the first attention module and the second attention module, the key channel information and spatial information in the image are captured. During the training process, the triple loss and cross entropy loss functions are used to optimize the model parameters. The triplet loss formula is: ; in Representation sample and The Euclidean distance of is the anchor image, i.e. the reference image selected by the image library; is a positive sample image, that is, an image belonging to the same person as the anchor image; is a negative sample image, that is, an image belonging to a different person than the anchor image; is a hyperparameter, which indicates the minimum distance difference between positive samples and negative samples; The cross entropy loss function formula is: ; Where N represents the total number of images in the dataset; represents the true label of image i; , represents the predicted probability that the input image belongs to each category; U represents the total number of categories, Represents the characteristics of the input image; The final loss function formula is: ; in is the weight coefficient, balancing the two losses; The effect of the model is demonstrated by drawing a curve graph of the recognition accuracy index; the performance of the model on the test set is visualized by showing the test image and its top K matching results.

[0012] In a second aspect, the present invention provides a pedestrian re-identification system based on a dual attention network, the system comprising: A data processing module is used to extract image data from the pedestrian re-identification data set and perform standardization on the extracted image data to obtain a processed training set and test set; A model building module is used to build a basic convolutional network for pedestrian re-identification; build a dual attention module, which includes a first attention module and a second attention module; insert the dual attention module into the basic convolutional network to obtain a dual attention network; A model training module, used to train the dual attention network based on the processed training set; Results visualization module for validating the trained dual attention network for person re-ID with the processed test set.

[0013] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the pedestrian re-identification method based on a dual attention network in the first aspect or any possible implementation of the first aspect.

[0014] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to execute the pedestrian re-identification method based on a dual attention network in the first aspect or any possible implementation of the first aspect.

[0015] In the technical solution provided by the present invention, the method includes extracting image data from a pedestrian re-identification data set, and standardizing the extracted image data to obtain a processed training set and a test set; constructing a basic convolutional network for pedestrian re-identification; constructing a dual attention module, which includes a first attention module and a second attention module; inserting the dual attention module into the basic convolutional network to obtain a dual attention network; training the dual attention network according to the processed training set, and verifying the trained dual attention network through the processed test set to achieve pedestrian re-identification. The method introduces dual attention into the basic convolutional network, and enhances the key information in the picture multiple times, thereby improving the accuracy and stability of re-identification and comprehensively improving the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A flowchart of a pedestrian re-identification method provided by an embodiment of the present invention; Figure 2 A schematic diagram of a dual attention network provided by an embodiment of the present invention; Figure 3 A schematic diagram of a pedestrian re-identification system provided by an embodiment of the present invention; Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] It should be clear that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.

[0021] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0022] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0023] Figure 1 A flowchart of a pedestrian re-identification method provided by an embodiment of the present invention, such as Figure 1 As shown, the method includes: Step 1: Extract image data from the person re-identification dataset and standardize the extracted image data to obtain the processed training set and test set.

[0024] In the embodiment of the present invention, image data is extracted from the public pedestrian re-identification dataset Market-1501. The dataset consists of 1501 pedestrians captured by 6 cameras and 32668 detected pedestrian rectangles. Each image contains photos of pedestrians from different perspectives, different lighting conditions, and different backgrounds to ensure the diversity and generalization ability of the data.

[0025] Standardize the extracted image data to facilitate input into the model for training. Scale all images to a uniform size, and apply image enhancement techniques such as rotation, translation, flipping, and color jittering to increase data diversity and improve the robustness of the model. In order to reduce bias during training, use data standardization and normalization so that the mean of each pixel value or channel is zero and the variance is one, which helps the model converge faster.

[0026] Step 2: Build a basic convolutional network for person re-identification.

[0027] In the embodiment of the present invention, step 2 includes: The basic convolutional network is a deep residual network, which contains multiple residual blocks, each residual block contains multiple residual units, each residual unit includes two 3×3 convolutional layers, and is activated by ReLU activation function; the main structure of the basic convolutional network is input layer, initial convolution layer, maximum pooling layer, first residual block, second residual block, third residual block, fourth residual block, global average pooling layer, fully connected layer, output layer; the first residual block contains 3 residual units, each with an output channel of 256, the second residual block contains 4 residual units, each with an output channel of 512, the third residual block contains 6 residual units, each with an output channel of 1024, and the fourth residual block contains 3 residual units, each with an output channel of 2048.

[0028] Step 3: Construct a dual attention module, which includes a first attention module and a second attention module.

[0029] In the embodiment of the present invention, the first attention module in step 3 includes: First, input the image feature map x, and obtain the feature y through adaptive average pooling and 1×1 convolution operation, which is expressed as: ; Then the feature y is subjected to 1×1 convolution and activation function Softmax to obtain the channel weight coefficient A1, which is expressed as: ; Then, the unit matrix A0 and the parameter matrix A2 are constructed to adjust the weighting coefficients of each channel; the final weighting coefficient matrix is: ; Multiply the feature y with the weighting coefficient matrix, and then go through 1×1 convolution, activation function ReLU, 1×1 convolution and activation function Sigmoid to get the final weighting matrix , whose expression is: ; Finally, the feature map x is multiplied by the weighting matrix y' to obtain the final weighted feature x'.

[0030] In the embodiment of the present invention, the second attention module in step 3 includes channel attention and spatial attention; For spatial attention, the number of input channels d, the number of groups g, and the image After that, the number of channels of the image is first grouped according to the preset number of groups through the grouping operation. The number of channels processed in each group is the number of input channels d divided by the number of groups g, and the grouped features are obtained. ; Then, the images of each group are globally averaged pooled in the height and width directions respectively. The specific formula is: ,in represents the input feature of the cth channel, H and W represent the height and width of the feature respectively; the two pooling results are concatenated together, and two channel weighting coefficients are generated through 1×1 convolution, activation function Sigmoid and group normalization operations, which are and , which is used to indicate the importance of the image in the height and width directions; finally, the weighted coefficient is multiplied by the original feature map to obtain a new weighted feature map , the specific formula is: ; For channel attention, the grouped features Processed by 3×3 convolution to extract the feature map of the image Next, the feature maps and feature map Perform global average pooling and Softmax operations to obtain feature maps and feature map , where the specific formula of Softmax is: ,in is the pooling result of the cth channel, is the pooling result of the i-th channel; then, the feature map and feature map Perform reshape operation to get feature map and feature map ; and combine the coefficients by matrix multiplication to generate the weighting matrix , the specific formula is: ; Finally, the obtained weighted matrix Features after activation function Sigmoid and grouping Multiply and reshape to get the final weighted feature map .

[0031] Step 4: Insert the dual attention module into the basic convolutional network to obtain the dual attention network.

[0032] In the embodiment of the present invention, Figure 2 As shown, step 4 includes: The number of input channels of the first attention module is set to 256 and 2048, respectively, and they are inserted into the first residual block and the fourth residual block of the basic convolutional network respectively; the input channels of the second attention module are set to 512 and 1024, respectively, and they are inserted into the second residual block and the third residual block of the basic convolutional network respectively; the basic convolutional network after inserting the first attention module and the second attention module is the final dual attention network, and pedestrian re-identification is achieved through the dual attention network.

[0033] Step 5: Train the dual attention network according to the processed training set, and verify the trained dual attention network through the processed test set to achieve pedestrian re-identification.

[0034] In the embodiment of the present invention, step 5 includes: The processed training set is input into the dual attention network for training. Through the combination of the first attention module and the second attention module, the key channel information and spatial information in the image are captured. During the training process, the triple loss and cross entropy loss functions are used to optimize the model parameters. The triplet loss formula is: ; in Representation sample and The Euclidean distance of is the anchor image, i.e. the reference image selected by the image library; is a positive sample image, that is, an image belonging to the same person as the anchor image; is a negative sample image, that is, an image belonging to a different person than the anchor image; is a hyperparameter, which indicates the minimum distance difference between positive samples and negative samples; The cross entropy loss function formula is: ; Where N represents the total number of images in the dataset; represents the true label of image i; , represents the predicted probability that the input image belongs to each category; U represents the total number of categories, Represents the characteristics of the input image; The final loss function formula is: ; in is the weight coefficient to balance the two losses. The loss function and gradient descent algorithm are used to continuously optimize the model and improve performance.

[0035] The effect of the model is demonstrated by drawing a curve graph of the recognition accuracy index; the performance of the model on the test set is visualized by showing the test image and its top K matching results.

[0036] In the embodiment of the present invention, the present method is compared with other methods, and the final experimental results are shown in Table 1, where Bottom-up Clustering Approach (BUC), Softened Similarity Learning (SSL), Generative and Contrastive Learning (GCL), and Camera-Aware Proxies (CAP) are all pedestrian re-identification algorithms, mAP represents the average accuracy of the returned results matching the query target when the evaluation model queries a large number of pedestrian images, and R1, R5, and R10 represent the probability of the existence of the retrieval target in the top 1, 5, and 10 in the sorted list returned by the algorithm. As can be seen from Table 1, compared with other methods, the method of the present invention has significant superiority.

[0037] Table 1 .

[0038] Figure 3 A schematic diagram of a pedestrian re-identification system provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the system includes: a data processing module, a model building module, a model training module and a result visualization module.

[0039] The data processing module is connected to the model building module; the model building module is connected to the model training module; the model training module is connected to the result visualization module.

[0040] The data processing module is used to extract image data from the pedestrian re-identification dataset and standardize the extracted image data to obtain the processed training set and test set; the model building module is used to build a basic convolutional network for pedestrian re-identification; build a dual attention module, which includes a first attention module and a second attention module; insert the dual attention module into the basic convolutional network to obtain a dual attention network; the model training module is used to train the dual attention network according to the processed training set; the result visualization module is used to verify the trained dual attention network through the processed test set to achieve pedestrian re-identification.

[0041] Traditional convolutional neural networks usually ignore the important spatial and channel information in the enhanced image, or only introduce single attention, which fails to capture important information more comprehensively. The present invention introduces dual attention and inserts different residual blocks in the deep residual network to enhance the key information in the image multiple times, thereby comprehensively improving the model performance.

[0042] Each step of the embodiment of the present invention may be performed by an electronic device, which includes but is not limited to a mobile phone, a tablet computer, a portable PC, a desktop computer, etc.

[0043] In the technical solution provided by the present invention, the method includes extracting image data from a pedestrian re-identification data set, and standardizing the extracted image data to obtain a processed training set and a test set; constructing a basic convolutional network for pedestrian re-identification; constructing a dual attention module, which includes a first attention module and a second attention module; inserting the dual attention module into the basic convolutional network to obtain a dual attention network; training the dual attention network according to the processed training set, and verifying the trained dual attention network through the processed test set to achieve pedestrian re-identification. The method introduces dual attention into the basic convolutional network, and enhances the key information in the picture multiple times, thereby improving the accuracy and stability of re-identification and comprehensively improving the model performance.

[0044] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the electronic device where the computer-readable storage medium is located is controlled to execute the above-mentioned embodiment of the pedestrian re-identification method based on the dual attention network.

[0045] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention, such as Figure 4 As shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, the pedestrian re-identification method based on the dual attention network in the embodiment is implemented. To avoid repetition, they are not described one by one here.

[0046] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 21 and does not constitute a limitation of the electronic device 21. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0047] The processor 211 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0048] The memory 212 may be an internal storage unit of the electronic device 21, such as a hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card (FlashCard), etc. equipped on the electronic device 21. Further, the memory 212 may also include both an internal storage unit of the electronic device 21 and an external storage device. The memory 212 is used to store computer programs and other programs and data required by network devices. The memory 212 may also be used to temporarily store data that has been output or is to be output.

[0049] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0050] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A pedestrian re-identification method based on dual attention network, characterized in that: The method comprises: Step 1: extract image data from the person re-identification dataset, and perform standardization on the extracted image data to obtain the processed training set and test set; Step 2: Build a basic convolutional network for person re-identification; Step 3: construct a dual attention module, which includes a first attention module and a second attention module; Step 4: Insert the dual attention module into the basic convolutional network to obtain a dual attention network; Step 5: Train the dual attention network according to the processed training set, and verify the trained dual attention network through the processed test set to achieve pedestrian re-identification.

2. The method according to claim 1, characterized in that The step 2 comprises: The basic convolutional network is a deep residual network, which contains multiple residual blocks, each residual block contains multiple residual units, each residual unit includes two 3×3 convolutional layers, and is activated by ReLU activation function; the main structure of the basic convolutional network is input layer, initial convolution layer, maximum pooling layer, first residual block, second residual block, third residual block, fourth residual block, global average pooling layer, fully connected layer, output layer; the first residual block contains 3 residual units, each with an output channel of 256, the second residual block contains 4 residual units, each with an output channel of 512, the third residual block contains 6 residual units, each with an output channel of 1024, and the fourth residual block contains 3 residual units, each with an output channel of 2048.

3. The method according to claim 1, characterized in that The first attention module in step 3 includes: First, input the image feature map x, and obtain the feature y through adaptive average pooling and 1×1 convolution operation, which is expressed as: ; Then the feature y is subjected to 1×1 convolution and activation function Softmax to obtain the channel weight coefficient A1, which is expressed as: ; Then, the unit matrix A0 and the parameter matrix A2 are constructed to adjust the weighting coefficients of each channel; the final weighting coefficient matrix is: ; Multiply the feature y with the weighting coefficient matrix, and then go through 1×1 convolution, activation function ReLU, 1×1 convolution and activation function Sigmoid to get the final weighting matrix , whose expression is: ; Finally, the feature map x is multiplied by the weighting matrix y' to obtain the final weighted feature x'.

4. The method according to claim 1, characterized in that: The second attention module in step 3 includes channel attention and spatial attention; For spatial attention, the number of input channels d, the number of groups g, and the image After that, the number of channels of the image is first grouped according to the preset number of groups through the grouping operation. The number of channels processed in each group is the number of input channels d divided by the number of groups g, and the grouped features are obtained. ; Then, the images of each group are globally averaged pooled in the height and width directions respectively. The specific formula is: ,in represents the input feature of the cth channel, H and W represent the height and width of the feature respectively; the two pooling results are concatenated together, and two channel weighting coefficients are generated through 1×1 convolution, activation function Sigmoid and group normalization operations, which are and , which is used to indicate the importance of the image in the height and width directions; finally, the weighted coefficient is multiplied by the original feature map to obtain a new weighted feature map , the specific formula is: ; For channel attention, the grouped features Processed by 3×3 convolution to extract the feature map of the image Next, we analyze the feature maps and feature map Perform global average pooling and Softmax operations to obtain feature maps and feature map , where the specific formula of Softmax is: ,in is the pooling result of the cth channel, is the pooling result of the i-th channel; then, the feature map and feature map Perform a reshape operation to obtain the feature map and feature map ; and combine the coefficients by matrix multiplication to generate the weighting matrix , the specific formula is: ; Finally, the obtained weighted matrix Features after activation function Sigmoid and grouping Multiply and reshape to get the final weighted feature map .

5. The method according to claim 1, characterized in that The step 4 comprises: The number of input channels of the first attention module is set to 256 and 2048, respectively, and they are inserted into the first residual block and the fourth residual block of the basic convolutional network respectively; the input channels of the second attention module are set to 512 and 1024, respectively, and they are inserted into the second residual block and the third residual block of the basic convolutional network respectively; the basic convolutional network after inserting the first attention module and the second attention module is the final dual attention network, and pedestrian re-identification is achieved through the dual attention network.

6. The method according to claim 1, characterized in that The step 5 comprises: The processed training set is input into the dual attention network for training. Through the combination of the first attention module and the second attention module, the key channel information and spatial information in the image are captured. During the training process, the triple loss and cross entropy loss functions are used to optimize the model parameters. The triplet loss formula is: ; in Representation sample and The Euclidean distance of is the anchor image, i.e. the reference image selected by the image library; is a positive sample image, that is, an image belonging to the same person as the anchor image; is a negative sample image, that is, an image belonging to a different person than the anchor image; is a hyperparameter, which indicates the minimum distance difference between positive samples and negative samples; The cross entropy loss function formula is: ; Where N represents the total number of images in the dataset; represents the true label of image i; , represents the predicted probability that the input image belongs to each category; U represents the total number of categories, Represents the characteristics of the input image; The final loss function formula is: ; in is the weight coefficient, balancing the two losses; The effect of the model is demonstrated by drawing a curve graph of the recognition accuracy index; the performance of the model on the test set is visualized by showing the test image and its top K matching results.

7. A pedestrian re-identification system based on dual attention network, characterized in that: The system comprises: A data processing module is used to extract image data from the pedestrian re-identification data set and perform standardization on the extracted image data to obtain a processed training set and test set; A model building module is used to build a basic convolutional network for pedestrian re-identification; build a dual attention module, which includes a first attention module and a second attention module; insert the dual attention module into the basic convolutional network to obtain a dual attention network; A model training module, used to train the dual attention network based on the processed training set; Results visualization module for validating the trained dual attention network for person re-ID with the processed test set.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the pedestrian re-identification method based on a dual attention network according to any one of claims 1 to 6.

9. An electronic device, characterized in that: include: one or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to perform the pedestrian re-identification method based on the dual attention network as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Truck picture re-identification method based on double-layer attention network

    CN115035476A

  • Method of semantically segmenting input image, apparatus for semantically segmenting input image, method of pre-training apparatus for semantically segmenting input image, training apparatus for pre-training apparatus for semantically segmenting input image, and computer-program product

    US20210406582A1