A Pedestrian Attribute Recognition Method Based on a Dual-Branch Self-Attention Network

The dual-branch self-attention network enhances person attribute recognition by capturing high-order attribute information and contextual relationships, addressing limitations in existing technologies to improve accuracy and expand application scenarios.

CN115439884BActive Publication Date: 2025-07-15SHANDONG UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210978456.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2025-07-15
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

The prior art fails to fully utilize the correlation between attributes and the contextual relationship of image areas in pedestrian attribute recognition, resulting in insufficient recognition accuracy, especially in complex backgrounds and occlusions.

Method used

Using a method based on a dual-branch self-attention network, the attribute correlation is mined through the second-order self-attention module and the attribute self-attention module, and the long-term dependency is captured in combination with the context self-attention module, and a dual-branch network is constructed for feature extraction and classification.

Benefits of technology

It improves the accuracy and robustness of pedestrian attribute recognition, broadens the application scenarios, and is suitable for pedestrian attribute recognition in large-scale monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115439884B_ABST
    Figure CN115439884B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian attribute recognition method based on a dual-branch self-attention network, belonging to the technical field of pattern recognition, which includes the following steps: image data collection and processing, constructing and dividing a data set; image feature extraction; constructing a dual-branch self-attention pedestrian attribute recognition network model to obtain image attribute-related information and context region information; training to output a dual-branch self-attention network model with good performance; and real-time collecting pedestrian images through a surveillance video, and automatically recognizing pedestrian attributes by using the trained two-branch self-attention network model. The present invention uses a dual-branch self-attention network to obtain attribute-related information and context relationships, and combines constraint losses, etc. to restrict the classification of attribute features, improving the attribute classification performance and enabling stable pedestrian attribute recognition in large-scale surveillance scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a pedestrian attribute recognition method based on a dual-branch self-attention network. Background Art

[0002] Pedestrian attributes are a series of high-level visual semantic features of humans, including demographic information (such as gender, age, etc.) and appearance attributes (such as hairstyle, hair color, clothing type and color), etc. The main content of the pedestrian attribute recognition task is to describe the features of a person given an image of a person from a predefined list of attributes, which is of great significance for pedestrian analysis and detection. Pedestrian attribute recognition can be applied in many fields. For example, in urban security and security, key targets can be quickly found from a large number of surveillance videos, and attributes such as gender, age, clothing, and walking posture can be analyzed; in commercial applications, modern urban service providers rely on information technologies such as big data and are gradually providing intelligent and personalized services for each person, and matching more accurate applicable products from each person's appearance and clothing style, etc.; in image retrieval, due to the increasing number of cameras in modern cities, a large amount of picture and video data is generated every day. How to achieve classification storage and image retrieval from these data faces huge challenges. Therefore, relevant attribute information can be used for automatic annotation and classification, providing an important basis for alleviating the data storage pressure and efficiently retrieving images.

[0003] Pedestrian attribute recognition is still a challenging task in real surveillance scenarios, where noises such as occlusion, complex background, and various views will reduce the recognition accuracy. The general process of the pedestrian attribute recognition classification algorithm based on pictures is as follows: 1) data partitioning, cropping pictures into a picture set with unified pixels and partitioning the data set; 2) inputting pictures, using backbone network model algorithms such as ResNet to extract pedestrian image features, and using a classifier to classify attribute features; 3) performing iterative training to find the optimal value and saving the model parameters. Currently, most attribute recognition technologies are designed based on standard convolutional neural networks. By collecting pedestrian samples obtained in surveillance scenarios and manually assigning labels, the recognition model is trained to enable the model to learn useful appearance expressions and action features from the samples and be able to recognize based on these features.

[0004] Previous work mainly solved the task of pedestrian attribute recognition from the following aspects:

[0005] 1) In the field of pedestrian attribute recognition, dozens of attributes usually need to be analyzed simultaneously. Among these attributes, some are closely related. For example, when the attributes of "skirt" and "long hair" appear, they are often associated with the attribute of "female gender". The attribute of clothing type can provide certain information for age judgment. By exploring the correlation between different attributes, the performance of attribute recognition can be effectively improved, which most previous methods have ignored.

[0006] 2) On the other hand, exploring the spatial context relationship in different image regions is also helpful for attribute recognition. An imaginable example is that when identifying the gender of a pedestrian, people tend to focus on multiple regions, such as around the head, the area of clothing and carried items, etc., that is, the regional context relationship existing in the picture needs to be considered. Although deep convolutional networks have achieved great success in pedestrian attribute recognition, the context relationship has not been fully utilized. This is because the receptive field of the units in the deep convolutional network is severely limited, and it may not be able to understand the global background and capture the long-distance dependencies of different regions. Summary of the Invention

[0007] To solve the above problems, the present invention proposes a pedestrian attribute recognition method based on a dual-branch self-attention network. First, it mines the high-order information between attributes, combines the first-order information, and uses the attribute self-attention module and the constraint function to obtain the attribute correlation information. Then, it uses the aggregated context information and the context self-attention module to capture the long-term dependencies of different regions, and realizes the pedestrian attribute recognition with high performance from two aspects of obtaining the attribute correlation features and the attribute context relationship. While improving the detection accuracy, it broadens the application scenarios of attribute recognition and is expected to create considerable economic value.

[0008] The technical solution of the present invention is as follows:

[0009] A pedestrian attribute recognition method based on a dual-branch self-attention network, comprising the following steps:

[0010] Step 1, Image data collection and processing, constructing and dividing the data set;

[0011] Step 2, Image feature extraction;

[0012] Step 3, Constructing a dual-branch self-attention pedestrian attribute recognition network model to obtain image attribute-related information and context region information. The dual-branch includes an attribute branch and a context branch. The attribute branch includes a second-order self-attention module and an attribute self-attention module. The context branch includes a region feature mapping module and a context self-attention module;

[0013] Step 4. Training to output a dual-branch self-attention network model with good performance;

[0014] Step 5: Collect pedestrian images in real time through surveillance videos, and use the trained two-branch self-attention network model to automatically identify pedestrian attributes.

[0015] Further, the specific process of step 1 is as follows: Extract pedestrian images from the surveillance video, and perform attribute annotation and cropping; uniformly crop the images into pictures with a size of 256×128 pixels to form a picture dataset D, and divide the dataset D into a training set D train and a test set D test .

[0016] Further, the specific process of step 2 is as follows: Use ResNet50 as the backbone network, and batch-enter the pictures using the batch processing method to obtain a feature map X∈R C×H×W , where H, W, and C represent the length, width, and dimension of the feature map respectively.

[0017] Further, the specific process of step 3 is as follows:

[0018] Step 3.1: Calculate the predicted value of the attribute branch based on the second-order self-attention module and the attribute self-attention module ;

[0019] Step 3.2: Calculate the predicted value of the context branch based on the context self-attention module ;

[0020] Step 3.3: The final classification prediction result is expressed as and the average value of, and use Sigmoid for weighted processing to obtain the final attribute classification result. To correspond to the instance label value, take 1 for the final attribute classification result greater than 0.5 and take 0 for the result less than or equal to 0.5.

[0021] Further, the specific process of step 3.1 is as follows:

[0022] The calculation process of the second-order self-attention module is as follows:

[0023] Step 3.1.1: The feature map X is convolved by 1×1 to obtain a three-dimensional tensor with a dimension of , and then the dimension of this tensor is changed to a two-dimensional matrix Q = H×W, and the same operation is repeated three times to generate three projections of the feature map X, namely K S , Q S and V S , and the dimensions are all Among them, the input channel is C-dimensional, and the output channel is -dimensional, and r represents the sampling reduction ratio;

[0024] Step 3.1.2: Use the projection KS and the projection Q S Calculate the covariance matrix As shown in Equation (1),

[0025]

[0026] where I and 1 are the Q-dimensional identity matrix and the all-ones matrix, respectively;

[0027] Step 3.1.3: Process the covariance matrix Σ using the Softmax function and use Q as the scaling factor of the covariance matrix;

[0028] Step 3.1.4: Dot-multiply the result obtained in Step 3.1.3 with V S to obtain as shown in Equation (2), and expand into a three-dimensional tensor with a shape of ;

[0029]

[0030] Step 3.1.5: Finally, splice and the feature map X through a 1×1 convolution to obtain a dimension of first-order features together as the input to the subsequent attribute self-attention module;

[0031] The calculation process of the attribute self-attention module is as follows:

[0032] Step 3.1.6: The input feature map with a shape of is convolved through different 1×1 convolutions and the last two dimensions of the data are transformed into one dimension to obtain K A , Q A and V A , where K A , Q A and V A represent the three input projections of the attribute self-attention module, where Q A , N H and M are the number of attention heads and the number of attributes, respectively, and D A represents the dimension of the attribute feature map;

[0033] Step 3.1.7: According to Equation (3), multiply the matrix K A by the transpose of the matrix Q A and obtain the attention scores of each attribute through the Sigmoid operation This score represents the probability that the input contains a certain attribute, where M represents the number of attributes;

[0034]

[0035] Step 3.1.8. Multiply the above attention scores by V A to obtain the predicted values corresponding to each attention head

[0036] Step 3.1.9. Then sum along the N H dimension of and stretch it into a preliminary prediction result of the attribute self-attention module with a dimension of M

[0037] Step 3.1.10. Design a constraint loss function to limit the prediction scores, as shown in Equation (4),

[0038]

[0039] where ω j represents the weight of the j-th attribute in the training dataset, and M represents the number of attributes; p ij , y ij represent the predicted value and the label value of the j-th attribute of the i-th sample, respectively;

[0040] Step 3.1.11. Finally, perform linearization processing on the preliminary prediction result and add it to K A to obtain the final prediction result of the attribute branch expressed as Equation (5),

[0041]

[0042] where W A ∈R M×M represents the linearization processing classifier parameters

[0043] Furthermore, the specific process of Step 3.2 is as follows:

[0044] Step 3.2.1. First, use a tokenization scheme to aggregate the feature map into K compact visual tokens, where K << H × W; for the input feature map X ∈ R H×W×C , perform token soft assignment through a local aggregation descriptor vector calculation kernel, and calculate the k-th visual token T k ∈R K×C , as shown in Equation (6),

[0045]

[0046] where α k (x l ) represents the soft assignment of the l-th local feature x lThe weighted value assigned to the k-th visual marker, c k is the k-th learnable anchor point;

[0047] Step 3.2.2: Use the self-attention module to capture the context relationships between different visual markers; adopt the multi-head self-attention layer and the feed-forward neural network to propagate messages among all visual markers, and their states are updated through Equation (7).

[0048]

[0049] where, d1 represents the adjustment factor; Q c , is obtained by performing a linear transformation on the output global feature T k , and Q c , K c , V c represent the three input projections of the context self-attention; W T ∈R M×M represents the classifier parameter in the linearization process;

[0050] Step 3.2.3: Then, perform matrix multiplication on Q c , K c , and then through the Softmax operation and the Dropout operation, randomly prune 50% of the parameters to obtain the context attention scores, and introduce the residual structure by adding them to V c to accelerate convergence; finally, obtain the context branch prediction value through the linear layer and the batch normalization operation As shown in Equation (8),

[0051]

[0052] where, BN represents the batch normalization operation, and W C ∈R M×C represents the classifier parameter in the linearization process.

[0053] Furthermore, the specific process of Step 4 is as follows:

[0054] First, use the training set D train to train the model, set the learning rate to 0.0001, the number of iterations to 30, use the Adam optimizer, and input 64 images per iteration;

[0055] Then calculate the two-branch loss functions and the constraint loss, and obtain the total loss , and minimize the loss value; among them, the two-branch loss functions and both adopt the weighted cross-entropy loss function as shown in Equation (9).

[0056]

[0057] Among them, ω j represents the weight occupied by the j-th attribute in the training dataset; M represents the number of attributes; p ij , y ij represent the predicted value and the label value of the j-th attribute of the i-th sample;

[0058] The total loss function of the entire dual-branch self-attention network model is expressed as the following formula (10). According to the obtained total loss minimize the loss value,

[0059]

[0060] where λ1, λ2, and λ3 are the weights of the two-branch loss function and the constraint loss function respectively;

[0061] Finally, use D test to test the model. After each training, test on the test set D test and compare the test results, and save the network model parameters with the best test set results.

[0062] The beneficial technical effects brought by the present invention:

[0063] The present invention uses a two-branch self-attention network to obtain attribute-related information and context relationships, and combines constraint losses, etc. to restrict the attribute feature classification, thereby improving the attribute classification performance. The present invention can stably achieve pedestrian attribute recognition in large-scale monitoring scenarios, and can be applied to fields such as personnel image retrieval, security and security detection, and commercial advertising placement, improving the performance and practicality of attribute recognition technology, and is of great significance for accelerating scientific and technological development, improving people's living standards, and promoting the improvement of social productivity. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a flowchart of the pedestrian attribute recognition method based on the dual-branch self-attention network of the present invention;

[0065] Figure 2 is a schematic diagram of the overall structure of the dual-branch self-attention network model of the present invention;

[0066] Figure 3 is a schematic diagram of the calculation process of the second-order self-attention module model of the present invention;

[0067] Figure 4 is a schematic diagram of the calculation process of the attribute self-attention module of the present invention;

[0068] Figure 5 is a schematic diagram of the calculation process of the context self-attention module of the present invention. Detailed implementation manners

[0069] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners:

[0070] In the field of pedestrian attribute recognition, it is often necessary to centrally analyze several attributes such as gender, age, sunglasses, clothing type, hairstyle, etc. Among these attributes, some are closely related. For example, the "skirt" attribute is usually associated with the "female" attribute, and the clothing type attribute can provide certain information to judge age. Therefore, exploring the relationships between attributes helps to improve the performance of attribute recognition. Exploring the context relationships of different image regions also helps attribute recognition. For example, when recognizing the gender of a pedestrian, people tend to focus on multiple regions, such as the region around the head, the body wear, etc., and consider their context relationships. Therefore, the present invention describes a pedestrian attribute recognition method based on a dual-branch self-attention network, which comprehensively obtains the information related to the attributes of the input picture and the context region information, and high-performance realizes the key algorithms such as "picture feature extraction" and "attribute feature classification" necessary for pedestrian attribute recognition.

[0071] The present invention proposes a novel dual-branch network (i.e., an attribute branch and a context branch) for pedestrian attribute recognition. The attribute branch proposes a second-order self-attention module to make full use of the limited feature dimension information and further improve the feature representation ability; the context branch uses a tokenization scheme to aggregate the feature maps and proposes a context self-attention module to explore the context relationships based on multiple visual tokens.

[0072] As Figure 1 shown, a pedestrian attribute recognition method based on a dual-branch self-attention network includes the following steps:

[0073] Step 1, Image data collection and processing, constructing and dividing the dataset. Extract pedestrian images from the surveillance video, and perform attribute annotation and cropping; uniformly crop the images into pictures with a size of 256×128 pixels to form a picture dataset D, and divide the dataset D into a training set D train and a test set D test .

[0074] Step 2, Image feature extraction. Use ResNet50 as the backbone network, and batch input the pictures using the batch processing method to obtain a feature map X∈R C×H×W , where H, W, and C respectively represent the length, width, and dimension of the feature map, and are respectively set to 8, 4, and 2048 in the embodiment of the present invention. Alternatively, use a ResNet101 network model with deeper layers and more parameters for image feature extraction to achieve better recognition accuracy.

[0075] Step 3, Construct as Figure 2The shown dual-branch self-attention pedestrian attribute recognition network model obtains image attribute-related information and context region information; the dual-branch includes an attribute branch and a context branch, the attribute branch includes a second-order self-attention module and an attribute self-attention module, and the context branch includes a region feature mapping module and a context self-attention module. The total loss of the model consists of three parts, where the loss of the attribute branch and the loss of the context branch are represented by respectively, and the constraint loss in the attribute branch is represented by respectively.

[0076] Step 3.1: Calculate the predicted value of the attribute branch based on the second-order self-attention module and the attribute self-attention module; the specific process is as follows:

[0077] The calculation process of the second-order self-attention module is as Figure 3 shown,

[0078] Step 3.1.1: The feature map X passes through a 1×1 convolution (the input channel is C-dimensional, and the output channel is -dimensional, where r represents the sampling reduction ratio, and r is set to 8 in the embodiment of the present invention) to obtain a tensor of dimension , and then the dimension of this tensor is changed to a two-dimensional tensor Q = H×W, and the same operation is repeated three times to generate three projections of the feature map X, denoted as K S , Q S and V S , all with dimensions of

[0079] Step 3.1.2: Use the projection K S and the projection Q S to calculate the covariance matrix as shown in Equation (1),

[0080]

[0081] where I and 1 are the Q-dimensional identity matrix and the all-one matrix respectively;

[0082] Step 3.1.3: Process the covariance matrix Σ using the Softmax function and use Q as the scaling factor of the covariance matrix, and this step can play a role in adjustment;

[0083] Step 3.1.4: Dot-multiply the result obtained in Step 3.1.3 with V S to obtain the second-order self-attention value as shown in Equation (2), and expand into a tensor with a shape of ;

[0084]

[0085] Step 3.1.5, finally and the feature map X are convolved by 1×1 convolution to obtain a dimension of The first-order features are concatenated together and jointly used as the input of the subsequent attribute self-attention module.

[0086] The output of the above second-order self-attention module will be used as the input of the attribute self-attention module introduced below. The calculation process of the attribute self-attention module is as Figure 4 shown

[0087] Step 3.1.6, the output feature map of Step 3.1.5 (with a shape of ) is convolved by different 1×1 convolutions and the last two dimensions of the data are transformed (reshaped) into one dimension to obtain K A , Q A and V A Three matrices, K A , Q A and V A respectively represent the three input projections of the attribute self-attention module, where Q A , N H and M are the number of attention heads and the number of attributes respectively, D A represents the dimension of the attribute feature map, which is set to 256 in the embodiments of the present invention;

[0088] Step 3.1.7, according to Equation (3), multiply the transpose of matrix K A and matrix Q A , and obtain the attention scores of each attribute through the Sigmoid operation This score represents the probability that the input contains a certain attribute. In the formula, M represents the number of attributes;

[0089]

[0090] Step 3.1.8, multiply the above attention scores by V A to obtain the predicted values corresponding to each attention head

[0091] Step 3.1.9, then sum along the N H dimension for and stretch it into a preliminary prediction result of an attribute branch with a dimension of M

[0092] Step 3.1.10, in order to ensure the learning of attribute-specific features, a constraint loss function is designed to limit the prediction scores, as shown in Equation (4),

[0093]

[0094] Among them, ω j represents the weight occupied by the j-th attribute in the training dataset, and M represents the number of attributes; p ij , y i j represent the predicted value and the label value of the j-th attribute of the i-th sample respectively;

[0095] Step 3.1.11. Finally, perform a linearization process on the preliminary prediction result to improve the robustness of the model and add it to K A to obtain the final prediction result of the attribute branch which can be expressed as Equation (5),

[0096]

[0097] where W A ∈R M×M represents the classifier parameter for the linearization process.

[0098] Step 3.2. Calculate the predicted value of the context branch based on the context self-attention module;

[0099] Due to the influence of the monitoring camera's perspective in the real scenario, images often get distorted. However, there is often a certain relationship between the positions of body parts and the positions of attached items. Therefore, it is necessary to explore the context region relationship. In the context branch, visual tokens are extracted from the feature map and further used to explore the context relationship between different regions. The specific process is as follows.

[0100] Step 3.2.1. First, adopt a tokenization scheme to aggregate the feature map into K compact visual tokens, where K << H×W. For the input feature map X∈R H×W×C , calculate the kernel (VLAD core) through the Vector of Locally Aggregated Descriptors (VLAD) for soft assignment of tokens and calculate the k-th visual token T k ∈R K×C , as shown in Equation (6).

[0101]

[0102] Among them, α k (x l ) represents the weighted value of assigning the l-th local feature x l to the k-th visual token, and c k is the k-th learnable anchor point.

[0103] Step 3.2.2. As attachedFigure 5 As shown, a self-attention module is used to capture the context relationship between different visual tokens. A multi-head self-attention layer and a feed-forward neural network (FFN) are used to propagate messages among all visual tokens, and their states are updated through Equation (7).

[0104]

[0105] Among them, d1 represents a regulation factor, which refers to the input dimension divided by the number of multi-head attention heads, and is 256 in the embodiment of the present invention; Q c , is obtained by performing a linear transformation on the output global feature T k . The intermediate feature dimension n c1 in the embodiment of the present invention is 256, n c2 = 64, Q c , K c , V c represent three input projections of context self-attention; W T ∈R M×M represents the classifier parameters in the linearization process.

[0106] Step 3.2.3. Then, use the matrix multiplication of Q c , K c and the Softmax operation, and randomly crop 50% of the parameters through Dropout to obtain the context attention score. By adding it to V c , a residual structure is introduced to accelerate convergence. Finally, the context branch prediction value is obtained by using a linear layer (FC) and a batch normalization operation (BN) As shown in Equation (8),

[0107]

[0108] Among them, BN represents the batch normalization operation, and W C ∈R M×C represents the classifier parameters in the linearization process.

[0109] Step 3.3. The final classification prediction result is expressed as and The average value of is weighted by Sigmoid to obtain the final attribute classification result. In order to correspond to the instance label value, the final attribute classification result greater than 0.5 is taken as 1, and the one less than or equal to 0.5 is taken as 0.

[0110] Step 4. Train a double-branch self-attention network model with good output performance. The present invention searches for the optimal value of the model through iterative training. The specific process is as follows:

[0111] First, use the training set D trainTrain the model with a learning rate of 0.0001, 30 iterations, and the Adam optimizer. Input 64 images per iteration.

[0112] Then calculate the loss functions of the two branches and the constraint loss to obtain the total loss , and minimize the loss value. Among them, the loss functions of the two branches and both adopt the weighted cross - entropy loss function shown in Equation (9).

[0113]

[0114] where ω j represents the weight of the j - th attribute in the training dataset; M represents the number of attributes; p ij , y ij represent the predicted value and label value of the j - th attribute of the i - th sample.

[0115] The total loss function of the entire double - branch self - attention network model can be expressed as Equation (10) below. Minimize the loss value according to the obtained total loss ,

[0116]

[0117] where λ1, λ2, λ3 are the weights of the two - branch loss function and the constraint loss function respectively. In the embodiments of the present invention, λ1 = 1, λ2 = 1, λ3 = 0.1.

[0118] Finally, use D test to test the model. Test on the test set D test after each training, compare the test results, and save the network model parameters with the best test set results.

[0119] Alternatively, the AdamW optimizer algorithm can also be used to further accelerate the iteration process.

[0120] Step 5: Real - time collect pedestrian images through the monitored video, and use the trained two - branch self - attention network model for automatic recognition of pedestrian attributes.

[0121] To prove the feasibility and superiority of the present invention, comparative experiments were conducted on three commonly used attribute recognition datasets (PETA, PA00K, RAP). The benchmark models used were ResNet50 and a linear classifier. The recognition accuracies of the test results of the present model on the above datasets reached 87.70%, 82.27%, and 83.68% respectively, which were 2.59%, 2.89%, and 5.20% higher than those of the benchmark models respectively, fully demonstrating that the present invention can effectively improve the application effect of the existing human attribute recognition methods.

[0122] Certainly, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.

Claims

1. A pedestrian attribute recognition method based on a dual-branch self-attention network, characterized in that It includes the following steps: Step 1: Image data acquisition and processing, constructing and dividing the dataset; Step 2: Image feature extraction; Step 3: Constructing a dual-branch self-attention pedestrian attribute recognition network model to obtain image attribute-related information and context region information. The dual-branch includes an attribute branch and a context branch. The attribute branch includes a second-order self-attention module and an attribute self-attention module, and the context branch includes a context self-attention module; The specific process is as follows: Step 3.1: Calculate the predicted value of the attribute branch based on the second-order self-attention module and the attribute self-attention module The calculation process of the second-order self-attention module is as follows: Step 3.1.

1. The feature map X is convolved by a 1×1 convolution to obtain a three-dimensional tensor with a dimension of , and then the dimension of this tensor is changed to a two-dimensional matrix Q = H × W. The same operation is repeated three times to generate three projection matrices of the feature map X, namely K S , Q S and V S , and the dimensions are all Among them, the input channel is C-dimensional, and the output channel is -dimensional, and r represents the sampling reduction ratio; Step 3.1.2, use projection K S and projection Q S to calculate the covariance matrix As shown in Equation (1), Among them, I and 1 are the Q-dimensional identity matrix and the all-ones matrix respectively; Step 3.1.3: Processing the covariance matrix Σ using the Softmax function and using Q as the scaling factor of the covariance matrix; Step 3.1.4: Multiply the result obtained in Step 3.1.3 with V S by dot product to obtain as shown in Equation (2), and expand into a tensor with a shape of ; Step 3.1.

5. Finally, and the feature map X are concatenated with the first-order features obtained by 1×1 convolution with a dimension of as the input to the subsequent attribute self-attention module; The calculation process of the attribute self-attention module is as follows: Step 3.1.

6. The three-dimensional feature map with the input shape of is subjected to different 1×1 convolutions, and the last two dimensions of the data are transformed into one dimension to obtain K A , Q A and V A , which respectively represent the three input projection matrices of the attribute self-attention module. Among them, Q A , N H and M are respectively the number of attention heads and the number of attributes, and D A represents the dimension of the attribute feature map; Step 3.1.7: According to Equation (3), multiply the matrix Q A by the transpose of the matrix K A , and then obtain the attention scores of each attribute through the Sigmoid operation This score represents the probability that the input contains a certain attribute, where M represents the number of attributes; Step 3.1.

8. Multiply the above attention scores by V A to obtain the predicted values corresponding to each attention head Step 3.1.9, then sum along the N H dimension pair to obtain a preliminary prediction result of the attribute self-attention module with dimension M by stretching it Step 3.1.10, Design Constraint Loss Function to limit the prediction score, as shown in Equation (4), where, ω j represents the weight of the j-th attribute in the training dataset, and pij and yij represent the predicted value and the labeled value of the j-th attribute of the i-th sample, respectively; Step 3.1.

11. Finally, perform linearization on the preliminary prediction result and add it to K A to obtain the final prediction result of the attribute branch which is expressed as Equation (5). Among them, W A ∈ R M×M represents the parameters of the linearized classifier; Step 3.2: Calculate the predicted value of the context branch based on the context self-attention module Step 3.3: The final classification prediction result is expressed as and The average value of is weighted by Sigmoid to obtain the final attribute classification result. For the final attribute classification result, if it is greater than 0.5, take 1; if it is less than or equal to 0.5, take 0. Step 4: Training a dual-branch self-attention network model with good output performance; Step 5: Real-time collecting pedestrian images through a monitored video and automatically recognizing pedestrian attributes using the trained dual-branch self-attention network model.

2. The pedestrian attribute recognition method based on a dual-branch self-attention network according to claim 1, wherein The specific process of the said step 1 is as follows: extract pedestrian images from the surveillance video, and perform attribute annotation and cropping; uniformly crop the images into pictures with a size of 256×128 pixels to form a picture dataset D, and divide the dataset D into a training set D train and a test set D test .

3. The pedestrian attribute recognition method based on a dual-branch self-attention network according to claim 1, characterized in that The specific process of the said step 2 is as follows: Using ResNet50 as the backbone network, batch input images by means of the batch processing method to obtain a feature map X ∈ R C×H×W , where H, W, and C respectively represent the length, width, and dimension of the feature map.

4. The pedestrian attribute recognition method based on a dual-branch self-attention network according to claim 1, wherein, The specific process of Step 3.2 is as follows: Step 3.2.

1. First, a tokenization scheme is adopted to aggregate the feature map into K compact visual tokens, where K << H×W; for the input feature map X ∈ R H×W×C , a kernel for calculating the local aggregation descriptor vector is used for soft assignment of tokens, and the k-th visual token T k ∈ R K×C is calculated as shown in Equation (6). Among them, α k (x l ) represents the weighted value assigned to the k-th visual token for the l-th local feature x l , and c k is the k-th learnable anchor point; Step 3.2.2: Using the self-attention module to capture the context relationship between different visual tokens; using the multi-head self-attention layer and the feed-forward neural network to propagate messages among all visual tokens, and their states are updated through Equation (7), Among them, d1 represents a regulation factor; is obtained by performing a linear transformation on the k-th visual token T k output, where Q c , K c , V c represent three input projections of the context self-attention; W T ∈R M×M represents the linearized processing classifier parameters; Step 3.2.3, then, for Q c , K c , perform matrix multiplication, then through the Softmax operation and the Dropout operation, randomly crop 50% of the parameters to obtain the context attention score, and introduce a residual structure by adding it to V c to accelerate convergence; finally, obtain the context branch prediction value through a linear layer and a batch normalization operation As shown in Equation (8), Among them, BN represents the batch normalization operation, and W C ∈R M×C represents the parameter in the linear layer.

5. The pedestrian attribute recognition method based on a dual-branch self-attention network according to claim 1, wherein The specific process of Step 4 is as follows: First, use the training set D train to train the model. Set the learning rate to 0.0001, the number of iterations to 30, use the Adam optimizer, and input 64 images per iteration; Then calculate the two branch loss functions and the constraint loss to obtain the total loss Minimize the loss value; among them, the two branch loss functions and both adopt the weighted cross-entropy loss function shown in Equation (9), where ω j represents the weight of the j-th attribute in the training dataset; M represents the number of attributes; pij and yij represent the predicted value and the labeled value of the j-th attribute of the i-th sample; The total loss function of the entire double-branch self-attention network model is expressed as Equation (10) below. According to the obtained total loss minimize the loss value Among them, λ1, λ2, and λ3 are the weights of the loss functions of the two branches and the constraint loss function respectively; Finally, use D test to test the model. After each training, test on the test set D test and compare the test results. Then save the network model parameters with the best test set results.

Citation Information

Patent Citations

  • Pedestrian re-identification method based on natural language description

    CN110909673A