Knowledge distillation-based low-resolution pedestrian attribute identification method, system and device
By constructing a pedestrian attribute recognition model based on knowledge distillation, combined with the gradient rotation loss function, the gradient conflict problem of pedestrian attribute recognition in low-resolution images is solved, and the recognition accuracy is improved, and it is suitable for monitoring scenarios.
Patent Information
- Application Number
- CN202510234863.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing pedestrian attribute recognition method has gradient conflict problems under low-resolution image conditions, resulting in insufficient recognition accuracy, especially in monitoring scenarios.
Using a knowledge distillation method, by constructing the first and second pedestrian attribute pre-training models, multi-level and full-process knowledge distillation is carried out, and combined with the gradient rotation loss function, model training is optimized to improve the recognition performance of low-resolution images.
Without adding network parameters, the accuracy of pedestrian attribute recognition in low-resolution images is improved, and is suitable for low-resolution scenarios such as monitoring.
Smart Images

Figure CN120339933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of attribute recognition, and particularly relates to a low-resolution pedestrian attribute recognition method, system and device based on knowledge distillation. Background Art
[0002] With the continuous development of monitoring technology, extracting pedestrian attribute information from video monitoring has received wide attention. However, the performance limitations of monitoring devices result in relatively low-resolution monitoring images (such as security monitoring and unmanned retail). At the same time, complex environments such as low light and occlusion also pose great challenges to pedestrian attribute recognition.
[0003] Currently, there have been many research works on pedestrian attribute recognition, including supervised pedestrian attribute recognition methods, weakly supervised pedestrian attribute recognition methods, and pedestrian attribute recognition methods based on relationship modeling. Among them, supervised pedestrian attribute recognition methods include: decomposing pedestrian images into components using a poset, and classifying pedestrian attributes at different levels (poset, person, context) in combination with an SVM classifier; using a graph neural network and knowledge transfer to extract local region features from a pre-trained human parsing model, and then realizing pedestrian attribute recognition. However, supervised pedestrian attribute recognition methods do not fully consider low-resolution video monitoring scenarios and rely on manual annotation, resulting in limitations in applicable scenarios. Weakly supervised pedestrian attribute recognition methods introduce an attention mechanism, use an attention mask to remove interference, and improve and optimize attribute recognition through a combination of coarse and fine attention mechanisms. In addition, a spatial position memory bank is generated based on a spatial consistency module and dynamically updated to maintain the accuracy of attribute recognition. Although such methods have been improved, there are still limitations in terms of image resolution; pedestrian attribute recognition methods based on relationship modeling group the entire attribute list, and then use a human body region generation module to realize pedestrian attribute recognition. In addition, this method obtains spatial relationships and potential semantic relationships through a graph convolutional neural network, and improves the recognition result by integrating prediction information and prior information. Since such methods do not consider the diversity of the number and types of pedestrian attributes, gradient conflict problems are likely to occur during network optimization, which in turn affects the accuracy of pedestrian attribute recognition results. In addition, existing pedestrian attribute recognition methods do not solve the gradient conflict problem under low-resolution images, which affects the accuracy of pedestrian attribute recognition results. The present invention optimizes the problems of existing methods through multi-level and whole-process knowledge distillation and constructing a gradient rotation loss function for model training. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention provides a low-resolution pedestrian attribute recognition method, system and device based on knowledge distillation.
[0005] To solve the above technical problems, the present invention is solved by the following technical solutions:
[0006] A low-resolution pedestrian attribute recognition method based on knowledge distillation, comprising the following steps:
[0007] Obtain the original pedestrian image set, downsample the original pedestrian images to obtain the downsampled pedestrian image set, and preprocess the original pedestrian image set and the downsampled pedestrian image set respectively to obtain the pedestrian image set and the low-resolution pedestrian image set;
[0008] Group the pedestrian attributes in the pedestrian image set and the low-resolution pedestrian image set respectively to form pedestrian attribute groups;
[0009] Construct a first pedestrian attribute pre-training model, including a first backbone network module and several first sub-task modules, and the number of the first sub-task modules is equal to the number of the pedestrian attribute groups. The first backbone network module receives the pedestrian image set and performs feature extraction to obtain first pedestrian feature data. Each first sub-task module performs attention feature extraction on the first pedestrian feature data and performs attribute recognition for the corresponding pedestrian attribute group to obtain a first prediction result. Train the first pedestrian attribute pre-training model based on the pedestrian image set to obtain a first pedestrian attribute recognition model;
[0010] Construct a second pedestrian attribute pre-training model, including a second backbone network module and several second sub-task modules, and the number of the second sub-task modules is the same as the number of the pedestrian attribute groups. The second backbone network module receives the low-resolution pedestrian image set and performs feature extraction to obtain second pedestrian feature data. Each second sub-task module performs gradient optimization on the second pedestrian feature data, performs attention feature extraction, and performs attribute recognition for the corresponding pedestrian attribute group to obtain a second prediction result. Train the second pedestrian attribute pre-training model based on the low-resolution pedestrian image set to obtain a second pedestrian attribute recognition model;
[0011] Distill the knowledge of the first pedestrian attribute recognition model hierarchically and throughout the process and apply it to the second pedestrian attribute recognition model, and train the second pedestrian attribute recognition model to obtain a pedestrian attribute recognition model;
[0012] Analyze the pedestrian image to be recognized through the pedestrian attribute recognition model to obtain the pedestrian attribute recognition result.
[0013] As an implementable manner, the preprocessing includes horizontal flipping, cropping adjustment, tensor conversion, and normalization processing.
[0014] As an implementable manner, the pedestrian attribute group is obtained through the following steps:
[0015] Obtain the spatial position and semantic logic of the pedestrian attributes, divide the pedestrian attributes based on the spatial position and semantic logic to obtain the initial pedestrian attribute groups;
[0016] Obtain the gradient distance between the remaining pedestrian attributes and the initial pedestrian attribute group, which is expressed as follows:
[0017] Dist j,k = cos(g j , g k )
[0018] Among them, Dist j,k represents the gradient distance, g j represents the j-th pedestrian attribute, and g k represents the k-th pedestrian attribute;
[0019] For a preset distance threshold, divide the pedestrian attributes whose gradient distances meet the distance threshold into the corresponding initial pedestrian attribute groups until all pedestrian attributes are grouped to obtain the pedestrian attribute groups.
[0020] As an implementable manner, the training process of the first pedestrian attribute pre-training model is as follows:
[0021] The first backbone network module receives the pedestrian image set and extracts features from the pedestrian images to obtain the first pedestrian feature data, which is expressed as follows:
[0022] z T = θ T (I HR )
[0023] Each first sub-task module includes a first spatial attention layer, a first attribute classification layer, and a first result layer. The first spatial attention layer receives the first pedestrian feature data, performs pooling and convolution operations on the first pedestrian feature data to obtain a first spatial attention map, and combines the first pedestrian feature data for element-wise dot product to obtain a first attention feature map. The first spatial attention map and the first attention feature map are expressed as follows:
[0024]
[0025]
[0026] The first attribute classification layer receives the first attention feature map and makes a prediction to obtain a first prediction vector, which is expressed as follows:
[0027]
[0028] The first result layer splices and combines the first prediction vectors corresponding to all pedestrian attributes to obtain a first prediction result;
[0029] Based on the first prediction result and the true attribute annotation of the pedestrian image set, construct a first attribute loss function, and train and tune the parameters of the first pedestrian attribute pre-training model based on the pedestrian image set and in combination with the first attribute loss function to obtain a first pedestrian attribute recognition model. The first attribute loss function is expressed as follows:
[0030]
[0031] Among them, θ T represents the network parameters of the first backbone network module, I HR represents the pedestrian image, z T represents the first pedestrian feature data, represents the i-th first spatial attention map, represents the i-th first spatial attention layer, Conv 1×1 represents a 1×1 convolutional layer, Pool Max represents the global max pooling layer, Pool Avg represents the global average pooling layer, ⊕ represents the connection operation between feature maps, represents the i-th first attention feature map, represents the i-th first pedestrian feature data, ⊙ represents the element-wise dot product, represents the i-th first prediction vector, f i T represents the i-th first attribute classification layer, L1 represents the first attribute loss function, y i represents the i-th true attribute annotation, BCE represents the binary cross-entropy loss function, and N represents the number of pedestrian attribute groups.
[0032] As an implementable manner, the training process of the second pedestrian attribute pre-training model is as follows:
[0033] The second backbone network module receives the low-resolution pedestrian image and performs feature extraction to obtain the second pedestrian feature data, which is expressed as follows:
[0034] z S = θ S (L LR )
[0035] Each second sub-task module includes a gradient rotation layer, a second spatial attention layer, a second attribute classification layer, and a second result layer. The second spatial attention layer receives the second pedestrian feature data, performs pooling and convolution operations on the second pedestrian feature data to obtain a second spatial attention map, and analyzes the second spatial attention map in combination with the gradient rotation matrix of the gradient rotation layer and the second pedestrian feature data to obtain a second attention feature map. The second spatial attention map and the second attention feature map are expressed as follows:
[0036]
[0037] The second attribute classification layer receives the second attention feature map for prediction to obtain a second prediction vector, which is expressed as follows:
[0038]
[0039] The second result layer concatenates and combines the second prediction vectors corresponding to all pedestrian attributes to obtain a second prediction result;
[0040] Based on the second prediction result and the true attribute annotation of the low-resolution pedestrian image set, an attribute classification loss function is constructed. Based on the low-resolution pedestrian image set and combined with the attribute classification loss function, the second pedestrian attribute pre-training model is trained and tuned to obtain a second pedestrian attribute recognition model. The attribute classification loss function is expressed as follows:
[0041]
[0042] Among them, L cls represents the attribute classification loss function, BCE represents the binary cross-entropy loss function, represents the second prediction result, y represents the corresponding true attribute annotation, f i S represents the i-th second attribute classification layer, represents the i-th second attention feature map, represents the i-th second prediction vector, y i represents the corresponding i-th true attribute annotation, z S represents the second pedestrian feature data, θ S represents the network parameters of the second backbone network module, L LR represents the low-resolution pedestrian image, The i-th second spatial attention map, Conv 1×1 represents a 1×1 convolutional layer, Pool Max represents the global max pooling layer, Pool Avg represents the global average pooling layer, ⊕ represents the connection operation between feature maps, represents the i-th second spatial attention layer, ⊙ represents the element-wise dot product, represents the i-th second pedestrian feature data, R i represents the i-th gradient rotation matrix, represents the i-th second prediction vector, f i S represents the i-th second attribute classification layer.
[0043] As an implementable manner, the distillation of the knowledge of the first pedestrian attribute recognition model at multiple levels and throughout the process includes the following steps:
[0044] Construct a feature map distillation loss function to make the first pedestrian feature data in the first pedestrian attribute recognition model close to the second pedestrian feature data in the second pedestrian attribute recognition model. The feature map distillation loss function is expressed as follows:
[0045]
[0046] Construct a spatial attention distillation loss function to make the first spatial attention map in the first pedestrian attribute recognition model close to the second spatial attention map in the second pedestrian attribute recognition model. The spatial attention distillation loss function is expressed as follows:
[0047]
[0048] Construct a prediction result distillation loss function to make the first prediction vector and the second prediction vector in the first pedestrian attribute recognition model close to each other. The prediction result distillation loss function is expressed as follows:
[0049]
[0050] Among them, L perce represents the feature map distillation loss function, H1 represents the height of the first pedestrian feature data or the second pedestrian feature data, W1 represents the width of the first pedestrian feature data or the second pedestrian feature data, and z T (h1, w1) represents the feature value at the position (h1, w1) in the first pedestrian feature data, and z S (h1, w1) represents the feature value at the position (h1, w1) in the second pedestrian feature data. (h1, w1) represents the position in the first pedestrian feature data or the second pedestrian feature data, and L spatial represents the spatial attention distillation loss function, represents the feature value at the position (h2, w2) in the i-th first spatial attention map, represents the feature value at the position (h2, w2) in the i-th second spatial attention map, N represents the number of the first sub-task modules or the number of the second sub-task modules, H2 represents the height of the first spatial attention map or the second spatial attention map, W2 represents the width of the first spatial attention map or the second spatial attention map, and L logits represents the prediction result distillation loss function, represents the first prediction vector, represents the second prediction vector, and i represents the i-th first sub-task module or the i-th second sub-task module.
[0051] As an implementable manner, when training the second pedestrian attribute recognition model, it also includes a process of adjusting parameters through a gradient rotation loss function. Specifically:
[0052] Based on the gradient of the pedestrian attribute group, obtain the average gradient of the pedestrian attribute group, which is expressed as follows:
[0053]
[0054] Perform gradient rotation on the average gradient of the pedestrian attribute group through a learnable gradient rotation matrix, and obtain a gradient rotation loss function based on the gradient of the pedestrian attribute group after gradient rotation, and adjust the parameters based on the gradient rotation loss function. The gradient rotation loss function is expressed as follows:
[0055]
[0056] where g i represents the gradient of the i-th pedestrian attribute group, g mean represents the average gradient of the pedestrian attribute group, N represents the number of pedestrian attribute groups, L roto represents the gradient rotation loss function, Rotate represents the gradient rotation operation, and R i represents the gradient rotation matrix.
[0057] As an implementable manner, the training process of the pedestrian attribute recognition model further includes the following steps:
[0058] Construct an attribute classification loss function through the second prediction result and the corresponding true attribute annotation, which is expressed as follows:
[0059]
[0060] Based on the attribute classification loss function, the feature map distillation loss function, the spatial attention distillation loss function, and the prediction result distillation loss function, construct an attribute recognition loss function, which is expressed as follows:
[0061] L total = L cls + α1L perce + α2L spatial + α3L logits
[0062] Train the second pedestrian attribute recognition model through the attribute recognition loss function and the gradient rotation loss function, and based on the low-resolution pedestrian image set and the pedestrian attribute group, to obtain the pedestrian attribute recognition model;
[0063] where L cls represents the attribute classification loss function, BCE represents the binary cross-entropy loss function, represents the second prediction result, y represents the corresponding true attribute annotation, and f i S represents the i-th second attribute classification layer, denotes the i-th second attention feature map, denotes the i-th second prediction vector, y i denotes the corresponding i-th true attribute annotation, L total denotes the attribute recognition loss function, L perce denotes the feature map distillation loss function, L spatial denotes the spatial attention distillation loss function, L logits denotes the prediction result distillation loss function, where α1, α2, α3 denote weight coefficients and N denotes the number of pedestrian attribute groups.
[0064] A low-resolution pedestrian attribute recognition system based on knowledge distillation, comprising an image preprocessing module, an attribute grouping module, a first model construction module, a second model construction module, a model training module, and a model analysis module;
[0065] The image preprocessing module obtains the original pedestrian image set, downsamples the original pedestrian images to obtain the downsampled pedestrian image set, and preprocesses the original pedestrian image set and the downsampled pedestrian image set respectively to obtain the pedestrian image set and the low-resolution pedestrian image set;
[0066] The attribute grouping module respectively groups the pedestrian attributes in the pedestrian image set and the low-resolution pedestrian image set to form pedestrian attribute groups;
[0067] The first model construction module constructs a first pedestrian attribute pre-training model, including a first backbone network module and a number of first sub-task modules, and the number of first sub-task modules is equal to the number of pedestrian attribute groups. The first backbone network module receives the pedestrian image set and performs feature extraction to obtain the first pedestrian feature data. Each first sub-task module performs attention feature extraction on the first pedestrian feature data and performs attribute recognition for the corresponding pedestrian attribute group to obtain the first prediction result. The first pedestrian attribute pre-training model is trained based on the pedestrian image set to obtain the first pedestrian attribute recognition model;
[0068] The second model construction module constructs a second pedestrian attribute pre-training model, including a second backbone network module and a number of second sub-task modules, and the number of second sub-task modules is the same as the number of pedestrian attribute groups. The second backbone network module receives the low-resolution pedestrian image set and performs feature extraction to obtain the second pedestrian feature data. Each second sub-task module performs gradient optimization on the second pedestrian feature data, performs attention feature extraction, and performs attribute recognition for the corresponding pedestrian attribute group to obtain the second prediction result. The second pedestrian attribute pre-training model is trained based on the low-resolution pedestrian image set to obtain the second pedestrian attribute recognition model;
[0069] The model training module distills the knowledge of the first pedestrian attribute recognition model at multiple levels and throughout the process and applies it to the second pedestrian attribute recognition model to train the second pedestrian attribute recognition model, thereby obtaining a pedestrian attribute recognition model;
[0070] The model analysis module analyzes the pedestrian image to be recognized through the pedestrian attribute recognition model to obtain a pedestrian attribute recognition result.
[0071] A low-resolution pedestrian attribute recognition device based on knowledge distillation, comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the following method is implemented:
[0072] Obtain an original pedestrian image set, downsample the original pedestrian images to obtain a downsampled pedestrian image set, and preprocess the original pedestrian image set and the downsampled pedestrian image set respectively to obtain a pedestrian image set and a low-resolution pedestrian image set;
[0073] Group the pedestrian attributes in the pedestrian image set and the low-resolution pedestrian image set respectively to form pedestrian attribute groups;
[0074] Construct a first pedestrian attribute pre-training model, including a first backbone network module and several first sub-task modules, and the number of first sub-task modules is equal to the number of pedestrian attribute groups. The first backbone network module receives the pedestrian image set and performs feature extraction to obtain first pedestrian feature data. Each first sub-task module performs attention feature extraction on the first pedestrian feature data and performs attribute recognition for the corresponding pedestrian attribute group to obtain a first prediction result. Based on the pedestrian image set, train the first pedestrian attribute pre-training model to obtain a first pedestrian attribute recognition model;
[0075] Construct a second pedestrian attribute pre-training model, including a second backbone network module and several second sub-task modules, and the number of second sub-task modules is the same as the number of pedestrian attribute groups. The second backbone network module receives the low-resolution pedestrian image set and performs feature extraction to obtain second pedestrian feature data. Each second sub-task module performs gradient optimization on the second pedestrian feature data, performs attention feature extraction, and performs attribute recognition for the corresponding pedestrian attribute group to obtain a second prediction result. Based on the low-resolution pedestrian image set, train the second pedestrian attribute pre-training model to obtain a second pedestrian attribute recognition model;
[0076] Distill the knowledge of the first pedestrian attribute recognition model at multiple levels and throughout the process and apply it to the second pedestrian attribute recognition model to train the second pedestrian attribute recognition model, thereby obtaining a pedestrian attribute recognition model;
[0077] Analyze the pedestrian image to be recognized through the pedestrian attribute recognition model to obtain a pedestrian attribute recognition result.
[0078] Due to the adoption of the above technical solutions, the present invention has remarkable technical effects:
[0079] The present invention obtains the first pedestrian attribute recognition model and the second pedestrian attribute recognition model through model construction and model training, and improves the knowledge distillation efficiency between models through a multi-level and whole-process distillation method. By establishing a gradient rotation loss function, the conflict between different pedestrian attribute recognitions is reduced. Without increasing the network parameters, the pedestrian attribute recognition performance of low-resolution pedestrian images is ensured, which is applicable to pedestrian attribute recognition in low-resolution scenarios such as monitoring. Brief Description of the Drawings
[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0081] Figure 1 is a schematic flowchart of the method of the present invention;
[0082] Figure 2 is a schematic diagram of the overall system of the present invention;
[0083] Figure 3 is a schematic diagram of the spatial attention layer structure of the present invention;
[0084] Figure 4 is a schematic diagram of the network structure and the knowledge distillation process of the present invention. Detailed Embodiments
[0085] The following will further elaborate on the present invention in conjunction with embodiments. The following embodiments are explanations of the present invention, and the present invention is not limited to the following embodiments.
[0086] Embodiment 1:
[0087] A method for low-resolution pedestrian attribute recognition based on knowledge distillation, as Figure 1 shown, includes the following steps:
[0088] S100. Obtain the original pedestrian image set, downsample the original pedestrian images to obtain the downsampled pedestrian image set, and preprocess the original pedestrian image set and the downsampled pedestrian image set respectively to obtain the pedestrian image set and the low-resolution pedestrian image set;
[0089] S200. Group the pedestrian attributes in the pedestrian image set and the low-resolution pedestrian image set respectively to form pedestrian attribute groups;
[0090] S300. Construct a first pedestrian attribute pre-training model, which includes a first backbone network module and several first sub-task modules, and the number of the first sub-task modules is equal to the number of pedestrian attribute groups. The first backbone network module receives a pedestrian image set and performs feature extraction to obtain first pedestrian feature data. Each first sub-task module performs attention feature extraction on the first pedestrian feature data and conducts attribute recognition for the corresponding pedestrian attribute group to obtain a first prediction result. Based on the pedestrian image set, the first pedestrian attribute pre-training model is trained to obtain a first pedestrian attribute recognition model;
[0091] S400. Construct a second pedestrian attribute pre-training model, which includes a second backbone network module and several second sub-task modules, and the number of the second sub-task modules is the same as the number of pedestrian attribute groups. The second backbone network module receives a low-resolution pedestrian image set and performs feature extraction to obtain second pedestrian feature data. Each second sub-task module performs gradient optimization on the second pedestrian feature data, then conducts attention feature extraction, and conducts attribute recognition for the corresponding pedestrian attribute group to obtain a second prediction result. Based on the low-resolution pedestrian image set, the second pedestrian attribute pre-training model is trained to obtain a second pedestrian attribute recognition model;
[0092] S500. Distill the knowledge of the first pedestrian attribute recognition model at multiple levels and throughout the process and apply it to the second pedestrian attribute recognition model, and train the second pedestrian attribute recognition model to obtain a pedestrian attribute recognition model;
[0093] S600. Analyze the pedestrian image to be recognized through the pedestrian attribute recognition model to obtain a pedestrian attribute recognition result.
[0094] Through data acquisition and data preprocessing, the present invention obtains a pedestrian image set and a low-resolution pedestrian image set, constructs a first pedestrian attribute pre-training model and a second pedestrian attribute pre-training model, and trains them respectively based on the pedestrian image set and the low-resolution pedestrian image set to obtain a first pedestrian attribute recognition model and a second pedestrian attribute recognition model. By distilling and applying the knowledge of the first pedestrian attribute recognition model and training the second pedestrian attribute recognition model, a pedestrian attribute model is obtained. Furthermore, the pedestrian attribute model is used to analyze the pedestrian image to be recognized to obtain a pedestrian attribute recognition result. Through knowledge distillation and model training, the present invention ensures the pedestrian attribute recognition performance of low-resolution pedestrian images and has a good application effect in low-resolution pedestrian image scenarios such as monitoring.
[0095] In this embodiment, the method for downsampling the original pedestrian image set is interpolation with a magnification factor of 1. The original pedestrian image set and the downsampled pedestrian image set are preprocessed, including horizontal flipping, cropping adjustment, tensor conversion, and normalization of the images. In this embodiment, the set image cropping size is 144 pixels in width and 288 pixels in height for the image, and the image is subjected to tensor conversion and normalization operations. Among them, the mean and variance in the normalization process are [0.485, 0.456, 0.406] and [0.229, 0.224, 0.225], respectively. After the preprocessing steps, a pedestrian image set and a low-resolution pedestrian image set are obtained.
[0096] Predefine K pedestrian attributes A = {a1, a2,... a K}, a i ∈ {1, 0} K , and group the pedestrian attributes of the pedestrian image set and the low-resolution pedestrian image set. In this embodiment, specifically: First, through the spatial position, the similar partial pedestrian attributes are divided into the same pedestrian attribute group. For example, "short sleeves" and "long sleeves" are divided into the "upper body" pedestrian attribute group; Second, for pedestrian attributes with a higher degree of abstraction or that are not easily divided directly by spatial position, they are grouped through semantic logic. For example, the pedestrian attributes "gender" and "age" are divided into the "gender + age" pedestrian attribute group. Through the above pedestrian attribute grouping, an initial pedestrian attribute group is obtained. For the remaining pedestrian attributes, according to the gradient distance from the already determined initial pedestrian attribute group, pedestrian attributes with similar gradient distances are considered to be closer in the gradient dimension and are divided into the initial pedestrian attribute group with the closest gradient distance. For attributes with a large gradient distance, they are divided into different pedestrian attribute groups. For example, through gradient distance analysis, the attribute "holding" is divided into the "bag" pedestrian attribute group until all pedestrian attributes are grouped, and N pedestrian attribute groups G = {G1, G2,..., G N} are obtained, where the gradient distance is expressed as follows:
[0097] Dist j,k = cos(g j , g k )
[0098] Among them, Dist j,k represents the gradient distance, g j represents the j-th pedestrian attribute, and g k represents the k-th pedestrian attribute.
[0099] Using the attribute grouping method in this embodiment, attribute grouping is performed on three general datasets in the field of pedestrian attribute recognition, namely Market-1501, DukeMTMC-reID, and PA-100K, to obtain pedestrian attribute groups respectively. Specifically, the Market-1501 dataset contains 32,668 images for model training and 13,328 images for model testing, and a total of 27 pedestrian attributes; the DukeMTMC-reID dataset includes 16,522 images for model training and 19,889 images for model testing, and a total of 23 pedestrian attributes; the PA-100K dataset includes 90,000 images for model training and 10,000 images for model testing, and a total of 26 pedestrian attributes. By using the attribute grouping method in this embodiment to group the three datasets, the obtained pedestrian attribute groups are shown in Table 1 below:
[0100] Table 1
[0101]
[0102] Construct a first pedestrian attribute pre-training model, including a first backbone network module and a first sub-task module. The first backbone network module receives a pedestrian image set and performs feature extraction to obtain first pedestrian feature data. Among them, the first pedestrian feature data is expressed as follows:
[0103] z T = θ T (I HR )
[0104] The first sub-task module in the first pedestrian attribute pre-training model corresponds one-to-one with the pedestrian attribute groups, that is, the first pedestrian attribute pre-training model includes N first sub-task modules. The first pedestrian feature data is copied into N copies and input into the first sub-task modules respectively. The first sub-task module includes a first spatial attention layer, a first attribute classification layer, and a first result layer. The first spatial attention layer has a structure as Figure 3As shown, attention feature extraction is performed on the first pedestrian feature data. Among them, the attention feature extraction obtains the first spatial attention map through pooling operations and convolutional operations. The pooling operations and convolutional operations are the core operations of the neural network, which are used to further extract data features, including local features such as edges, textures, and shapes in the image. By combining the pooling operations and convolutional operations, the number of parameters of the model is reduced while obtaining image features. The pooling operations in this embodiment include global max pooling operations and average pooling operations. In this process, the first pedestrian feature data remains H×W unchanged. Through convolutional operations, the first spatial attention map related to the current first sub-task module is obtained. The first spatial attention map is input into the Sigmoid activation function for activation and weighted to the first pedestrian feature data through dot product operations to obtain the first attention feature map. The first attention feature map will activate the regions in the first pedestrian feature data related to the pedestrian attribute group corresponding to the current first sub-task module and suppress irrelevant regions, enabling the model to capture the global features of the image with fewer parameters. The first spatial attention map and the first attention feature map are specifically represented as follows:
[0105]
[0106]
[0107] The first attribute classification layer f in the first sub-task module i T ∈F={f1,f2,...,f N}, performs attribute recognition for the corresponding pedestrian attribute group to obtain the first prediction vector, which is represented as follows:
[0108]
[0109] The first result layer splices and combines the first prediction vectors obtained by all the first sub-task modules to obtain the first prediction result
[0110] Based on the true pedestrian attributes in the training set of the pedestrian image set for annotation, the true attribute annotation is obtained, such as pedestrian attributes like "wearing glasses". By measuring the difference between the first prediction result and the true attribute annotation, and based on the binary cross-entropy loss function, the first attribute loss function is constructed. In addition, the first training loss function can also be obtained through mean square error loss, root mean square error loss, mean absolute error loss, and focal loss, etc. Based on the first training loss function and the pedestrian image set, the network parameters in the first pedestrian attribute pre-training model are updated and trained until the model performance meets the preset performance threshold to obtain the first pedestrian attribute recognition model. Among them, the first attribute loss function is represented as follows:
[0111]
[0112] where θ T represents the network parameters of the first backbone network module, I HR represents the pedestrian image, z T represents the first pedestrian feature data, represents the i-th first spatial attention map, represents the i-th first spatial attention layer, Conv 1×1 represents a 1×1 convolutional layer, Pool Max represents the global max pooling layer, Pool Avg represents the global average pooling layer, ⊕ represents the connection operation between feature maps, represents the i-th first attention feature map, represents the i-th first pedestrian feature data, ⊙ represents the element-wise dot product, represents the i-th first prediction vector, f i T represents the i-th first attribute classification layer, L1 represents the first attribute loss function, y i represents the i-th true attribute annotation, BCE represents the binary cross-entropy loss function, and N represents the number of pedestrian attribute groups.
[0113] Construct a second pedestrian attribute pre-training model. The second pedestrian attribute pre-training model includes a second backbone network module and a second sub-task module. The second backbone network module extracts features from the low-resolution pedestrian image to obtain the second pedestrian feature data, which is expressed as follows:
[0114] z S = θ S (L LR )
[0115] The second sub-task module in the second pedestrian attribute pre-training model corresponds one-to-one with the pedestrian attribute groups. That is, the second pedestrian attribute pre-training model includes N second sub-task modules. The second sub-task module includes a gradient rotation layer, a second spatial attention layer, a second attribute classification layer, and a second result layer. Copy the second pedestrian feature data N times, perform gradient rotation through the gradient rotation layer, and then input it to the second spatial attention layer. The structure of the second spatial attention layer is as Figure 3As shown, pooling and convolution operations are performed on the received second pedestrian feature data. The pooling operations include global max pooling and global average pooling. During this process, the size H×W of the second pedestrian feature data remains unchanged. By performing a convolution operation on the pooled matrix, a second spatial attention map related to the pedestrian attribute group corresponding to the current second sub-task module is obtained. The second spatial attention map is activated through the Sigmoid function and weighted to the second pedestrian feature data through element-wise dot product to obtain a second attention feature map. The second attention feature map activates the regions related to the pedestrian attribute group corresponding to the current second sub-task module in the second pedestrian feature data and suppresses the irrelevant regions, extracting deeper features of the low-resolution pedestrian image. The second spatial attention map and the second attention feature map are expressed as follows:
[0116]
[0117] The second prediction vector is obtained by predicting the second attention feature map through the second attribute classification layer, which is expressed as follows:
[0118]
[0119] The second result layer concatenates and combines the second prediction vectors obtained from all the second sub-task modules to obtain the second prediction result
[0120] Based on the second prediction result and the true attribute annotation of the low-resolution pedestrian image set, in this embodiment, an attribute classification loss function is constructed through the binary cross-entropy loss function, and the network parameters of the second pedestrian attribute pre-training model are trained and updated until the model converges to obtain the second pedestrian attribute recognition model. Among them, the attribute classification loss function is expressed as follows:
[0121]
[0122] Among them, z S represents the second pedestrian feature data, θ S represents the network parameters of the second backbone network module, L LR represents the low-resolution pedestrian image, the i-th second spatial attention map, Conv 1×1 represents a 1×1 convolutional layer, Pool Max represents the global max pooling layer, Pool Avg represents the global average pooling layer, ⊕ represents the connection operation between feature maps, represents the i-th second spatial attention layer, represents the i-th second attention feature map, ⊙ represents the element-wise dot product, represents the i-th second pedestrian feature data, R idenotes the $i$-th gradient rotation matrix, denotes the $i$-th second prediction vector, $f$ i S denotes the $i$-th second attribute classification layer, $L$ cls denotes the attribute classification loss function, and BCE denotes the binary cross-entropy loss function, denotes the second prediction result, $y$ denotes the corresponding true attribute annotation, denotes the $i$-th second prediction vector, $N$ denotes the number of pedestrian attribute groups, $y$ i denotes the corresponding $i$-th true attribute annotation.
[0123] Knowledge distillation, as a technique for transferring the knowledge of a large and complex model to a small and lightweight model, is used to reduce the computational amount and storage requirements of the model, and retain the performance of the original model as much as possible. Through the knowledge transfer between models, the model can significantly reduce the number of parameters and computational complexity while maintaining high performance. In this embodiment, the knowledge of the first pedestrian attribute recognition model is distilled and applied to the second pedestrian attribute recognition model, so that the second pedestrian attribute recognition model learns from the first pedestrian attribute recognition model at multiple levels and throughout the process. Among them, the structural diagrams of the first pedestrian attribute recognition model, the second pedestrian attribute recognition model, and the process of knowledge distillation are as Figure 4 shown. In this embodiment, the multi-level and full-process knowledge distillation includes: using the feature map distillation loss function to make the second backbone network module in the second pedestrian attribute recognition model imitate the first backbone network module in the first pedestrian attribute recognition model; using the spatial attention distillation loss function to make the second spatial attention map in the second pedestrian attribute recognition model imitate the first spatial attention map in the first pedestrian attribute recognition model; using the prediction result distillation loss function to make the second prediction vector in the second pedestrian attribute recognition model imitate the first prediction vector in the first pedestrian attribute recognition model. It is specifically implemented through the following steps:
[0124] Reduce the dimensions of the first pedestrian feature data and the second pedestrian feature data from $H\times W\times C$ to $H\times W$, and construct a feature map distillation loss function based on the pixel-level loss between the feature data, so that the first pedestrian feature data in the first pedestrian attribute recognition model is close to the second pedestrian feature data in the second pedestrian attribute recognition model, that is, the second pedestrian feature data imitates the first pedestrian feature data, and the distance between the pedestrian image and the low-resolution pedestrian image in the feature space is shortened. Among them, the feature map distillation loss function is expressed as follows:
[0125]
[0126] To impose spatial consistency constraints between the first pedestrian attribute recognition model and the second pedestrian attribute recognition model, and to help the second pedestrian attribute recognition model acquire better spatial information capture ability during the mid-stage of learning, a spatial attention distillation loss function is constructed through the pixel-level loss between the first spatial attention map and the second spatial attention map, so that the first spatial attention map in the first pedestrian attribute recognition model is close to the second spatial attention map in the second pedestrian attribute recognition model, that is, the second spatial attention map mimics the first spatial attention map. Among them, the spatial attention distillation loss function is expressed as follows:
[0127]
[0128] To ensure that the prediction results of the first pedestrian attribute recognition model and the second pedestrian attribute recognition model are consistent in the later stage, a prediction result distillation loss function is constructed through the mean square error between the first prediction vector and the second prediction vector, so that the first prediction vector and the second prediction vector in the first pedestrian attribute recognition model are close to each other, that is, the second prediction vector mimics the first prediction vector. Among them, the prediction result distillation loss function is expressed as follows:
[0129]
[0130] Among them, L perce represents the feature map distillation loss function, H1 represents the height of the first pedestrian feature data or the second pedestrian feature data, W1 represents the width of the first pedestrian feature data or the second pedestrian feature data, and z T (h1, w1) represents the feature value at the position (h1, w1) in the first pedestrian feature data, and z S (h1, w1) represents the feature value at the position (h1, w1) in the second pedestrian feature data. (h1, w1) represents the position in the first pedestrian feature data or the second pedestrian feature data. L spatial represents the spatial attention distillation loss function, represents the feature value at the position (h2, w2) in the i-th first spatial attention map, represents the feature value at the position (h2, w2) in the i-th second spatial attention map. N represents the number of the first sub-task modules or the number of the second sub-task modules. H2 represents the height of the first spatial attention map or the second spatial attention map, and W2 represents the width of the first spatial attention map or the second spatial attention map. L logits represents the prediction result distillation loss function, represents the first prediction vector, represents the second prediction vector, and i represents the i-th first sub-task module or the i-th second sub-task module.
[0131] In this embodiment, to resolve the gradient conflict between subtasks, the lowest points of each gradient descent are unified. The average gradient of the pedestrian attribute group is obtained through the gradient of the pedestrian attribute group, and the average gradient of the pedestrian attribute group is rotated using a gradient rotation matrix. A gradient rotation loss function is constructed through the gradient of the pedestrian attribute group after gradient rotation. Among them, the average gradient of the pedestrian attribute group and the gradient rotation loss function are expressed as follows:
[0132]
[0133] Among them, g i represents the gradient of the i-th pedestrian attribute group, g mean represents the average gradient of the pedestrian attribute group, N represents the number of pedestrian attribute groups, L roto represents the gradient rotation loss function, Rotate represents the gradient rotation operation, and R i represents the gradient rotation matrix.
[0134] To improve the performance of the pedestrian attribute recognition model, the training and optimization of the second pedestrian attribute recognition model in this embodiment include two parts of optimization objectives. One part constructs an attribute recognition loss function based on the attribute classification loss function, feature map distillation loss function, spatial attention distillation loss function, and prediction result distillation loss function to achieve multi-level and whole-process knowledge distillation and attribute classification. Among them, the attribute recognition loss function is expressed as follows:
[0135] L total = L cls + α1L perce + α2L spatial + α3L logits
[0136] The other part optimizes the gradient rotation module of different pedestrian attribute groups based on the gradient rotation loss function. Through the low-resolution pedestrian image set and the pedestrian attribute group, combined with the gradient rotation loss function, the second pedestrian attribute recognition model is trained through iterative update until the loss function is minimized to obtain the pedestrian attribute recognition model. The two parts of the model training objectives in this embodiment are expressed as follows:
[0137]
[0138] Among them, L cls represents the attribute classification loss function, BCE represents the binary cross-entropy loss function, represents the second prediction result, y represents the corresponding true attribute annotation, f i S represents the i-th second attribute classification layer, represents the i-th second attention feature map, denotes the i-th second prediction vector, y i denotes the corresponding i-th true attribute annotation, L total denotes the attribute recognition loss function, L perce denotes the feature map distillation loss function, L spatial denotes the spatial attention distillation loss function, L logits denotes the prediction result distillation loss function, α1, α2, α3 denote weight coefficients, θ denotes the network parameters in the second pedestrian attribute recognition model, and R denotes the gradient rotation matrix.
[0139] In this embodiment, the model is trained on an NVIDIA GeForce GTX 1080Ti graphics card. The data batch size during training is set to 8, and the SGD training optimizer is used. Two learning rates are set during training. One is set to 0.0085 for the first backbone network module and the second backbone network module to ensure that it reaches model convergence faster; the other is set to 0.0008 for the gradient rotation layer to finely adjust the values of the gradient rotation matrix. In the SGD training optimizer, the momentum is 0.9, the weight decay is set to 1e-3, and Nesterov acceleration is enabled. To better control and adjust the learning rate, a learning rate decay measurement is adopted. In this embodiment, the StepLR scheduler is used, and the learning rate is reduced to 0.1 times the current learning rate every 30 iterations, which helps the model converge better during training and improves the performance of pedestrian attribute recognition.
[0140] Through the above settings of the model training process, a pedestrian attribute recognition model is obtained through model training. In this embodiment, the average attribute accuracy (Accuracy%) and the average instance accuracy (Precision%) are used as the evaluation criteria for the pedestrian attribute recognition results. The larger (Accuracy%) and (Precision%) are, the higher the recognition accuracy. On the three datasets of Market-1501, DukeMTMC-reID, and PA-100K, the pedestrian attribute recognition method of this embodiment is verified. At the same time, it is compared with three methods of SRMAR, ALM, and GSR-MAR. The (Accuracy%) and (Precision%) test results of the method of this embodiment and the three comparison methods are shown in Table 2 below:
[0141] Table 2
[0142]
[0143] As can be seen from Table 2, the (Accuracy%) and (Precision%) of the method in this embodiment in the three datasets are higher than those of the three comparison methods, indicating that the performance of the pedestrian attribute recognition result of this method has been greatly improved.
[0144] Example 2:
[0145] A low-resolution pedestrian attribute recognition system based on knowledge distillation, as Figure 2 shown, includes an image preprocessing module 100, an attribute grouping module 200, a first model construction module 300, a second model construction module 400, a model training module 500, and a model analysis module 600;
[0146] The image preprocessing module 100 obtains the original pedestrian image set, downsamples the original pedestrian images to obtain a downsampled pedestrian image set, and preprocesses the original pedestrian image set and the downsampled pedestrian image set respectively to obtain a pedestrian image set and a low-resolution pedestrian image set;
[0147] The attribute grouping module 200 respectively groups the pedestrian attributes in the pedestrian image set and the low-resolution pedestrian image set to form pedestrian attribute groups;
[0148] The first model construction module 300 constructs a first pedestrian attribute pre-training model, including a first backbone network module and a number of first sub-task modules, and the number of first sub-task modules is equal to the number of pedestrian attribute groups. The first backbone network module receives the pedestrian image set and performs feature extraction to obtain first pedestrian feature data. Each first sub-task module performs attention feature extraction on the first pedestrian feature data and performs attribute recognition for the corresponding pedestrian attribute group to obtain a first prediction result. The first pedestrian attribute pre-training model is trained based on the pedestrian image set to obtain a first pedestrian attribute recognition model;
[0149] The second model construction module 400 constructs a second pedestrian attribute pre-training model, including a second backbone network module and a number of second sub-task modules, and the number of second sub-task modules is the same as the number of pedestrian attribute groups. The second backbone network module receives the low-resolution pedestrian image set and performs feature extraction to obtain second pedestrian feature data. Each second sub-task module performs gradient optimization on the second pedestrian feature data, performs attention feature extraction, and performs attribute recognition for the corresponding pedestrian attribute group to obtain a second prediction result. The second pedestrian attribute pre-training model is trained based on the low-resolution pedestrian image set to obtain a second pedestrian attribute recognition model;
[0150] The model training module 500 distills the knowledge of the first pedestrian attribute recognition model at multiple levels and throughout the process and applies it to the second pedestrian attribute recognition model to train the second pedestrian attribute recognition model to obtain a pedestrian attribute recognition model;
[0151] The model analysis module 600 analyzes the to-be-recognized pedestrian image through the pedestrian attribute recognition model to obtain a pedestrian attribute recognition result.
[0152] All changes and modifications made without departing from the spirit and scope of the present invention, and all equivalent technical solutions also fall within the scope of the present invention.
[0153] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0154] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0155] The present invention is described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0158] It should be noted that:
[0159] The "one embodiment" or "embodiments" mentioned in the specification means that the specific features, structures or characteristics described in connection with the embodiments are included in at least one embodiment of the present invention. Therefore, the phrases "one embodiment" or "embodiments" that appear throughout the specification do not necessarily all refer to the same embodiment.
[0160] In addition, it should be noted that for the specific embodiments described in this specification, the shapes of their components, the names taken, etc. may be different. Any equivalent or simple changes made to the structure, features and principles according to the inventive concept of the present invention are included in the protection scope of the present invention. Those skilled in the art to which the present invention pertains can make various modifications, supplements or use similar ways of substitution to the specific embodiments described, as long as they do not deviate from the structure of the present invention or exceed the scope defined by the claims, they should fall within the protection scope of the present invention.
Claims
1. A low-resolution pedestrian attribute recognition method based on knowledge distillation, characterized in that Including the following steps: Obtain the original pedestrian image set, downsample the original pedestrian images to obtain the downsampled pedestrian image set, and preprocess the original pedestrian image set and the downsampled pedestrian image set respectively to obtain the pedestrian image set and the low-resolution pedestrian image set; Group the pedestrian attributes in the pedestrian image set and the low-resolution pedestrian image set respectively to form pedestrian attribute groups; Construct a first pedestrian attribute pre-training model, including a first backbone network module and a number of first sub-task modules, and the number of first sub-task modules is equal to the number of pedestrian attribute groups. The first backbone network module receives the pedestrian image set and performs feature extraction to obtain first pedestrian feature data. Each first sub-task module performs attention feature extraction on the first pedestrian feature data and performs attribute recognition for the corresponding pedestrian attribute group to obtain a first prediction result. Train the first pedestrian attribute pre-training model based on the pedestrian image set to obtain a first pedestrian attribute recognition model; Construct a second pedestrian attribute pre-training model, including a second backbone network module and a number of second sub-task modules, and the number of second sub-task modules is the same as the number of pedestrian attribute groups. The second backbone network module receives the low-resolution pedestrian image set and performs feature extraction to obtain second pedestrian feature data. Each second sub-task module performs gradient optimization on the second pedestrian feature data, performs attention feature extraction, and performs attribute recognition for the corresponding pedestrian attribute group to obtain a second prediction result. Train the second pedestrian attribute pre-training model based on the low-resolution pedestrian image set to obtain a second pedestrian attribute recognition model; Distill the knowledge of the first pedestrian attribute recognition model at multiple levels and throughout the process and apply it to the second pedestrian attribute recognition model, and train the second pedestrian attribute recognition model to obtain a pedestrian attribute recognition model; Analyze the pedestrian image to be recognized through the pedestrian attribute recognition model to obtain the pedestrian attribute recognition result.
2. The low-resolution pedestrian attribute recognition method based on knowledge distillation according to claim 1, wherein The preprocessing includes horizontal flipping, cropping adjustment, tensor conversion, and normalization processing.
3. The low-resolution pedestrian attribute recognition method based on knowledge distillation according to claim 1, characterized in that The pedestrian attribute group is obtained through the following steps: Obtain the spatial position and semantic logic of the pedestrian attributes, and divide the pedestrian attributes based on the spatial position and semantic logic to obtain the initial pedestrian attribute group; Obtain the gradient distance between the remaining pedestrian attributes and the initial pedestrian attribute group, which is expressed as follows: Dist j,k = cos(g j , g k ) Among them, Dist j,k represents the gradient distance, and g j represents the j-th pedestrian attribute, and g k represents the k-th pedestrian attribute; Set a distance threshold, and divide the pedestrian attributes whose gradient distance meets the distance threshold into the corresponding initial pedestrian attribute group until all pedestrian attributes are grouped to obtain the pedestrian attribute group.
4. The method for low-resolution pedestrian attribute recognition based on knowledge distillation according to claim 1, wherein, The training process of the first pedestrian attribute pre-training model is as follows: The first backbone network module receives the pedestrian image set and performs feature extraction on the pedestrian images to obtain first pedestrian feature data, which is expressed as follows: z T = θ T (I HR ) Each first sub-task module includes a first spatial attention layer, a first attribute classification layer, and a first result layer. The first spatial attention layer receives the first pedestrian feature data, performs pooling and convolution operations on the first pedestrian feature data to obtain a first spatial attention map, combines the first pedestrian feature data for element-wise dot product to obtain a first attention feature map. The first spatial attention map and the first attention feature map are expressed as follows: The first attribute classification layer receives the first attention feature map and makes predictions to obtain a first prediction vector, which is expressed as follows: The first result layer concatenates and combines the first prediction vectors corresponding to all pedestrian attributes to obtain a first prediction result; Based on the first prediction result and the true attribute annotations of the pedestrian image set, a first attribute loss function is constructed. Based on the pedestrian image set and in combination with the first attribute loss function, the first pedestrian attribute pre-training model is trained and parameter-tuned to obtain a first pedestrian attribute recognition model. The first attribute loss function is expressed as follows: Among them, θ T represents the network parameters of the first backbone network module, I HR represents the pedestrian image, z T represents the first pedestrian feature data, represents the i-th first spatial attention map, represents the i-th first spatial attention layer, Conv 1×1 represents a 1×1 convolutional layer, Pool Max represents the global max pooling layer, Pool Avg represents the global average pooling layer, represents the connection operation between feature maps, represents the i-th first attention feature map, represents the i-th first pedestrian feature data, ⊙ represents the element-wise dot product, represents the i-th first prediction vector, f i T represents the i-th first attribute classification layer, L1 represents the first attribute loss function, y i represents the i-th true attribute annotation, BCE represents the binary cross-entropy loss function, and N represents the number of pedestrian attribute groups.
5. The method for low-resolution pedestrian attribute recognition based on knowledge distillation according to claim 1, characterized in that The training process of the second pedestrian attribute pre-training model is as follows: The second backbone network module receives a low-resolution pedestrian image and performs feature extraction to obtain second pedestrian feature data, which is expressed as follows: z S = θ S (L LR ) Each second sub-task module includes a gradient rotation layer, a second spatial attention layer, a second attribute classification layer, and a second result layer. The second spatial attention layer receives the second pedestrian feature data, performs pooling and convolution operations on the second pedestrian feature data to obtain a second spatial attention map, and analyzes it in combination with the gradient rotation matrix of the gradient rotation layer and the second pedestrian feature data to obtain a second attention feature map. The second spatial attention map and the second attention feature map are expressed as follows: The second attribute classification layer receives the second attention feature map and makes predictions to obtain a second prediction vector, which is expressed as follows: The second result layer concatenates and combines the second prediction vectors corresponding to all pedestrian attributes to obtain a second prediction result; Based on the second prediction result and the true attribute annotations of the low-resolution pedestrian image set, an attribute classification loss function is constructed. Based on the low-resolution pedestrian image set and in combination with the attribute classification loss function, the second pedestrian attribute pre-training model is trained and parameter-tuned to obtain a second pedestrian attribute recognition model. The attribute classification loss function is expressed as follows: Among them, L cls represents the attribute classification loss function, and BCE represents the binary cross - entropy loss function. represents the second prediction result, y represents the corresponding true attribute annotation, and f i S represents the i - th second attribute classification layer. represents the i - th second attention feature map. represents the i - th second prediction vector, and y i represents the corresponding i - th true attribute annotation, and z S represents the second pedestrian feature data, and θ S represents the network parameters of the second backbone network module, and L LR represents the low - resolution pedestrian image. The i - th second spatial attention map, Conv 1×1 represents a 1×1 convolutional layer, and Pool Max represents the global max - pooling layer, and Pool Avg represents the global average - pooling layer. represents the connection operation between feature maps. represents the i - th second spatial attention layer, and ⊙ represents the element - wise dot product. represents the i - th second pedestrian feature data, and R i represents the i - th gradient rotation matrix. represents the i - th second prediction vector, and f i S represents the i - th second attribute classification layer.
6. The method for low-resolution pedestrian attribute recognition based on knowledge distillation according to claim 1, characterized in that The multi-level and full-process distillation of the knowledge of the first pedestrian attribute recognition model includes the following steps: Construct a feature map distillation loss function to make the first pedestrian feature data in the first pedestrian attribute recognition model close to the second pedestrian feature data in the second pedestrian attribute recognition model. Among them, the feature map distillation loss function is expressed as follows: Construct a spatial attention distillation loss function to make the first spatial attention map in the first pedestrian attribute recognition model close to the second spatial attention map in the second pedestrian attribute recognition model. Among them, the spatial attention distillation loss function is expressed as follows: Construct a prediction result distillation loss function to make the first prediction vector and the second prediction vector in the first pedestrian attribute recognition model close to each other. Among them, the prediction result distillation loss function is expressed as follows: Among them, L perce represents the feature map distillation loss function, H1 represents the height of the first pedestrian feature data or the second pedestrian feature data, W1 represents the width of the first pedestrian feature data or the second pedestrian feature data, z T (h1, w1) represents the feature value at the position (h1, w1) in the first pedestrian feature data, z S (h1, w1) represents the feature value at the position (h1, w1) in the second pedestrian feature data, (h1, w1) represents the position in the first pedestrian feature data or the second pedestrian feature data, L spatial represents the spatial attention distillation loss function, represents the feature value at the position (h2, w2) in the i-th first spatial attention map, M i S (h2, w2) represents the feature value at the position (h2, w2) in the i-th second spatial attention map, N represents the number of the first sub-task modules or the number of the second sub-task modules, H2 represents the height of the first spatial attention map or the second spatial attention map, W2 represents the width of the first spatial attention map or the second spatial attention map, L logits represents the prediction result distillation loss function, represents the first prediction vector, represents the second prediction vector, and i represents the i-th first sub-task module or the i-th second sub-task module.
7. The method for low-resolution pedestrian attribute recognition based on knowledge distillation according to claim 1, wherein When training the second pedestrian attribute recognition model, it also includes a process of parameter tuning through a gradient rotation loss function. Specifically: Based on the gradients of the pedestrian attribute groups, the average gradient of the pedestrian attribute groups is obtained, which is expressed as follows: Perform gradient rotation on the average gradient of the pedestrian attribute group through a learnable gradient rotation matrix, and obtain a gradient rotation loss function based on the gradients of the pedestrian attribute group after gradient rotation, and perform parameter tuning based on the gradient rotation loss function. The gradient rotation loss function is expressed as follows: Among them, g i represents the gradient of the i-th pedestrian attribute group, and g mean represents the average gradient of the pedestrian attribute group. N represents the number of pedestrian attribute groups, and L roto represents the gradient rotation loss function, Rotate represents the gradient rotation operation, and R i represents the gradient rotation matrix.
8. The method for low-resolution pedestrian attribute recognition based on knowledge distillation according to claim 1, wherein The training process of the pedestrian attribute recognition model further includes the following steps: Construct an attribute classification loss function through the second prediction result and the corresponding true attribute annotation, which is expressed as follows: Based on the attribute classification loss function, the feature map distillation loss function, the spatial attention distillation loss function, and the prediction result distillation loss function, construct an attribute recognition loss function, which is expressed as follows: L total = L cls + α1L perce + α2L spatial + α3L logits Train the second pedestrian attribute recognition model through the attribute recognition loss function and the gradient rotation loss function, and based on the low-resolution pedestrian image set and the pedestrian attribute group, to obtain the pedestrian attribute recognition model; Among them, L cls represents the attribute classification loss function, and BCE represents the binary cross-entropy loss function. represents the second prediction result, y represents the corresponding true attribute annotation, and f i S represents the i-th second attribute classification layer. represents the i-th second attention feature map. represents the i-th second prediction vector, and y i represents the corresponding i-th true attribute annotation, and L total represents the attribute recognition loss function, and L perce represents the feature map distillation loss function, and L spatial represents the spatial attention distillation loss function, and L logits represents the prediction result distillation loss function. α1, α2, and α3 represent weight coefficients, and N represents the number of pedestrian attribute groups.
9. A low-resolution pedestrian attribute recognition system based on knowledge distillation, characterized in that, It includes an image preprocessing module, an attribute grouping module, a first model construction module, a second model construction module, a model training module, and a model analysis module; The image preprocessing module obtains the original pedestrian image set, downsamples the original pedestrian images to obtain a downsampled pedestrian image set, and preprocesses the original pedestrian image set and the downsampled pedestrian image set respectively to obtain a pedestrian image set and a low-resolution pedestrian image set; The attribute grouping module respectively groups the pedestrian attributes in the pedestrian image set and the low-resolution pedestrian image set to form pedestrian attribute groups; The first model construction module constructs a first pedestrian attribute pre-training model, which includes a first backbone network module and a number of first sub-task modules, and the number of first sub-task modules is equal to the number of pedestrian attribute groups. The first backbone network module receives the pedestrian image set and performs feature extraction to obtain first pedestrian feature data. Each first sub-task module performs attention feature extraction on the first pedestrian feature data and performs attribute recognition for the corresponding pedestrian attribute group to obtain a first prediction result. Train the first pedestrian attribute pre-training model based on the pedestrian image set to obtain a first pedestrian attribute recognition model; The second model construction module constructs a second pedestrian attribute pre-training model, which includes a second backbone network module and a number of second sub-task modules, and the number of second sub-task modules is the same as the number of pedestrian attribute groups. The second backbone network module receives the low-resolution pedestrian image set and performs feature extraction to obtain second pedestrian feature data. Each second sub-task module performs gradient optimization on the second pedestrian feature data, performs attention feature extraction, and performs attribute recognition for the corresponding pedestrian attribute group to obtain a second prediction result. Train the second pedestrian attribute pre-training model based on the low-resolution pedestrian image set to obtain a second pedestrian attribute recognition model; The model training module distills the knowledge of the first pedestrian attribute recognition model at multiple levels and throughout the process and applies it to the second pedestrian attribute recognition model, and trains the second pedestrian attribute recognition model to obtain the pedestrian attribute recognition model; The model analysis module analyzes the pedestrian image to be recognized through the pedestrian attribute recognition model to obtain the pedestrian attribute recognition result.
10. A low-resolution pedestrian attribute recognition device based on knowledge distillation, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.