Multi-scale pedestrian attribute positioning identification method suitable for embedded hardware
By applying a multi-scale pedestrian attribute positioning recognition method on embedded hardware, using the attribute area positioning network and multi-scale bidirectional feature fusion structure, the problems of high power consumption and poor portability of pedestrian attribute recognition algorithm in the prior art are solved, and the pedestrian attribute recognition effect with high real-time and low power consumption is achieved.
Patent Information
- Application Number
- CN202510178404.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
AI Technical Summary
The existing pedestrian attribute recognition algorithm runs in a server environment, and there are problems of high power consumption and poor portability, making it difficult to realize embedded hardware applications with high real-time and low power consumption.
A multi-scale pedestrian attribute positioning recognition method suitable for embedded hardware is proposed. Through a flexible attribute area positioning network and a multi-scale two-way feature fusion structure, the most discriminant area is discovered, and the cross-scale information supplement and optimization is realized through the bidirectional flowing feature information, improving the discrimination ability of attribute features.
It realizes more accurate identification of pedestrian location attributes in complex scenarios, reduces power consumption, improves portability and real-time, and can operate without network conditions on embedded hardware.
Smart Images

Figure CN120014558A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a multi-scale pedestrian attribute positioning and recognition method suitable for embedded hardware. Background Art
[0002] In recent years, with the rapid development of my country's economy and technology, cameras have been widely used in various fields of life. Pedestrian information in the video can effectively reflect human behavior or intention in a specific scene, and pedestrian information is the easiest information to obtain in public places. Pedestrian attribute recognition is to analyze the attribute characteristics of pedestrians in surveillance videos, such as gender, age, hair length, clothing type and color, etc. It can effectively improve the accuracy of technologies such as pedestrian retrieval and pedestrian re-identification in video surveillance. Therefore, pedestrian attribute recognition technology has broad application prospects.
[0003] In real-world scenarios, images in surveillance videos usually contain multiple pedestrians, complex backgrounds, and various obstructions (such as vehicles, buildings, etc.); in crowded scenes, images may contain multiple pedestrians, and there may be overlap or obstruction between pedestrians. Directly identifying pedestrian attributes from the entire image will cause serious interference. Through target detection technology, the system can focus on the local area of each pedestrian, ensuring that the system analyzes each pedestrian individually, thereby improving the accuracy of attribute recognition.
[0004] Pedestrian attributes vary. Some require shallow features, some require deep features, some require local features, and some require global features. It is crucial to extract a feature that contains different attributes. At the same time, for some fine-grained attributes, such as hats and hairstyles, detailed positioning is required. It is particularly important to accurately position small features. At present, most pedestrian attribute recognition algorithms are developed for general parallel computing hardware in server environments, which have the disadvantages of high power consumption and poor portability. Based on this, it is urgent to develop a pedestrian attribute recognition method based on an embedded hardware platform. This method can have high real-time and portability, lower power consumption, higher energy efficiency, and can run offline without relying on the network, which greatly improves the efficiency of pedestrian attribute recognition. Summary of the invention
[0005] To solve the above problems, the present invention aims to propose a multi-scale pedestrian attribute positioning and recognition method suitable for embedded hardware, which adaptively discovers the most discriminative areas through a flexible attribute region positioning network; in addition, a multi-scale bidirectional feature fusion structure is introduced, which not only fuses high-level global features with low-level local features, but also realizes cross-scale information supplementation and optimization through bidirectional flow of feature information, improves the discriminative ability of attribute features, and realizes more accurate recognition of pedestrian part attributes in complex scenes. The trained target detection and pedestrian attribute recognition models can be deployed in series on the embedded development board to realize edge AI.
[0006] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0007] A multi-scale pedestrian attribute positioning and recognition method suitable for embedded hardware includes the following steps:
[0008] S1: Input the target image to be detected into the pre-trained target detection model to obtain the bounding box coordinates and confidence score of the target image;
[0009] S2: The object detection model is detected through multi-scale features, which includes the fusion of feature maps of multiple scales. The open source dataset is used as training data to train the object detection model for single-class pedestrian detection.
[0010] S3: The embedded BN_inception network is used as the backbone network of the pedestrian attribute recognition network model to adapt to the lightweight network and improve the operator matching support on the board side. The pedestrian attribute features are extracted through training to obtain feature maps of different scales.
[0011] S4: Capture the feature maps CM1, CM2, and CM3 of different scales output by BN_inception respectively, use a 1×1 convolution layer to adjust the number of channels of the input feature map, use a bidirectional feature fusion method to downsample and upsample the original feature maps CM1, CM2, and CM3 to match the scales of feature maps at different levels, downsample the high-level feature maps to the low-level scale, and then upsample the low-level feature maps to the high-level scale to achieve feature alignment between different scales; then use a weighted fusion strategy to weight the feature maps of different scales to make fuller use of the semantic information and spatial details at different levels, connect the obtained feature maps with the residual module, and add them to the original input feature maps to obtain the weighted feature maps of each layer;
[0012] S5: Input the fused weighted feature map into the attribute region localization network, and find the most discriminative region through the adaptive mechanism to improve the classification and recognition accuracy of pedestrian attributes;
[0013] S6: Use open source datasets for training to learn to locate and identify various attributes of pedestrians based on image features at different scales;
[0014] S7: Deploy the model to the embedded development board through format conversion and quantization process to realize edge AI.
[0015] Furthermore, the target image to be detected in step S1 is a captured image in an actual scene.
[0016] Furthermore, the method further includes: preprocessing the captured image, specifically:
[0017] The formats of the acquired target images to be detected are uniformly changed to JPG format, and the sizes of the target images are uniformly set to 640*640 pixels, and the sizes of the attribute recognition images are uniformly set to 256*128 pixels.
[0018] Furthermore, the pedestrian attribute recognition network model in step S3 includes a backbone network BN_inception, a multi-scale component network for bidirectional feature fusion, and an attribute region positioning network ARL;
[0019] The backbone network BN_inception includes a 7×7 convolutional layer Conv1 and multiple inception blocks of different feature levels, and uses three different levels of incep_3b, incep_4d and incep_5b blocks to output three feature maps CM1, CM2 and CM3 respectively;
[0020] The multi-scale component network uses a bidirectional feature fusion method to downsample and upsample the original feature maps CM1, CM2, and CM3 to match the scales of feature maps at different levels; then a weighted fusion strategy is used to weight the feature maps of different scales, and the weighted feature maps M1, M2, and M3 of each layer are obtained by adding them to the original input feature maps. The fused feature maps are used as new input features x1, x2, and x3, and then three separate prediction vectors y1, y2, and y3 are obtained after the attribute localization network ARL is passed through;
[0021] The attribute region localization network ARL is used to operate the multi-scale fused feature map as new input features x1, x2 and x3, automatically discover the discriminant region of each attribute in a weakly supervised manner, and obtain three separate prediction vectors y1, y2, y3.
[0022] Further, the attribute region location network ARL includes a spatial transformer derived from STN;
[0023] The spatial transformer treats the attribute region as a simple bounding box and is trained end-to-end without region annotations; this is achieved by the following transformation:
[0024]
[0025] Among them, sx and sy are scaling parameters, tx and ty are translation parameters, and the desired bounding box can be obtained through these four parameters. and are the source coordinates and target coordinates of the I-th pixel respectively.
[0026] Furthermore, the input feature X i After a series of linear and nonlinear layers, a weight vector is generated to recalibrate cross-channel features and an additional residual link is used to maintain complementary information. Finally, the region-based features sampled by bilinear interpolation are used for attribute classification. The prediction of the mth attribute in the i-th layer is simply expressed as:
[0027]
[0028] A deep supervision mechanism is applied for training, where four individual predictions are directly supervised by the ground truth labels; during inference, multiple prediction vectors are aggregated through an efficient voting scheme to produce the maximal response at different feature levels.
[0029] Furthermore, the prediction vector is attribute feature information determined by a loss function of pedestrian attribute recognition;
[0030] Among them, the loss function of pedestrian attribute recognition adopts the weighted binary cross entropy loss function, which is expressed as:
[0031]
[0032] Among them, γ m =e -am is the loss weight of the mth attribute, am is the prior class distribution of the mth attribute, m is the number of attributes, i represents the i-th branch, where i∈{1,2,3,4}, and σ represents the sigmoid activation function; the total training loss is calculated by the sum of the four individual losses
[0033] Furthermore, the attribute region localization network ARL obtains four separate prediction vectors y1, y2, y3, and y4 from three ARL groups and one global branch; a deep supervision mechanism is applied to training, in which the four separate predictions are directly supervised by the basic truth labels. During the inference process, multiple prediction vectors are aggregated through an effective voting scheme to produce the maximum response at different feature levels, and a maximum voting scheme is used to select the best prediction from the most accurate different levels of the attribute region; thereby, the trained target detection and pedestrian attribute recognition models are deployed to the embedded system to realize edge AI.
[0034] Furthermore, the target detection and pedestrian attribute recognition models need to quantify the network model in the format supported by the hardware device during deployment and application, including:
[0035] Convert the float32 data type in the network model to the int8 data type for quantization. Assuming that the normalized 32-bit floating-point data is D, the quantization process of data D is expressed as follows:
[0036]
[0037] Among them, S q and ZP represents the asymmetric quantization parameter of the input tensor, D q Indicates quantized int8 data, round indicates rounding, and the function clamp indicates:
[0038]
[0039] Among them, the clamp function limits the randomly changing values to a given interval, a and b represent constants, and x represents a variable.
[0040] Furthermore, the open source dataset in step S2 is COCO2017, the open source dataset in step S6 is PA-100K and PETA, and the target detection model is the yolov5 network model.
[0041] Beneficial effects: The present invention improves the existing target detection model into a single-category detection model, and detects pedestrians. For the pedestrian attribute recognition model, the present invention adds an attribute region positioning network (ARL) on its basis to adaptively discover the most discriminative area. In addition, a multi-scale bidirectional feature fusion structure is introduced, which not only fuses high-level global features with low-level local features, but also realizes cross-scale information supplementation and optimization through bidirectional flow of feature information, improves the discriminative ability of attribute features, and realizes more accurate recognition of pedestrian part attributes in complex scenes. The improved network model is then converted and quantified to reduce the size of the model and improve the portability of the model while ensuring accuracy. The present invention uses a series mode to deploy it on an embedded system; finally, embedded hardware is used to accelerate the reasoning of the model, which has high practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0043] Figure 1 It is a main algorithm flow chart of a multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of a framework of a target detection model in a multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a framework of a pedestrian attribute recognition network model in a multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to an embodiment of the present invention;
[0046] Figure 4 A schematic diagram of the structure of an attribute region positioning network ARL in a multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to an embodiment of the present invention;
[0047] Figure 5 Two groups of attribute analysis result diagrams when the PETA data set is used in a multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to an embodiment of the present invention;
[0048] Figure 6 These are two sets of attribute analysis result diagrams when the PA-100K data set is used in a multi-scale pedestrian attribute positioning and recognition method suitable for embedded hardware described in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0050] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0051] The technical concept of the present invention is: first, the target detection model is trained as a single category, and the pedestrian attribute recognition model is trained as a multi-label model, so as to train the general feature extraction ability of the network; because two different models need to be trained separately, the training data sets and image resolutions are also different. When deployed on the board side, the two models need to be deployed in series, and the input target image also needs to be processed to a certain extent, as follows: the target pedestrian is detected by the target detection model, and the image frame coordinates of the pedestrian are obtained and cropped, and sent to the pedestrian attribute model for attribute label recognition. The present invention designs an attribute region positioning network (ARL) and a multi-scale bidirectional feature fusion structure, which not only fuses high-level global features with low-level local features, but also realizes cross-scale information supplementation and optimization through bidirectional flow of feature information, improves the discrimination ability of attribute features, and realizes more accurate recognition of pedestrian part attributes in complex scenes.
[0052] Example 1
[0053] Based on the above technical concept, see Figure 1-6 : A multi-scale pedestrian attribute positioning and recognition method suitable for embedded hardware in this embodiment includes the following steps:
[0054] S1: Input the target image to be detected into the pre-trained target detection model to obtain the bounding box coordinates and confidence score of the target image;
[0055] In actual scenes, images usually contain multiple pedestrians, complex backgrounds, and various obstructions (such as vehicles, buildings, etc.); in crowded scenes, images may contain multiple pedestrians, and there may be overlap or obstruction between pedestrians; therefore, in this embodiment, through the target detection technology (target detection model), the system can focus on the local area of each pedestrian, ensuring that the system analyzes each pedestrian separately, thereby improving the accuracy of attribute recognition; the target detection model of this embodiment adopts the yolov5 network model;
[0056] S2: The object detection model is detected through multi-scale features, which includes the fusion of feature maps of multiple scales. The open source dataset is used as training data to train the object detection model for single-class pedestrian detection.
[0057] The dataset of the target detection model of this embodiment uses the open source dataset coco2017 as training data, and performs single-class training on the target detection model. Obtain images containing pedestrians from the original dataset and divide them into a dataset and a validation set in proportion, and annotate the images of the dataset and the validation set with pedestrians; the annotation format includes: the bounding box coordinates of the target in the image, a single category label (unified as a category ID); preprocess the annotated image according to the following steps: scale the image to the target model input size (640x640) to ensure that the relative proportion of the target remains unchanged; normalize the pixel values to the range of [0,1], or standardize according to the mean and standard deviation of the pre-trained model; make it conform to the data format of the yolov5 network model;
[0058] S3: The embedded BN_inception network is used as the backbone network of the pedestrian attribute recognition network model to adapt to the lightweight network and improve the operator matching support on the board side. The pedestrian attribute features are extracted through training to obtain feature maps of different scales.
[0059] S4: Capture the feature maps CM1, CM2, and CM3 of different scales output by BN_inception respectively, use a 1×1 convolution layer to adjust the number of channels of the input feature map, use a bidirectional feature fusion method to downsample and upsample the original feature maps CM1, CM2, and CM3 to match the scales of feature maps at different levels, downsample the high-level feature maps to the low-level scale, and then upsample the low-level feature maps to the high-level scale to achieve feature alignment between different scales; then use a weighted fusion strategy to weight the feature maps of different scales to make fuller use of the semantic information and spatial details at different levels, connect the obtained feature maps with the residual module, and add them to the original input feature maps to obtain the weighted feature maps of each layer;
[0060] S5: Input the fused weighted feature map into the attribute region localization network, and find the most discriminative region through the adaptive mechanism to improve the classification and recognition accuracy of pedestrian attributes;
[0061] S6: Use open source datasets for training to learn to locate and identify various attributes of pedestrians based on image features at different scales;
[0062] The data set of the pedestrian attribute recognition network model of this embodiment uses the open source data set PA-100K and PETA as training data to train the pedestrian attribute recognition network model; the pedestrian attribute recognition training set includes pedestrian images and corresponding pedestrian attribute labels, wherein the categories of all pedestrian attributes constitute the overall pedestrian attribute set, which is divided into upper and lower body pedestrian attribute sets according to the upper and lower body parts of the pedestrians; the training images used are images taken by multiple cameras with non-overlapping fields of view in real scenes, and are images containing most parts of pedestrians obtained by pedestrian detector detection or manual calibration, and the pedestrian attribute labels are manually calibrated, and the attribute recognition image size is unified to 256*128 pixels;
[0063] S7: Deploy the model to the embedded development board through format conversion and quantization process to realize edge AI.
[0064] This embodiment adaptively discovers the most discriminative areas through a flexible attribute region positioning network. In addition, a multi-scale bidirectional feature fusion structure is introduced, which not only fuses high-level global features with low-level local features, but also realizes cross-scale information supplementation and optimization through bidirectional flow of feature information, improves the discriminative ability of attribute features, and realizes more accurate recognition of pedestrian part attributes in complex scenes. The trained target detection and pedestrian attribute recognition models can be deployed in series on the embedded development board to realize edge AI.
[0065] It should be noted that, considering the embedded board deployment, the model of this embodiment removes the attention mechanism component, and adopts the attribute region location network ARL component and the bidirectional feature fusion multi-scale component to match the operator support problem of the board and reduce the number of model parameters, making the model more lightweight.
[0066] In the specific implementation, the construction of the target detection model
[0067] Target detection process: Input the converted dataset and validation set images into Figure 2 Train the yolov5 network model, get the training results and obtain the model weight file.
[0068] In this embodiment, the input of the yolov5 network model is an image, and the output of the model is (x, y, w, h, c), which respectively represent the x and y coordinates of the prediction box in the image coordinate system, the width and height of the rectangle, and the confidence. In order to ensure that all targets are detected, multiple targets will be output as much as possible, and then the error correction in the later stage will be used to remove the wrong prediction results. The error correction method used is the non-maximum suppression method NMS, which uses the non-maximum suppression method to filter out the prediction results with higher overlap. Finally, the image is cropped according to the coordinate position information output by the model.
[0069] In a specific example, the target image to be detected in step S1 is a captured image in an actual scene.
[0070] In a specific example, the method further includes: preprocessing the captured image, specifically:
[0071] The formats of the acquired target images to be detected are uniformly changed to JPG format, and the sizes of the target images are uniformly set to 640*640 pixels, and the sizes of the attribute recognition images are uniformly set to 256*128 pixels.
[0072] In the specific implementation, the construction of the pedestrian attribute recognition network model
[0073] For pedestrian attribute recognition, many existing methods treat pedestrian attribute recognition as a multi-label problem and only extract various attribute features from a picture. Such methods rely on overall features, but pedestrian attributes are different. Some require shallow features, some require deep features, some require local features, and some require global features. Regional features are more useful for fine-grained attribute classification. In order to predict the existence of a specific attribute, it is necessary to locate the area related to the attribute.
[0074] In a specific example, Figure 3 As shown: the pedestrian attribute recognition network model in step S3 includes a backbone network BN_inception, a multi-scale component network with bidirectional feature fusion, and an attribute region positioning network ARL;
[0075] The backbone network BN_inception includes a 7×7 convolutional layer Conv1 and multiple inception blocks of different feature levels, and uses three different levels of incep_3b, incep_4d and incep_5b blocks to output three feature maps CM1, CM2 and CM3 respectively;
[0076] The multi-scale component network uses a bidirectional feature fusion method to downsample and upsample the original feature maps CM1, CM2, and CM3 to match the scales of feature maps at different levels; then a weighted fusion strategy is used to weight the feature maps of different scales, and the weighted feature maps M1, M2, and M3 of each layer are obtained by adding them to the original input feature maps. The fused feature maps are used as new input features x1, x2, and x3, and then three separate prediction vectors y1, y2, and y3 are obtained after the attribute localization network ARL is passed through;
[0077] The attribute region localization network ARL is used to operate the multi-scale fused feature map as new input features x1, x2 and x3, automatically discover the discriminant region of each attribute in a weakly supervised manner, and obtain three separate prediction vectors y1, y2, y3.
[0078] In a specific example, Figure 4 As shown: the attribute region localization network ARL includes a spatial transformer derived from STN;
[0079] The spatial transformer treats the attribute region as a simple bounding box and is trained end-to-end without region annotations; this is achieved by the following transformation:
[0080]
[0081] Among them, sx and sy are scaling parameters, tx and ty are translation parameters, and the desired bounding box can be obtained through these four parameters. and are the source coordinates and target coordinates of the I-th pixel respectively.
[0082] In a specific example, the input feature X i After a series of linear and nonlinear layers, a weight vector is generated to recalibrate cross-channel features and an additional residual link is used to maintain complementary information. Finally, the region-based features sampled by bilinear interpolation are used for attribute classification. The prediction of the mth attribute in the i-th layer is simply expressed as:
[0083]
[0084] A deep supervision mechanism is applied for training, where four individual predictions are directly supervised by the ground truth labels; during inference, multiple prediction vectors are aggregated through an efficient voting scheme to produce the maximal response at different feature levels.
[0085] In a specific example, the prediction vector is attribute feature information determined by a loss function of pedestrian attribute recognition;
[0086] Among them, the loss function of pedestrian attribute recognition adopts the weighted binary cross entropy loss function, which is expressed as:
[0087]
[0088] Among them, γ m =e -am is the loss weight of the mth attribute, am is the prior class distribution of the mth attribute, m is the number of attributes, i represents the i-th branch, where i∈{1,2,3,4}, and σ represents the sigmoid activation function; the total training loss is calculated by the sum of the four individual losses
[0089] In a specific example, the attribute region localization network ARL obtains four separate prediction vectors y1, y2, y3, and y4 from three ARL groups and one global branch; a deep supervision mechanism is applied to training, where the four separate predictions are directly supervised by the ground truth labels; during the inference process, multiple prediction vectors are aggregated through an effective voting scheme to generate the maximum response at different feature levels, and a maximum voting scheme is used to select the best prediction from the most accurate different levels of the attribute region; thereby, the trained target detection and pedestrian attribute recognition models are deployed to embedded systems to realize edge AI.
[0090] Experimental analysis of the pedestrian attribute recognition network model in this embodiment
[0091] Experimental analysis of the above methods is carried out:
[0092] The graphics card used in the relevant training process in this embodiment is NVIDIA GeForce RTX 3090; the processor is Intel(R) Xeon(R) Silver4114CPU@2.20GHz; the training software environment is Ubuntu18.04, CUDA Version:11.2, Pytorch1.10.0, Python3.7.
[0093] In this example, we use the adaptive learning rate (ADAM) method as the optimizer, with the initial learning rate set to 0.001, and use the StepLR strategy to decay by 0.1 times every 20 epochs. Set the momentum to 0.9 and the weight decay to 0.0005. Perform 60 rounds of training, and use the early stopping method to terminate the training early based on the performance of the validation set. Set the batch size to 32 on each GPU and adjust it accordingly based on the number of GPUs used.
[0094] In order to verify the effectiveness of the proposed algorithm, this embodiment uses five evaluation criteria, namely label-based average accuracy (mA) and instance-based accuracy (Accu), precision (Prec), recall (recall) and F1-score (F1-score), to compare the improved algorithm with the original algorithm using the PETA and PA-100K datasets.
[0095] (1) Analysis of PETA data set results
[0096] The PETA dataset was proposed by Deng et al. from the Department of Information Engineering at the Chinese University of Hong Kong. It consists of 8 outdoor scenes and 2 indoor scenes, including 8705 pedestrians and a total of 19,000 images. Its resolution range is large, consisting of images ranging from 17*39 to 169*365. Each pedestrian is annotated with 61 binary and 4 multi-category attributes. We selected 35 attributes with an accuracy rate greater than 5% for evaluation. Figure 5 As shown in the figure, two sets of attribute analysis results are shown when using the PETA dataset. The result of pedestrian attribute analysis is shown on the right side of the picture. Figure 5 The recognition result in a is a short-haired male with the age between 31 and 45 and wearing shoes; Figure 5 The recognition result in b is a short-haired male between the ages of 16 and 30 wearing jeans. Figure 5 Like a, the gender attribute is the default attribute and is not displayed.
[0097] Table 1 Performance analysis using PETA dataset
[0098]
[0099] Table 1 shows the algorithm proposed by the present invention and the baseline algorithm, and gradually adds each component network: bidirectional feature fusion multi-scale component, ARL component and attention mechanism component, and compares them with each other at the same time. The experimental comparison shows that by adding each component module, each evaluation index has been improved to a certain extent, but the number of parameters has also increased. Considering the deployment of the embedded board, we removed the attention mechanism component and used the ARL component and bidirectional feature fusion multi-scale component. Through the comparison of experimental data, we reduced the number of parameters by 3.36M, and only the precision (Prec) indicator showed a slight decline. Other indicators also have a certain improvement, which is reflected in the good effect of the algorithm on the PETA dataset.
[0100] (2) Experimental comparison on the PA-100K dataset
[0101] PA-100K was proposed by Liu et al. As a large-scale pedestrian attribute dataset, PA-100K contains 100,000 pedestrian images taken in 598 scenes. In the PA-100K dataset, the attributes are set to 26 types, including gender, age, and object attributes such as handbags, clothing, etc. Compared with other public datasets, PA-100K provides a wide range of pedestrian attribute datasets. Figure 6 As shown in the figure, two sets of pedestrian attribute analysis results are shown on the right side of the picture when the PA-100K dataset is used. Figure 6The recognition result in c is a male aged between 18 and 60, wearing short sleeves and pants. The gender attribute is used as the default attribute and is not displayed. Figure 6 The recognition result in d is a female between the ages of 18 and 60, wearing glasses and long sleeves and pants.
[0102] Table 2 Performance analysis using the PA-100K dataset
[0103]
[0104] As can be seen from Table 2, in the PA-100K dataset, the situation of each evaluation indicator is roughly the same as that on the PETA dataset. Compared with the baseline algorithm, each evaluation indicator has a certain amount of improvement. The mA score increased by 2.47%, and the Recall score increased by 2.25%.
[0105] Model deployment and application process
[0106] Embedded software deployment mainly includes model file conversion and quantization, development board environment configuration, and inference verification. In order to adapt to the characteristics of embedded hardware, the trained model needs to be converted and quantized. First, the trained yolov5 network model and pedestrian attribute recognition model are converted to ONNX format; the network model in the format supported by the hardware device is quantized, and then the quantized network model is deployed on the edge device and loaded.
[0107] Model quantization uses uint8 (asymmetric quantization). Under the premise of ensuring accuracy, model quantization can effectively reduce the model size, reduce storage space, and speed up reasoning.
[0108] In a specific example, the target detection and pedestrian attribute recognition model needs to quantize the network model in the format supported by the hardware device during the deployment and application process, including:
[0109] Convert the float32 data type in the network model to the int8 data type for quantization. Assuming that the normalized 32-bit floating-point data is D, the quantization process of data D is expressed as follows:
[0110]
[0111] Among them, S q and ZP represents the asymmetric quantization parameter of the input tensor, D q Indicates quantized int8 data, round indicates rounding, and the function clamp indicates:
[0112]
[0113] Among them, the clamp function limits the randomly changing values to a given interval, a and b represent constants, and x represents a variable.
[0114] Then, the quantized target detection model and attribute recognition model are deployed in series on the embedded development board and the quantized network model is loaded.
[0115] The board-side inference process is as follows: the image to be predicted is input into the quantized network model, and the quantized network model is used to preprocess the image to be predicted. The image preprocessing part mainly includes operations such as the size of the input image and the channel order, which must be adjusted according to the input format of the model;
[0116] Then, the quantized target detection model is used to perform inference and prediction on the image, and the model weight file of the target detection model is loaded to obtain the output result (x, y, w, h). The image is cropped according to the corresponding rectangular coordinates of the output result, and the cropped image is sent to the quantized pedestrian attribute recognition network model. According to the attribute recognition model weight file, the label attribute information of the pedestrian target image is obtained, and the output result such as the pedestrian detection box and the pedestrian attribute information is represented on the original image to obtain a visual prediction result.
[0117] In summary, the method of the present invention performs pedestrian attribute recognition based on the yolov5 detection model and the improved attribute region positioning network (ARL) and the multi-scale method of bidirectional feature fusion, and trains two network models; then the trained network model is converted and quantized to reduce the size of the model and improve the portability of the model while ensuring accuracy; and the trained target detection and pedestrian attribute recognition models are finally deployed on the embedded system in a series form, which has high practicality. The experimental results of the improved attribute recognition network model and the baseline model on the PETA and PA-100K data sets show that the attribute recognition model proposed by the present invention has better performance while reducing the size of the network model.
[0118] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-scale pedestrian attribute positioning and recognition method suitable for embedded hardware, characterized in that: The process includes the following steps: S1: Input the target image to be detected into the pre-trained target detection model to obtain the bounding box coordinates and confidence score of the target image; S2: The object detection model is detected through multi-scale features, which includes the fusion of feature maps of multiple scales. The open source dataset is used as training data to train the object detection model for single-class pedestrian detection. S3: The embedded BN_inception network is used as the backbone network of the pedestrian attribute recognition network model to adapt to the lightweight network and improve the operator matching support on the board side. The pedestrian attribute features are extracted through training to obtain feature maps of different scales. S4: Capture the feature maps CM1, CM2, and CM3 of different scales output by BN_inception respectively, use a 1×1 convolution layer to adjust the number of channels of the input feature map, and use the bidirectional feature fusion method to downsample and upsample the original feature maps CM1, CM2, and CM3 to match the scales of feature maps at different levels. Downsample the high-level feature map to the low-level scale, and then upsample the low-level feature map to the high-level scale to achieve feature alignment between different scales. Then, a weighted fusion strategy is adopted to perform weighted processing on feature maps of different scales to make fuller use of semantic information and spatial details at different levels. The obtained feature maps are connected with the original input feature maps using residual modules and added to obtain the weighted feature maps of each layer. S5: The fused weighted feature map is input into the attribute region localization network, and the most discriminative region is found through an adaptive mechanism to improve the classification and recognition accuracy of pedestrian attributes; S6: Use open source datasets for training to learn to locate and identify various attributes of pedestrians based on image features at different scales; S7: Deploy the model to the embedded development board through format conversion and quantization process to realize edge AI.
2. The multi-scale pedestrian attribute positioning and identification method applicable to embedded hardware according to claim 1 is characterized in that: The target image to be detected in step S1 is a captured image in an actual scene.
3. The multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to claim 2 is characterized in that: Also includes: The captured image is preprocessed, specifically: The formats of the acquired target images to be detected are uniformly changed to JPG format, and the sizes of the target images are uniformly set to 640*640 pixels, and the sizes of the attribute recognition images are uniformly set to 256*128 pixels.
4. The multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to claim 1, characterized in that: The pedestrian attribute recognition network model in step S3 includes a backbone network BN_inception, a multi-scale component network with bidirectional feature fusion, and an attribute region positioning network ARL; The backbone network BN_inception includes a 7×7 convolutional layer Conv1 and multiple inception blocks of different feature levels, and uses three different levels of incep_3b, incep_4d and incep_5b blocks to output three feature maps CM1, CM2 and CM3 respectively; The multi-scale component network uses a bidirectional feature fusion method to downsample and upsample the original feature maps CM1, CM2, and CM3 to match the scales of feature maps at different levels; then a weighted fusion strategy is used to weight the feature maps of different scales, and the weighted feature maps M1, M2, and M3 of each layer are obtained by adding them to the original input feature maps. The fused feature maps are used as new input features x1, x2, and x3, and then three separate prediction vectors y1, y2, and y3 are obtained after the attribute localization network ARL is passed; The attribute region localization network ARL is used to operate the multi-scale fused feature map as new input features x1, x2 and x3, automatically discover the discriminant region of each attribute in a weakly supervised manner, and obtain three separate prediction vectors y1, y2, y3.
5. The multi-scale pedestrian attribute positioning and identification method applicable to embedded hardware according to claim 4 is characterized in that: The attribute region positioning network ARL includes a spatial transformer derived from STN; The spatial transformer treats the attribute region as a simple bounding box and is trained end-to-end without region annotations; this is achieved by the following transformation: Among them, sx and sy are scaling parameters, tx and ty are translation parameters, and the desired bounding box can be obtained through these four parameters. and are the source coordinates and target coordinates of the I-th pixel respectively.
6. The multi-scale pedestrian attribute positioning and identification method applicable to embedded hardware according to claim 4 is characterized in that: The input feature X i After a series of linear and nonlinear layers, a weight vector is generated to recalibrate cross-channel features and an additional residual link is used to maintain complementary information. Finally, the region-based features sampled by bilinear interpolation are used for attribute classification. The prediction of the mth attribute in the i-th layer is simply expressed as: A deep supervision mechanism is applied for training, where four individual predictions are directly supervised by the ground truth labels; during inference, multiple prediction vectors are aggregated through an efficient voting scheme to produce the maximal response at different feature levels.
7. The multi-scale pedestrian attribute positioning and identification method applicable to embedded hardware according to claim 4 is characterized in that: The prediction vector is attribute feature information determined by the loss function of pedestrian attribute recognition; Among them, the loss function of pedestrian attribute recognition adopts the weighted binary cross entropy loss function, which is expressed as: Among them, γ m =e -am is the loss weight of the mth attribute, am is the prior class distribution of the mth attribute, m is the number of attributes, i represents the i-th branch, where i∈{1,2,3,4}, and σ represents the sigmoid activation function; the total training loss is calculated by the sum of the four individual losses 8. The multi-scale pedestrian attribute positioning and identification method applicable to embedded hardware according to claim 4 is characterized in that: The attribute region localization network ARL obtains four separate prediction vectors y1, y2, y3, and y4 from three ARL groups and one global branch; a deep supervision mechanism is applied to training, where the four separate predictions are directly supervised by the ground truth labels. During the inference process, multiple prediction vectors are aggregated through an effective voting scheme to produce the maximum response at different feature levels, and a maximum voting scheme is used to select the best prediction from the most accurate different levels of the attribute region; thereby the trained target detection and pedestrian attribute recognition models are deployed to embedded systems to realize edge AI.
9. The multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to claim 8, characterized in that: The target detection and pedestrian attribute recognition models need to quantify the network model in the format supported by the hardware device during deployment and application, including: Convert the float32 data type in the network model to the int8 data type for quantization. Assuming that the normalized 32-bit floating-point data is D, the quantization process of data D is expressed as follows: Among them, S q and ZP represents the asymmetric quantization parameter of the input tensor, D q Indicates quantized int8 data, round indicates rounding, and the function clamp indicates: Among them, the clamp function limits the randomly changing values to a given interval, a and b represent constants, and x represents a variable.
10. The multi-scale pedestrian attribute positioning and recognition method applicable to embedded hardware according to claim 1, characterized in that: The open source dataset in step S2 is COCO2017, the open source dataset in step S6 is PA-100K and PETA, and the target detection model is the yolov5 network model.
Citation Information
Patent Citations
Human body attribute recognition method based on attention mechanism and multi-task learning
CN111597870A
Pedestrian attribute identification method combining global and local features
CN116311372A