Pedestrian Attribute Recognition Method Based on Prompt-Finetuning of Pre-Trained Large Models
Through a pre-trained large model based on tip fine-tuning, image and attribute features are extracted using CLIP vision and text encoder, and feature fusion is performed through multimodal Transformer module, the problems of low pedestrian attribute recognition accuracy and poor generalization ability in the prior art are solved, and more efficient pedestrian attribute recognition is achieved.
Patent Information
- Application Number
- CN202310081570.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-01-16
AI Technical Summary
The existing pedestrian attribute recognition methods fail to make full use of the relationship between pedestrian images and attribute labels, resulting in low recognition accuracy and poor generalization ability.
A pre-trained large model based on cues fine-tuning is adopted, and image and attribute features are extracted using CLIP vision and text encoder, and feature fusion is performed through multimodal Transformer module, and prediction is performed by combining a feedforward network.
It improves the accuracy and generalization ability of pedestrian attribute recognition, makes full use of attribute semantic information, and reduces the increase in computational volume.
Smart Images

Figure CN116259075B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a pedestrian attribute recognition method based on prompt fine-tuning of a pre-trained large model. Background Art
[0002] The goal of pedestrian attribute recognition is to describe the intermediate-level semantic information of people using a set of predefined attributes (such as age, height, hairstyle, clothing). It plays an important role in the field of computer vision, especially in intelligent video surveillance and autonomous driving, and also promotes the research of other visual tasks, including pedestrian re-identification, pedestrian search, and pedestrian detection. With the help of artificial intelligence, such as CNN (Convolutional Neural Network) and RNN (Recursive Neural Network), this research field has received extensive attention and made great progress. However, due to poor imaging quality in extreme cases (including motion blur, shadows, occlusions, low resolution, multi-views, and nighttime), pedestrian attribute recognition remains a challenging task.
[0003] Most existing pedestrian attribute methods are based on CNN and RNN networks, making it difficult to utilize the high-level semantic information of pedestrians, resulting in low recognition accuracy. Moreover, CNN-based methods do not consider the semantic relevance of pedestrian attributes, leading to suboptimal performance, while RNN-based methods overly rely on the artificially predefined attribute order and are difficult to achieve the best performance. For example, in the paper "Deep-camp: Deep convolutional action & attribute mid-level patterns", a part-based model and a CNN-based pedestrian attribute recognition are combined, and the training of CNN is accelerated to learn stronger normalized features from a smaller dataset. However, this method using CNN as the backbone network has defects. Since there are internal associations between pedestrian attributes, such as the two attributes "long hair" and "female" being highly correlated, using such pure visual pedestrian attribute methods will ignore the semantic information of attributes and lead to suboptimal problems. Although there are existing works based on Transformer that integrate visual and text information and solve the above problems to a certain extent, due to the use of independently pre-trained visual and text encoders, there are significant differences between the visual and text features, which may limit the subsequent visual and text modality fusion process and reduce the recognition accuracy. In addition, most existing pedestrian attribute recognition methods use models pre-trained on single-modal datasets, which results in poor generalization ability of the models and significant differences between image and text features. Summary of the Invention
[0004] The objective of the present invention is to design a pedestrian attribute recognition method based on prompt fine-tuning of a pre-trained large model, so as to solve the problems of sub-optimality and poor generalization ability caused by the failure to fully utilize the relationship between pedestrian images and attribute labels in the prior art.
[0005] The present invention solves the above technical problems through the following technical solutions:
[0006] A pedestrian attribute recognition method based on a pre-trained large model with added prompt fine-tuning, wherein the pre-trained large model includes: a CLIP visual encoder, a CLIP text encoder, a multi-modal Transformer module, and a classifier module; the CLIP visual encoder and the CLIP text encoder are visual and text feature extractors of the visual language model CLIP; the multi-modal Transformer module performs adaptive fusion and long-distance modeling on attributes through a multi-head self-attention mechanism, and obtains the fused features after passing through multiple layers of Transformer encoder layers; the classifier module uses FFN to obtain the scores of each attribute and output the classification result;
[0007] The pedestrian attribute recognition method includes the following steps:
[0008] Step 1: Preprocess the input pedestrian image to be classified and the pedestrian attributes to be evaluated;
[0009] Step 2: Respectively input the pedestrian image to be classified and the pedestrian attributes to be evaluated into the pre-trained large model to obtain visual features and text features respectively;
[0010] Step 3: Connect the obtained visual features and text features and input them into the multi-modal Transformer module to perform modal fusion and information interaction on the connected visual features and text features to obtain the fused and interacted features;
[0011] Step 4: Extract the fused tokens (Tokens) corresponding to the positions of the text features, and input them into the classifier to obtain the scores of each attribute;
[0012] Step 5: Determine whether the score is greater than the threshold. If the score is greater than the threshold, the attribute is considered to exist; otherwise, it is considered to not exist. After comparing each attribute with the threshold, the prediction result is output.
[0013] Further, the CLIP visual encoder uses ResNet or a visual Transformer encoder; the CLIP text encoder is designed based on a Transformer encoder and uses the model parameters of CLIP ViT-L / 14.
[0014] Further, the method for preprocessing the input pedestrian image to be classified and the pedestrian attributes to be evaluated in step one is as follows: Preprocess the input pedestrian image: Fill the black edges of the pedestrian image in advance to prevent distortion of pedestrian features during subsequent resizing, resize the pedestrian image to 224*224, and perform data augmentation of random horizontal flipping and random cropping during training; Preprocess the input pedestrian image: Use a prompt template to expand the attribute phrases in the input pedestrian attribute set into language descriptions to adapt to the CLIP text encoder.
[0015] Further, the training method of the pre-trained large model in step two is as follows: The CLIP visual encoder and the CLIP text encoder load the model parameters of CLIP ViT-L / 14, and the multi-modal Transformer module loads the model parameters of ViT-B / 16 pre-trained on the ImageNet-21K dataset and fine-tuned on the ImageNet-1K dataset.
[0016] Further, the method for obtaining visual features in step two is as follows: Add multiple learnable prompt tokens to the input tokens of each layer of the Transformer encoder layer in the CLIP visual encoder, located between the classification token and the image patch tokens, so as to fine-tune the CLIP visual encoder, and obtain visual features after passing through multiple layers of Transformer encoder layers.
[0017] Further, the method for obtaining text features in step two is as follows: After tokenizing the segmented and augmented attribute sentences, obtain the text embeddings after passing through the embedding layer and send them into the CLIP text encoder. Add multiple learnable prompt tokens to the input tokens of each layer of the Transformer encoder layer in the CLIP text encoder, located after the text tokens, so as to fine-tune the CLIP text encoder, and obtain the text features after passing through multiple layers of Transformer encoder layers.
[0018] The advantages of the present invention are as follows:
[0019] (1) In view of the characteristics that the existing pedestrian attribute recognition methods cannot make full use of attribute semantic information and have poor generalization ability, the present invention uses the visual and text encoders of CLIP to extract image features and attribute features. After fusing the two-modal features through a multi-modal Transformer module, the prediction result is obtained through a feed-forward network. By modeling the pedestrian attribute recognition problem as a visual-language fusion problem, a pre-trained large visual-language model is used as the backbone network to extract visual and text features with better inter-modal connections, and then the connection between vision and text is modeled through a multi-modal Transformer, making full use of the attribute semantic information. It can be seen that the pre-trained large model's good generalization ability is retained through prompt tuning, and the model has stronger practicality.
[0020] (2) The method of the present invention fuses the connected visual-text features through the global modeling ability of the Transformer, making better use of the semantic information of the attributes.
[0021] (3) The method of the present invention selects to use the CLIP large model pre-trained on 400 million image-text pairs to alleviate these problems. However, using a large model as the backbone network will bring an increase in computational complexity. The method of introducing prompt tuning is used to reduce the number of parameters to be adjusted. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flowchart of the pedestrian attribute recognition method based on prompt-tuning pre-trained large model in Embodiment 1 of the present invention;
[0023] Figure 2 is a schematic diagram of the network model structure of the pedestrian attribute recognition method based on prompt-tuning pre-trained large model in Embodiment 1 of the present invention;
[0024] Figure 3 is an experimental result of the pedestrian attribute recognition method based on prompt-tuning pre-trained large model in Embodiment 1 of the present invention when tested on the PETA and PA100k pedestrian attribute datasets and a comparison chart with other methods;
[0025] Figure 4 is an experimental result of the pedestrian attribute recognition method based on prompt-tuning pre-trained large model in Embodiment 1 of the present invention when tested on the RAPv1 and RAPv2 pedestrian attribute datasets and a comparison chart with other methods;
[0026] Figure 5 is an experimental result of the pedestrian attribute recognition method based on prompt-tuning pre-trained large model in Embodiment 1 of the present invention when tested on the WIDER pedestrian attribute dataset and a comparison chart with other methods;
[0027] Figure 6It is the experimental results of the pedestrian attribute recognition method based on prompt fine-tuning of the pre-trained large model in Embodiment 1 of the present invention tested on the PETA-ZS and RAP-ZS pedestrian attribute datasets and the comparison graph with other methods. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments:
[0030] Embodiment 1
[0031] As Figure 1 shown, it is the flowchart of the pedestrian attribute recognition method based on prompt fine-tuning of the pre-trained large model in the embodiment of the present invention, including the following steps:
[0032] Step 1: Preprocess the input pedestrian image to be classified and the pedestrian attributes to be evaluated;
[0033] Step 2: Send the pedestrian image to be classified and the pedestrian attributes to be evaluated into the pre-trained large model with a CLIP visual encoder and a CLIP text encoder added with prompts respectively, so as to obtain visual features and text features respectively;
[0034] Step 3: After connecting the obtained visual features and text features, send them into the multi-modal Transformer module to perform modal fusion and information interaction on the connected visual features and text features, and obtain the features after fusion and interaction;
[0035] Step 4: Take out the features after fusion at the position of the text features, and send them into the classifier to obtain the scores of each attribute;
[0036] Step 5: Determine whether the score is greater than the threshold. The attribute with a score greater than the threshold is regarded as the existence of the attribute, otherwise it is regarded as the non-existence of the attribute. After each attribute is compared with the threshold, the prediction result is output.
[0037] As Figure 2 shown, it is the schematic diagram of the network model structure adopted by the present invention. The network model includes: a CLIP visual encoder, a CLIP text encoder, a multi-modal Transformer module, and an FFN (feed-forward neural network) module; Figure 2The medium attribute set is a list of attributes to be evaluated. F.E is the feature embedding, P.E is the position embedding, the prompt is the added learnable prompt vector. The CLIP visual encoder and the CLIP text encoder are the visual and text feature extractors of the vision-language model CLIP, where the CLIP visual encoder uses ResNet or a vision Transformer; the CLIP text encoder is designed based on the Transformer encoder and uses the model parameters of CLIP ViT-L / 14; the multimodal Transformer module is a 12-layer Transformer; Add&Norm is the residual connection and layer normalization; the CLIP (Contrastive Language-Image Pre-Training) model is a pre-trained neural network model released by OpenAI in early 2021 for matching images and texts.
[0038] The training process and the testing process of the model are as follows:
[0039] (1) Training process
[0040] 1) The CLIP visual encoder and the CLIP text encoder load the model parameters of CLIP ViT-L / 14, and the multimodal Tranformer module loads the model parameters of ViT-B / 16 that are pre-trained on the ImageNet-21K dataset and fine-tuned on the ImageNet-1K dataset.
[0041] 2) Preprocess the input pedestrian image. Fill the black edges of the pedestrian image in advance to prevent distortion of pedestrian features during subsequent resizing. Resize the pedestrian image to 224*224 and perform data augmentation of random horizontal flipping and random cropping during the training process. Split and augment the input pedestrian attribute set to obtain attribute sentences to adapt to the CLIP text encoder.
[0042] 3) The preprocessed pedestrian image is embedded to obtain the image embedding, which is then fed into the CLIP visual encoder. The embedding layer includes a feature embedding F.E and a position embedding P.E. 25 learnable prompt tokens are added to the input tokens of each layer of the Transformer encoder layer in the CLIP visual encoder, and the position is between the classification token and the image patch tokens, so as to fine-tune the CLIP visual encoder. After passing through 24 layers of the Transformer encoder layer, the image features are obtained. At the same time, the segmented and augmented attribute sentences are tokenized, and after passing through the embedding layer, the text embedding is obtained and fed into the CLIP text encoder. 3 learnable prompt tokens are added to the input tokens of each layer of the Transformer encoder layer in the CLIP text encoder, and the position is after the text tokens, so as to fine-tune the CLIP text encoder. After passing through 12 layers of the Transformer encoder layer, the text features are obtained.
[0043] 4) The image features and the text features are concatenated and fed into the multi-modal Transformer module for modality fusion and information interaction. Through the multi-head self-attention mechanism, adaptive fusion and long-range modeling of the attributes are performed. After passing through 12 layers of the Transformer encoder layer, the fused features are obtained. Finally, the tokens at the corresponding positions of the text features are fed into the FFN to obtain the scores of each attribute and output the classification results.
[0044] 5) Only the prompt tokens and the FFN in the model are trained, and the model parameters of the remaining parts are kept frozen. The prompt tokens are randomly initialized, and the random gradient descent optimizer is used to train for 20 epochs on all datasets. Based on the cosine learning rate scheduler, the warm-up process is set to 5 epochs. During the warm-up period, the initial learning rate decreases at a rate of 0.01, and the weight decay is 0.0001. The batch size is set to 16. On the PETA, PA100k, RAPv1, and RAPv2 datasets, the learning rate of 0.016 is used for the prompt tokens, and the learning rate of 0.008 is used for the FFN. On the WIDER, PETA-ZS, and RAP-ZS datasets, the learning rate of 0.002 is used for the prompt tokens, and the learning rate of 0.001 is used for the FFN;
[0045] 6) Finally, the model is saved for the testing process.
[0046] (2) Testing process
[0047] 1) Let the CLIP visual and text encoders load the model parameters of CLIP ViT-L / 14, and the multi-modal Transformer load the model parameters of ViT-B / 16 pre-trained on the ImageNet-21K dataset and fine-tuned on the ImageNet-1K dataset. Load the prompt tokens and FFN parameters saved during the training phase.
[0048] 2) Preprocess the input pedestrian image by padding it with black borders, resizing the pedestrian image to 224*224, segmenting and augmenting the input pedestrian attributes to obtain an attribute sentence to adapt to the text encoder of CLIP.
[0049] 3) Send the preprocessed pedestrian image and the pedestrian attributes to be evaluated into the CLIP visual encoder and text encoder with the loaded prompt parameters respectively to obtain visual and text features. After connecting the obtained visual and text features and sending them into the multi-modal Transformer for fusion, the interacted features are obtained. The part corresponding to the text features is sent into the FFN to obtain the scores of each attribute and output the classification results.
[0050] Experimental results
[0051] Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 are the experimental results of the method of the present invention and the comparison graphs with other methods. The tests are respectively carried out on 5 main pedestrian attribute datasets, namely PETA and PA100k, RAPv1 and RAPv2, WIDER, PETA-ZS and RAP-ZS. Among them, PETA-ZS and RAP-ZS are the datasets of the PETA and RAPv2 datasets under the zero-shot segmentation method. The test results are evaluated with other pedestrian attribute recognition methods in terms of mA (average precision of all attributes), Acc (average precision of all samples), Prec (accuracy), Recall (recall rate) and F1 score. Among them, PromptPAR represents the evaluation result of the present invention, and its classification accuracy has achieved good results.
[0052] In the present invention, by treating pedestrian attribute recognition as a visual-language fusion problem and making full use of the relationship between pedestrian images and attributes, the attribute phrases are first extended into sentences, and a pre-trained visual-language model is used as the backbone network to extract features of images and attributes. The CLIP model that conducts contrastive learning on the image-text pair dataset well connects the visual and language modalities in the feature space, and the Vision Transformer used in CLIP well models the long-range relationships of pixels. Then, a multi-modal Transformer is used to effectively fuse the features of the two modalities, and a feed-forward network is used for attribute prediction. To effectively optimize the framework, a prompt fine-tuning technique is adopted, only the prompt vector and the classification head are adjusted, and the parameters of the visual-language model and the multi-modal Transformer module are fixed, effectively reducing the number of parameters to be adjusted; by using the method of prompt fine-tuning to fine-tune the pre-trained large model to narrow the gap between visual-language features, improve the model generalization ability, and by using the multi-modal Transformer to model the connection between vision and text, the attribute semantic information is fully utilized.
[0053] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A pedestrian attribute recognition method for fine-tuning a pre-trained large model based on prompts, characterized in that The pre-trained large model includes: a CLIP visual encoder, a CLIP text encoder, a multi-modal Transformer module, and a classifier module; the CLIP visual encoder and the CLIP text encoder are the visual and text feature extractors of the visual language model CLIP; the multi-modal Transformer module adaptively fuses and performs long-range modeling on attributes through the multi-head self-attention mechanism, and obtains the fused features after passing through multiple Transformer encoder layers; the classifier module uses FFN to obtain the scores of each attribute and output the classification result; The pedestrian attribute recognition method includes the following steps: Step 1: Preprocess the input pedestrian image to be classified and the pedestrian attributes to be evaluated; Step 2: Send the pedestrian image to be classified and the pedestrian attributes to be evaluated into the pre-trained large model respectively, so as to obtain visual features and text features respectively; The method for obtaining the visual features is as follows: Add multiple learnable prompt tokens to the input tokens of each layer of the Transformer encoder layer of the CLIP visual encoder, and the position is between the classification token and the image patch token, so as to fine-tune the CLIP visual encoder, and obtain visual features after passing through multiple Transformer encoder layers; The method for obtaining the text features is as follows: After tokenizing the segmented and augmented attribute sentences, obtain the text embedding after passing through the embedding layer and send it into the CLIP text encoder. Add multiple learnable prompt tokens to the input tokens of each layer of the Transformer encoder layer of the CLIP text encoder, and the position is after the text tokens, so as to fine-tune the CLIP text encoder, and obtain text features after passing through multiple Transformer encoder layers; Step 3: Connect the obtained visual features and text features and send them into the multi-modal Transformer module to perform modal fusion and information interaction on the connected visual features and text features, and obtain the fused and interactive features; Step 4: Take out the fused token Token corresponding to the position of the text features, send it into the classifier, and obtain the scores of each attribute; Step 5: Judge whether the score is greater than the threshold. The attribute with a score greater than the threshold is regarded as the existence of the attribute, otherwise it is regarded as the non-existence of the attribute. Each attribute is compared with the threshold and the prediction result is output.
2. The pedestrian attribute recognition method based on prompt fine-tuning of a pre-trained large model according to claim 1, wherein The CLIP visual encoder uses ResNet or a visual Transformer encoder; the CLIP text encoder is designed based on the Transformer encoder and uses the model parameters of CLIP ViT-L / 14.
3. The pedestrian attribute recognition method based on prompt fine-tuning of a pre-trained large model according to claim 1, wherein, The method for preprocessing the input pedestrian image to be classified and the pedestrian attributes to be evaluated in Step 1 is as follows: Preprocess the input pedestrian image: Fill the black edges of the pedestrian image in advance to prevent distortion of pedestrian features during subsequent size adjustment. Resize the pedestrian image to 224*224, and perform data augmentation of random horizontal flipping and random cropping during training; Preprocess the input pedestrian image: Use the prompt template to expand the attribute phrases in the input pedestrian attribute set into language descriptions.
4. The pedestrian attribute recognition method based on prompt fine-tuning of a pre-trained large model according to claim 3, wherein The training method of the pre-trained large model described in Step 2 is as follows: The CLIP visual encoder and the CLIP text encoder load the model parameters of CLIP ViT-L / 14, and the multi-modal Transformer module loads the model parameters of ViT-B / 16 that are pre-trained on the ImageNet-21K dataset and fine-tuned on the ImageNet-1K dataset.