Pedestrian attribute identification method based on open vocabulary adaptive positioning

Through the cross-modal pedestrian attribute recognition framework, combined with OpenPose and SAM algorithms for pose estimation and semantic segmentation, the problems of insufficient utilization of single modality information and unknown attribute recognition in existing technologies are solved, and efficient and accurate recognition is achieved in complex environments.

CN120599657APending Publication Date: 2025-09-05HENAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510684296.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing pedestrian attribute recognition technologies mainly rely on single-modal information, which makes it difficult to maintain efficiency and accuracy in complex scenarios, unable to handle unknown attributes, and existing methods have shortcomings in cross-modal feature alignment.

Method used

A cross-modal pedestrian attribute recognition framework is constructed, combining OpenPose and SAM algorithms for pose estimation and semantic segmentation. Cross-modal alignment of image and text features is achieved through multi-head attention mechanism and knowledge distillation mechanism, and a cross-modal alignment loss function is introduced to optimize feature matching.

Benefits of technology

The accuracy and robustness of pedestrian attribute recognition are improved, the system can handle unknown attributes in complex environments, and the adaptability and generalization ability of the system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599657A_ABST
    Figure CN120599657A_ABST
Patent Text Reader

Abstract

The invention provides a pedestrian attribute recognition method based on open vocabulary adaptive positioning, which comprises the following steps of: preprocessing an image in an image data set to obtain a preprocessed image, and extracting text features of an input prompt language text; extracting skeleton key points of all pedestrians by using a key point extraction technology to obtain a predicted heat map, determining a key point information set, and selecting a target pedestrian closest to the center of the image; key point screening is carried out to obtain prompt points, and a final key point set is generated; performing target detection on pedestrians to obtain a prompt box, converting a final key point set into a prompt vector, and performing segmentation by using an SAM algorithm to obtain a segmentation mask graph; generating a posture region positioning map; and inputting the posture region positioning map and the text for describing the pedestrian into a cross-modal open vocabulary pedestrian attribute identification module to obtain the specific attribute of the pedestrian. According to the method, the problem that a traditional method cannot cope with new attributes and complex scenes is solved, and the robustness and adaptability of a pedestrian attribute recognition system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian attribute recognition, and in particular to a pedestrian attribute recognition method based on open vocabulary adaptive positioning. Background Art

[0002] With the rapid development of intelligent video surveillance, public safety systems, and smart transportation, pedestrian attribute recognition (PAR) technology has garnered widespread attention as a key technology. Pedestrian attribute recognition aims to leverage deep learning models to quickly and accurately identify various attributes of target pedestrians, including gender, age, clothing, and belongings, from large-scale image or video data based on specific query criteria (such as pedestrian images, descriptive text, or attribute features). This technology is of great significance for applications such as criminal investigation, pedestrian behavior understanding, and pedestrian tracking. Pedestrian attribute text is manually pre-roughly located based on empirical experience. Current Pedestrian Attribute Recognition technology only considers single-modal information input and is unable to integrate different types of information. This renders trained models ineffective in new scenarios.

[0003] Existing pedestrian attribute recognition systems have many shortcomings. On the one hand, most systems rely on single-modal information, such as images, text descriptions, or attribute features. Single-modal methods make insufficient use of information in complex scenes. For example, when the image is blurred or occluded, the performance of systems that rely solely on visual information will drop significantly, and text descriptions or attribute features alone are also difficult to fully characterize the target pedestrian. On the other hand, existing pedestrian attribute recognition technology can only recognize predefined fixed attributes and cannot handle new unknown attributes that appear in actual scenes, which limits the applicability and scalability of the system.

[0004] In recent years, a variety of innovative methods have emerged in the fields of human pose estimation, pedestrian retrieval, and behavior recognition. These methods have different core objectives and implementation paths, often striking different balances between model optimization, real-time performance, and accuracy. First, a real-time, lightweight human pose estimation method based on OpenPose introduces Bezier curves to optimize joint motion trajectories, improving the robustness of dynamic pose recognition and making it suitable for motion analysis scenarios. However, this method has shortcomings in utilizing deep keypoint information in static images. Second, a real-time, lightweight 2D human pose estimation method optimizes the model's real-time performance and inference speed by streamlining the network structure, making it suitable for resource-constrained devices. However, it is limited in fine-grained attribute recognition and has difficulty in deeply mining keypoint information. For pedestrian retrieval, a text-driven pedestrian retrieval method based on SAM uses preset cue points and a cross-modal loss function to align text and images, improving retrieval accuracy. However, its fixed template design limits its generalization ability in complex scenes. In terms of human action recognition, the human action recognition method based on multi-stream three-dimensional adaptive graph convolution improves the accuracy of action recognition by modeling spatiotemporal features through three-dimensional GCN, but does not consider the fine-grained analysis of pedestrian attributes, making it difficult to meet the needs of multimodal attribute retrieval. The visual text attribute alignment method based on semantic segmentation improves the accuracy of cross-modal matching by segmenting pedestrian body parts and aligning them with text attributes, but it relies too much on manually divided semantic regions and has difficulty processing the global semantic association between images and text. Finally, the real-time sign language intelligent recognition method focuses on the real-time recognition of sign language movements based on the spatiotemporal attention mechanism and graph convolutional network. Although it has achieved good results in the field of sign language parsing, the application scenario of this method is significantly different from pedestrian attribute analysis, and it cannot be directly transferred to the tasks of pedestrian action recognition or attribute analysis.

[0005] Therefore, to address the problem of handling newly emerged unknown attributes in real-world scenarios and bridge the gap between traditional person re-identification (PDR) techniques and practical applications, the Towards Open Vocabulary Pedestrian Attribute Recognition (POAR) framework was proposed. The POAR algorithmic framework uses an image encoder and a text encoder as its foundational models for open-vocabulary Pedestrian Attribute Recognition. The image encoder segments the input image into image blocks and uses a multi-head attention mechanism to extract global visual features, outputting a high-dimensional feature vector. Simultaneously, the text encoder converts attribute descriptions into semantic features and uses a multi-head attention mechanism to model semantic relationships within the text, generating an embedding representation aligned with the visual features. The visual and text features are then aligned using a cross-modal encoder, and the multi-head attention mechanism models the relationship between the two modalities, ensuring optimized representation within a unified feature space. Furthermore, to enhance semantic consistency between the visual and text encoders and improve the recognition of unseen attributes, POAR introduces a knowledge distillation mechanism. During training, a visual encoder pre-trained with CLIP (Contrastive Language-Image Pre-Training) is used as the teacher model to guide POAR's image encoder in learning a feature distribution that better matches the text encoder. The knowledge distillation mechanism uses the Kullback-Leibler (KL) divergence loss to bring the visual features generated by the student model closer to those of the teacher model, thereby narrowing the distribution gap between the visual and textual modalities and improving cross-modal matching capabilities. The feature vectors output by the final model are classified and tokenized for attribute prediction. Attribute prediction uses the cross-modal alignment loss (MTMC Loss) to optimize the matching ability of the two modalities by minimizing the cosine distance between visual and textual features.

[0006] Compared to traditional attribute recognition methods, POAR, by introducing multimodal information fusion and optimization, can handle attribute descriptions in open vocabulary, while also improving the accuracy and robustness of attribute recognition. This parameter-free cross-modal alignment loss function optimization effect is significant and can converge quickly, making it suitable for complex open scene pedestrian attribute recognition tasks.

[0007] Current pedestrian attribute recognition technology mainly relies on a single modality (such as images, text descriptions or attribute labels), and has obvious limitations in information expression and adaptability. Single-modality methods are difficult to fully utilize the diversity and detailed features among pedestrians. For example, the performance of a system that relies solely on image modality drops significantly under blur, occlusion or complex backgrounds, and using text or attribute modality alone is also difficult to fully characterize the target pedestrian, resulting in insufficient information utilization and poor system robustness. In addition, existing open pedestrian attribute recognition still relies on experience to perform rough positioning of the image first. When the attribute distribution deviates from the training data, it is difficult to achieve accurate alignment of attribute text features and image part features based on the rough positioning of the current method, which limits the performance of the model for open vocabulary pedestrian attribute recognition.

[0008] Among the current mainstream technical solutions, OpenPose is an open source real-time multi-person pose estimation system that can detect and locate key points of the human body, hands, and face. Although the pose estimation algorithm based on OpenPose has achieved remarkable results in optimizing pedestrian detection speed and reducing computational overhead, it generally lacks in-depth analysis and application of skeletal key point information. SAM (Segment Anything Model) is a general image segmentation model that aims to achieve high-precision segmentation of any object in the image by accepting multiple types of input prompts (such as points, bounding boxes, or coarse masks). However, the solution based on the SAM algorithm is limited by the static characteristics of the preset prompt points, and has the defect of insufficient accuracy in locating pedestrian images in complex backgrounds. Summary of the Invention

[0009] To address the technical challenges of existing pedestrian attribute recognition, which suffer from insufficient accuracy and generalizability, this paper proposes a pedestrian attribute recognition method based on open vocabulary adaptive positioning. This method adjusts the existing pedestrian attribute recognition framework to meet two new requirements: first, building a cross-modal pedestrian attribute recognition framework, and second, introducing open vocabulary attribute recognition technology, enabling the system to better cope with complex and changing real-world environments. This paper improves the accuracy of pedestrian attribute recognition in complex environments and addresses the limitations of existing technologies in single-modality and unknown attribute recognition.

[0010] To achieve the above object, the technical solution of the present invention is implemented as follows: a pedestrian attribute recognition method based on open vocabulary adaptive positioning, the steps of which are as follows:

[0011] Step 1: preprocess the original image in the image dataset to obtain a preprocessed image, and extract the text features of the input prompt text;

[0012] Step 2: Use OpenPose-based key point extraction technology to extract the skeleton key points of all pedestrians in the image dataset, obtain the predicted heat map of each joint point, determine the key point information set with the help of partial affinity fields, and perform center point calculation to select the target pedestrian closest to the image center;

[0013] Step 3: Using the text features of the prompt, the key point information set is screened to obtain prompt points, and the prompt points are combined with the key point information set to generate the final key point set;

[0014] Step 4: Detect the pedestrians in the preprocessed image and obtain the hint box closest to the image center. Convert the final set of key points into a hint vector. The SAM algorithm uses the hint vector and hint box to segment the target pedestrian in the preprocessed image and obtain a segmentation mask for the entire pedestrian. The entire segmentation mask is combined with the preprocessed image to generate a posture region localization map.

[0015] Step 5: Multimodal fusion and recognition: The posture region localization map and the prompt text describing the pedestrian are simultaneously input into the cross-modal open vocabulary pedestrian attribute recognition module to obtain the specific attributes of the pedestrian.

[0016] Preferably, the preprocessing in step 1 and the SAM algorithm in step 4 are both implemented by using an image encoder of a CLIP model of an open vocabulary pedestrian attribute recognition module;

[0017] The method for extracting the text features of the input prompt text in step 1 and converting the final key point set into the prompt vector in step 4 are both implemented by using the text encoder of the CLIP model of the open vocabulary pedestrian attribute recognition module.

[0018] The image encoder is based on the Vision Transformer image encoder, including a 12-layer Transformer architecture, which uses multiple attribute tokens to represent different attributes in the image, and each attribute token corresponds to a different area of ​​the image;

[0019] The text encoder converts the input prompt text into a fixed-dimensional representation, and the image encoder maps the input original image to the corresponding same embedding space through linear transformation, so that its dimension is consistent with the text features;

[0020] Contrastive learning is used to maximize the similarity between preprocessed image and text feature pairs in the same embedding space and minimize the similarity between false pairs.

[0021] Preferably, the method for determining the key point information set is: inputting the pre-processed image into the key point extraction module based on OpenPose, extracting the feature map through the C' layer convolutional neural network In the last layer of convolutional neural network, the initial predicted heat map of each key point is generated based on the extracted feature map F. The multiple convolution blocks and convolutions of the multi-level convolutional neural network process the preprocessed image I, predict the direction vector between each pair of joint points, calculate the partial affinity field, obtain the connection relationship of the key points of the human body, and perform multiple iterations to integrate them into the complete skeleton points of a person. The text features of the prompt text and the generated initial predicted heat map are input into the CLIP model, and the self-attention mechanism of the Transformer model is used to capture the attributes of the prompt text according to the text features, focusing on the local part of the pedestrian image to obtain the final predicted heat map. Get the key point set K=[k1,k2,...,k N ], where k i =[k i1 ,k i2 ,....,k iQ ] represents the key point set of the i-th person, and represents the jth key point information of the i-th person, represents the key point k ij The x-coordinate of represents the key point k ij The y coordinate of represents the key point k ij The confidence of is obtained through the final prediction heat map, N represents the number of pedestrian images in the input dataset, Q is the total number of preset human key points, and each key point number j∈Q corresponds to a human body part.

[0022] Preferably, the center prompt point module is implemented by: calculating the center point of the pre-processed image I; calculating the geometric center of each pedestrian according to the key point coordinates; calculating the Euclidean distance d between the geometric center of the pedestrian and the center point of the pre-processed image I. i ; According to the Euclidean distance d i Select the target pedestrian closest to the center of the image.

[0023] Preferably, the method for generating the final key point set is:

[0024] ④ The text features of the prompt text t p Convert to semantic vector z S , each key point k ij The information is converted into a feature vector Calculate eigenvectors With the semantic vector z S The relevance weight wij ;

[0025] ⑤Find effective key points: retain the relevance weight higher than the threshold τ ij =τ c (1+αw ij )’s key point k ij is a valid key point, where α and τ c is the correlation parameter;

[0026] ⑥Generate prompt points: Generate the prompt points of the i-th pedestrian:

[0027]

[0028] Among them, x 颈部 、x 右髋 、x 左髋 、x 颈部 、y 颈部 、y 右髋 、y 左髋 are the x- and y-coordinates of the valid key points respectively;

[0029] Get the cue point set of the image dataset in, represents the set of cue points of the i-th pedestrian;

[0030] Gather cue points Add to the key point information set K = [k1, k2, ..., k N ], updated to the final set of key points

[0031] Preferably, the implementation method of step 4 is as follows: the pre-processed image I is input into the image encoder, and after being processed by the multi-layer Transformer, the multi-level feature is embedded in the prompt encoder of the SAM algorithm according to the final key point information set. Generate a hint vector p; add a segmentation hint point module to the hint encoder. The mask decoder of the SAM algorithm combines the extracted feature embeddings f1-f4, the hint vector p, and the hint box closest to the image center to generate one or more segmentation masks for the pedestrian area and a confidence score for each mask. The mask with the highest confidence score is selected as the segmentation mask map M.

[0032] The segmentation mask image M is directly overlaid on the pixels with a pixel value of 1 in the preprocessed image I, and the color image containing only the target area is obtained as the posture area positioning image I segmented .

[0033] Preferably, the size is H raw ×W raw ×3 original image Input into the image encoder, scaled to the target size H×W, and normalized to obtain the preprocessed image

[0034] Set the prompt text T p Input into the text encoder, the attribute label is first combined with the prompt template to form a sentence, and then converted into a text token through linear transformation, and input together with the position code into the text encoder to generate the text feature of the prompt. D represents the dimension of the feature vector after the text prompt is converted;

[0035] The prompt template is set in the text encoder. Position encoding gives each text token a position information, and the order relationship of attribute tokens is clarified through position encoding.

[0036] The implementation method of the contrastive learning is: calculating the cosine similarity Determine the degree of match between the two; where |||| represents the L2 norm; calculate the contrast loss function through cosine similarity to maximize the similarity between correct visual-text pairs while minimizing the similarity between irrelevant visual-text pairs.

[0037] The loss function of the predicted heat map is: Among them, H true (x,y) represents the real key point heat map;

[0038] The loss function of the direction vector is: Where V i true (x,y) represents the true direction vector based on manual annotation, V i pred (x,y) is the direction vector predicted by the multi-level convolutional neural network;

[0039] The mask decoder is implemented based on the Transformer structure + multi-layer attention mechanism.

[0040] Preferably, the method for obtaining the specific attributes of the pedestrian is:

[0041] Step S1: Position the posture area in the image I segmented Input the image encoder of the open vocabulary pedestrian attribute recognition module, and extract the pedestrian's visual feature representation v through the image encoder poar ;

[0042] Step S2: Convert the attribute categories of the prompt text describing the pedestrian into natural language sentences and input them into the text encoder for encoding to generate text embedding t;

[0043] Step S3: Calculate the visual feature representation v of the imagepoar The similarity between the text embedding t and the posture region positioning map I is determined based on the similarity segmented The most matching attribute description is used to realize attribute identification.

[0044] Preferably, the cross-modal alignment capability is optimized by distillation learning, where the teacher model uses the CLIP pre-trained image encoder to generate stable visual features v teacher , the student model uses the image encoder of the open vocabulary pedestrian attribute recognition module to generate features v student , gradually approaching the visual features v of the teacher model through distillation learning during the training process teacher distribution to improve the cross-modal feature alignment capability.

[0045] Preferably, the image encoder of the open vocabulary pedestrian attribute recognition module positions the posture region in the image I segmented Divide the image into fixed-size patches and convert them into embedding vectors. Use multiple learnable attribute tokens to focus on the visual features of specific parts, and use the Transformer's attention mechanism to make each attribute token focus only on a specific part of the image.

[0046] The image encoder uses a multi-layer Transformer to model global information. At each layer, the Transformer uses a self-attention mechanism to calculate the similarity between image patches and image patches, and between image patches and attribute tokens to determine which areas require more attention.

[0047] The method for generating the text embedding t is as follows: the input text feature T is converted into attribute tokens using the Byte-Pair Encoding method; the attribute tokens are input into the Transformer for multi-layer self-attention calculation, and each attribute token is respectively calculated using different weight matrices to represent the information of the query, key, and value; after the multi-layer Transformer, a sequence containing the semantic information of each attribute token is finally output, and an [EOS] token is added at the end of the sequence to represent the semantics of the entire sentence, which is the text embedding t;

[0048] The similarity in step S3 is achieved by using dot product calculation;

[0049] MTMC loss function is used Optimize the semantic alignment performance of the image encoder and text encoder of the open vocabulary pedestrian attribute recognition module: the visual-to-text contrastive loss is:

[0050]

[0051] The contrastive loss from text to vision is:

[0052]

[0053] Among them, t b represents the b-th text embedding vector, v a represents the ath visual embedding vector, τ represents the temperature parameter in contrastive learning; the superscript + represents the embedding of positive samples that match the text or vision; T represents the matrix transpose; G represents the total number of text embeddings in the batch; G a Representation and visual embedding vector v a The corresponding number of positive sample text embeddings; Q1 represents the number of images in the batch, Q b Represents the text embedding vector t b The corresponding number of positive visual embeddings;

[0054] The loss function of the distillation learning is Indicates the calculation of the KL divergence loss between the teacher model and the student model.

[0055] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention focuses on a cross-modal pedestrian attribute recognition framework, mainly solving the problem that it is difficult to achieve semantic alignment between open vocabulary text and images, constructs an open vocabulary adaptive positioning framework, and achieves more accurate attribute recognition through multimodal feature fusion. The present invention innovatively designs an Adaptive Coupling Localization Module, which generates dynamic prompt points by deeply fusing the SAM algorithm and the OpenPose posture estimation algorithm, and combines the key point screening mechanism to achieve accurate spatial positioning of open attribute prompts, combined with semantic understanding. The present invention inputs the text and the regional image into the text encoder and the image encoder respectively, and realizes open vocabulary-image semantic alignment through training through the loss function. This multimodal fusion strategy effectively solves the problem of inaccurate alignment of text features and visual features in traditional methods, thereby significantly improving the recognition performance of the system.

[0056] The present invention uses an adaptive coupled positioning module combined with pose estimation and semantic segmentation technology to accurately locate pedestrian areas and extract key point information. The adaptive coupled positioning module is optimized using the OpenPose algorithm and the SAM algorithm, combined with cue point screening, and can still accurately locate pedestrians in complex backgrounds and occluded scenes. Next, the image and text information are input into the image encoder and text encoder respectively, and cross-modal feature loss is used to ensure semantic alignment between the image and the open vocabulary text. The proposed method can handle unknown attributes not seen in the training phase and optimizes them using cross-modal matching loss, thereby improving the ability to capture fine-grained attributes. It demonstrates strong generalization and high accuracy in dynamically changing scenes, solving the problem that traditional methods are unable to cope with new attributes and complex scenes, and improving the robustness and adaptability of pedestrian attribute recognition systems.

[0057] The present invention uses open vocabulary attribute recognition technology, extracting visual features through an image encoder and combining it with a text encoder to generate semantic embeddings, thereby achieving cross-modal alignment. With the help of a cross-modal matching loss, it can identify unknown attributes not seen during the training phase and enhance the recognition ability of fine-grained attributes, thereby demonstrating stronger generalization and higher accuracy in dynamically changing scenarios. The present invention is applicable to applications such as criminal investigation, pedestrian behavior understanding, and pedestrian tracking. For example, it can assist security agencies in supervision and criminal investigation work, quickly locate target pedestrians under a large number of video surveillance cameras, and feedback their location and pedestrian-related attribute feature information, thereby greatly reducing the waste of time and resources caused by manual investigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0059] Figure 1 Flowchart of the present invention.

[0060] Figure 2 for Figure 1 The figure shows a schematic diagram of the adaptive coupled positioning module processing pedestrian images.

[0061] Figure 3 for Figure 1 Detailed flow chart of the adaptive coupling positioning module is shown.

[0062] Figure 4 Schematic diagram of the multimodal alignment loss of the present invention.

[0063] Figure 5 Schematic diagram of the model attention areas for different pedestrian attributes according to the present invention.

[0064] Figure 6 This is a test result diagram of the present invention based on text search image on the public dataset PETA.

[0065] Figure 7 This is a schematic diagram of the human body parts corresponding to the key point information proposed in the present invention.

[0066] Figure 8 Diagram of the distillation learning framework proposed in this invention. DETAILED DESCRIPTION

[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0068] like Figure 1 As shown in the figure, the present invention proposes a pedestrian attribute recognition method based on open vocabulary adaptive positioning, which combines the two methods of posture estimation and semantic segmentation to achieve comprehensive analysis and recognition of pedestrian attributes. First, the adaptive coupled positioning module is used to estimate the posture of the input pedestrian image, and the key point extraction module based on OpenPose is used to generate the corresponding predicted heat map and partial affinity fields (PAFs). The PAFs are then extracted by applying a specific loss function (such as the confidence loss function L conf and partial affinity field loss function L PAF ) is optimized to extract the key point information of the human body. Secondly, the cue point screening module is used to process all the key point information, and segmentation cue points are dynamically generated according to the key point information. The SAM module is used to accurately segment the pedestrian area, and the image encoder of the open vocabulary pedestrian attribute recognition module is used to generate image embedding. The cue encoder is combined with the convolution module to generate an accurate segmentation mask. Finally, the image encoder and text encoder are used to perform cross-modal feature fusion on the segmented pedestrian area. By introducing the cross-modal matching loss (MTMCLoss), as shown in Figure 4 As shown in the figure, the semantic alignment performance of visual encoding and text encoding is further optimized to improve the accuracy of pedestrian attribute recognition, thereby achieving accurate recognition of pedestrian attributes (such as gender, age group, clothing color, etc.). The overall architecture of the system is as follows Figure 1This invention not only addresses the significant limitations of existing pedestrian attribute recognition technologies in terms of information utilization, adaptability, and scalability, but also demonstrates greater robustness, flexibility, and accuracy in practical application scenarios. The proposed framework enables efficient training and testing on a single NVIDIA GeForce RTX 3090 GPU.

[0069] The present invention aims to achieve accurate identification of pedestrian attributes by integrating multimodal information (such as posture features, semantic information and visual features). The specific steps of the present invention are:

[0070] Step 1: Preprocess the original image in the image dataset to obtain a preprocessed image, and extract the text features of the input prompt text.

[0071] The overall outline of the present invention is as follows: First, an adaptive coupled localization module is used to locate the local area of ​​pedestrians in the image by receiving the prompt text. Then, a cross-modal open-vocabulary pedestrian attribute recognition module is used to achieve cross-modal alignment of the localized pedestrian image region features and attribute text features, and identify attributes that were not seen during the training phase. Finally, a distillation learning method is used to transfer the generalization performance of CLIP to the proposed model, thereby achieving accurate recognition of open-vocabulary attributes, such as Figure 1 shown.

[0072] The adaptive coupling positioning module consists of a key point extraction module, an image segmentation module, a cue point screening module, a center cue point module, and a segmentation cue point module. The key point extraction module is connected to the cue point screening module and the center cue point module respectively. The cue point screening module and the segmentation cue point module are both connected to the image segmentation module. The preprocessed image is input into the key point extraction module to extract the key point set, and the center cue point module is used to ignore the interference of non-core characters. Secondly, the cue point screening module is combined with the prompt text to generate cue point information and improve the key point set. Then, the image and the key point set are input into the image segmentation module, and the segmentation cue point module is used to constrain the segmentation area to generate an accurate character mask. At the same time, the image and character mask are merged to generate a posture area positioning map, such as Figure 2 and Figure 3As shown in the figure, the OpenPose module first performs pose estimation on the input pedestrian image dataset and extracts a set of human keypoint information. Secondly, the keypoint screening module combines the prompt text and the keypoint information set to generate cue point information to improve the keypoint information set. Then, the image segmentation module based on the SAM algorithm segments the pedestrian region, uses its image encoder to generate image embeddings, and combines the keypoint information set with the convolution module to generate a precise segmentation mask. This is then combined with the input image to generate a pose region localization map. Finally, the image encoder and text encoder of the open vocabulary pedestrian attribute recognition module (POAR) are used to perform cross-modal feature alignment on the segmented pedestrian region, thereby accurately identifying pedestrian attributes (such as gender, age group, clothing color, etc.). The adaptive coupled localization module extracts the pedestrian's skeletal keypoint information and dynamically generates task-related cue points to accurately locate the target area. This design enhances the ability to capture local details of pedestrians, especially when pedestrian features are blurred or occluded in the detection image, demonstrating higher robustness.

[0073] The pre-processing method is: the original image Is a size H raw ×W raw ×3 RGB image is input into the image encoder of the convolutional neural network (CNN), the image is scaled to the target size of 224×224, and the preprocessed image is obtained by standardization. Facilitates subsequent processing.

[0074] Set the prompt text T p Input into the text encoder based on the Transformer model, the attribute label is first combined with the prompt template to form a sentence, and then converted into a text token through a linear transformation. Together with the position encoding, it is input into the Transformer model to generate the text features of the prompt. This facilitates subsequent processing. D represents the dimension of the feature vector after the text prompt is converted.

[0075] The image encoder and text encoder used in the above application are derived from the CLIP model, which is currently the most popular and classic image-text processing model.

[0076] The image encoder of the present invention adopts a Transformer model with a 12-layer Transformer architecture. Unlike the commonly used Transformer model which has only one classification token, multiple classification tokens are used here for attribute recognition. The traditional Transformer model has only one classification token, such as the [CLS] token in ViT, which is used to represent the features of the entire image. The image encoder of the present invention uses multiple attribute tokens (set as [ATT] token, which is a learnable token in the following text, a total of 8) to represent different attributes in the image (such as upper body, lower body, accessories, hair, etc., such as Figure 4 Each attribute token corresponds to a different area of ​​the image, helping the Transformer model better focus on different parts of the pedestrian, thereby performing more fine-grained attribute recognition.

[0077] Attribute labels are predefined. The public dataset contains both images and attribute labels describing them. Prompt templates are set in the text encoder; they are included with the public dataset. Each image has a matching description, which is categorized as a natural language description + label. Positional encoding assigns a position to each text token, clarifying the order of tokens. The text encoder uses sinusoidal positional encoding, which is automatically generated when text is input into the text encoder.

[0078] The input prompt text is converted into a fixed-dimensional representation. The input visual features are mapped to the same embedding space through a linear transformation, making their dimensions consistent with the text features. The text tokens are processed by the text encoder to output high-dimensional feature vectors; the visual tokens are processed by the image encoder and can be used for cross-modal alignment learning.

[0079] Use contrastive learning to maximize the correct pre-processed image I and text features t p The similarity between pairs in the shared embedding space is calculated by calculating the cosine similarity. Determine the degree of matching between the two, where |||| represents the L2 norm. p , guiding the CLIP model to focus on image regions relevant to the prompt, thereby increasing the confidence in these regions. The contrastive loss function is calculated using cosine similarity to maximize the similarity between correct visual-text pairs while minimizing the similarity between unrelated visual-text pairs. Through this mapping and contrastive learning, the CLIP model achieves consistent embedding of text and images.

[0080] Step 2: Use OpenPose-based key point extraction technology to extract the skeletal key points of all pedestrians in the image dataset, obtain the predicted heat map of each joint point, and use partial affinity fields to determine the key point information set. Then, perform center point calculation to select the target pedestrian closest to the image center.

[0081] Obtain image key point information: Input the images in the image dataset into the key point extraction module based on OpenPose, and extract feature maps through the C' layer convolutional neural network In the last layer of convolutional neural network, the initial prediction heat map H of each joint point (i.e. key point) is generated based on the extracted feature map F. pred Keypoints usually correspond to important joints of the human body, such as shoulders, elbows, wrists, knees, ankles, head, neck, etc., see Table 1.

[0082] Table 1 Number-part comparison table: the names and serial numbers of the extracted key points

[0083]

[0084] Figure 1 The key point extraction module uses a dual-branch multi-level CNN architecture. The first part is responsible for generating the initial prediction heat map; the second part is responsible for extracting some affinity fields, that is, the connection relationship between the key points of the human body, and integrating them into the complete skeleton points of a character through multiple iterations. The text features of the prompt text and the generated initial prediction heat map are input into the CLIP model, and the self-attention mechanism of the Transformer model is used to pay more attention to the semantic area and improve the accuracy of key point prediction. The self-attention mechanism is used to allow the text encoder to capture the attributes of the prompt text after receiving the text features, to focus on the local part of the pedestrian image, and to obtain the final prediction heat map. Represents the predicted heatmap for keypoint i at position (x, y). For example, if we input "This person has short black hair," the text encoder will capture the attribute of "short hair" and focus more on the token [hair], which is the area in the upper half of the image.

[0085] The loss function for calculating the predicted heatmap is used to guide the network to optimize the key point detection accuracy. The loss function for calculating the confidence is as follows: Among them, H true (x,y) represents the real key point heat map, which is used to supervise the convolutional neural network to learn the standard reference of human key points, obtained from the image dataset with manually annotated key points.

[0086] From the prediction heatmap H predA series of coordinates are known, but the specific joint names represented by each coordinate are unknown. To more accurately predict the spatial relationship between each pair of joint points, partial affinity fields (PAFs) are calculated. The preprocessed image I is processed through multiple convolutional blocks and convolutions of a multi-stage CNN to predict the direction vector between each pair of joint points. The direction vector represents the direction and strength of the connection between the two key points. Compared with the fuzzy information of the predicted heat map, the direction vector can more accurately represent the relationship between key points. The loss function for calculating the confidence is as follows: Where V i true (x,y) represents the true direction vector based on manual annotation, V i pred (x,y) is the direction vector predicted by multi-level CNN.

[0087] The key point extraction module uses OpenPose technology to obtain the key point location and confidence score. The confidence score is obtained through the network output of the model. When generating the prediction heat map, the pixel value on the prediction heat map represents the probability of whether a certain area has the joint. The higher the value, the more likely the location is to be the joint. Finally, the key point set K = [k1, k2, ..., k N ], where k i =[k i1 ,k i2 ,....,k iQ ] represents the key point set of the i-th person, specifically expressed as Respectively represent the j-th key point information of the i-th person, represents the key point k ij The x-coordinate of represents the key point k ij The y coordinate of Indicates its confidence, N represents the number of pedestrian images in the input dataset, and Q is the total number of preset human key points. Each key point number j∈Q corresponds to a human body part, as shown in Table 1, such as 0 for the nose, 1 for the neck, etc. Figure 7 The numbered positions in the skeleton key point diagram.

[0088] In actual scenes, complex environments (such as insufficient light, occlusion, and blur) may affect pedestrian detection. When there are multiple pedestrians in the image, directly using OpenPose for key point detection will detect the key point information of multiple individuals. Therefore, a center cue point module is set up to select the target pedestrian closest to the center of the image among all detected pedestrians to help OpenPose focus only on the most central pedestrian at a time, and only on one person. The role of the center cue point module is to solve the problem of multiple pedestrian interference. The specific implementation method is as follows:

[0089] 1. Calculate the center point of the preprocessed image I H and W are the width and height of the image, respectively, both are 224.

[0090] 2. Use the calculation of the geometric center of each pedestrian Among them, μ i represents the center point coordinates of the i-th pedestrian, k i is the key point set of the i-th person, is the coordinate information of the jth key point of the i-th pedestrian. This formula shows that the average of the x and y coordinates of all the key points of the i-th person is calculated to represent the center point of the pedestrian.

[0091] 3. Calculate the Euclidean distance between the pedestrian and the image center: μ i,x ,μ i,y Represents the x, y coordinates of the geometric center of the i-th pedestrian, μ image,x and μ image,y Represent the x and y coordinates of the geometric center of the preprocessed image I respectively.

[0092] 4. Select the target pedestrian closest to the center of the image.

[0093] Step 3: Use the text features of the prompt to filter the key points of the key point information set to obtain prompt points. The prompt points and the key point information set are combined to generate the final key point set.

[0094] Because the information quality of the key point information set K is not high enough, the cue point screening module will screen out low-quality key points and generate additional cue points based on the existing key point information as new key points.

[0095] The prompt point screening module uses the obtained key point information set K and prompt words to generate the prompt point set K p ,like Figure 3 shown.

[0096] The specific process is as follows:

[0097] ⑦ Detect the correlation between text and prompt points: The text features of the prompt t p Convert to semantic vector z S , each key point k ij The information is converted into a feature vector

[0098] Calculate the semantic relevance weight of each key point and the prompt:

[0099]

[0100] ⑧Find effective key points: retain weights above the threshold τ ijkey points.

[0101] τij=τ c (1+αw ij ), where τ c =0.5,α=0.3.

[0102] only When k is the key point ij is an effective key point.

[0103] ⑨Generate cue points: Generate cue points for the i-th pedestrian using the following formula:

[0104]

[0105] Among them, x 颈部 、x 右髋 、x 左髋 、x 颈部 、y 颈部 、y 右髋 、y 左髋 are the x-coordinate and y-coordinate of the valid key points respectively.

[0106] The final set of cue points after processing N pedestrian images is in, A set of cue points suitable for representing the i-th pedestrian.

[0107] Then we will set the cue points Add to the previously extracted key point information set K = [k1, k2, ..., k N ], then the updated is the final set of key points.

[0108] Step 4: Perform target detection on the pedestrians in the preprocessed image to obtain the prompt box closest to the image center, convert the final key point set into a prompt vector, use the SAM algorithm, prompt vector, and prompt box to segment the target pedestrians in the preprocessed image to obtain the entire pedestrian segmentation mask map; combine the entire pedestrian segmentation mask map with the preprocessed image to generate a posture area positioning map.

[0109] The preprocessed image I and the final key point information set The image segmentation module is input to locate and segment the local attributes of pedestrians. The image segmentation module is built based on the SAM algorithm and includes a prompt encoder and a mask encoder. The pre-processed image I is input to the image encoder based on the Vision Transformer. After multi-layer Transformer processing, feature embeddings f1-f4 are obtained. Multi-level feature embeddings f1-f4 are extracted from the input pre-processed image I, which are the outputs of different levels of the image encoder. The prompt encoder is based on the final key point information set. Generate a prompt vector p. Finally, use the mask decoder to segment the pedestrians in the image and generate a mask. Simultaneously, the present invention adds a segmentation prompt point module to the prompt encoder to help SAM better segment the person. The segmentation prompt point module detects pedestrians in the image using the YOLO object detection algorithm, selects the pedestrian at the center of the image, generates a prompt box, and passes it to the prompt encoder, which can better optimize the segmentation effect.

[0110] The mask decoder combines the feature embedding f1-f4 extracted by the image encoder and the hint vector p converted by the hint encoder, and then adds the segmentation hint point module to obtain the hint box closest to the center of the image. It generates one or more segmentation masks of the pedestrian area and the confidence score of each mask, and selects the mask with the highest confidence score as the segmentation mask map. The mask decoder is implemented based on the Transformer structure + multi-layer attention mechanism.

[0111] Because pedestrian images have low pixel counts and complex scenes, OpenPose may extract overlapping keypoint information (such as two heads), which can lead to incorrect segmentation of overlapping areas of multiple bodies during segmentation. Therefore, a segmentation hint module is implemented to constrain the boundaries of the target object and ensure more accurate segmentation. This is just a small module that helps the SAM algorithm perform initial processing; keypoint information will be combined later for further segmentation.

[0112] Get the posture area positioning map: I segmented =I⊙M, ⊙ represents the element-by-element product. That is, the segmentation mask image M is directly overlaid on the pixels with a pixel value of 1 (or other values ​​representing the target area) in the preprocessed image I, and the pixel values ​​of the corresponding positions in the preprocessed image I are replaced with their pixel values ​​to obtain a color image containing only the target area, which is the posture area positioning image I. segmented ,like Figure 1 shown.

[0113] Step 5: Multimodal fusion and recognition: Positioning the posture region in Figure I segmented The text features of the prompt describing the pedestrian are simultaneously input into the cross-modal open vocabulary pedestrian attribute recognition module POAR to obtain the specific attributes of the pedestrian.

[0114] The specific implementation steps are:

[0115] Step S1: Position the posture area in the image I segmented The image encoder of the open vocabulary pedestrian attribute recognition module POAR is input, and the visual feature representation v of the pedestrian is extracted through the image encoder. poar The image encoder is the encoder to be trained, which is based on the ViT (Vision Transformer) model. The image encoder takes the input image I segmented Divide into fixed-size patches and convert them into embedding vectors. Use multiple learnable attribute tokens (such as "head", "upper body", "lower body", etc.) Figure 5 The embedding vector is pulled into one dimension, which can be understood as a 1-row n-column format. For example, a 224x224 pixel image can be divided into 14x14 patches, flattened into a one-dimensional vector of 1 row and 196 columns. Then 0-13 represents the first row of the image, 14-27 represents the second row of the image, and so on. Then the vectors 0-97 represent the upper part of the image, and 98-195 represent the lower part of the image. According to the sequence number of the embedding vector, the posture area positioning diagram I segmented Then set multiple learnable attribute tokens and use the Transformer's attention mechanism (selectively focusing on the most relevant information by calculating the relationship between different images) to let each token focus only on a specific part of the image. For example, [head] only focuses on the upper half of the patches, because the heads of pedestrians in the image are often in the upper half of the image, so [head] only focuses on the upper half of the patches. Similarly, [lower body] only focuses on the lower half of the patches. Its attention range is as follows Figure 5 As shown in the figure, the white blocks are where the Transformer focuses its attention. Different attribute tokens focus on specific areas (e.g., the [head] token focuses only on the head), reducing interference and focusing on the visual features of specific parts.

[0116] The image encoder uses a multi-layer Transformer to model global information. At each layer, the Transformer uses a self-attention mechanism to calculate the similarity between image blocks and image blocks, and between image blocks and attribute tokens to determine which areas need more attention. This process is carried out in each layer of the Transformer, so as the number of layers increases, the model can gradually capture higher-level global information. After the Transformer calculation, the final visual feature extracted from each attribute token is represented as v poar . Figure 5Represents the visual attention area in the image encoder based on the ViT model in the attribute recognition module, and the attention area of ​​different attribute tokens in the image.

[0117] Figure 4 I1-I N Represents different pedestrian image samples, which is the posture area positioning map I obtained in step 5 segmented ,These images will be input into the image encoder for feature extraction, ,converted into high-dimensional feature vectors, and cross-modally aligned with the text ,features generated by the text encoder.,The visual features extracted from the image are used to judge the ,degree of match between the image and the text.

[0118] Step S2: Prompt text T describing the pedestrian p Attribute categories are converted into natural language sentences and input into a text encoder for encoding, generating a text embedding t. This is achieved by using the Byte-Pair encoding method to convert the input prompt text T into tokens. The tokens are then input into a Transformer for multi-layer self-attention calculations. Each token is passed through a different weight matrix to calculate the information representations Q (query), K (key), and V (value). Each Transformer layer calculates a new token representation, and each token "sees" the entire input sequence, capturing semantic information. This generates a text embedding t. The multiple layers of Transformers above ultimately output a sequence containing the semantic information of each token. At the end of this sequence, a special [EOS] (End of Sentence) token represents the semantics of the entire sentence, called the text embedding t, which is used for subsequent visual-text matching. The image encoder is the ViT (Vision Transformer) model, and the text encoder is based on the text encoder of the CLIP model.

[0119] Step S3: Calculate the similarity between the visual embedding of the image and the attribute text embedding. The specific calculation method is as follows: the image encoder extracts the visual embedding features corresponding to different attribute areas of the input image, and the text encoder encodes the text descriptions of multiple attribute categories into text embedding features. The dot product is used to calculate the similarity between the visual embedding and the text embedding. Based on the similarity, the most matching attribute description in the input image is determined to achieve attribute recognition, such as Figure 4 shown.

[0120] This paper uses the Many-to-Many Contrastive (MTMC) loss function to optimize the semantic alignment performance of the image encoder and the text encoder:

[0121] The temperature parameter τ is set to 5, so that the similarity between the visual embedding and the corresponding text embedding is as high as possible, and the similarity with the non-corresponding text embedding is as low as possible. Then the visual-to-text contrast loss is:

[0122]

[0123] The same goal is to maximize the similarity between text embedding and corresponding visual embedding, and minimize the similarity with non-corresponding visual embedding. The text-to-visual contrast loss is:

[0124]

[0125] Among them, t b represents the b-th text embedding vector, v a represents the a-th visual embedding vector, and τ represents the temperature parameter in contrastive learning, which is set to 5 here. Represents the dot product similarity between text embedding and visual embedding. Superscript +: represents the embedding of the positive sample that matches text t. Superscript T represents the transpose of the matrix. G represents the number of texts in the batch. G a Representation and visual embedding vector v a The number of corresponding positive sample text embeddings. Q represents the number of images in the batch. b Represents the text embedding vector t b The corresponding number of positive visual embeddings.

[0126] Then the MTMC loss function is:

[0127] Training phase: By calculating the similarity between image embeddings and text embeddings, increasing the similarity for correctly matched sample pairs and reducing the similarity for incorrectly matched sample pairs, the model's representational capabilities are optimized so that visual and text embeddings can be aligned in the same semantic space. Batch construction: Each batch contains multiple images and their corresponding attribute texts; the image encoder and text encoder are used to extract visual embeddings and text embeddings respectively; all image-text pairs are constructed, and the dot product is calculated to form a similarity matrix. The MTMC loss function is applied to "closer" the similarity of positive sample pairs and "difference" the similarity of negative sample pairs; the encoder parameters are optimized through backpropagation. Testing phase: The visual embedding of the input image is calculated and the similarity is calculated with all possible text descriptions. The text description with the highest similarity is selected as the predicted attribute of the image.

[0128] Step S4: In order to improve the generalization ability of the image encoder in the open attribute scenario, a distillation learning mechanism is introduced, such as Figure 8As shown in Figure 2, distillation learning is used to optimize cross-modal alignment capabilities. Distillation learning is a model compression technique that allows a small model (i.e., student model) to learn the knowledge of a large model (i.e., teacher model) to reduce computational costs and improve inference efficiency while maintaining model performance as much as possible. The teacher model uses a CLIP pre-trained visual encoder to generate stable visual features v teacher The student model uses POAR’s visual encoder to generate features v student , through distillation learning during the training process, gradually approaching the visual features v of the teacher model teacher distribution to improve the cross-modal feature alignment capability.

[0129] The distillation losses are as follows: Calculate the KL divergence loss between the teacher model and the student model, and optimize the generated feature v student , making it closer to the visual feature v teacher , ensuring that POAR can better understand open vocabulary properties.

[0130] Figure 6 After inputting the natural language description text, find the pedestrian images that meet the description conditions in the PETA dataset. Figure 6 It can be seen that the effect of the present invention is that pedestrian images that meet the description can be found from the PETA dataset based on the input text (natural language sentence).

[0131] Some image segmentation tasks can be achieved through other alternatives (such as SoloV2), but there is still a significant gap in performance and applicability compared to the SAM (Segment Anything Model) algorithm. SoloV2 can directly segment pedestrian areas, but its segmentation quality depends on image clarity and a fixed prior framework, and its ability to recognize pedestrians in complex backgrounds or dynamic scenes is limited. The SAM algorithm, on the other hand, combines an image encoder and a hint encoder to not only generate accurate segmentation masks using key points of the human body, but also dynamically adjust the segmentation range to adapt to the diversity of pedestrians in different scenarios, and perform better in occlusion and blur.

[0132] Furthermore, SAM is more compatible with multimodal fusion frameworks. By integrating key points generated by OpenPose, the SAM algorithm's segmentation accuracy is effectively improved through a prompting mechanism, and it seamlessly integrates with the subsequent POAR attribute recognition module. However, SoloV2 lacks multimodal support, making it difficult to meet the demands of subsequent tasks. In summary, the SAM algorithm significantly outperforms SoloV2 in dynamic adaptability, segmentation accuracy, and multimodal support, making it the optimal choice for pedestrian attribute recognition and identification in complex environments.

[0133] Different from existing methods, this paper focuses on the problem of unknown attribute recognition in complex environments. The contributions of this paper are mainly reflected in the following aspects:

[0134] 1. Adaptive positioning and optimization: In response to the rough processing problem of attribute text and image positioning in the existing technology, the present invention designs an adaptive coupled positioning module. By combining the OpenPose posture estimation algorithm and the SAM semantic segmentation algorithm, it can accurately locate the pedestrian area in the image and dynamically adjust the positioning effect according to different prompts. Because pedestrian images are affected by factors such as posture, lighting, and occlusion, the state of pedestrians in each image is different. The SAM and OpenPose models are robust for pedestrian-related processing. Here, OpenPose is used to generate relatively stable skeleton point information, and prompts are used to identify the corresponding positions. Then, the parts are located according to the segmentation map of the SAM algorithm; the whole process acts alternately, thereby realizing the automatic positioning of pedestrians according to images with different attribute texts. This innovative design significantly improves the ability to recognize pedestrian features in blurred, occluded or complex backgrounds.

[0135] 2. Open vocabulary cross-modal alignment: Traditional pedestrian attribute recognition methods often rely on fixed attribute descriptions or modal inputs, and are difficult to deal with situations with unknown attributes. The cross-modal alignment framework proposed in the present invention realizes open vocabulary attribute recognition through multimodal information fusion. The framework can automatically adjust in unknown attribute scenarios. Specifically, an adaptive coupling positioning module is first used to locate the local area of ​​the pedestrian in the image by receiving the prompt language. It is worth mentioning that the proposed adaptive coupling positioning module does not require training and can accurately locate attributes that have not been seen in the training set; then, a cross-modal open vocabulary pedestrian attribute recognition module is used to achieve cross-modal alignment of the located pedestrian image area features and attribute text features; finally, a distillation learning method is used to transfer the generalization performance of CLIP to the proposed model, thereby achieving accurate recognition of open vocabulary attributes.

[0136] Through these innovations, the present invention demonstrates strong robustness and adaptability in complex real-world scenarios, addresses the shortcomings of existing methods in complex environments and new attribute recognition, and further promotes the advancement of pedestrian attribute recognition technology.

[0137] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A pedestrian attribute recognition method based on open vocabulary adaptive positioning, characterized in that: The steps are as follows: Step 1: preprocess the original image in the image dataset to obtain a preprocessed image, and extract the text features of the input prompt text; Step 2: Use OpenPose-based key point extraction technology to extract the skeleton key points of all pedestrians in the image dataset, obtain the predicted heat map of each joint point, determine the key point information set with the help of partial affinity fields, and perform center point calculation to select the target pedestrian closest to the image center; Step 3: Using the text features of the prompt, the key point information set is screened to obtain prompt points, and the prompt points are combined with the key point information set to generate the final key point set; Step 4: Detect the pedestrians in the preprocessed image and obtain the hint box closest to the image center. Convert the final set of key points into a hint vector. The SAM algorithm uses the hint vector and hint box to segment the target pedestrian in the preprocessed image and obtain a segmentation mask for the entire pedestrian. The entire segmentation mask is combined with the preprocessed image to generate a posture region localization map. Step 5: Multimodal fusion and recognition: The posture region localization map and the prompt text describing the pedestrian are simultaneously input into the cross-modal open vocabulary pedestrian attribute recognition module to obtain the specific attributes of the pedestrian.

2. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 1 is characterized in that: The implementation method of the preprocessing in step 1 and the SAM algorithm in step 4 for extracting multi-level feature embedding from the preprocessed image is both implemented by using the image encoder of the CLIP model of the open vocabulary pedestrian attribute recognition module; The method for extracting the text features of the input prompt text in step 1 and converting the final key point set into the prompt vector in step 4 are both implemented by using the text encoder of the CLIP model of the open vocabulary pedestrian attribute recognition module; The image encoder is based on the Vision Transformer image encoder, including a 12-layer Transformer architecture, which uses multiple attribute tokens to represent different attributes in the image, and each attribute token corresponds to a different area of ​​the image; The text encoder converts the input prompt text into a fixed-dimensional representation, and the image encoder maps the input original image to the corresponding same embedding space through linear transformation, so that its dimension is consistent with the text features; Contrastive learning is used to maximize the similarity between preprocessed image and text feature pairs in the same embedding space and minimize the similarity between false pairs.

3. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 1 or 2, characterized in that: The method for determining the key point information set is as follows: inputting the preprocessed image into the key point extraction module based on OpenPose, and extracting the feature map through the C' layer convolutional neural network In the last layer of convolutional neural network, the initial predicted heat map of each key point is generated based on the extracted feature map F. The multiple convolutional blocks and convolutions of the multi-level convolutional neural network process the preprocessed image I, predict the direction vector between each pair of joint points, calculate the partial affinity field, obtain the connection relationship of the key points of the human body, and perform multiple iterations to integrate them into the complete skeleton points of a person. The text features of the prompt text and the generated initial predicted heat map are input into the CLIP model. The self-attention mechanism of the Transformer model is used to capture the attributes of the prompt text based on the text features, focusing on the local part of the pedestrian image, and obtaining the final predicted heat map H. pred (x, y); get the key point set K = [k1, k2, ..., k N ], where k i =[k i1 ,k i2 ,....,k iQ ] represents the key point set of the i-th person, and represents the jth key point information of the i-th person, represents the key point k ij The x-coordinate of represents the key point k ij The y coordinate of represents the key point k ij The confidence of is obtained through the final prediction heat map, N represents the number of pedestrian images in the input dataset, Q is the total number of preset human key points, and each key point number j∈Q corresponds to a human body part.

4. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 3 is characterized in that: The implementation method of the center prompt point module is as follows: calculate the center point of the preprocessed image I; calculate the geometric center of each pedestrian according to the key point coordinates; calculate the Euclidean distance d between the geometric center of the pedestrian and the center point of the preprocessed image I i ; According to the Euclidean distance d i Select the target pedestrian closest to the center of the image.

5. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 3 is characterized in that: The method for generating the final key point set is: ① The text feature t of the prompt text p Convert to semantic vector z S , each key point k ij The information is converted into a feature vector Calculate eigenvectors With the semantic vector z S The relevance weight w ij ; ② Find effective key points: retain the relevance weight above the threshold τ ij =τ c (1+αw ij )’s key point k ij is a valid key point, where α and τ c is the correlation parameter; ③Generate prompt points: Generate the prompt points of the i-th pedestrian: Among them, x 颈部 、x 右髋 、x 左髋 、x 颈部 、y 颈部 、y 右髋 、y 左髋 are the x- and y-coordinates of the valid key points respectively; Get the cue point set of the image dataset in, represents the set of cue points of the i-th pedestrian; Gather cue points Add to the key point information set K = [k1, k2, ..., k N ], updated to the final set of key points 6. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 5 is characterized in that: The implementation method of step 4 is as follows: the pre-processed image I is input into the image encoder, and after being processed by the multi-layer Transformer, the multi-level feature embedding in the SAM algorithm is obtained. The prompt encoder is based on the final key point information set. Generate a hint vector p; add a segmentation hint point module to the hint encoder. The mask decoder of the SAM algorithm combines the extracted feature embeddings f1-f4, the hint vector p, and the hint box closest to the image center to generate one or more segmentation masks for the pedestrian area and a confidence score for each mask. The mask with the highest confidence score is selected as the segmentation mask map M. The segmentation mask image M is directly overlaid on the pixels with a pixel value of 1 in the preprocessed image I, and the color image containing only the target area is obtained as the posture area positioning image I segmented .

7. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 6, characterized in that: Size H raw ×W raw ×3 original image Input into the image encoder, scaled to the target size H×W, and normalized to obtain the preprocessed image Set the prompt text T p Input into the text encoder, the attribute label is first combined with the prompt template to form a sentence, and then converted into a text token through linear transformation, and input together with the position code into the text encoder to generate the text feature of the prompt. D represents the dimension of the feature vector after the text prompt is converted; The prompt template is set in the text encoder. Position encoding gives each text token a position information, and the order relationship of attribute tokens is clarified through position encoding. The implementation method of the contrastive learning is: calculating the cosine similarity Determine the degree of match between the two; where |||| represents the L2 norm; calculate the contrast loss function through cosine similarity to maximize the similarity between correct visual-text pairs while minimizing the similarity between unrelated visual-text pairs; The loss function of the predicted heat map is: Among them, H true (x,y) represents the real key point heat map; The loss function of the direction vector is: Where V i true (x,y) represents the true direction vector based on manual annotation, V i pred (x,y) is the direction vector predicted by the multi-level convolutional neural network; The mask decoder is implemented based on the Transformer structure + multi-layer attention mechanism.

8. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to any one of claims 4 to 7, characterized in that: The method for obtaining the specific attributes of the pedestrian is: Step S1: Position the posture area in the image I segmented Input the image encoder of the open vocabulary pedestrian attribute recognition module, and extract the pedestrian's visual feature representation v through the image encoder poar ; Step S2: Convert the attribute categories of the prompt text describing the pedestrian into natural language sentences and input them into the text encoder for encoding to generate text embedding t; Step S3: Calculate the visual feature representation v of the image poar The similarity between the text embedding t and the posture region positioning map I is determined based on the similarity segmented The most matching attribute description is used to realize attribute identification.

9. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 8, characterized in that: The cross-modal alignment capability is optimized through distillation learning, where the teacher model uses the CLIP pre-trained image encoder to generate stable visual features v teacher , the student model uses the image encoder of the open vocabulary pedestrian attribute recognition module to generate features v student , gradually approaching the visual features v of the teacher model through distillation learning during the training process teacher distribution to improve the cross-modal feature alignment capability.

10. The pedestrian attribute recognition method based on open vocabulary adaptive positioning according to claim 9, characterized in that: The image encoder of the open vocabulary pedestrian attribute recognition module locates the posture region in the image I segmented Divide the image into fixed-size patches and convert them into embedding vectors. Use multiple learnable attribute tokens to focus on the visual features of specific parts, and use the Transformer's attention mechanism to make each attribute token focus only on a specific part of the image. The image encoder uses a multi-layer Transformer to model global information. At each layer, the Transformer uses a self-attention mechanism to calculate the similarity between image patches and image patches, and between image patches and attribute tokens to determine which areas require more attention. The method for generating the text embedding t is as follows: the text features of the prompt text are converted into attribute tokens using the Byte-Pair Encoding method; the attribute tokens are input into the Transformer for multi-layer self-attention calculation, and each attribute token is respectively calculated using different weight matrices to represent the information of the query, key, and value; after the multi-layer Transformer, a sequence containing the semantic information of each attribute token is finally output, and an [EOS] token is added at the end of the sequence to represent the semantics of the entire sentence, which is the text embedding t; The similarity in step S3 is achieved by using dot product calculation; MTMC loss function is used Optimize the semantic alignment performance of the image encoder and text encoder of the open vocabulary pedestrian attribute recognition module: the visual-to-text contrastive loss is: The contrastive loss from text to vision is: Among them, t b represents the b-th text embedding vector, v a represents the ath visual embedding vector, τ represents the temperature parameter in contrastive learning; the superscript + represents the embedding of positive samples that match the text or vision; T represents the matrix transpose; G represents the total number of text embeddings in the batch; G a Representation and visual embedding vector v a The corresponding number of positive sample text embeddings; Q1 represents the number of images in the batch, Q b Represents the text embedding vector t b The corresponding number of positive visual embeddings; The loss function of the distillation learning is Indicates the calculation of the KL divergence loss between the teacher model and the student model.