Vision-language bimodal-based multi-label pedestrian attribute identification method and system
By adopting the visual-language bimodal multi-label recognition method in pedestrian attribute recognition technology, combining image chunking and dual-modal feature aggregation, the accuracy and robustness of pedestrian attribute recognition in complex environments is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202411930194.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-09
AI Technical Summary
Existing pedestrian attribute recognition technology is difficult to accurately identify pedestrian attributes in complex environments (such as motion blur, shadow, occlusion, low resolution, multi-view and nighttime, etc.), and insufficient mining of deep semantic information of attribute phrases affects the recognition accuracy.
The multi-label pedestrian attribute recognition method based on vision-language dual mode is adopted, and visual features are processed through image chunking and specific encoder, combined with text feature processing of fixed and dynamic prompts, and the dual-modal feature aggregation is realized by operations such as feature cascade and full connection layer, and the binary cross entropy is optimized as a loss function.
In complex environments, the accuracy and robustness of pedestrian attribute recognition can be significantly improved, and pedestrian attribute recognition can be more accurately and comprehensively identified, improving recognition accuracy.
Smart Images

Figure CN119964232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pedestrian attribute recognition, and in particular to a multi-label pedestrian attribute recognition method and system based on vision-language dual modality. Background Art
[0002] As a key technology, pedestrian attribute recognition (PAR) is committed to using advanced technologies such as deep learning, especially convolutional neural networks (CNN) and recurrent neural networks (RNN), to accurately capture and analyze the multi-dimensional feature information of pedestrians from complex scenes. Although this field has made significant progress driven by intelligent technology, its accuracy and robustness are still facing severe challenges in practical applications, such as motion blur in dynamic scenes, shadow interference caused by changes in ambient light, object occlusion, insufficient image resolution, diversity of viewing angles, and difficulty in night imaging. These challenges require researchers to continuously explore innovative methods and optimize algorithm design to improve the performance of PAR systems in complex environments and ensure accurate recognition and efficient analysis of pedestrian attributes.
[0003] Person Attribute Recognition (PAR) aims to predict attribute information of a person by extracting semantic information. With the help of artificial intelligence, such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), this research field has received widespread attention and made great progress. However, it remains challenging due to the poor imaging quality in extreme conditions, including motion blur, shadows, occlusions, low resolution, multiple views, and nighttime. Existing methods rarely mine the semantic information of attribute phrases, which is very important for attribute recognition. The relationship between semantic attributes and visual representations has not been fully explored and utilized. Existing works usually adopt CNN for feature learning and only encode local relations, but some attribute information heavily relies on long-range pixel relationship modeling.
[0004] In the field of pedestrian attribute recognition, a significant shortcoming is that the deep semantic information of attribute phrases is not fully mined, which is particularly important in improving recognition accuracy. The close connection between semantic attributes and visual representations, as a key link connecting high-level semantics with low-level visual features, has not been fully explored and utilized. The current mainstream methods mostly use convolutional neural networks (CNNs) for feature learning. This method focuses on the encoding of local visual patterns, but it is limited when dealing with attributes that are highly dependent on the global structure of the image or the interaction between remote pixels. Specifically, the recognition of attributes such as posture, clothing style, or behavioral patterns in a specific context often requires comprehensive consideration of the complex interactions between widely distributed visual elements in the image, rather than just being limited to local areas. Therefore, it is difficult to fully capture the global, structural, and semantic characteristics of these attributes by relying solely on the local feature extraction capabilities of CNNs. Summary of the invention
[0005] In order to solve the above technical problems existing in the prior art, the present invention proposes a multi-label pedestrian attribute recognition method and system based on vision-language dual modality to solve the above technical problems.
[0006] According to a first aspect of the present invention, a multi-label pedestrian attribute recognition method based on vision-language dual modality is proposed, comprising:
[0007] S1: Receives image x in the dataset i , the image x i Divide the image into multiple image blocks, add learnable markers to the image blocks, and combine them into the input of the image encoder ViT to extract visual feature information;
[0008] S2: prepare the input of the text encoder, including fixed prompts and dynamic prompts;
[0009] S3: Send the fixed prompts and dynamic prompts to the text encoder for text information feature extraction, and use feature cascade method, through a fully connected layer and a ReLU activation function, to obtain the final text feature information;
[0010] S4: The fully connected layer transforms the dimension of text feature information into the same dimension as the visual feature, and uses feature cascading to aggregate the visual feature information and text feature information. The information is sent to the encoding layer of the image encoder ViT for refined fusion, and pedestrian attributes are predicted based on the fused features.
[0011] In some specific embodiments, in S1, the image is divided into Z image blocks of size P*P*C, Z = (H*W) / (P*P), wherein Z represents the number of image blocks obtained after the image is divided, P represents the pixel size of each image block in the horizontal and vertical directions, C represents the number of channels of the image, H represents the height of the original input image, and W represents the width of the original input image.
[0012] In some specific embodiments, the fixed prompt is "This photo contain[class]" and the dynamic prompt is "a pedestrian with a XXXX[class]", where "XXXX" is a trainable parameter used to refine the text information describing the attributes in each picture based on the visual features encoded in the image.
[0013] In some specific embodiments, the entire pedestrian attribute recognition is a multi-label classification task, and the binary cross entropy is used as the loss function, which is expressed as: in, Represents the image x i The predicted probability value of the jth attribute in , r j is the proportion of positive samples of the jth attribute in the training set.
[0014] In some specific embodiments, CLIP is used as the backbone network.
[0015] According to a second aspect of the present invention, a computer-readable storage medium is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above method is implemented.
[0016] According to a third aspect of the present invention, a multi-label pedestrian attribute recognition system based on vision-language dual modality is proposed, comprising:
[0017] Image input and preprocessing unit, configured to receive images x in the dataset i , the image x i Divide the image into multiple image blocks, add learnable markers to the image blocks, and combine them into the input of the image encoder ViT to extract visual feature information;
[0018] A text encoder input building unit configured to prepare an input of a text encoder, including fixed prompts and dynamic prompts;
[0019] A text feature extraction and fusion unit is configured to send the fixed prompt words and the dynamic prompt words to the text encoder to extract text information features, and adopt a feature cascade method, through a fully connected layer and a ReLU activation function, to obtain the final text feature information by fusion;
[0020] The visual-linguistic modality feature aggregation and prediction unit is configured to transform the text feature information dimension into the same dimension as the visual feature through a fully connected layer, aggregate the visual feature information and the text feature information using a feature cascade method, and send them to the encoding layer of the image encoder ViT for refined fusion, and predict pedestrian attributes based on the fused features.
[0021] In some specific embodiments, the image input and preprocessing unit divides the image into Z image blocks of size P*P*C, Z = (H*W) / (P*P), wherein Z represents the number of image blocks obtained after the image is divided, P represents the pixel size of each image block in the horizontal and vertical directions, C represents the number of channels of the image, H represents the height of the original input image, and W represents the width of the original input image.
[0022] In some specific embodiments, the fixed prompt is "This photo contain[class]" and the dynamic prompt is "a pedestrian with a XXXX[class]", where "XXXX" is a trainable parameter used to refine the text information describing the attributes in each picture based on the visual features encoded in the image.
[0023] In some specific embodiments, the entire pedestrian attribute recognition is a multi-label classification task, and the binary cross entropy is used as the loss function, which is expressed as: in, Represents the image x i The predicted probability value of the jth attribute in , r j is the proportion of positive samples of the jth attribute in the training set.
[0024] In some specific embodiments, CLIP is used as the backbone network.
[0025] The present invention proposes a multi-label pedestrian attribute recognition method and system based on vision-language bimodality, which integrates visual and language bimodal information, uses image segmentation and a specific encoder to process visual features, combines text feature processing with fixed and dynamic prompts, and uses feature cascades, fully connected layers and other operations to achieve bimodal feature aggregation, while using binary cross entropy as the loss function for optimization. This method can effectively deal with the problem of intra-class attribute changes in multi-label pedestrian attribute recognition, and can more accurately and comprehensively identify pedestrian attributes in complex environments (such as motion blur, shadows, occlusions, low resolution, multiple views and night imaging) compared to existing technologies, thereby improving the accuracy and robustness of pedestrian attribute recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and together with the description are used to explain the principles of the present invention. Other embodiments and many expected advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. Other features, objects and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments made with reference to the following drawings:
[0027] Figure 1 is a flowchart of a multi-label pedestrian attribute recognition method based on vision-language dual modality according to an embodiment of the present application;
[0028] Figure 2This is a framework diagram of a multi-label pedestrian attribute recognition algorithm based on vision-language dual modality according to an embodiment of the present application;
[0029] Figure 3 This is an architecture diagram of a multi-label pedestrian attribute recognition system based on vision-language dual modality according to an embodiment of the present application;
[0030] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION
[0031] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0032] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0033] Figure 1 FIG. 1 shows a flow chart of a multi-label pedestrian attribute recognition method based on vision-language dual modality according to an embodiment of the present application. Figure 1 As shown, the method comprises the following steps:
[0034] S1: Receives image x in the dataset i , the image x i The image is divided into multiple image blocks, and learnable markers are added to the image blocks, which are combined into the input of the image encoder ViT to extract visual feature information.
[0035] In a specific embodiment, the image is divided into Z image blocks of size P*P*C, Z = (H*W) / (P*P), wherein Z represents the number of image blocks obtained after the image is divided, P represents the pixel size of each image block in the horizontal and vertical directions, C represents the number of channels of the image, H represents the height of the original input image, and W represents the width of the original input image.
[0036] S2: Prepare the input of the text encoder, including fixed prompts and dynamic prompts. The fixed prompt is "This photo contain [class]" and the dynamic prompt is "a pedestrian with a XXXX [class]", where "XX XX" is a trainable parameter used to refine the text information describing the attributes in each image based on the visual features encoded in the image.
[0037] S3: The fixed prompts and dynamic prompts are sent to the text encoder for text information feature extraction. The feature cascade method is adopted, and after a fully connected layer and a ReLU activation function, the final text feature information is fused.
[0038] In a specific embodiment, the entire pedestrian attribute recognition is a multi-label classification task, and the binary cross entropy is used as the loss function, which is expressed as: in, Represents the image x i The predicted probability value of the jth attribute in , r j is the proportion of positive samples of the jth attribute in the training set.
[0039] S4: The fully connected layer transforms the dimension of text feature information into the same dimension as the visual feature, and uses feature cascading to aggregate the visual feature information and text feature information. The information is sent to the encoding layer of the image encoder ViT for refined fusion, and pedestrian attributes are predicted based on the fused features.
[0040] Continue to refer Figure 2 , Figure 2 FIG. 1 shows a framework diagram of a multi-label pedestrian attribute recognition algorithm based on vision-language dual modality according to an embodiment of the present application. Figure 2 As shown, the whole method mainly includes an image encoder (taking ViT as an example), a text encoder (taking ViT as an example), text feature fusion and visual-language modality feature aggregation. In order to realize a multi-label pedestrian attribute recognition algorithm based on visual-language dual modality, the present invention mainly includes the following steps:
[0041] Step S1: Receive any image x in the dataset D i As input, the purpose is to identify multiple pedestrian attributes. The image x i The corresponding label is y i ∈{0,1} M , M represents the number of categories of pedestrian attributes, and the corresponding y i The image is divided into Z blocks of size P*P*C, where Z = (H*W) / (P*P), P represents the pixel size of each block in the horizontal and vertical directions, C represents the number of channels of the image, H represents the height of the original input image, and W represents the width of the original input image; plus a learnable marker, it is combined into the input of the image encoder ViT Extract visual feature information.
[0042] Step S2: The entire pedestrian attribute recognition problem can be regarded as a multi-label classification task, using binary cross entropy (BCELoss) as the loss function. The loss function is expressed as follows: in, To predict the value Mapped to a probability value between 0 and 1, Represents the image x i The predicted probability value of the jth attribute in r j is the proportion of positive samples of the jth attribute in the training set, and N represents the total number of images in the dataset.
[0043] Step S3: The input of the text encoder mainly includes two prompts, one is a fixed prompt "This photo contains [class]", and the other is a dynamic prompt "a pedestrian with a XXXX [class]". Fixed prompts are used to learn the common feature information between attributes and are not trainable.
[0044] Step S4: Since the same attribute may have different appearances, which will cause large changes within the class, the fixed prompt template cannot obtain refined semantic information. Therefore, this patent introduces dynamic prompts, and the initialized prompt template is: "a pedestrian with a XXXX[class]", where "XXXX" is a trainable parameter, which is used to refine the text information of the attribute in each picture according to the visual features of the image encoding.
[0045] Step S5: Send both the fixed prompt and the dynamic prompt to the text encoder to extract text information features and fuse them using the text modality feature fusion module, that is, use a feature cascade method, then pass through a fully connected layer and a ReLU activation function to obtain the final text feature information.
[0046] Step S6: In order to make the two modal feature information be used together for the final attribute prediction, this solution uses an encoding layer of ViT to aggregate the visual-linguistic modal feature information. First, a fully connected layer is used to transform the dimension of the text feature information obtained in step S5 into the same dimension as the visual feature, and then the visual feature information and the text feature information are aggregated by feature cascading, and then sent to the encoding layer of ViT for refined fusion.
[0047] Figure 3 FIG. 1 shows an architecture diagram of a multi-label pedestrian attribute recognition system based on vision-language dual modality according to an embodiment of the present application. Figure 3As shown, the system includes an image input and preprocessing unit 301, a text encoder input construction unit 302, a text feature extraction and fusion unit 303, and a visual-language modality feature aggregation and prediction unit 304. The image input and preprocessing unit 301 is configured to receive an image x in a data set. i , the image x i The image is divided into multiple image blocks, and learnable markers are added to the image blocks to form the input of the image encoder ViT to extract visual feature information; the text encoder input construction unit 302 is configured to prepare the input of the text encoder, including fixed prompts and dynamic prompts; the text feature extraction and fusion unit 303 is configured to send the fixed prompts and dynamic prompts to the text encoder for text information feature extraction, and adopts a feature cascade method, through a fully connected layer and a ReLU activation function, to obtain the final text feature information through fusion; the visual-language modality feature aggregation and prediction unit 304 is configured to transform the text feature information dimension into the same dimension as the visual feature through a fully connected layer, and adopt a feature cascade method to aggregate the visual feature information and the text feature information, and send them to the encoding layer of the image encoder ViT for refined fusion, and predict pedestrian attributes based on the fused features.
[0048] The present application designs a multi-label classification pedestrian attribute recognition algorithm based on vision-language bimodality. The design scheme mainly introduces information modeling of language modality to obtain rich semantic information, and uses CLIP as the backbone network to extract bimodal feature information of vision and language. At the same time, in multi-label tasks, the same attribute category may have different appearance forms (i.e., attribute changes within the class). In addition to fixed prompts, the present invention also designs a dynamic prompt to deal with the problem of large category changes in multi-label tasks.
[0049] In a specific embodiment, the model is evaluated using multi-label classification evaluation indicators such as accuracy, recall, and F1-score. On the PETA dataset, compared with other single-modal or non-bimodal fusion and refined feature processing pedestrian attribute recognition algorithms, the method of this application has significantly improved the accuracy and robustness of pedestrian attribute recognition in complex environments (such as occlusion, low resolution, etc.). For example, the accuracy rate is increased from 70% to 80%, and the F1-score is increased from 0.65 to 0.75.
[0050] Reference below Figure 4 , which shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. Figure 4 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0051] like Figure 4 As shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage part 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0052] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage section 408 as needed.
[0053] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.
[0054] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0055] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0056] The modules involved in the embodiments of the present application may be implemented by software or by hardware.
[0057] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device: receives an image x in a data set; i , the image x iThe image is divided into multiple image blocks, and learnable markers are added to the image blocks to form the input of the image encoder ViT to extract visual feature information; the input of the text encoder is prepared, including fixed prompts and dynamic prompts; the fixed prompts and dynamic prompts are sent to the text encoder for text information feature extraction, and the final text feature information is obtained by fusion through a fully connected layer and a ReLU activation function using a feature cascade method; the text feature information dimension is transformed into the same dimension as the visual feature through a fully connected layer, and the visual feature information and text feature information are aggregated using a feature cascade method, and sent to the encoding layer of the image encoder ViT for refined fusion, and pedestrian attributes are predicted based on the fused features.
[0058] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A multi-label pedestrian attribute recognition method based on vision-language dual modality, characterized in that: include: S1: Receives image x in the dataset i , the image x i Divide the image into multiple image blocks, add learnable markers to the image blocks, and combine them into the input of the image encoder ViT to extract visual feature information; S2: prepare the input of the text encoder, including fixed prompts and dynamic prompts; S3: sending the fixed prompt and the dynamic prompt into the text encoder to extract text information features, using a feature cascade method, passing through a fully connected layer and a ReLU activation function, and fusing to obtain final text feature information; S4: The dimension of the text feature information is converted into the same dimension as the visual feature through the fully connected layer, the visual feature information and the text feature information are aggregated by feature cascading, and sent to the encoding layer of the image encoder ViT for fusion, and pedestrian attributes are predicted based on the fused features.
2. The multi-label pedestrian attribute recognition method based on vision-language dual modality according to claim 1 is characterized in that: In S1, the image is divided into Z image blocks of size P*P*C, Z=(H*W) / (P*P), wherein Z represents the number of image blocks obtained after the image is divided, P represents the pixel size of each image block in the horizontal and vertical directions, C represents the number of channels of the image, H represents the height of the original input image, and W represents the width of the original input image.
3. The multi-label pedestrian attribute recognition method based on vision-language dual modality according to claim 1 is characterized in that: The fixed prompt is "This photo contain[class]" and the dynamic prompt is "a pedestrian witha XXXX[class]", where "XXXX" is a trainable parameter used to refine the text information describing the attributes in each picture according to the visual features encoded in the image.
4. The multi-label pedestrian attribute recognition method based on vision-language dual modality according to claim 1 is characterized in that: The entire pedestrian attribute recognition is a multi-label classification task, using binary cross entropy as the loss function, which is expressed as: in, Represents the image x i The predicted probability value of the jth attribute in , r j is the proportion of positive samples of the jth attribute in the training set.
5. The multi-label pedestrian attribute recognition method based on vision-language dual modality according to claim 1 is characterized in that: CLIP is used as the backbone network.
6. A computer-readable storage medium having one or more computer programs stored thereon, characterized in that: When the one or more computer programs are executed by a computer processor, the method according to any one of claims 1 to 5 is implemented.
7. A multi-label pedestrian attribute recognition system based on vision-language dual modality, characterized by: include: Image input and preprocessing unit, configured to receive images x in the dataset i , the image x i Divide the image into multiple image blocks, add learnable markers to the image blocks, and combine them into the input of the image encoder ViT to extract visual feature information; A text encoder input building unit configured to prepare an input of a text encoder, including fixed prompts and dynamic prompts; A text feature extraction and fusion unit is configured to send the fixed prompt and the dynamic prompt to the text encoder to extract text information features, adopt a feature cascade method, pass through a fully connected layer and a ReLU activation function, and fuse to obtain the final text feature information; The visual-linguistic modality feature aggregation and prediction unit is configured to transform the text feature information dimension into the same dimension as the visual feature through the fully connected layer, aggregate the visual feature information and the text feature information by feature cascading, send them to the encoding layer of the image encoder ViT for refined fusion, and predict pedestrian attributes based on the fused features.
8. The multi-label pedestrian attribute recognition system based on vision-language dual modality according to claim 7 is characterized in that: The image input and preprocessing unit divides the image into Z image blocks of size P*P*C, where Z=(H*W) / (P*P), wherein Z represents the number of image blocks obtained after the image is divided, P represents the pixel size of each image block in the horizontal and vertical directions, C represents the number of channels of the image, H represents the height of the original input image, and W represents the width of the original input image.
9. The multi-label pedestrian attribute recognition system based on vision-language dual modality according to claim 7 is characterized in that: The fixed prompt is "This photo contain[class]" and the dynamic prompt is "a pedestrian witha XXXX[class]", where "XXXX" is a trainable parameter used to refine the text information describing the attributes in each picture according to the visual features encoded in the image.
10. The multi-label pedestrian attribute recognition system based on vision-language dual modality according to claim 7 is characterized in that: The entire pedestrian attribute recognition is a multi-label classification task, using binary cross entropy as the loss function, which is expressed as: in, Represents the image x i The predicted probability value of the jth attribute in , r j is the proportion of positive samples of the jth attribute in the training set.
11. The multi-label pedestrian attribute recognition system based on vision-language dual modality according to claim 7 is characterized in that: CLIP is used as the backbone network.