Method and device for realizing pedestrian attribute identification and evaluation based on multi-modal large model, processor and computer readable storage medium thereof
By using the collaborative work of multimodal large model and CLIP model in pedestrian attribute recognition, the shortcomings of traditional methods in accuracy and robustness are solved, and pedestrian attribute recognition with high accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510188165.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
The traditional pedestrian attribute recognition method has shortcomings in terms of accuracy and robustness, especially under the influence of different angles, postures, lighting conditions and occlusion factors, it is difficult to ensure the reliability of the recognition.
Using a method based on a multimodal large model, the accurate identification and evaluation of pedestrian attributes are achieved through the collaborative work of the human body recognition module, pedestrian attribute analysis module, output quality evaluation module and accuracy evaluation module. This method utilizes the powerful image resolution and natural language processing capabilities of multimodal large models, and combines the pre-trained CLIP model to evaluate output quality and accuracy.
It improves the accuracy and robustness of pedestrian attribute recognition, ensures the reliability and quality of identification results, and is suitable for application scenarios such as intelligent safety monitoring systems.
Smart Images

Figure CN120107894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, in particular to the field of computer vision and pattern recognition, and specifically refers to a method, device, processor and computer-readable storage medium thereof for realizing pedestrian attribute recognition and evaluation based on a multimodal large model. Background Art
[0002] With the development of artificial intelligence technology, especially the progress in computer vision, pedestrian attribute recognition has become an important research direction in many fields such as intelligent monitoring and safety supervision. Pedestrian attribute recognition refers to the automatic extraction of pedestrian appearance feature information from images or videos, such as gender, age, clothing color, and items carried. It is of great significance to improve the intelligence level of monitoring systems and achieve accurate crowd management and public safety management.
[0003] However, traditional pedestrian attribute recognition methods face many challenges. On the one hand, due to the diversity and complexity of pedestrian appearance, such as differences in angles, postures, lighting conditions, and occlusion, the accuracy of pedestrian attribute recognition is difficult to guarantee. On the other hand, traditional recognition algorithms often rely on manually designed features, which not only increases the complexity of the algorithm, but also limits the improvement of recognition performance. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a method, device, processor and computer-readable storage medium for pedestrian attribute recognition and evaluation based on a multimodal large model that has high reliability, good accuracy and a wide range of applications.
[0005] In order to achieve the above objectives, the method, device, processor and computer-readable storage medium for realizing pedestrian attribute recognition and evaluation based on a multimodal large model of the present invention are as follows:
[0006] The method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model has the following main features:
[0007] (1) The human body recognition module identifies and confirms whether there is a human body object in the input image and locates the human body target;
[0008] (2) The pedestrian attribute analysis module is fine-tuned based on the multimodal large model, analyzes the attributes of pedestrians in the image, generates descriptions in a predetermined format, and outputs text;
[0009] (3) The output quality evaluation module calculates the cosine similarity between the output text of the multimodal large model and the original image through the pre-trained pedestrian attribute image-text matching clip model;
[0010] (4) Accuracy evaluation module: checks the accuracy of pedestrian attributes in the model output text and obtains its detection accuracy.
[0011] Preferably, the step (1) specifically comprises the following steps:
[0012] (1.1) Directly extract pedestrian images using visual detection models;
[0013] (1.2) If pedestrian data is difficult to extract, the large model can be fine-tuned by building a self-built difficult sample pedestrian detection dataset.
[0014] Preferably, the step (2) specifically comprises the following steps:
[0015] (2.1) Determine the attributes of pedestrians that need to be identified;
[0016] (2.2) Use the self-built dataset method to construct a fine-tuning dataset, and output the text format of pedestrian images and their corresponding description texts in a unified manner;
[0017] (2.3) Fine-tune the large multimodal model and conduct subsequent evaluation on the output model after fine-tuning.
[0018] Preferably, the step (3) specifically comprises the following steps:
[0019] (3.1) Construct a clip model and train it on the same training data set to obtain a clip model that can simultaneously calculate the feature vectors of images and texts;
[0020] (3.2) Use the CLIP model to extract the feature vectors of the output text and the original image at the same time, and use the trained clip model to calculate the feature vector similarity of the image-text pair output by the multimodal large model to obtain the similarity score of the image-text pair;
[0021] (3.3) Simultaneously, the feature vector similarity is evaluated with the original data image-text pair to obtain the image-text pair similarity score;
[0022] (3.4) The image-text pair similarity score is used as the standard to calculate whether there is a deviation between the image-text pair similarity score output by the large model and the image-text pair similarity score. If so, the model output is judged to be a hallucination; otherwise, the model output is not a hallucination.
[0023] Preferably, the step (4) specifically comprises the following steps:
[0024] (4.1) Check whether the pedestrian attributes mentioned in the output text are consistent with the actual situation, evaluate whether the output text content of the large model is consistent with the original data label, and calculate the accuracy of the large model attribute recognition;
[0025] (4.2) Find the images where the large model attribute recognition errors occurred and save them in a separate folder.
[0026] The system for realizing pedestrian attribute recognition and evaluation based on a multimodal large model is characterized in that the system comprises:
[0027] A human body recognition module is used to identify and confirm whether there is a human body object in the input image and locate the human body target;
[0028] A pedestrian attribute analysis module, connected to the human body recognition module, is used to perform fine-tuning on the basis of the multimodal large model, analyze various attributes of pedestrians in the image, generate descriptions in a predetermined format, and output text;
[0029] An output quality evaluation module, connected to the pedestrian attribute analysis module, is used to calculate the cosine similarity between the output text of the multimodal large model and the original image through a pre-trained pedestrian attribute image-text matching clip model to verify the consistency and rationality of the output;
[0030] The accuracy assessment module is connected to the output quality evaluation module and is used to evaluate whether the text content output by the large model is consistent with the original data label to obtain its detection accuracy.
[0031] The device for realizing pedestrian attribute recognition and evaluation based on a multimodal large model is characterized in that the device comprises:
[0032] a processor configured to execute computer-executable instructions;
[0033] The memory stores one or more computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the method for realizing pedestrian attribute recognition and evaluation based on the multimodal large model are implemented.
[0034] The processor for realizing pedestrian attribute recognition and evaluation based on a multimodal large model is characterized in that the processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the above-mentioned method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model are realized.
[0035] The computer-readable storage medium is characterized in that a computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model.
[0036] The method, device, processor and computer-readable storage medium of the present invention for realizing pedestrian attribute recognition and evaluation based on a multimodal large model are adopted. Through the coordinated work of the human body recognition module, the pedestrian attribute analysis module, the output quality evaluation module and the accuracy evaluation module, not only the pedestrian attributes can be accurately identified, but also the quality of the recognition results can be effectively evaluated, thus ensuring the reliability and accuracy of the output, and providing strong technical support for the application and development of intelligent safety monitoring systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 The present invention is a flowchart of a method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model.
[0038] Figure 2 Schematic diagram of the working process of the human body recognition module in the present invention.
[0039] Figure 3 It is a schematic diagram of the workflow of the pedestrian attribute analysis module based on the multimodal large model in the present invention.
[0040] Figure 4 Schematic diagram of the workflow of the output quality assessment module based on the CLIP model in the present invention.
[0041] Figure 5 Schematic diagram of the workflow of the accuracy assessment module in the present invention. DETAILED DESCRIPTION
[0042] In order to more clearly describe the technical content of the present invention, further description is given below in conjunction with specific embodiments.
[0043] The method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model of the present invention comprises the following steps:
[0044] (1) The human body recognition module identifies and confirms whether there is a human body object in the input image and locates the human body target;
[0045] (2) The pedestrian attribute analysis module is fine-tuned based on the multimodal large model, analyzes the attributes of pedestrians in the image, generates descriptions in a predetermined format, and outputs text;
[0046] (3) The output quality evaluation module calculates the cosine similarity between the output text of the multimodal large model and the original image through the pre-trained pedestrian attribute image-text matching clip model;
[0047] (4) Accuracy evaluation module: checks the accuracy of pedestrian attributes in the model output text and obtains its detection accuracy.
[0048] As a preferred embodiment of the present invention, the step (1) specifically comprises the following steps:
[0049] (1.1) Directly extract pedestrian images using visual detection models;
[0050] (1.2) If pedestrian data is difficult to extract, the large model can be fine-tuned by building a self-built difficult sample pedestrian detection dataset.
[0051] As a preferred embodiment of the present invention, the step (2) specifically comprises the following steps:
[0052] (2.1) Determine the attributes of pedestrians that need to be identified;
[0053] (2.2) Use the self-built dataset method to construct a fine-tuning dataset, and output the text format of pedestrian images and their corresponding description texts in a unified manner;
[0054] (2.3) Fine-tune the large multimodal model and conduct subsequent evaluation on the output model after fine-tuning.
[0055] As a preferred embodiment of the present invention, the step (3) specifically comprises the following steps:
[0056] (3.1) Construct a clip model and train it on the same training data set to obtain a clip model that can simultaneously calculate the feature vectors of images and texts;
[0057] (3.2) Use the CLIP model to extract the feature vectors of the output text and the original image at the same time, and use the trained clip model to calculate the feature vector similarity of the image-text pair output by the multimodal large model to obtain the similarity score of the image-text pair;
[0058] (3.3) Simultaneously, the feature vector similarity is evaluated with the original data image-text pair to obtain the image-text pair similarity score;
[0059] (3.4) The image-text pair similarity score is used as the standard to calculate whether there is a deviation between the image-text pair similarity score output by the large model and the image-text pair similarity score. If so, the model output is judged to be a hallucination; otherwise, the model output is not a hallucination.
[0060] As a preferred embodiment of the present invention, the step (4) specifically comprises the following steps:
[0061] (4.1) Check whether the pedestrian attributes mentioned in the output text are consistent with the actual situation, evaluate whether the output text content of the large model is consistent with the original data label, and calculate the accuracy of the large model attribute recognition;
[0062] (4.2) Find the images where the large model attribute recognition errors occurred and save them in a separate folder.
[0063] The system for realizing pedestrian attribute recognition and evaluation based on a multimodal large model of the present invention comprises:
[0064] A human body recognition module is used to identify and confirm whether there is a human body object in the input image and locate the human body target;
[0065] A pedestrian attribute analysis module, connected to the human body recognition module, is used to perform fine-tuning on the basis of the multimodal large model, analyze various attributes of pedestrians in the image, generate descriptions in a predetermined format, and output text;
[0066] An output quality evaluation module, connected to the pedestrian attribute analysis module, is used to calculate the cosine similarity between the output text of the multimodal large model and the original image through a pre-trained pedestrian attribute image-text matching clip model to verify the consistency and rationality of the output;
[0067] The accuracy assessment module is connected to the output quality evaluation module and is used to evaluate whether the text content output by the large model is consistent with the original data label to obtain its detection accuracy.
[0068] The device for realizing pedestrian attribute recognition and evaluation based on a multimodal large model of the present invention, wherein the device comprises:
[0069] a processor configured to execute computer-executable instructions;
[0070] The memory stores one or more computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the method for realizing pedestrian attribute recognition and evaluation based on the multimodal large model are implemented.
[0071] The processor of the present invention is used to implement pedestrian attribute recognition and evaluation based on a multimodal large model, wherein the processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the above-mentioned method for implementing pedestrian attribute recognition and evaluation based on a multimodal large model are implemented.
[0072] The computer-readable storage medium of the present invention stores a computer program thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model.
[0073] In recent years, the rise of deep learning technology, especially multimodal large models, has provided new ideas for solving the above problems. Multimodal large models can process multiple types of data such as images and texts at the same time. Through large-scale data training, they have powerful image understanding and natural language processing capabilities. It is based on this background that the present invention proposes a pedestrian attribute recognition and evaluation method based on a multimodal large model, aiming to solve the problems of low accuracy and insufficient robustness of pedestrian attribute recognition in the prior art. In the field of intelligent security monitoring, pedestrian attribute recognition is a technical challenge. The present invention makes full use of the powerful image parsing capabilities of the multimodal large model to achieve accurate recognition of pedestrian attributes, thereby enhancing the overall effectiveness of the intelligent security monitoring system.
[0074] In a specific embodiment of the present invention, Figure 1 As shown, the embodiment of the present invention proposes a method for monitoring and evaluating pedestrian attributes based on a multimodal large model, including the following modules:
[0075] A human body recognition module is used to identify and confirm whether there is a human body object in the input image;
[0076] The pedestrian attribute analysis module is fine-tuned based on the multimodal large model so that it can analyze the attributes of pedestrians in the image and generate descriptions in a predetermined format;
[0077] The output quality evaluation module uses a pre-trained clip model for pedestrian attribute image-text matching to calculate the cosine similarity between the output text of the multimodal large model and the original image to verify the consistency and rationality of the output, so as to ensure that the large model will not have abnormal output and avoid the model from producing erroneous or illogical results;
[0078] The accuracy evaluation module is used to test the accuracy of pedestrian attributes in the model output text and obtain its detection accuracy to measure the performance of the model.
[0079] Human Recognition Module: This module is responsible for detecting and locating human targets from the input image, and is the basis for subsequent pedestrian attribute analysis. Through an efficient and accurate human detection algorithm, it can ensure that only the human part in the image is analyzed for attributes, improving the pertinence and efficiency of processing.
[0080] Pedestrian attribute analysis module: Based on the multimodal large model, this module enables the model to understand and extract various attribute information of pedestrians in the image (such as gender, age, clothing color, etc.) through fine-tuning, and convert these attributes into structured text descriptions. The powerful representation and generalization capabilities of the multimodal large model enable it to accurately identify pedestrian attributes in a variety of scenarios.
[0081] Output quality evaluation module: In order to ensure the output quality of the pedestrian attribute analysis module, the present invention introduces a pre-trained pedestrian attribute text-image correspondence CLIP model. This model can calculate the cosine similarity between the text description generated by the multimodal large model and the original image, so as to evaluate the consistency between the output text and the image content and prevent the occurrence of abnormal output.
[0082] Accuracy evaluation module: Finally, the pedestrian attributes output by the multimodal large model are verified through the accuracy evaluation module. Specifically, it checks whether the pedestrian attributes mentioned in the output text are consistent with the actual situation, so as to calculate the recognition accuracy. This process helps to continuously optimize the model performance and ensure its reliability and effectiveness in practical applications.
[0083] The execution process of the human body recognition module specifically includes the following steps:
[0084] Pedestrian images can be directly extracted using visual detection models such as yolo or dino;
[0085] When encountering data areas where it is difficult to extract pedestrians, you can fine-tune the large model by building your own difficult sample pedestrian detection dataset to enable the large model to have the ability to extract difficult pedestrian samples.
[0086] The execution process of the pedestrian attribute analysis module specifically includes the following steps:
[0087] First, determine the specific attributes of the pedestrian to be identified. In this solution, the pedestrian's age, gender, upper body color and style, lower body color and style, whether he is carrying a hat, shoulder bag, backpack, umbrella, handbag and other attributes are selected for identification;
[0088] After confirming the pedestrian attribute categories, we built a large fine-tuning dataset using our own dataset. The content of the dataset mainly consists of pedestrian images and their corresponding description text pairs. We need to unify the description text paradigm so that the multimodal large model has a unified output text format.
[0089] The multimodal large model is then fine-tuned, and after fine-tuning the output model can be evaluated in the next step.
[0090] The execution process of the output quality evaluation module specifically includes the following steps:
[0091] In order to eliminate the large deviation of the output of the large model (hallucination), a clip model was constructed to train the same training data set, and a clip model that can simultaneously calculate the feature vectors of images and texts was obtained;
[0092] The trained clip model is used to calculate the feature vector similarity of the image-text pair output by the multimodal large model to obtain the similarity score of the image-text pair;
[0093] At the same time, the feature vector similarity is evaluated with the original data image-text pair to obtain the image-text pair similarity score. This score is used as the standard to calculate whether there is a large deviation between the image-text pair similarity score output by the large model and this score. If there is a large deviation, the model output is judged to be an illusion, thereby achieving output quality evaluation of the large model.
[0094] The execution process of the accuracy assessment module specifically includes the following steps:
[0095] Use statistical methods to identify attributes that are inconsistent between the output of the large model and the original text, and calculate the objective accuracy of the large model attribute recognition;
[0096] At the same time, the corresponding image path is found and the images with errors in large model attribute recognition are saved in a separate folder. It is convenient to check at any time which pedestrian image has deviation in large model attribute recognition, so as to realize subjective evaluation.
[0097] The following lists the preferred embodiments of the low-quality data evaluation system and method for intelligent safety monitoring systems in industrial safety monitoring systems to clearly illustrate the contents of the present invention. It should be clear that the contents of the present invention are not limited to the following embodiments, and other improvements through conventional technical means of ordinary technicians in this field are also within the scope of the concept of the present invention.
[0098] S100 builds a human recognition module, which uses the target detection capability of the large model to detect pedestrians in the input image and identify whether the video image contains human targets. The module is mainly divided into the following steps:
[0099] S101 training data preparation, obtain actual scene data from surveillance cameras, use open source human detection model to crop the area of interest containing human body as positive sample, and the area without human body as negative sample. The data is divided into training set (accounting for 80% of the total), validation set (accounting for 10% of the total) and test set (accounting for 10% of the total).
[0100] S102 human body detection function training, converting the collected data into jsonl format and inputting it into the large model for training,
[0101] S103 Model testing and iterative update: Use the test set data to test the model effect. If the effect is not ideal, analyze the cause and supplement the data to retrain the model until the model effect meets the expected requirements.
[0102] like Figure 2 As shown, the working process of the human body recognition module of the present invention is as follows:
[0103] (1.1) Input image;
[0104] (1.2) Extract image features (such as color, texture, shape, etc.);
[0105] (1.3) Use human detection algorithms (such as YOLO, etc.) to detect human targets;
[0106] (1.4) Determine whether a human body is detected. If so, locate the human body target; otherwise, output information that there is no human object.
[0107] like Figure 3 As shown, the workflow of the pedestrian attribute analysis module in the present invention is as follows:
[0108] (2.1) Obtaining a human body image located by a human body recognition module;
[0109] (2.2) Load pre-trained multimodal large models (such as DeepSeek-VL2, Qwen2-VL and other series models);
[0110] (2.3) Fine-tune the model using pedestrian attribute-related datasets;
[0111] (2.4) Input human image to the fine-tuned model;
[0112] (2.5) The model analyzes the pedestrian's attributes (such as gender, age, clothing, belongings, etc.);
[0113] (2.6) Output pedestrian attribute analysis results.
[0114] like Figure 4 As shown, the workflow of the output quality assessment module of the present invention is as follows:
[0115] (3.1) Obtain the text output by the pedestrian attribute analysis module and the original input image;
[0116] (3.2) Load the pre-trained CLIP model for pedestrian attribute image-text matching;
[0117] (3.3) Use the CLIP model to extract the feature vector of the output text, and use the CLIP model to extract the feature vector of the original image;
[0118] (3.4) Calculate the cosine similarity between the text and image feature vectors;
[0119] (3.5) Output similarity results and evaluate output quality.
[0120] like Figure 5 As shown, the workflow of the accuracy assessment module of the present invention is as follows:
[0121] (4.1) Obtain the attribute results output by the pedestrian attribute analysis module and the corresponding real pedestrian attribute labels;
[0122] (4.2) Determine whether the evaluation of all attribute categories is completed. If so, continue with step (4.3); otherwise, continue with step (4.5);
[0123] (4.3) Calculate the overall accuracy based on the accuracy of each attribute category;
[0124] (4.4) Output accuracy evaluation results;
[0125] (4.5) Select the next attribute category;
[0126] (4.6) Calculate the accuracy of the current attribute category (number of correct recognitions / total number of recognitions).
[0127] The present invention has the following characteristics:
[0128] Multi-module collaborative working architecture: The present invention proposes a complete system including human body recognition, pedestrian attribute analysis, output quality evaluation and precision assessment. Each module cooperates with each other to ensure the accuracy and reliability of pedestrian attribute recognition. This is a key innovation point not mentioned in the comparative document.
[0129] Output quality evaluation method: The present invention particularly emphasizes the calculation of the cosine similarity between the output text of the multimodal large model and the original image through the pre-trained pedestrian attribute image-text matching clip model to evaluate the consistency and rationality of the output. This point is not specifically reflected in other comparative documents. There are obvious differences in the cosine similarity matching between different texts and images. Based on this, the influence of the illusion of the multimodal large model on the final output pedestrian attribute text can be significantly eliminated.
[0130] Specific operation of accuracy assessment: The present invention designs a special accuracy assessment module to check the accuracy of pedestrian attributes in the model output text and obtain the detection accuracy. This meticulous error checking and feedback mechanism is the unique contribution of the present invention. Special processing of the human body recognition module: Pedestrian images are directly extracted using the visual detection model. When pedestrian data is difficult to extract, the large model is fine-tuned by self-building a difficult sample pedestrian detection data set. This flexible processing method helps to improve the accuracy and adaptability of human body recognition.
[0131] The fine-tuning process of the pedestrian attribute analysis module: by determining the pedestrian attributes that need to be identified, constructing a fine-tuning dataset from a self-built dataset, fine-tuning the multimodal large model, and subsequently evaluating the output model, this series of unique fine-tuning processes can make the model more suitable for pedestrian attribute recognition tasks and is creative.
[0132] Innovative evaluation method of the output quality evaluation module: The output quality evaluation module calculates the cosine similarity between the output text of the multimodal large model and the original image through the pre-trained pedestrian attribute image-text matching clip model, and uses this to determine whether the model output is a hallucination. This method can effectively evaluate the consistency and rationality of the model output text and the original image, and is an innovative evaluation method.
[0133] The accuracy evaluation module can calculate the accuracy and save the error images: The accuracy evaluation module can not only check whether the pedestrian attributes mentioned in the output text are consistent with the actual situation, evaluate whether the output text content of the large model is consistent with the original data label, and calculate the accuracy of the large model attribute recognition, but also find the images with errors in the large model attribute recognition and save them in a separate folder. This precise calculation of the recognition accuracy and the saving of error images facilitate further optimization and improvement of the model, and improve the reliability and practicality of the model.
[0134] In summary, the present invention is not only innovative in technical details, but also shows significant progress in overall architecture design and practical application effects. These unique technical features bring obvious creativity and practicality to this application, making it of great value in the application field of intelligent safety monitoring systems.
[0135] Compared with the prior art, the present invention has the following significant advantages:
[0136] Higher recognition accuracy: By utilizing the powerful image parsing capabilities of the multimodal large model, the present invention can more accurately identify pedestrian attributes, such as gender, age, clothing color, etc. Compared with traditional methods, the multimodal model can better capture pedestrian features and reduce misjudgment by fusing image and text information.
[0137] Stronger robustness: The human recognition module and pedestrian attribute analysis module of the present invention both use advanced deep learning algorithms, which can maintain high recognition performance under different lighting conditions, angle changes, occlusion and other factors. In addition, the output quality evaluation module further ensures the stability and reliability of the model output by calculating the cosine similarity.
[0138] Comprehensive evaluation system: This invention not only provides pedestrian attribute recognition function, but also designs a dedicated output quality evaluation module and accuracy evaluation module to form a complete evaluation system. This enables the system to automatically detect and correct potential errors, continuously optimize model performance, and ensure long-term stable high-precision recognition.
[0139] Strong adaptability: The multimodal large model has strong generalization ability and can adapt to pedestrian attribute recognition tasks in various scenarios, without the need to retrain the model for each specific scenario. This greatly reduces the deployment cost and maintenance difficulty of the system and improves the flexibility and breadth of application.
[0140] Structured output: The pedestrian attribute analysis module outputs the recognition results in a structured text format, which is convenient for integration with other systems or applications, improving the usability and interoperability of the data. This structured output method is also helpful for subsequent data analysis and processing.
[0141] Innovative technical solution: This invention creatively applies a multimodal large model to the field of pedestrian attribute recognition, and solves many problems existing in the prior art through a series of innovative module designs. This new technical solution has brought revolutionary changes to application scenarios such as intelligent safety monitoring systems.
[0142] The specific implementation scheme of this embodiment can refer to the relevant description in the above embodiment, which will not be repeated here.
[0143] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0144] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" refers to at least two.
[0145] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0146] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0147] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the corresponding program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.
[0148] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0149] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0150] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0151] The method, device, processor and computer-readable storage medium of the present invention for realizing pedestrian attribute recognition and evaluation based on a multimodal large model are adopted. Through the coordinated work of the human body recognition module, the pedestrian attribute analysis module, the output quality evaluation module and the accuracy evaluation module, not only the pedestrian attributes can be accurately identified, but also the quality of the recognition results can be effectively evaluated, thus ensuring the reliability and accuracy of the output, and providing strong technical support for the application and development of intelligent safety monitoring systems.
[0152] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it is apparent that various modifications and variations may be made without departing from the spirit and scope of the present invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method for pedestrian attribute recognition and evaluation based on a multimodal large model, characterized in that: The method comprises the following steps: (1) The human body recognition module identifies and confirms whether there is a human body object in the input image and locates the human body target; (2) The pedestrian attribute analysis module is fine-tuned based on the multimodal large model, analyzes the attributes of pedestrians in the image, generates descriptions in a predetermined format, and outputs text; (3) The output quality evaluation module calculates the cosine similarity between the output text of the multimodal large model and the original image through the pre-trained pedestrian attribute image-text matching clip model; (4) Accuracy evaluation module: checks the accuracy of pedestrian attributes in the model output text and obtains its detection accuracy.
2. The method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model according to claim 1 is characterized in that: The step (1) specifically comprises the following steps: (1.1) Directly extract pedestrian images using visual detection models; (1.2) If pedestrian data is difficult to extract, the large model can be fine-tuned by building a self-built difficult sample pedestrian detection dataset.
3. The method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model according to claim 1 is characterized in that: The step (2) specifically comprises the following steps: (2.1) Determine the attributes of pedestrians that need to be identified; (2.2) Use the self-built dataset method to construct a fine-tuning dataset, and output the text format of pedestrian images and their corresponding description texts in a unified manner; (2.3) Fine-tune the large multimodal model and conduct subsequent evaluation on the output model after fine-tuning.
4. The method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model according to claim 1 is characterized in that: The step (3) specifically comprises the following steps: (3.1) Construct a clip model and train it on the same training data set to obtain a clip model that can simultaneously calculate the feature vectors of images and texts; (3.2) Use the CLIP model to extract the feature vectors of the output text and the original image at the same time, and use the trained clip model to calculate the feature vector similarity of the image-text pair output by the multimodal large model to obtain the similarity score of the image-text pair; (3.3) Simultaneously, the feature vector similarity is evaluated with the original data image-text pair to obtain the image-text pair similarity score; (3.4) The image-text pair similarity score is used as the standard to calculate whether there is a deviation between the image-text pair similarity score output by the large model and the image-text pair similarity score. If so, the model output is judged to be a hallucination; otherwise, the model output is not a hallucination.
5. The method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model according to claim 1 is characterized in that: The step (4) specifically comprises the following steps: (4.1) Check whether the pedestrian attributes mentioned in the output text are consistent with the actual situation, evaluate whether the output text content of the large model is consistent with the original data label, and calculate the accuracy of the large model attribute recognition; (4.2) Find the images where the large model attribute recognition errors occurred and save them in a separate folder.
6. A system for realizing pedestrian attribute recognition and evaluation based on a multimodal large model for implementing the method described in claim 1, characterized in that: The system comprises: A human body recognition module is used to identify and confirm whether there is a human body object in the input image and locate the human body target; A pedestrian attribute analysis module, connected to the human body recognition module, is used to perform fine-tuning on the basis of the multimodal large model, analyze various attributes of pedestrians in the image, generate descriptions in a predetermined format, and output text; An output quality evaluation module, connected to the pedestrian attribute analysis module, is used to calculate the cosine similarity between the output text of the multimodal large model and the original image through a pre-trained pedestrian attribute image-text matching clip model to verify the consistency and rationality of the output; The accuracy assessment module is connected to the output quality evaluation module and is used to evaluate whether the text content output by the large model is consistent with the original data label to obtain its detection accuracy.
7. A device for realizing pedestrian attribute recognition and evaluation based on a multimodal large model, characterized in that: The device comprises: a processor configured to execute computer executable instructions; A memory storing one or more computer executable instructions, wherein when the computer executable instructions are executed by the processor, each step of the method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model as described in any one of claims 1 to 6 is implemented.
8. A processor for realizing pedestrian attribute recognition and evaluation based on a multimodal large model, characterized in that: The processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the method for realizing pedestrian attribute recognition and evaluation based on a multimodal large model as described in any one of claims 1 to 6.