Labeled Data Scoring Method and False Positive Labeled Data Identification Method Based on the Same
By constructing multiple-choice questions and judgment questions, using visual language models for question-and-answer questions, calculating the scores of each target box labeling data, and identifying false positive labeling data, the model training problems caused by false positive labeling in manual labeling are solved, and the efficiency and quality of data set production are improved.
Patent Information
- Application Number
- CN202510412530.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-02
AI Technical Summary
False positive marks often appear during manual labeling, resulting in reduced model training performance, overfitting and unstable training. Manually checking false positive marks is costly and inefficient, which cannot meet the needs of rapid iteration.
A method of scoring label data is provided. By obtaining target images, constructing multiple-choice questions and judgment questions, using visual language models for question-and-answer questions, calculating the number of times the selection of answers and judging the correct answers of each visual language model, and comprehensively obtaining the score of each target box label data, thereby identifying false positive label data.
It improves the accuracy of labeled data scores, ensures the comprehensiveness and accuracy of evaluation results, reduces misjudgment, improves the efficiency and quality of target detection data set production, and reduces the cost and workload of manual inspection.
Smart Images

Figure CN119942315B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method for scoring labeled data and a method for identifying false positive labeled data based on the same. Background Art
[0002] With the development of deep learning, the demand for high-quality labeled data in object detection tasks has increased. However, false positive labels (mistakenly labeling non-target objects as targets) often occur during the manual labeling process, which can lead to problems such as decreased performance, overfitting, and unstable training during model training. Manually checking false positive labels is costly and inefficient, and cannot meet the needs of rapid iteration.
[0003] Therefore, how to automatically score labeled data efficiently and accurately is an urgent problem to be solved currently, in order to provide support for subsequent identification of false positive labels, thereby improving the efficiency and quality of object detection dataset production. Summary of the Invention
[0004] This application provides a method for scoring labeled data and a method for identifying false positive labeled data based on the same, achieving the technical effect of being able to improve the accuracy of scoring labeled data and ensure the comprehensiveness and precision of evaluation results.
[0005] To achieve the above object, the main technical solutions adopted in this application include:
[0006] In the first aspect, an embodiment of this application provides a method for scoring labeled data, and the scoring method includes:
[0007] Obtain a target image containing a target object, where the target image is labeled with target box annotation data corresponding to the target object, and the target box annotation data includes a target category name;
[0008] Repeatedly construct multiple-choice questions based on the target category name and a set of interference category names, and input the multiple-choice questions and the target image into at least one vision-language model for question answering to determine the number of correct selection answers corresponding to the vision-language model;
[0009] Repeatedly generate true / false questions based on the target category name and the target image from a pre-configured true / false question template pool, and input the true / false questions and the target image into at least one vision-language model for question answering to determine the number of correct judgment answers corresponding to the vision-language model;
[0010] Based on the number of correct selection answers and the number of correct judgment answers, determine the average score of all the vision-language models, and use the average score as the target score corresponding to the target box annotation data in the target image.
[0011] A method for scoring labeled data provided in this embodiment improves the accuracy of target recognition by combining a vision-language model. Specifically, the vision-language model is used to answer multiple-choice questions and true / false questions for a target image, and based on the results of these answers, the recognition accuracy of the target bounding box data is verified. In this way, misjudgments are reduced. Finally, by calculating the number of correct answers of each vision-language model in multiple-choice questions and true / false questions, the score of the labeled data for each target box is comprehensively obtained, providing a more reliable basis for the quality assessment of target detection labeling.
[0012] In one embodiment, the target box labeled data includes a target bounding box; the steps of repeatedly constructing multiple-choice questions based on the target category name and a set of interference category names, and inputting the multiple-choice questions and the target image into at least one vision-language model for answering to determine the number of correct answers of the selected answers corresponding to the vision-language model include:
[0013] Steps for constructing multiple-choice questions: Using the coordinate information of the target bounding box in the target image and the target image to construct a selection question, and using the target category name and the set of interference category names to construct answer options, obtaining a multiple-choice question including the selection question and the answer options;
[0014] Steps for answering multiple-choice questions: Inputting the multiple-choice question and the target image into at least one vision-language model, and obtaining the selected answer corresponding to each vision-language model;
[0015] Repeatedly execute the steps for constructing multiple-choice questions and the steps for answering multiple-choice questions until the selected answers of at least one vision-language model in at least one multiple-choice question are obtained, and obtain the number of correct answers of the selected answers corresponding to each vision-language model.
[0016] This embodiment generates multiple-choice questions based on the coordinate information of the image and the target image, etc. The content of the questions will be set according to the description in the image. The answer options include the correct category and interference categories, and the order of the answer options will be randomized. Then, the multiple-choice questions and the target image are input into the vision-language model together. The model analyzes the image content and selects the most appropriate answer. Finally, by constructing and answering multiple-choice questions multiple times, the answering situation of the model is recorded.
[0017] In one embodiment, the method for obtaining the set of interference category names includes:
[0018] Obtaining an image dataset including the target image; wherein, the image dataset includes a set of category names of different target images;
[0019] Selecting at least one interference category name different from the target category name from the set of category names; wherein, multiple interference category names are different from each other;
[0020] Determine the preset independent category name option as the interference category name;
[0021] Integrate all the interference category names to obtain the set of interference category names.
[0022] In this embodiment, data is obtained from an image dataset containing various target images and category names. Then, interference category names different from the target category are selected. These interference category names may be similar to the target category name but not the same. Next, the preset independent category name option is determined as the interference category name, further enriching the selection of interference category names. Finally, these interference category names are integrated into a set of interference category names for subsequent scoring.
[0023] In one implementation, the target box annotation data includes target bounding boxes; the step of repeatedly generating true / false questions based on the target category name and the target image from a pre-configured true / false question template pool, and inputting the true / false questions and the target image into at least one vision-language model for question answering to determine the number of correct judgment answers corresponding to the vision-language model includes:
[0024] True / false question construction step: Generate true / false questions from a pre-constructed true / false question template pool based on the coordinate information of the target bounding box in the target image, the target image, and the target category name;
[0025] True / false question answering step: Input the true / false questions and the target image into at least one vision-language model to obtain the judgment answers corresponding to each vision-language model;
[0026] Repeat the true / false question construction step and the true / false question answering step until at least one vision-language model's judgment answers in at least one true / false question are obtained, and obtain the number of correct judgment answers corresponding to each vision-language model.
[0027] In this embodiment, true / false questions are generated from a pre-constructed true / false question template pool based on the coordinate information in the target image, the target image itself, and the target category name. Then, these true / false questions and the target image are input into at least one vision-language model, and the model generates judgment answers based on the image content. This process is repeated multiple times until the answers of each vision-language model in multiple rounds of true / false questions are obtained, and the number of correct answers of each model is counted. This improves the accuracy of evaluating the target box annotation data.
[0028] In one implementation, determining the average score of all the vision-language models based on the number of correct selection answers and the number of correct judgment answers includes:
[0029] For any visual language model, based on the number of correct selected answers, the number of correct judged answers, and the number of repetitions, determine the comprehensive score corresponding to any visual language model;
[0030] Based on the comprehensive scores, determine the average score of all the visual language models.
[0031] In this embodiment, the total number of correct times of the visual language model can be obtained by summing the number of correct selected answers and the number of correct judged answers. By dividing by the total number of repetitions (i.e., the sum of the number of repetitions of multiple-choice questions and the number of repetitions of true / false questions), the comprehensive score corresponding to each visual language model is obtained. Averaging the comprehensive scores of all visual language models gives the target score for the target box annotation data. This average score can reflect the overall judgment of all visual language models on the target box annotation data and provide a more accurate basis for subsequent false positive evaluation.
[0032] In one implementation, the determining the average score of all the visual language models includes:
[0033]
[0034] where z is the average score; n is the number of visual language models; x i is the number of correct selected answers corresponding to the i-th visual language model; w i is the number of correct judged answers corresponding to the i-th visual language model; k is the number of repetitions of multiple-choice questions; j is the number of repetitions of true / false questions.
[0035] In a second aspect, an embodiment of the present application provides a method for identifying false positive annotation data, and the identification method includes:
[0036] Obtain an image data set; wherein, the image data set includes a plurality of target images, and each target image is annotated with target box annotation data to be recognized corresponding to a target object;
[0037] For the target box annotation data in any target image, use the scoring method as described above to score and obtain the target score corresponding to the target box annotation data in the target image;
[0038] In the case where the target score is lower than a preset score threshold, determine the target box annotation data corresponding to the target score as false positive annotation data.
[0039] A method for identifying false positive annotation data provided in this embodiment includes obtaining an image dataset that contains multiple target images, and each image is annotated with target box annotation data to be identified. Score each target box annotation data to obtain a target score. If the target score is lower than a preset score threshold, it is determined as false positive annotation data. This automated process not only reduces the cost and workload of manual inspection, but also greatly improves the identification efficiency of false positives. Especially in the process of processing large-scale datasets, it can quickly complete the screening to ensure the high quality of the dataset. At the same time, it can effectively support the production of large-scale object detection datasets, improve the dataset production efficiency, and provide more accurate and reliable data for subsequent model training.
[0040] In a third aspect, an embodiment of the present application provides a scoring device for annotation data, and the device includes:
[0041] An image acquisition unit, configured to acquire a target image containing a target object, where the target image is annotated with target box annotation data corresponding to the target object, and the target box annotation data includes a target category name;
[0042] A multiple-choice question construction and answering unit, configured to repeatedly construct multiple-choice questions based on the target category name and a set of interference category names, and input the multiple-choice questions and the target image into at least one vision-language model for answering to determine the number of correct answers for the corresponding multiple-choice questions of the vision-language model;
[0043] A true / false question construction and answering unit, configured to repeatedly generate true / false questions based on the target category name and the target image from a pre-configured true / false question template pool, and input the true / false questions and the target image into at least one vision-language model for answering to determine the number of correct answers for the corresponding true / false questions of the vision-language model;
[0044] A model scoring unit, configured to determine the average score of all the vision-language models based on the number of correct answers for the multiple-choice questions and the number of correct answers for the true / false questions, and use the average score as the target score corresponding to the target box annotation data in the target image.
[0045] In a fourth aspect, an embodiment of the present application provides a computer device, including:
[0046] A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the above-mentioned scoring method or the above-mentioned identification method.
[0047] Fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the above-mentioned scoring method or execute the above-mentioned recognition method. Description of the Drawings
[0048] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a flowchart of a method for scoring labeled data provided by an embodiment of the present application;
[0050] Figure 2 It is a flowchart of step S3 provided by an embodiment of the present application;
[0051] Figure 3 It is a flowchart of a method for obtaining a set of interference category names provided by an embodiment of the present application;
[0052] Figure 4 It is a flowchart of step S5 provided by an embodiment of the present application;
[0053] Figure 5 It is a flowchart of step S7 provided by an embodiment of the present application;
[0054] Figure 6 It is a flowchart of a method for identifying false positive labeled data provided by an embodiment of the present application;
[0055] Figure 7 It is a block diagram of a device for scoring labeled data provided by an embodiment of the present application;
[0056] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed Embodiments
[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0058] With the rapid development of deep learning in fields such as computer vision, natural language processing, and speech recognition, the demand for training data is increasing day by day. Especially in object detection tasks, high-quality labeled data is crucial. The manual annotation of object detection datasets usually includes steps such as the recognition of target objects, localization, bounding box drawing, and the assignment of target category labels. However, errors may occur during the annotation process, such as mislabeling non-target objects as targets (false positive annotations) or misclassifying target objects. These problems will introduce noise and affect the model training effect.
[0059] Annotation noise has a negative impact on model training, including misleading the model to learn incorrect decision boundaries, resulting in performance degradation, or causing overfitting, where the model performs well on the training data but has poor generalization ability. In addition, noisy labels may also increase the instability of the training process, increasing costs and difficulties.
[0060] Although the process of manually checking false positive annotations is necessary, it has high costs and inefficiencies. Manual checking requires a large amount of human resources, recruitment, training, and management, and it takes a long time, unable to meet the needs of rapid iteration, thus affecting the research and development progress.
[0061] Based on the above technical problems, according to the embodiments of the present application, an embodiment of a method for scoring annotated data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0062] In this embodiment, a method for scoring annotated data is provided. Figure 1 The following is a flowchart of a method for scoring annotated data provided by the embodiments of the present application. As Figure 1 shown, the process includes the following steps:
[0063] Step S1, obtain a target image containing a target object. The target image is annotated with target box annotation data corresponding to the target object, and the target box annotation data includes the target category name.
[0064] Specifically, obtain a target image containing a target object. Ensure that the target image is annotated with target box annotation data corresponding to the target object, including the target category name.
[0065] Step S3, repeatedly construct multiple-choice questions based on the target category name and a set of interference category names, and input the multiple-choice questions and the target image into at least one vision-language model for question answering to determine the number of correct selection answers corresponding to the vision-language model.
[0066] Specifically, select distracting class names different from the target class name from the target detection dataset. Based on the target class name and the set of distracting class names, construct multiple-choice questions. Input the multiple-choice questions and the target image into at least one vision-language model to obtain the answers of the vision-language model. Determine the number of times the vision-language model selects the correct answer. If the vision-language model correctly selects the target class name, it is regarded as a correct answer; if it selects a distracting class, it is regarded as an incorrect answer. By counting the number of correct answers of the vision-language model, further analyze the credibility of the target bounding box annotation data.
[0067] Step S5: Repeatedly generate true / false questions based on the target class name and the target image from the pre-configured true / false question template pool, and input the true / false questions and the target image into at least one vision-language model for question answering to determine the number of correct judgment answers of the vision-language model.
[0068] Specifically, randomly select a template from the pre-configured true / false question template pool. Based on the target class name and the target image, use the template to generate true / false questions. Input the true / false questions and the target image into at least one vision-language model to obtain the answers of the vision-language model. Determine the number of times the vision-language model gives a correct judgment. By counting the number of correct answers of the vision-language model, the credibility of the target bounding box annotation data can be evaluated. If the vision-language model answers correctly multiple times, it indicates that the target bounding box annotation data is relatively accurate and has a high credibility; conversely, if the number of incorrect answers is large, it indicates that the target bounding box annotation data may have a high degree of inaccuracy or errors.
[0069] Step S7: Based on the number of correct selected answers and the number of correct judgment answers, determine the average score of all vision-language models, and use the average score as the target score corresponding to the target bounding box annotation data in the target image.
[0070] Specifically, calculate the comprehensive score of each vision-language model based on the number of correct answers to the multiple-choice questions and true / false questions. Take the average of the comprehensive scores of all vision-language models to obtain the final average score. Use the average score as the target score corresponding to the target bounding box annotation data in the target image. Through the average score, a unified quantitative index for the accuracy of the target bounding box annotation data is provided.
[0071] A method for scoring annotation data provided in this embodiment improves the accuracy of target recognition by combining vision-language models. Specifically, use the vision-language model to conduct question answering on multiple-choice questions and true / false questions for the target image, and verify the recognition accuracy of the target annotation bounding box data based on the results of these question answers. In this way, misjudgments are reduced. Finally, by calculating the number of correct times of each vision-language model in multiple-choice questions and true / false questions, the score of each target bounding box annotation data is comprehensively obtained, thereby providing a more reliable basis for the quality assessment of the annotation for target detection.
[0072] Figure 2 This is the flowchart of step S3 provided by the embodiment of the present application. The target box annotation data includes the target bounding box. This process may include the following steps:
[0073] Step S31, multiple-choice question construction step: Construct a selection question based on the coordinate information of the target bounding box in the target image and the target image, and construct answer options based on the target category name and the set of interfering category names, to obtain a multiple-choice question including the selection question and the answer options.
[0074] Specifically, based on the coordinate information of the target bounding box in the target image and the target image, generate a clear selection question asking which target category the image content belongs to. For example, "In image I, which of the following categories does the content within the coordinate range (100, 100, 300, 300) belong to?" Then use the target category name H 0 and the set of interfering category names {H 1 , H 2 ,..., H m} to construct the answer options. Provide multiple options. The set of interfering category names includes answer options different from the target category name H 0 , and each interfering category name in the set of interfering category names is different from each other. In addition, it also includes an independent category name option, such as "does not belong to any of the other options".
[0075] It should be noted here that the order of the answer options should be randomly shuffled when generating the multiple-choice question. By shuffling the order of the answer options, fairness can be improved. Assume that the target category name H 0 in the answer options is "car", the interfering category name H 1 is "bicycle", H 2 is "person", H 3 is "bicycle", H 4 is "does not belong to any of the other options", and after random shuffling, it may be {"person", "car", "does not belong to any of the other options", "bicycle"}.
[0076] Once a clear selection and randomly ordered answer options are generated, they can be combined into a complete multiple-choice question.
[0077] Step S33, multiple-choice question answering step: Input the multiple-choice question and the target image into at least one vision-language model to obtain the selection answer corresponding to each vision-language model.
[0078] Specifically, the constructed multiple-choice questions (including the question and answer options) are converted into text form, and the target image is passed as input to the vision-language model. The vision-language model generates an answer based on the input image and text question. The vision-language model combines the content of the target image and the question text, and performs reasoning through the internal vision encoder and language decoder. The token sequence generated by the vision-language model is decoded into natural language text to obtain the selection answer of the vision-language model to the multiple-choice question. Each vision-language model generates an answer for one input multiple-choice question. Preferably, the vision-language model can adopt models such as the Qwen-VL series, Intern-VL series, and deepseek-VL series.
[0079] Step S35, repeatedly execute the multiple-choice question construction step and the multiple-choice question answering step until at least one selection answer of the vision-language model in at least one multiple-choice question is obtained, and obtain the number of correct selection answers corresponding to each vision-language model.
[0080] Specifically, by repeatedly executing the multiple-choice question construction and answering steps, multiple answers can be obtained, thereby improving the accuracy of the evaluation results.
[0081] In this embodiment, multiple-choice questions are generated based on the coordinate information of the image and the content of the target image, etc. The content of the question is set according to the description in the image. The answer options include the correct category and the interference category, and the order of the answer options is randomized. Then, the multiple-choice questions and the target image are input into the vision-language model together. The model analyzes the image content and selects the most appropriate answer. Finally, by constructing and answering multiple-choice questions multiple times, the answering situation of the model is recorded.
[0082] Figure 3 It is a flowchart of the acquisition method of the set of interference category names provided by the embodiment of the present application. This process may include the following steps:
[0083] Step S301, obtain an image dataset containing the target image; wherein, the image dataset includes a set of category names of different target images.
[0084] Specifically, the image dataset is the core of the object detection task. Each target image in the image dataset is crucial label information in the object detection task. These label information clarify the position of the target object in the target image and its corresponding target category name. The set of category names is the set of all different target category names in the image dataset, which constitutes the label space of the object detection task.
[0085] Step S303, select at least one interference category name different from the target category name from the set of category names; wherein, the multiple interference category names are different from each other.
[0086] Specifically, the interference category name is a category different from the target category name and is used to construct multiple-choice questions. Specifically, the interference category name must be different from the target category name. Multiple interference category names cannot be repeated to ensure that each interference item is independent. The interference category name can include categories similar to and dissimilar from the target category name. Preferably, the selection of the interference category name can be made randomly or based on heuristic rules.
[0087] Step S305: Determine the preset independent category name option as the interference category name.
[0088] Specifically, introduce the preset independent category name option, which can be "not belonging to any of the other option categories".
[0089] Step S307: Combine all the interference category names to obtain an interference category name set.
[0090] Specifically, by combining all the interference category names to obtain an interference category name set, a complete multiple-choice question can be constructed. This not only increases the difficulty and diversity of the questions and answers but also improves the ability to judge false positive annotations in real scenarios.
[0091] This embodiment obtains data from an image dataset containing various target images and category names. Then, select interference category names different from the target category. These interference category names may be similar to the target category name but not the same. Next, the preset independent category name option will be determined as the interference category name to further enrich the selection of interference category names. Finally, these interference category names are integrated into an interference category name set for subsequent scoring.
[0092] Figure 4 This is the flowchart of step S5 provided by the embodiment of the present application. The target box annotation data includes the target bounding box. This process may include the following steps:
[0093] Step S51: True / False question construction step: Generate true / false questions from a pre-constructed true / false question template pool based on the coordinate information of the target bounding box in the target image, the target image, and the target category name.
[0094] Specifically, the true / false question template pool is a set containing multiple different true / false question templates. Each template represents a way of asking questions, forming different judgment questions for the coordinate information, the target image, and the target category name. The design of different templates can examine the content of the target category from multiple perspectives and avoid a single way of asking questions. For example, one template may directly ask whether the target image conforms to the target category name, while another template may ask questions by emphasizing the exclusion of other categories.
[0095] For example, Template 1: "Does the content in the target image I belong to the target category name?" This question directly asks whether the content of the target image conforms to a certain category. Template 2: "Based on the target image I, determine whether the content of the image is consistent with the target category name, rather than other categories?"
[0096] Step S53, True / False question answering step: Input the true / false question and the target image into at least one vision-language model to obtain the judgment answer corresponding to each vision-language model.
[0097] Specifically, convert the constructed true / false question into text form and use it as input together with the target image to the vision-language model. The vision-language model generates an answer based on the input image and text question. The vision-language model combines the content of the target image and the question text and performs reasoning through the internal vision encoder and language decoder. Decode the token sequence generated by the vision-language model into natural language text to obtain the selection answer of the vision-language model to the true / false question. Each vision-language model generates one answer for one input true / false question. Preferably, the vision-language model can adopt models such as the Qwen-VL series, Intern-VL series, and deepseek-VL series.
[0098] Step S55, Repeat the true / false question construction step and the true / false question answering step until at least one vision-language model's judgment answer in at least one true / false question is obtained, and obtain the number of correct judgment answers corresponding to each vision-language model.
[0099] Specifically, by repeating the true / false question construction and answering steps, multiple answers can be obtained, thereby improving the accuracy of the evaluation result.
[0100] In this embodiment, based on the coordinate information in the target image, the target image itself, and the target category name, true / false questions are generated from a pre-constructed true / false question template pool. Then, these true / false questions and the target image are input into at least one vision-language model, and the model generates a judgment answer based on the image content. This process is repeated multiple times until the answers of each vision-language model in multiple rounds of true / false questions are obtained, and the number of correct answers of each model is counted. Improve the accuracy of the evaluation of the target box annotation data.
[0101] Figure 5 It is a flowchart of step S7 provided by the embodiment of the present application. This process may include the following steps:
[0102] Step S71, For any vision-language model, based on the number of correct selection answers, the number of correct judgment answers, and the number of repetitions, determine the comprehensive score corresponding to any vision-language model.
[0103] Specifically, the comprehensive score can be determined by the number of correct answer selections, the number of correct answer judgments, and the number of repetitions. By adding the number of correct answers for multiple-choice questions and true / false questions and dividing by the total number of repetitions, a score between 0 and 1 can be obtained, which is used to evaluate the accuracy of the target box annotation data. The higher the data, the higher the correctness of the target box annotation data; conversely, it represents lower correctness of the target box annotation data.
[0104] Step S73: Based on the comprehensive score, determine the average score of all visual language models.
[0105] Specifically, the average score of all visual language models is obtained by calculating the average of the comprehensive scores of all visual language models.
[0106] In some preferred embodiments, determining the average score of all visual language models includes:
[0107]
[0108] where z is the average score; n is the number of visual language models; x i is the number of correct answer selections corresponding to the i-th visual language model; w i is the number of correct answer judgments corresponding to the i-th visual language model; k is the number of repetitions of multiple-choice questions; j is the number of repetitions of true / false questions.
[0109] In this embodiment, the total number of correct answers for the visual language model can be obtained by summing the number of correct answer selections and the number of correct answer judgments. By dividing by the total number of repetitions (i.e., the sum of the number of repetitions of multiple-choice questions and true / false questions), the comprehensive score corresponding to each visual language model is obtained. The comprehensive scores of all visual language models are averaged to obtain the target score for the target box annotation data. This average score can reflect the overall judgment of all visual language models on the target box annotation data and provide a more accurate basis for subsequent false positive evaluation.
[0110] In this embodiment, a method for identifying false positive annotation data is provided. Figure 6 This is a flowchart of a method for identifying false positive annotation data provided by an embodiment of the present application. As Figure 6 shown, the process includes the following steps:
[0111] Step S2: Obtain an image data set; where the image data set includes multiple target images, and each target image is annotated with target box annotation data corresponding to the target object to be recognized.
[0112] Step S4: For the target box annotation data in any target image, use the scoring method in Steps S1 - S7 to obtain the target score corresponding to the target box annotation data in the target image.
[0113] Step S6, in the case where the target score is lower than the preset score threshold, determine the target box annotation data corresponding to the target score as false positive annotation data.
[0114] Specifically, provide a labeled image dataset for subsequent scoring and false positive annotation detection. For each target image in the image dataset, extract its target box annotation data, and use the scoring method of Steps S1 to S7 to score each target box annotation data to obtain the target score corresponding to each target box annotation data. Set a preset score threshold (such as 0.5 or 0.6). For the target box annotation data whose target score is lower than the preset score threshold, determine it as false positive annotation data.
[0115] A method for identifying false positive annotation data provided in this embodiment includes obtaining an image dataset containing multiple target images, and each image is labeled with target box annotation data to be recognized. Score each target box annotation data to obtain a target score. If the target score is lower than the preset score threshold, determine it as false positive annotation data. This automated process not only reduces the cost and workload of manual inspection, but also greatly improves the recognition efficiency of false positive annotations. Especially in the process of processing large-scale datasets, it can quickly complete the screening to ensure the high quality of the dataset. At the same time, it can effectively support the production of large-scale object detection datasets, improve the dataset production efficiency, and provide more accurate and reliable data for subsequent model training.
[0116] Correspondingly, please refer to Figure 7 which is a block diagram of a scoring device for annotation data provided in an embodiment of this application. The terminal includes:
[0117] An image acquisition unit 101, configured to acquire a target image including a target object, where the target image is labeled with target box annotation data corresponding to the target object, and the target box annotation data includes a target category name;
[0118] A multiple-choice question construction and question-answering unit 103, configured to repeatedly construct multiple-choice questions based on the target category name and a set of interference category names, and input the multiple-choice questions and the target image into at least one vision-language model for question-answering to determine the number of correct selection answers corresponding to the vision-language model;
[0119] A true / false question construction and question-answering unit 105, configured to repeatedly generate true / false questions based on the target category name and the target image from a pre-configured true / false question template pool, and input the true / false questions and the target image into at least one vision-language model for question-answering to determine the number of correct judgment answers corresponding to the vision-language model;
[0120] The model scoring unit 107 is used to determine the average score of all visual language models based on the number of correct selected answers and the number of correct judged answers, and use the average score as the target score corresponding to the target box annotation data in the target image.
[0121] In some alternative embodiments, the target box annotation data includes a target bounding box; the multiple-choice question construction and answering unit 103 includes:
[0122] Multiple-choice question construction step: Construct a multiple-choice question with the coordinate information of the target bounding box in the target image and the target image, and construct answer options with the target category name and a set of interference category names to obtain a multiple-choice question including the multiple-choice question and the answer options;
[0123] Multiple-choice question answering step: Input the multiple-choice question and the target image into at least one visual language model to obtain the selected answer corresponding to each visual language model;
[0124] Repeat the multiple-choice question construction step and the multiple-choice question answering step until the selected answers of at least one visual language model in at least one multiple-choice question are obtained, and obtain the number of correct selected answers corresponding to each visual language model.
[0125] In some alternative embodiments, the method for obtaining the set of interference category names includes:
[0126] Obtain an image data set including the target image; wherein, the image data set includes a set of category names of different target images;
[0127] Select at least one interference category name different from the target category name from the set of category names; wherein, the multiple interference category names are different from each other;
[0128] Determine the preset independent category name option as the interference category name;
[0129] Integrate all the interference category names to obtain the set of interference category names.
[0130] In some alternative embodiments, the target box annotation data includes a target bounding box; the true / false question construction and answering unit 105 includes:
[0131] True / false question construction step: Generate a true / false question from a pre-constructed true / false question template pool based on the coordinate information of the target bounding box in the target image, the target image and the target category name;
[0132] True / false question answering step: Input the true / false question and the target image into at least one visual language model to obtain the judgment answer corresponding to each visual language model;
[0133] Repeat the steps of constructing the judgment questions and the steps of answering the judgment questions until at least one visual language model obtains a judgment answer in at least one judgment question, and obtain the number of correct judgment answers corresponding to each visual language model.
[0134] In some alternative embodiments, the model scoring unit 107 includes:
[0135] For any visual language model, determine the comprehensive score corresponding to any visual language model based on the number of correct selected answers, the number of correct judgment answers, and the number of repetitions;
[0136] Based on the comprehensive scores, determine the average score of all visual language models.
[0137] In some alternative embodiments, determining the average score of all visual language models includes:
[0138]
[0139] where z is the average score; n is the number of visual language models; x i is the number of correct selected answers corresponding to the i-th visual language model; w i is the number of correct judgment answers corresponding to the i-th visual language model; k is the number of repetitions of the multiple-choice questions; j is the number of repetitions of the judgment questions.
[0140] The further functional descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0141] A scoring device for labeled data in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0142] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As shown in Figure 8As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 8 In the figure, a processor 10 is taken as an example.
[0143] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0144] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0145] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0146] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.
[0147] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.
[0148] Embodiments of the present application also provide a computer-readable storage medium. The methods according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0149] The devices or units illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0150] For convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0151] Those skilled in the art should understand that the embodiments of the present application can be provided as methods and devices. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and devices according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a dedicated computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the processFigure 1 one process or multiple processes and / or boxes Figure 1 a device for the functions specified in one box or multiple boxes.
[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the operations in the process Figure 1 one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes.
[0155] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.
[0156] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0157] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
[0158] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for scoring labeled data, characterized in that: The scoring method includes: Acquire a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, and the target frame annotation data includes a target category name; Repeatedly construct multiple-choice questions based on the target category name and the set of interference category names, and input the multiple-choice questions and the target image into at least one visual language model for question and answer, and determine the number of correct answers corresponding to the visual language model; Repeatedly generating a judgment question based on the target category name and the target image from a pre-configured judgment question template pool, and inputting the judgment question and the target image into at least one visual language model for question answering, and determining the number of correct judgment answers corresponding to the visual language model; Based on the number of correct answer selections and the number of correct answer determinations, an average score of all the visual language models is determined, and the average score is used as a target score corresponding to the target box annotation data in the target image.
2. The scoring method according to claim 1, characterized in that: The target box annotation data includes a target bounding box; repeatedly constructing a multiple-choice question based on the target category name and the interference category name set, inputting the multiple-choice question and the target image into at least one visual language model for question answering, and determining the number of correct answers corresponding to the visual language model, including: Multiple-choice question construction step: constructing a multiple-choice question with the coordinate information of the target bounding box in the target image and the target image, constructing answer options with the target category name and the interference category name set, and obtaining a multiple-choice question including the multiple-choice question and the answer options; Multiple-choice question answering step: inputting the multiple-choice question and the target image into at least one visual language model, and obtaining a multiple-choice answer corresponding to each visual language model; Repeat the multiple-choice question construction step and the multiple-choice question answering step until a selection answer of at least one visual language model in at least one multiple-choice question is obtained, and the number of correct selection answers corresponding to each visual language model is obtained.
3. The scoring method according to claim 1 or 2, characterized in that: The method for obtaining the interference category name set includes: Acquire an image data set containing the target image; wherein the image data set includes a set of category names of different target images; Selecting at least one interference category name different from the target category name from the category name set; wherein the multiple interference category names are different from each other; Determine a preset independent category name option as the interference category name; All the interference category names are combined to obtain the interference category name set.
4. The scoring method according to claim 1, characterized in that: The target box annotation data includes a target bounding box; repeatedly generating a judgment question based on the target category name and the target image from a pre-configured judgment question template pool, inputting the judgment question and the target image into at least one visual language model for question answering, and determining the number of correct judgment answers corresponding to the visual language model, including: True or False question construction step: generating a true or false question from a pre-constructed true or false question template pool based on the coordinate information of the target bounding box in the target image, the target image and the target category name; Judgment question answering step: inputting the judgment question and the target image into at least one visual language model, and obtaining a judgment answer corresponding to each visual language model; Repeat the judgment question construction step and the judgment question answering step until a judgment answer of at least one visual language model in at least one judgment question is obtained, and the number of correct judgment answers corresponding to each visual language model is obtained.
5. The scoring method according to claim 1, characterized in that: The determining, based on the number of correct answer selections and the number of correct answer determinations, an average score of all the visual language models comprises: For any visual language model, based on the number of correct selected answers, the number of correct judged answers and the number of repetitions, determine a comprehensive score corresponding to any visual language model; Based on the comprehensive score, an average score of all the visual language models is determined.
6. The scoring method according to claim 1, characterized in that: Determining the average score of all the visual language models comprises: Among them, z is the average score; n is the number of visual language models; x i is the number of correct answers selected by the i-th visual language model; w i is the number of correct answers corresponding to the i-th visual language model; k is the number of repetitions of the multiple-choice question; j is the number of repetitions of the judgment question.
7. A method for identifying false positive labeled data, characterized in that: The identification method comprises: Acquire an image data set; wherein the image data set includes a plurality of target images, each of the target images is annotated with target frame annotation data to be identified corresponding to the target object; For the target frame annotation data in any target image, scoring is performed using the scoring method according to any one of claims 1 to 6 to obtain a target score corresponding to the target frame annotation data in the target image; When the target score is lower than a preset score threshold, the target frame annotation data corresponding to the target score is determined as false positive annotation data.
8. A scoring device for annotated data, characterized in that: The device comprises: An image acquisition unit, configured to acquire a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, wherein the target frame annotation data includes a target category name; A multiple-choice question construction and question-answering unit, used to repeatedly construct multiple-choice questions based on the target category name and the interference category name set, and input the multiple-choice questions and the target image into at least one visual language model for question-answering, and determine the number of correct answers corresponding to the visual language model; A judgment question construction and question-answering unit, used to repeatedly generate judgment questions based on the target category name and the target image from a pre-configured judgment question template pool, and input the judgment question and the target image into at least one visual language model for question-answering, and determine the number of correct judgment answers corresponding to the visual language model; The model scoring unit is used to determine the average score of all the visual language models based on the number of correct answer selections and the number of correct answer judgments, and use the average score as the target score corresponding to the target box annotation data in the target image.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the scoring method according to any one of claims 1 to 6 or the identification method according to claim 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the scoring method according to any one of claims 1 to 6 or the identification method according to claim 7.
Citation Information
Patent Citations
Question identification method
CN118781612A
Saliency for anchor-based object detection
US20240054754A1