Label data scoring method and false positive label data identification method based on label data scoring method

By constructing multiple-choice questions and judgment questions and entering a visual language model for question-and-answer questions, the accuracy of the target box labeling data was evaluated, and the model performance degradation caused by false positive marking in manual labeling was solved, and efficient and accurate labeling data scoring and false positive recognition were achieved.

CN119942315AActive Publication Date: 2025-05-06ZHEJIANG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510412530.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-06
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

False positive marking often occurs during manual labeling, resulting in reduced performance, overfitting and unstable training during model training. Manually checking false positive marking is costly and inefficient, which cannot meet the needs of rapid iteration.

Method used

Provide a method of scoring the annotation data, by obtaining the target image, constructing multiple-choice questions and judgment questions, and inputting them into the visual language model for question-and-answer questions, and calculating the average score based on the correct number of times to evaluate the accuracy of the annotation data of the target box.

Benefits of technology

It improves the accuracy of target recognition, reduces misjudgment, provides a more reliable basis for the quality evaluation of target detection annotations, and improves the recognition efficiency of false positive annotations through automated processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942315A_ABST
    Figure CN119942315A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses an annotation data scoring method and a false positive annotation data identification method based on the same, and the scoring method comprises the steps: obtaining a target image containing a target object; repeatedly constructing a selection question based on the target category name and the interference category name set, and inputting the selection question and the target image into at least one visual language model for question and answer; judging questions based on the target category names and the target images are repeatedly generated from a pre-configured judging question template pool, and the judging questions and the target images are input into at least one visual language model for question answering; and determining an average score of all the visual language models on the basis of the correct answer selection times and the correct answer judgment times, and taking the average score as a target score corresponding to the target frame annotation data in the target image. According to the technical scheme provided by the invention, the accuracy of marking data scoring can be improved, and the comprehensiveness and accuracy of an evaluation result are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for scoring labeled data and a method for identifying false positive labeled data based thereon. Background Art

[0002] With the development of deep learning, the demand for high-quality labeled data in target detection tasks has increased. However, false positive annotations (mistakenly labeling non-target objects as targets) often occur during manual annotation, which can lead to performance degradation, overfitting, and unstable training during model training. Manual inspection of false positive annotations is costly and inefficient, and cannot meet the needs of rapid iteration.

[0003] Therefore, how to automatically score the labeled data efficiently and accurately is an urgent problem to be solved in order to provide support for subsequent false positive annotation identification, thereby improving the efficiency and quality of target detection dataset production. Summary of the invention

[0004] The present application provides a method for scoring labeled data and a method for identifying false positive labeled data based thereon, which achieves the technical effect of improving the accuracy of labeling data scoring and ensuring the comprehensiveness and accuracy of evaluation results.

[0005] In order to achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, an embodiment of the present application provides a method for scoring labeled data, the scoring method comprising: Acquire a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, and the target frame annotation data includes a target category name; Repeatedly construct multiple-choice questions based on the target category name and the set of interference category names, and input the multiple-choice questions and the target image into at least one visual language model for question and answer, and determine the number of correct answers corresponding to the visual language model; Repeatedly generating a judgment question based on the target category name and the target image from a pre-configured judgment question template pool, and inputting the judgment question and the target image into at least one visual language model for question answering, and determining the number of correct judgment answers corresponding to the visual language model; Based on the number of correct answer selections and the number of correct answer determinations, an average score of all the visual language models is determined, and the average score is used as a target score corresponding to the target box annotation data in the target image.

[0006] The present embodiment provides a method for scoring annotated data, which improves the accuracy of target recognition by combining a visual language model. Specifically, the visual language model is used to answer multiple-choice questions and true-or-false questions for the target image, and the recognition accuracy of the target annotation frame data is verified based on the results of these questions and answers. In this way, misjudgments are reduced. Finally, by calculating the number of correct answers of each visual language model in the multiple-choice questions and true-or-false questions, the score of each target frame annotation data is comprehensively obtained, thereby providing a more reliable basis for the quality assessment of target detection annotation.

[0007] In one embodiment, the target box annotation data includes a target bounding box; repeatedly constructing a multiple-choice question based on the target category name and the interference category name set, inputting the multiple-choice question and the target image into at least one visual language model for question answering, and determining the number of correct answers corresponding to the visual language model, comprises: Multiple-choice question construction step: constructing a multiple-choice question with the coordinate information of the target bounding box in the target image and the target image, constructing answer options with the target category name and the interference category name set, and obtaining a multiple-choice question including the multiple-choice question and the answer options; Multiple-choice question answering step: inputting the multiple-choice question and the target image into at least one visual language model, and obtaining a multiple-choice answer corresponding to each visual language model; Repeat the multiple-choice question construction step and the multiple-choice question answering step until a selection answer of at least one visual language model in at least one multiple-choice question is obtained, and the number of correct selection answers corresponding to each visual language model is obtained.

[0008] This embodiment generates multiple-choice questions based on the coordinate information of the image and the target image. The question content is set according to the description in the image. The answer options include the correct category and the interference category, and the order of the answer options is randomized. Then, the multiple-choice question and the target image are input into the visual language model together. The model selects the most appropriate answer by analyzing the image content. Finally, by constructing and answering the multiple-choice questions multiple times, the model's answers are recorded.

[0009] In one embodiment, the method for acquiring the interference category name set includes: Acquire an image data set containing the target image; wherein the image data set includes a set of category names of different target images; Selecting at least one interference category name different from the target category name from the category name set; wherein the multiple interference category names are different from each other; Determine a preset independent category name option as the interference category name; All the interference category names are combined to obtain the interference category name set.

[0010] This embodiment obtains data from an image dataset containing multiple target images and category names. Then, interference category names different from the target category are selected. These interference category names may be similar to the target category name, but not the same. Then, the preset independent category name options are determined as interference category names to further enrich the selection of interference category names. Finally, these interference category names are integrated into a set of interference category names for subsequent scoring.

[0011] In one embodiment, the target box annotation data includes a target bounding box; repeatedly generating a judgment question based on the target category name and the target image from a preconfigured judgment question template pool, inputting the judgment question and the target image into at least one visual language model for question answering, and determining the number of correct judgment answers corresponding to the visual language model, comprises: True or False question construction step: generating a true or false question from a pre-constructed true or false question template pool based on the coordinate information of the target bounding box in the target image, the target image and the target category name; Judgment question answering step: inputting the judgment question and the target image into at least one visual language model, and obtaining a judgment answer corresponding to each visual language model; Repeat the judgment question construction step and the judgment question answering step until a judgment answer of at least one visual language model in at least one judgment question is obtained, and the number of correct judgment answers corresponding to each visual language model is obtained.

[0012] This embodiment generates judgment questions from a pre-built judgment question template pool based on the coordinate information in the target image, the target image itself, and the target category name. Then, these judgment questions and the target image are input into at least one visual language model, and the model generates judgment answers based on the image content. This process is repeated multiple times until the answers of each visual language model are obtained in multiple rounds of judgment questions, and the number of correct answers of each model is counted. Improve the accuracy of the evaluation of the target box annotation data.

[0013] In one embodiment, determining the average score of all the visual language models based on the number of correct answer selections and the number of correct answer determinations includes: For any visual language model, based on the number of correct selected answers, the number of correct judged answers and the number of repetitions, determine a comprehensive score corresponding to any visual language model; Based on the comprehensive score, an average score of all the visual language models is determined.

[0014] In this embodiment, the total correct number of visual language models can be obtained by adding up the number of correct answer selections and the number of correct judgment answers. The comprehensive score corresponding to each visual language model is obtained by dividing by the total number of repetitions (i.e., the sum of the number of repetitions of the multiple-choice questions and the number of repetitions of the judgment questions). The comprehensive scores of all visual language models are averaged to obtain the target score for the target box annotation data. This average score can reflect the overall judgment of all visual language models on the target box annotation data, providing a more accurate basis for subsequent false positive evaluation.

[0015] In one embodiment, determining the average score of all the visual language models includes: Among them, z is the average score; n is the number of visual language models; x i is the number of correct answers selected by the i-th visual language model; w i is the number of correct answers corresponding to the i-th visual language model; k is the number of repetitions of the multiple-choice question; j is the number of repetitions of the judgment question.

[0016] In a second aspect, an embodiment of the present application provides a method for identifying false positive labeled data, the identification method comprising: Acquire an image data set; wherein the image data set includes a plurality of target images, each of the target images is annotated with target frame annotation data to be identified corresponding to the target object; For the target frame annotation data in any target image, scoring is performed using the scoring method described above to obtain a target score corresponding to the target frame annotation data in the target image; When the target score is lower than a preset score threshold, the target frame annotation data corresponding to the target score is determined as false positive annotation data.

[0017] The present embodiment provides a method for identifying false positive annotated data, in which the acquired image data set contains multiple target images, each image is annotated with target frame annotation data to be identified. Each target frame annotation data is scored to obtain a target score. If the target score is lower than a preset score threshold, it is determined to be false positive annotation data. This automated process not only reduces the cost and workload of manual inspection, but also greatly improves the recognition efficiency of false positive annotations, especially in the processing of large-scale data sets, it can quickly complete screening and ensure the high quality of the data set. At the same time, it can effectively support the production of large-scale target detection data sets, improve the efficiency of data set production, and provide more accurate and reliable data for subsequent model training.

[0018] In a third aspect, an embodiment of the present application provides a scoring device for labeled data, the device comprising: An image acquisition unit, configured to acquire a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, wherein the target frame annotation data includes a target category name; A multiple-choice question construction and question-answering unit, used to repeatedly construct multiple-choice questions based on the target category name and the interference category name set, and input the multiple-choice questions and the target image into at least one visual language model for question-answering, and determine the number of correct answers corresponding to the visual language model; A judgment question construction and question-answering unit, used to repeatedly generate judgment questions based on the target category name and the target image from a pre-configured judgment question template pool, and input the judgment question and the target image into at least one visual language model for question-answering, and determine the number of correct judgment answers corresponding to the visual language model; The model scoring unit is used to determine the average score of all the visual language models based on the number of correct answer selections and the number of correct answer judgments, and use the average score as the target score corresponding to the target box annotation data in the target image.

[0019] In a fourth aspect, an embodiment of the present application provides a computer device, including: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the above-mentioned scoring method or the above-mentioned identification method by executing the computer instructions.

[0020] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the above-mentioned scoring method or the above-mentioned identification method. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A flowchart of a method for scoring labeled data provided in an embodiment of the present application; Figure 2 A flowchart of step S3 provided in an embodiment of the present application; Figure 3 A flowchart of a method for obtaining a set of interference category names provided in an embodiment of the present application; Figure 4 A flowchart of step S5 provided in an embodiment of the present application; Figure 5 A flowchart of step S7 provided in an embodiment of the present application; Figure 6 A flowchart of a false positive labeled data identification method provided in an embodiment of the present application; Figure 7 A block diagram of a scoring device for annotated data provided in an embodiment of the present application; Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0024] With the rapid development of deep learning in computer vision, natural language processing, and speech recognition, the demand for training data is increasing. Especially in the task of object detection, high-quality labeled data is crucial. Manual labeling of object detection datasets usually includes steps such as target object identification, positioning, bounding box drawing, and target category label assignment. However, errors may occur in the labeling process, such as mislabeling non-target objects as targets (false positive labeling) or misclassifying target objects. These problems will introduce noise and affect the model training effect.

[0025] Label noise has a negative impact on model training, including misleading the model to learn the wrong decision boundary, causing performance degradation, or causing overfitting, where the model performs well on the training data but has poor generalization ability. In addition, noisy labels may also increase the instability of the training process, increasing costs and difficulty.

[0026] Although the process of manually checking false positive annotations is necessary, it is costly and inefficient. Manual checking requires a lot of human resources, recruitment, training and management, and is time-consuming. It cannot meet the needs of rapid iteration, which in turn affects the progress of research and development.

[0027] Based on the above technical problems, according to an embodiment of the present application, an embodiment of a method for scoring labeled data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.

[0028] In this embodiment, a method for scoring labeled data is provided. Figure 1 A flowchart of a method for scoring labeled data provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the process includes the following steps: Step S1, obtaining a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, and the target frame annotation data includes a target category name.

[0029] Specifically, a target image containing a target object is obtained, and target box annotation data corresponding to the target object is ensured to be annotated in the target image, including the target category name.

[0030] Step S3, repeatedly construct multiple-choice questions based on the target category name and the interference category name set, and input the multiple-choice questions and the target image into at least one visual language model for question and answer, and determine the number of correct answers corresponding to the visual language model.

[0031] Specifically, a distractor category name that is different from the target category name is selected from the target detection dataset. Based on the target category name and the distractor category name set, a multiple-choice question is constructed. The multiple-choice question and the target image are input into at least one visual language model to obtain the answer of the visual language model. The number of times the visual language model selects the correct answer is determined. If the visual language model correctly selects the target category name, it is considered a correct answer; if the distractor category is selected, it is considered an incorrect answer. By counting the number of correct answers of the visual language model, the credibility of the target box annotation data is further analyzed.

[0032] Step S5, repeatedly generate judgment questions based on target category names and target images from a pre-configured judgment question template pool, and input the judgment questions and target images into at least one visual language model for question answering, and determine the number of correct judgment answers corresponding to the visual language model.

[0033] Specifically, a template is randomly selected from a pre-configured judgment question template pool. Based on the target category name and the target image, a judgment question is generated using the template. The judgment question and the target image are input into at least one visual language model to obtain the answer of the visual language model. Determine the number of correct judgments given by the visual language model. By counting the number of correct answers given by the visual language model, the credibility of the target box annotation data can be evaluated. If the visual language model gives correct answers many times, it means that the target box annotation data is relatively accurate and has a high credibility; conversely, if the number of incorrect answers is large, it means that the target box annotation data may have high inaccuracy or errors.

[0034] Step S7, based on the number of correct answer selections and the number of correct answer judgments, determine the average score of all visual language models, and use the average score as the target score corresponding to the target box annotation data in the target image.

[0035] Specifically, based on the number of correct answers to multiple-choice questions and true-or-false questions, the comprehensive score of each visual language model is calculated. The comprehensive scores of all visual language models are averaged to obtain the final average score. The average score is used as the target score corresponding to the target box annotation data in the target image. Through the average score, a unified quantitative indicator for the accuracy of the target box annotation data is provided.

[0036] The present embodiment provides a method for scoring annotated data, which improves the accuracy of target recognition by combining a visual language model. Specifically, the visual language model is used to answer multiple-choice questions and true-or-false questions for the target image, and the recognition accuracy of the target annotation frame data is verified based on the results of these questions and answers. In this way, misjudgments are reduced. Finally, by calculating the number of correct answers of each visual language model in the multiple-choice questions and true-or-false questions, the score of each target frame annotation data is comprehensively obtained, thereby providing a more reliable basis for the quality assessment of annotations for target detection.

[0037] Figure 2 In the flowchart of step S3 provided in the embodiment of the present application, the target box annotation data includes a target bounding box; the process may include the following steps: Step S31, multiple-choice question construction step: construct a multiple-choice question with the coordinate information of the target bounding box in the target image and the target image, construct answer options with the target category name and the interference category name set, and obtain a multiple-choice question including a multiple-choice question and answer options.

[0038] Specifically, based on the coordinate information of the target bounding box in the target image and the target image, an explicit selection question is generated to ask which target category the image content belongs to. For example, “In image I, the content within the coordinate range (100, 100, 300, 300) belongs to which of the following categories?” Then, the target category name H0 and the interference category name set {H1, H2, ..., H m}Construct answer options. Provide multiple options. The set of distractor category names includes answer options that are different from the target category name H0, and the distractor category names in the set are different from each other. In addition, it also includes independent category name options, such as "does not belong to any category in other options."

[0039] It should be noted here that the order of the answer options should be randomly shuffled when generating multiple-choice questions. By shuffling the order of the answer options, fairness can be improved. Assuming that the target category name H0 in the answer options is car, the interference category name H1 is bicycle, H2 is person, H3 is bicycle, and H4 is a category that does not belong to any of the other options, the random shuffling may be {"person", "car", "does not belong to any of the other options", "bicycle"}.

[0040] Once the clear choice and random order answer options have been generated, they can be combined into a complete multiple choice question.

[0041] Step S33, multiple-choice question answering step: input the multiple-choice question and the target image into at least one visual language model, and obtain the multiple-choice answer corresponding to each visual language model.

[0042] Specifically, the constructed multiple-choice questions (including multiple-choice questions and answer options) are converted into text form and passed to the visual language model as input with the target image. The visual language model generates answers based on the input image and text questions. The visual language model combines the content of the target image and the question text, and performs reasoning through the internal visual encoder and language decoder. The word sequence generated by the visual language model is decoded into natural language text to obtain the visual language model's choice answers to the multiple-choice questions. Each visual language model generates an answer for a multiple-choice question input once. Preferably, the visual language model can adopt models such as the Qwen-VL series, the Intern-VL series, and the deepseek-VL series.

[0043] Step S35, repeatedly executing the multiple-choice question construction step and the multiple-choice question answering step until obtaining the selection answer of at least one visual language model in at least one multiple-choice question, and obtaining the number of correct selection answers corresponding to each visual language model.

[0044] Specifically, by repeatedly executing the multiple-choice question construction and question-answering steps, multiple responses can be obtained, thereby improving the accuracy of the evaluation results.

[0045] This embodiment generates multiple-choice questions based on the coordinate information of the image and the target image. The question content is set according to the description in the image. The answer options include the correct category and the interference category, and the order of the answer options is randomized. Then, the multiple-choice question and the target image are input into the visual language model together. The model selects the most appropriate answer by analyzing the image content. Finally, by constructing and answering the multiple-choice questions multiple times, the model's answers are recorded.

[0046] Figure 3 A flowchart of a method for obtaining a set of interference category names provided in an embodiment of the present application, the process may include the following steps: Step S301, obtaining an image data set containing target images; wherein the image data set includes a set of category names of different target images.

[0047] Specifically, the image dataset is the core of the target detection task. Each target image in the image dataset is a crucial label information in the target detection task. These label information clearly defines the location of the target object in the target image and its corresponding target category name. The category name set is the set of all different target category names in the image dataset, which constitutes the label space of the target detection task.

[0048] Step S303: selecting from the category name set at least one interference category name different from the target category name; wherein the multiple interference category names are different from each other.

[0049] Specifically, the interference category name is a category different from the target category name, and is used to construct the multiple-choice question. Specifically, the interference category name must be different from the target category name. Multiple interference category names cannot be repeated to ensure that each interference item is independent. The interference category name can include categories that are similar to the target category name and categories that are not similar. Preferably, the selection of the interference category name can be random selection or selection based on heuristic rules.

[0050] Step S305: determine the preset independent category name option as the interference category name.

[0051] Specifically, a preset independent category name option is introduced, and the independent category name option may be "not belonging to any category of other options".

[0052] Step S307: All interference category names are combined to obtain an interference category name set.

[0053] Specifically, by synthesizing all interference category names and obtaining a set of interference category names, a complete multiple-choice question can be constructed, which not only increases the difficulty and diversity of the questions and answers, but also improves the ability to judge false positive labels in real scenarios.

[0054] This embodiment obtains data from an image dataset containing multiple target images and category names. Then, interference category names different from the target category are selected. These interference category names may be similar to the target category name, but not the same. Then, the preset independent category name options are determined as interference category names to further enrich the selection of interference category names. Finally, these interference category names are integrated into a set of interference category names for subsequent scoring.

[0055] Figure 4 In the flowchart of step S5 provided in the embodiment of the present application, the target box annotation data includes a target bounding box; the process may include the following steps: Step S51, judgment question construction step: based on the coordinate information of the target bounding box in the target image, the target image and the target category name, generate a judgment question from a pre-constructed judgment question template pool.

[0056] Specifically, the judgment question template pool is a collection that contains multiple different judgment question templates. Each template represents a questioning method, which forms different judgment questions for coordinate information, target image, and target category name. The design of different templates can examine the content of the target category from multiple angles and avoid a single questioning method. For example, one template may directly ask whether the target image meets the target category name, while another template asks questions by emphasizing the exclusion of other categories.

[0057] For example, Template 1: “Does the content in target image I belong to the target category name?” This question directly asks whether the content of the target image fits a certain category. Template 2: “Based on target image I, determine whether the content of the image is consistent with the target category name, rather than other categories?” Step S53, judgment question answering step: input the judgment question and the target image into at least one visual language model, and obtain the judgment answer corresponding to each visual language model.

[0058] Specifically, the constructed judgment question is converted into text form and passed to the visual language model as input with the target image. The visual language model generates an answer based on the input image and text question. The visual language model combines the content of the target image and the question text, and performs reasoning through the internal visual encoder and language decoder. The word sequence generated by the visual language model is decoded into natural language text to obtain the visual language model's selection answer to the judgment question. Each visual language model generates an answer for a judgment question input once. Preferably, the visual language model can adopt models such as the Qwen-VL series, the Intern-VL series, and the deepseek-VL series.

[0059] Step S55, repeatedly executing the judgment question construction step and the judgment question questioning step until obtaining the judgment answer of at least one visual language model in at least one judgment question, and obtaining the number of correct judgment answers corresponding to each visual language model.

[0060] Specifically, by repeatedly executing the judgment question construction and question-answering steps, multiple answers can be obtained, thereby improving the accuracy of the evaluation results.

[0061] This embodiment generates judgment questions from a pre-built judgment question template pool based on the coordinate information in the target image, the target image itself, and the target category name. Then, these judgment questions and the target image are input into at least one visual language model, and the model generates judgment answers based on the image content. This process is repeated multiple times until the answers of each visual language model are obtained in multiple rounds of judgment questions, and the number of correct answers of each model is counted. Improve the accuracy of the evaluation of the target box annotation data.

[0062] Figure 5 The flowchart of step S7 provided in the embodiment of the present application may include the following steps: Step S71, for any visual language model, based on the number of correct answer selections, the number of correct answer judgments and the number of repetitions, a comprehensive score corresponding to any visual language model is determined.

[0063] Specifically, the comprehensive score can be determined by the number of correct choices, the number of correct judgments, and the number of repetitions. By adding the number of correct answers to multiple-choice questions and true-or-false questions and dividing by the total number of repetitions, a score between 0 and 1 can be obtained to evaluate the accuracy of the target box annotation data. The higher the score, the higher the accuracy of the target box annotation data. Conversely, the lower the accuracy of the target box annotation data.

[0064] Step S73: Determine the average score of all visual language models based on the comprehensive score.

[0065] Specifically, the average score of all visual language models is obtained by calculating the average of the comprehensive scores of all visual language models.

[0066] In some preferred embodiments, determining the average score of all visual language models includes: Among them, z is the average score; n is the number of visual language models; x i is the number of correct answers selected by the i-th visual language model; w i is the number of correct answers corresponding to the i-th visual language model; k is the number of repetitions of the multiple-choice question; j is the number of repetitions of the judgment question.

[0067] In this embodiment, the total correct number of visual language models can be obtained by adding up the number of correct answer selections and the number of correct judgment answers. The comprehensive score corresponding to each visual language model is obtained by dividing by the total number of repetitions (i.e., the sum of the number of repetitions of the multiple-choice questions and the number of repetitions of the judgment questions). The comprehensive scores of all visual language models are averaged to obtain the target score for the target box annotation data. This average score can reflect the overall judgment of all visual language models on the target box annotation data, providing a more accurate basis for subsequent false positive evaluation.

[0068] In this embodiment, a method for identifying false positive labeled data is provided. Figure 6 A flowchart of a false positive labeling data identification method provided in an embodiment of the present application is as follows: Figure 6 As shown, the process includes the following steps: Step S2, obtaining an image data set; wherein the image data set includes a plurality of target images, each target image is annotated with target frame annotation data to be identified corresponding to the target object.

[0069] Step S4, for the target frame annotation data in any target image, score using the scoring method of steps S1 to S7 to obtain a target score corresponding to the target frame annotation data in the target image.

[0070] Step S6: when the target score is lower than a preset score threshold, the target frame annotation data corresponding to the target score is determined as false positive annotation data.

[0071] Specifically, a labeled image dataset is provided for subsequent scoring and false positive labeling detection. For each target image in the image dataset, its target frame labeling data is extracted, and each target frame labeling data is scored using the scoring method of steps S1 to S7 to obtain a target score corresponding to each target frame labeling data. A preset score threshold (such as 0.5 or 0.6) is set, and the target frame labeling data whose target score is lower than the preset score threshold is determined as false positive labeling data.

[0072] The present embodiment provides a method for identifying false positive annotated data, in which the acquired image data set contains multiple target images, each image is annotated with target frame annotation data to be identified. Each target frame annotation data is scored to obtain a target score. If the target score is lower than a preset score threshold, it is determined to be false positive annotation data. This automated process not only reduces the cost and workload of manual inspection, but also greatly improves the recognition efficiency of false positive annotations, especially in the processing of large-scale data sets, it can quickly complete screening and ensure the high quality of the data set. At the same time, it can effectively support the production of large-scale target detection data sets, improve the efficiency of data set production, and provide more accurate and reliable data for subsequent model training.

[0073] Accordingly, please refer to Figure 7 A block diagram of a scoring device for annotated data provided in an embodiment of the present application, the terminal includes: An image acquisition unit 101 is used to acquire a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, and the target frame annotation data includes a target category name; A multiple-choice question construction and question-answering unit 103 is used to repeatedly construct multiple-choice questions based on a set of target category names and interference category names, and input the multiple-choice questions and target images into at least one visual language model for question-answering, and determine the number of correct answers corresponding to the visual language model; The judgment question construction and question-answering unit 105 is used to repeatedly generate judgment questions based on target category names and target images from a pre-configured judgment question template pool, and input the judgment questions and target images into at least one visual language model for question-answering, and determine the number of correct judgment answers corresponding to the visual language model; The model scoring unit 107 is used to determine the average score of all visual language models based on the number of correct answer selections and the number of correct answer judgments, and use the average score as the target score corresponding to the target box annotation data in the target image.

[0074] In some optional implementations, the target box annotation data includes a target bounding box; the multiple-choice question construction and question-answering unit 103 includes: Multiple-choice question construction step: construct a multiple-choice question using the coordinate information of the target bounding box in the target image and the target image, and construct answer options using the target category name and the interference category name set to obtain a multiple-choice question including a multiple-choice question and answer options; Multiple-choice question answering step: input the multiple-choice question and the target image into at least one visual language model, and obtain the corresponding multiple-choice answer of each visual language model; Repeat the multiple-choice question construction step and the multiple-choice question answering step until a selection answer of at least one visual language model in at least one multiple-choice question is obtained, and the number of correct selection answers corresponding to each visual language model is obtained.

[0075] In some optional implementations, the method for obtaining the interference category name set includes: Acquire an image dataset containing target images; wherein the image dataset includes a set of category names of different target images; Selecting from the category name set at least one interference category name different from the target category name; wherein the multiple interference category names are different from each other; Determine the preset independent category name option as the interference category name; All interference category names are combined to obtain an interference category name set.

[0076] In some optional implementations, the target box annotation data includes a target bounding box; the judgment question construction and question answering unit 105 includes: True or False Question Construction Step: Generate a true or false question from a pre-constructed true or false question template pool based on the coordinate information of the target bounding box in the target image, the target image, and the target category name; True or False Question Answering Step: Input the true or false question and the target image into at least one visual language model, and obtain the true or false answer corresponding to each visual language model; Repeat the judgment question construction step and the judgment question questioning step until a judgment answer of at least one visual language model in at least one judgment question is obtained, and the number of correct judgment answers corresponding to each visual language model is obtained.

[0077] In some optional implementations, the model scoring unit 107 includes: For any visual language model, based on the number of correct answer selections, the number of correct answer judgments, and the number of repetitions, a comprehensive score corresponding to any visual language model is determined; Based on the comprehensive score, the average score of all visual language models is determined.

[0078] In some optional implementations, determining the average score of all visual language models includes: Among them, z is the average score; n is the number of visual language models; x i is the number of correct answers selected by the i-th visual language model; w i is the number of correct answers corresponding to the i-th visual language model; k is the number of repetitions of the multiple-choice question; j is the number of repetitions of the judgment question.

[0079] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0080] In this embodiment, a scoring device for annotated data is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0081] See also Figure 8 , Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 10 is taken as an example.

[0082] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0083] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0084] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0085] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0086] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0087] The embodiment of the present application also provides a computer-readable storage medium. The above method according to the embodiment of the present application can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0088] The devices or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0089] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0090] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods and devices. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0091] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices and apparatuses according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0092] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0094] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0095] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0096] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

[0097] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for scoring labeled data, characterized in that: The scoring method includes: Acquire a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, and the target frame annotation data includes a target category name; Repeatedly construct multiple-choice questions based on the target category name and the set of interference category names, and input the multiple-choice questions and the target image into at least one visual language model for question and answer, and determine the number of correct answers corresponding to the visual language model; Repeatedly generating a judgment question based on the target category name and the target image from a pre-configured judgment question template pool, and inputting the judgment question and the target image into at least one visual language model for question answering, and determining the number of correct judgment answers corresponding to the visual language model; Based on the number of correct answer selections and the number of correct answer determinations, an average score of all the visual language models is determined, and the average score is used as a target score corresponding to the target box annotation data in the target image.

2. The scoring method according to claim 1, characterized in that: The target box annotation data includes a target bounding box; repeatedly constructing a multiple-choice question based on the target category name and the interference category name set, inputting the multiple-choice question and the target image into at least one visual language model for question answering, and determining the number of correct answers corresponding to the visual language model, including: Multiple-choice question construction step: constructing a multiple-choice question with the coordinate information of the target bounding box in the target image and the target image, constructing answer options with the target category name and the interference category name set, and obtaining a multiple-choice question including the multiple-choice question and the answer options; Multiple-choice question answering step: inputting the multiple-choice question and the target image into at least one visual language model, and obtaining a multiple-choice answer corresponding to each visual language model; Repeat the multiple-choice question construction step and the multiple-choice question answering step until a selection answer of at least one visual language model in at least one multiple-choice question is obtained, and the number of correct selection answers corresponding to each visual language model is obtained.

3. The scoring method according to claim 1 or 2, characterized in that: The method for obtaining the interference category name set includes: Acquire an image data set containing the target image; wherein the image data set includes a set of category names of different target images; Selecting at least one interference category name different from the target category name from the category name set; wherein the multiple interference category names are different from each other; Determine a preset independent category name option as the interference category name; All the interference category names are combined to obtain the interference category name set.

4. The scoring method according to claim 1, characterized in that: The target box annotation data includes a target bounding box; repeatedly generating a judgment question based on the target category name and the target image from a pre-configured judgment question template pool, inputting the judgment question and the target image into at least one visual language model for question answering, and determining the number of correct judgment answers corresponding to the visual language model, including: True or False question construction step: generating a true or false question from a pre-constructed true or false question template pool based on the coordinate information of the target bounding box in the target image, the target image and the target category name; Judgment question answering step: inputting the judgment question and the target image into at least one visual language model, and obtaining a judgment answer corresponding to each visual language model; Repeat the judgment question construction step and the judgment question answering step until a judgment answer of at least one visual language model in at least one judgment question is obtained, and the number of correct judgment answers corresponding to each visual language model is obtained.

5. The scoring method according to claim 1, characterized in that: The determining, based on the number of correct answer selections and the number of correct answer determinations, an average score of all the visual language models comprises: For any visual language model, based on the number of correct selected answers, the number of correct judged answers and the number of repetitions, determine a comprehensive score corresponding to any visual language model; Based on the comprehensive score, an average score of all the visual language models is determined.

6. The scoring method according to claim 1, characterized in that: Determining the average score of all the visual language models comprises: Among them, z is the average score; n is the number of visual language models; x i is the number of correct answers selected by the i-th visual language model; w i is the number of correct answers corresponding to the i-th visual language model; k is the number of repetitions of the multiple-choice question; j is the number of repetitions of the judgment question.

7. A method for identifying false positive labeled data, characterized in that: The identification method comprises: Acquire an image data set; wherein the image data set includes a plurality of target images, each of the target images is annotated with target frame annotation data to be identified corresponding to the target object; For the target frame annotation data in any target image, scoring is performed using the scoring method according to any one of claims 1 to 6 to obtain a target score corresponding to the target frame annotation data in the target image; When the target score is lower than a preset score threshold, the target frame annotation data corresponding to the target score is determined as false positive annotation data.

8. A scoring device for annotated data, characterized in that: The device comprises: An image acquisition unit, configured to acquire a target image containing a target object, wherein the target image is annotated with target frame annotation data corresponding to the target object, wherein the target frame annotation data includes a target category name; A multiple-choice question construction and question-answering unit, used to repeatedly construct multiple-choice questions based on the target category name and the interference category name set, and input the multiple-choice questions and the target image into at least one visual language model for question-answering, and determine the number of correct answers corresponding to the visual language model; A judgment question construction and question-answering unit, used to repeatedly generate judgment questions based on the target category name and the target image from a pre-configured judgment question template pool, and input the judgment question and the target image into at least one visual language model for question-answering, and determine the number of correct judgment answers corresponding to the visual language model; The model scoring unit is used to determine the average score of all the visual language models based on the number of correct answer selections and the number of correct answer judgments, and use the average score as the target score corresponding to the target box annotation data in the target image.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the scoring method according to any one of claims 1 to 6 or the identification method according to claim 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the scoring method according to any one of claims 1 to 6 or the identification method according to claim 7.

Citation Information

Patent Citations

  • Question identification method

    CN118781612A

  • Saliency for anchor-based object detection

    US20240054754A1

  • A method for determining a label of a fall event

    WO2024146800A1