Three-dimensional visual anaphora data quality evaluation method based on multi-role multi-view reasoning

By employing a multi-role, multi-perspective reasoning method, a multimodal large-scale model thinking chain of jury-judge is constructed to solve the problems of logical paradoxes and referential ambiguities in 3D visual referential data. This enables efficient and automated data quality assessment, improving the quality of dataset annotation and the reliability of model training.

CN120833341AInactive Publication Date: 2025-10-24ZHEJIANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511340702.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-10-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing 3D visual reference datasets contain a large number of toxic samples, mainly caused by logical paradoxes and ambiguities in reference. These samples are difficult to automatically identify and remove, leading to misleading model training and evaluation. Furthermore, manual review is time-consuming and labor-intensive.

Method used

Employing a multi-role, multi-perspective reasoning approach, this study constructs a multimodal large-scale model of jury-judge thinking chain through logic, descriptive consistency, distinguishability, and ambiguity analysis. This model enables multi-stage evaluation and re-evaluation, eliminating toxic samples and improving data quality.

Benefits of technology

It effectively identifies and eliminates logical paradoxes and referential ambiguities, improves the quality of dataset annotation and the reliability of model training, and approaches the level of human expert judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833341A_ABST
    Figure CN120833341A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional visual anaphora data quality evaluation method based on multi-role multi-view reasoning, and the method comprises the steps: constructing a partner-judge chain reasoning framework of a similar judicial trial mechanism, carrying out the multi-angle analysis of a sample through a plurality of partner units from four core dimensions, namely logicality, consistency, distinguishability and ambiguity, and forming multiple judgment opinions; the regional judge units integrate the accompanying results in respective analysis regions to ensure the consistency and stability of judgment; and finally, arbitrating global opinions by a final examination judge unit in combination with an evidence-based refining strategy, adaptively recombining visual and text evidences, and correcting inference uncertainty caused by incomplete observation or information deviation. According to the method, high-accuracy data quality judgment can be realized in a complex three-dimensional scene without depending on fine adjustment of a specific task or external perception, the judgment capability equivalent to manual judgment is realized in a three-dimensional visual anaphora task, and the performance and robustness of a model on a screened data set can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of three-dimensional visual understanding and data quality evaluation, and particularly relates to a three-dimensional visual reference data quality evaluation method based on multi-role and multi-view reasoning. BACKGROUND

[0002] Three-dimensional visual reference (3D Visual Grounding, 3DVG) aims to locate a unique target object in a complex three-dimensional scene according to a natural language description, and is an important task connecting human language understanding and machine environment perception. The task is widely applied in fields such as robot interaction, augmented reality, virtual reality and embodied intelligence, and often relies on large-scale artificial annotation datasets (such as ScanRefer) for model training and evaluation.

[0003] In a typical 3DVG dataset, a scene is usually composed of multi-view images and sparse point clouds, and an annotator needs to write a natural language description for a target object under limited visibility and incomplete three-dimensional observation conditions. However, due to heavy annotation burden, insufficient scene observation and imperfect language expression, there are inevitably a large number of “toxic samples” in the dataset, i.e. false annotations that mislead model training and evaluation. Such toxic samples mainly come from two aspects: (1) logical paradox: the description has internal contradictions or symmetry ambiguity, making the target unable to be uniquely determined; (2) reference ambiguity: the description is too general or lacks key information, resulting in multiple similar objects simultaneously satisfying the description conditions. However, further manual review of complex 3DVG data will consume a lot of manpower and have low economic benefits. Therefore, an automatic evaluation method based on multi-modal large models is urgently needed, which can comprehensively consider multi-view images and natural language descriptions at the scene level, and effectively identify and remove toxic samples through multi-role and multi-stage reasoning, to improve the annotation quality of 3DVG datasets and the reliability of model training, and solve the problem of automatic identification of toxic samples caused by annotation logical paradox and reference ambiguity in the prior art. SUMMARY

[0004] To solve the problems of the prior art and achieve efficient evaluation of complex three-dimensional visual reference information, the application adopts the following technical solutions:

[0005] A 3D visual referent data quality assessment method based on multi-role multi-view reasoning obtains 3D visual referent data, including multi-view 3D images and text descriptions of target objects in the images. A reasoning model is constructed based on the multiple roles. A first role performs multi-view analysis based on the logic and description consistency of the text descriptions and the distinguishability and ambiguity of the images. A second role performs auxiliary information splicing and re-evaluation on suspicious results of the distinguishability analysis, re-evaluates suspicious 3D visual referent data from the ambiguity analysis, and makes a comprehensive judgment based on the results of the logic and description consistency analysis, as well as the re-evaluation results of the distinguishability and ambiguity analysis, to obtain the quality and reasoning interpretation of the 3D visual referent data. The comprehensive judgment does not directly access the original multi-view images or text descriptions to maintain the interpretability of the reasoning chain and avoid information interference. The second role performs local refinement and global judgment based on the first role's evaluation. This allows the reasoning model to perform comprehensive and integrated judgment on scene-level high-proximity multi-frame images and their corresponding complex natural language referent information for 3D visual referent tasks by simply injecting prompt words for few-shot case learning and specifying the input and output formats, thereby achieving targeted complex information evaluation.

[0006] Furthermore, the logic analysis of the first role is based on a contextual learning strategy, and uses text logic prompts to perform a logic completeness analysis on the text description to detect logic problems and obtain a logic analysis result including a logic judgment result and a reasoning explanation;

[0007] The description consistency analysis of the first role is to perform consistency analysis on multiple text descriptions of the target object, construct a description group for multiple text descriptions of the same target object, perform logical screening within the description group to eliminate low-confidence text descriptions, and then perform cross-consistency verification on the current text description and the screened description group;

[0008] The distinguishability analysis of the first role is performed on a set of images from the positive sample perspectives of the target object, determining the presence of the target object in each perspective image and filtering out low-confidence perspective images, and then determining whether the target object exists uniquely based on the remaining perspective images as positive samples;

[0009] The ambiguity analysis of the first role is to perform existence determination based on the interference perspective image set to determine whether there are other objects of the same category as the target object in the scene, so as to detect potential object reference ambiguity.

[0010] Furthermore, the second role identifies suspicious samples and high-confidence auxiliary samples from the distinguishability analysis results to construct suspicious auxiliary sample pairs, splices and / or combines the visual images of the suspicious samples and auxiliary samples to complete the occluded or missing parts, and then re-evaluates the existence and / or uniqueness to improve the evaluation accuracy.

[0011] Further, the second role directly screens suspicious samples from the ambiguity analysis result, and re-evaluates the existence evaluation result and the corresponding reason obtained by introducing the corroborative refinement mechanism to replace the original existence evaluation result and the corresponding reason.

[0012] Further, the comprehensive determination of the second role is based on the description consistency analysis result and the distinguishability analysis result of the first role, and the re-evaluation result of the distinguishability analysis result and the ambiguity analysis result of the first role by the second role, combined with the comprehensive determination prompt word, to obtain the final data quality result and the push interpretation. Since the original multi-view image or text description is not directly accessed, the explainability of the reasoning chain can be maintained and information interference can be avoided.

[0013] Further, the consistency analysis of the first role is based on the prompt word of the text consistency evaluation and the plurality of text descriptions of the target object, to evaluate the consistency of the target object to obtain the consistency evaluation result and the corresponding reason. High-reliability text descriptions are screened from the consistency evaluation result by a confidence threshold to construct a set of filtered text descriptions of the target object. The evaluation of the target object is not limited to the judgment of a single text, but obtains core information from the observations of multiple annotators to reduce misjudgment caused by the expression errors of a single annotator. The text description of the current target object is evaluated based on the prompt word in the cross-validation process and the set of text descriptions to obtain the consistency evaluation result of the text description of the current target object and the corresponding reason.

[0014] Further, the distinguishability analysis of the first role is based on the positive sample visual image and the existence prompt word to evaluate the existence of the target object of the positive sample visual image to obtain the existence evaluation result, and the positive sample visual image with high existence is screened by an existence judgment threshold. The uniqueness of the target object visual image is evaluated based on the filtered positive sample visual image and the uniqueness prompt word to obtain the uniqueness evaluation result and the corresponding reason.

[0015] Further, the ambiguity analysis of the first role is based on the interference visual image and the existence prompt word to evaluate the existence of the target object of the interference visual image to obtain the existence evaluation result and the corresponding reason.

[0016] Further, the distinguishability analysis of the second role is performed by introducing a corroborative refining mechanism, constructing a suspicious auxiliary sample pair, and re-evaluating by splicing operation to replace the original existence and uniqueness evaluation results; then, the updated existence evaluation results are combined with the updated uniqueness evaluation results, the updated existence evaluation reasons are combined with the updated uniqueness evaluation reasons, and the final distinguishability evaluation results and corresponding reasons of the positive sample visual image are obtained. That is, the partial result judgment error or illusion caused by the partial occlusion of the view angle and the incomplete shooting of the multi-modal large model in the judgment process can be effectively reduced, and the influence of the singular result on the final score can be reduced.

[0017] The advantages and beneficial effects of the present application are:

[0018] The present application constructs a jury-judge multi-role collaborative thinking chain reasoning mechanism, combines multi-dimensional evaluations such as logic, consistency, distinguishability, and ambiguity, and the corroborative refining and re-evaluation strategy of the regional judge, and realizes the accurate identification of the potential logical paradox and reference ambiguity in the 3D visual reference data. The present method effectively improves the data labeling quality and evaluation robustness, and establishes a new paradigm for data quality discrimination for 3D visual reference tasks. The experimental results show that the 3DVG sample screening effect of the present application is close to the judgment level of human experts. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 The figure is a schematic diagram of the method in the embodiment of the present application.

[0020] Figure 2 The figure is a schematic diagram of the jury thinking chain constructed in the embodiment of the present application.

[0021] Figure 3 The figure is a schematic diagram of the judge thinking chain constructed in the embodiment of the present application.

[0022] Figure 4 The figure is an example effect diagram in the embodiment of the present application. DETAILED DESCRIPTION

[0023] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0024] The three-dimensional visual reference data quality evaluation method based on multi-role and multi-view reasoning includes the following steps:

[0025] Step 1: Obtain a three-dimensional visual reference (3DVG) data sample, the sample including a multi-view scene image set and a natural language description of the target object ; wherein each image Corresponding to a viewing area in a three-dimensional scene, Used to refer to a unique target object ,This step provides the input data basis for subsequent multimodal reasoning.

[0026] Step 2: Construct a jury-judge multimodal large model thinking chain reasoning model The system includes multiple juror roles and multiple judge roles based on the multimodal large model (MLLM); specifically, there are four types of jurors: With three judge characters Each juror receives input samples in parallel and conducts a multi-perspective analysis based on four dimensions: logic, descriptive consistency, distinguishability, and ambiguity. The judge, acting as the judge, makes local refinements and global decisions based on the jurors' evaluations. The jury-judge multimodal large-scale model, a chain of reasoning, requires no fine-tuning. Simply injecting prompts for few-shot case learning and specifying input and output formats allows for comprehensive and integrated judgment of scene-level, highly similar multi-frame images and their corresponding complex natural language referential information in 3DVG tasks, enabling targeted evaluation of complex information.

[0027] Step 3: The text logic juror performs a logical completeness analysis on the natural language description, generating a logical judgment score and reasoning explanation. The logic juror uses a contextual learning strategy to analyze the logical completeness of the natural language description, detecting logical problems such as self-contradiction and mutually exclusive conditions, and outputs the corresponding logical judgment score and reasoning explanation. The logic juror can make judgments directly based on the input description, without relying on external perception modules.

[0028] Specifically, the text logic juror Natural language description Perform a logical integrity analysis to determine whether there are any self-contradictions, circular references, or logical invalidities. is a logical evaluation prompt, the logical decision process is:

[0029]

[0030] in, Represents a logical decision score, Represents the generated inference explanation.

[0031] Step 4: The text consistency jurors conduct a consistency assessment on multiple descriptions of the target object, including logical screening within the description group and cross-validation with the current description. The text consistency jurors conduct a two-stage assessment of the description group consisting of multiple descriptions related to the target object. In the first stage, the description group is logically screened to eliminate low-confidence descriptions. In the second stage, the current description is cross-validated with the filtered description group to determine the consistency between the target description and the descriptions in the same batch.

[0032] Specifically, the text consistency jurors Multiple descriptions of the target object To perform consistency assessment, first filter the submodules Perform logical filtering within the description group to remove low-confidence descriptions:

[0033]

[0034] in, Represents a set of consistency evaluation results for multiple descriptions of the target object , A set of reasons for each evaluation , Indicates the prompt word for text consistency assessment, It means to filter out the high reliability expressions that exceed the confidence threshold in the evaluation results corresponding to each description. represents the confidence threshold for reliability screening, Represents the filtered target object description set;

[0035] Through the text consistency assessment step, the evaluation of the target object's expression is not limited to the judgment of a single text, but can better integrate other data in the dataset, obtain core information from the observations of multiple annotators, and reduce misjudgments caused by errors in the expression of a single annotator.

[0036] Then, by cross-validating the submodule Compare current descriptions With the filtered description group consistency:

[0037]

[0038] in, Indicates the final evaluation result of the text consistency of the current evaluation statement. Indicate the reasons for their evaluation of Represents the prompt word during the cross-validation process.

[0039] Step 5: Two-stage evaluation by distinguishability jurors based on the positive sample view set containing the target object, the first stage determines the existence of the object in the view, which needs to determine the existence of each view and filter low-confidence views; the second stage determines its uniqueness, whether the target object exists uniquely in the remaining view set, to comprehensively form the distinguishability score.

[0040] Specifically, the distinguishability jurors based on the positive sample view set containing the target object perform two-stage evaluation;

[0041] The first stage is existence determination:

[0042]

[0043] wherein, represents the object existence evaluation of each picture in the positive sample view picture set, represents the corresponding reason, represents the existence thought chain part executed by the distinguishability jurors, represents the corresponding prompt word, represents the i-th target object positive sample view image, represents the existence evaluation score of a single picture, represents the threshold value through existence determination, represents the positive sample view set after existence determination;

[0044] The second stage is uniqueness determination:

[0045]

[0046] wherein, represents the object uniqueness evaluation of each picture in the updated positive sample view picture set, represents the corresponding reason, represents the uniqueness thought chain part executed by the distinguishability jurors, represents the corresponding prompt word.

[0047] Step 6: Ambiguity jurors determine whether there are other objects in the scene that meet the description conditions based on the interference view set, to detect potential referential ambiguity; the ambiguity jurors need to determine the existence based on the interference view picture set, which contains other objects of the same class as the target object, to detect whether the description is also true for non-target objects, thereby identifying potential referential ambiguity.

[0048] Specifically, the ambiguity jurors Based on the set of interference perspectives Detect whether there are other objects that meet the description condition. Its decision formula is:

[0049]

[0050] wherein, represents the object existence evaluation of each picture in the interference perspective picture set, represents the corresponding reason, represents the existence prompt word, represents the image of the i-th interference perspective.

[0051] Step 7: The outputs of distinguishability and ambiguity jurors are refined by regional magistrates. The distinguishability regional magistrate performs secondary processing on suspicious evaluation results through a corroborative refinement mechanism, and the ambiguity regional magistrate directly re-evaluates suspicious samples.

[0052] The distinguishability regional magistrate performs secondary processing on suspicious evaluation results through a corroborative refinement mechanism, and the ambiguity regional magistrate directly re-evaluates suspicious samples.

[0053] Specifically, the regional magistrates refine the outputs of the jurors; for the outputs of the distinguishability jurors, the distinguishability regional magistrates Introduce a corroborative refinement mechanism to construct a set of suspicion-assistance sample pairs , and replace the original results by splicing operations and re-evaluation:

[0054]

[0055] wherein, represents the updated existence and uniqueness scores, represents the updated score reasons, represents the replacement of the original evaluation by the refined data, represents the original existence and uniqueness scores, represents the corresponding reasons, represents the existence and uniqueness scores obtained by the corroborative refinement, represents the corresponding reasons.

[0056] Then, the refined existence and uniqueness results are combined to obtain the final distinguishability score and explanation. By this method, the partial result judgment errors or hallucinations caused by the occlusion of some perspectives and incomplete shooting of the multi-modal large model in the judgment process can be effectively reduced, and the influence of the singular results on the final score is reduced. The formula is as follows:

[0057]

[0058] Wherein, represents the distinguishability score of each picture in the positive sample view set, represents the corresponding reason, represents the updated existence score, represents the updated uniqueness score, and respectively represent the corresponding reason.

[0059] For the ambiguous jury output, the ambiguous region judge directly screens suspicious samples and re-evaluates the original results:

[0060]

[0061] Wherein, represents the object existence evaluation of each picture in the updated interference perspective picture set, represents the updated existence evaluation reason, represents the corresponding existence evaluation obtained by the evidence-based refinement, represents the corresponding existence evaluation reason after refinement.

[0062] Step 8: The district judge synthesizes the judgment results of each jury and region judge to output the final quality score and explanation of the data sample. The district judge only synthesizes the evaluation results of each jury and region judge, and does not directly access the original multi-perspective image or text description, in order to maintain the explainability of the reasoning chain and avoid information interference.

[0063] Specifically, the district judge synthesizes the results of each jury and region judge to form the final quality score and explanation:

[0064]

[0065] Wherein, represents the district prompt, represents the final quality score, represents the final reasoning explanation.

[0066] Embodiments:

[0067] To verify the effectiveness of the present method, we use the public large-scale 3DVG dataset ScanRefer to illustrate the method. The overall overview of the whole method is shown in Figure 1 , which includes the following steps:

[0068] Step 1: Data sample acquisition

[0069] Read the sample data of the three-dimensional visual reference (3DVG) task from the ScanRefer dataset, each sample including a multi-view scene image set and a corresponding natural language description. The multi-view scene image set provides observations of the target environment from different perspectives, and the description corresponds to the target object to be located. This step provides complete multi-modal input for subsequent multi-dimensional reasoning evaluation (see Figure 2 ).

[0070] Step 2: Build a jury-judge multi-role reasoning system

[0071] The multi-role reasoning system built by the present invention is driven by a multi-modal large model (MLLM) and includes multiple jury roles and multiple judge roles (see Figure 1 ). The jury analyzes different dimensional features of the sample in parallel, and the judge refines and synthesizes the jury results, finally outputting the sample quality judgment. The structure of the jury chain of thought reasoning is shown in Figure 2 , and the structure of the judge chain of thought is shown in Figure 3 .

[0072] Step 3: Logical jury evaluation

[0073] The logical jury receives the natural language description and performs logical completeness analysis to determine whether there are self-contradictions, logical paradoxes, or other logical defects that cannot uniquely refer. The output includes a logical judgment score and detailed reasoning explanation, providing basic information at the text semantic level for subsequent judgment (see Figure 2 orange part).

[0074] Step 4: Consistency jury evaluation

[0075] The consistency jury performs consistency analysis on multiple different descriptions (description group) of the target object existing in the dataset. First, the screening submodule eliminates logically inconsistent or low-confidence descriptions; then the verification submodule cross- verifies the current description with the filtered description group, and outputs the consistency judgment score and its reasoning explanation (see Figure 2 blue part).

[0076] Step 5: Distinctiveness jury evaluation

[0077] The distinguishability jurors perform two-stage evaluation based on the positive sample view set containing the target object: the first stage is to determine whether the target object exists in each view (existence determination); the second stage is to further determine whether the object is unique in the view where the existence determination passes (uniqueness determination). This step outputs the determination scores and explanations in the existence and uniqueness dimensions (see Figure 2 Green part).

[0078] Step 6: Ambiguity Juror Evaluation

[0079] The ambiguity jurors perform existence determination based on the interference view set (containing other objects of the same category as the target object), and if these interference objects are highly confident to be considered as meeting the description condition, it means that there is potential referential ambiguity in the description. This step outputs the ambiguity determination score and explanation (see Figure 2 Red part).

[0080] Step 7: Area Judge Refinement Evaluation

[0081] After the jurors complete the preliminary determination, the area judges refine the results.

[0082] The distinguishability area judges introduce a corroborative refinement mechanism, and for suspicious determination results, select appropriate "suspicious view - auxiliary view" pairs to improve the reliability of the determination by visual information splicing and re-evaluation (see Figure 3 Green part).

[0083] The ambiguity area judges directly re-evaluate suspicious ambiguity determination samples, thereby reducing false positives. This step can effectively alleviate the instability caused by missing or occluded multi-view observations (see Figure 3 Red part).

[0084] Step 8: Final Judge Decision

[0085] The final judge synthesizes the determination results of all jurors and area judges, performs the final sample quality decision, and outputs the comprehensive score and detailed explanation. This step does not directly access the original input data, but performs high-level aggregation analysis based on the reasoning results of the previous steps, thereby maintaining the explainability of the decision-making process (see Figure 3 Gray part).

[0086] The typical examples of the ScanRefer dataset (as shown in Figure 4 ) show that the method of the present application can effectively identify various types of toxic samples, including logical paradoxes and referential ambiguities, significantly improving the annotation quality and training stability of the dataset.

[0087] The above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for three-dimensional visual reference data quality evaluation based on multi-role multi-view reasoning, characterized in that: Acquire three-dimensional visual reference data, including multi-view three-dimensional images and text descriptions of target objects in the images; Construct a reasoning model based on multiple roles, perform multi-view analysis on the text descriptions through the first role based on logical consistency, distinguishability and ambiguity dimensions of the images, the second role performs auxiliary information splicing and re-evaluation on suspicious results of distinguishability analysis, re-evaluates suspicious three-dimensional visual reference data of ambiguity analysis, and comprehensively determines based on the analysis results of logical consistency and distinguishability, and the re-evaluation results of distinguishability and ambiguity analysis, to obtain the quality of three-dimensional visual reference data and the interpretation of the reasoning.

2. The method of claim 1, wherein the method is based on multi-role and multi-view reasoning. The logical analysis of the first role is based on a context learning strategy, and the logical completeness of the text description is analyzed through text logic to detect logical problems and obtain logical analysis results. The description consistency analysis of the first role is to analyze the consistency of multiple text descriptions of the target object, to construct a description group for multiple text descriptions of the same target object, to perform logical screening on the description group to remove low-confidence text descriptions, and to cross-verify the current text description with the screened description group. The distinguishability analysis of the first role is based on the image set of the positive sample view of the target object to analyze the distinguishability, to judge the existence of the target object in each view image and filter low-confidence view images, and to judge whether the target object exists uniquely based on the remaining view images as positive samples. The ambiguity analysis of the first role is based on the interference view image set to determine whether there are other objects in the scene that are the same as the target object to detect potential object reference ambiguity.

3. The method of claim 2, wherein the method further comprises: The second role identifies suspicious samples and high-confidence auxiliary samples from the distinguishability analysis results to construct a suspicious auxiliary sample pair, and re-evaluates the existence and / or uniqueness of the visual images of the suspicious samples and auxiliary samples after splicing and / or combining them.

4. The method of claim 2, wherein the method further comprises: The second role directly screens suspicious samples from the ambiguity analysis results and re-evaluates them through the introduction of a corroborative refinement mechanism to obtain existence evaluation results and corresponding reasons refined by the corroborative refinement mechanism to replace the original existence evaluation results and corresponding reasons.

5. The method of claim 2, wherein the method further comprises: The comprehensive determination of the second role is based on the description consistency analysis results and distinguishability analysis results of the first role, and the re-evaluation results of the distinguishability analysis results and ambiguity analysis results of the first role by the second role to obtain the final data quality results and the interpretation of the reasoning.

6. The method of claim 2, wherein the method further comprises: The consistency analysis of the first role evaluates the consistency of the target object based on the prompt words of the text consistency evaluation and the multiple text descriptions of the target object, obtains consistency evaluation results and corresponding reasons, and selects high-reliability text descriptions from the consistency evaluation results through a confidence threshold to construct a set of filtered text descriptions of the target object; based on the prompt words in the cross-validation process and the set of text descriptions, the current text description of the target object is evaluated to obtain the consistency evaluation results and corresponding reasons of the current text description of the target object.

7. The method of claim 2, wherein the method further comprises: The distinguishability analysis of the first role is based on the positive sample visual image and the existence prompt word, and existence evaluation is performed on the target object of the positive sample visual image to obtain an existence evaluation result. The positive sample visual image with high existence is filtered through an existence judgment threshold; Based on the filtered positive sample visual image and the uniqueness prompt word, the uniqueness of the target object visual image is evaluated to obtain a uniqueness evaluation result and a corresponding reason.

8. The method of claim 7, wherein the method further comprises: The ambiguity analysis of the first role is based on the interference visual image and the existence prompt word, and existence evaluation is performed on the target object of the interference visual image to obtain an existence evaluation result and a corresponding reason. 9.The method of claim 7, wherein the method further comprises: The distinguishability analysis of the second role is based on the introduction of a supporting type refining mechanism, construction of a suspicious auxiliary sample pair, and reevaluation through splicing operation to replace the original existence and uniqueness evaluation results. Then, the updated existence evaluation result is combined with the updated uniqueness evaluation result, the updated existence evaluation reason is combined with the updated uniqueness evaluation reason, and the final distinguishability evaluation result and the corresponding reason of the positive sample visual image are obtained.

Citation Information

Patent Citations

  • Training data set labeling method and device for AI visual identification

    CN119399579A

  • Quality evaluation method and system for text-generated three-dimensional content, medium and terminal

    CN119516566A

  • Intelligent content evaluation and optimization method and system based on multi-standard preference learning

    CN120494074A

  • Agent-based multi-dimensional data quality intelligent evaluation method

    CN120496720A

  • Automatic scene calibration method based on machine vision

    CN120635605A