Multimodal large model hallucination suppression contrast description data generation method and device
By constructing a verifiable set of visual facts for multimodal large models and a comparative descriptive data generation method, the problem of hallucination suppression in open-world scenarios for multimodal large models is solved, achieving highly reliable and interpretable data generation and improving the hallucination suppression capability of multimodal large models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-04-17
AI Technical Summary
Multimodal large models are prone to producing illusions in open-world scenarios. Existing technologies struggle to automatically and scalably generate verifiable comparative training data, resulting in high noise levels and limited coverage, especially for typical illusion types such as long-tailed objects, fine-grained attributes/relationships, and images and text.
By acquiring raw multimodal samples, a detection query set is generated, and an open vocabulary object detection model is used to obtain the object set and its spatial positioning information. A verifiable set of visual facts is constructed, comparative descriptive data is generated, and a highly credible comparative descriptive training dataset is output through multi-source evidence consistency verification and quality screening.
It enables large-scale, evidence-traceable data generation in open-world scenarios, improves the interpretability and verifiability of the generated data, overcomes the problem of insufficient hallucination suppression in existing technologies, and outputs highly credible and controllable comparative descriptive training datasets.
Smart Images

Figure CN121686147B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimodal large model training data generation technology, and in particular to a method and apparatus for generating contrastive descriptive data for hallucination suppression in multimodal large models. Background Technology
[0002] Large Vision-Language Models (LVLMs / MLLMs) excel in open-ended generative tasks such as image captioning, visual question answering, and text-to-image dialogue. However, they are prone to "illusions" in open-world scenarios, generating objects, attributes, relationships, or textual information that are inconsistent with the facts of the image. This problem not only reduces the reliability of interactions but also introduces systematic noise when using generated data to feed back into training.
[0003] Existing engineering practices for reducing illusions typically rely on manual annotation or simple negative sample construction: either these methods are costly and difficult to scale, or they lack verifiable evidence, resulting in noisy samples with limited coverage, particularly failing to cover typical illusion types such as long-tail objects, fine-grained attributes / relationships, and images with text. Therefore, there is an urgent need for a comparative descriptive data generation method that can automatically generate, verify, audit, and possess hard-example mining capabilities in open-world conditions. Summary of the Invention
[0004] The purpose of this invention is to provide a method and apparatus for generating contrastive descriptive data for suppressing illusions in a multimodal large model, in order to solve the problems in the prior art such as the difficulty in automatically scaling up the generation of contrastive training data, the lack of verifiable evidence for negative samples leading to high noise, the ignoring of multiple types of illusions such as object attributes, object relationships, and image and text elements while only covering the existence of objects, and the lack of a difficult example screening mechanism leading to unstable training gains.
[0005] According to a first aspect of the embodiments of this application, a method for generating multimodal large-model hallucination suppression contrastive description data is provided, comprising:
[0006] Obtain raw multimodal samples, wherein the raw multimodal samples include at least images and image-related text information;
[0007] Based on the image and the text information associated with the image, a detection query set is generated;
[0008] The detection query set and the image are input into the open vocabulary target detection model to obtain the object set and the spatial positioning information of the object;
[0009] Based on the object set and the spatial positioning information corresponding to the objects, a verifiable set of visual facts is constructed. The set of visual facts consists of multiple atomic facts. Each atomic fact is the smallest fact unit that describes the existence of an object, the attribute of an object, the relationship between objects, or the image and text elements. The set of visual facts includes at least two of the above types of atomic facts.
[0010] Comparative descriptive data is generated based on the visual fact set. The comparative descriptive data includes positive sample descriptions that are consistent with the visual fact set and negative sample descriptions that conflict with the visual fact set. The negative sample descriptions are generated by a comparative perturbation operator that performs minimal editing on at least one atomic fact and satisfies surface similarity constraints.
[0011] Perform multi-source evidence consistency verification and quality screening on the comparative description data to obtain a high-confidence comparative description training dataset;
[0012] The high-confidence comparative description training dataset is subjected to hard example filtering to obtain the final training dataset.
[0013] According to a second aspect of the embodiments of this application, a multimodal large-model hallucination suppression contrast description data generation apparatus is provided, comprising:
[0014] The data acquisition module is used to acquire raw multimodal samples, which include at least images and text information related to the images;
[0015] The query generation module is used to generate a set of detection queries based on the image and image-related text information;
[0016] The target extraction module is used to input the detection query set and the image into an open vocabulary target detection model to obtain a set of objects and the spatial positioning information of the objects.
[0017] The fact construction module is used to construct a verifiable set of visual facts based on the set of objects and the spatial positioning information corresponding to the objects. The set of visual facts consists of multiple atomic facts. Each atomic fact is the smallest fact unit that describes the existence of an object, the attribute of an object, the relationship between objects, or the image and text elements. The set of visual facts includes at least two of the above types of atomic facts.
[0018] A comparison generation module is used to generate comparative descriptive data based on the visual fact set. The comparative descriptive data includes positive sample descriptions that are consistent with the visual fact set and negative sample descriptions that conflict with the visual fact set. The negative sample descriptions are generated by a comparison perturbation operator that performs minimal editing on at least one atomic fact and satisfies surface similarity constraints.
[0019] The verification and filtering module is used to perform multi-source evidence consistency verification and quality filtering on the comparative description data to obtain a highly reliable comparative description training dataset.
[0020] The difficult example filtering and output module is used to perform difficult example filtering on the high-confidence comparative description training dataset to obtain the final training dataset.
[0021] According to a third aspect of the embodiments of this application, an electronic device is provided, comprising:
[0022] One or more processors;
[0023] Memory, used to store one or more programs;
[0024] When the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in the first aspect.
[0025] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method as described in the first aspect.
[0026] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects:
[0027] This invention first generates a detection query set based on images and their related text, and then inputs the queries and images into an open vocabulary target detection model to obtain a set of objects and their spatial location information. This allows all subsequent descriptions to be based on visual evidence that the objects can be located, overcoming the problems of unstable object sources, insufficient coverage of long-tail concepts, and difficulty in automatically verifying the existence of objects in existing automatic data generation, which leads to noise accumulation. This enables large-scale, traceable data generation capabilities for open-world scenarios.
[0028] Furthermore, since this invention constructs a verifiable visual fact set composed of atomic facts based on object sets and spatial positioning information, and uniformly organizes and binds at least two types of facts from object existence, object attributes, object relationships, and image and text elements to corresponding evidence, thus forming a closed loop of "description-fact-evidence", it overcomes the problems of existing technologies that only cover a single type of fact, lack systematic modeling of common illusions such as attributes / relationships / images and text, and have incomplete evidence chains that make auditing and verification difficult. In this way, it can consistently express facts and constrain evidence for multiple types of illusions, and improve the interpretability and verifiability of the generated data.
[0029] Based on this, the present invention generates positive sample descriptions by using the set of visual facts as constraints, and generates negative sample descriptions that conflict with the facts by performing minimal editing on at least one atomic fact through contrast perturbation. At the same time, it applies surface similarity constraints, and combines multi-source evidence consistency verification and quality screening, as well as hard example screening based on confusion score, to overcome the problems of arbitrary construction of negative samples, excessive differences from positive samples or multiple simultaneous errors leading to unfocused training signals, and semantic drift and instability caused by the mixing of unverifiable samples. In this way, it outputs a highly reliable, controllable difficulty and more realistic hallucination form of contrastive description training dataset, thereby effectively supporting the improvement of hallucination suppression capabilities of multimodal large models. Attached Figure Description
[0030] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0031] Figure 1 This is a flowchart illustrating a method for generating contrastive descriptive data to reduce multimodal large model illusion, according to an exemplary embodiment.
[0032] Figure 2 This is a block diagram illustrating a contrastive descriptive data generation apparatus for reducing multimodal large model illusion, according to an exemplary embodiment.
[0033] Figure 3 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0034] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0035] Figure 1 This is a flowchart illustrating a method for generating contrastive descriptive data for multimodal large-model hallucination suppression according to an exemplary embodiment. (Reference) Figure 1 The method of this invention embodiment may include the following steps:
[0036] S1: Obtain the original multimodal samples, which include at least images and text information related to the images;
[0037] Specifically, taking a campus laboratory security inspection scenario as an example, a raw multimodal sample is obtained, including images. The images are photographs taken during on-site inspections of laboratories or corridors. Text information is used to describe, inquire about, or verify the image content. Samples (I,T) or (I,Q,A) are obtained, where I is the image, T is the image description or retrieval text, Q is the image-related instruction or question, and A is a reference answer or candidate answer. The text information consists of at least one of T, Q, and A. To reduce noise, the text can be deduplicated, cleaned, segmented, and normalized to obtain normalized text T′.
[0038] S2: Based on the image and image-related text information, generate a detection query set; this step may include the following sub-steps:
[0039] S21: Based on the text information, extract noun phrases, proper noun phrases, or quantity phrases as object candidates, and perform word form normalization and synonym mapping on the object candidates to obtain a set of object phrases;
[0040] Specifically, taking the scenario of campus laboratory safety inspection as an example, the aforementioned standard text This can include statements such as "There is a fire extinguisher at the entrance of the corridor, an EXIT sign above the door, and goggles on the lab table." Based on the text information, the sentences are segmented and tagged with parts of speech. Dependency parsing or phrase chunking methods are used to extract noun phrases, proper noun phrases, or quantifier phrases as candidate objects. Case unification, word form restoration, plural normalization, and stop word filtering are performed on the candidate objects. Synonym mapping and normalization are completed by combining a preset thesaurus or vector similarity to obtain a set of object phrases. In one embodiment, the compound phrase "object + attribute" is decomposed into object phrases and attribute phrases as query items to reduce missed detections caused by compound phrases and improve open vocabulary detection recall.
[0041] S22: Perform optical character recognition on the image to obtain text strings and their location regions, and perform noise reduction, merging and normalization on the recognition results to obtain a set of text element phrases;
[0042] Specifically, taking a campus laboratory security inspection scenario as an example, the image The image can be a photo taken during a patrol at a corridor entrance, showing an "EXIT" sign or chemical label above the door. Optical character recognition (OCR) is performed on the image to obtain a text string, character confidence level, and corresponding location region. Noise reduction processing is then applied to the recognition results, including deleting low-confidence strings, filtering meaningless symbols and abnormally long segments. Adjacent or overlapping text regions are merged and aggregated at the line level. The strings are then subjected to full-width / half-width character unification, whitespace normalization, case normalization, and common character correction to obtain a set of text element phrases for subsequent image text element fact construction and text element perturbation.
[0043] S23: Sample long-tail concepts from an open vocabulary object library and expand and filter the object phrase set using a language model to obtain an expanded object phrase set;
[0044] Specifically, taking a campus laboratory safety inspection scenario as an example, the object phrase set may include phrases such as "fire extinguisher," "EXIT sign," and "goggles." Long-tail concepts are selected from an open vocabulary object library according to a preset sampling strategy as candidate object phrases, and the object phrase set and the long-tail concepts are input into a language model to generate an extended object phrase set. The extension includes synonym rewriting, hypernym / hyponym expansion, and morphological variant generation. The filtering includes deleting concepts that are irrelevant to the image domain, deleting concepts that are highly repetitive with existing object phrases, and deleting concepts that contain sensitive or undetectable words, thereby obtaining the extended object phrase set.
[0045] S24: Perform deduplication and merging on the object phrase set, the text element phrase set, and the extended object phrase set to form a detection query set;
[0046] Specifically, taking a campus laboratory safety inspection scenario as an example, the object phrase set may include "fire extinguisher" and "goggles," the text element phrase set may include "EXIT," and the extended object phrase set may include "eyewash station" and "emergency sprinkler," etc.; the object phrase set, the text element phrase set, and the extended object phrase set are subjected to deduplication and merging processing. Deduplication includes deduplication based on complete string matching and near-synonym merging based on synonym mapping or vector similarity threshold; the merged phrases are subjected to length pruning and format unification, and a source identifier (text extraction, OCR, object library sampling, or language model extension) can be written for each phrase, finally forming a detection query set for subsequent target localization and object set construction in open vocabulary target detection.
[0047] S3: Input the detection query set and the image into the open vocabulary target detection model to obtain the object set and the spatial positioning information of the object;
[0048] Specifically, taking a campus laboratory security inspection scenario as an example, the image The image can be a photo taken during a patrol at the entrance of a corridor. The detection query set can include query phrases such as "fire extinguisher," "EXIT sign," and "goggles." The detection query set and the image are then input into an open vocabulary object detection model (such as Grounding DINO) to obtain a candidate object set. ,in Spatial location information for the object (including bounding box coordinates or the region defined by the bounding box). For the object phrase that matches the detection query set, The confidence level is used to perform threshold filtering on the candidate object set to remove those with a confidence level lower than a preset threshold. The candidate results are selected, and non-maximum suppression is used to remove redundancy from overlapping candidate regions; further, the results are analyzed based on a thesaurus or phrase similarity. Synonym merging and normalization are performed to obtain a set of objects and their corresponding spatial location information, which are used for the subsequent construction of a set of visual facts.
[0049] S4: Construct a verifiable set of visual facts based on the object set and the spatial positioning information corresponding to the objects. The set of visual facts consists of multiple atomic facts, where each atomic fact is the smallest fact unit describing the existence of an object, the attribute of an object, the relationship between objects, or the element of an image or text. The set of visual facts includes at least two of the above types of atomic facts. This step may include the following sub-steps:
[0050] S41: For each object in the object set, generate an object existence fact and use the object phrase, its corresponding spatial location information, and confidence level as evidence of existence;
[0051] Specifically, taking a campus laboratory safety inspection scenario as an example, the image may contain objects such as "fire extinguisher," "EXIT sign," and "goggles"; for the object set and its spatial positioning information obtained in step S3, each detected object is traversed, and the first object is recorded as... One object for:
[0052]
[0053] in Spatial positioning information (e.g., bounding box coordinates). For object phrases, The confidence level is used to generate an object existence fact based on the object phrase, spatial location information, and confidence level, and the spatial location information is written into the existence fact entry as a verifiable evidence field.
[0054]
[0055] in At least include the bounding box It may further include at least one of the following: an object region image obtained by bounding box cropping, and an object mask obtained by a segmentation model, for subsequent consistency verification and audit traceability of object existence. Optionally, objects with a confidence level below a preset threshold are first removed or marked as low-confidence objects to reduce noise in subsequent fact-building.
[0056] S42: Perform attribute extraction or attribute verification on the object region in the object set to generate object attribute facts, and bind a corresponding evidence region to each attribute fact;
[0057] Specifically, taking a campus laboratory safety inspection scenario as an example, attributes such as the color of a fire extinguisher and the color of an "EXIT" sign can be extracted or verified; based on the spatial positioning information of the object... A target region is identified, and attribute extraction or attribute validation is performed on the target region to obtain attribute names and attribute confidence scores; the attributes include at least one of color, material, action, orientation, and quantity. In one embodiment, attribute extraction can be obtained by predicting the target region using an attribute classification model; in another embodiment, attribute validation can be performed by a visual question-and-answer validation model to determine whether the object has attributes. The determination is made by binding the attribute name, attribute confidence level, and evidence region to construct the object attribute facts. :
[0058]
[0059] in For attribute confidence, The evidence region includes at least one of the following: an object clipping region, an object mask, or a region defined by a bounding box. To improve verifiability, the evidence region can simultaneously record its source (bounding box clipping / segmentation mask) and region scale information, thereby supporting subsequent consistency verification modules in reviewing the attribute facts.
[0060] S43: Calculate geometric or spatial relationships based on spatial location information between objects, generate object relationship facts, and provide geometric consistency verification results as evidence for each relationship fact;
[0061] Specifically, taking a campus laboratory safety inspection scenario as an example, spatial relationships such as "EXIT sign above the door" and "fire extinguisher beside the door" can be constructed; for object pairs in an object set... Based on their spatial positioning information Computational geometric or spatial relations The relationship includes at least one of left_of, right_of, above, below, inside, and overlap. In one embodiment, the relationship can be determined by statistical measures such as the relative position of the center points of the object bounding boxes, the horizontal / vertical projection overlap rate, the intersection-over-union ratio (IoU), the inclusion rate, or a distance threshold; and the geometric consistency verification result is output as evidence of the relationship to construct the object relationship fact. :
[0062]
[0063] in, and For object indexes, used to indicate object pairs ; and Each is an object With object The object phrase; Spatial relationship category; This is the relational evidence field, used to represent the geometric consistency verification result of the spatial relationship:
[0064]
[0065] For the geometric consistency verification results, It can be a Boolean decision marker or a continuous statistic (such as IoU, coverage ratio, centroid offset). By binding the relationship decision to the evidence field, it can be directly reused in subsequent validation and screening stages. Consistency verification of relational facts reduces the probability of "fabricated relations" illusion samples entering the training set.
[0066] S44: Perform optical character recognition on the image to generate image text element facts, and bind the location region and recognition confidence level to each text element fact as evidence;
[0067] Specifically, taking a campus laboratory security inspection scenario as an example, the image may contain identifying text such as "EXIT". Optical character recognition (OCR) is performed on the image to obtain the text string, location region, and recognition confidence level. The recognition results are then subjected to denoising, merging, and normalization processing to obtain a stable set of text elements. The denoising includes deleting low-confidence text, filtering abnormal symbols and meaningless fragments; the merging includes row-level aggregation of adjacent or overlapping text regions; and the normalization includes unifying full-width and half-width characters, standardizing whitespace characters, and normalizing capitalization. Image text element facts are constructed for each text element. :
[0068]
[0069] in, An index for text elements in an image, used to indicate the first... OCR text elements, It is a text string. For location area, To identify confidence levels; At least include location area and its confidence level At least one of the following, and optionally including a text region cropped image, for subsequent consistency verification and comparison perturbation generation to determine whether the text in the image has been fabricated / tampered with.
[0070] S45: Write at least two facts from S41 to S44 above as atomic facts into the same set of visual facts, and attach verifiable visual evidence to each atomic fact, thereby obtaining the set of verifiable visual facts;
[0071] Specifically, at least two categories of facts obtained in steps S41–S44—object existence facts, object attribute facts, object relationship facts, and image / text element facts—are written into the same visual fact set. Each fact is treated as an atomic fact and uniformly numbered or indexed, forming a traceable "fact-evidence" chain. The visual fact set can be represented as:
[0072]
[0073] For each atomic fact in the set Bind the corresponding evidence field (For example ),in, In order to be consistent with atomic facts The corresponding verifiable visual evidence field; A field that represents evidence that an object exists; Evidence field representing the facts of an object's attributes; Evidence fields representing the facts of object relationships; The evidence field represents the facts of image and text elements, and can record atomic fact type tags and associated object indexes, so that subsequent steps can quickly retrieve and perform comparison perturbations and consistency checks according to "object-attribute-relationship-text element". Through the above writing, indexing and evidence binding operations, the verifiable set of visual facts is finally obtained.
[0074] S5: Generate comparative descriptive data based on the visual fact set. The comparative descriptive data includes positive sample descriptions that are consistent with the visual fact set and negative sample descriptions that conflict with the visual fact set. The negative sample descriptions are generated by a comparative perturbation operator that performs minimal editing on at least one atomic fact and satisfies surface similarity constraints.
[0075] Specifically, taking the scenario of campus laboratory security inspection as an example, the set of visual facts... It can support facts such as "there is an EXIT sign above the door", "there is a fire extinguisher at the door", and "there are safety goggles on the lab bench", based on the set of visual facts. Constructing comparative descriptive data pairs First, generate positive sample descriptions. In one embodiment, the standard text is... Perform fact constraint cleaning, deleting key object phrases or text elements that cannot be grounded in the visual fact set; in another embodiment, use templates or language models to generate descriptive text, and constrain the key objects, attributes, relationships, or text elements in the generated text to be grounded in the visual fact set. This provides support, thereby obtaining positive sample descriptions consistent with the set of visual facts. Then negative sample descriptions are generated. From the set of visual facts Select a subset of atomic facts to be perturbed (the aforementioned) Typically contains only one key atomic fact), and regarding the stated Applying the contrast perturbation operator yields the perturbed set of facts:
[0076]
[0077] Based on Will The text fragment corresponding to the key atomic fact is replaced with the one that is... Consistent expression, thus obtaining negative sample descriptions that conflict with the set of visual facts. (For example, "There is an EXIT sign above the door" can be minimized to "There is an ENTRANCE sign above the door" or "There is a fire extinguisher at the door" can be minimized to "There is no fire extinguisher at the door"); In order to improve the deceptiveness of negative samples, the perturbation object can be selected using a confusion perception strategy, so that the replacement item is semantically similar to the real object and has a high co-occurrence probability, so as to more closely resemble the illusion generated by the multimodal large model.
[0078] The minimum edit includes: modifying only one key atomic fact in the positive sample description while keeping the rest of the facts unchanged; and the surface similarity constraint includes at least one of an edit distance threshold or a semantic similarity threshold.
[0079] The contrast perturbation operators include at least: non-existent object perturbation, attribute flip perturbation, relation flip perturbation, count perturbation, and text element perturbation; the count perturbation is used to rewrite the number of objects or quantifiers that are inconsistent with visual facts, and the text element perturbation is used to introduce or rewrite text, signs, or numbers in an image incorrectly.
[0080] S6: Perform multi-source evidence consistency verification and quality screening on the comparative description data to obtain a high-confidence comparative description training dataset;
[0081] Specifically, for each set of comparative descriptive data Taking a campus laboratory security inspection scenario as an example, the image... Photos of inspections at the entrance of the corridor; positive sample description. A possible negative sample description is "There is an EXIT sign above the door, and a fire extinguisher is located at the doorway". This can be a minimally edited description of one of the key facts (e.g., "There is an ENTRANCE sign above the door" or "There is no fire extinguisher at the door"). First, based on the visual fact set and corresponding evidence fields, key object phrases, key attributes / relationships, and key text elements are determined, and the corresponding text fragments are located in the positive and negative sample descriptions respectively (for example, "EXIT sign / ENTRANCE sign" is located as a key text element, "fire extinguisher" is located as a key object phrase, and "at the door / above the door" is located as a key relationship fragment). Then, a multi-source evidence consistency check is performed on the positive and negative sample descriptions, requiring that the key facts in the positive sample descriptions are all supported by image evidence (for example, OCR evidence supports "EXIT", and target location evidence supports "fire extinguisher"), and that at least one perturbed key fact in the negative sample description cannot be supported by image evidence or conflicts with image evidence. After passing the consistency check, a quality screening is further performed, which includes at least: checking whether the negative sample meets the minimum edit constraint (only one key atomic fact is perturbed and the rest of the facts remain unchanged), and checking whether the positive and negative samples meet the surface similarity constraint (edit distance threshold or semantic similarity threshold). Samples that fail the consistency check or quality screening are removed or regenerated to obtain a high-confidence comparative description training dataset.
[0082] The multi-source evidence consistency verification includes:
[0083] Target feasibility verification: Perform secondary phrase localization on the object phrases in the description and calculate the overlap or confidence level;
[0084] Attribute or relation consistency verification: Perform attribute classification verification or geometric relation verification on the attributes or relations in the description;
[0085] Visual implication verification: The visual implication model is used to determine whether the description of a positive sample is implied by the image, and whether the description of a negative sample contradicts the image or is not implied.
[0086] Logical closed-loop consistency verification: Generate logically related queries for key objects and key attributes in the description and check the consistency of the closed-loop responses to identify potential illusory descriptions.
[0087] S7: Perform hard example filtering on the high-confidence comparative description training dataset to obtain the final training dataset;
[0088] Specifically, taking a campus laboratory safety inspection scenario as an example, a confusion score is calculated for samples that have passed the multi-source evidence consistency verification and quality screening. And based on this, more deceptive negative samples are selected; wherein, the confusion score It is obtained by weighting at least one of the following: the semantic similarity between positive and negative samples, the visual compatibility between negative samples and the image, and the trigger signal of the target multimodal model for negative samples (e.g., positive samples). For "the door has an EXIT sign above it", negative samples The phrase "there is an ENTRANCE sign above the door" is semantically similar to "the sign above the door," but the negative sample conflicts with the textual evidence in the image. (Higher); Priority retention Samples exceeding a preset threshold are selected, and the retained samples are stratified and sampled according to the type and difficulty range of negative sample perturbation to control the data distribution. Subsequently, the selected samples are output in a structured format as the final training dataset, and at least one of the following is written for each sample: object set, visual fact set or its index, corresponding visual evidence field, perturbed atomic fact identifier, and verification result, for subsequent training, evaluation and auditing.
[0089] The difficult example screening includes: retaining more confusing negative samples based on confusion scores, and performing stratified sampling based on the type and difficulty of negative samples;
[0090] The training dataset is output in at least one of the following formats: comparative selection sample, true / false discrimination sample, or preference pair sample, and for each sample, at least one of the following is written: object set, visual evidence field, perturbed atomic fact identifier, and verification result.
[0091] As can be seen from the above technical solutions, the present invention first generates a detection query set based on the image and its related text, and inputs the query and the image together into an open vocabulary target detection model to obtain the object set and its spatial positioning information. This enables all subsequent descriptions to be based on visual evidence that the object can be located, overcoming the problems of unstable object sources, insufficient coverage of long-tail concepts, and difficulty in automatically verifying whether the object actually exists in existing automatic data generation, which leads to noise accumulation. Thus, it realizes the ability to generate large-scale, traceable data for open-world scenarios.
[0092] Furthermore, since this invention constructs a verifiable visual fact set composed of atomic facts based on object sets and spatial positioning information, and uniformly organizes and binds at least two types of facts from object existence, object attributes, object relationships, and image and text elements to corresponding evidence, thus forming a closed loop of "description-fact-evidence", it overcomes the problems of existing technologies that only cover a single type of fact, lack systematic modeling of common illusions such as attributes / relationships / images and text, and have incomplete evidence chains that make auditing and verification difficult. In this way, it can consistently express facts and constrain evidence for multiple types of illusions, and improve the interpretability and verifiability of the generated data.
[0093] Based on this, the present invention generates positive sample descriptions by using the set of visual facts as constraints, and generates negative sample descriptions that conflict with the facts by performing minimal editing on at least one atomic fact through contrast perturbation. At the same time, it applies surface similarity constraints, and combines multi-source evidence consistency verification and quality screening, as well as hard example screening based on confusion score, to overcome the problems of arbitrary construction of negative samples, excessive differences from positive samples or multiple simultaneous errors leading to unfocused training signals, and semantic drift and instability caused by the mixing of unverifiable samples. In this way, it outputs a highly reliable, controllable difficulty and more realistic hallucination form of contrastive description training dataset, thereby effectively supporting the improvement of hallucination suppression capabilities of multimodal large models.
[0094] Corresponding to the above-described embodiments of the method for generating contrastive descriptive data for multimodal large model hallucination suppression, this application also provides embodiments of an apparatus for generating contrastive descriptive data for multimodal large model hallucination suppression.
[0095] Figure 2 This is a block diagram illustrating a multimodal large-model hallucination suppression contrast description data generation device according to an exemplary embodiment. (Refer to...) Figure 2 The device includes:
[0096] Data acquisition module 1 is used to acquire raw multimodal samples, wherein the raw multimodal samples include at least images and text information related to the images;
[0097] Query generation module 2 is used to generate a detection query set based on the image and image-related text information;
[0098] Target extraction module 3 is used to input the detection query set and the image into an open vocabulary target detection model to obtain a set of objects and the spatial positioning information corresponding to the objects;
[0099] The fact construction module 4 is used to construct a verifiable set of visual facts based on the set of objects and the spatial positioning information corresponding to the objects. The set of visual facts consists of multiple atomic facts. Each atomic fact is the smallest fact unit that describes the existence of an object, the attribute of an object, the relationship between objects, or the image and text elements. The set of visual facts includes at least two of the above types of atomic facts.
[0100] The comparison generation module 5 is used to generate comparative description data based on the visual fact set. The comparative description data includes positive sample descriptions that are consistent with the visual fact set and negative sample descriptions that conflict with the visual fact set. The negative sample descriptions are generated by a comparison perturbation operator that performs minimal editing on at least one atomic fact and satisfies surface similarity constraints.
[0101] The verification and filtering module 6 is used to perform multi-source evidence consistency verification and quality filtering on the comparative description data to obtain a high-confidence comparative description training dataset.
[0102] The difficult example filtering and output module 7 is used to perform difficult example filtering on the high-confidence comparative description training dataset to obtain the final training dataset.
[0103] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0104] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0105] Accordingly, this application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the multimodal large-model hallucination suppression contrast description data generation method as described above. Figure 3 The diagram shown is a hardware structure diagram of any device with data processing capabilities, which is a multimodal large-model hallucination suppression contrast description data generation device provided in an embodiment of the present invention. (Except for...) Figure 3In addition to the processor, memory, DMA controller, disk, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0106] Accordingly, this application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the multimodal large-model illusion suppression contrast description data generation method described above. The computer-readable storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data-processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.
[0107] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0108] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A multi-modal large model hallucination suppression contrastive description data generation method, characterized in that, include: Obtain raw multimodal samples, wherein the raw multimodal samples include at least images and image-related text information; Based on the image and the text information associated with the image, a detection query set is generated; The detection query set and the image are input into the open vocabulary target detection model to obtain the object set and the spatial positioning information of the object; Based on the object set and the spatial positioning information corresponding to the objects, a verifiable set of visual facts is constructed. The set of visual facts consists of multiple atomic facts. Each atomic fact is the smallest fact unit that describes the existence of an object, the attribute of an object, the relationship between objects, or the image and text elements. The set of visual facts includes at least two of the above types of atomic facts. Comparative descriptive data is generated based on the visual fact set. The comparative descriptive data includes positive sample descriptions that are consistent with the visual fact set and negative sample descriptions that conflict with the visual fact set. The negative sample descriptions are generated by a comparative perturbation operator that performs minimal editing on at least one atomic fact and satisfies surface similarity constraints. Perform multi-source evidence consistency verification and quality screening on the comparative description data to obtain a high-confidence comparative description training dataset; The high-confidence comparative description training dataset is subjected to hard example filtering to obtain the final training dataset.
2. The method of claim 1, wherein, Based on the image and image-related text information, a detection query set is generated, including: Based on the text information, noun phrases, proper noun phrases, or quantity phrases are extracted as object candidates, and word form normalization and synonym mapping are performed on the object candidates to obtain a set of object phrases; Optical character recognition is performed on the image to obtain text strings and their location regions, and the recognition results are denoised, merged and normalized to obtain a set of text element phrases; Long-tail concepts are sampled from an open vocabulary object library, and the object phrase set is expanded and filtered by a language model to obtain an expanded object phrase set; The object phrase set, the text element phrase set, and the extended object phrase set are deduplicated and merged to form a detection query set.
3. The method of claim 1, wherein, Based on the object set and the spatial positioning information corresponding to the objects, a verifiable set of visual facts is constructed, including: (1) For each object in the object set, generate the existence fact of the object, and use the object phrase and its corresponding spatial location information and confidence level as evidence of existence; (2) Perform attribute extraction or attribute verification on the object regions in the object set to generate object attribute facts, and bind the corresponding evidence region to each attribute fact; (3) Calculate geometric or spatial relationships based on spatial positioning information between objects, generate object relationship facts, and provide geometric consistency verification results for each relationship fact as evidence; (4) Perform optical character recognition on the image to generate image text element facts, and bind the location region and recognition confidence level to each text element fact as evidence; Write at least two of the facts in (1) to (4) above as atomic facts into the same set of visual facts, and attach verifiable visual evidence to each atomic fact to obtain the set of verifiable visual facts.
4. The method according to claim 1, characterized in that, The minimum edit includes: modifying only one key atomic fact in the positive sample description while keeping the rest of the facts unchanged; and the surface similarity constraint includes at least one of an edit distance threshold or a semantic similarity threshold.
5. The method according to claim 1, characterized in that, The contrast perturbation operators include at least: non-existent object perturbation, attribute flip perturbation, relation flip perturbation, count perturbation, and text element perturbation; the count perturbation is used to rewrite the number of objects or quantifiers that are inconsistent with visual facts, and the text element perturbation is used to introduce or rewrite text, signs, or numbers in an image incorrectly.
6. The method according to claim 1, characterized in that, The multi-source evidence consistency verification includes: Target feasibility verification: Perform secondary phrase localization on the object phrases in the description and calculate the overlap or confidence level; Attribute or relation consistency verification: Perform attribute classification verification or geometric relation verification on the attributes or relations in the description; Visual implication verification: The visual implication model is used to determine whether the description of a positive sample is implied by the image, and whether the description of a negative sample contradicts the image or is not implied. Logical closed-loop consistency verification: Generate logically related queries for key objects and key attributes in the description and check the consistency of the closed-loop responses to identify potential illusory descriptions.
7. The method according to claim 1, characterized in that, The difficult example screening includes: retaining more confusing negative samples based on confusion scores, and performing stratified sampling based on the type and difficulty of negative samples; The training dataset is output in at least one of the following formats: comparative selection sample, true / false discrimination sample, or preference pair sample, and for each sample, at least one of the following is written: object set, visual evidence field, perturbed atomic fact identifier, and verification result.
8. A device for generating contrastive descriptive data for multimodal large-scale hallucination suppression, characterized in that, include: The data acquisition module is used to acquire raw multimodal samples, which include at least images and text information related to the images; The query generation module is used to generate a set of detection queries based on the image and image-related text information; The target extraction module is used to input the detection query set and the image into an open vocabulary target detection model to obtain a set of objects and the spatial positioning information of the objects. The fact construction module is used to construct a verifiable set of visual facts based on the set of objects and the spatial positioning information corresponding to the objects. The set of visual facts consists of multiple atomic facts. Each atomic fact is the smallest fact unit that describes the existence of an object, the attribute of an object, the relationship between objects, or the image and text elements. The set of visual facts includes at least two of the above types of atomic facts. A comparison generation module is used to generate comparative descriptive data based on the visual fact set. The comparative descriptive data includes positive sample descriptions that are consistent with the visual fact set and negative sample descriptions that conflict with the visual fact set. The negative sample descriptions are generated by a comparison perturbation operator that performs minimal editing on at least one atomic fact and satisfies surface similarity constraints. The verification and filtering module is used to perform multi-source evidence consistency verification and quality filtering on the comparative description data to obtain a highly reliable comparative description training dataset. The difficult example filtering and output module is used to perform difficult example filtering on the high-confidence comparative description training dataset to obtain the final training dataset.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Question and answer information generation method and device based on large model, electronic equipment and medium
CN118568240A
Comparison decoding illusion mitigation method and device based on multiple modes, and terminal
CN118966387A