Multi-modal data construction method based on VLM-LLM cooperative verification and instruction driving

Through the collaborative verification design mechanism between VLM and LLM, the multimodal data set is generated using instruction drivers, which solves the problems of high cost and low efficiency in the existing technology, and achieves the goal of efficiently building high-quality multimodal data sets.

CN120494114AActive Publication Date: 2025-08-15HANGZHOU MOREDIAN TECH CO LTD

Patent Information

Application Number
CN202510976674.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-15
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

It is difficult for the existing technology to efficiently construct high-quality multimodal data sets, and there are problems such as high cost, large quality fluctuations, low iteration efficiency, insufficient semantic understanding, modal fragmentation and poor dynamic adaptability.

Method used

Through the collaborative verification design mechanism between VLM and LLM, the instruction-driven mechanism is used to obtain multiple sets of graphic descriptions, perform logical judgment and cross-check, generate detailed question-and-answer pairs of different styles, depths and structures, and build multimodal data sets.

Benefits of technology

Reduce manual labeling, improve data construction efficiency, reduce error labeling rate, achieve automatic adaptation to new fields, improve the flexibility and adaptability of data sets, and provide high-quality multimodal data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494114A_ABST
    Figure CN120494114A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal data construction method based on VLM-LLM cooperative verification and instruction driving, and the method achieves the full-process automation through a VLM and LLM cooperative verification design mechanism, reduces the manual labeling link, and enables the construction efficiency of complex scene data to be improved. Furthermore, through a multi-crossing rechecking mechanism of LLM, the self-error correction capability is established, the error labeling rate is remarkably reduced, and the later correction cost is greatly reduced; by automatically scheduling the expert identity in the LLM, the excellent domain migration capability is realized, the problem of artificial expert dependence is effectively solved, the method can automatically adapt to the new image field, manual intervention is not needed, and the flexibility and adaptability of multi-modal data construction are improved. Compared with a traditional manual labeling scheme and a semi-automatic labeling scheme, the technical disadvantages of insufficient semantic understanding, modal splitting, poor dynamic adaptability and the like are overcome, the manpower disadvantages of high cost, large quality fluctuation, low iteration efficiency and the like are overcome, and a solution is provided for efficiently producing a high-quality multi-modal data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of large model application technology, and in particular to a multimodal data construction method, apparatus, system, computer equipment and computer-readable storage medium based on VLM-LLM collaborative verification and instruction-driven. Background Art

[0002] With the rapid development of artificial intelligence (AI), large multimodal models are increasingly being used in various fields. However, the construction of high-quality multimodal datasets has always been a key bottleneck restricting model performance. Currently, the construction of multimodal datasets relies primarily on three methods: manual annotation, semi-automatic annotation, and automatic generation of large models.

[0003] In traditional manual annotation solutions, annotators need to understand the image content and generate corresponding image-text descriptions and question-answer pairs. However, this method still requires a large amount of manual participation, which is not only costly, but also greatly affects the annotation quality due to the annotator's professional level, resulting in problems such as insufficient semantic understanding and modality fragmentation. Semi-automatic labeling solutions have improved efficiency to a certain extent by introducing automated tools to assist manual labeling; however, such methods still have technical disadvantages such as rough generated content and lack of multimodal verification, as well as high costs for subsequent corrections and difficulties in domain migration.

[0004] In recent years, with the development of big model technology, large-model-based automatic data construction methods have gradually emerged. These methods use large AI models to perform deep learning and semantic understanding of expert knowledge text features to generate question-answer pairs. However, existing large-model solutions still suffer from a series of technical issues: first, the separation of description and reasoning leads to a lack of deep connection between the generated graphic descriptions and the question-answer content; second, insufficient data diversity makes it difficult to meet different scenarios and user needs; third, the lack of effective verification mechanisms easily leads to the risk of error propagation, and the problem of model hallucination is particularly prominent.

[0005] Furthermore, existing technical solutions generally struggle to achieve fine-grained control over the data generation process, resulting in a lack of diversity in the style, depth, and structure of the generated multimodal datasets. In summary, how to efficiently and accurately construct multimodal datasets for training optimization models has become a pressing technical challenge. Summary of the Invention

[0006] The present application provides a method for constructing multimodal data based on VLM-LLM collaborative verification and instruction-driven, the method comprising: Obtain multiple sets of graphic and text descriptions of the image to be processed through the VLM (Visual-Language Model, abbreviated as VLM); Using the Large Language Model (LLM) model, a logical judgment is performed on the image and text description to obtain a target image and text description that is consistent with logic and common sense. Using a command-driven mechanism, controllable QA generation is performed based on the target image and text description and an external domain knowledge base, resulting in multiple sets of detailed question-answer pairs with different styles, depths, and structures. Perform a cross-second verification based on the graphic description, the external domain knowledge base, and multiple sets of detail question-answer pairs to obtain target detail question-answer pairs that meet preset requirements; Based on the target-detail question-answer pairs, the target graphic description, and the image to be processed, a multimodal dataset for model training and optimization is constructed.

[0007] In some embodiments, performing logical judgment on the graphic description using the LLM model to obtain a target graphic description that conforms to logic and common sense includes: Taking multiple groups of image and text descriptions as input, the LLM model is instructed by the second prompt to determine whether the semantic similarity between any group of image and text descriptions is greater than the preset similarity threshold. If so, the group of image and text descriptions is established as the target image and text description; if not, the group of image and text descriptions is discarded.

[0008] In some embodiments, the second prompt is configured to utilize the language analysis expert capabilities of the LLM model to construct content similarity analysis logic, and obtain abnormal graphic and text descriptions caused by model hallucinations through similarity analysis and threshold judgment.

[0009] In some embodiments, a command-driven mechanism is used to perform controllable QA generation based on the target graphic description and an external domain knowledge base, thereby obtaining multiple sets of detailed question-answer pairs of different styles, depths, and structures, including: Instructing the LLM model through a third prompt to use the MOE expert capability to perform domain recognition on the target graphic description and match the domain expert role corresponding to the target graphic description; Utilize the capabilities of the domain expert role and the external domain knowledge base, and construct multiple sets of detailed question-answer pairs based on the prompt instruction and the target graphic description, where: In the process of constructing the detailed question and answer pair, the third prompt is used to instruct the LLM model to dynamically adjust the question and answer style, question and answer depth, and question and answer format.

[0010] In some embodiments, the method further comprises: Inputting the detailed question-answer pairs, target graphic descriptions and external domain knowledge base data into the LLM model; The fourth prompt instructs the LLM model to perform secondary cross-checking and fact verification on the input data in the role of a logical reasoning expert. If the verification passes, the detailed question and answer pair is output.

[0011] In some embodiments, constructing the multimodal dataset based on the target-detail question-answer pair, the target text-image description, and the image to be processed includes: Build a candidate pool of image and text descriptions based on the target image and text descriptions, and build a candidate pool of question and answer based on the target detail question and answer pairs; Through a random selection mechanism, a group of picture and text descriptions are randomly selected from the picture and text description candidate pool, and a group of detail question and answer pairs are randomly selected from the question and answer candidate pool; The randomly selected graphic and text descriptions are combined with detail question-answer pairs, and the multimodal dataset is constructed based on multiple combination results.

[0012] In some embodiments, obtaining multiple sets of graphic and text descriptions of the image to be processed through the VLM model includes: Preprocessing the image to be processed, and instructing the VLM model through a first prompt to process the image to be processed in parallel based on the preprocessed image to obtain multiple groups of graphic and text descriptions; The first VLM model is used to extract global semantics, and the second VLM model is used to extract local detail semantics, respectively obtaining two sets of differentiated initial image and text descriptions. The initial graphic and text descriptions are combined to obtain the graphic and text descriptions.

[0013] In a second aspect, an embodiment of the present application provides a multimodal data construction system based on VLM-LLM collaborative verification and instruction-driven, the system comprising an acquisition module, a construction module and a combination module, wherein: The acquisition module is used to acquire multiple groups of graphic and text descriptions of the image to be processed through the VLM model; The construction module is used to perform logical judgment on the graphic description through the LLM model to obtain a target graphic description that conforms to logic and common sense; Furthermore, using a command-driven mechanism, controllable QA generation is performed based on the target graphic description and an external domain knowledge base, obtaining multiple sets of detailed question-answer pairs with different styles, depths, and structures; Furthermore, a cross-check is performed based on the graphic description, the external domain knowledge base, and multiple groups of detail question-answer pairs to obtain target detail question-answer pairs that meet preset requirements; The combination module is used to construct a multimodal dataset for model training and optimization based on the target detail question-answer pairs, the target graphic description, and the image to be processed.

[0014] In a third aspect, an embodiment of the present application provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect above when executing the computer program.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect above.

[0016] Compared with related technologies, the multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven provided by the embodiment of the present application reduces the manual annotation process through the collaborative verification design mechanism of VLM and LLM, thereby improving the efficiency of complex scene data construction; further, through the multiple cross-review mechanism of LLM, a self-correction capability is established, which significantly reduces the error annotation rate and greatly reduces the cost of later corrections; by automatically scheduling the identity of experts in LLM, excellent domain migration capabilities are achieved, effectively solving the problem of dependence on manual experts, and can automatically adapt to new image fields without manual intervention, thereby improving the flexibility and adaptability of multimodal dataset construction. Compared with traditional manual annotation schemes and semi-automatic annotation schemes, the present invention overcomes technical disadvantages such as insufficient semantic understanding, modal fragmentation, and poor dynamic adaptability, as well as human disadvantages such as high cost, large quality fluctuations, and low iteration efficiency, and provides a solution for the efficient production of high-quality multimodal datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 This is a flowchart of a multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven according to an embodiment of the present application; Figure 2 is a flowchart of another multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven according to an embodiment of the present application; Figure 3 This is a structural block diagram of a multimodal data construction system based on VLM-LLM collaborative verification and instruction-driven according to an embodiment of the present application; Figure 4 Schematic diagram of the internal structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative efforts are within the scope of protection of this application.

[0019] Obviously, the drawings described below are merely examples or embodiments of the present application. Those skilled in the art can, without inventive effort, apply the present application to other similar scenarios based on these drawings. Furthermore, it is also understood that, although the effort involved in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, changes in design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as an insufficiency of the content disclosed in this application.

[0020] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.

[0021] Unless otherwise defined, technical or scientific terms used herein shall have the ordinary meaning as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "an," "the," and similar expressions used herein do not denote quantitative limitations and may refer to either the singular or the plural. The terms "comprise," "include," "have," and any variations thereof, used herein, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules (units) is not limited to the listed steps or units but may also include steps or units not listed, or may include other steps or units inherent to the process, method, product, or apparatus. The terms "connected," "connected," "coupled," and similar expressions used herein are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used herein, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" may mean: A exists alone; A and B exist simultaneously; or B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0022] This application provides a multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven, Figure 1 Flowchart of a multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven according to an embodiment of the present application, such as Figure 1 As shown in the figure, this method builds a high-quality multimodal dataset by combining the collaborative work of the visual language model (VLM) and the large language model (LLM) with an instruction-driven mechanism. The specific implementation steps include: S101, obtaining multiple groups of graphic and text descriptions of the image to be processed through the VLM model.

[0023] The images to be processed are preprocessed, including image resizing, color balancing, and noise removal. The resolution of the preprocessed images is uniformly adjusted to 1024×1024 pixels to ensure consistency in subsequent processing. The first prompt instructs the VLM model to process the preprocessed images in parallel, generating multiple sets of graphic and text descriptions.

[0024] Optionally, the exemplary content of the first prompt is: "Please describe the content of this picture in detail, including the main object, scene environment, color characteristics, spatial relationship and possible activities." Among them, two different VLM models are used for parallel processing: the first VLM model (such as Gemma3) is used to perform global semantic extraction to generate descriptions focusing on the overall scene and main objects; the second VLM model (such as InternVL2.5) is used to perform local detail semantic extraction to generate descriptions focusing on detail features and attributes; these two models respectively obtain two sets of differentiated initial graphic and text descriptions.

[0025] In one exemplary embodiment: An example of a global semantic description generated by the first VLM model: "This is a photo of an outdoor park. In the center of the picture, there is a young woman sitting on a bench reading a book. She is surrounded by lush green vegetation and open lawn areas. The sky is clear, and the sun shines through the leaves, creating mottled shadows." An example of a local detail description generated by the second VLM model: "The woman in the picture is wearing a light blue dress, her long brown hair tied in a ponytail, and holding a book with a red cover. The bench is made of wood and is dark brown. Three tall oak trees can be seen in the background, and there is a small fountain in the distance on the right. The ground is paved with gray stone slabs." Furthermore, in the present application, the two sets of initial graphic and text descriptions are combined to obtain a complete graphic and text description. Optionally, the combination adopts the principle of semantic complementarity to ensure that both global information and local details are preserved.

[0026] The resulting image description example reads: "This is a photo of an outdoor park. In the center of the frame, a young woman sits on a bench and reads a book. The woman is wearing a light blue dress and has long brown hair tied in a ponytail. She is holding a book with a red cover. The bench is made of wood and is dark brown in color. It is surrounded by lush green vegetation and open lawn areas. Three tall oak trees are visible in the background. There is a small fountain in the distance on the right, and the ground is paved with gray stone slabs. The sky is clear, and the sunlight filters through the leaves, creating mottled light and shadows." In practice, we repeat the above process 3-5 times for each image to be processed, using slightly different variations of the prompt word each time, to obtain multiple sets of distinct image and text descriptions. These descriptions form the basic dataset for subsequent processing.

[0027] S102, using the LLM model, logically judge the graphic description to obtain a target graphic description that is consistent with logic and common sense.

[0028] Taking multiple groups of image and text descriptions as input, the LLM model is instructed by the second prompt to determine whether the semantic similarity between any group of image and text descriptions is greater than the preset similarity threshold. If so, the group of image and text descriptions is established as the target image and text description; if not, the group of image and text descriptions is discarded.

[0029] Among them, the second prompt is set to use the language analysis expert capabilities of the LLM model to build content similarity analysis logic, and obtain abnormal graphic and text descriptions caused by model hallucinations through similarity analysis and threshold judgment.

[0030] The specific second prompt example content is: "As a language analysis expert, please analyze the following multiple sets of descriptions about the same image. Calculate the semantic similarity between them and identify inconsistent descriptions that may be caused by model hallucinations. If the semantic similarity between any pair of descriptions is less than 0.75 (with 1 as the full score), please mark the description that may be hallucinated and explain the reason." It should be noted that the preset similarity threshold is set to 0.75, a value determined through experimentation to effectively balance accuracy and diversity. LLM models (such as GPT-4) receive multiple sets of image and text descriptions and calculate the semantic similarity between them. This calculation method is based on a comprehensive score of three dimensions: key entity recognition, attribute matching, and scene consistency.

[0031] Furthermore, for each pair of image-text descriptions, the LLM model extracts core entities (e.g., "young woman," "bench," "park"), attributes (e.g., "light blue dress," "red-cover book"), and scene elements (e.g., "oak tree," "fountain," "cobblestone path"), and cross-references these with other descriptions. If a description contains key elements that are inconsistent with other descriptions (e.g., mentioning "man" instead of "woman," or "cafe" instead of "park"), it is flagged as a possible hallucination.

[0032] Finally, the LLM model outputs the semantic similarity analysis results and labels each description group with a credibility score. Description groups with similarity exceeding a preset threshold are retained as target image-text descriptions, while descriptions that do not meet the requirements are discarded. This step effectively filters out hallucinations that may be generated by the VLM model, improving the data quality for subsequent processing.

[0033] S103 uses an instruction-driven mechanism to perform controllable QA generation based on the target graphic description and external domain knowledge base, obtaining multiple groups of detailed question-answer pairs with different styles, depths, and structures.

[0034] Specifically, the LLM model is instructed through the third prompt to use the MOE expert capability to perform domain recognition on the target graphic description and match the domain expert role corresponding to the target graphic description.

[0035] An example of the third prompt reads: "Please analyze the following image descriptions and identify the primary areas of interest (e.g., natural scenery, urban architecture, human activities, artwork, etc.). As an expert in this field, generate 10 in-depth question-and-answer pairs based on the descriptions. These questions and answers cover different cognitive levels (from basic description to in-depth analysis) and include a variety of question-and-answer formats (e.g., open-ended, closed-ended, comparative, etc.)." In its implementation, the LLM model first analyzes the subject and content of the target graphic description to identify its domain. For example, the aforementioned park scene description identified two key areas: "urban public space" and "leisure activities." Based on this identification, the model then selects the most relevant expert roles from a pre-defined expert role library, such as "urban planning expert" and "leisure sociology expert." External domain knowledge bases are then accessed, including relevant knowledge on urban park design principles, public space usage patterns, and the psychology of leisure activities.

[0036] Furthermore, leveraging the capabilities of the domain expert role and external domain knowledge base, multiple sets of detailed question-answer pairs are constructed based on prompt instructions and target image and text descriptions. During this process, a third prompt instructs the LLM model to dynamically adjust the question-answer style, depth, and format.

[0037] Optionally: Question and answer styles include four basic styles: academic, educational, conversational and exploratory.

[0038] Academic Q&A focuses on professional terminology and theoretical frameworks; educational Q&A focuses on knowledge transfer and concept explanation; conversational Q&A simulates natural communication; and exploratory Q&A encourages creative thinking and hypothesis analysis.

[0039] The depth of question-answering is divided into four levels: foundation, understanding, application, and analysis. The foundation level focuses on directly observable facts; the understanding level explores the principles behind phenomena; the application level discusses the practical application of knowledge; and the analysis level involves critical thinking and multi-perspective evaluation.

[0040] The optional question-answer format includes five types: open-ended, closed-ended, comparative, hypothetical, and evaluative. Open-ended questions have no fixed answers; closed-ended questions have clear answers; comparative questions require comparative analysis; hypothetical questions explore possibilities; and evaluative questions require value judgments.

[0041] Specifically, the generated question-answer pairs are as follows: 1. Basic level, closed, educational type: Question: "What material is the bench the woman is sitting on in the picture made of?" A: "The bench in the picture is made of wood and has a dark brown color. This material is common in park settings because wood has a natural beauty that blends in with the surroundings while providing a comfortable sitting surface." 2. Understanding level, openness, and academic type: Question: "From an urban planning perspective, what are the design considerations for the placement of these park benches?" A: "From an urban planning perspective, the layout of park benches usually takes into account multiple factors: first, the visual landscape. This bench is located in a setting surrounded by greenery, offering a good view. Second, functional zoning. It is located at the transition zone between activity areas and quiet areas. Third, microclimate considerations. The shade of the trees provides natural shade. Finally, social distancing design provides privacy without complete isolation. This arrangement embodies the 'human-scale' design concept, emphasizing the user experience." 3. Application level, comparative, and conversational: Question: “In the digital age, what do you think are the differences between reading a paper book in a park like the one pictured here and reading on a digital device?” A: "Reading a print book in the park is a very different experience compared to using an electronic device. Printed books allow you to fully immerse yourself in the natural environment, without the glare of screens or the worry of battery drain. The tactile experience and the sound of turning pages enhance the sensory experience when reading a physical book, making it easier to become one with your surroundings. While electronic devices are portable and can carry a large number of books, they can easily introduce distractions like social media, and the screens can be difficult to read outdoors in sunlight. Which reading method do you prefer?" 4. Analytical level, evaluative, exploratory: Question: “How would you evaluate the potential impact of the public space design shown in the figure on promoting the mental health of urban residents?” Answer: "The impact of public space design in the figure on urban residents' mental health can be assessed from multiple dimensions: First, the presence of greenery provides a 'natural recovery' environment, which studies have shown can reduce stress hormone levels. Second, the mixed layout of open and semi-private spaces meets different social needs, supporting solitude while not excluding interaction. Third, multi-sensory experiences (visual beauty, natural sounds, fresh air) promote psychological relaxation. Finally, facilities that support sedentary activities such as reading encourage deep reflection and mindfulness. This design embodies the concept of 'healing landscapes' and may be an effective buffer against the stresses of urban life. It is worth considering how to make such spaces more inclusive and serve people of all ages and abilities." It can be understood that in this solution, 20-30 sets of question-answer pairs with different styles, depths, and forms are generated for each image, forming a rich pool of detailed question-answer candidates.

[0042] S104: Perform a cross-second verification based on the graphic description, external domain knowledge base, and multiple sets of detail question and answer pairs to obtain target detail question and answer pairs that meet the preset requirements.

[0043] Specifically, the LLM model is fed with detailed question-answer pairs, target image-text descriptions, and data from an external domain knowledge base. The fourth prompt instructs the LLM model to perform a secondary cross-verification and fact-check on the input data using the logic reasoning expert role. If the verification passes, the detailed question-answer pairs are output.

[0044] An example of the fourth prompt reads: "As a logical reasoning expert, please conduct a rigorous factual and logical review of the following question-answer pairs. Check whether the Q&A content is consistent with the image description, conforms to common sense and domain knowledge, and contains any logical contradictions. For each Q&A pair, please give a 'pass' or 'fail' assessment and explain your reasons. For failed Q&A pairs, please provide correction suggestions." The LLM model performs the following verification process: 1. Fact consistency check: Verify that the factual statements in the question-answer pair are consistent with the target image-text description. For example, if the question-answer pair states "the woman is wearing a red dress," but the image-text description states "a light blue dress," the result is marked as inconsistent.

[0045] 2. Logical coherence check: Verify whether there are any logical contradictions within the question-answer pair. For example, if the question asks "Why did this woman choose to read indoors?" and the image description clearly shows an outdoor scene, it will be marked as a logical error.

[0046] 3. Knowledge Accuracy Check: Verify that the expertise stated in the question-answer pair is consistent with the external domain knowledge base. For example, if the answer mentions "This park design style belongs to the Baroque style," the system will check the information about park design styles in the knowledge base for verification.

[0047] 4. Reasoning plausibility check: Verify whether the reasoning process in the question-answer pair is reasonable. For example, if the answer directly infers "she is a literature student" from "the woman is reading", it will be marked as excessive reasoning.

[0048] For each question-answer pair, the LLM model provides an evaluation result and reasoning. Question-answer pairs that pass the verification are marked as target detail question-answer pairs; those that fail the verification are discarded or re-verified after adjustments based on the correction suggestions.

[0049] Verification example: The question is: "The fountain in the picture is Baroque, which suggests that the park may have been built in the 17th or 18th century." Verification result: Failed, reason: 1) The caption only mentions a small fountain in the distance on the right, without specifying its style; 2) Inferring the age of an entire park from the style of a single fountain is a logical leap; 3) There is insufficient visual detail to support a Baroque style judgment. Suggested corrections: Clearly label speculative content as hypothetical, or simply describe the visible fountain features without making a stylistic judgment.

[0050] It can be understood that through this secondary cross-verification process, high-quality target detail question and answer pairs are screened out to ensure their factual accuracy, logical coherence and knowledge reliability.

[0051] S105: Construct a multimodal dataset for model training and optimization based on the target detail question and answer pairs, target text and image descriptions, and the images to be processed.

[0052] The target image and text descriptions are used to construct a candidate pool of image and text descriptions, and the target detail question and answer pairs are used to construct a candidate pool of question and answer pairs. The image and text description candidate pool includes all target image and text descriptions that pass the logical judgment, and the question and answer candidate pool includes all target detail question and answer pairs that pass the cross-secondary verification.

[0053] Through the random selection mechanism, a group of image and text descriptions are randomly selected from the image and text description candidate pool, and a group of detail question and answer pairs are randomly selected from the question and answer candidate pool.

[0054] The randomly selected image and text descriptions are combined with the detail question and answer pairs, and a multimodal dataset is constructed based on multiple combination results. Specifically, in this embodiment, the combination methods include but are not limited to: Basic combination: directly associate the image, text description, and question-answer pair to form triple data.

[0055] Hierarchical combination: The image is used as the root node, the image and text description as the first-level child nodes, and the question-answer pairs as the second-level child nodes to form tree-structured data.

[0056] Interactive combination: Integrate text and image descriptions into question-answer pairs to form context-enhanced question-answer data.

[0057] For each image to be processed, this solution generates 5-10 different combinations to increase data diversity. Optionally, the final multimodal dataset contains the following fields: image ID: unique identifier, image path: original image storage location, image feature vector: pre-calculated image feature representation, image and text description: selected target image and text description, question and answer pair set: selected target detail question and answer pair set and metadata: including data generation time, model version used, verification status, etc.

[0058] The constructed multimodal dataset can be used for model training and optimization of various downstream tasks, including but not limited to: multimodal dialogue system training, visual question-answering model optimization, image description generation model training, cross-modal retrieval system development, and multimodal understanding ability evaluation. The dataset also includes quality scores and difficulty level annotations to facilitate the selection of appropriate training samples for different training stages.

[0059] also, Figure 2 It is a flowchart of another multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven according to an embodiment of the present application.

[0060] Through the above steps S101 to S105, the multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven provided by the embodiment of the present application reduces the manual annotation link through the collaborative verification design mechanism of VLM and LLM, so that the efficiency of complex scene data construction is improved; further, through the multiple cross-review mechanism of LLM, a self-correction capability is established, which significantly reduces the error annotation rate and greatly reduces the cost of later corrections; through the automatic scheduling of expert identities in LLM, excellent domain migration capabilities are achieved, effectively solving the problem of dependence on manual experts, and can automatically adapt to new image fields without manual intervention, thereby improving the flexibility and adaptability of multimodal dataset construction. Compared with traditional manual annotation schemes and semi-automatic annotation schemes, the present invention overcomes technical disadvantages such as insufficient semantic understanding, modal fragmentation, and poor dynamic adaptability, as well as human disadvantages such as high cost, large quality fluctuations, and low iteration efficiency, and provides a solution for the efficient production of high-quality multimodal datasets.

[0061] This application also provides a multimodal data construction system based on VLM-LLM collaborative verification and instruction-driven, Figure 3 This is a structural block diagram of a multimodal data construction system based on VLM-LLM collaborative verification and instruction-driven according to an embodiment of the present application, such as Figure 3 As shown, the system includes: an acquisition module 30, a construction module 31 and a combination module 32, wherein: The acquisition module 30 is used to acquire multiple groups of graphic and text descriptions of the image to be processed through the VLM model; The images to be processed are preprocessed, including image resizing, color balancing, and noise removal. The resolution of the preprocessed images is uniformly adjusted to 1024×1024 pixels to ensure consistency in subsequent processing. The first prompt instructs the VLM model to process the preprocessed images in parallel, generating multiple sets of graphic and text descriptions. The first prompt reads: "Please describe the content of this image in detail, including the subject, scene environment, color characteristics, spatial relationships, and possible activities." During this process, the system employs two different VLM models for parallel processing: the first VLM (e.g., CLIP-ViT-L-14) extracts global semantics, generating descriptions focused on the overall scene and key objects; the second VLM (e.g., BLIP-2) extracts local semantics, generating descriptions focused on detailed features and attributes. These two models produce two distinct sets of initial image-text descriptions.

[0062] The construction module 31 is used to perform logical judgment on the graphic description through the LLM model to obtain a target graphic description that conforms to logic and common sense; and to use the instruction-driven mechanism to perform controllable QA generation based on the target graphic description and the external domain knowledge base to obtain multiple sets of detail question and answer pairs with different styles, depths, and structures; and to perform cross-secondary verification based on the graphic description, the external domain knowledge base, and the multiple sets of detail question and answer pairs to obtain target detail question and answer pairs that meet preset requirements. Among them, multiple groups of image and text descriptions are used as input, and the LLM model is instructed by the second prompt to determine whether the semantic similarity between any group of image and text descriptions is greater than a preset similarity threshold. If so, the group of image and text descriptions is established as the target image and text description; if not, the group of image and text descriptions is discarded.

[0063] This module instructs the LLM model through the third prompt to leverage the MOE expert capabilities to identify the domain of the target image and text description and match it with a corresponding domain expert role. It then accesses an external domain knowledge base and, leveraging the capabilities of the domain expert role and the external domain knowledge base, constructs multiple sets of detailed question-and-answer pairs based on the prompt instructions and the target image and text description. During this process, the third prompt instructs the LLM model to dynamically adjust the question-and-answer style, depth, and format.

[0064] The detailed question-answer pairs, target image and text descriptions, and external domain knowledge base data are then input into the LLM model. The fourth prompt instructs the LLM model to perform a secondary cross-check and fact-check on the input data using the logic reasoning expert role. If the verification passes, the detailed question-answer pairs are output.

[0065] The combination module 32 is used to construct a multimodal dataset for model training and optimization based on the target detail question-answer pairs, the target graphic description, and the image to be processed.

[0066] The target image and text descriptions are used to construct a candidate pool of image and text descriptions, and the target detail question and answer pairs are used to construct a candidate pool of question and answer pairs. The image and text description candidate pool includes all target image and text descriptions that pass the logical judgment, and the question and answer candidate pool includes all target detail question and answer pairs that pass the cross-secondary verification.

[0067] A random selection mechanism randomly selects a set of descriptions from the candidate pool of descriptions and a set of question-answer pairs from the candidate pool of questions and answers. These randomly selected descriptions are combined with the question-answer pairs, and a multimodal dataset is constructed based on the resulting combinations. The random selection mechanism uses a weighted random sampling algorithm, with weights based on the completeness score of the descriptions and the quality score of the question-answer pairs.

[0068] Through this system, through the collaborative verification design mechanism of VLM and LLM, full process automation is achieved, manual labeling links are reduced, and the efficiency of complex scene data construction is improved; further, through the multiple cross-checking mechanism of LLM, self-correction capabilities are established, which significantly reduces the error labeling rate and greatly reduces the cost of later corrections; through the automatic scheduling of expert identities in LLM, excellent domain migration capabilities are achieved, effectively solving the problem of dependence on manual experts, and can automatically adapt to new image fields without manual intervention, thereby improving the flexibility and adaptability of multimodal dataset construction. Compared with traditional manual labeling solutions and semi-automatic labeling solutions, the present invention overcomes technical disadvantages such as insufficient semantic understanding, modality fragmentation, and poor dynamic adaptability, as well as human disadvantages such as high cost, large quality fluctuations, and low iteration efficiency, providing a solution for the efficient production of high-quality multimodal datasets.

[0069] In one embodiment, Figure 4 is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application, such as Figure 4 As shown, an electronic device is provided, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 4 As shown. The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected via an internal bus, wherein the non-volatile memory stores an operating system, a computer program, and a database. The processor is used to provide computing and control capabilities, the network interface is used to communicate with external terminals via a network connection, the internal memory is used to provide an environment for the operation of the operating system, the computer program, when executed by the processor, implements a multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven, and the database is used to store data.

[0070] Those skilled in the art will understand that Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. Specifically, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0071] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements any one of the multimodal data construction methods based on VLM-LLM collaborative verification and instruction-driven in the above embodiments.

[0072] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0073] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A multimodal data construction method based on VLM-LLM collaborative verification and instruction-driven, characterized in that: The method comprises: Obtain multiple sets of graphic and text descriptions of the image to be processed through the VLM model; Through the LLM model, the graphic description is logically judged to obtain the target graphic description that conforms to logic and common sense; Using a command-driven mechanism, controllable QA generation is performed based on the target image and text description and an external domain knowledge base, resulting in multiple sets of detailed question-answer pairs with different styles, depths, and structures. Perform a cross-second verification based on the graphic description, the external domain knowledge base, and multiple sets of detail question-answer pairs to obtain target detail question-answer pairs that meet preset requirements; Based on the target-detail question-answer pairs, the target graphic description, and the image to be processed, a multimodal dataset for model training and optimization is constructed.

2. The method according to claim 1, characterized in that By using the LLM model, the graphic description is logically judged to obtain the target graphic description that conforms to logic and common sense, including: Taking multiple groups of image and text descriptions as input, the LLM model is instructed by the second prompt to determine whether the semantic similarity between any group of image and text descriptions is greater than the preset similarity threshold. If so, the group of image and text descriptions is established as the target image and text description; if not, the group of image and text descriptions is discarded.

3. The method according to claim 2, characterized in that The second prompt is configured to utilize the language analysis expert capabilities of the LLM model to construct content similarity analysis logic, and obtain abnormal graphic and text descriptions caused by model hallucinations through similarity analysis and threshold judgment.

4. The method according to claim 1, wherein Using a command-driven mechanism, we perform controllable QA generation based on the target image and text description and an external domain knowledge base, obtaining multiple sets of detailed question-answer pairs with different styles, depths, and structures, including: Instructing the LLM model through a third prompt to use the MOE expert capability to perform domain recognition on the target graphic description and match the domain expert role corresponding to the target graphic description; Utilize the capabilities of the domain expert role and the external domain knowledge base, build multiple sets of detailed question-answer pairs based on the prompt instruction and the target graphic description, where: In the process of constructing the detailed question and answer pair, the third prompt is used to instruct the LLM model to dynamically adjust the question and answer style, question and answer depth, and question and answer format.

5. The method according to claim 4, characterized in that The method further comprises: Inputting the detailed question-answer pairs, target graphic descriptions and external domain knowledge base data into the LLM model; The fourth prompt instructs the LLM model to perform secondary cross-checking and fact verification on the input data in the role of a logical reasoning expert. If the verification passes, the detailed question and answer pair is output.

6. The method according to claim 1, characterized in that Constructing the multimodal dataset based on the target-detail question-answer pair, the target graphic-text description, and the image to be processed includes: Build a candidate pool of image and text descriptions based on the target image and text descriptions, and build a candidate pool of question and answer based on the target detail question and answer pairs; Through a random selection mechanism, a group of picture and text descriptions are randomly selected from the picture and text description candidate pool, and a group of detail question and answer pairs are randomly selected from the question and answer candidate pool; The randomly selected graphic and text descriptions are combined with detail question-answer pairs, and the multimodal dataset is constructed based on multiple combination results.

7. The method according to claim 1, characterized in that Obtaining multiple sets of graphic and text descriptions of the image to be processed through the VLM model includes: Preprocessing the image to be processed, and instructing the VLM model through a first prompt to process the image to be processed in parallel based on the preprocessed image to obtain multiple groups of graphic and text descriptions; The first VLM model is used to extract global semantics, and the second VLM model is used to extract local detail semantics, respectively obtaining two sets of differentiated initial image and text descriptions. The initial graphic and text descriptions are combined to obtain the graphic and text descriptions.

8. A multimodal data construction system based on VLM-LLM collaborative verification and instruction drive, characterized in that: The system includes an acquisition module, a construction module and a combination module, wherein: The acquisition module is used to acquire multiple groups of graphic and text descriptions of the image to be processed through the VLM model; The construction module is used to perform logical judgment on the graphic description through the LLM model to obtain a target graphic description that conforms to logic and common sense; Furthermore, using a command-driven mechanism, controllable QA generation is performed based on the target graphic description and an external domain knowledge base, obtaining multiple sets of detailed question-answer pairs with different styles, depths, and structures; Furthermore, a cross-check is performed based on the graphic description, the external domain knowledge base, and multiple groups of detail question-answer pairs to obtain target detail question-answer pairs that meet preset requirements; The combination module is used to construct a multimodal dataset for model training and optimization based on the target detail question-answer pairs, the target graphic description, and the image to be processed.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Electric power defect image detection method based on image-text question-answer multi-modal model

    CN117763107A

  • Visual language model construction method based on machine learning

    CN118470158A

  • Multi-modal large model-based sequence character bill image question and answer data generation method

    CN119169650A

  • System and method transforming visual commonsense reasoning as commonsense reasoning and visual recognition with large language models

    US20250078462A1

  • Chat system and method based on visual content dialogue

    WO2024230846A1

Cited By

  • Large model workflow-based data processing and fine tuning data synthesis method, system and device, and medium

    CN121053667A