Knowledge-intensive visual question and answer automatic data generation method and device

By employing a collaborative mechanism between the main agent and domain expert agents, and multi-agent role-playing, combined with deep semantic similarity algorithms and multimodal reward model evaluation, the shortcomings of visual question answering systems in terms of professionalism and accuracy are addressed, and high-quality visual question answering data generation is achieved.

CN120930746APending Publication Date: 2025-11-11TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202511043738.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing visual question answering systems suffer from insufficient professionalism, fluctuating accuracy, lack of logical consistency, and limited interpretability when faced with specific domain expertise and complex knowledge scenarios. They struggle to generate content with professional depth and are prone to 'illusion' phenomena that do not match the given visual content.

Method used

By employing a dynamic collaboration mechanism between the main agent and domain expert agents, a three-level prompting system is constructed, which includes domain knowledge, evaluation criteria, and generation specifications. Combined with a keyframe sampling algorithm based on deep semantic similarity and a multimodal reward model evaluation method, data generation for multi-agent role-playing is performed. Furthermore, an adversarial generation strategy is used to construct interference samples, thereby achieving the automated generation of high-quality data.

Benefits of technology

It significantly improves the professionalism, accuracy, and logical consistency of the generated data, ensuring that the generated content has both professional depth and diversity, and solves quality defects such as insufficient professionalism, poor semantic accuracy, and lack of contextual coherence in the generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930746A_ABST
    Figure CN120930746A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge-intensive visual question and answer automatic data generation method and device, and the method comprises the steps: constructing an original visual data set containing the professional knowledge of a target domain according to a static image, a video stream and multimedia content; extracting a representative frame sequence, converting the audio information into text information, and extracting character information in the static image to construct a structured visual instance database; according to the prompt text meeting the preset professional depth condition, establishing a three-level prompt system containing domain knowledge, an evaluation standard and a generation specification; generating a corresponding visual question and answer pair data set according to the dynamic cooperation of the main agent and the domain expert agent; generating a multi-agent quality evaluation system according to the quality evaluation result; and designing a difficulty grading mechanism according to the negative example sample. According to the method, the professionality, the accuracy and the diversity of the visual question and answer data are remarkably improved, and reliable data support is provided for training and evaluation of a multi-modal large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of multimodal artificial intelligence and intelligent data processing technology, and in particular to a knowledge-intensive visual question-answering automated data generation method and apparatus. Background Technology

[0002] With the rapid development of knowledge-enhanced multimodal large-scale modeling technology, significant breakthroughs have been achieved in cross-modal knowledge understanding capabilities, encompassing both visual and linguistic aspects. Modern systems can now generate fluent and natural text responses with deep knowledge based on user-provided images or videos. However, traditional visual question-answering benchmarks typically only include short answers consisting of 1-5 keywords, a format that cannot meet the needs of current generative dialogue systems for evaluating complex knowledge representations and long texts from specialized domains. While manual annotation can produce free-form responses based on professional knowledge, its high knowledge acquisition cost severely restricts large-scale application. This situation has driven the transformation of visual question-answering tasks from manual annotation to automation.

[0003] Currently, automated data generation methods based on knowledge-enhanced multimodal large language models demonstrate the ability to simultaneously parse image semantic information and generate natural language responses through a pre-trained knowledge framework that integrates visual encoders and language decoders. While knowledge-driven multimodal large language models have shown application potential in scenarios such as image description generation and question-answering reasoning, these technologies still face systemic challenges when dealing with domain-specific expertise and complex knowledge scenarios. These challenges include insufficient specialization, fluctuating accuracy, lack of logical consistency, and limited knowledge interpretability. These issues urgently necessitate technological improvements to build a knowledge-complete, efficient, and reliable automated annotation system.

[0004] Current visual question answering systems mainly rely on manual rule design and crowdsourced annotation for data labeling. This method is not only costly and difficult to standardize, but also fails to meet the efficiency requirements of large-scale data processing.

[0005] While automated annotation techniques based on multimodal large models have alleviated the pressure of manual annotation to some extent, the question-answer pairs they generate often have significant limitations due to the lack of domain expertise. On the one hand, they struggle to produce content with professional depth; on the other hand, they are prone to producing "illusion" phenomena that do not match the given visual content. These problems manifest as quality defects such as insufficient professionalism in generated content, poor semantic accuracy, and lack of contextual coherence, which have been addressed through several generations of improvements. Summary of the Invention

[0006] This application provides a knowledge-intensive visual question-answering automated data generation method and apparatus to solve the problems of related technologies, which on the one hand are difficult to produce content with professional depth, and on the other hand are prone to the "illusion" phenomenon that does not match the given visual content. Specifically, the generated content has quality defects such as insufficient professionalism, poor semantic accuracy, and lack of contextual coherence.

[0007] The first aspect of this application provides a knowledge-intensive visual question-answering automated data generation method, comprising the following steps: collecting static images, video streams, and multimedia content containing audio from multiple agents, and constructing an original visual dataset containing target domain expertise based on the static images, video streams, and multimedia content; extracting keyframes from the video data of the original visual dataset, and based on the keyframes, converting audio information into text information, and extracting text information from the static images, to construct a structured visual instance database that satisfies preset complete semantic representation conditions based on the text information; establishing a three-level prompting system containing domain knowledge, evaluation criteria, and generation specifications based on the structured visual instance database and prompt text that satisfies preset professional depth conditions; and generating knowledge-intensive visual question-answering automated data based on the three-level prompting system and the dynamic collaboration between the main agent and the domain expert agent.

[0008] Optionally, in one embodiment of this application, after generating knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the dynamic collaboration between the main agent and the domain expert agent, the method further includes: generating a visual question-answering pair dataset based on the knowledge-intensive visual question-answering automation data, and obtaining the quality assessment results of the visual question-answering pair dataset to generate a quality assessment system for the multi-agent based on the quality assessment results; based on the quality assessment system and at least one positive sample, establishing at least one negative sample for the multi-agent through semantic perturbation, and designing a difficulty grading mechanism based on the at least one negative sample to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the difficulty grading mechanism.

[0009] Optionally, in one embodiment of this application, the step of extracting keyframes from the video data of the original visual dataset includes: extracting high-dimensional feature vectors of the image using a pre-trained visual model; calculating cosine similarity between adjacent frames based on the high-dimensional feature vectors to obtain deep semantic similarity between consecutive frames based on the cosine similarity, and determining whether the semantic similarity between adjacent frames is lower than a preset dynamic threshold; if the semantic similarity between adjacent frames is lower than the preset dynamic threshold, then determining the keyframes of the representative frame sequence based on the adjacent frames.

[0010] Optionally, in one embodiment of this application, the step of establishing a three-level prompting system containing domain knowledge, evaluation criteria, and generation specifications based on prompt text that meets preset professional depth conditions includes: constructing generation guidance prompts containing the domain knowledge based on the professional characteristics of the target application scenario, and designing professional skill description texts for each role in the multi-agent system based on the guidance prompts; writing standard data to generate standard examples, and designing data perturbation examples containing typical interference factors; and establishing the three-level prompting system based on the professional skill description texts, the standard examples, and the data perturbation examples.

[0011] Optionally, in one embodiment of this application, the step of generating knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the three-level prompting system and the dynamic collaboration between the main agent and the domain expert agent includes: generating preliminary data annotations based on preprocessed structured visual data; evaluating the quality of the annotations based on the preliminary data annotations, generating annotation quality evaluation results, and determining whether the content complexity of the annotation quality evaluation results exceeds a preset range; if the content complexity exceeds the preset range, iteratively integrating the annotation opinions of the main agent and the domain expert agent until the annotation opinions reach the target quality standard, and then combining the annotation opinions to generate the knowledge-intensive visual question-answering automation data for multi-agent role-playing.

[0012] Optionally, in one embodiment of this application, obtaining the quality assessment result of the visual question-answering dataset and generating the quality assessment system of the multi-agent based on the quality assessment result includes: using a pre-built multimodal reward model to perform a quality assessment on the visual question-answering data to generate a score for the visual question-answering data; quantifying the score into a reward value for the visual question-answering data, and comparing the reward value with a preset reward threshold to generate the quality assessment system.

[0013] Optionally, in one embodiment of this application, the step of establishing at least one negative example sample of the multi-agent through semantic perturbation, and designing a difficulty grading mechanism based on the at least one negative example sample to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing according to the difficulty grading mechanism, includes: preprocessing structured visual data, and applying at least one of the following multi-level semantic interference operations to the preprocessed structured visual data: replacement of key visual elements, distortion of semantic relationships, and interference with textual information, to create at least one negative example sample of the multi-agent; based on the at least one negative example sample, inputting noisy data into a pre-constructed multi-agent generation framework to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing.

[0014] A second aspect of this application provides a knowledge-intensive visual question-answering automated data generation device, comprising: a data acquisition module for acquiring static images, video streams, and multimedia content containing audio from multiple agents, and constructing an original visual dataset containing target domain expertise based on the static images, video streams, and multimedia content; an extraction module for extracting keyframes from the video data of the original visual dataset, converting audio information into text information based on the keyframes, and extracting text information from the static images to construct a structured visual instance database that meets preset complete semantic representation conditions based on the text information; a building module for establishing a three-level prompting system containing domain knowledge, evaluation criteria, and generation specifications based on the structured visual instance database and prompt text that meets preset professional depth conditions; and a generation module for generating knowledge-intensive visual question-answering automated data based on the three-level prompting system and the dynamic collaboration between the main agent and the domain expert agent.

[0015] Optionally, in one embodiment of this application, it further includes: an acquisition module, configured to generate a visual question-answering pair dataset based on the knowledge-intensive visual question-answering automation data for multi-agent role-playing after generating such data based on the dynamic collaboration between the main agent and the domain expert agent, and to acquire the quality assessment results of the visual question-answering pair dataset, so as to generate a quality assessment system for the multi-agent based on the quality assessment results; and a design module, configured to establish at least one negative sample of the multi-agent based on the quality assessment system and at least one positive sample through semantic perturbation, and to design a difficulty grading mechanism based on the at least one negative sample, so as to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the difficulty grading mechanism.

[0016] Optionally, in one embodiment of this application, the extraction module includes: an extraction unit, configured to extract high-dimensional feature vectors of the image using a pre-trained visual model; a judgment unit, configured to calculate cosine similarity between adjacent frames based on the high-dimensional feature vectors, to obtain deep semantic similarity between consecutive frames based on the cosine similarity, and to determine whether the semantic similarity between adjacent frames is lower than a preset dynamic threshold; and a determination unit, configured to determine keyframes of the representative frame sequence based on the adjacent frames when the semantic similarity between adjacent frames is lower than the preset dynamic threshold.

[0017] Optionally, in one embodiment of this application, the establishment module includes: a construction unit, configured to construct generation guidance prompts containing domain knowledge based on the professional characteristics of the target application scenario, and design professional skill description texts for each role in the multi-agent system based on the guidance prompts; a design unit, configured to write standard data to generate standard examples, and design data perturbation examples containing typical interference factors; and an establishment unit, configured to establish a three-level prompt system based on the professional skill description texts, the standard examples, and the data perturbation examples.

[0018] Optionally, in one embodiment of this application, the generation module includes: a data annotation generation unit, used to generate preliminary data annotations based on preprocessed structured visual data; an evaluation unit, used to evaluate the quality of the annotations based on the preliminary data annotations, how to generate annotation quality evaluation results, and determine whether the content complexity of the annotation quality evaluation results exceeds a preset range; and a generation unit, used to iteratively integrate the annotation opinions of the main intelligent agent and the domain expert intelligent agent when the content complexity exceeds the preset range, until the annotation opinions reach the target quality standard, and then synthesize the annotation opinions to generate the knowledge-intensive visual question-answering automation data for multi-agent role-playing.

[0019] Optionally, in one embodiment of this application, the acquisition module includes: a score generation unit, used to perform quality assessment on the visual question-answering data using a pre-built multimodal reward model to generate a score for the visual question-answering data; and a comparison unit, used to quantify the score into a reward value for the visual question-answering data and compare the reward value with a preset reward threshold to generate the quality assessment system.

[0020] Optionally, in one embodiment of this application, the design module includes: a preprocessing unit, configured to preprocess structured visual data and apply at least one of the following multi-level semantic interference operations to the preprocessed structured visual data: replacement of key visual elements, distortion of semantic relationships, and interference with textual information, to create at least one negative example sample of the multi-agent; and an input unit, configured to input noisy data into a pre-constructed multi-agent generation framework based on the at least one negative example sample, to generate knowledge-intensive visual question-answering automation data for the multi-agent role-playing.

[0021] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the knowledge-intensive visual question-answering automated data generation method as described in the above embodiments.

[0022] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described knowledge-intensive visual question-answering automated data generation method.

[0023] This application's embodiments achieve iterative optimization of annotation quality through a dynamic collaboration mechanism between the main intelligent agent and domain expert intelligent agents, significantly improving the professionalism, accuracy, and logical consistency of the generated data. The method creatively constructs a three-tiered professional prompting engineering system encompassing domain knowledge points, evaluation criteria, and generation specifications. Combined with expert-designed standard examples and perturbation examples, it ensures that the generated data possesses both professional depth and diversity. In terms of technical implementation, this application designs a dynamic threshold keyframe sampling algorithm based on deep semantic similarity, achieving intelligent filtering through a pre-trained visual model; develops a multimodal reward model evaluation method, jointly evaluating from three dimensions: semantic consistency, logical coherence, and linguistic fluency; and innovatively employs an adversarial generation strategy to construct interference samples with apparent plausibility. These technological innovations collectively constitute a complete end-to-end automated data generation process, from multimodal data acquisition, preprocessing, intelligent generation to quality evaluation and interference sample generation, achieving a dual improvement in the efficiency and quality of visual question-answering data generation, providing reliable data support for the training and evaluation of large multimodal models. This solves the problem that related technologies struggle to produce content with professional depth on the one hand, and are prone to producing "illusion" phenomena that do not match the given visual content on the other. Specifically, this manifests as quality defects such as insufficient professionalism of the generated content, poor semantic accuracy, and lack of contextual coherence.

[0024] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0025] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0026] Figure 1 This is a flowchart of a knowledge-intensive visual question-answering automated data generation method according to an embodiment of this application;

[0027] Figure 2 This is a schematic diagram illustrating the specific process of a knowledge-intensive visual question-answering automated data generation method according to an embodiment of this application;

[0028] Figure 3 This is a flowchart illustrating a knowledge-intensive visual question-answering automated data generation method according to an embodiment of this application;

[0029] Figure 4This is an implementation diagram of a specialized prompting engineering method for a knowledge-intensive visual question answering automated data generation method according to an embodiment of this application;

[0030] Figure 5 This is an implementation diagram of a specialized prompting engineering method for a knowledge-intensive visual question answering automated data generation method according to an embodiment of this application;

[0031] Figure 6 This is a schematic diagram of the structure of a knowledge-intensive visual question-answering automated data generation device according to an embodiment of this application;

[0032] Figure 7 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation

[0033] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0034] The following description, with reference to the accompanying drawings, illustrates a knowledge-intensive visual question-answering automated data generation method and apparatus according to embodiments of this application. Addressing the aforementioned background technologies, which struggle to produce content with professional depth and are prone to producing "illusion" content that doesn't match the given visual content (specifically manifested as insufficient professionalism, poor semantic accuracy, and lack of contextual coherence), this application provides a knowledge-intensive visual question-answering automated data generation method. This method utilizes a dynamic collaboration mechanism between the main intelligent agent and a domain expert intelligent agent to iteratively optimize annotation quality, significantly improving the professionalism, accuracy, and logical consistency of the generated data. This method creatively constructs a three-tiered professional prompting engineering system encompassing domain knowledge points, evaluation criteria, and generation specifications. Combined with expert-designed standard examples and perturbation examples, it ensures that the generated data possesses both professional depth and diversity. In terms of technical implementation, this invention designs a dynamic threshold keyframe sampling algorithm based on deep semantic similarity, achieving intelligent filtering through a pre-trained visual model; it develops a multimodal reward model evaluation method, jointly evaluating from three dimensions: semantic consistency, logical coherence, and linguistic fluency; and it innovatively employs an adversarial generation strategy to construct interference samples with apparent plausibility. These technological innovations collectively constitute a complete end-to-end automated data generation process, from multimodal data acquisition, preprocessing, intelligent generation to quality evaluation and interference sample generation, achieving a dual improvement in the efficiency and quality of visual question-answering data generation, providing reliable data support for the training and evaluation of large multimodal models. This solves the problems of related technologies, which on the one hand struggle to produce content with professional depth, and on the other hand are prone to producing "illusion" phenomena that do not match the given visual content, specifically manifested as insufficient professionalism in generated content, poor semantic accuracy, and lack of contextual coherence.

[0035] Specifically, Figure 1 This is a flowchart illustrating a knowledge-intensive visual question-answering automated data generation method provided in an embodiment of this application.

[0036] like Figure 1 As shown, this knowledge-intensive visual question-answering automated data generation method includes the following steps:

[0037] In step S101, static images, video streams, and multimedia content containing audio of the multi-agent system are collected, and an original visual dataset containing target domain expertise is constructed based on the static images, video streams, and multimedia content.

[0038] In actual implementation, such as Figure 2As shown, embodiments of this application can perform multimodal data acquisition: constructing raw visual datasets containing domain-specific expertise, with acquisition objects including still images, video streams, and multimedia content containing audio. To enhance the authenticity and professionalism of the generated data, during the collection of specified raw visual content, it is also necessary to collect corresponding metadata, such as image / video captions, background information, category, source, and other textual information.

[0039] In step S102, keyframes of representative frame sequences are extracted from the video data of the original visual dataset. Based on the keyframes, audio information is converted into text information, and text information in static images is extracted to construct a structured visual instance database that meets the preset complete semantic representation conditions.

[0040] In actual implementation, the embodiments of this application can perform data preprocessing and feature extraction: keyframes of representative frame sequences are extracted from the video data of the original visual dataset using a keyframe sampling algorithm to ensure the integrity of important visual information. Based on the keyframes, audio information is converted into text information using speech recognition technology, and text information in the image is extracted using OCR technology. A structured visual instance database that meets the preset complete semantic representation conditions is constructed based on the text information. For example, in the structured processing stage, the system associates and integrates the extracted visual features and text features according to the smallest semantic unit. The video data uses complete narrative segments as the basic content unit, and each instance contains a visual feature vector, a text feature description, and a corresponding semantic boundary marker, thereby constructing a structured visual instance database with complete semantic representation capabilities.

[0041] This application presents an embodiment of a dynamic threshold keyframe sampling algorithm based on deep semantic similarity, which achieves intelligent filtering through a pre-trained visual model.

[0042] Optionally, in one embodiment of this application, extracting keyframes from the video data of the original visual dataset includes: extracting high-dimensional feature vectors of the image using a pre-trained visual model; calculating the cosine similarity between adjacent frames based on the high-dimensional feature vectors to obtain the deep semantic similarity between consecutive frames based on the cosine similarity, and determining whether the semantic similarity between adjacent frames is lower than a preset dynamic threshold; if the semantic similarity between adjacent frames is lower than the preset dynamic threshold, then determining the keyframes of the representative frame sequence based on the adjacent frames.

[0043] Specifically, the preprocessing process for the original visual content in this embodiment is as follows: After the original visual content is acquired, the system automatically performs key information and feature extraction operations. The keyframe sampling algorithm achieves intelligent filtering by calculating the deep semantic similarity between consecutive frames. This similarity is obtained by calculating the cosine similarity of the high-dimensional feature vectors extracted by the pre-trained visual model. When the semantic similarity between adjacent frames is lower than a preset dynamic threshold, the system will simultaneously retain the adjacent frame as a keyframe to ensure the integrity of important visual information.

[0044] In step S103, based on the structured visual instance database, a three-level prompting system is established according to the prompting text that meets the preset professional depth conditions, which includes domain knowledge, evaluation criteria and generation specifications.

[0045] In actual implementation, the embodiments of this application can implement a professional prompting project: based on a structured visual instance database, domain experts write prompting texts with professional depth, and establish a three-level prompting system that includes domain knowledge, evaluation criteria and generation specifications.

[0046] This application's embodiments creatively construct a three-tiered professional prompting engineering system that includes domain knowledge points, evaluation criteria, and generation specifications. Combined with expert-designed standard examples and perturbation examples, it ensures that the generated data possesses both professional depth and diversity.

[0047] It should be noted that the preset professional depth conditions can be set by those skilled in the art according to the actual situation, and no specific restrictions are imposed here.

[0048] Optionally, in one embodiment of this application, a three-level prompting system comprising domain knowledge, evaluation criteria, and generation specifications is established based on prompting text that meets preset professional depth conditions. This includes: constructing generation guidance prompts containing domain knowledge based on the professional characteristics of the target application scenario, and designing professional skill description texts for each role in the multi-agent system based on the guidance prompts; compiling standard data to generate standard examples, and designing data perturbation examples containing typical interference factors; and establishing a three-level prompting system based on the professional skill description texts, standard examples, and data perturbation examples.

[0049] The specific implementation process of the specialized prompting project in this application embodiment includes:

[0050] (1) Domain experts first construct generation guidance prompts containing key domain knowledge points based on the professional characteristics of the target application scenario, and design professional skill description texts for each role in the multi-agent system.

[0051] (2) Domain experts develop standard data generation examples to establish a question-answer pair paradigm that meets professional requirements, including question expression, answer depth, and standardization of professional terminology.

[0052] (3) To improve the robustness of the generated data, domain experts also designed data perturbation examples containing typical interference factors. These examples simulate semantic shifts, professional expression deviations and other situations that may occur in real scenarios, providing a benchmark for subsequent quality assessment.

[0053] The embodiments of this application can be constructed through three levels of prompting engineering to ensure that the generated data has both professional depth and robustness, and to ensure that the generated data has the necessary diversity and reliability while maintaining professional depth.

[0054] In step S104, based on the three-level prompting system, knowledge-intensive visual question-answering automation data for multi-agent role-playing is generated according to the dynamic collaboration between the main agent and the domain expert agent.

[0055] In actual implementation, the embodiments of this application can perform multi-agent collaborative data generation: based on a three-level prompting system, a data generation algorithm based on multi-agent role-playing is designed and implemented to automatically generate visual question-answer pairs. That is, the embodiments of this application can generate knowledge-intensive visual question-answering automated data based on the dynamic collaboration between the main agent and the domain expert agent.

[0056] The embodiments of this application can develop a technical solution that can automatically generate high-quality professional visual question-answering data, and subsequently establish a rapid quality assessment system to support it, thereby significantly improving the professionalism, accuracy and logical consistency of the generated data.

[0057] Optionally, in one embodiment of this application, based on a three-level prompting system, knowledge-intensive automated visual question-answering data for multi-agent role-playing is generated according to the dynamic collaboration between the main agent and the domain expert agent. This includes: generating preliminary data annotations based on preprocessed structured visual data; evaluating the quality of the annotations based on the preliminary data annotations, generating annotation quality evaluation results, and determining whether the content complexity of the annotation quality evaluation results exceeds a preset range; if the content complexity exceeds the preset range, iteratively integrating the annotation opinions of the main agent and the domain expert agent until the annotation opinions reach the target quality standard, and then combining the annotation opinions to generate knowledge-intensive automated visual question-answering data for multi-agent role-playing.

[0058] As one possible approach, the multi-agent collaborative data generation method achieves automated generation of high-quality visual question-answer pairs through a role-playing mechanism. Its specific implementation process includes three stages:

[0059] (1) Initialization phase: A master agent generates preliminary data annotations based on preprocessed structured visual data. These annotations include basic visual feature descriptions and initial question-answer pairs.

[0060] (2) Multi-Agent Collaboration Stage: The system employs a dynamic role allocation and knowledge fusion mechanism to iteratively optimize annotation quality. Specifically, the Master Agent performs multi-dimensional analysis of the current annotations using the judgment rules written by domain experts in the aforementioned specialization suggestion project. When the accuracy of professional terms is detected to be lower than a set threshold or the semantic complexity exceeds a predetermined range, the expert collaboration mechanism is triggered. This mechanism first generates corresponding expert profiles from the professional skill descriptions written by domain experts in the aforementioned specialization suggestion project, and then instantiates corresponding expert agent instances. Each expert agent is equipped with a professional domain knowledge base and a domain-specific annotation rule engine. The expert agent generates new data annotations based on its own professional skills and sends them to the Master Agent. The Master Agent uses an attention-based fusion algorithm to integrate the opinions of multiple experts, generate new data annotations, and re-evaluates the overall quality indicators after each iteration. The collaboration process terminates when any of the following conditions are met: all data quality reaches the preset standard; the quality improvement in three consecutive iterations is less than the set threshold; or the maximum number of collaboration rounds is reached.

[0061] (3) The main agent integrates the annotation opinions of all collaborating agents and outputs the final data annotation result after multi-level optimization.

[0062] The embodiments of this application achieve iterative optimization of annotation quality based on dynamic role allocation and collaborative decision-making mechanisms, effectively improving the performance of generated data in terms of professionalism, accuracy and logical consistency.

[0063] Optionally, in one embodiment of this application, after generating knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the dynamic collaboration between the main agent and the domain expert agent, the method further includes: generating a visual question-answering pair dataset based on the knowledge-intensive visual question-answering automation data, and obtaining the quality assessment results of the visual question-answering pair dataset to generate a quality assessment system for multi-agents based on the quality assessment results; based on the quality assessment system and at least one positive sample, establishing at least one negative sample for multi-agents through semantic perturbation, and designing a difficulty grading mechanism based on at least one negative sample to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the difficulty grading mechanism.

[0064] In practical implementation, embodiments of this application can generate a visual question-answering pair dataset based on knowledge-intensive automated visual question-answering data, and perform quality assessment on the generated dataset to obtain quality assessment results. Based on these results, a multi-agent quality assessment system can be generated. Embodiments of this application can also generate interference samples: negative examples are generated based on positive samples through semantic perturbation, and a difficulty grading mechanism is designed.

[0065] Optionally, in one embodiment of this application, obtaining the quality assessment results of the visual question answering dataset and generating a multi-agent quality assessment system based on the quality assessment results includes: using a pre-built multimodal reward model to assess the quality of the visual question answering data to generate a score for the visual question answering data; quantifying the score into a reward value for the visual question answering data, and comparing the reward value with a preset reward threshold to generate a quality assessment system.

[0066] It is understood that this application proposes an innovative multimodal reward model evaluation method for automatically detecting the quality of visual question-answering data. This method, through a multimodal reward model constructed using a deep neural network, can simultaneously jointly model visual content, question text, and generated answers, outputting a comprehensive quality score.

[0067] Specifically, when given a complete question-and-answer pair containing visual context, a question, and a corresponding answer, the reward model evaluates the pair from the following dimensions: (1) semantic consistency: the degree of matching between the answer and the visual content; (2) logical coherence: the reasonableness of the causal relationship between the answer and the question; and (3) linguistic fluency: the accuracy and fluency of natural language expression. The model quantifies the above evaluation dimensions into reward values ​​between 0 and 1. When the reward value is lower than a preset threshold (e.g., 0.6), the system automatically marks the question-and-answer pair as potentially having quality problems, including but not limited to: errors in understanding visual content, logical contradictions, lack of professional knowledge, or incomplete language expression.

[0068] The embodiments of this application are based on an automated quality detection method with multi-dimensional joint evaluation, which performs joint evaluation from three dimensions: semantic consistency, logical coherence and language fluency, significantly improving the reliability and usability of large-scale visual question answering data generation.

[0069] Optionally, in one embodiment of this application, at least one negative example sample of a multi-agent is established through semantic perturbation, and a difficulty grading mechanism is designed based on the at least one negative example sample to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing. This includes: preprocessing structured visual data, and applying at least one of the following multi-level semantic interference operations to the preprocessed structured visual data: replacement of key visual elements, distortion of semantic relationships, and interference with textual information, to create at least one negative example sample of a multi-agent; and inputting noisy data into a pre-constructed multi-agent generation framework based on the at least one negative example sample to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing.

[0070] It is understood that the embodiments of this application propose a method for constructing interference samples based on an adversarial generation strategy, which creates highly deceptive negative samples by systematically introducing semantic perturbations.

[0071] In actual implementation, the embodiments of this application can process structured visual data and apply multi-level semantic interference operations to the preprocessed structured visual data, including but not limited to (1) replacement of key visual elements: replacing specific objects or attributes in the image while maintaining the coherence of the scene context; (2) distortion of semantic relationships: modifying the spatial or logical relationships between visual elements; (3) interference with text information: performing operations such as synonym replacement, negative word insertion, or logical connector tampering on the question or answer text. These interference operations are not completely random, but are based on the deep understanding of the original data by the multimodal large model, ensuring that the generated noise data contains substantial semantic deviations under the surface rationality. Subsequently, the processed noise data is input into the multi-agent generation framework in the above-mentioned multi-agent collaborative data generation step, and question-answer pairs with professional appearance but substantial errors are generated through the same generation process. In particular, the interference samples generated by this method have three typical characteristics: surface rationality (conforming to professional expression norms), semantic deception (concealed contradictions with visual content), and difficulty gradation (controllable interference intensity).

[0072] This application's embodiments generate seemingly plausible interference samples through systematic semantic perturbation operations, providing high-quality negative examples for model testing. It innovatively employs an adversarial generation strategy to construct seemingly plausible interference samples, and the systematic interference sample generation method provides a high-quality negative example data foundation for subsequent model robustness testing and evaluation system construction.

[0073] This application achieves automated generation of high-quality visual question-answering data by constructing a complete end-to-end processing flow. The innovations together constitute a complete end-to-end automated data generation process, from multimodal data acquisition, preprocessing, intelligent generation to quality assessment and interference sample generation, achieving a dual improvement in the efficiency and quality of visual question-answering data generation, and providing reliable data support for the training and evaluation of large multimodal models.

[0074] Specifically, it can be combined with Figures 3 to 5 As shown, the working principle of the knowledge-intensive visual question-answering automated data generation method in this application is explained in detail with a specific embodiment.

[0075] like Figure 3 As shown, embodiments of this application may include the following steps:

[0076] Step S301: Multimodal data collection.

[0077] Step S302: Data preprocessing and feature extraction.

[0078] Step S303: Specialized prompting project.

[0079] Step S304: Data generation through multi-agent collaboration.

[0080] Step S305: Quality assessment system.

[0081] Step S306: Generation of interference samples.

[0082] Furthermore, such as Figure 4 As shown, embodiments of this application may include the following steps:

[0083] Step S401: Construct generation guidance prompts containing key domain knowledge points and design professional skill description text for multi-agent systems.

[0084] Step S402: Write a standard data generation example and establish a question-and-answer pair paradigm that meets professional requirements.

[0085] Step S403: Design a typical example of data perturbation for interference factors.

[0086] like Figure 5 As shown, embodiments of this application may include the following steps:

[0087] Step S501: The main intelligent agent generates preliminary data annotations.

[0088] Step S502: The main agent performs dynamic role allocation, initializes the expert agent, and the expert agent generates data labels.

[0089] Step S503: The main agent integrates the annotation opinions of all expert agents and outputs the optimized final data annotation results.

[0090] The knowledge-intensive automated visual question-answering data generation method proposed in this application achieves iterative optimization of annotation quality through a dynamic collaboration mechanism between the main intelligent agent and the domain expert intelligent agent, significantly improving the professionalism, accuracy, and logical consistency of the generated data. This method creatively constructs a three-level professional prompting engineering system encompassing domain knowledge points, evaluation criteria, and generation specifications. Combined with expert-designed standard examples and perturbation examples, it ensures that the generated data possesses both professional depth and diversity. In terms of technical implementation, this invention designs a dynamic threshold keyframe sampling algorithm based on deep semantic similarity, achieving intelligent filtering through a pre-trained visual model; develops a multimodal reward model evaluation method, jointly evaluating from three dimensions: semantic consistency, logical coherence, and linguistic fluency; and innovatively employs an adversarial generation strategy to construct interference samples with apparent plausibility. These technological innovations collectively constitute a complete end-to-end automated data generation process, from multimodal data acquisition, preprocessing, intelligent generation to quality evaluation and interference sample generation, achieving a dual improvement in the efficiency and quality of visual question-answering data generation, and providing reliable data support for the training and evaluation of large multimodal models. This solves the problem that related technologies struggle to produce content with professional depth on the one hand, and are prone to producing "illusion" phenomena that do not match the given visual content on the other. Specifically, this manifests as quality defects such as insufficient professionalism in the generated content, poor semantic accuracy, and lack of contextual coherence.

[0091] Next, with reference to the accompanying drawings, a knowledge-intensive visual question-answering automated data generation apparatus according to an embodiment of this application is described.

[0092] Figure 6 This is a schematic diagram of the structure of the knowledge-intensive visual question-answering automated data generation device according to an embodiment of this application.

[0093] like Figure 6 As shown, the knowledge-intensive visual question-answering automated data generation device 10 includes: a data acquisition module 100, an extraction module 200, a data creation module 300, and a data generation module 400.

[0094] Specifically, the acquisition module 100 is used to acquire static images, video streams, and multimedia content containing audio from multiple agents, and to construct an original visual dataset containing target domain expertise based on the static images, video streams, and multimedia content.

[0095] The extraction module 200 is used to extract keyframes of representative frame sequences from the video data of the original visual dataset, and based on the keyframes, convert audio information into text information and extract text information from static images, so as to construct a structured visual instance database that meets the preset complete semantic representation conditions based on the text information and the text information.

[0096] Module 300 is established to create a three-level prompting system based on a structured visual instance database and prompting text that meets preset professional depth conditions. This system includes domain knowledge, evaluation criteria, and generation specifications.

[0097] The generation module 400 is used to generate knowledge-intensive visual question-answering automation data based on a three-level prompting system and the dynamic collaboration between the main agent and the domain expert agent.

[0098] Optionally, in one embodiment of this application, the knowledge-intensive visual question-answering automated data generation device 10 further includes an acquisition module and a design module.

[0099] The acquisition module is used to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the dynamic collaboration between the main agent and the domain expert agent, generate a visual question-answering pair dataset based on the knowledge-intensive visual question-answering automation data, and obtain the quality evaluation results of the visual question-answering pair dataset, so as to generate a multi-agent quality evaluation system based on the quality evaluation results.

[0100] The design module is used to establish at least one negative sample for multi-agent roles based on a quality assessment system and at least one positive sample, through semantic perturbation, and to design a difficulty grading mechanism based on at least one negative sample, so as to generate knowledge-intensive visual question answering automation data for multi-agent role-playing according to the difficulty grading mechanism.

[0101] Optionally, in one embodiment of this application, the extraction module 200 includes: an extraction unit, a judgment unit, and a determination unit.

[0102] The extraction unit is used to extract high-dimensional feature vectors from images using a pre-trained visual model.

[0103] The judgment unit is used to calculate the cosine similarity between adjacent frames based on the high-dimensional feature vector, so as to obtain the deep semantic similarity between consecutive frames based on the cosine similarity, and to determine whether the semantic similarity between adjacent frames is lower than the preset dynamic threshold.

[0104] The determination unit is used to determine the key frames of a representative frame sequence based on adjacent frames when the semantic similarity between adjacent frames is lower than a preset dynamic threshold.

[0105] Optionally, in one embodiment of this application, the establishment module 300 includes: a construction unit, a design unit, and an establishment unit.

[0106] The construction unit is used to build generation guidance prompts containing domain knowledge based on the professional characteristics of the target application scenario, and to design professional skill description texts for each role in the multi-agent system based on the guidance prompts.

[0107] The design unit is used to compile standard data to generate standard examples and to design data perturbation examples that include typical interference factors.

[0108] Establish a unit to create a three-level prompting system based on professional skill description text, standard examples, and data perturbation examples.

[0109] Optionally, in one embodiment of this application, the generation module 400 includes: a data annotation generation unit, an evaluation unit, and a generation unit.

[0110] The data annotation generation unit is used to generate preliminary data annotations based on the preprocessed structured visual data.

[0111] The evaluation unit is used to evaluate the quality of annotations based on preliminary data, to generate annotation quality evaluation results, and to determine whether the complexity of the annotation quality evaluation results exceeds a preset range.

[0112] The generation unit is used to iteratively integrate the annotation opinions of the main intelligent agent and the domain expert intelligent agent when the content complexity exceeds the preset range, until the annotation opinions reach the target quality standard, and then synthesize the annotation opinions to generate knowledge-intensive visual question-answering automation data with multi-agent role-playing.

[0113] Optionally, in one embodiment of this application, the acquisition module includes a scoring generation unit and a comparison unit.

[0114] The scoring generation unit is used to evaluate the quality of visual question-answering data using a pre-built multimodal reward model in order to generate a score for the visual question-answering data.

[0115] The comparison unit is used to quantify the score into a reward value for the visual question-answering data and compare the reward value with a preset reward threshold to generate a quality assessment system.

[0116] Optionally, in one embodiment of this application, the design module includes a preprocessing unit and an input unit.

[0117] The preprocessing unit is used to preprocess the structured visual data and apply at least one of the following multi-level semantic interference operations to the preprocessed structured visual data: replacement of key visual elements, distortion of semantic relationships, and interference with textual information, in order to create at least one negative sample of the multi-agent system.

[0118] The input unit is used to input noisy data into a pre-built multi-agent generation framework based on at least one negative sample to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing.

[0119] It should be noted that the foregoing explanation of the embodiment of the knowledge-intensive visual question answering automated data generation method also applies to the knowledge-intensive visual question answering automated data generation device of this embodiment, and will not be repeated here.

[0120] The knowledge-intensive automated visual question-answering data generation device proposed in this application achieves iterative optimization of annotation quality through a dynamic collaboration mechanism between the main intelligent agent and the domain expert intelligent agent, significantly improving the professionalism, accuracy, and logical consistency of the generated data. This method creatively constructs a three-level professional prompting engineering system encompassing domain knowledge points, evaluation criteria, and generation specifications. Combined with expert-designed standard examples and perturbation examples, it ensures that the generated data possesses both professional depth and diversity. In terms of technical implementation, this invention designs a dynamic threshold keyframe sampling algorithm based on deep semantic similarity, achieving intelligent filtering through a pre-trained visual model; develops a multimodal reward model evaluation method, jointly evaluating from three dimensions: semantic consistency, logical coherence, and linguistic fluency; and innovatively employs an adversarial generation strategy to construct interference samples with apparent plausibility. These technological innovations collectively constitute a complete end-to-end automated data generation process, from multimodal data acquisition, preprocessing, intelligent generation to quality evaluation and interference sample generation, achieving a dual improvement in the efficiency and quality of visual question-answering data generation, and providing reliable data support for the training and evaluation of large multimodal models. This solves the problem that related technologies struggle to produce content with professional depth on the one hand, and are prone to producing "illusion" phenomena that do not match the given visual content on the other. Specifically, this manifests as quality defects such as insufficient professionalism in the generated content, poor semantic accuracy, and lack of contextual coherence.

[0121] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:

[0122] The memory 701, the processor 702, and the computer program stored on the memory 701 and executable on the processor 702.

[0123] When the processor 702 executes the program, it implements the knowledge-intensive visual question-answering automated data generation method provided in the above embodiments.

[0124] Furthermore, electronic devices also include:

[0125] Communication interface 703 is used for communication between memory 701 and processor 702.

[0126] The memory 701 is used to store computer programs that can run on the processor 702.

[0127] The memory 701 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0128] If the memory 701, processor 702, and communication interface 703 are implemented independently, then the communication interface 703, memory 701, and processor 702 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0129] Optionally, in a specific implementation, if the memory 701, processor 702, and communication interface 703 are integrated on a single chip, then the memory 701, processor 702, and communication interface 703 can communicate with each other through an internal interface.

[0130] The processor 702 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0131] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described knowledge-intensive visual question-answering automated data generation method.

[0132] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0133] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0134] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0135] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0136] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0137] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0138] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0139] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A knowledge-intensive visual question-answering automated data generation method, characterized in that, Includes the following steps: Collect static images, video streams, and multimedia content containing audio from multiple agents, and construct an original visual dataset containing target domain expertise based on the static images, video streams, and multimedia content; Keyframes of representative frame sequences are extracted from the video data of the original visual dataset. Based on the keyframes, audio information is converted into text information, and text information is extracted from the static images. A structured visual instance database that meets the preset complete semantic representation conditions is constructed based on the text information. Based on the structured visual instance database, a three-level prompting system is established according to the prompting text that meets the preset professional depth conditions, which includes domain knowledge, evaluation criteria and generation specifications. Based on the aforementioned three-level prompting system, knowledge-intensive automated visual question-answering data for multi-agent role-playing is generated through dynamic collaboration between the main agent and the domain expert agent.

2. The method according to claim 1, characterized in that, After generating knowledge-intensive, automated visual question-answering data for multi-agent role-playing based on the dynamic collaboration between the main agent and the domain expert agent, it also includes: A visual question-answering pair dataset is generated based on the knowledge-intensive automated visual question-answering data, and the quality assessment results of the visual question-answering pair dataset are obtained, so as to generate the quality assessment system of the multi-agent based on the quality assessment results; Based on the quality assessment system and at least one positive sample, at least one negative sample of the multi-agent is established through semantic perturbation, and a difficulty grading mechanism is designed according to the at least one negative sample to generate knowledge-intensive visual question answering automation data for multi-agent role-playing according to the difficulty grading mechanism.

3. The method according to claim 1, characterized in that, The extraction of keyframes from the video data of the original visual dataset includes: High-dimensional feature vectors of the image are extracted using a pre-trained visual model; The cosine similarity between adjacent frames is calculated based on the high-dimensional feature vector, and the deep semantic similarity between consecutive frames is obtained based on the cosine similarity. It is then determined whether the semantic similarity between adjacent frames is lower than a preset dynamic threshold. If the semantic similarity between adjacent frames is lower than the preset dynamic threshold, then the key frames of the representative frame sequence are determined based on the adjacent frames.

4. The method according to claim 1, characterized in that, The three-level prompting system, which includes domain knowledge, evaluation criteria, and generation specifications, is established based on prompt text that meets preset professional depth conditions. This includes: Based on the professional characteristics of the target application scenario, a generation guidance prompt containing the domain knowledge is constructed, and based on the guidance prompt, a professional skill description text for each role in the multi-agent system is designed. Compile standard data to generate standard examples, and design data perturbation examples that include typical interference factors; Establish a three-level prompting system based on the professional skill description text, the standard example, and the data perturbation example.

5. The method according to claim 1, characterized in that, The aforementioned three-level prompting system generates knowledge-intensive automated visual question-answering data based on the dynamic collaboration between the main agent and the domain expert agent, including: Preliminary data annotations are generated based on the preprocessed structured visual data; Based on the preliminary data annotation, the quality of the annotation is evaluated, how to generate the annotation quality evaluation result, and whether the content complexity of the annotation quality evaluation result exceeds a preset range; If the complexity of the content exceeds the preset range, the annotation opinions of the main agent and the domain expert agent are iteratively integrated until the annotation opinions reach the target quality standard. Then, the annotation opinions are combined to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing.

6. The method according to claim 2, characterized in that, The step of obtaining the quality assessment results of the visual question-answering pair dataset, and generating the quality assessment system of the multi-agent system based on the quality assessment results, includes: The visual question-answering data is evaluated for quality using a pre-built multimodal reward model to generate a score for the visual question-answering data; The score is quantified into a reward value for the visual question-answering data, and the reward value is compared with a preset reward threshold to generate the quality assessment system.

7. The method according to claim 2, characterized in that, The process involves establishing at least one negative example sample for the multi-agent system through semantic perturbation, and designing a difficulty grading mechanism based on the at least one negative example sample to generate knowledge-intensive visual question-answering automation data for multi-agent role-playing based on the difficulty grading mechanism, including: Preprocess structured visual data and apply at least one of the following multi-level semantic interference operations to the preprocessed structured visual data: replacement of key visual elements, distortion of semantic relationships, and interference with textual information, to create at least one negative example sample of the multi-agent. Based on the at least one negative sample, noisy data is input into a pre-built multi-agent generation framework to generate knowledge-intensive automated visual question answering data for multi-agent role-playing.

8. A knowledge-intensive visual question-answering automated data generation device, characterized in that, include: The acquisition module is used to acquire static images, video streams, and multimedia content containing audio from multiple agents, and to construct a raw visual dataset containing target domain expertise based on the static images, video streams, and multimedia content. The extraction module is used to extract keyframes of representative frame sequences from the video data of the original visual dataset, and based on the keyframes, convert audio information into text information and extract text information from the static images, so as to construct a structured visual instance database that meets the preset complete semantic representation conditions based on the text information and the text information. A module is established to create a three-level prompting system based on the structured visual instance database and prompting text that meets preset professional depth conditions, which includes domain knowledge, evaluation criteria, and generation specifications. The generation module is used to generate knowledge-intensive visual question-answering automation data based on the three-level prompting system and the dynamic collaboration between the main agent and the domain expert agent.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the knowledge-intensive visual question-answering automated data generation method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the knowledge-intensive visual question-answering automated data generation method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Answer quality evaluation method and system based on divide-and-conquer agent

    CN118551851A

  • Large language model knowledge question-answering method and system fused with multi-modal knowledge graph

    CN118627628A

  • Question answering system construction method based on knowledge graph and related equipment

    CN118839756A

  • ME-RAG-based data center large model intelligent operation and maintenance method and system

    CN120011523A

  • Autonomous decision-making method of learning main body based on deep reinforcement learning

    JP2024028097A

Cited By

  • Test set construction method and device based on multi-modal knowledge tree, equipment and medium

    CN121278391A

  • Real machine data processing method of intelligent agent and related equipment thereof

    CN121366302A

  • Real-machine data processing methods and related equipment for intelligent agents

    CN121366302B

  • Method and system for automatically generating video based on multi-agent unstructured knowledge

    CN121418636A

  • Method and device for generating medical question and answer pairs

    CN121561117A