Multi-agent system, report generation method and device for vertical domain scenario
By leveraging the division of labor and collaboration within a multi-agent system and utilizing a vertical domain large language model and a multimodal reasoning model, the accuracy problem in generating cross-modal data analysis reports was solved, enabling efficient and accurate data analysis and report generation in vertical domain scenarios.
Patent Information
- Application Number
- CN202511461660.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing technologies have low accuracy in cross-modal data fusion and analysis result report generation, and cannot effectively handle multimodal data in complex and professional vertical scenarios, resulting in insufficient breadth of data utilization, depth of analysis and process efficiency.
A multi-agent system is adopted, including a multimodal retrieval enhancement agent, a report generation agent, and a vertical domain multimodal data analysis agent. Through division of labor and cooperation, the entire process from user input to data analysis report generation is automated. The interaction and analysis are carried out using a vertical domain large language model and a multimodal reasoning model.
It improves data processing efficiency and report generation accuracy, and is suitable for the analysis needs of complex multimodal data in vertical domain scenarios, realizing the automated and customized generation of data analysis reports.
Smart Images

Figure CN120929780B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular, to a multi-agent system for vertical domain scenarios, a report generation method and device. BACKGROUND
[0002] In today's information age, various vertical domain scenarios (such as smart city, emergency safety, intelligent transportation, etc.) have accumulated massive cross-modal data. With the growth of cross-modal data, analysis and reporting of cross-modal data are crucial.
[0003] In the prior art, an agent system based on a large language model has shown great potential in natural language interaction and automatic task execution, providing a train of thought for the analysis and reporting of cross-modal data. However, the prior art has low accuracy in processing cross-modal data fusion and generating analysis result reports.
[0004] Therefore, how to analyze cross-modal data and improve the accuracy of generating data analysis reports is a problem that needs to be solved. SUMMARY
[0005] Embodiments of the present application provide a multi-agent system for vertical domain scenarios, a report generation method and device to analyze cross-modal data and improve the accuracy of generating data analysis reports.
[0006] In a first aspect, embodiments of the present application provide a multi-agent system for vertical domain scenarios, the multi-agent system comprising: a multi-modal retrieval enhanced agent, a report generation agent, and a vertical multi-modal data analysis agent;
[0007] The multi-modal retrieval enhanced agent is configured to determine a target agent based on input information of a user, and in response to the target agent being the report generation agent, send the input information to the report generation agent.
[0008] The report generation agent is configured to determine retrieval reference data and a data analysis task based on the input information, and send the retrieval reference data and the data analysis task to the vertical multi-modal data analysis agent.
[0009] The vertical multi-modal data analysis agent is configured to retrieve based on the retrieval reference data to obtain to-be-analyzed data, perform data analysis on the to-be-analyzed data based on the data analysis task to obtain a first data analysis result, and feed back the to-be-analyzed data and the first data analysis result to the report generation agent.
[0010] The report generation agent is configured to generate a data analysis report based on the to-be-analyzed data and the first data analysis result.
[0011] In a possible implementation, the report generation agent is further configured to:
[0012] Before the data analysis report is generated based on the data to be analyzed and the first data analysis result, a target report template is determined based on at least one of the input information, the retrieved reference data, and the first data analysis result.
[0013] The report generation agent is specifically configured to fill the data to be analyzed and the first data analysis result into the target report template to obtain the data analysis report.
[0014] In a possible implementation, the vertical multi-modal data analysis agent includes a vertical multi-modal inference model and a vertical large language model, and is specifically configured to:
[0015] When the modality of the input information is a text modality, at least one round of interaction is performed between the vertical large language model and the user.
[0016] When either of the first case and the second case exists, interaction is performed between the vertical multi-modal inference model and the user, the first case is that the modality of the input information includes a non-text modality, and the second case is that the modality of the data analysis report includes a non-text modality.
[0017] In a possible implementation, the report generation agent is further configured to:
[0018] After the data analysis report is generated based on the data to be analyzed and the first data analysis result, the data analysis report is sent to the multi-modal retrieval enhancement agent.
[0019] The multi-modal retrieval enhancement agent is further configured to output the data analysis report.
[0020] In a possible implementation, the vertical multi-modal data analysis agent is further configured to:
[0021] Before the data to be analyzed is obtained by retrieval based on the retrieved reference data, a target database is determined from a plurality of candidate databases based on the retrieved reference data.
[0022] In response to the target database being a relational database, a database query statement is generated based on the retrieved reference data by using a code generation model.
[0023] The vertical multi-modal data analysis agent is specifically configured to query the relational database based on the database query statement to obtain the data to be analyzed.
[0024] In a possible implementation, the multi-modal retrieval enhanced agent is further configured to:
[0025] In response to the target agent being the vertical multi-modal data analysis agent, send the input information to the vertical multi-modal data analysis agent.
[0026] The vertical multi-modal data analysis agent is further configured to:
[0027] based on the modality of the input information, perform retrieval on the input information to obtain a retrieval result, and based on the retrieval result, perform data analysis to obtain a second data analysis result; and send the second data analysis result to the multi-modal retrieval enhanced agent.
[0028] In a second aspect, a report generation method is provided. The method is applied to a multi-agent system, and the multi-agent system includes a multi-modal retrieval enhanced agent, a report generation agent, and a vertical multi-modal data analysis agent. The method includes:
[0029] based on input information of a user, determining a target agent by using the multi-modal retrieval enhanced agent; and in response to the target agent being the report generation agent, sending the input information to the report generation agent.
[0030] based on the input information, determining retrieval reference data and a data analysis task by using the report generation agent; and sending the retrieval reference data and the data analysis task to the vertical multi-modal data analysis agent.
[0031] based on the retrieval reference data, performing retrieval to obtain to-be-analyzed data by using the vertical multi-modal data analysis agent; based on the data analysis task, performing data analysis on the to-be-analyzed data to obtain a first data analysis result; and feeding back the to-be-analyzed data and the first data analysis result to the report generation agent.
[0032] based on the to-be-analyzed data and the first data analysis result, generating a data analysis report by using the report generation agent.
[0033] In a third aspect, a report generation apparatus is provided. The apparatus is applied to a multi-agent system, and the multi-agent system includes a multi-modal retrieval enhanced agent, a report generation agent, and a vertical multi-modal data analysis agent. The apparatus includes:
[0034] The first determining module is configured to determine a target intelligent agent based on the input information of the user by the multi-modal retrieval enhanced intelligent agent, and send the input information to the report generation intelligent agent in response to the target intelligent agent being the report generation intelligent agent.
[0035] The second determining module is configured to determine retrieval reference data and a data analysis task based on the input information by the report generation intelligent agent, and send the retrieval reference data and the data analysis task to the vertical multi-modal data analysis intelligent agent.
[0036] The analysis module is configured to perform retrieval based on the retrieval reference data by the vertical multi-modal data analysis intelligent agent to obtain to-be-analyzed data, perform data analysis on the to-be-analyzed data based on the data analysis task to obtain a first data analysis result, and feed back the to-be-analyzed data and the first data analysis result to the report generation intelligent agent.
[0037] The generation module is configured to generate a data analysis report based on the to-be-analyzed data and the first data analysis result by the report generation intelligent agent.
[0038] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor.
[0039] The memory stores computer execution instructions.
[0040] The processor executes the computer execution instructions stored in the memory, so that the processor executes the method of any one of the second aspect.
[0041] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by a processor to implement the method of any one of the second aspect.
[0042] In a sixth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the method of any one of the second aspect.
[0043] The multi-agent system, the report generation method and the device for the vertical field scene provided by the embodiments of the present application can enhance the target agent determination and the input information forwarding of the agent through multi-modal retrieval, determine the retrieval reference data and the task of the report generation agent and send them to the vertical multi-modal data analysis agent, retrieve data, analyze and feed back the results through the vertical multi-modal data analysis agent, and the report generation agent can also generate a data analysis report. Through the agent architecture of division of labor and cooperation, the multi-agent system realizes the full-process automatic processing from user input to data analysis report generation, improves the data processing efficiency and the accuracy of report generation, and is suitable for the analysis needs of complex multi-modal data in the vertical field scene. BRIEF DESCRIPTION OF DRAWINGS
[0044] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0046] Figure 1 A schematic diagram of a multi-agent system for a vertical field scene provided by the embodiments of the present application;
[0047] Figure 2 A flowchart of a vertical large language model training process provided by the embodiments of the present application;
[0048] Figure 3 A flowchart of a method for obtaining an image-text pair provided by the embodiments of the present application;
[0049] Figure 4 A model architecture diagram of a preset model provided by the embodiments of the present application;
[0050] Figure 5 A schematic diagram of a training process of a vertical multi-modal reasoning large model provided by the embodiments of the present application;
[0051] Figure 6 A schematic diagram of constructing a multi-modal database and a domain knowledge graph provided by the embodiments of the present application;
[0052] Figure 7 A flowchart of constructing a domain knowledge graph provided by the embodiments of the present application;
[0053] Figure 8 A schematic diagram of a relational database query service provided by the embodiments of the present application;
[0054] Figure 9 is a flow diagram of a report generation method provided by an embodiment of the present application;
[0055] Figure 10 is an example architecture diagram of a multi-agent system in a vertical domain scenario provided by an embodiment of the present application;
[0056] Figure 11 is an architecture diagram of a retrieval model provided by an embodiment of the present application;
[0057] Figure 12 is a schematic diagram of a training process of a multi-modal embedding model provided by an embodiment of the present application;
[0058] Figure 13 is a structure diagram of a report generation apparatus provided by the present application;
[0059] Figure 14 is a structure diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0060] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is with reference to the drawings, in which like numerals represent like elements, unless otherwise described in connection with the following figures. The following exemplary embodiments described in the following examples are not representative of all embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application, as detailed in the appended claims.
[0061] In the present application, the term "comprising" and its variants can refer to non-limiting inclusion; the term "or" and its variants can refer to "and / or". In the present application, the terms "first", "second", etc. are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. In the present application, "a plurality of" means two or more. "And / or", which describes the relationship between the associated objects, means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects.
[0062] With the rapid development of artificial intelligence technology, various industries (i.e., "vertical scenarios") have accumulated massive amounts of data assets during their production and operation processes. This data is diverse, including structured data stored in traditional relational databases (such as user information, device status, and transaction records), as well as increasingly abundant unstructured or semi-structured data such as images, videos, text logs, and various IoT time-series sensor data. This data spans multiple modalities, suffers from low utilization rates, and exhibits significant data silo problems. How to efficiently and intelligently integrate and analyze this multimodal data to extract deeper value and drive business decisions has become a core challenge.
[0063] In recent years, agent systems based on large language models (LLMs) have shown great potential in natural language interaction and automated task execution, providing new solutions to the aforementioned challenges. However, existing technical solutions still have significant shortcomings in terms of data utilization breadth, analysis depth, process efficiency, and development costs when dealing with complex and specialized vertical scenarios, which seriously restricts their large-scale application and implementation.
[0064] Current large-scale model-based data analysis agents use Text to Structured Query Language (Text2SQL) technology. This technology, combined with system prompts or relevant knowledge bases, first converts query commands into Structured Query Language (SQL) statements before performing database queries. This process can only query relational data and cannot query other modal data such as images and videos. This greatly limits the application scenarios and allows for only simple analysis, without the ability to perform in-depth professional analysis using specialized algorithms or tools.
[0065] For example, current AI applications in various vertical scenarios suffer from lengthy processes and unsatisfactory results. Taking structured services for people and vehicles as an example, it involves multiple steps such as object detection, pedestrian / vehicle attribute recognition, SQL query command construction, and relational database query. Errors accumulate gradually during this process, leading to poor final application results. Furthermore, the requirements for AI implementation vary across different vertical scenarios, necessitating the construction of various neural network models and the creation of corresponding datasets, resulting in a huge workload for development.
[0066] Therefore, the multi-agent system, report generation method, and apparatus for vertical domain scenarios provided in this application can determine the target agent and forward input information through a multimodal retrieval augmentation agent, determine the retrieval reference data and tasks through a report generation agent and send them to the vertical domain multimodal data analysis agent, retrieve data, analyze it, and provide feedback results through the vertical domain multimodal data analysis agent, and the report generation agent can also generate a data analysis report. This multi-agent system, through a collaborative agent architecture, achieves fully automated processing from user input to data analysis report generation, improving data processing efficiency and report generation accuracy, and is suitable for the analysis needs of complex multimodal data in vertical domain scenarios.
[0067] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0068] Figure 1 This is a schematic diagram of a multi-agent system in a vertical domain scenario provided in an embodiment of this application. Figure 1 As shown, the system includes: a multimodal retrieval enhancement agent, a report generation agent, and a vertical domain multimodal data analysis agent.
[0069] It should be noted that the intelligent agents can connect through an interface, through the same network, or through other communication methods; this application does not limit this.
[0070] The multimodal retrieval augmented agent can be used to determine the target agent based on user input information. Furthermore, in response to the target agent being a report-generating agent, the multimodal retrieval augmented agent can send the input information to the report-generating agent.
[0071] Optionally, the multimodal retrieval augmented agent can have the ability to receive user input information and determine the target agent based on the information. The user input information can be information provided by the user to the multi-agent system to initiate a task or express a need, and can include various forms such as text and images.
[0072] Optionally, the multimodal retrieval augmentation agent can determine which agent should perform the subsequent tasks based on the user input information.
[0073] The report generation agent is used to determine the retrieval reference data and data analysis tasks based on the input information; the retrieval reference data and data analysis tasks are then sent to the vertical domain multimodal data analysis agent.
[0074] Optionally, the report-generating agent can organize the data analysis results into a data analysis report. Reference data retrieval can guide the vertical-domain multimodal data analysis agent to retrieve data relevant to the current task from massive datasets. The data analysis task clarifies the specific analytical operations and objectives that the vertical-domain multimodal data analysis agent needs to perform on the data to be analyzed, such as analyzing data trends and correlations.
[0075] The vertical domain multimodal data analysis agent can be used to retrieve data based on reference data to obtain the data to be analyzed; perform data analysis on the data to be analyzed based on the data analysis task to obtain the first data analysis result; and feed the data to be analyzed and the first data analysis result back to the report generation agent.
[0076] Optionally, the vertical domain multimodal data analysis agent can perform in-depth analysis of multimodal data to extract valuable information from the data. The data to be analyzed can be a set of data that needs in-depth analysis, retrieved from the data source based on the reference data. The first data analysis result can be the preliminary analysis result obtained by the vertical domain multimodal data analysis agent after processing the data to be analyzed according to the data analysis task.
[0077] The report-generating agent can also be used to generate data analysis reports based on the data to be analyzed and the results of initial data analysis. Optionally, the data analysis report can be a document that presents the data analysis process and results in a specific format, used to convey the conclusions and recommendations of the data analysis to users.
[0078] The multi-agent system provided in this application can determine the target agent and forward input information through a multimodal retrieval augmentation agent, determine the retrieval reference data and tasks through a report generation agent and send them to the vertical domain multimodal data analysis agent, which then retrieves, analyzes, and provides feedback on the data. The report generation agent can also generate a data analysis report. This multi-agent system, through a collaborative agent architecture, achieves fully automated processing from user input to data analysis report generation, improving data processing efficiency and report generation accuracy, and is suitable for analyzing complex multimodal data in vertical domain scenarios.
[0079] In one implementation, the report generating agent is further configured to determine a target report template based on at least one of input information, retrieved reference data, and the first data analysis result before generating a data analysis report based on the data to be analyzed and the first data analysis result, and then fill the target report template with the data to be analyzed and the first data analysis result to obtain the data analysis report.
[0080] Optionally, the target report template can be a template used to standardize the format and content of data analysis reports. Different templates may correspond to different report styles, structures, or presentation focuses.
[0081] Optionally, the report generation agent can determine the target report template based on any one of the input information, retrieved reference data, and the first data analysis result. Alternatively, the report generation agent can determine the target report template based on any two of the input information, retrieved reference data, and the first data analysis result. Or, the report generation agent can determine the target report template based on all three: input information, retrieved reference data, and the first data analysis result.
[0082] Optionally, the report generating agent can accurately place the data to be analyzed and the results of the first data analysis into the corresponding areas of the template according to the format and position specified in the target report template to complete the filling.
[0083] The embodiments of this application can determine the target report template based on input information, retrieved reference data, or first data analysis results, and can generate customized reports according to different user needs and data characteristics, thereby improving the applicability of the report.
[0084] After generating a data analysis report based on the data to be analyzed and the results of the initial data analysis, the report-generating agent can send the report to the multimodal retrieval augmentation agent. The multimodal retrieval augmentation agent can also be used to output the data analysis report.
[0085] Optionally, the multimodal retrieval augmentation agent can present the received data analysis reports to users or other systems in an appropriate manner. For example, the output methods may include displaying them on an interface, sending them to a designated email address, or any one or more of these methods.
[0086] The embodiments of this application can enhance the unified output of data analysis reports through multimodal retrieval, thereby achieving standardization and centralized management of data analysis report output, simplifying the process for users to obtain reports, and improving user experience.
[0087] In one implementation, the vertical domain multimodal data analysis agent may include a vertical domain multimodal inference model, a vertical domain large language model, and the vertical domain multimodal data analysis agent itself. The vertical domain multimodal inference model is capable of understanding and processing data of multiple modalities and performing inference and analysis based on this data. The vertical domain large language model possesses powerful natural language processing capabilities, enabling it to understand, generate, and respond to text information relevant to the domain, allowing for fluent text interaction with users, answering user questions in the domain, or providing relevant information.
[0088] A vertical-domain multimodal data analysis agent can interact with the user at least once through a vertical-domain large language model when the input information is in a text modality. When either the first or second scenario exists, it interacts with the user through a vertical-domain multimodal inference model. The first scenario is when the input information modality includes non-textual modalities. The second scenario is when the data analysis report modality includes non-textual modalities.
[0089] Optionally, the text modality can be a form of information expression that exists in text form, such as articles, dialogues, instructions, and other textual content. The non-text modality can be other forms of information expression besides text, such as images (e.g., photos, charts), videos, or any one or more of these.
[0090] Optionally, the vertical domain multimodal data analysis agent can interact with the user through a vertical domain multimodal inference model when the modality of the input information includes non-textual modalities. The vertical domain multimodal data analysis agent can also interact with the user through a vertical domain multimodal inference model when the modality of the data analysis report includes non-textual modalities. Furthermore, the vertical domain multimodal data analysis agent can interact with the user through a vertical domain multimodal inference model when both the modality of the input information and the modality of the data analysis report include non-textual modalities.
[0091] Optionally, the interaction can be an information exchange and feedback process between the vertical-domain large language model or the vertical-domain multimodal reasoning model and the user. The vertical-domain large language model or the vertical-domain multimodal reasoning model responds based on the user's input information, and the user continues to input information based on the response to achieve interaction.
[0092] Optionally, the vertical domain large language model can be a large language model capable of in-depth analysis and processing of text data within a vertical domain, providing text interaction services. For example, the vertical domain large language model can be trained based on vertical domain data, building upon a large language model. The vertical domain can be a specific vertical field, such as emergency safety.
[0093] For example, Figure 2 This is a schematic diagram illustrating the training process of a large-scale language model in a vertical domain, as provided in an embodiment of this application. Figure 2 As shown, the training process of this vertical domain large language model is as follows:
[0094] Step one involves fine-tuning instructions on a vertical domain plain text dataset using a supervised fine-tuning (SFT) dataset. This step enables the large language model in this vertical domain to master the specialized terminology and expressions of the vertical domain, thus establishing a foundation of domain knowledge.
[0095] For example, an SFT dataset may be a set of multiple instructions and a set of paired responses to those instructions that conform to the vertical domain specification.
[0096] Step two involves fine-tuning the first segment of the model based on a vertical domain plain text preference dataset, using reinforcement learning based on human feedback. This step aligns the model's output with human preferences.
[0097] Optionally, the vertical domain plain text preference dataset can be obtained by reconstructing the vertical domain plain text dataset using existing techniques, and there are no restrictions on this.
[0098] Step three, building upon step two, adds a function calling dataset and performs fine-tuning using reinforcement learning based on human feedback in the second stage. This step enables the large-scale language model in this vertical domain to possess function calling capabilities.
[0099] Optionally, during the training process of step three of this vertical domain large language model, the ratio of Functioncalling data to vertical domain plain text data in each batch of training data can be dynamically adjusted using the cosine annealing algorithm. The ratio is 1:4 at the beginning of training, and increases to 1:1 when the training batch reaches halfway, and remains unchanged thereafter.
[0100] The following section provides a detailed introduction to the vertical multimodal inference model.
[0101] For example, an electronic device may acquire a vertical multimodal sample dataset, which includes: vertical multimodal sample cues, and vertical sample responses.
[0102] Optionally, the vertical multimodal sample cue and the vertical sample response can be an image-text pair, wherein the image in the image-text pair corresponds to the vertical sample response, and the text in the image-text pair corresponds to the vertical multimodal sample cue. The electronic device can directly obtain the image-text pair from an open-source database, or it can obtain the image-text pair by processing the image, table, or video to be processed.
[0103] Taking the processing of images, tables, and videos by electronic devices to obtain image-text pairs as an example, Figure 3 This is a flowchart illustrating a method for obtaining image-text pairs provided in an embodiment of this application, as shown below. Figure 3 As shown, the method for obtaining image-text pairs is as follows:
[0104] 1. Use the existing multimodal large model to process all the images, tables and videos to obtain the corresponding detailed text descriptions, that is, obtain image-text pairs.
[0105] 2. Use another existing multimodal large model to score the image-text pairs generated in step 1, and retain image-text pairs with a text description matching score higher than a threshold. For example, retain image-text pairs with a text description matching score higher than 8.
[0106] 3. Use existing large language models to analyze the text descriptions in all image-text pairs and filter out image-text pairs that do not match the target domain.
[0107] 4. Use the visual encoder in the existing multimodal large model as the feature extractor for image data, extract the representation vectors only for the image data in all image-text pairs, and store the representation vectors of these images in the vector database.
[0108] 5. Filter vectors in the vector database based on similarity. For example, only one image is kept for each representation vector with a similarity greater than 0.95, and 50% of the images with a similarity between 0.85 and 0.95 are randomly removed, resulting in high-quality, non-repeating image-text pairs for the target domain.
[0109] After acquiring a multimodal sample dataset from a specific vertical domain, the electronic device can train a pre-defined model to obtain a multimodal inference model for that domain. This pre-defined model may include a hidden layer, a text output module, and a multimodal feature vector output module. Figure 4 This is a schematic diagram of a preset model provided in an embodiment of this application, such as... Figure 4 As shown, the preset model may include image preprocessing, a visual encoder, such as a Vision Transformer (a model architecture for image processing), pixel unshuffle, a projector, a prompt, a text tokenizer, a Large Language Model (LLM), a language model head (LMHead), an embedding projector, an output text module, and a multi-vectors module.
[0110] The preset model may also include a hidden layer. Figure 4 Not shown in the diagram. For example, the hidden layer may include multiple convolutional layers, or the structure of the hidden layer may refer to any existing semantic feature extraction model. This hidden layer can be used for vertical multimodal sample cues (e.g., Figure 4The hidden layer extracts semantic features from the prompts in the sample (using the prompt words in the text) to obtain semantic features that characterize the prompts for multimodal samples in the vertical domain. This hidden layer can be used to extract semantic features based on the vertical sample response (e.g., ... Figure 4 Semantic features are extracted from images, videos, and tables to obtain semantic features that characterize the responses of vertical samples.
[0111] The text output module can be used to output predicted text describing the vertical multimodal sample prompts and responses based on the aforementioned semantic features. In other words, the text output module can rewrite the semantic features representing the vertical sample responses and the semantic features representing the vertical multimodal sample prompts into predicted text, thus realizing the rewriting of vertical sample responses and vertical multimodal sample prompts into predicted text. Optionally, this text output module can be a prediction module in any existing modality rewriting model for rewriting multimodal data into text modality data (used to predict text output based on features extracted from the hidden layer). This predicted text can be used for the next round of training of the preset model.
[0112] The multimodal feature vector output module is used to output predicted multimodal features (also known as multimodal embedding vectors) based on semantic features, representing "vertical multimodal sample cues and vertical sample responses." These predicted multimodal features can be used for training the next round of the pre-defined model. Optionally, the multimodal feature vector output module can refer to any existing multimodal feature extraction model capable of encoding semantic features to obtain multimodal features, such as the existing hierarchical image captioning (HI-Captioner) model based on the Swing Transformer (a visual model name).
[0113] For example, Figure 5 A schematic diagram illustrating the training process of a large-scale multimodal inference model in a vertical domain, provided for an embodiment of the application, is shown below. Figure 5 As shown, the training process of this large-scale multimodal inference model is as follows:
[0114] 1. Freeze the Embedding task on the vertical multimodal sample dataset and train the LM Head (language model head) only on the vertical multimodal dataset through single-turn dialogue.
[0115] 2. On the vertical multimodal sample dataset, freeze the Embedding task and train the LM Head only on the vertical multimodal dataset through multi-turn dialogue.
[0116] 3. Freeze the Embedding task on the vertical multimodal sample preference dataset and train the LM Head only on the vertical multimodal dataset using reinforcement learning from human feedback (RLHF).
[0117] 4. Train the LM Head using reinforcement learning with human feedback on the vertical multimodal sample preference dataset and function calling data.
[0118] Optionally, the vertical multimodal sample preference dataset is a reconstruction of the vertical multimodal sample dataset. This reconstruction method can refer to the existing reconstruction method of sample data in reinforcement learning training. For example, the ratio of Functioncalling data to vertical multimodal data can be dynamically varied using cosine, starting at 1:4 at the beginning of training, increasing to 1:1 when half of the training batches are completed, and then remaining unchanged thereafter.
[0119] In this embodiment, when the input is text-based, interaction can be achieved through a large language model. When the input or report contains non-textual modalities (such as images / videos), processing switches to a multimodal inference model. The multimodal inference model supports multimodal input and output, enhancing the processing capability for non-textual data, enabling the understanding and analysis of cross-modal information, and improving the adaptability and interactive experience of the multi-agent system in complex vertical domain scenarios.
[0120] It should be noted that the embodiments of this application can also switch models in response to user operations. When the multimodal inference model is switched to the vertical domain large language model, non-textual modal data can be placed using placeholders.
[0121] In one implementation, the vertical domain multimodal data analysis agent, before retrieving the data to be analyzed based on the retrieval reference data, is further configured to determine a target database from multiple candidate databases based on the retrieval reference data. Optionally, the candidate databases may include any one or more of relational databases, multimodal databases, and domain knowledge graph databases.
[0122] The following section provides a detailed explanation of how to construct a multimodal database and domain knowledge graph. Figure 6 This application provides a schematic diagram illustrating the construction of a multimodal database and a domain knowledge graph, as shown in the embodiments of this application. Figure 6As shown, images and videos can be used to obtain a multimodal database based on a multimodal embedding model and system prompts. This multimodal database can include multi-vectors (multimodal feature vectors) and text descriptions. Image-text documents and text documents can be used to construct a domain knowledge graph based on a rich text document parsing module and a multimodal embedding model.
[0123] Rich text documents can be document types that, in addition to text content, can also contain rich formatting and elements such as font styles, images, tables, charts, audio and video, and hyperlinks. The parsing method for rich text documents can be, for example, as described in the following embodiments.
[0124] Figure 7 This application provides a schematic diagram of a process for constructing a domain knowledge graph, as illustrated in the embodiments of this application. Figure 7 As shown, image and text documents and text documents can include Portable Document Format (PDF), Office (a document format), and plain text documents.
[0125] Electronic devices can use a pre-set PDF converter to convert all the aforementioned domain knowledge documents into PDF format, and then use an image converter to convert the PDF into an image. This image is in Joint Photographic Experts Group (JPG) format. A layout element extraction model can identify layout elements in the image, extract element numbers, and divide the image into blocks based on areas such as images, tables, body text, and titles. This layout element extraction model can, for example, refer to any existing model capable of recognizing text information from images.
[0126] Optionally, the layout element extraction model can automatically identify and extract various layout elements from the image converted from the rich text document. This model can employ existing deep learning models. For example, the layout elements may include one or more of the following: figures, tables, chapter names, text paragraphs, figure / table titles, figure / table captions, etc. The attribute information of the layout elements can be information describing their characteristics and properties; for example, the attribute information may include the layout element number.
[0127] Optionally, the layout element extraction model can first perform multiple batch processing operations on all the images converted from rich text documents until all the converted images are processed. Secondly, based on the recognition results of each page, the model can segment out the regions of all elements on all pages, obtaining images corresponding to all elements in all rich text documents and all pages. Thirdly, the model can divide the images corresponding to all elements into two main categories, corresponding to different system prompts: text elements and chart elements. Finally, the model can group elements by image size within each category, reducing size differences within the same batch and avoiding unnecessary computational increases caused by alignment issues.
[0128] The multimodal large model (i.e., the aforementioned multimodal embedding model) can extract text / knowledge from each element by batch processing text elements and graph / table elements. The multimodal large model (i.e., the aforementioned multimodal embedding model) can extract embedding vectors from sliced text, including recombining the recognition results of each element into a Markdown (a markup language) plain text document, and extracting vectors based on document slices. Optionally, the electronic device can input a preset prompt word template into the aforementioned multimodal large model to enable it to achieve the functions described above. Optionally, the electronic device can write the text information corresponding to all elements into the document in a top-to-bottom, left-to-right order, generating a text document corresponding to the original rich text document.
[0129] In summary, electronic devices can construct a domain knowledge graph from page element numbers, document slice extraction vectors, and document metadata. The electronic device can construct this domain knowledge graph, for example, by referring to the method described in the following embodiments:
[0130] In one implementation, the domain knowledge graph can include six entity types and six relation types, as shown below:
[0131] Domain (Emergency Safety Sub-sectors): Gas, Water, Petrochemical, Transportation, etc.
[0132] DocCategory (Document Type): Enterprise production documents, academic literature, etc.
[0133] Document: papers, regulations, management rules, etc.
[0134] Page: page1, page2, page3, etc.
[0135] ContentElement (elements): Image elements, table elements, heading elements, etc.
[0136] KnowledgeChunk: Document Chunk 1, Document Chunk 2, Document Chunk 3, etc.
[0137] The page entity contains four attributes: original document identifier (ID), number of images, number of tables, and page overview. The element entity contains five attributes: original document ID, original page ID, save path, text information, and representation vector. The knowledge fragment entity corresponds to each document slice, and each document slice corresponds to one or more element entities. The knowledge fragment entity contains three attributes: representation vector, text information, and number of elements.
[0138] The relation types are as follows:
[0139] BELONGS_TO (belongs to): describes which category a document belongs to. For example, "Gas Management Regulations" BELONGS_TO "Management Regulations" indicates that "Gas Management Regulations" belongs to the Management Regulations category.
[0140] IN_DOMAIN (Domain): Describes which sub-domain of emergency safety a document type belongs to, for example, "Management Regulations" IN_DOMAIN "Gas".
[0141] HAS_PAGE (Containing Page): Describes which pages a document includes, for example, "Report" HAS_PAGE Page 1.
[0142] LOCATED_IN (located in): Describes which page an element is located on, for example, Table 3 LOCATED_IN Page 5.
[0143] COMPRISED_OF (Composition): Describes the elements contained in a knowledge fragment, such as knowledge fragment ACOMPRISED_OF Table 3 + Text Paragraph 5.
[0144] CITES (Citation): Describes which documents the current document references, such as an academic paper referencing another document in CITES.
[0145] This knowledge graph can trace the source of each knowledge fragment retrieved to the specific document, page, and location of that fragment, facilitating easy querying and confirmation.
[0146] It should be noted that this multimodal embedding model can be trained on rich text documents, primarily PDFs, using YOLOv8 (a deep learning model architecture) as training data. This results in a rich text document element extraction model that extracts various elements from PDF documents, such as figures, tables, chapter titles, text paragraphs, figure titles, and figure annotations. Compared to the existing Retrieval-Augmented Generation (RAG) systems, which suffer from inaccurate processing of chart information in rich text documents, the embodiments of this application can improve the accuracy of extracting domain knowledge from charts.
[0147] Optionally, the multimodal large model can obtain corresponding text descriptions and representation vectors based on data to be processed, such as images, tables, and videos. For example, the multimodal large model can generate text descriptions of the data to be processed using the output head of a language model, and can also extract features from the data to be processed using a visual encoder to obtain the representation vectors of the data to be processed.
[0148] For example, a 10-second clip of an anomaly or sudden event can be extracted at 1-3 frames per second. Multiple images of the resulting anomaly or sudden event are then input into a multimodal large-scale model. This model generates a text description and a representation vector. The text description and representation vector are then associated and stored to create a multimodal database.
[0149] The vertical multimodal data analysis agent can respond to the target database being a relational database by generating database query statements based on retrieved reference data through a code generation model, and then querying the relational database based on the database query statements to obtain the data to be analyzed.
[0150] Optionally, the code generation model can be based on a deep learning model to convert retrieved reference data into database query statements. Figure 8 As shown in Figure 8, which illustrates a relational database query service provided in this application embodiment, the input text can be retrieval reference data. The vertical multimodal data analysis agent can generate a database query statement, such as SQL, based on a Code Large Language Model (Code LLM) of the retrieval reference data. The relational database is then queried through the SQL database query interface to obtain the data to be analyzed.
[0151] This application embodiment achieves efficient retrieval of various types of databases by dynamically selecting databases and automatically generating query statements, thereby improving the accuracy and efficiency of data acquisition and reducing the cost and error rate of manually writing query statements.
[0152] The above describes the detailed content of how a multimodal retrieval enhancement agent can respond to a target agent as a report generating agent. In one embodiment, the multimodal retrieval enhancement agent can also respond to a target agent as a vertical domain multimodal data analysis agent by sending input information to the vertical domain multimodal data analysis agent.
[0153] The vertical domain multimodal data analysis agent can be used to perform retrieval based on input information, obtain retrieval results, and conduct data analysis based on the retrieval results to obtain a second data analysis result. The second data analysis result is then sent to the multimodal retrieval enhancement agent.
[0154] Optionally, the second data analysis result can be the final analysis result obtained by the vertical domain multimodal data analysis agent after performing data analysis based on the retrieval results, which can reflect key features, trends, relationships, and other information in the retrieval results. Optionally, the vertical domain multimodal data analysis agent can include any existing multimodal data analysis model. Alternatively, the vertical domain multimodal data analysis agent can also be a vertical domain multimodal data analysis model trained on an existing multimodal large language model using vertical domain data.
[0155] The embodiments of this application can support data retrieval and analysis directly through a vertical domain multimodal data analysis intelligent agent, simplifying the process, improving the real-time performance and flexibility of data processing, and are suitable for vertical domain scenarios that require rapid response.
[0156] This application also provides a report generation method. This method is applied to multi-agent systems as described in any of the foregoing embodiments. Figure 9 This is a flowchart illustrating a report generation method provided in an embodiment of this application, such as... Figure 9 As shown, the method includes:
[0157] S101. Enhance the agent through multimodal retrieval, determine the target agent based on the user's input information; in response to the target agent being the report generating agent, send the input information to the report generating agent.
[0158] S102. The report generation agent determines the retrieval reference data and data analysis task based on the input information; and sends the retrieval reference data and data analysis task to the vertical domain multimodal data analysis agent.
[0159] S103. The vertical domain multimodal data analysis agent retrieves the data to be analyzed based on the reference data; it performs data analysis on the data to be analyzed based on the data analysis task to obtain the first data analysis result; and it feeds back the data to be analyzed and the first data analysis result to the report generation agent.
[0160] S104. Generate an intelligent agent through the report, and generate a data analysis report based on the data to be analyzed and the results of the first data analysis.
[0161] The specific implementation of the report generation method provided in this application can be referred to the multi-agent system for the vertical domain scenario described in the foregoing embodiments, and will not be repeated here.
[0162] The multi-agent system provided in this application can determine the target agent and forward input information through a multimodal retrieval augmentation agent, determine the retrieval reference data and tasks through a report generation agent and send them to the vertical domain multimodal data analysis agent, which then retrieves, analyzes, and provides feedback on the data. The report generation agent can also generate a data analysis report. This multi-agent system, through a collaborative agent architecture, achieves fully automated processing from user input to data analysis report generation, improving data processing efficiency and report generation accuracy, and is suitable for analyzing complex multimodal data in vertical domain scenarios.
[0163] Figure 10 This application provides an example architecture diagram of a multi-agent system in a vertical domain scenario, as shown in the embodiments below. Figure 10 As shown, the multi-agent system includes a multimodal agent retrieval-augmented generation (Agentic RAG), which is the multimodal retrieval-augmented agent in the above embodiment; a security report generation agent, which is the report generation agent in the above embodiment; and a multimodal data analysis agent, which is the vertical domain multimodal data analysis agent in the above embodiment.
[0164] The multimodal agentic RAG can include a large-scale language model for a specific domain, a large-scale multimodal reasoning model for a specific domain, and contextual information that can share user input. For example, user input can be an image and / or dialogue content. The dialogue content could be something like, "Are there any security risks this week?".
[0165] Multimodal data analysis agents can perform multimodal retrieval services based on the Model Context Protocol server (MCP server). They can also perform relational data query services based on the MCP server. During data analysis, the multimodal data analysis agent can invoke various external tools to provide professional algorithm services based on the MCP server. For example, this may include specialized mechanism algorithms 1, traditional computer vision (CV) algorithms 2, and traditional machine learning algorithms 3.
[0166] Based on the above embodiments, Figure 11 This is a schematic diagram of the architecture of a retrieval model provided in an embodiment of this application, such as... Figure 11 As shown, the retrieval model can respond to user input text and / or input images. The input text is rewritten into multiple queries by the Query (query) and intent recognition model, such as query1 and query2.
[0167] The rewritten input text and / or the input image are then input into the multimodal embedding model, the architecture of which can be referenced above. Figure 4 It can include a text tokenizer and ViT (a model architecture for image processing). Figure 12 A schematic diagram illustrating the training process of the multimodal embedding model provided in this application embodiment, as shown below. Figure 12 As shown, the training process for a multimodal embedding model can be achieved through the following steps:
[0168] 1. Freeze the Embedding task and train the LM Head (language model head) only on the vertical multimodal dataset. It should be noted that the vertical multimodal dataset can be a dataset of vertical multimodal samples.
[0169] 2. Freeze the text generation task and train the Embedding task only on the vertical multimodal dataset.
[0170] 3. For the same batch of training data, there is a 50% chance of randomly starting either the generation or embedding task, and a 50% chance of performing both tasks simultaneously. This training data may include a vertical-domain multimodal dataset and an Optical Character Recognition (OCR) dataset. The vertical-domain multimodal dataset can account for up to 70%, and the OCR dataset can account for up to 30%.
[0171] 4. All batches of training data are used simultaneously for two tasks, but only one batch (epoch) is executed in this stage. This training data is consistent with the training data in step 3.
[0172] The multimodal embedding model can provide cross-modal retrieval services based on the Model Context Protocol (MCP), including: basic semantic retrieval, cross-modal semantic retrieval, sparse retrieval / keyword retrieval, and hybrid retrieval. This cross-modal retrieval service can perform searches based on multimodal databases and domain knowledge graphs to ultimately obtain search results.
[0173] The above are the method embodiments provided in this application. The apparatus provided in this application will be described below.
[0174] Figure 13 A schematic diagram of a report generation device provided in this application is shown below. Figure 13 As shown, the report generation device 400 is applied to a multi-agent system, which includes: a multimodal retrieval enhancement agent, a report generation agent, and a vertical domain multimodal data analysis agent. The report generation device 400 provided in this embodiment includes:
[0175] The first determining module 401 is used to enhance the agent through multimodal retrieval, determine the target agent based on the user's input information, and send the input information to the report generating agent in response to the target agent being the report generating agent.
[0176] The second determining module 402 is used to generate an intelligent agent through a report, determine the retrieval reference data and data analysis task based on the input information, and send the retrieval reference data and data analysis task to the vertical domain multimodal data analysis intelligent agent.
[0177] The analysis module 403 is used by a vertical domain multimodal data analysis agent to retrieve data to be analyzed based on reference data; based on the data analysis task, it performs data analysis on the data to be analyzed to obtain the first data analysis result. The data to be analyzed and the first data analysis result are then fed back to the report generation agent.
[0178] The generation module 404 is used to generate an intelligent agent through a report, which generates a data analysis report based on the data to be analyzed and the results of the first data analysis.
[0179] The report generation device 400 provided in this embodiment can execute the methods provided in any of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0180] Figure 14 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 14As shown, the electronic device 500 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 500 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus 504.
[0181] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.
[0182] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0183] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0184] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0185] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0186] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0187] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0188] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0189] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0190] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0191] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0192] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0193] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0194] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0195] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A multi-agent system for a vertical domain scenario, characterized in that, The multi-agent system includes: a multimodal retrieval enhancement agent, a report generation agent, and a vertical domain multimodal data analysis agent; The multimodal retrieval enhancement agent is used to determine a target agent based on user input information; in response to the target agent being the report generating agent, the input information is sent to the report generating agent. The report generating agent is used to determine retrieval reference data and data analysis tasks based on the input information; and to send the retrieval reference data and the data analysis tasks to the vertical domain multimodal data analysis agent; the retrieval reference data is used to guide the vertical domain multimodal data analysis agent to perform data retrieval. The vertical domain multimodal data analysis agent is used to retrieve data to be analyzed from multiple candidate databases based on the retrieval reference data; perform data analysis on the data to be analyzed based on the data analysis task to obtain a first data analysis result; and feed back the data to be analyzed and the first data analysis result to the report generation agent; the multiple candidate databases include at least one of relational databases, multimodal databases, and domain knowledge graph databases; The report generating agent is used to generate a data analysis report based on the data to be analyzed and the results of the first data analysis.
2. The multi-agent system according to claim 1, characterized in that, The report-generating agent is also used for: Before generating a data analysis report based on the data to be analyzed and the first data analysis result, a target report template is determined based on at least one of the input information, the retrieved reference data, and the first data analysis result; The report generation agent is specifically used to fill the target report template with the data to be analyzed and the first data analysis results to obtain the data analysis report.
3. The multi-agent system according to claim 2, characterized in that, The vertical domain multimodal data analysis agent includes: a vertical domain multimodal inference model and a vertical domain large language model. Specifically, the vertical domain multimodal data analysis agent is used for: When the input information is in text mode, at least one round of interaction is conducted with the user through the vertical domain large language model; When either the first or the second scenario exists, the user is interacted with through the vertical multimodal reasoning model. The first scenario is that the modality of the input information includes a non-textual modality; the second scenario is that the modality of the data analysis report includes a non-textual modality.
4. The multi-agent system according to any one of claims 1-3, characterized in that, The report-generating agent is also used for: After generating a data analysis report based on the data to be analyzed and the first data analysis result, the data analysis report is sent to the multimodal retrieval enhancement agent; The multimodal retrieval enhancement agent is also used to output the data analysis report.
5. The multi-agent system according to any one of claims 1-3, characterized in that, The vertical multimodal data analysis intelligent agent is also used for: Before performing the search based on the search reference data to obtain the data to be analyzed, the target database is determined from multiple candidate databases based on the search reference data; In response to the fact that the target database is a relational database, a database query statement is generated based on the retrieval reference data using a code generation model; The vertical domain multimodal data analysis agent is specifically used to query the relational database based on the database query statement to obtain the data to be analyzed.
6. The multi-agent system according to any one of claims 1-3, characterized in that, The multimodal retrieval enhancement agent is also used for: In response to the target agent being the vertical multimodal data analysis agent, the input information is sent to the vertical multimodal data analysis agent; The vertical multimodal data analysis intelligent agent is also used for: Based on the modality of the input information, and the input information, a search is performed to obtain search results, and based on the search results, data analysis is performed to obtain a second data analysis result; The second data analysis result is sent to the multimodal retrieval enhancement agent.
7. A report generation method, characterized in that, The method is applied to a multi-agent system, which includes: a multimodal retrieval enhancement agent, a report generation agent, and a vertical domain multimodal data analysis agent; the method includes: The multimodal retrieval enhancement agent determines a target agent based on user input information; in response to the target agent being the report generating agent, the input information is sent to the report generating agent. The report generates an intelligent agent that, based on the input information, determines retrieval reference data and data analysis tasks; the retrieval reference data and the data analysis tasks are then sent to the vertical domain multimodal data analysis intelligent agent; the retrieval reference data is used to guide the vertical domain multimodal data analysis intelligent agent in data retrieval. The vertical domain multimodal data analysis agent retrieves data to be analyzed from multiple candidate databases based on the retrieval reference data; performs data analysis on the data to be analyzed based on the data analysis task to obtain a first data analysis result; and feeds back the data to be analyzed and the first data analysis result to the report generation agent; the multiple candidate databases include at least one of relational databases, multimodal databases, and domain knowledge graph databases. The report generates an intelligent agent that generates a data analysis report based on the data to be analyzed and the results of the first data analysis.
8. A report generation device, characterized in that, The device is applied to a multi-agent system, which includes: a multimodal retrieval enhancement agent, a report generation agent, and a vertical domain multimodal data analysis agent; the device includes: The first determining module is configured to determine a target agent based on user input information through the multimodal retrieval augmentation agent; and in response to the target agent being the report generating agent, send the input information to the report generating agent. The second determining module is used to generate an intelligent agent through the report, determine retrieval reference data and data analysis tasks based on the input information, and send the retrieval reference data and the data analysis tasks to the vertical domain multimodal data analysis intelligent agent; the retrieval reference data is used to guide the vertical domain multimodal data analysis intelligent agent to perform data retrieval; The analysis module is used to retrieve data to be analyzed from multiple candidate databases based on the retrieval reference data through the vertical multimodal data analysis agent; perform data analysis on the data to be analyzed based on the data analysis task to obtain a first data analysis result; and feed back the data to be analyzed and the first data analysis result to the report generation agent; the multiple candidate databases include at least one of the following: relational databases, multimodal databases, and domain knowledge graph databases; The generation module is used to generate an intelligent agent through the report, and generate a data analysis report based on the data to be analyzed and the first data analysis results.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in claim 7.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in claim 7.
Citation Information
Patent Citations
System and method for constructing intelligent report robot based on multi-dimensional data
CN111159353A
Complex information retrieval system and method
CN118779364A