Database interaction method and device based on multi-agent cooperation, equipment and medium

By employing a multi-agent collaborative architecture and a multi-verification mechanism, the shortcomings of existing database interaction technologies in multimodal data processing and task planning are addressed, enabling efficient and accurate database interaction and data interpretation, thereby improving the system's intelligence level and user experience.

CN120821738AActive Publication Date: 2025-10-21XIAMEN UNIV OF TECH +2

Patent Information

Application Number
CN202511317091.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-21
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing database interaction technologies have limited preprocessing capabilities when processing multimodal inputs, making it difficult to fully capture the deep semantic features of the data, resulting in information loss or omissions, low query efficiency and feedback accuracy, lack of flexibility in task planning, and inability to effectively integrate user needs and multi-source retrieval results. The generated query instructions have poor accuracy and execution success rate.

Method used

It adopts a multi-agent collaborative architecture, performs semantic segmentation and vectorization processing through a format adapter, combines intelligent search and knowledge base retrieval, decomposes complex tasks into atomic subtasks, and uses a dual-agent collaboration and three-level matching mechanism for database interaction, and introduces a multi-verification mechanism to ensure the accuracy of SQL query statements.

Benefits of technology

It improves the parallel response capability of multimodal data processing, enhances the accuracy of semantic similarity retrieval and the comprehensiveness of information recall, ensures the robustness and efficient execution of the system in complex scenarios, and provides an intuitive, accurate and efficient data interpretation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821738A_ABST
    Figure CN120821738A_ABST
Patent Text Reader

Abstract

The invention discloses a database interaction method, device and equipment based on multi-agent cooperation and a medium, and relates to the technical field of database interaction. The method comprises the steps of obtaining a query request, performing semantic partitioning and vectorization, and obtaining a retrieval vector set; and selecting a search source and searching to obtain an information block. And obtaining candidate knowledge through knowledge base retrieval. And integrating the information blocks and the candidate knowledge to obtain a candidate information set. And constructing a dynamic demand understanding model according to the query request and the candidate information set, and obtaining a task state matrix and a tool calling task. And calling a corresponding auxiliary tool and operating, updating the task state matrix according to an operation result and returning the task state matrix to the task planning agent to update the tool calling task until the tool calling task is no longer output, and obtaining a task result. And when the database interaction module is called, generating a chart according to an output result of the database interaction module, and outputting the chart and a task result together. Otherwise, directly outputting a task result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database interaction technology, and in particular to a database interaction method, apparatus, device and medium based on multi-agent collaboration. Background Art

[0002] The fields of computer science and artificial intelligence have a significant demand for efficient and intelligent database interaction systems. In real-world applications, users often need to process multimodal data input, including text, structured data, and multimedia files, to support complex queries and decision analysis.

[0003] For example, in scenarios like intelligent customer service and data analysis platforms, systems must quickly respond to diverse query requests while ensuring information integrity and semantic depth. These requirements require database interaction technology to possess efficient data preprocessing capabilities, precise semantic parsing mechanisms, and dynamic task scheduling strategies to address the growing challenges of processing multi-source, heterogeneous data.

[0004] Existing database interaction technologies have obvious deficiencies in multiple key areas. When processing multimodal inputs, the system has limited pre-processing capabilities, which often leads to information loss or omissions, making it difficult to fully capture the deep semantic features of the data, thereby affecting the accuracy of subsequent vectorization and similarity retrieval. In the information retrieval process, existing methods over-rely on simple keyword matching and fail to effectively utilize high-quality embedding conversion and semantic segmentation technology, resulting in retrieval results that are difficult to accurately respond to users' actual needs, and query efficiency and feedback accuracy are severely limited. In addition, there is a lack of flexibility in task planning, and it is impossible to effectively integrate user needs, historical interaction records, and multi-source retrieval results. Complex query tasks are difficult to break down into independently executable sub-tasks, and the tool calling process appears to be less intelligent and adaptive, significantly reducing overall execution efficiency and task completion.

[0005] These flaws are further reflected in the database interaction process. Existing technologies, from parsing input data to generating SQL queries, lack proper matching of tables, columns, and specific constraints. Furthermore, they lack multiple validation mechanisms covering syntax, semantics, and results. This results in poor accuracy and execution success rates for the resulting queries. These issues not only limit the system's robustness in diverse data scenarios, but also hinder user experience and data insight, hindering the development of intelligent database management technology. Summary of the Invention

[0006] The present invention provides a database interaction method, apparatus, device and medium based on multi-agent collaboration to improve at least one of the above technical problems.

[0007] In a first aspect, the present invention provides a database interaction method based on multi-agent collaboration, which includes steps S1 to S5.

[0008] S1. Obtain the query request, then perform semantic segmentation through the format adapter, and then vectorize it into a vector representation to obtain the retrieval vector set.

[0009] S2. Based on the search vector set, the search source selection agent selects a search source and uses a meta-search engine to search to obtain information blocks, and then searches the knowledge base to obtain candidate knowledge. The information blocks and candidate knowledge are then integrated to obtain a candidate information set.

[0010] S3. A task planning agent is used to construct a dynamic demand understanding model based on the query request and the candidate information set to generate atomically executable subtasks and obtain a task status matrix and a tool calling task.

[0011] S4. Call the corresponding auxiliary tool according to the tool calling task and perform the operation, update the task status matrix according to the operation result, and return the updated task status matrix to the task planning agent to update the tool calling task until the task planning agent no longer outputs the tool calling task and obtains the task result.

[0012] S5. When the task planning agent calls the database interaction module, a chart is generated based on the output of the database interaction module and output together with the task result. Otherwise, the task result is directly output.

[0013] In a second aspect, the present invention provides a database interaction device based on multi-agent collaboration, which includes a query acquisition module, a search source module, a task planning module, a task execution module and an output module.

[0014] The query acquisition module is used to obtain the query request, then perform semantic segmentation through the format adapter, and then vectorize it into a vector representation to obtain the retrieval vector set.

[0015] The search source module is configured to select a search source through a search source selection agent based on the retrieval vector set, search using a meta-search engine, obtain information blocks, and obtain candidate knowledge through a knowledge base search. The information blocks and candidate knowledge are then integrated to obtain a candidate information set.

[0016] The task planning module is used to build a dynamic demand understanding model based on the query request and the candidate information set through a task planning agent to generate subtasks that can be executed atomically, obtain a task status matrix and a tool call task.

[0017] The task execution module is used to call the corresponding auxiliary tool and perform operations according to the tool call task, update the task status matrix according to the operation results, and return the updated task status matrix to the task planning agent to update the tool call task until the task planning agent no longer outputs the tool call task and obtains the task result.

[0018] The output module is used to generate a chart based on the output results of the database interaction module when the task planning agent calls the database interaction module, and output it together with the task result. Otherwise, the task result is directly output.

[0019] In a third aspect, the present invention provides a database interaction device based on multi-agent collaboration, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the database interaction method based on multi-agent collaboration as described in any paragraph of the first aspect.

[0020] In a fourth aspect, the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a database interaction method based on multi-agent collaboration as described in any paragraph of the first aspect.

[0021] By adopting the above technical solution, the present invention can achieve the following technical effects: Compared with existing technologies, this invention deeply optimizes key aspects such as multimodal data processing, intelligent retrieval, and task planning, and introduces a series of innovative mechanisms to enhance the system's intelligence, retrieval accuracy, and interactive experience. Specifically, it includes the following aspects: 1. Adopt multimodal data preprocessing and format adapter parsing technology to separate real-time query and knowledge base construction, improve the system's parallel response capability, and reduce the complexity of processing data in multiple formats.

[0022] 2. By integrating multimodal embedding transformation and semantic segmentation processing, the deep semantic features of diverse information such as text and images are captured to achieve more precise semantic similarity retrieval and obtain higher matching accuracy.

[0023] 3. Integrate intelligent search and knowledge base retrieval processes in parallel. Through automatic keyword generation, cosine similarity matching and re-ranking mechanisms, information blocks that are highly consistent with user queries can be extracted more quickly in large-scale searches, thereby improving the comprehensiveness and accuracy of information recall.

[0024] 4. Use the task planning agent to perform multi-dimensional integration of user queries, historical interactions, preset preferences and search results, and decompose complex tasks into atomic sub-tasks. By implementing refined task scheduling and dynamic tool calls, the system can ensure robustness and efficient execution in complex scenarios.

[0025] 5. In the database interaction link, dual-agent collaboration and a three-level matching mechanism are adopted to perform precise matching from tables, columns to specific value types, and supplemented by a closed-loop error correction mechanism with triple verification (syntax, semantics, and result levels) to ensure that the generated SQL query statements comply with ANSI standards and have a high execution success rate.

[0026] 6. In the visualization stage, by combining intelligent decision trees with hierarchical summary architecture, we can match users' explicit needs, implicit preferences, and data features to achieve automated visualization selection and semantic alignment, thereby providing an intuitive, accurate, and efficient data interpretation experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is an architectural diagram of the database interaction method.

[0028] Figure 2 It is a framework diagram of the database interaction method.

[0029] Figure 3 It is a flowchart of the database interaction method.

[0030] Figure 4 It is a visual example of the query results returned by the database interaction module. DETAILED DESCRIPTION

[0031] Example 1, please refer to Figures 1 to 4 The present invention provides a database interaction method based on multi-agent collaboration, which can be performed by a database interaction device based on multi-agent collaboration (hereinafter referred to as the database interaction device). In particular, one or more processors in the database interaction device are used to implement steps S0 to S5.

[0032] S0. Pre-built Database and Knowledge Base: Based on the structured data uploaded by users, natural language structure information at the table, column, and value levels is extracted, vectorized, and stored in a vector database. Heterogeneous files uploaded by users are converted to a unified processing format, and data is segmented and vectorized to obtain a searchable content mapping relationship library. Heterogeneous files are unstructured or semi-structured data.

[0033] It should be noted that the knowledge base and database can directly use existing databases without having to build them themselves. Therefore, the construction of the knowledge base and database is not an essential technical feature of the present invention. In this embodiment, the database is constructed based on structured data. The knowledge base is constructed based on unstructured or semi-structured data.

[0034] Preferably, step S0 includes steps S01 to S06.

[0035] S01. During the database pre-build process, the user's uploaded database file or database configuration information, as well as the database structure information file, are first obtained. Natural language structure information at the table, column, and value levels is extracted and generated. After extracting the natural language structure information, it is vectorized (embedding) and ultimately stored in a vector database for semantic search. The database file can be SQLite or a user-configured MySQL or PostgreSQL database. The following describes the logic for generating structure information at the table, column, and value levels.

[0036] The table-level description for each table is defined as: .

[0037] Where: Indicates table-level description, For the table name, express Description of usage, Indicates the primary key field Description, Indicates the Description of each field (field name, type, meaning), is the number of fields.

[0038] The column-level description for any column in a table is defined as: .

[0039] Where: Indicates column-level description, For field names, For the table name, )for Data type, Flag for the key type (primary key or foreign key, or omitted if none). A description of the meaning of the field.

[0040] For a specific value, its natural language description is: .

[0041] Where: For value level description, For the value of a specific field, For field names, For the table name, A description of the foreign key field correspondence (if any), A description of the meaning of the field.

[0042] S02. Perform vectorization (embedding) operation on the generated structural information.

[0043] The natural language description text generated above Mapping to high-dimensional vectors .

[0044] .

[0045] Among them, the function Usually it is OpenAI's embedding model (such as text-embedding-ada-002), which generally produces 768 or 1536-dimensional vectors. For example: , , Where, for vector, for vector, for vector.

[0046] S03. Store the vectorization (embedding) results in a vector database.

[0047] Each vector database record is: Where: Records in the database, is the unique identifier of the vector, Describe the text in the original natural language, The generated high-dimensional vector, is the record type (table, column or value), The name of the table to which it belongs, The column name to which it belongs (required only at the column level and value level), is the actual value (value level only).

[0048] When users or subsequent tools need to search using natural language, they can use cosine similarity to search in the vector space. This method finds the most similar natural language description record and quickly locates the corresponding table, column, or specific value in the database.

[0049] S04. During the knowledge base construction process, the heterogeneous files uploaded by the user are first parsed through the format adapter to convert data in different formats into a unified processing format.

[0050] Specifically, unified processing of heterogeneous files is achieved by relying on existing format parsing and adaptation technologies. The existing technologies used include multi-format file parsing libraries (such as Apache Tika for parsing text files such as PDF and DOCX, Pillow for processing image formats, and PyPDF2 for PDF text extraction), data format conversion frameworks (such as Pandas for processing semi-structured data such as CSV / Excel), and multimedia parsing tools (such as FFmpeg for extracting audio metadata).

[0051] S05. Chunking the parsed data to ensure appropriate granularity during knowledge base retrieval. In practical applications, the system typically uses simple paragraph-based chunking rules, such as using two line breaks (\n\n) as delimiters to divide the text into several semantic chunks.

[0052] S06. Vectorize and encode these knowledge blocks to obtain a searchable content mapping relationship library. Specifically, each block Encoding functions by vectorization Convert to a high-dimensional vector , and finally form a searchable content mapping relationship library. Indicates the blocks, express vector.

[0053] S1. Query request preprocessing: Obtain the query request, then perform semantic segmentation through a format adapter, and then vectorize it into a vector representation to obtain a retrieval vector set.

[0054] In addition to text content, the query request can also contain content in various formats, such as text, structured data, and multimedia files. In this embodiment, semantic segmentation through a format adapter is the same as S04 to S05 in S0. Specifically, when semantic segmentation is performed through a format adapter: for text type information, no additional preprocessing is performed. For structured data, additional parsing is required to extract key fields and convert them into a unified format to align with other data types. Audio data is directly converted into text to simplify the subsequent processing flow so that it can be consistent with other text data. Image data is stored in its original format to ensure that it can be matched and associated with other information during subsequent retrieval.

[0055] After semantic segmentation using a format adapter, a search term extraction agent is invoked to extract search keywords from the semantic segments. These search keywords are then vectorized to obtain a set of search vectors. This search term extraction agent is fine-tuned using prompts from an existing large language model (e.g., as a professional information retrieval assistant, you are responsible for extracting the most effective query phrases from a given text to find relevant information in a knowledge base system or search engine). The search keywords are converted into high-dimensional vector representations using vectorization technology, which are then used for similarity matching of the search information segments.

[0056] For example, a user submits a query: "I'd like to travel during the May Day holiday. Please recommend a city that's best suited for short-term travel, taking into account the weather conditions over the past week, the cost of living in the city, and the cultural atmosphere. Finally, output the conclusion in Chinese, along with a brief English summary." The search term extraction agent phrases the key information and generates a list of search terms. The extracted search terms are as follows: ["Weather conditions over the past week", "Cost of living in the city", "Cultural atmosphere in the city", cities suitable for short-term travel"]. After extracting the corresponding phrase information, the system vectorizes each search term for similarity matching in subsequent search information blocks.

[0057] S2. Information Retrieval: Based on the retrieval vector set, a search source selection agent selects a search source and uses a meta-search engine to search for information blocks. Candidate knowledge is retrieved through a knowledge base search. The information blocks and candidate knowledge are then integrated to obtain a candidate information set.

[0058] The core goal of the system is to extract the information blocks that are most relevant to the user query to support subsequent intelligent task planning and reasoning generation. Preferably, the information retrieval of the embodiment of the present invention adopts a dual retrieval strategy, which includes an intelligent search process and a knowledge base retrieval process. On the one hand, the intelligent search process is used to automatically parse the user query and generate retrieval keywords, and the candidate information blocks are preliminarily screened out through the meta search engine. On the other hand, the knowledge base retrieval technology is used to perform cosine similarity matching on the pre-processed user query embedding and the pre-vectorized document data in the knowledge base. After matching, the system will re-sort the retrieved information blocks and screen out the Top K candidate sets with the highest semantic relevance. Ultimately, the results of the two retrieval strategies will be integrated into a unified candidate information set and passed to the task planning agent as background knowledge support. Specifically, step S2 includes steps S21 to S27.

[0059] S21. Based on the retrieval vector, a search source selection agent is used to select a search source. Specifically, the system passes the retrieval vector to the search source selection agent, which selects an appropriate retrieval source based on its semantic information. The search source selection agent is fine-tuned using prompt words (e.g., "You are a task classifier. Your task is to analyze the user's query, determine whether a search operation is required, and, if so, select the appropriate search type") on an existing large language model.

[0060] S22: Based on the search vector and the selected search source, a meta-search engine is used to search and obtain information blocks related to the user's needs. This step is a preliminary screening.

[0061] S23: Vectorize the search results to convert the text information into vector representation. Specifically, vectorize the information blocks related to the user's needs.

[0062] S24. Compare the vector representation of the search results with the search vector using a semantic similarity algorithm, and select the top-K information blocks with the highest semantic relevance, thereby ensuring that the returned content is highly consistent with the user query at the semantic level. The semantic similarity algorithm may be an existing algorithm such as cosine similarity.

[0063] Specifically, before the actual search begins, each search phrase is fed into a context-sensitive "search source selection agent." This agent's goal is to determine whether the query requires a search and further determine the appropriate search type. The system offers four search types: academic research, social community, natural science, and general internet. For example, for "city living costs," the classifier model outputs the general internet category. Regardless of the search type, the result format remains consistent: {title, content, link, source}. The system takes the search source and search term as input and invokes the corresponding search engine, returning the following example result: JSON text [{Title: 2025 Cost of Living Ranking of Major Chinese Cities. Content: According to the latest 2025 data, Beijing, Shanghai, and Shenzhen still rank highest in terms of living costs, while cities like City A and Xi'an are favored for their lower costs and good living conditions. Link: https: / / . Source: Data Source Example.}].

[0064] As shown above, search results typically include a summary, title, link, source, and publication date. The system converts these structured results into JSON text and performs embedded vectorization. These vectors are then compared for similarity with the semantic vectors of the original search terms, extracting the top-K most relevant results as the output information chunks for intelligent search. Cosine similarity is typically used for similarity calculations.

[0065] S26: Perform similarity matching between the search vector set and the pre-built knowledge base, identify the most similar document segments, and obtain candidate knowledge. The knowledge base has document data pre-stored in blocks.

[0066] Specifically, the entire knowledge base content is first vector-embedded, and then phrase vectors are extracted and compared one by one. The system then uses a semantic matching mechanism to identify the document paragraphs or knowledge nodes that are most similar to the phrase, and selects the candidate knowledge with the highest similarity score as the knowledge base retrieval output.

[0067] Example output for the search phrase "cities suitable for short-term travel": json text [{Title: Recommended five domestic destinations suitable for short-distance travel. Content: For short-term trips of three to five days, City F, City C, City B, City A, and City D are widely recommended. These cities not only have convenient transportation but also offer a balanced combination of cultural experience and consumption, making them suitable for holiday relaxation. Source: "2024 Urban Culture and Tourism White Paper." File number: .... Chapter identifier: ...}]. This information block comes from a document in the knowledge base. Through semantic matching, the system determines that it is highly relevant to "cities suitable for short-term travel" and returns it as a knowledge base search result.

[0068] Ultimately, the candidate knowledge will be used together with the information block as the semantic context support input for the city recommendation generation task. Each retrieval vector performs an intelligent search and a knowledge base match, so the entire process requires multiple information extractions.

[0069] S27. Incorporate the information blocks and candidate knowledge into the candidate information set.

[0070] The candidate information set serves as background knowledge for subsequent task planning agents, assisting the model in providing more accurate, detailed, and reliable supporting information during reasoning and generation. By integrating intelligent search and knowledge base retrieval strategies in parallel, the system can mine and extract potential key information on a larger scale, thereby improving the overall accuracy and efficiency of the system. This process ensures that the system not only retrieves content that matches keywords but also accurately understands user needs at a semantic level, thereby improving the accuracy and reliability of overall intelligent decision-making.

[0071] S3. Task planning: A task planning agent is used to build a dynamic demand understanding model based on the query request and the candidate information set to generate atomically executable subtasks, obtain a task status matrix and a tool call task.

[0072] The task planning agent was created by fine-tuning an existing large language model with prompts. (For example, the prompts might include the role setting: "You are a task planning agent. Your core responsibility is to decompose complex user query tasks into a series of executable subtasks that meet atomic constraints (i.e., the single tool call principle and the single call principle). Dynamically optimize and adjust the task sequence based on user input, knowledge blocks, and task status. Based on real-time updates to the task status matrix (S), re-engage in task scheduling when necessary to ensure effective task progress and maximize task utility.")

[0073] During the task planning phase, the task planning agent, the system's core scheduling module, implements demand analysis through multi-dimensional information fusion. This agent first integrates query requests, candidate information sets, user preferences (such as data filtering criteria), historical interaction records (e.g., common query patterns), and personalized requirements (such as the priority of specific fields) to construct a dynamic demand understanding model.

[0074] Based on this model, the agent conducts in-depth analysis of complex tasks and breaks them down into a series of atomically executable subtasks. This splitting logic strictly adheres to the principle of "a single tool, a single call, can complete the task." For example, when searching a database and calling a weather interface simultaneously, the system generates two independent subtasks to ensure independence and non-interference between the functions.

[0075] During the task execution phase, the system dynamically calls upon a list of pre-set tools (including database query tools, code interpretation interfaces, weather interfaces, and other auxiliary tools). After each tool call, the system updates the task status matrix in real time, detailing the completion status of each subtask, the returned results, and any exception codes. If incomplete results or tool errors are encountered, the agent automatically initiates a loop iteration mechanism, reassessing remaining requirements, adjusting the order of tool calls, or adding supplementary conditions to continuously advance the task until the preset completion threshold is met, or triggering an error termination after the maximum number of iterations has been reached. This mechanism is introduced to improve task completion and result accuracy, while ensuring the system's robustness and efficiency in complex scenarios.

[0076] The Task Planning Agent (TPA) is designed to decompose complex user query tasks into a series of executable atomic subtasks, and dynamically adjust the task execution strategy using multi-source heterogeneous information to ensure the smooth completion of the task.

[0077] The entire process can be expressed as the following formula: Where, For task sequence, Planning agents for tasks, Query information for users, is the candidate information set, Preset preferences for users, For conversation history information, To meet the personalized needs of users, is the task state matrix, Indicates the subtasks.

[0078] During the task decomposition phase, the task planning agent needs to decompose the complex query task into several atomic subtasks: Where, Indicates the subtasks, For the The name of the tool that needs to be called for each subtask, For the The calling tool parameters of each subtask. The following constraints are met: Single tool principle: . Single call principle: Where, For the tool name, For atomicity, Indicates the value is true.

[0079] According to the task sequence, a task status matrix is ​​generated. Task status matrix Used to track and provide feedback on the execution status of each subtask in real time: Where, For the tool name, To call the tool parameters, For execution status, Return results for the tool, and The start and end timestamps of the subtask execution.

[0080] The task status matrix is ​​updated in real time: Where, is the updated task state matrix, Indicates update, is the time step The task status matrix at For the currently executing task, For the current calling tool parameter information, The feedback content returned by the tool. Contains error messages, supplementary requests, additional data, etc.

[0081] When a subtask fails to execute or the result is incomplete, the system will trigger the loop iteration mechanism and re-call To modify the task: Where, It means the The task status matrix at the iteration time.

[0082] The loop continues until any of the following conditions are met: 1. The completion threshold is reached: 2. The system determines that the task cannot be completed and triggers an error termination: 3. Complete the task. Indicates that the system evaluates the completion of the entire task based on the completeness, consistency, and tool call quality of the results, with a value range of [0,1]. Indicates the preset completion threshold (such as 0.95). Represents a termination condition judgment function, such as user interruption, excessive number of consecutive failures, and resource exhaustion.

[0083] During the task planning phase, the task planning agent receives input from the previous steps, combines several pieces of information, performs task planning, and returns a corresponding tool call list. The task planning agent primarily receives the following input: user query information, candidate information set, user preset preferences, conversation history, user personalized requirements, and the real-time task status matrix.

[0084] Let's take a specific example. The input information to the task planning agent is as follows: 1. User query information: "I want to travel during the May Day holiday. Please recommend a city that is most suitable for short-term travel, taking into account the weather conditions in the past week, the cost of living in the city, and the cultural atmosphere. Finally, output the conclusion in Chinese, with a brief English summary." 2. Retrieved knowledge blocks: candidate information set. 3. User preset preference: ["Answer in Chinese"]. 4. Conversation history information: [Specific historical information]. 5. User personalized needs: ["Prefer travel environments with pleasant climate and sunny weather", "Prefer coastal cities or cities with natural scenery"] 6. Real-time task state matrix: Empty (empty for the first input; in subsequent iterations, the results of the previous iteration are used as input).

[0085] After receiving the above user input, the task planning agent performs a long thinking and outputs a task state matrix.

[0086] Table 1 Task status matrix (example)

[0087] S4. Task execution: Call the corresponding auxiliary tool according to the tool call task and perform the operation, update the task status matrix according to the operation result, and return the updated task status matrix to the task planning agent to update the tool call task until the task planning agent no longer outputs the tool call task and obtains the task result.

[0088] During mission planning and execution, the system includes auxiliary tools such as database interaction module, calculator, code interpreter, weather interface and translation interface.

[0089] When the Task Planning Agent (TPA) calls the auxiliary tool, the tool receives the tool call parameters from the TPA ( ), and execute the corresponding tool call, feed back the execution results to TPA, and update the task status matrix at the same time ( ). The specific formula of this step is: .

[0090] Where: The execution result returned by the tool, For tool actuators, ∈{calculator, code interpreter, weather interface, translation interface} represents the type of tool called, Specific parameters for tool calls.

[0091] The database interaction module utilizes a multi-agent collaborative architecture to achieve high-precision query construction and full-process supervision. When the task planning agent issues a subtask, the system first relies on a semantic parsing engine to deeply analyze the input content, automatically extracting key entities (such as the table name "order"), attributes (such as the column name "amount"), and constraints (such as the numeric range ">1000"). Then, combined with a pre-built database, a three-level matching process is performed: first, the target data source is precisely determined through table mapping. Second, specific fields are located using column association information. Finally, value type validation ensures the validity of each constraint. Successfully matched elements are encoded as structured context, which in turn drives the SQL generation agent to construct ANSI-compliant query statements using a thought chain reasoning mechanism and a few-shot example library. Furthermore, the task correction agent employs a triple validation mechanism to rigorously monitor and correct the entire process: At the syntactic level, the system verifies the validity of the SQL structure using the abstract syntax tree (AST). At the semantic level, execution plan rehearsals detect potential logical inconsistencies (such as missing fields or data mismatches). At the result level, threshold alerts are set for empty result sets or exceeding limit results. When an anomaly is detected, the correction agent dynamically injects specific prompts (such as supplementing missing JOIN conditions or correcting GROUP BY grouping logic) and automatically falls back to the associated structure information (schema) extraction or SQL generation stage for re-execution, thus forming a closed-loop error correction mechanism. The semantic parsing engine is an existing large language model (LLM).

[0092] When the mission planning agent invokes the database interaction module, during the information extraction and schema association phase, the module first receives the tool call parameters from the mission planning agent (the actual input is the natural language description of the query database after being processed by the mission planning agent). It then uses the Large Language Model (LLM) to perform preliminary analysis and extraction of the input information, specifically, extracting keywords from the query information.

[0093] Record user query information as , the information extraction model is The extracted information is , then the information extraction stage model is: .

[0094] Before matching the database structure information, some preprocessing of the current database is required. The original database structure information is: . .in: For the original database schema, For the Tables, for Middle columns, for No. values.

[0095] The specific operations for preprocessing database structure information include: vectorized computation representation of table names, column names, and typical values ​​to facilitate semantic matching. Let the vectorized operation be function , then: The vectorized representation of the table is: The vectorized representation of the columns is: The vectorized representation of the value is: Where, Represents a vector.

[0096] In the structural information association extraction stage, the vector of the user question is semantically matched with the vector of the structural information to obtain the degree of association between tables, columns, and values: In step S1, the vector representation of the query request has been obtained through vectorization: The correlation scoring formula between the query request and each table is: . The degree of correlation between the question and the column: The degree of correlation between the problem and the specific value: Where, Score the association, is the cosine similarity.

[0097] Then the models whose correlation degree exceeds the threshold are used as correlation structure information. Table selection: . Column selection: . Value selection: Where, For the association table, For the associated columns, For the associated value, is the table association threshold, is the threshold value of the correlation degree of the column, is the correlation threshold of the value.

[0098] After the above structural information preprocessing and correlation formula derivation, the extracted structural information set is: Where, For the structure information collection, For the value of a specific field, For field names, is the table name.

[0099] In the SQL statement generation stage, after matching the information and structure information, this information will be passed to a large language model, which generates structured SQL statements by using the chain of thought (CoT) and few-shot examples (FewShot) method. Let the generated SQL be , the corresponding model is ,but: When generating SQL, LLM combines the paradigm of few-shot learning to imitate prior examples, thereby obtaining more accurate and standardized SQL expressions.

[0100] In the task correction and feedback iteration phase: After the SQL is generated, it is executed and the corresponding execution results are generated ( After execution, the database returns a specific result set or error message). The correction agent will check the execution results and compare them with the semantics of the original question to determine the correctness or deviation of the SQL statement: .in, The task correction agent compares the execution results with the original semantics and draws a conclusion, indicating the consistency or discrepancy between the current execution results and the user's intent. Revise corresponds to the task correction agent.

[0101] If a problem is found, the task correction intelligence will Pass to , SQL will be corrected based on these contexts and new SQL statements will be generated: .

[0102] The above process will be iterated until the result of the SQL statement execution meets the problem semantic requirements or the upper limit of the number of iterations is reached: .

[0103] in, The final output, Indicates the current number of iterations, Indicates the maximum number of iterations allowed.

[0104] When the number of iterations reaches the upper limit or the execution result is correct, the task correction agent will give the execution result and a summary to the task planning agent and update the corresponding task state matrix.

[0105] The Calculator is used to perform calculations on input data.

[0106] The code interpreter is used to execute the input code and return the execution results and status.

[0107] The Weather API is used to obtain weather data based on input parameters.

[0108] The Translation API is used to translate input text.

[0109] After each tool call, the task state matrix is ​​updated. For example, the update example after the calculator call is: .

[0110] For example: According to the tool call task of the first task planning, the weather interface is first called to process the task, a return result is obtained, and the task state matrix is ​​updated according to the result, and then the updated task state matrix is ​​returned to the task planning agent.

[0111] The task planning agent re-executes step S3, performs a long reflection after receiving the updated input, and outputs a new tool call task. For example, the following JSON text [{Tool: Database Interaction Module. Tool Parameters: {Query: Query cost of living comparison data for six cities: City A, City B, City C, City D, City E, and City F}}] is used.

[0112] Then, step S4 is entered again and the database interaction module is called. At this stage, the database interaction module accepts the tool parameters passed in from the previous step and passes them to the information extraction agent, resulting in a corresponding return result: ["City A", "City B", "City C", "City D", "City E", "City F", "City", "Cost of Living", "Comparative Data"]. Next, each keyword undergoes text vectorization, resulting in a vector representation with the same dimensionality as the embedded database structure. These keyword vectors are then subjected to cosine similarity calculations with the table-, column-, and value-level natural language description vectors stored in the vector database to identify the most relevant database entities. Based on this information, the following complete structure of the associated database tables / columns / values ​​is obtained. The SQL generation agent is called with the aforementioned structural information and the tool parameters: "Query comparative cost of living data for six cities: City A, City B, City C, City D, City E, and City F" as input. This generates a SQL statement. The semantic parsing engine, SQL generation agent, and task correction agent are all derived by fine-tuning the prompt words of an existing large language model. An example of a prompt for the semantic parsing engine is: From the following sentence, extract all keywords related to database structure matching. An example of a prompt for the SQL generation agent is: You are an efficient, accurate, and professional SQL generation assistant. Your task is to infer and generate accurate, standardized, and executable SQL query statements based on the user's natural language query and the database's structural information (schema). An example of a prompt for the task correction agent is: As a full-process supervisor, please use a triple verification mechanism (abstract syntax tree AST verification, semantic preview detection, and result threshold alarm) to perform closed-loop error correction on the SQL generation process. When an anomaly is found, prompts will be automatically injected and the process will fall back to the associated pattern extraction or SQL generation link. Finally, the task status matrix is ​​updated based on the results output by the database interaction module, and the updated task status matrix is ​​then returned to the task planning agent.

[0113] Repeat the above operation until the task planning agent no longer outputs the tool call task, and provide the complete output of the task planning agent as the final result. In this embodiment, when the task planning agent does not output the tool list called, it is regarded as an end mark.

[0114] S5. Visualization and result synthesis: When the task planning agent calls the database interaction module, a chart is generated based on the output of the database interaction module and output together with the task result. Otherwise, the task result is directly output.

[0115] Preferably, when a database interaction module is invoked during a task, the system first performs a compliance check on the results returned by the database. It then determines the optimal presentation method based on visualization decision logic (e.g., identifying trend data as preferentially displayed in a line chart, while time series data is better suited for a time series chart). This process takes into account both user needs (e.g., a line chart is preferred when the query contains "trend") and the characteristics of the data itself (e.g., time series data is more suitable for a time series chart).

[0116] Finally, the system synthesizes the output results using a hierarchical summary architecture. This structured presentation of data improves readability and analysis efficiency. The bottom layer generates data snapshots by extracting key metrics (such as maximum values ​​and outliers). The middle layer matches appropriate analysis templates based on the query type, for example, highlighting the percentage of data change in comparative analysis tasks. The top layer performs semantic verification and intent alignment on the final results, based on the original requirements from the task planning phase.

[0117] Ultimately, the system integrates visualizations, structured data summaries, and natural language explanations into Markdown rich text format, using an optimized spatial layout to align with user cognitive habits, providing a clear and efficient data interpretation experience and ensuring that users can intuitively and accurately understand the query results.

[0118] Preferably, step S5 specifically includes steps S51 to S55.

[0119] S51. Check the task status matrix to obtain the calling status of the database interaction module.

[0120] S52: When it is detected that a database interaction module is called during the task iteration process, query results returned by each database interaction module are obtained respectively. Otherwise, the task result is directly output.

[0121] S53. After the system obtains the database query results, it generates a structured summary of the result data and then preliminarily categorizes the field semantics using preprocessing rules. Preprocessing rules include: time fields, categorical fields, numerical fields, and / or ratio fields. The structured summary includes each field's attribute name, data type, unique value statistics, frequency distribution, missing rate, and example values.

[0122] The specific rules of preprocessing are as follows: 1. Time field identification (time / date / timestamp): The data type is time-based, such as datetime, date, and timestamp. Field names contain keywords such as "date," "time," "timestamp," "created," "updated," "year," and "month." Example value formats match time, such as 2023-01-01 and 2024 / 05 / 12 08:00:00. The data type must be high in unique values ​​and have an orderly or periodic distribution.

[0123] 2. Categorical field identification (category / label / enumeration): The data type is string or integer, but the number of unique values ​​is small (e.g., <100). The field name contains keywords such as "type", "category", "label", "status", "gender", and "region". The frequency distribution has a long tail or a significant concentration. High repetition rate: the first few values ​​cover the majority of the sample (e.g., the top 5 values ​​account for more than 80%).

[0124] 3. Numeric field identification (continuous / discrete): The data type is numeric, such as int, float, or decimal. The field name contains keywords, such as "count," "number," "age," "amount," and "score." The field has a high number of unique values ​​(with a threshold set based on the scenario, such as >100). The values ​​are reasonably distributed, without obvious class labels.

[0125] 4. Ratio field identification (proportion / percentage / ratio): Field names contain keywords such as "rate", "ratio", "percent", and "pct". The value range is [0, 1] or [0%, 100%]. The data type is float or string (with a % sign). Precision may be limited or the mantissa may be repeated (e.g., 0.25, 0.5, 0.75).

[0126] 5. Identifier / primary key field identification: Field names contain keywords such as "id", "uuid", "key", and "code". Unique values ​​are equal to the number of samples, with no duplicates. Fields are used to index or link to other tables.

[0127] 6. Text field identification (description / content): Field names containing keywords, such as "description," "comment," "remark," "text," and "content." The average character length is relatively long (e.g., >50 characters). The data type is string / text. The content lacks obvious structured information (not suitable for classification).

[0128] S54: The summary information about the query result extracted based on the preceding operation is provided to a data analysis agent as context, and the data analysis agent selects the required charts for visualization based on preset discrimination rules.

[0129] The data analysis agent was fine-tuned using an existing large language model to generate prompts. For example, the prompts for the data analysis agent might read: "Task Description: Based on the provided database query result summary and the user's original requirement (origin_query), use the Chain-of-Thought reasoning method to systematically derive and determine the most appropriate visualization chart type and related parameters, and output the recommended chart in ECharts' JSON format."

[0130] The default discrimination rules include: When the query involves a combination of time series and numeric fields, a line chart or area chart is recommended (for example, date + sales). When the query involves a combination of category and numeric fields, a column chart, bar chart, or pie chart is recommended (for example, product category + amount). When the query involves proportional data, a pie chart or donut chart is selected (for example, product category + sales volume share). When the query involves a combination of multiple numeric fields, a radar chart or parallel column chart is recommended (for example, date + pageviews + click-through rate + conversions). For difficult-to-determine field combinations, the query results are simply listed without additional visualization.

[0131] Specifically, in the chart type decision process, the model not only considers the type and combination of fields, but also combines the implicit semantic goals in the user's query (such as whether to emphasize trends, whether to highlight comparisons, and whether multi-dimensional combination analysis is required). At the same time, in this process, the key parameter information required for visualization is automatically derived, such as the main axis field, comparison dimension, aggregation method, time granularity, etc., and the corresponding chart parameters are output as the rendering input for subsequent visualization. Ultimately, through this visualization generation logic, the system ensures intelligence from structured query summaries to automatic chart configuration, and has the ability to independently determine chart types, automatically aggregate multi-dimensional data, and adapt visualizations, providing users with data presentation results with semantic adaptability and visual expression accuracy.

[0132] S55. Pass the structured summary to the data analysis agent. After receiving the above input, the data analysis agent will think for a long time and output the json text of the chart.

[0133] S56. Generate a visual chart based on the selected chart type and the JSON text of the chart.

[0134] S57: Integrate the visualization chart and the final results returned by the task planning agent into Markdown rich text format, obtain the final output results, display them to the user, and then end.

[0135] After the entire process is complete, the system also records the status and feedback information of each link in real time, forming a closed-loop optimization mechanism. By counting the execution results of each subtask and generating a task completion matrix, the system dynamically adjusts the collaborative strategies between agents based on user interaction feedback and execution logs. When an anomaly is detected, the system automatically triggers the feedback correction mechanism, adding supplementary conditions or re-adjusting the tool call sequence, and re-executing the task until the desired goal is achieved. Through the above-mentioned processing and feedback optimization measures, this embodiment aims to improve the accuracy and response speed of database queries, enhance the user experience and the system's data insight capabilities.

[0136] Specifically, step S5 primarily visualizes the query results returned by the database interaction module in the preceding step, aiming to make the information presentation more intuitive and clear. This step is not necessarily triggered; it is only triggered when the preceding task planning agent calls a tool that includes the database interaction module and the tool successfully calls and returns the corresponding query results. This visualization step is only initiated after the two preceding steps are completed, and is only called once, ultimately concluding.

[0137] If, by checking the task status matrix, it is found that multiple database interaction modules have been called and the query results have been correctly returned, then their records need to be sorted out and summaries extracted for each of them.

[0138] Extract summary format description: The summary is returned in the format of "list[dictionary]", where each item in the list represents the query result returned by a database interaction module call. Each dictionary (dict) contains three fields: the input query (origin_query) used to call this database interaction module, the markdown format of the returned result (result), and detailed information for each column of the returned result (columns). Columns includes the column name (column_name), column data type (data_type), number of unique values ​​(unique_values_count), example values ​​(example_values, randomly extracted), unique values ​​and their frequency (value_frequencies, extracting the top three with the highest frequency, sorted in descending order), number of missing values ​​(missing_count), and missing rate (missing_rate).

[0139] After that, the format of the summary is processed and passed to the data analysis agent: After receiving the above input, the data analysis agent thinks for a long time and outputs the json text of the chart: The system uses the echarts tool to visualize the json text of the chart. After visualization, Figure 4 Finally, the final output of the previous task planning agent and the chart here will be displayed to the user and then end.

[0140] The database interaction method based on multi-agent collaboration of the present invention aims to achieve efficient processing of data in different formats and improved query accuracy through improved data preprocessing, information retrieval mechanism, intelligent task planning, database interaction and final result visualization synthesis, so as to reduce redundant information transmission and enhance user experience. The database interaction method has the ability to process multi-format input data, can significantly improve the efficiency and accuracy of information retrieval, and realize intelligent and adaptive task scheduling. The database interaction method improves the response speed to user query requests and enhances the flexibility and accuracy of database queries by optimizing the collaboration strategy between multiple agents. At the same time, it introduces a visualization optimization solution in the interactive interface design to further enhance the system's human-computer interaction experience and data insight capabilities.

[0141] Embodiment 2: The present invention provides a database interaction device based on multi-agent collaboration, which includes a query acquisition module, a search source module, a task planning module, a task execution module and an output module.

[0142] The query acquisition module is used to obtain the query request, then perform semantic segmentation through the format adapter, and then vectorize it into a vector representation to obtain the retrieval vector set.

[0143] The search source module is configured to select a search source through a search source selection agent based on the retrieval vector set, search using a meta-search engine, obtain information blocks, and obtain candidate knowledge through a knowledge base search. The information blocks and candidate knowledge are then integrated to obtain a candidate information set.

[0144] The task planning module is used to build a dynamic demand understanding model based on the query request and the candidate information set through a task planning agent to generate subtasks that can be executed atomically, obtain a task status matrix and a tool call task.

[0145] The task execution module is used to call the corresponding auxiliary tool and perform operations according to the tool call task, update the task status matrix according to the operation results, and return the updated task status matrix to the task planning agent to update the tool call task until the task planning agent no longer outputs the tool call task and obtains the task result.

[0146] The output module is used to generate a chart based on the output results of the database interaction module when the task planning agent calls the database interaction module, and output it together with the task result. Otherwise, the task result is directly output.

[0147] Embodiment 3: The present invention provides a database interaction device based on multi-agent collaboration, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the database interaction method based on multi-agent collaboration as described in any paragraph of Embodiment 1.

[0148] Embodiment 4. The present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a database interaction method based on multi-agent collaboration as described in any paragraph of Embodiment 1.

[0149] Obviously, the embodiments described above are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0150] In the several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.

[0151] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0152] If the functions are implemented as software modules and sold or used as standalone products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for causing a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, removable hard drives, read-only memories, random access memories, magnetic disks, or optical disks. It should be noted that, as used herein, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. Without further limitation, the phrase "comprises a..." does not preclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0153] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0154] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0155] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0156] The references to "first" and "second" in the embodiments merely distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or precedence of "first" and "second" can be interchanged where appropriate. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.

[0157] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A database interaction method based on multi-agent collaboration, characterized in that: Include: Get the query request, then perform semantic segmentation through the format adapter, vectorize it into vector representation, and obtain the retrieval vector set; According to the retrieval vector set, a search source is selected by a search source selection agent and a meta-search engine is used to search to obtain information blocks, and candidate knowledge is obtained by searching a knowledge base; then, the information blocks and the candidate knowledge are integrated to obtain a candidate information set; A task planning agent is used to construct a dynamic demand understanding model based on the query request and the candidate information set to generate atomically executable subtasks, obtain a task status matrix and a tool call task; Calling the corresponding auxiliary tool according to the tool calling task and performing an operation, updating the task state matrix according to the operation result, and returning the updated task state matrix to the task planning agent to update the tool calling task until the task planning agent no longer outputs the tool calling task and obtains the task result; When the task planning agent calls the database interaction module, a chart is generated according to the output result of the database interaction module and output together with the task result; otherwise, the task result is directly output.

2. The database interaction method based on multi-agent collaboration according to claim 1, characterized in that: The task planning agent is obtained by fine-tuning the existing large language model with prompt words; The task planning agent is designed to decompose the user's complex query task into a series of executable atomic subtasks, and dynamically adjust the task execution strategy by utilizing multi-source heterogeneous information to ensure the smooth completion of the task; The entire process can be expressed as the following formula: Where, For task sequence, Planning agents for tasks, Query information for users, is the candidate information set, Preset preferences for users, For conversation history information, To meet the personalized needs of users, is the task state matrix, Indicates the subtasks; During the task decomposition phase, the task planning agent needs to decompose the complex query task into several atomic subtasks: Where, Indicates the subtasks, For the The name of the tool that needs to be called for each subtask, For the The calling tool parameters of each subtask; The following constraints are met: Single tool principle: ; Single call principle: Where, For the tool name, For atomicity, Indicates that the value is true; According to the task sequence, a task state matrix is ​​generated; the task state matrix Used to track and provide feedback on the execution status of each subtask in real time: Where, For the tool name, To call the tool parameters, For execution status, Return results for the tool, and The start and end timestamps of the subtask execution.

3. The database interaction method based on multi-agent collaboration according to claim 1 is characterized in that: During mission planning and execution, the system's tools include: database interaction module, calculator, code interpreter, weather interface, and translation interface; The database interaction module adopts a multi-agent collaborative architecture to achieve high-precision query construction and full-process supervision. When the task planning agent issues a subtask, the system first relies on the semantic parsing engine to deeply analyze the input content and automatically extract key entities, attributes, and constraints. Then, combined with the pre-built database, it performs three-level matching: first, the target data source is accurately determined through association table mapping; second, the specific field is located with the help of column association information; finally, the validity of each constraint is ensured through value type verification. The successfully matched elements are encoded as structured context, which in turn drives the SQL generation agent to use the thinking chain reasoning mechanism and the few-sample example library to construct query statements that meet the ANSI standard. In addition, the task correction agent uses a triple verification mechanism to strictly supervise and correct the entire process: at the grammatical level, the system verifies the legitimacy of the SQL structure through the abstract syntax tree; at the semantic level, the execution plan preview is used to detect potential logical contradictions. At the result level, threshold alarms are set for empty result sets or out-of-limit results. When an abnormality is detected, the correction intelligent body dynamically injects specific prompts and automatically falls back to the associated structure information extraction or SQL generation stage for re-execution, thus forming a closed-loop error correction mechanism.

4. The database interaction method based on multi-agent collaboration according to claim 1, characterized in that: When the task planning agent calls the database interaction module, a chart is generated according to the output result of the database interaction module and output together with the task result; Otherwise, directly output the task result, including: Check the task status matrix to obtain the call status of the database interaction module; When it is detected that a database interaction module is called during the task iteration process, the query results returned by each database interaction module are obtained respectively; otherwise, the task result is directly output; After the system obtains the database query results, it generates a structured summary of the result data and then uses preprocessing rules to preliminarily categorize the field semantics. The preprocessing rules include: time fields, categorical fields, numerical fields, and / or ratio fields. The structured summary includes the attribute name, data type, unique value statistics, frequency distribution, missing rate, and example value of each field. Based on the summary information about the query results extracted by the previous operation, the user's query information is provided as context to a data analysis agent; the data analysis agent will select the required visualization charts based on the preset discrimination rules; among which the preset discrimination rules include: when the query involves a combination of time series and numerical type fields, it is recommended to choose a line chart or area chart; when the query involves a combination of category type and numerical type fields, it is recommended to choose a column chart, bar chart or pie chart; when the query involves proportional data, the system selects a pie chart or a donut chart; when the query involves a combination of multiple numerical type fields, it is recommended to choose a radar chart or a parallel column chart; for some field combinations that are difficult to distinguish, the query return results are only listed without additional visualization processing; Pass the structured summary to the data analysis agent. After receiving the input, the data analysis agent performs a long thinking and outputs the JSON text of the chart. Generate a visual chart based on the selected chart type and the chart's JSON text; The visualization chart and the final results returned by the task planning agent are integrated and output in Markdown rich text format. The final output results are obtained and displayed to the user before the end.

5. The database interaction method based on multi-agent collaboration according to any one of claims 1 to 4, characterized in that: According to the retrieval vector set, a search source is selected by a search source selection agent and a meta-search engine is used to search to obtain information blocks, and candidate knowledge is obtained by searching a knowledge base; then, the information blocks and the candidate knowledge are integrated to obtain a candidate information set, specifically including: According to the retrieval vector, the search source is selected by the search source selection agent; Based on the retrieval vector and the selected search source, a meta-search engine is used to search and obtain information blocks related to the user's needs; Perform vectorization on search results and convert text information into vector representation; The vector representation of the search results is compared with the retrieval vector through the semantic similarity algorithm, and the top-K information blocks with the highest semantic relevance are selected; Performing similarity matching between the retrieval vector set and a pre-built knowledge base search to identify the most similar document fragments and obtain candidate knowledge; wherein the knowledge base has document data pre-blocked and stored; The information blocks and candidate knowledge are unified into the candidate information set.

6. The database interaction method based on multi-agent collaboration according to any one of claims 1 to 4, characterized in that: Get the query request, then perform semantic segmentation through the format adapter, vectorize it into vector representation, and obtain the retrieval vector set, which includes: Get query request; The query request is semantically segmented using a format adapter. No additional preprocessing is performed on text-based information. Structured data requires additional parsing to extract key fields and convert them into a unified format to align with other data types. Audio data is directly converted to text to simplify subsequent processing and ensure consistency with other text data. Image data is stored in its original format to ensure that it can be matched and associated with other information during subsequent retrieval. The search term extraction agent extracts search keywords from the semantic blocks; wherein the search term extraction agent is obtained by fine-tuning the prompt words; The search keywords are vectorized to obtain a search vector set.

7. The database interaction method based on multi-agent collaboration according to any one of claims 1 to 4, characterized in that: It also includes pre-built databases and knowledge bases: Based on the structured data uploaded by users, natural language structure information at the table, column, and value levels is extracted, vectorized, and then stored in a vector database; based on the heterogeneous files uploaded by users, they are converted into a unified processing format, and data is processed and vectorized to obtain a searchable content mapping relationship library; Pre-construction of database and knowledge base, including: During the database pre-construction process, the database file or database configuration information uploaded by the user, as well as the database structure information file, is first obtained, and natural language structure information at the table level, column level, and value level is extracted and generated from them; Perform vectorized operations on the generated structural information; The vectorization results are stored in the vector database: each record in the vector database is: Where: Records in the database, is the unique identifier of the vector, Describe the text in the original natural language, The generated high-dimensional vector, For record type, The name of the table to which it belongs, For the column name, is the actual value; In the process of building the knowledge base, the format adapter is used to parse the heterogeneous files uploaded by users to convert data in different formats into a unified processing format. The parsed data is processed in blocks; These knowledge blocks are vectorized and encoded to obtain a searchable content mapping relationship library.

8. A database interaction device based on multi-agent collaboration, characterized in that: Include: The query acquisition module is used to obtain the query request, then perform semantic segmentation through the format adapter, and then vectorize it into a vector representation to obtain the retrieval vector set; A search source module is configured to select a search source through a search source selection agent according to the retrieval vector set and use a meta-search engine to search to obtain information blocks, and to obtain candidate knowledge through a knowledge base search; and then integrate the information blocks and candidate knowledge to obtain a candidate information set; A task planning module is used to build a dynamic demand understanding model based on the query request and the candidate information set through a task planning agent to generate atomically executable subtasks, obtain a task status matrix and a tool call task; A task execution module is used to call the corresponding auxiliary tool according to the tool call task and perform an operation, update the task state matrix according to the operation result, and return the updated task state matrix to the task planning agent to update the tool call task until the task planning agent no longer outputs the tool call task and obtains the task result; The output module is used to generate a chart based on the output results of the database interaction module when the task planning intelligent agent calls the database interaction module, and output it together with the task result; otherwise, the task result is directly output.

9. A database interaction device based on multi-agent collaboration, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement the database interaction method based on multi-agent collaboration as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the database interaction method based on multi-agent collaboration as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Natural language intelligent query method and device based on multi-agent interaction

    CN118012900A

  • Multi-agent collaboration method and system for engineering construction project

    CN119443147A

  • Mine prospecting prediction method based on multi-agent technology

    CN120234387A

  • Intelligent agent-based big language model retrieval enhancement generation system and method

    CN120470088A

  • Natural language query processing

    US12346315B1

Cited By

  • Graph query processing method and system based on multi-agent collaboration

    CN121210525A

  • Method and system for automatically inspecting semiconductor factory by intelligent agent

    CN121235680A

  • Multi-agent data query method and device, electronic equipment and storage medium

    CN121434260A

  • Database question and answer assistant construction method and system and central air conditioner display terminal

    CN121456089A

  • Multi-agent collaborative reasoning voice search system based on user characteristics

    CN121479017A