Multi-agent collaborative data query method and device
Through the data query method of multi-agent collaborative, data query tasks are split and optimized, and the optimal resource is selected to perform sub-tasks, solving the problem of low data query efficiency and achieving efficient data query.
Patent Information
- Application Number
- CN202510894794.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In the prior art, the data query method has caused serious query burden on multiple data interfaces, resulting in low data query efficiency.
The data query method of multi-agent collaborative is adopted, and the data query task is split into multiple subtasks, and a multi-layer scheduling model is used to select the optimal resource from multiple agents, data query services and data interfaces to perform subtasks, and the subtask results are finally fused to optimize resource allocation and parallel processing.
It significantly improves the efficiency of data query, optimizes resource allocation, improves the success rate of task execution, and reduces query time.
Smart Images

Figure CN120470019A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a multi-agent collaborative data query method and device. Background Art
[0002] Data querying refers to the process of retrieving specific information from a database or other data storage system. It is typically implemented using a specific query language (such as SQL or GraphQL). Data querying includes data retrieval: extracting required information from large amounts of data; data analysis: supporting data statistics, analysis, and reporting; and decision support: providing data support for business decisions.
[0003] In the related art, data queries typically involve executing a data query task separately across multiple data interfaces, then fusing the query results from each interface to produce the final query result. As the volume of query data increases, this query method in the related art places a heavy query burden on each data interface, resulting in low data query efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a multi-agent collaborative data query method, device, electronic device, computer-readable storage medium and computer program product, which can effectively improve data query efficiency.
[0005] The technical solution of the embodiment of the present application is implemented as follows: The present invention provides a multi-agent collaborative data query method, including: In response to the received data query instruction, construct a data query task corresponding to the data query instruction, and split the data query task to obtain multiple sub-data query tasks of the data query task; For each of the sub-data query tasks, determining a target selection service for the sub-data query task from a plurality of agent selection services based on the first scheduling model; For each of the sub-data query tasks, based on the second scheduling model, a target data query service corresponding to the sub-data query task is determined from a plurality of data query services corresponding to the target selection service of the sub-data query task; and through the target data query service, based on the third scheduling model, a target data interface service corresponding to the sub-data query task is determined from a plurality of data interface services corresponding to the target data query service; and through the target data interface service, based on the fourth scheduling model, a target data interface corresponding to the sub-data query task is determined from a plurality of data interfaces corresponding to the target data interface service; Executing the sub-data query tasks corresponding to each target data interface through the target data interface corresponding to each sub-data query task, and obtaining the sub-data query results corresponding to each sub-data query task; The sub-data query results corresponding to the multiple sub-data query tasks are merged to obtain a data query result.
[0006] The present invention provides a multi-agent collaborative data query device, comprising: a splitting module configured to construct, in response to a received data query instruction, a data query task corresponding to the data query instruction, and split the data query task into a plurality of sub-data query tasks of the data query task; and determine, for each sub-data query task, a target selection service for the sub-data query task from a plurality of agent selection services based on a first scheduling model; a scheduling module for determining, for each sub-data query task, a target data query service corresponding to the sub-data query task from a plurality of data query services corresponding to a target selection service of the sub-data query task based on a second scheduling model; and determining, through the target data query service, a target data interface service corresponding to the sub-data query task from a plurality of data interface services corresponding to the target data query service based on a third scheduling model; and determining, through the target data interface service, a target data interface corresponding to the sub-data query task from a plurality of data interfaces corresponding to the target data interface service based on a fourth scheduling model; A query module, configured to execute the sub-data query task corresponding to each target data interface through the target data interface corresponding to each sub-data query task, and obtain the sub-data query result corresponding to each sub-data query task; The fusion module is used to fuse the sub-data query results corresponding to the multiple sub-data query tasks to obtain a data query result.
[0007] An embodiment of the present application provides an electronic device, including: a memory for storing computer-executable instructions or computer programs; The processor is used to implement the multi-agent collaborative data query method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0008] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute and implement the multi-agent collaborative data query method provided in an embodiment of the present application.
[0009] The present application provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to execute the multi-agent collaborative data query method described above in the present application.
[0010] The embodiments of the present application have the following beneficial effects: In response to data query instructions, the first scheduling model splits complex data query tasks into multiple subtasks. By analyzing the complexity and resource requirements of the tasks, the task size is appropriately allocated, enabling each subtask to execute independently, thus providing a foundation for parallel processing. Next, for each subtask, the second scheduling model selects the most suitable service from multiple data query services. This selection is based on factors such as service load, performance, and availability, ensuring efficient resource utilization. The third scheduling model further refines this selection process, determining the target data interface service for each subtask from the multiple data interface services corresponding to the target data query service. The fourth scheduling model selects the most suitable interface from the multiple specific data interfaces corresponding to the target data interface service to execute the subtask. This process ensures that each subtask is efficiently executed on the most appropriate data interface. Through this layered scheduling process, each subtask is executed in the most optimized environment, generating sub-data query results. These results are then merged to produce the complete data query result. This hierarchical scheduling mechanism not only optimizes resource allocation and improves task execution success rate, but also significantly reduces query time through parallel processing, thereby significantly improving overall data query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a schematic diagram of the architecture of the data query system provided in an embodiment of the present application; Figure 2 Schematic diagram of the structure of an electronic device for data query provided by an embodiment of the present application; Figure 3 This is a flow chart of the multi-agent collaborative data query method provided in the embodiment of the present application. Figure 1 ; Figure 4 This is a flow chart of the multi-agent collaborative data query method provided in the embodiment of the present application. Figure 2 ; Figure 5 This is a flow chart of the multi-agent collaborative data query method provided in the embodiment of the present application. Figure 3 ; Figure 6This is a schematic diagram of the principle of the multi-agent collaborative data query method provided in the embodiment of the present application Figure 1 ; Figure 7 This is a schematic diagram of the principle of the multi-agent collaborative data query method provided in the embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0012] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0013] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0014] In the following description, the terms "first\second\third" are only used to distinguish similar objects and do not represent a specific order of the objects. It can be understood that "first\second\third" can be interchanged with the specific order or sequence when permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0016] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0017] 1) Data query: This refers to the process of retrieving specific information from a database or other data storage system. It is typically implemented using a specific query language (such as SQL, GraphQL, etc.). Data query functions include data retrieval: extracting required information from large amounts of data; data analysis: supporting data statistics, analysis, and reporting; and decision support: providing data support for business decisions.
[0018] 2) Database: A database is an organized collection of data, typically stored electronically in a computer system. It supports the storage, retrieval, management, and updating of data. Relational databases, such as MySQL, PostgreSQL, and Oracle, use a tabular structure to store data. Non-relational databases, such as MongoDB (a document database), Redis (a key-value store), and Cassandra (a column store), are examples.
[0019] 3) Large Language Model (LLM): A large language model (LLM) is a deep learning-based AI model that typically has billions or even hundreds of billions of parameters and is capable of generating natural language text. These models are trained on large amounts of text data to learn the patterns and structure of language.
[0020] 4) Data Interface: A data interface is an interface that allows data exchange between different systems or components. It defines the data format, transmission method, and operation rules. Application Programming Interface (API): APIs such as REST APIs and GraphQL APIs are used for data exchange between applications. Database interfaces such as JDBC (Java Database Connectivity) and ODBC (Open Database Connectivity) are used for interaction between applications and databases.
[0021] 5) The Model Context Protocol (MCP) is a communication protocol for standardized tool invocation and agent collaboration. It supports models dynamically accessing external APIs or database tools during reasoning while maintaining consistent context. MCP defines a standardized way for models to call external tools. These tools can be APIs, databases, or other services. This standardized approach allows models to interact with external systems more flexibly. During reasoning, models must maintain consistent context. MCP uses a protocol mechanism to ensure that models correctly pass and maintain context information when calling external tools. This helps models maintain consistent and accurate reasoning in complex tasks. MCP supports dynamic access to external tools during reasoning. This means that models can call required tools in real time based on the current reasoning state, eliminating the need to predefine all tool calls at the start of reasoning. In multi-agent systems, multiple models need to work together. MCP provides a mechanism for these models to interact through a standardized communication protocol, enabling efficient collaboration.
[0022] 6) Agent: An intelligent module capable of autonomously completing specific tasks, typically possessing state perception, decision-making and planning capabilities, and tool-invoking capabilities. Agents can form multi-agent systems with other agents, working collaboratively to complete complex tasks. Agents are able to perceive the state of their environment. This means they can gather and understand information about it to make appropriate decisions. This state perception enables agents to respond to environmental changes in real time, ensuring the adaptability and effectiveness of their behavior. Based on this perceived state information, agents can make decisions and perform planning. This includes selecting the optimal action path, allocating resources, and adjusting strategies. This decision-making and planning capability enables agents to effectively solve complex problems and achieve their goals.
[0023] 7) A2A (Agent-to-Agent Protocol): This protocol is an open standard initiated by Google that aims to enable communication and interoperability between different AI agent systems. The core goal of the A2A protocol is to break down the barriers between different AI agents, enabling them to communicate and collaborate effectively in a dynamic multi-agent ecosystem. This allows AI agents to collaborate on complex tasks, for example, by delegating subtasks from one agent to another.
[0024] 8) NL2SQL (Natural Language to SQL): This technology automatically translates human natural language input into executable database query statements. It is a key capability commonly used in intelligent question-answering systems, converting users' natural language questions into structured SQL queries to retrieve relevant information from the database. The primary goal of NL2SQL is to enable users to query databases in natural language without having to write complex SQL statements. This significantly lowers the barrier to entry for database queries, improves the user experience, and enables even non-technical users to easily access data. The NL2SQL system must be able to understand the user's natural language input, including the intent of the question and key information. It then converts the understood natural language question into a structured SQL query statement. The generated SQL query statement is sent to the database for execution, returning the query results. The query results are presented in a user-friendly format, such as a table or chart.
[0025] During the implementation of the embodiments of this application, the applicant discovered that the related technology has the following problems: In the related art, data queries typically involve executing a data query task separately across multiple data interfaces, then fusing the query results from each interface to produce the final query result. As the volume of query data increases, this query method in the related art places a heavy query burden on each data interface, resulting in low data query efficiency.
[0026] The embodiments of the present application provide a multi-agent collaborative data query method, device, electronic device, computer-readable storage medium and computer program product, which can effectively improve the efficiency of data query. The following describes an exemplary application of the data query system provided by the embodiments of the present application.
[0027] See also Figure 1 , Figure 1 1 is a schematic diagram of the architecture of a data query system 100 provided in an embodiment of the present application. A terminal (terminal 400 is shown as an example) is connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0028] The terminal 400 is used for the user to use the client 410 and display the data query result on the graphical interface 410 - 1 (graphic interface 410 - 1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.
[0029] In some embodiments, the server 200 can be an independent physical server, or a server cluster or business system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart TV, smart watch, car terminal, etc., but is not limited to this. The electronic device provided in the embodiment of the present application can be implemented as a terminal or as a server. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application.
[0030] See also Figure 2 , Figure 2 is a structural diagram of an electronic device 500 for data query provided in an embodiment of the present application, wherein: Figure 2 The electronic device 500 shown may be Figure 1 The server 200 or the terminal 400 in Figure 2 The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2Various buses are labeled as bus system 440 .
[0031] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0032] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 430.
[0033] Memory 450 includes volatile memory or nonvolatile memory, or may include both volatile and nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0034] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0035] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0036] In some embodiments, the multi-agent collaborative data query device provided in the embodiments of the present application can be implemented in software. Figure 2A multi-agent collaborative data query device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a splitting module 4551, a scheduling module 4552, a query module 4553, and a fusion module 4554. These modules are logical and can be arbitrarily combined or further split according to the functions they implement. The functions of each module will be described below.
[0037] In other embodiments, the multi-agent collaborative data query device provided in the embodiments of the present application can be implemented in hardware. As an example, the multi-agent collaborative data query device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the multi-agent collaborative data query method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0038] In some embodiments, a terminal or server can implement the multi-agent collaborative data query method provided in the embodiments of the present application by running a computer program or computer-executable instructions. For example, the computer program can be a native program (e.g., a dedicated data query program) or a software module in the operating system, such as a data query module that can be embedded in any program (e.g., an instant messaging client, a photo album program, an electronic map client, or a navigation client); for example, it can be a native application (APP), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be any form of application, module, or plug-in.
[0039] The multi-agent collaborative data query method provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server or terminal provided in the embodiment of the present application.
[0040] See also Figure 3 , Figure 3 This is a flow chart of the multi-agent collaborative data query method provided in the embodiment of the present application. Figure 1 , will combine Figure 3Steps 101 to 108 are shown for illustration. The multi-agent collaborative data query method provided in the embodiment of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following will be illustrated by taking the server alone as an example.
[0041] In step 101, in response to a received data query instruction, a data query task corresponding to the data query instruction is constructed.
[0042] In some embodiments, a data query instruction refers to a request issued by a user or system to query specific data or information. It typically includes query parameters, conditions, and the target data source. Data query instruction sources include: User input: A user enters a query request through an interface (such as a webpage or application), such as entering keywords into a search engine. System call: A query request issued by other modules or services within the system, such as a query operation in a database management system.
[0043] In some embodiments, a data query instruction refers to a request issued by a user or system containing natural language semantics. Its characteristics include: Semantic input: The user expresses the query intent in natural language (for example, "Query the operating risks of a certain company over the past three years"). In the context of big data, users cannot know what query information and intent will accurately match the fields in the database. The multi-layer retrieval and multi-layer scheduling process in this application is precisely the process of hierarchically parsing this ambiguous query intent to improve the hit rate. A data query instruction refers to the query intent expressed by the user in natural language, rather than directly using structured query statements. This query method is more intuitive and natural, suitable for non-technical users. In a big data environment, users often find it difficult to precisely specify query conditions. Therefore, the system uses a multi-layer retrieval and multi-layer scheduling process to gradually parse the user's query intent and generate precise query statements, thereby improving the query hit rate. This not only improves the user experience but also enhances the flexibility and adaptability of the system. A data query instruction refers to a request issued by a user or system containing natural language semantics to retrieve information from a database. This instruction is characterized by the user expressing the query intent in natural language, rather than directly using structured query statements (such as SQL).
[0044] In some embodiments, the data query instruction carries at least one query statement, and the above step 101 can be implemented as follows: for each query statement, from the preset statement-execution logic mapping relationship, query the index entry including the query statement, and determine the execution logic in the index entry as the target execution logic corresponding to the query statement; through the task construction model, based on the target execution logic corresponding to each query statement, generate a data query task corresponding to the data query instruction.
[0045] In some embodiments, the query statement carried in the data query instruction refers to a specific query request, which is usually expressed in a query language (such as SQL, GraphQL, etc.). Each query statement includes specific query parameters and conditions.
[0046] For example, query statement 1: SELECT * FROM table1 WHERE condition1; query statement 2: SELECT * FROM table2 WHERE condition2.
[0047] In some embodiments, the statement-execution logic mapping relationship is a preset mapping table used to map query statements to specific execution logic. This mapping relationship is usually pre-defined by the system to quickly match query statements and execution logic.
[0048] For example, index entries: Each query statement has a corresponding index entry in the mapping relationship, which contains the query statement identifier and the corresponding execution logic. Execution logic: The execution logic in each index entry defines how to execute the query statement, including the data source, query path, result processing, etc.
[0049] In some embodiments, index entries are queried and target execution logic is determined. For each query statement, each query statement in the data query instruction is processed one by one. An index entry containing the query statement is searched from a preset statement-execution logic mapping relationship. The execution logic in the found index entry is determined as the target execution logic corresponding to the query statement.
[0050] For example, query statement 1: SELECT * FROM table1 WHERE condition1; index entry 1: {"query": "SELECT * FROM table1 WHERE condition1", "executionLogic": "Logic1"}; target execution logic 1: Logic1. Query statement 2: SELECT * FROM table2 WHERE condition2; index entry 2: {"query": "SELECT * FROM table2 WHERE condition2", "executionLogic": "Logic2"}; target execution logic 2: Logic2.
[0051] In some embodiments, a task construction model is used to generate a complete data query task based on target execution logic. The task construction model receives the target execution logic corresponding to each query statement and encapsulates the target execution logic into an executable data query task, including query parameters, data source information, result processing logic, etc.
[0052] As an example, the target execution logic of query statement 1 is: Logic1; the generated task 1 is: {"query": "SELECT * FROM table1 WHERE condition1", "dataSource": "table1", "executionLogic": "Logic1"}; the target execution logic of query statement 2 is: Logic2; the generated task 2 is: {"query": "SELECT * FROMtable2 WHERE condition2", "dataSource": "table2", "executionLogic": "Logic2"}.
[0053] In some embodiments, the task construction model includes: a logic parsing unit, used to parse the target execution logic and extract key information in the execution logic; a data source binding unit, used to bind relevant data source information according to the query statement and target execution logic; a task generation unit, used to encapsulate the query statement, target execution logic and data source information into the data query task.
[0054] In some embodiments, the logic parsing unit parses the target execution logic and extracts key information in the execution logic, such as the query path, query operation and result processing logic. The data source binding unit binds the relevant data source information according to the query statement and the target execution logic, and determines which data source to obtain the data from. The task generation unit encapsulates the query statement, target execution logic and data source information into the data query task. The data query task also includes a formatting setting for the query result, which is used to output the query result in a preset format. The task construction model ultimately generates a data query task corresponding to the data query instruction. The data query task includes a query statement, target execution logic and data source information related to the query statement to ensure the integrity and executability of the task.
[0055] In this way, for each query statement, the system queries the corresponding index entry from the preset statement-execution logic mapping relationship. This process, based on the preset mapping relationship, can quickly locate the execution logic of each query statement, avoiding complex real-time parsing and logical deduction, thereby significantly reducing the time complexity of query statement processing. By determining the execution logic in the index entry as the target execution logic, the system can accurately assign an appropriate execution strategy to each query statement. Subsequently, the task construction model generates data query tasks based on these target execution logics. The introduction of the task construction model enables the system to effectively integrate the logic of the query statement with the actual data source information, query parameters, etc., to generate structured and executable data query tasks. Not only does this improve the efficiency of task generation, but it also ensures the accuracy and consistency of the generated tasks through the preset mapping relationship and modeling processing.
[0056] In step 102, the data query task is split to obtain multiple sub-data query tasks of the data query task.
[0057] In some embodiments, a data query task refers to a query operation that the system needs to perform, typically involving retrieving data from a database or other data source. A query task may contain multiple query statements, involve multiple data sources, or involve complex query logic. For large datasets or complex query logic, a single query task may require significant computing resources and time to complete. Therefore, splitting a query task into multiple subtasks can improve execution efficiency and system scalability.
[0058] In some embodiments, the first scheduling model is a preset algorithm or set of rules used to determine how to split a data query task into multiple subtasks. The scheduling model typically considers factors such as task complexity, data source distribution, and resource availability. By rationally allocating subtasks to different resources (such as CPU, memory, and network bandwidth), resource utilization is improved. This ensures even distribution of system load, preventing overloading of certain resources while leaving others idle. Multiple subtasks can be executed simultaneously, thereby shortening overall query time.
[0059] In some embodiments, it is necessary to analyze the structure and requirements of the original data query task. This includes identifying query statements, data sources, query logic, etc. According to the rules of the first scheduling model, the data query task is split into multiple subtasks. The scheduling model may consider the following factors: Complexity of the query statement: split complex query statements into multiple simple sub-queries. Distribution of data sources: split query tasks involving multiple data sources into subtasks for each data source. Resource availability: reasonably allocate the number and size of subtasks based on the resource status of the system. Each subtask is an independent query task that contains part of the logic and data source information of the original query task. Subtasks can be executed in parallel or sequentially, depending on the design of the scheduling model.
[0060] In some embodiments, see Figure 4 , Figure 4 This is a flow chart of the multi-agent collaborative data query method provided in the embodiment of the present application. Figure 2 , Figure 3 Step 102 shown may be performed by Figure 4 Steps 1021 to 1023 are implemented as shown.
[0061] In step 1021 , based on the data query instruction, a first quantity of the plurality of sub-data query tasks and a first probability indicating that the data query task can be executed are predicted.
[0062] In some embodiments, a data query task refers to a query operation that the system needs to perform, typically involving retrieving data from a database or other data source. A query task may contain multiple query statements, involve multiple data sources, or involve complex query logic. For large datasets or complex query logic, a single query task may require significant computing resources and time to complete. Therefore, splitting a query task into multiple subtasks can improve execution efficiency and system scalability. The first scheduling model is a preset algorithm or set of rules that determines how to split a data query task into multiple subtasks. The scheduling model typically considers factors such as task complexity, data source distribution, and resource availability. By rationally allocating subtasks to different resources (such as CPU, memory, and network bandwidth), resource utilization is improved. This ensures even distribution of system load, preventing some resources from being overloaded while others are idle. Multiple subtasks can be executed simultaneously, thereby shortening overall query time. The structure and requirements of the original data query task need to be analyzed. This includes identifying the query statement, data source, and query logic. The data query task is then split into multiple subtasks based on the rules of the first scheduling model. The scheduling model may consider the following factors: Query statement complexity: Complex queries are split into multiple simpler subqueries. Data source distribution: Split query tasks involving multiple data sources into subtasks specific to each data source. Resource availability: Appropriately allocate the number and size of subtasks based on system resource availability. Each subtask is an independent query task, incorporating some of the logic and data source information of the original query task. Subtasks can be executed in parallel or sequentially, depending on the design of the scheduling model.
[0063] In some embodiments, the probability of successful execution of a data query task is predicted based on the complexity of the data query instruction and the rules of the scheduling model. Historical data analysis: Analyze historical execution data to understand the success rate of similar query tasks. Resource assessment: Evaluate the current system resource status, including CPU, memory, network bandwidth, etc. Complexity assessment: Evaluate the difficulty of task execution based on the complexity of the query statement. Apply the scheduling model: Based on the rules of the first scheduling model, comprehensively consider the above factors and predict the probability of successful task execution.
[0064] As an example, the query instruction involves two tables: table1 and table2. The query logic includes joint query and conditional filtering. Evaluate complexity: The query involves two data sources and contains joint query logic. Apply scheduling model: According to the rules of the first scheduling model, it is predicted that the query task will be split into two subtasks: Subtask 1: Retrieve data that meets the conditions from table1. Subtask 2: Retrieve data that meets the conditions from table2. Prediction result: The predicted number of subtasks is 2. Historical data analysis: Analyzing historical execution data, it was found that the success rate of similar query tasks was 90%. Resource evaluation: The current system resources are sufficient, and the CPU and memory usage are low. Complexity assessment: The query logic is complex, but the system resources are sufficient to cope with it. Apply scheduling model: Taking all the above factors into consideration, the probability of successful task execution is predicted to be 85%.
[0065] In some embodiments, the above-mentioned first scheduling model includes a feature extraction layer, a first prediction layer and a second prediction layer. The above-mentioned step 1021 can be implemented in the following manner: through the feature extraction layer, feature extraction is performed on the data query instruction to obtain query instruction features; through the first prediction layer, based on the query instruction features, the first number of the multiple sub-data query tasks is predicted; through the second prediction layer, based on the query instruction features, a first probability for indicating that the data query task can be executed is predicted.
[0066] In some embodiments, the first scheduling model consists of the following: a feature extraction layer, which extracts key features from data query instructions. These features may include query statement structure, data source type, and historical execution time. A first prediction layer, which predicts the number of subtasks after splitting the data query task based on the extracted features. A second prediction layer, which predicts the probability of successful execution of the data query task based on the extracted features.
[0067] In some embodiments, feature extraction involves extracting information useful for prediction from data query instructions. This information can include the length of the query statement, the number of data sources involved, historical execution time, etc. The extracted features are used in the subsequent prediction layer to help the model make more accurate predictions. The first prediction layer is a prediction model used to predict the number of subtasks after the data query task is split. Input: Query instruction features extracted by the feature extraction layer. Output: Predicted number of subtasks. Implementation: The first prediction layer can be implemented using machine learning algorithms (such as linear regression, decision trees, neural networks, etc.).
[0068] In some embodiments, the second prediction layer is a prediction model used to predict the probability of successful execution of a data query task. Input: Query instruction features extracted by the feature extraction layer. Output: Predicted execution probability. Implementation: The second prediction layer can be implemented using machine learning algorithms (such as logistic regression, random forest, or deep learning).
[0069] In some embodiments, data query instructions are analyzed to extract features such as query statement structure, data source type, and historical execution time. For example, for the query statement SELECT * FROM table1 WHERE condition1, the extracted features may include: the amount of data in table1, the complexity of condition1, etc.
[0070] In some embodiments, the extracted features are input into a model at the first prediction layer. Based on the relationship between historical data and the features, the model predicts the number of subtasks after the data query task is split. For example, if historical data shows that similar query statements are typically split into three subtasks, the model will predict the number of subtasks to be three.
[0071] In some embodiments, the extracted features are input into a model in the second prediction layer. Based on the relationship between historical data and the features, the model predicts the probability of successful execution of the data query task. For example, if historical data shows a 90% success rate for similar query statements, the model will predict a 90% probability of execution.
[0072] For example, on a large e-commerce platform, users can query order information through the system. Due to the large volume and complexity of order data, the query task may involve multiple database tables (such as the order table, user table, and product table), and may require complex joint queries. To optimize query performance, the system needs to predict the query complexity (such as the number of subtasks) and the probability of success before executing the query. The feature extraction layer inputs: user-submitted order query instructions, such as "Query order information for all products purchased by user A in the past month." Feature extraction: Query statement structure: Extracts structural features of the query statement, such as whether it involves joint queries and whether it contains time range filters. Data source information: Identifies the data sources involved in the query, such as the order table, user table, and product table. Historical execution data: Extracts information such as execution time and resource consumption of similar queries based on historical query records. User behavior characteristics: Analyzes user A's historical query behavior, such as query frequency and query time range. Output: A set of query instruction features, such as query statement structure features, including joint queries and time range filters. Data source information: Involves the order table, user table, and product table. Historical execution time: The average execution time is 2 seconds. User behavior characteristics: User A queries orders 3 times per month.
[0073] Continuing with the previous example, the first prediction layer receives input: the query instruction feature set obtained from the feature extraction layer. Prediction process: Based on the query statement structure, the query complexity is determined. For example, join queries typically need to be split into multiple subtasks. Based on data source information, the number of tables involved is determined. For example, a query involving three tables may need to be split into three subtasks. The predicted number of subtasks is adjusted based on historical execution data. For example, if historical data shows that similar queries are typically split into four subtasks, the predicted number of subtasks is four. Output: The predicted number of subtasks. For example, the predicted number of subtasks is four. The second prediction layer receives input: the query instruction feature set obtained from the feature extraction layer. Prediction process: Based on the query statement structure, the query complexity is determined. For example, complex join queries may reduce the probability of success. Based on data source information, the number of tables involved and the amount of data are determined. For example, queries involving multiple large tables may reduce the probability of success. Based on historical execution data, the success probability is adjusted. For example, if historical data shows an 85% success rate for similar queries, the predicted probability of success is 85%. Output: The predicted probability of success. For example, the probability of successful execution is predicted to be 85%.
[0074] For example, suppose user A submits a query: "Query order information for all items purchased by user A in the past month." The feature extraction layer extracts features including: query structure (including a joint query (order and item tables) and a time range filter (past month). Data source information: Involves the order, user, and item tables. Historical execution time: Average execution time is 2 seconds. User behavior characteristics: User A queries orders three times per month. The first prediction layer predicts the number of subtasks to be four based on the query structure and data source information: Subtask 1: Extract user A's order information from the order table. Subtask 2: Extract user A's detailed information from the user table. Subtask 3: Extract item information related to the order from the item table. Subtask 4: Perform a joint query on the results of the above subtasks to generate the final query result. The second prediction layer predicts an 85% probability of successful execution based on the query structure, data source information, and historical execution data. This query involves a joint query of multiple tables and is relatively complex. Historical data shows an 85% success rate for similar queries. The current system resources are sufficient to support the query execution.
[0075] In this way, the feature extraction layer analyzes the data query instructions and extracts key features. These features reflect important information such as query complexity, data source distribution, and historical execution history. This provides a data foundation for subsequent predictions, ensuring accuracy and reliability. Next, the first prediction layer uses the extracted features to predict the number of subtasks after the data query task is split. By analyzing the relationship between features and the number of subtasks, the model can quickly and accurately estimate the complexity of the task, providing a basis for resource allocation and task scheduling. Simultaneously, the second prediction layer uses the same features to predict the probability of successful execution of the data query task. This probabilistic prediction enables the system to assess task feasibility in advance, optimize execution strategies, and avoid resource waste. Through this hierarchical prediction mechanism, the system not only efficiently handles complex query tasks but also dynamically adjusts resource allocation and execution strategies based on the prediction results, significantly improving overall system performance and reliability.
[0076] In step 1022, the first probability is compared with a probability threshold to obtain a first comparison result.
[0077] In some embodiments, the first comparison result is used to indicate whether the first probability is greater than a probability threshold. When the first probability is not greater than the probability threshold, a prompt message is output, where the prompt message is used to indicate that the data query has failed.
[0078] In some embodiments, the first probability: This is predicted by the second prediction layer and is used to indicate the probability that the data query task can be successfully executed. It is a value between 0 and 1, indicating the possibility of task success. Probability threshold: This is a preset threshold used to determine whether the possibility of successful task execution is high enough. The specific value of the threshold can be set according to the actual needs and reliability requirements of the system, for example, to 0.5, 0.7 or 0.9. First comparison result: By comparing the first probability with the probability threshold, a Boolean value result (i.e., "yes" or "no") is obtained. If the first probability is greater than the probability threshold, the comparison result is "yes"; otherwise, it is "no".
[0079] In some embodiments, if the first probability is greater than the probability threshold, this means the prediction model determines that the data query task has a high probability of success. In this case, the system can continue executing the query task or further optimize the query strategy to improve execution efficiency. If the first probability is not greater than the probability threshold, this means the prediction model determines that the data query task has a low probability of success. This may be due to the query task being too complex, involving too many data sources, insufficient system resources, or other potential issues.
[0080] In some embodiments, when the first probability is not greater than the probability threshold, the system outputs a prompt message to inform the user that the data query task may fail. The prompt message can be text, a warning icon, or other forms of user interface feedback. Purpose: The purpose of the prompt message is to let the user know in advance that the query task may not be successfully executed, thereby avoiding long wait times and wasting system resources. The user can adjust the query conditions, optimize the query statement, or select another query method based on the prompt message.
[0081] For example, suppose a user submits a complex order query request in a large e-commerce system, seeking information on all orders for a specific product within the past year. The system's second prediction layer predicts a 0.3 (30%) probability of successful execution for this query. If the preset probability threshold is 0.5 (50%), then the first probability (0.3) is no greater than the probability threshold (0.5). In this case, the system will output a prompt message, such as: "The query may not succeed. Please check the query conditions or try again later." This allows users to be informed in advance of the possibility of query failure, avoiding lengthy waits for unsuccessful queries.
[0082] In step 1023 , if the first comparison result indicates that the first probability is greater than the probability threshold, the data query task is split into the multiple sub-data query tasks based on the first quantity.
[0083] In some embodiments, if the first probability is greater than the probability threshold, this means that the prediction model believes that the data query task has a high probability of being successfully executed. This indicates that the complexity and resource requirements of the query task are within the processing capabilities of the system. In this case, the system can continue to execute the query task and split the query task into multiple subtasks based on the predicted number of subtasks (first number). First number: This is predicted by the first prediction layer and represents the number of subtasks after the data query task is split. Analyze the query task: The system first analyzes the structure and requirements of the original data query task and identifies the parts that can be executed independently. Split subtasks: Split the query task into multiple subtasks based on the predicted first number. Each subtask contains part of the logic and data source information of the original query task. Generate subtasks: Each subtask is an independent query task that can be executed in parallel or sequentially. The generation of subtasks ensures the integrity and executability of the task.
[0084] For example, suppose a user submits a complex order query request on a large e-commerce platform, seeking information on all orders for a specific product within the past year. The system's second prediction layer predicts a probability of 0.8 (80%) for successful execution of this query task, with a preset probability threshold of 0.5 (50%). Furthermore, the system's first prediction layer predicts that this query task can be divided into four subtasks. The first probability (0.8) exceeds the probability threshold (0.5), so the first comparison result is "yes." Based on the predicted first quantity (four subtasks), the query task is divided into the following subtasks: Subtask 1: Extract the user's order information from the order table within the past year. Subtask 2: Extract information about a specific product from the product table. Subtask 3: Extract detailed user information from the user table. Subtask 4: Combine the results of these subtasks to generate the final query result.
[0085] For example, the first scheduler, given a user query for "Company's Legal Risks," and the A2A tools included in this hierarchy—"Shell Company Identification System," "Enterprise Legal Litigation Analysis System," "Enterprise ESG Analyzer," and "Equity Structure Query System"—and a first quantity of 1, decides to activate the "Enterprise Legal Litigation Analysis System" for subsequent tasks. Based on the query "Company's Legal Risks" and an analysis of the various tools' functionality, the first scheduling model decides to activate the "Enterprise Legal Litigation Analysis System." This decision may be based on the tool's expertise and efficiency in handling legal risk-related queries, as well as the strong match between the query and its functionality. Furthermore, the first quantity of 1 may indicate that only one tool is needed to satisfy the query, or that the "Enterprise Legal Litigation Analysis System" has the highest priority among all candidate tools. By activating the "Enterprise Legal Litigation Analysis System," the first scheduling model ensures that subsequent tasks are more precisely focused on the "legal risks" aspect of the user's concern, providing more targeted and valuable information. This process not only demonstrates the intelligence and flexibility of the first scheduling model in task allocation, but also demonstrates the efficiency and complementarity of A2A tools in collaborative work.
[0086] In some embodiments, the data query task includes the target execution logic corresponding to each query statement carried in the data query instruction. The above-mentioned splitting of the data query task into the multiple sub-data query tasks based on the first quantity can be achieved in the following manner: determining the second quantity of the target execution logic in the data query task, and comparing the second quantity with the first quantity to obtain a second comparison result; if the second comparison result indicates that the first quantity is less than or equal to the second quantity, splitting the data query task into the first quantity of sub-data query tasks; if the second comparison result indicates that the first quantity is greater than the second quantity, splitting the data query task into the second quantity of sub-data query tasks, and the sub-data query task includes at least one of the target execution logic.
[0087] In some embodiments, a data query task includes multiple query statements, each of which has corresponding target execution logic. Target execution logic is obtained from a preset statement-execution logic mapping and is used to guide the execution of each query statement. A first quantity is predicted by the first prediction layer and indicates the number of subtasks into which the data query task can be split. A second quantity is obtained by analyzing the target execution logic in the data query task and indicates the actual number of target execution logics. A Boolean value is obtained by comparing the first quantity with the second quantity. The comparison result is used to determine how to split the data query task. If the first quantity is less than or equal to the second quantity, the predicted number of subtasks is within the range of the actual target execution logics. Therefore, the data query task can be split according to the predicted first quantity. If the first quantity is greater than the second quantity, the predicted number of subtasks exceeds the actual number of target execution logics. In this case, the data query task should be split according to the actual number of target execution logics. Each subtask contains at least one target execution logic. These subtasks can be executed independently, and the results merged upon completion.
[0088] As an example, assume that in a large database system, a user submits a data query task containing multiple query statements. The system predicts a first quantity of 4 through the first prediction layer and a second quantity of 3 through analysis of the query task. The first quantity (4) is greater than the second quantity (3), so the second comparison result indicates that the first quantity is greater than the second quantity. Based on the second comparison result, the data query task is split into a second number (3) of subtasks. Each subtask contains a target execution logic: Subtask 1: Target execution logic for executing the first query statement. Subtask 2: Target execution logic for executing the second query statement. Subtask 3: Target execution logic for executing the third query statement.
[0089] In this way, the first number (the predicted number of subtasks) obtained based on the prediction model is compared with the second number (the number of target execution logic) obtained from the actual analysis. This comparison process provides a decision-making basis for task splitting and ensures the rationality of the splitting strategy. When the first number is less than or equal to the second number, the system splits the task according to the predicted first number, fully leveraging the forward-looking nature of the prediction model, allowing task splitting to adapt in advance to changes in system resources and query complexity, thereby optimizing resource allocation and improving query efficiency. When the first number is greater than the second number, the system splits according to the actual number of target execution logic, avoiding resource waste and increased management overhead caused by excessive splitting. By ensuring that each sub-data query task contains at least one target execution logic, this method further guarantees the executable and independence of the subtasks, allowing the subtasks to be processed in parallel, further improving the system's throughput.
[0090] In this way, the first scheduling model predicts the first number of multiple sub-data query tasks and the first probability that the data query task can be executed based on the data query instruction. It utilizes the ability to analyze historical data and the current system status to quickly evaluate the feasibility and complexity of the task. By comparing the first probability with the preset probability threshold, the possibility of successful task execution can be determined. If the first probability is greater than the probability threshold, it means that the task has a high success rate. At this time, the data query task is split into multiple sub-tasks based on the predicted first number. By splitting complex tasks into multiple sub-tasks, these sub-tasks can be processed in parallel, thereby significantly improving query efficiency and resource utilization.
[0091] In step 103, for each of the sub-data query tasks, a target selection service for the sub-data query task is determined from a plurality of agent selection services based on the first scheduling model.
[0092] In some embodiments, in intelligent question-and-answer systems, especially those involving multi-agent collaboration, the selection of an agent selection service (A2A Server) is a critical step. An A2A Server is an agent that provides specific functions or services, while the A2A Host is the master agent responsible for coordinating and scheduling these services. The first scheduling model helps the A2A Host select the service most suitable for handling the current sub-data query task from multiple available A2A Servers. A2A Servers are agents that provide specific functions or services, such as: a shell company identification system for identifying shell companies; a corporate legal litigation analysis system for analyzing a company's legal litigation; an enterprise ESG analyzer for analyzing a company's environmental, social, and governance (ESG) performance; and an equity structure query system for querying a company's equity structure. These A2A Servers are available tools in the system, each with specific functions and applicable scenarios.
[0093] In some embodiments, the first scheduling model is a decision-making mechanism that helps the A2A Host select the most suitable service from multiple available A2A Servers to handle the current sub-data query task. The scheduling model typically considers the following factors: Task requirements: The specific requirements of the sub-data query task, such as the query type and required data fields. Service capabilities: The functionality and performance of each A2A Server, such as processing speed and accuracy. Resource availability: The load of each A2A Server in the system, ensuring that the selected service has sufficient resources to handle the task.
[0094] In some embodiments, for each sub-data query task, the first scheduling model (i.e., the agent described above) performs the following steps to determine the target service: It analyzes each sub-data query task, extracting key information and query intent. For example, for the sub-task "Query a company's legal litigation records over the past three years," key information extracted might include "legal litigation," "over the past three years," and "company." It evaluates the capabilities of each A2A server to determine which services can handle the current sub-task. For example, for the sub-task "Query a company's legal litigation records over the past three years," it evaluates which A2A servers have legal litigation analysis capabilities. It checks the current load of each A2A server to ensure that the selected service has sufficient resources to handle the task. For example, if an A2A server is currently overloaded, a less loaded service may be selected. Based on this evaluation, the first scheduling model determines which A2A server to route the sub-task to. For example, if the "Enterprise Legal Litigation Analysis System" can handle the sub-task "Query a company's legal litigation records over the past three years" and its current load is low, the first scheduling model will route the sub-task to the "Enterprise Legal Litigation Analysis System."
[0095] For example, suppose a user's query is "Query a company's legal risks over the past three years." The system breaks this query down into the following subtasks: Subtask 1: Query a company's legal litigation records over the past three years. Subtask 2: Query a company's compliance records over the past three years. For Subtask 1, "Query a company's legal litigation records over the past three years," Task Analysis: Extract key information: "Legal litigation," "Recent three years," and "Company." Service Assessment: Evaluate the capabilities of the A2A Server to determine if the "Enterprise Judicial Litigation Analysis System" can handle this subtask. Resource Check: Check the current load of the "Enterprise Judicial Litigation Analysis System" to confirm that it has sufficient resources. Decision: The first scheduling model decides to route Subtask 1 to the "Enterprise Judicial Litigation Analysis System." For Subtask 2, "Query a company's compliance records over the past three years," Task Analysis: Extract key information: "Compliance records," "Recent three years," and "Company." Service Assessment: Evaluate the capabilities of the A2A Server to determine if the "Enterprise ESG Analyzer" can handle this subtask. Resource Check: Check the current load of the "Enterprise ESG Analyzer" to confirm that it has sufficient resources. Decision: The first scheduling model decides to route subtask 2 to the “Enterprise ESG Analyzer”.
[0096] In step 104 , for each sub-data query task, based on the second scheduling model, a target data query service corresponding to the sub-data query task is determined from multiple data query services corresponding to the target selection service of the sub-data query task.
[0097] In some embodiments, sub-data query task: This is an independent query task split from the data query task, and each sub-task contains one or more query statements and its corresponding target execution logic. Purpose: By splitting complex data query tasks into multiple sub-tasks, query efficiency can be improved, resource allocation can be optimized, and parallel processing can be supported. The second scheduling model is an algorithm or rule set used to determine how to assign sub-data query tasks to different data query services. Function: Resource optimization: According to the characteristics of the sub-tasks and the resource status of the system, tasks are reasonably assigned to different services to improve resource utilization. Load balancing: Ensure that the system load is evenly distributed to avoid overloading of some services while other services are idle. Performance optimization: Select the service that is most suitable for executing each sub-task to improve query efficiency. Determine the target data query service.
[0098] In some embodiments, a target data query service is determined, and the input is: each sub-data query task and its characteristics (such as the complexity of the query statement, the type of data source, the expected execution time, etc.). Processing: Analyze sub-task characteristics: The second scheduling model first analyzes the characteristics of each sub-data query task to understand its resource requirements and execution characteristics. Evaluate service capabilities: Based on the characteristics of the sub-task, evaluate the capabilities of multiple data query services, including the load, performance, availability, etc. of the service. Select target service: Based on the above evaluation, select the data query service that is most suitable for executing the sub-task. The selection criteria may include: Current load of the service: Select a service with a lower load to avoid overload. Performance of the service: Select the service with the best performance to improve query efficiency. Availability of the service: Select the service with the highest availability to ensure the reliability of the task. Distribution of data sources: Select a service close to the data source to reduce data transmission time and cost. Output: Assign a target data query service to each sub-data query task.
[0099] In some embodiments, see Figure 5 , Figure 5 This is a flow chart of the multi-agent collaborative data query method provided in the embodiment of the present application. Figure 3 , Figure 3 Step 104 shown may be performed by Figure 5 Steps 1041 to 1043 are shown to be implemented.
[0100] In step 1041, the second scheduling model is used to predict, based on the sub-data query task, a first matching degree between the sub-data query task and each of the data query services, and a second probability indicating that the sub-data query task can be executed.
[0101] In some embodiments, for predicting the first matching degree, the first matching degree refers to the degree of adaptation between the sub-data query task and a certain data query service. The higher the matching degree, the more suitable the service is for executing the sub-task. Prediction process: Analyze sub-task characteristics: Extract the characteristics of the sub-data query task, such as the complexity of the query statement, the type of data source, the expected execution time, etc. Evaluate service characteristics: Evaluate the characteristics of each data query service, such as the current load, performance, availability, etc. Calculate matching degree: Calculate the matching degree between the sub-task and each service based on the sub-task characteristics and service characteristics. The calculation of matching degree can be based on multiple factors, such as: Current load of the service: Services with lower load have higher matching degree. Performance of the service: Services with higher performance have higher matching degree. Availability of the service: Services with higher availability have higher matching degree. Distribution of data sources: Services close to the data source have higher matching degree. Output: The matching degree between each sub-data query task and each data query service.
[0102] In some embodiments, the second probability refers to the probability that the sub-data query task can be successfully executed on a certain data query service. Analyze sub-task characteristics: Extract the characteristics of the sub-data query task, such as the complexity of the query statement, the type of data source, the historical execution success rate, etc. Evaluate service characteristics: Evaluate the characteristics of each data query service, such as the current load, performance, availability, etc. Calculate the execution probability: Based on the sub-task characteristics and service characteristics, predict the probability of successful execution of the sub-task on each service. The calculation of the execution probability can be based on a variety of factors, such as: Current load of the service: Services with lower load have a higher execution probability. Performance of the service: Services with higher performance have a higher execution probability. Availability of the service: Services with higher availability have a higher execution probability. Historical execution data: The historical execution success rate of similar tasks on the service. Output: The execution probability of each sub-data query task on each data query service.
[0103] As an example, feature extraction is performed on a sub-data query task to obtain sub-query task features. For example, for the sub-task "Extract a user's order information from the order table over the past year," extracted features may include: Query complexity: High; Data source type: Order table; Query condition: Time range (past year); Historical execution time: Average 2 seconds. The first prediction layer predicts the sub-data query task's match with each data query service based on the sub-query task features. For example, for Service A, Service B, and Service C: Service A's match is 0.9; Service B's match is 0.6; Service C's match is 0.3. The second prediction layer predicts the probability of successful execution of the sub-data query task on each data query service based on the sub-query task features. For example, for Service A, Service B, and Service C: Service A's execution probability is 0.95; Service B's execution probability is 0.75; Service C's execution probability is 0.50.
[0104] In some embodiments, the above-mentioned second scheduling model includes a feature extraction layer, a first prediction layer and a second prediction layer. The above-mentioned step 1041 can be implemented in the following manner: through the feature extraction layer, feature extraction is performed on the sub-data query task to obtain sub-query task features; through the first prediction layer, based on the sub-query task features, a first matching degree of the sub-data query task with each of the data query services is predicted; through the second prediction layer, based on the sub-query task features, a second probability is predicted to indicate that the sub-data query task can be executed.
[0105] In some embodiments, the second scheduling model consists of three main layers: A feature extraction layer extracts key features from sub-data query tasks. A first prediction layer predicts the matching degree between the sub-data query task and each data query service based on the extracted features. A second prediction layer predicts the probability of successful execution of the sub-data query task on each data query service based on the extracted features.
[0106] In some embodiments, key features are extracted from the sub-data query task, and these features will be used in the subsequent prediction process. Input: sub-data query task. Output: sub-query task features. Query statement analysis: Extract the structure, complexity, data source types involved, query conditions, etc. of the query statement. Data source information: Extract information such as the location and data volume of the data source involved in the query. Historical execution data: Extract data such as the historical execution time and resource consumption of similar query tasks. Other features: Extract other factors that may affect task execution, such as the length of the query statement and the number of tables involved.
[0107] In some embodiments, for the extracted sub-query task features, the matching degree between the sub-data query task and each data query service is predicted. Input: sub-query task features. Output: the first matching degree between the sub-data query task and each data query service. Matching degree calculation: based on the sub-query task features and the characteristics of the service (such as current load, performance, availability, etc.), the matching degree of each sub-task and each service is calculated. Matching degree evaluation: the higher the matching degree, the more suitable the service is for executing the sub-task. The matching degree calculation can be based on multiple factors, such as: the current load of the service: services with lower loads have higher matching degrees. the performance of the service: services with higher performance have higher matching degrees. the availability of the service: services with higher availability have higher matching degrees. the distribution of data sources: services close to the data sources have higher matching degrees.
[0108] In some embodiments, based on the extracted sub-query task features, the probability that the sub-data query task can be successfully executed on each data query service is predicted. Execution probability calculation: Based on the sub-query task features and the characteristics of the service (such as current load, performance, availability, etc.), the probability of each sub-task being successfully executed on each service is predicted. Probability evaluation: The higher the execution probability, the greater the possibility that the sub-task will be successfully executed on the service. The probability calculation can be based on a variety of factors, such as: The current load of the service: Services with lower loads have a higher probability of execution. The performance of the service: Services with higher performance have a higher probability of execution. The availability of the service: Services with higher availability have a higher probability of execution. Historical execution data: The historical execution success rate of similar tasks on the service.
[0109] As an example, the feature extraction layer extracts features from sub-data query tasks to obtain sub-query task features. For example, for the sub-task of extracting a user's order information from the order table over the past year, the extracted features may include: query complexity: high; data source type: order table; query condition: time range (past year); historical execution time: average 2 seconds. The first prediction layer predicts the matching degree of the sub-data query task with each data query service based on the sub-query task features. For example, for services A, B, and C: the matching degree of service A is 0.9; the matching degree of service B is 0.6; and the matching degree of service C is 0.3. The second prediction layer predicts the probability of successful execution of the sub-data query task on each data query service based on the sub-query task features. For example, for services A, B, and C: the execution probability of service A is 0.95; the execution probability of service B is 0.75; and the execution probability of service C is 0.50.
[0110] In this way, the feature extraction layer extracts features from the sub-data query task, generating sub-query task features. This process quickly identifies key attributes of the task, such as query complexity and the types of data sources involved, providing foundational data for subsequent predictions. Next, the first prediction layer uses these features to predict the first degree of compatibility between the sub-data query task and each data query service. This compatibility calculation comprehensively considers task characteristics and the current state of the service, such as load and performance, to accurately assess which service is most suitable for executing a specific sub-task. Simultaneously, the second prediction layer uses the same features to predict the second probability of successful execution of the sub-data query task on each service. This probabilistic prediction further enhances the system's decision-making capabilities, enabling it to not only select the most suitable service but also assess the probability of successful execution of the task on that service. Through this dual prediction mechanism, the system can make more informed decisions during resource allocation and task execution, improving the success rate and efficiency of task execution while optimizing resource utilization and avoiding resource waste.
[0111] In step 1042, the second probability is compared with a probability threshold to obtain a second comparison result.
[0112] In some embodiments, the second comparison result is used to indicate whether the second probability is greater than a probability threshold. If the second probability is not greater than the probability threshold, a prompt message is output, where the prompt message is used to indicate that the data query has failed.
[0113] In step 1043 , if the second comparison result indicates that the second probability is greater than the probability threshold, a target data query service corresponding to the sub-data query task is determined from the multiple data query services based on the first matching degree.
[0114] In some embodiments, the second probability is obtained by the second prediction layer and indicates the probability that the sub-data query task can be successfully executed on a data query service. A probability threshold is a preset threshold used to determine whether the probability of successful task execution is sufficiently high. A second comparison result is obtained by comparing the second probability with the probability threshold. If the second probability is greater than the probability threshold, the comparison result is "yes"; otherwise, the comparison result is "no".
[0115] In some embodiments, if the second probability is greater than the probability threshold, this means the prediction model determines that the sub-data query task has a high probability of being successfully executed on a particular data query service. In this case, the system can continue executing the query task and select the most appropriate service based on the first match. If the second probability is not greater than the probability threshold, this means the prediction model determines that the sub-data query task has a low probability of being successfully executed on a particular data query service. This may be due to the complexity of the query task, the large amount of data involved, the high service load, or other potential issues.
[0116] In some embodiments, if the second probability is not greater than the probability threshold, the system outputs a prompt message to inform the user that the data query task may fail. The prompt message can be text, a warning icon, or other forms of user interface feedback. The purpose of the prompt message is to let the user know in advance that the query task may not be successfully executed, thereby avoiding long wait times or wasting system resources. The user can adjust the query conditions, optimize the query statement, or choose another query method based on the prompt message.
[0117] In some embodiments, a first degree of match: This is predicted by the first prediction layer and is used to indicate the degree of adaptation of the sub-data query task to each data query service. Select the target service: If the second probability is greater than the probability threshold, the system selects the most suitable service from multiple data query services based on the first degree of match. The selection criteria generally include: Highest degree of match: Select the service with the highest degree of match because it can best meet the requirements of the task. Load balancing: Among multiple services with high degrees of match, select the service with the lower current load to optimize resource allocation. Optimal performance: Select the service with the highest performance to improve query efficiency.
[0118] For example, suppose a user submits a complex query task in a distributed database system. The system splits it into multiple subtasks and uses the second scheduling model to predict the execution probability and matching degree of each subtask on different services. For a particular subtask, the system predicts a second probability of 0.9 (90%) on service A, 0.6 (60%) on service B, and 0.3 (30%) on service C. Assume the probability threshold is 0.7 (70%). Comparison results: Service A's second probability (0.9) exceeds the probability threshold (0.7). Service B's second probability (0.6) is less than the probability threshold (0.7). Service C's second probability (0.3) is less than the probability threshold (0.7). For services B and C, since the second probabilities are less than the probability threshold, the system displays a prompt message: "The query task may fail on this service. Please select another service or optimize the query criteria." For service A, since the second probability is greater than the probability threshold, the system further evaluates the first matching degree: Service A's matching degree is 0.8 (high matching degree). Service B's matching degree is 0.6 (medium matching degree). The matching degree of service C is 0.4 (low matching degree). Service A, which has the highest matching degree, is selected as the target data query service to execute this subtask.
[0119] In this way, feature extraction is performed on each sub-data query task to obtain sub-query task features. These features reflect the complexity and resource requirements of the task. Next, the first prediction layer predicts the first degree of compatibility between the sub-task and each data query service based on these features. This process comprehensively considers factors such as the current load, performance, and availability of the service to ensure the rationality of task allocation. Simultaneously, the second prediction layer, based on the same features, predicts the second probability of successful execution of the sub-task on each service. This probabilistic prediction provides a reliability assessment for task execution. By comparing the second probability with a preset probability threshold, the system can filter out tasks with low success rates and avoid wasting resources on these tasks. When the second probability is greater than the probability threshold, the system further selects the most suitable service from multiple data query services based on the first degree of compatibility to execute the sub-task. This selection process not only considers the compatibility between the task and the service, but also takes into account the overall resource allocation and load balancing of the system.
[0120] In step 105, the target data query service is used to determine the target data interface service corresponding to the sub-data query task from among a plurality of data interface services corresponding to the target data query service based on a third scheduling model.
[0121] In some embodiments, the target data query service is the data query service that is most suitable for executing a sub-data query task, selected through the second scheduling model. The target data query service provides the computing resources and data access capabilities required to execute the sub-tasks. The third scheduling model is an algorithm or set of rules used to manage and optimize the further allocation of sub-data query tasks in the target data query service. It usually takes into account factors such as the characteristics of the task, the performance and load of the data interface service. Resource optimization: Improve resource utilization by reasonably allocating sub-tasks to different data interface services. Load balancing: Ensure that the load in the target data query service is evenly distributed to avoid overloading of certain data interface services. Performance optimization: Select the data interface service that is most suitable for executing each sub-task to improve query efficiency.
[0122] For example, subtask characteristics: complex query statements, involving large amounts of data, and expected to take a long time to execute. Interface 1: currently has a low load and high performance, making it suitable for complex queries. Interface 2: currently has a medium load and medium performance, making it suitable for medium-complexity queries. Interface 3: currently has a high load and low performance, making it suitable for simple queries. Selection Result: Based on the subtask characteristics and service features, Interface 1 is selected as the target data interface service because it currently has a low load and high performance, making it suitable for complex queries.
[0123] In some embodiments, the above-mentioned step 105 can be implemented as follows: through the third scheduling model, based on the sub-data query task, predict the second matching degree of the sub-data query task with each of the data interface services, and the third probability indicating that the sub-data query task can be executed; compare the third probability with the probability threshold to obtain a third comparison result; if the third comparison result indicates that the third probability is greater than the probability threshold, then through the target data query service, based on the second matching degree, determine the target data interface service corresponding to the sub-data query task from the multiple data interface services.
[0124] In some embodiments, the third scheduling model includes a feature extraction layer, a first prediction layer and a second prediction layer. The above-mentioned third scheduling model, based on the sub-data query task, predicts the second matching degree of the sub-data query task with each of the data interface services, and the third probability indicating that the sub-data query task can be executed. It can be achieved as follows: through the feature extraction layer, feature extraction is performed on the sub-data query task to obtain sub-query task features; through the first prediction layer, based on the sub-query task features, the second matching degree of the sub-data query task with each of the data interface services is predicted; through the second prediction layer, based on the sub-query task features, the second probability indicating that the sub-data query task can be executed is predicted.
[0125] In some embodiments, the third scheduling model: This is an algorithm or rule set used to manage and optimize the further allocation of sub-data query tasks in the target data query service. It generally takes into account factors such as the characteristics of the task, the performance and load of the data interface service. Function: Resource optimization: Improve resource utilization by rationally allocating sub-tasks to different data interface services. Load balancing: Ensure that the load in the target data query service is evenly distributed to avoid overloading of certain data interface services. Performance optimization: Select the data interface service that is most suitable for executing each sub-task to improve query efficiency.
[0126] In some embodiments, the second matching degree refers to the degree of adaptation between the sub-data query task and a certain data interface service. The higher the matching degree, the more suitable the service is for executing the sub-task. Analyze sub-task characteristics: extract the characteristics of the sub-data query task, such as the complexity of the query statement, the type of data source, the expected execution time, etc. Evaluate data interface service characteristics: evaluate the characteristics of each data interface service, such as current load, performance, availability, etc. Calculate matching degree: calculate the matching degree of the sub-task with each data interface service based on the sub-task characteristics and service characteristics. The calculation of matching degree can be based on multiple factors, such as: Current load of the service: Services with lower load have higher matching degree. Performance of the service: Services with higher performance have higher matching degree. Availability of the service: Services with higher availability have higher matching degree. Distribution of data sources: Services close to the data source have higher matching degree.
[0127] In some embodiments, the third probability refers to the probability that the sub-data query task can be successfully executed on a certain data interface service. Analyze sub-task characteristics: extract the characteristics of the sub-data query task, such as the complexity of the query statement, the type of data source, the historical execution success rate, etc. Evaluate data interface service characteristics: evaluate the characteristics of each data interface service, such as the current load, performance, availability, etc. Calculate the execution probability: predict the probability of successful execution of the sub-task on each data interface service based on the sub-task characteristics and service characteristics. The calculation of the execution probability can be based on a variety of factors, such as: The current load of the service: services with lower loads have a higher probability of execution. The performance of the service: services with higher performance have a higher probability of execution. The availability of the service: services with higher availability have a higher probability of execution. Historical execution data: the historical execution success rate of similar tasks on this service.
[0128] In some embodiments, a probability threshold is a preset threshold used to determine whether the probability of successful task execution is sufficiently high. A third comparison result is obtained by comparing the third probability with the probability threshold. If the third probability is greater than the probability threshold, the comparison result is yes; otherwise, it is no.
[0129] In some embodiments, the third probability is compared with a probability threshold: if the third probability is greater than the probability threshold, it indicates that the task has a higher success rate on the data interface service. Select the target data interface service: based on the second matching degree, select the most suitable service from multiple data interface services. The selection criteria generally include: Highest matching degree: Select the service with the highest matching degree because it best meets the requirements of the task. Load balancing: Among multiple services with high matching degrees, select the service with the lower current load to optimize resource allocation. Optimal performance: Select the service with the highest performance to improve query efficiency.
[0130] For example, assume that in a distributed database system, the target data query service (Service A) has multiple data interface services (Interface 1, Interface 2, and Interface 3), each with varying performance and load. The system needs to select the most suitable data interface service for a sub-data query task. Predict the second degree of match. Sub-task characteristics: The query statement is complex, involves a large amount of data, and is expected to take a long time to execute. The match for Interface 1 is 0.9 (high match). The match for Interface 2 is 0.6 (medium match). The match for Interface 3 is 0.3 (low match). Predict the third probability: the execution probability for Interface 1 is 0.95 (high probability), the execution probability for Interface 2 is 0.75 (medium probability), and the execution probability for Interface 3 is 0.50 (low probability). Comparison results: The third probability for Interface 1 (0.95) exceeds the probability threshold (0.7). The third probability for Interface 2 (0.75) exceeds the probability threshold (0.7). The third probability for Interface 3 (0.50) is less than the probability threshold (0.7). Selection result: Based on the second matching degree, interface 1 is selected as the target data interface service because it not only has the highest matching degree but also has an execution probability higher than the probability threshold.
[0131] In this way, based on the characteristics of the sub-data query task, the second degree of matching of each task with each data interface service is predicted, taking into account factors such as the complexity of the task, the type of data source, and the current load and performance of the service, so as to accurately evaluate which services are most suitable for executing specific sub-tasks. At the same time, the model also predicts the third probability that the sub-task can be successfully executed on each service. This probability prediction provides a reliability assessment for the execution of the task. By comparing the third probability with the preset probability threshold, the system can screen out tasks with a lower success rate and avoid wasting resources on these tasks. When the third probability is greater than the probability threshold, the system further selects the most suitable service from multiple data interface services based on the second degree of matching to execute the sub-task, which not only considers the adaptability of the task and the service, but also takes into account the overall resource allocation and load balancing of the system.
[0132] In step 106, the target data interface service is used to determine the target data interface corresponding to the sub-data query task from among the multiple data interfaces corresponding to the target data interface service based on a fourth scheduling model.
[0133] In some embodiments, a target data interface service is selected through a third scheduling model as the data interface service most suitable for executing a sub-data query task. The target data interface service provides the specific data access interfaces required to execute the sub-task, which connect to the actual data storage or data source. The fourth scheduling model is an algorithm or set of rules used to manage and optimize the further allocation of sub-data query tasks to target data interface services. It typically takes into account factors such as the characteristics of the task, the performance of the specific data interface, and the load.
[0134] In some embodiments, analyzing subtask characteristics: extracting characteristics of the sub-data query task, such as the complexity of the query statement, the type of data source, the expected execution time, etc. Evaluating specific data interface characteristics: evaluating the characteristics of each specific data interface in the target data interface service, such as the current load, performance, availability, etc. Selecting a target data interface: based on the above evaluation, selecting the specific data interface that is most suitable for executing the sub-task. The selection criteria may include: Current load of the interface: interfaces with lower loads have a higher matching degree. Performance of the interface: interfaces with higher performance have a higher matching degree. Availability of the interface: interfaces with higher availability have a higher matching degree. Distribution of data sources: interfaces close to the data source have a higher matching degree. Assign a target data interface to each sub-data query task.
[0135] In some embodiments, the above-mentioned step 106 can be implemented as follows: through the fourth scheduling model, based on the sub-data query task, predict the third matching degree between the sub-data query task and each of the data interfaces, and the fourth probability indicating that the sub-data query task can be executed; compare the fourth probability with the probability threshold to obtain a fourth comparison result; if the fourth comparison result indicates that the fourth probability is greater than the probability threshold, then through the target data interface service, based on the third matching degree, determine the target data interface corresponding to the sub-data query task from the multiple data interfaces.
[0136] In some embodiments, the fourth scheduling model includes a feature extraction layer, a first prediction layer and a second prediction layer. The above-mentioned fourth scheduling model, based on the sub-data query task, predicts the third matching degree between the sub-data query task and each of the data interfaces, and the fourth probability used to indicate that the sub-data query task can be executed. It can be achieved as follows: through the feature extraction layer, feature extraction is performed on the sub-data query task to obtain sub-query task features; through the first prediction layer, based on the sub-query task features, the third matching degree between the sub-data query task and each of the data interfaces is predicted; through the second prediction layer, based on the sub-query task features, the fourth probability used to indicate that the sub-data query task can be executed is predicted.
[0137] In some embodiments, the fourth scheduling model: This is an algorithm or rule set used to manage and optimize the further allocation of sub-data query tasks in the target data interface service. It generally takes into account factors such as the characteristics of the task, the performance and load of the specific data interface. Function: Resource optimization: Improve resource utilization by reasonably allocating sub-tasks to different specific data interfaces. Load balancing: Ensure that the load in the target data interface service is evenly distributed to avoid overloading of certain data interfaces. Performance optimization: Select the specific data interface that is most suitable for executing each sub-task to improve query efficiency.
[0138] In some embodiments, the third matching degree refers to the degree of adaptation between the sub-data query task and a specific data interface. The higher the matching degree, the more suitable the interface is for executing the sub-task. Extract the characteristics of the sub-data query task, such as the complexity of the query statement, the type of data source, the expected execution time, etc. Evaluate the characteristics of each specific data interface in the target data interface service, such as the current load, performance, availability, etc. Calculate the matching degree between the sub-task and each specific data interface based on the sub-task characteristics and the interface characteristics. The matching degree can be calculated based on multiple factors, such as: The current load of the interface: the interface with lower load has a higher matching degree. The performance of the interface: the interface with higher performance has a higher matching degree. The availability of the interface: the interface with higher availability has a higher matching degree. The distribution of data sources: the interface close to the data source has a higher matching degree.
[0139] In some embodiments, the fourth probability refers to the probability that the sub-data query task can be successfully executed on a specific data interface. Analyze sub-task characteristics: extract the characteristics of the sub-data query task, such as the complexity of the query statement, the type of data source, the historical execution success rate, etc. Evaluate specific data interface characteristics: evaluate the characteristics of each specific data interface, such as the current load, performance, availability, etc. Calculate the execution probability: predict the probability of successful execution of the sub-task on each specific data interface based on the sub-task characteristics and interface characteristics. The calculation of the execution probability can be based on a variety of factors, such as: The current load of the interface: the execution probability of the interface with lower load is higher. The performance of the interface: the execution probability of the interface with higher performance is higher. The availability of the interface: the execution probability of the interface with higher availability is higher. Historical execution data: the historical execution success rate of similar tasks on this interface.
[0140] In some embodiments, a probability threshold is a preset threshold used to determine whether the probability of successful task execution is sufficiently high. A fourth comparison result is obtained by comparing the fourth probability with the probability threshold. If the fourth probability is greater than the probability threshold, the comparison result is yes; otherwise, the comparison result is no.
[0141] In some embodiments, if the fourth probability is greater than a probability threshold, it indicates that the task has a higher success rate on the data interface. Based on the third degree of match, the most suitable data interface is selected from multiple specific data interfaces. The selection criteria generally include: Highest degree of match: Select the interface with the highest degree of match because it best meets the requirements of the task. Load balancing: Among multiple interfaces with high degrees of match, select the interface with the lowest current load to optimize resource allocation. Optimal performance: Select the interface with the highest performance to improve query efficiency. A target data interface is assigned to each sub-data query task.
[0142] For example, assume that in a distributed database system, the target data interface service (Service A) has multiple specific data interfaces (Interface 1, Interface 2, and Interface 3), each with varying performance and load. The system needs to select the most appropriate specific data interface for a sub-data query task. Predicting the third match-degree sub-task features: the query statement is complex, involves a large amount of data, and is expected to take a long time to execute. Interface 1's match-degree is 0.9 (high match). Interface 2's match-degree is 0.6 (medium match). Interface 3's match-degree is 0.3 (low match). Predicting the fourth probability: Interface 1's execution probability is 0.95 (high probability). Interface 2's execution probability is 0.75 (medium probability). Interface 3's execution probability is 0.50 (low probability). Interface 1's fourth probability (0.95) is greater than the probability threshold (0.7). Interface 2's fourth probability (0.75) is greater than the probability threshold (0.7). Interface 3's fourth probability (0.50) is less than the probability threshold (0.7). Based on the third matching degree, interface 1 is selected as the target data interface because it not only has the highest matching degree but also has an execution probability higher than the probability threshold.
[0143] In this way, the fourth scheduling model predicts the third degree of matching between each task and each data interface based on the characteristics of the sub-data query task. This process takes into account factors such as the complexity of the task, the type of data source, and the current load and performance of the interface, so that it can accurately evaluate which interfaces are most suitable for executing specific sub-tasks. At the same time, the model also predicts the fourth probability that the sub-task can be successfully executed on each interface. This probability prediction provides a reliability assessment for the execution of the task. By comparing the fourth probability with the preset probability threshold, the system can screen out tasks with a lower success rate and avoid wasting resources on these tasks. When the fourth probability is greater than the probability threshold, the system further selects the most suitable data interface from multiple data interfaces based on the third degree of matching to execute the sub-task. This selection process not only takes into account the adaptability of the task and the interface, but also takes into account the overall resource allocation and load balancing of the system.
[0144] In step 107, the sub-data query tasks corresponding to each target data interface are executed through the target data interface corresponding to each sub-data query task, and the sub-data query results corresponding to each sub-data query task are obtained.
[0145] In some embodiments, the target data interface is a specific data interface that is most suitable for executing a sub-data query task, selected through the fourth scheduling model. The target data interface provides the specific data access path required to execute the sub-task, and is connected to the actual data storage or data source. Task allocation: Assign each sub-data query task to its corresponding target data interface. Task execution: Execute each sub-data query task through the target data interface. The execution process includes: Data access: Access the actual data storage or data source through the target data interface. Query execution: Execute the query operation according to the query statement of the sub-data query task. Result generation: Generate the query result for each sub-data query task.
[0146] As an example, assume a distributed database system contains multiple sub-data query tasks. Each task is assigned to the most appropriate target data interface using the fourth scheduling model. The system needs to execute the query tasks and obtain the query results through these target data interfaces. Task assignment: Subtask 1: Assigned to target data interface 1. Subtask 2: Assigned to target data interface 2. Subtask 3: Assigned to target data interface 3.
[0147] Continuing with the previous example, Subtask 1: Access the data source through target data interface 1. Execute a query statement, such as SELECT * FROM table1 WHERE condition1. Generate query results, such as [{"id": 1, "value": "data1"}, {"id": 2, "value": "data2"}]. Subtask 2: Access the data source through target data interface 2. Execute a query statement, such as SELECT * FROM table2 WHERE condition2. Generate query results, such as [{"id": 3, "value": "data3"}, {"id": 4, "value": "data4"}]. Subtask 3: Access the data source through target data interface 3. Execute a query statement, such as SELECT * FROM table3 WHERE condition3. Generate query results, such as [{"id": 5, "value": "data5"}, {"id": 6, "value": "data6"}].
[0148] In step 108, the sub-data query results corresponding to the multiple sub-data query tasks are merged to obtain a data query result.
[0149] In some embodiments, sub-data query results refer to the query results generated after executing each sub-data query task. These results are typically partial data or intermediate results that require further processing to generate the final data query result. Result fusion refers to the merging of multiple sub-data query results into a complete data query result. This process typically involves operations such as merging, deduplication, sorting, and formatting data. This generates a complete and consistent data query result and provides it to the user or caller.
[0150] In some embodiments, all sub-data query results are merged into a unified data structure. For example, multiple result lists are merged into a single list. If duplicate data items exist in the sub-data query results, duplicate removal is performed. The merged data is sorted as needed, for example, by a specific field (such as ID). The merged data is formatted into the final query results, for example, by converting the data into JSON or a table format.
[0151] As an example, suppose that in a distributed database system, there are multiple sub-data query tasks, each of which is executed through the target data interface and generates partial query results. The system needs to merge these results to generate the final data query result (specifically, the target data interface executes and generates partial query results, which are passed up layer by layer in the system (see Figure 6 The system shown in the figure goes from the data table layer to the query construction layer, from the query construction layer to the agent layer, from the agent layer to the A2A agent selection layer, and from the A2A agent selection layer to the main agent layer (A2A host layer). The A2A host layer aggregates the query results. The final aggregated results are integrated through a large language model to provide semantic answers to the user's interactive questions. The results of subtask 1 are: [{"id": 1, "value": "data1"}, {"id": 2, "value": "data2"}]; the results of subtask 2 are: [{"id": 3, "value": "data3"}, {"id": 4, "value": "data4"}]; the results of subtask 3 are: [{"id": 5, "value": "data5"}, {"id": 6, "value": "data6"}]. All results are merged into a list, and duplicate data items are checked and removed (if any). The data is sorted by the id field. Format the merged data into the final query result.
[0152] In this way, after responding to a data query instruction, the first scheduling model splits the complex data query task into multiple subtasks. By analyzing the task's complexity and resource requirements, the task size is appropriately allocated, allowing each subtask to execute independently, providing a foundation for parallel processing. Next, for each subtask, the second scheduling model selects the most suitable service from multiple data query services. This selection is based on factors such as service load, performance, and availability, ensuring efficient resource utilization. The third scheduling model further refines this selection process, determining the target data interface service for each subtask from the multiple data interface services corresponding to the target data query service. The fourth scheduling model selects the most suitable interface from the multiple specific data interfaces corresponding to the target data interface service to execute the subtask. This process ensures that each subtask is efficiently executed on the most appropriate data interface. Through this layered scheduling process, each subtask is executed in the most optimized environment, generating sub-data query results. These results are then merged to produce the complete data query result. This hierarchical scheduling mechanism not only optimizes resource allocation and improves the success rate of task execution, but also significantly reduces query time through parallel processing, thereby significantly improving overall data query efficiency.
[0153] The following describes an exemplary application of the embodiment of the present application in an actual data query application scenario.
[0154] Large Language Models (LLMs) are a type of generative AI based on deep neural networks and self-supervised pre-training on a trillion-word corpus. They can understand and generate multimodal content, including multilingual text, code, and tables. Thanks to their large-scale parameters and contextual learning mechanisms, they demonstrate emergent capabilities such as reasoning, dialogue, programming, and enhanced retrieval, and are becoming the core engine of general intelligent systems.
[0155] In high-cost, sensitive data query services, primarily in the financial, government, and enterprise sectors, massive amounts of data in the billions and a callable ecosystem of hundreds of thousands of API tools make it difficult to achieve accurate, low-cost, and efficient data retrieval using traditional, manual methods. The continuous improvement of LLM foundation performance and the widespread adoption of the MCP (Model Context Protocol) have led to the emergence of architectures such as React and Planner-Execute. These architectures enable mainstream LLMs to call MCP tools, enabling code execution, voice conversion, and search enhancement, significantly expanding their functionality and forming intelligent entities.
[0156] NL2SQL technology, which converts natural language into database query commands, is also a key tool for the Multi-Channel Query Processing (MCP) required for LLM. Large-model-based NL2SQL tools can be embedded as MCP tools within database-oriented systems, often used in business question-answering agents for complex multi-table queries. This approach provides a viable path for MCP- and agent-based data query technology, but it still does not address the core issue of balancing cost, efficiency, and accuracy in data querying with massive amounts of data.
[0157] With Google's recent release of a standard A2A (Agent to Agent) protocol, heterogeneous AI agents will be able to collaborate efficiently while ensuring security. For massive amounts of data, applying more complex hierarchical data retrieval designs will significantly reduce the time complexity of data queries. The A2A protocol provides a technical foundation for the integration and collaboration of data query agents, which previously relied on single-step MCP calls. On the other hand, in data query systems built on A2A-MCP systems, errors are not "instantly exposed" like with single-layer calls; instead, they are amplified and accumulated layer by layer. Furthermore, as the layers deepen, queries become more expensive the further they go. If the same problem is resolved across redundant agents / tools, the bill increases exponentially. Therefore, the key to designing a robust A2A-MCP system lies in a supporting scheduling system based on reinforcement learning.
[0158] The embodiment of the present application proposes a hierarchical scheduling business query big data service based on the A2A-MCP framework. The embodiment of the present application minimizes the time complexity through more than ten layers of grading, and at the same time sets a scheduler at the key layer. In each layer of Agent / Tool selection process, multiple candidate query trajectories are synchronously generated, and reinforcement learning updates are performed based on the relative reward signal that comprehensively considers accuracy, latency and economic cost, thereby significantly reducing the overall call cost while ensuring retrieval accuracy. Through this method, the cross-agent query process in a massive data environment can be stably converged to the cost-effectiveness optimal path within a limited number of rounds, providing a feasible and scalable massive data intelligent retrieval solution for comprehensive business query scenarios.
[0159] Big data leads to an explosion in computational complexity: a complete query chain must go through Layer: Data → Table → DB → API → … → User. A set of optional actions for a layer The quantity is If brute force search is used, the computational complexity is equivalent to the full path combination (1)In the embodiment of this application, The key steps in the retrieval chain are A local MDP, the parent strategy only determines "which layer / which branch to choose" without enumerating the entire path; (2) Training the four-level scheduler Transformer encoding Make discrete decisions in the candidate subset, Reduced to (3) Each level of scheduler will assign corresponding instructions to the A2A, agent or MCP at that level, including routing, TOP-K and early stop probability. , thereby controlling the total number of query links and reducing the computational complexity to a calculable range. Errors accumulate deeply. The traditional strategy of using simple thresholds for Top-K screening and fixed pipeline design, because the thresholds and directions are simply specified, will introduce errors or uncertainties to each level of retrieval operations in the multi-level links. Since inaccurate information returned by a certain API will mislead subsequent decisions, these errors may gradually increase as the link length increases. If there is no effective control mechanism, the reliability of the final result will be difficult to guarantee. The previous method lacked this mechanism. In the embodiment of the present application, the Top-k threshold is adjusted in real time with the context and budget according to the network output of the scheduler; the stop threshold Guaranteed by Taylor upper bound At the same time, the embodiment of the present application introduces a hierarchical scheduling mechanism with feedback. The scheduler built on three levels performs policy control from top to bottom and feedbacks information such as accuracy, cost, and delay from bottom to top, thereby controlling the cumulative error transmitted to the deeper layers. Limiting Top-K only at the bottom level cannot guarantee the number of global API calls. Converging on budget In the embodiment of the present application, the optimization algorithm under the dual budget constraint ensures that it exits early when the local accuracy increment decreases when the additional computation is performed. At the same time, the optimization is performed under a constrained model, and its mathematical form is , while using Update Lagrange multipliers , thus ensuring the second constraint Always stuck on the budget line Nearby. Strategy oscillation: Multi-level retrieval often only obtains meaningful feedback (such as the accuracy of the final answer) in the final stage. The intermediate steps lack direct reward signals, which leads to extremely sparse rewards in the reinforcement learning training process. It is difficult for the intelligent agent to measure the contribution of each decision to the final result. If the classic PPO reinforcement learning strategy is adopted, the training of highly sparse rewards is extremely unstable and easily leads to KL divergence. In order to ensure the learnability and stable convergence of the strategy in extremely sparse reward scenarios, the training in the embodiment of the present application adopts a bottom-up freeze-thaw process: first lock the bottom-level strategy training after convergence, and then thaw and optimize the upper layer layer by layer to avoid inter-layer noise interference. GRPO is applied within each layer, and the advantage signal is generated by the relative ranking within the group and updated within the KL trust domain. Stable convergence can be achieved under sparse reward conditions without the need for a value network.
[0160] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the principle of the multi-agent collaborative data query method provided in the embodiment of the present application Figure 1 The overall retrieval link from user queries to the underlying data is expanded to a complex "11+4" structure consisting of an eleven-level directed acyclic tree graph and a four-level scheduler.
[0161] In some embodiments, see Figure 6 , the application embodiment designed an "11+4" retrieval chain structure consisting of an eleven-level hierarchical retrieval chain and a four-level scheduler, and used it as an important part of the patent protection scope. This structure subdivides a complete information retrieval process into eleven levels of operation units, among which the parameters of the four key routing levels are regulated by the corresponding four-level schedulers. The functions of each level operation and each scheduler are as follows: First level: user input layer (User), the user carries the query information. Second level: application interaction layer (Application), which should include a graphical interface, interact with the user, obtain query information, and pass it to the subsequent layer; the query data is then fed back to the user. Third level: main intelligent agent layer (A2A Host), the total scheduler (Scheduler 1) (that is, the first scheduling model described above) and Agent Host parse the user's query data respectively, and Scheduler 1 obtains the routing, Top-K and cutoff probabilities through the input query information through Transformer. The Agent Host then parses, infers, splits, and processes the user's query. The fourth level is the A2A selection layer (A2A Client). The A2A Host divides the overall query task into several subtasks and, in the optional A2A Server, selects the corresponding sub-A2AServer to process the different subtasks or stop the task according to the instructions given by Scheduler 1, thus achieving A2A collaboration. The fifth level is the agent selection layer (A2AServer). Each A2A Server to which the A2A Client connects has its own built-in secondary scheduler (Scheduler 2) (also known as the second scheduling model described above), which determines the scale of Agent resources to be used for each subtask. Simultaneously, the Agent on the A2A Server parses, infers, splits, and processes the task, and selects and activates the corresponding Agent according to the scheduler's instructions. The sixth level is the Agent layer. Each Agent connected to an A2A Server may have its own built-in three-level scheduler (Scheduler 3) (also known as the third scheduling model described above). However, for Agents without a scheduler or not created internally by the system, static scheduling parameters can be estimated based on their agent descriptions and call history. The Agent parses, infers, splits, and processes tasks, and selects and activates the corresponding type and number of MCP Clients according to the scheduler's instructions. The seventh level is the Query Construction Layer (MCPServer). Activated MCP Clients connect to their corresponding MCP Servers. These MCP Servers may have built-in four-level schedulers (also known as the fourth scheduling model described above). If not, static scheduling parameters are used. The scheduler determines the scale of API resources to be used for each specific subtask, and the MCP Server activates the corresponding API to complete each subtask. The eighth level is the Tool Call Layer (API). This layer executes specific API requests and sends the constructed query tasks to the target database interface, potentially involving command conversion such as NL2SQL. Asynchronous call mechanisms and fault tolerance are introduced at this layer to improve search efficiency and stability. Level 9: Database Layer (DB), which accepts query instructions and completes queries. Level 10: Table Layer, which stores data in the form of tables. Level 11: Data Layer (Data), which stores specific fields and data entries, which may be natural language, numerical, or other multimodal data. This eleven-level structure clearly defines the responsibilities of each stage in the search chain. Through layer-by-layer advancement and feedback, it ensures efficient and reliable search.
[0162] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of the principle of the multi-agent collaborative data query method provided in the embodiment of the present application Figure 2 , Figure 7 The scheduler described above is shown, that is, Figure 6 The structures of the first scheduling model, the second scheduling model, the third scheduling model and the fourth scheduling model are shown. The observed state in the hierarchical MDP is . It consists of a Encoding plus, three-head strategy network, together constitutes Figure 7 The following components are the output of the scheduler and can control the behavior of this layer: routing header : In the candidate subtask set The above gives a discrete probability distribution, which is used to select the next hop A2A / Agent / API and select the action. Activator-MDP Top-k head : Integer budget allocation, used to limit the number of branches activated simultaneously in this layer . Termination header : Bernoulli probability, controls whether to terminate the search chain in advance at this level, if two consecutive steps have an advantage , then The probability ends prematurely terminating the chain.
[0163] In some embodiments, for the training of the scheduler, the immediate reward consists of three parts: , in Represents improved accuracy, range ; Represents the number of new branches triggered; represents the new API call cost. The optimization problem with the early stopping mechanism is ,in Update Lagrange multipliers , thus ensuring the second constraint Always stuck on the budget line nearby.
[0164] It is understandable that in the embodiments of the present application, when data related to data query tasks is involved and is applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0165] The following continues to describe the exemplary structure of the multi-agent collaborative data query device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2As shown, the software modules stored in the multi-agent collaborative data query device 455 of the memory 450 may include: a splitting module, configured to construct a data query task corresponding to the data query instruction in response to a received data query instruction, and split the data query task based on a first scheduling model to obtain multiple sub-data query tasks of the data query task; a scheduling module for determining, for each sub-data query task, a target data query service corresponding to the sub-data query task from a plurality of data query services based on a second scheduling model, and determining, through the target data query service, a target data interface service corresponding to the sub-data query task from a plurality of data interface services corresponding to the target data query service based on a third scheduling model; and determining, through the target data interface service, a target data interface corresponding to the sub-data query task from a plurality of data interfaces corresponding to the target data interface service based on a fourth scheduling model; A query module, configured to execute the sub-data query task corresponding to each target data interface through the target data interface corresponding to each sub-data query task, and obtain the sub-data query result corresponding to each sub-data query task; The fusion module is used to fuse the sub-data query results corresponding to the multiple sub-data query tasks to obtain a data query result.
[0166] In some embodiments, the data query instruction carries at least one query statement, and the above-mentioned splitting module is also used to query the index entry including the query statement from the preset statement-execution logic mapping relationship for each query statement, and determine the execution logic in the index entry as the target execution logic corresponding to the query statement; through the task construction model, based on the target execution logic corresponding to each query statement, generate a data query task corresponding to the data query instruction.
[0167] In some embodiments, the above-mentioned splitting module is also used to predict, through the first scheduling model and based on the data query instruction, a first quantity of the multiple sub-data query tasks and a first probability indicating that the data query task can be executed; compare the first probability with the probability threshold to obtain a first comparison result; if the first comparison result indicates that the first probability is greater than the probability threshold, then based on the first quantity, split the data query task into the multiple sub-data query tasks.
[0168] In some embodiments, the data query task includes the target execution logic corresponding to each query statement carried in the data query instruction. The above-mentioned splitting module is also used to determine the second number of the target execution logic in the data query task, and compare the second number with the first number to obtain a second comparison result; if the second comparison result indicates that the first number is less than or equal to the second number, the data query task is split into the first number of sub-data query tasks; if the second comparison result indicates that the first number is greater than the second number, the data query task is split into the second number of sub-data query tasks, and the sub-data query task includes at least one of the target execution logic.
[0169] In some embodiments, the first scheduling model includes a feature extraction layer, a first prediction layer, and a second prediction layer. The above-mentioned splitting module is also used to extract features of the data query instruction through the feature extraction layer to obtain query instruction features; through the first prediction layer, based on the query instruction features, predict the first number of the multiple sub-data query tasks; through the second prediction layer, based on the query instruction features, predict a first probability indicating that the data query task can be executed.
[0170] In some embodiments, the above-mentioned scheduling module is also used to predict, through the second scheduling model and based on the sub-data query task, a first matching degree between the sub-data query task and each of the data query services, and a second probability indicating that the sub-data query task can be executed; compare the second probability with a probability threshold to obtain a second comparison result; if the second comparison result indicates that the second probability is greater than the probability threshold, then determine the target data query service corresponding to the sub-data query task from the multiple data query services based on the first matching degree.
[0171] In some embodiments, the second scheduling model includes a feature extraction layer, a first prediction layer and a second prediction layer. The above-mentioned scheduling module is also used to perform feature extraction on the sub-data query task through the feature extraction layer to obtain sub-query task features; through the first prediction layer, based on the sub-query task features, predict the first matching degree of the sub-data query task with each of the data query services; through the second prediction layer, based on the sub-query task features, predict a second probability indicating that the sub-data query task can be executed.
[0172] In some embodiments, the above-mentioned scheduling module is also used to predict, through the third scheduling model and based on the sub-data query task, a second matching degree between the sub-data query task and each of the data interface services, and a third probability indicating that the sub-data query task can be executed; compare the third probability with the probability threshold to obtain a third comparison result; if the third comparison result indicates that the third probability is greater than the probability threshold, then determine, through the target data query service and based on the second matching degree, from the multiple data interface services, the target data interface service corresponding to the sub-data query task.
[0173] In some embodiments, the above-mentioned scheduling module is also used to predict, through the fourth scheduling model and based on the sub-data query task, a third matching degree between the sub-data query task and each of the data interfaces, and a fourth probability indicating that the sub-data query task can be executed; compare the fourth probability with the probability threshold to obtain a fourth comparison result; if the fourth comparison result indicates that the fourth probability is greater than the probability threshold, then determine, through the target data interface service, based on the third matching degree, from the multiple data interfaces, the target data interface corresponding to the sub-data query task.
[0174] The present application provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to execute the multi-agent collaborative data query method described above in the present application.
[0175] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the multi-agent collaborative data query method provided by the embodiment of the present application, for example, Figure 3 The multi-agent collaborative data query method shown.
[0176] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various electronic devices including one or any combination of the above memories.
[0177] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0178] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in an HTML document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0179] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0180] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A multi-agent collaborative data query method, characterized in that: The method comprises: In response to the received data query instruction, construct a data query task corresponding to the data query instruction, and split the data query task to obtain multiple sub-data query tasks of the data query task; For each of the sub-data query tasks, determining a target selection service for the sub-data query task from a plurality of agent selection services based on the first scheduling model; For each of the sub-data query tasks, based on the second scheduling model, a target data query service corresponding to the sub-data query task is determined from a plurality of data query services corresponding to the target selection service of the sub-data query task; and through the target data query service, based on the third scheduling model, a target data interface service corresponding to the sub-data query task is determined from a plurality of data interface services corresponding to the target data query service; and through the target data interface service, based on the fourth scheduling model, a target data interface corresponding to the sub-data query task is determined from a plurality of data interfaces corresponding to the target data interface service; Executing the sub-data query tasks corresponding to each target data interface through the target data interface corresponding to each sub-data query task, and obtaining the sub-data query results corresponding to each sub-data query task; The sub-data query results corresponding to the multiple sub-data query tasks are merged to obtain a data query result.
2. The method according to claim 1, characterized in that The data query instruction carries at least one query statement, and the step of constructing a data query task corresponding to the data query instruction in response to the received data query instruction includes: For each query statement, query the index entry including the query statement from a preset statement-execution logic mapping relationship, and determine the execution logic in the index entry as the target execution logic corresponding to the query statement; By constructing a task model, based on the target execution logic corresponding to each query statement, a data query task corresponding to the data query instruction is generated.
3. The method according to claim 1, characterized in that The step of splitting the data query task to obtain multiple sub-data query tasks of the data query task includes: Predicting, based on the data query instruction, a first number of the plurality of sub-data query tasks and a first probability indicating that the data query task can be executed; The first probability is compared with a probability threshold to obtain a first comparison result. If the first comparison result indicates that the first probability is greater than the probability threshold, the data query task is split into the multiple sub-data query tasks based on the first quantity.
4. The method according to claim 3, characterized in that The data query task includes the target execution logic corresponding to each query statement carried in the data query instruction, and the data query task is split into the multiple sub-data query tasks based on the first quantity, including: Determine a second number of the target execution logics in the data query task, and compare the second number with the first number to obtain a second comparison result; If the second comparison result indicates that the first number is less than or equal to the second number, splitting the data query task into the first number of sub-data query tasks; If the second comparison result indicates that the first number is greater than the second number, the data query task is split into the second number of sub-data query tasks, each of which includes at least one of the target execution logics.
5. The method according to claim 3, characterized in that The first scheduling model includes a feature extraction layer, a first prediction layer, and a second prediction layer. The first scheduling model predicts, based on the data query instruction, a first number of the plurality of sub-data query tasks and a first probability indicating that the data query task can be executed, including: Performing feature extraction on the data query instruction through the feature extraction layer to obtain query instruction features; Predicting, by the first prediction layer, a first number of the plurality of sub-data query tasks based on the query instruction characteristics; Through the second prediction layer, based on the query instruction characteristics, a first probability is predicted to indicate that the data query task can be executed.
6. The method according to claim 1, characterized in that The determining, based on the second scheduling model, a target data query service corresponding to the sub-data query task from a plurality of data query services corresponding to the target selection service of the sub-data query task includes: Based on the sub-data query task, the second scheduling model predicts a first matching degree between the sub-data query task and each of the data query services, and a second probability indicating that the sub-data query task can be executed; The second probability is compared with a probability threshold to obtain a second comparison result. If the second comparison result indicates that the second probability is greater than the probability threshold, a target data query service corresponding to the sub-data query task is determined from the multiple data query services based on the first matching degree.
7. The method according to claim 6, characterized in that The second scheduling model includes a feature extraction layer, a first prediction layer, and a second prediction layer. The second scheduling model predicts, based on the sub-data query task, a first matching degree between the sub-data query task and each of the data query services, and a second probability indicating that the sub-data query task can be executed, including: Performing feature extraction on the sub-data query task through the feature extraction layer to obtain sub-query task features; Predicting, by the first prediction layer, a first matching degree between the sub-data query task and each of the data query services based on the sub-query task characteristics; Through the second prediction layer, based on the sub-query task characteristics, a second probability indicating that the sub-data query task can be executed is predicted.
8. The method according to claim 1, characterized in that The determining, by the target data query service and based on a third scheduling model, a target data interface service corresponding to the sub-data query task from a plurality of data interface services corresponding to the target data query service includes: Based on the sub-data query task, the third scheduling model is used to predict a second matching degree between the sub-data query task and each of the data interface services, and a third probability indicating that the sub-data query task can be executed; The third probability is compared with the probability threshold to obtain a third comparison result. If the third comparison result indicates that the third probability is greater than the probability threshold, the target data interface service corresponding to the sub-data query task is determined from the multiple data interface services through the target data query service based on the second matching degree.
9. The method according to claim 1, characterized in that Determining, by the target data interface service and based on a fourth scheduling model, a target data interface corresponding to the sub-data query task from a plurality of data interfaces corresponding to the target data interface service includes: Predicting, using the fourth scheduling model and based on the sub-data query task, a third degree of matching between the sub-data query task and each of the data interfaces, and a fourth probability indicating that the sub-data query task is executable; The fourth probability is compared with the probability threshold to obtain a fourth comparison result. If the fourth comparison result indicates that the fourth probability is greater than the probability threshold, the target data interface corresponding to the sub-data query task is determined from the multiple data interfaces through the target data interface service based on the third matching degree.
10. A multi-agent collaborative data query device, characterized in that: The device comprises: a splitting module configured to construct, in response to a received data query instruction, a data query task corresponding to the data query instruction, and split the data query task into a plurality of sub-data query tasks of the data query task; and determine, for each sub-data query task, a target selection service for the sub-data query task from a plurality of agent selection services based on a first scheduling model; a scheduling module for determining, for each sub-data query task, a target data query service corresponding to the sub-data query task from a plurality of data query services corresponding to a target selection service of the sub-data query task based on a second scheduling model; and determining, through the target data query service, a target data interface service corresponding to the sub-data query task from a plurality of data interface services corresponding to the target data query service based on a third scheduling model; and determining, through the target data interface service, a target data interface corresponding to the sub-data query task from a plurality of data interfaces corresponding to the target data interface service based on a fourth scheduling model; A query module, configured to execute the sub-data query task corresponding to each target data interface through the target data interface corresponding to each sub-data query task, and obtain the sub-data query result corresponding to each sub-data query task; The fusion module is used to fuse the sub-data query results corresponding to the multiple sub-data query tasks to obtain a data query result.
Citation Information
Patent Citations
Method and device for realizing interactive message sending based on large language model
CN118869649A
Data query method and related equipment
CN119336799A
Question and answer method based on thinking chain and intelligent agent and ChatBI system
CN120012760A
Multi-agent recommendation method and device combined with preference learning, electronic equipment, storage medium and program product
CN120045697A
Data query method and device, medium and product
CN120196651A