Method and device for implementing a data weaving system based on dialogue and multi-agent collaboration

Through the data weaving system that collaborates between the main agent and the domain agent, the large language model is used for intention recognition and task disassembly, the problems of low efficiency and insufficient security of the data weaving system in the existing technology are solved, and efficient and intelligent enterprise-level data management is achieved.

CN119537460BActive Publication Date: 2025-08-15HANGZHOU EASTCOM SOFTWARE TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411575224.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-08-15
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

Existing data weaving systems are inefficient and insecure when processing large-scale, multi-source heterogeneous data, and are cumbersome to expand and maintain, and lack automation and intelligence capabilities, making it difficult to meet enterprise-level data management needs.

Method used

A data weaving system based on dialogue and multi-agent collaboration is adopted, and data retrieval, calculation and fusion are realized through precise division of labor and efficient collaboration between the main agent and the domain agent, intent recognition and task disassembly are used to combine data retrieval, data computing agent and data governance agent for data management.

Benefits of technology

It simplifies the implementation process of data braiding system, improves the efficiency and intelligence of enterprise-level data braiding, ensures data security and consistency, supports easy expansion and adjustment, and improves system flexibility and fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537460B_ABST
    Figure CN119537460B_ABST
Patent Text Reader

Abstract

The present application provides a method for implementing a data weaving system based on dialogue and multi-agent collaboration, including a main agent identifying user intent, disassembling and allocating data processing tasks according to user intent, issuing data retrieval tasks to a data retrieval agent in a domain agent, the data retrieval agent executing the data retrieval task on the managed data domain, obtaining retrieval results, the main agent issuing data computing tasks to a data computing agent in a domain agent, the data computing agent executing the data computing task, obtaining a result data set, the main agent executing a data fusion task on the retrieval results and / or result data set, and returning the fused data set to the user. In the present invention, the precise division of labor and efficient collaboration mechanism between the main agent and the domain agent in the multi-agent system simplifies the implementation process of the data weaving system, makes the system easy to expand, and realizes the efficiency and intelligence of the enterprise-level data weaving process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of network technology, and more particularly, to a method and apparatus for implementing a data weaving system based on dialogue and multi-agent collaboration. Background Art

[0002] With the continuous advancement of informatization across various industries, enterprises and organizations have accumulated a variety of heterogeneous data, including structured and unstructured data. This data has become a vital asset for enterprises and organizations. The key to effectively managing and utilizing this data lies in breaking down data silos within enterprises and organizations and promoting rapid data integration and interoperability. Traditional data management methods face challenges such as low processing efficiency and insufficient security when processing large-scale, multi-source, and heterogeneous data. Furthermore, centralized data supply models also face challenges such as massive resource requirements, real-time computing bottlenecks, and poor adaptability to diverse data application scenarios.

[0003] To address these challenges, data weaving has emerged as an emerging cross-platform data integration and management concept. It focuses on data connectivity, integration, and unified views, aiming to enhance the seamless flow and integration of data. However, the implementation of current data weaving systems remains complex, requiring the construction of a global knowledge graph and data federation. Expansion and maintenance are cumbersome, and insufficient automation and intelligence capabilities limit their effectiveness in practical applications. Summary of the Invention

[0004] This application describes a method and device for implementing a data weaving system based on dialogue and multi-agent collaboration, which can solve the above technical problems.

[0005] According to a first aspect, a method for implementing a data weaving system based on dialogue and multi-agent collaboration is provided. The data weaving system includes a main agent and at least one domain agent. A domain agent is a collection of sub-agents that manage data assets of a data domain. A domain agent includes a data retrieval agent and a data calculation agent.

[0006] The method comprises:

[0007] The master agent identifies user intent and, based on the user intent, decomposes and allocates data processing tasks, wherein the decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks;

[0008] The master agent issues a data retrieval task to the data retrieval agent in the domain agent, wherein the data retrieval task is used to obtain storage information of the field to be retrieved in the data domain;

[0009] The data retrieval agent performs a data retrieval task on the managed data domain, obtains a retrieval result, and the retrieval result includes the table structure where the field to be retrieved is located and the storage information in the table structure, and returns the retrieval result to the master agent;

[0010] The master agent issues a data computing task to the data computing agent in the domain agent, wherein the data computing task is used to perform calculations based on the search results and / or one or more fields in the data domain;

[0011] The data computing agent executes the data computing task to obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or the fields in the data domain, and returns the result data set to the master agent;

[0012] The master agent performs a data fusion task on the retrieval results and / or the result data set, and returns the fused data set to the user.

[0013] According to a second aspect, a data weaving system based on dialogue and multi-agent collaboration is provided, wherein the data weaving system includes a main agent and at least one domain agent, wherein a domain agent is a collection of sub-agents that manage a data domain;

[0014] The main agent is used to identify user intentions and, based on the user intentions, to decompose and allocate data processing tasks, wherein the decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks;

[0015] The master agent is further configured to issue a data retrieval task to the data retrieval agent in the domain agent, wherein the data retrieval task is configured to obtain storage information of the field to be retrieved in the data domain;

[0016] The data retrieval agent is used to perform data retrieval tasks on the managed data domain, obtain retrieval results, including the table structure where the field to be retrieved is located and the storage information in the table structure, and return the retrieval results to the main agent;

[0017] The master agent is used to issue data computing tasks to the data computing agent in the domain agent, wherein the data computing tasks are used to perform calculations based on the search results and / or one or more fields in the data domain;

[0018] The data computing agent is configured to execute the data computing task, obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or the fields in the data domain, and return the result data set to the master agent;

[0019] The main agent is used to perform a data fusion task on the retrieval results and / or the result data set, and return the fused data set to the user.

[0020] According to a third aspect, a device for implementing a data weaving system based on dialogue and multi-agent collaboration is provided. The data weaving system includes a main agent and at least one domain agent. A domain agent is a collection of sub-agents that manage data assets of a data domain. A domain agent includes a data retrieval agent and a data calculation agent.

[0021] The first processing device is used for the master agent to identify user intent and, based on the user intent, to decompose and allocate data processing tasks, wherein the decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks; the master agent issues data retrieval tasks to the data retrieval agents in the domain agents, wherein the data retrieval tasks are used to obtain storage information of the fields to be retrieved in the data domain;

[0022] a second processing device configured to cause the data retrieval agent to perform a data retrieval task on the managed data domain, obtain a retrieval result including the table structure where the to-be-retrieved field is located and information stored in the table structure, and return the retrieval result to the master agent;

[0023] a third processing device, configured for the master agent to issue a data computing task to a data computing agent in a domain agent, wherein the data computing task is configured to perform calculations based on the search results and / or one or more fields in the data domain;

[0024] a fourth processing device configured to cause the data computing agent to execute the data computing task, obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or fields in the data domain, and return the result data set to the master agent;

[0025] The fifth processing device is used for the main agent to perform a data fusion task on the retrieval results and / or the result data set, and return the fused data set to the user.

[0026] In the above-mentioned system and method provided in the embodiments of this specification, the implementation process of the data weaving system is simplified by utilizing the precise division of labor and efficient collaboration mechanism of the main agent and domain agent in the multi-agent system. The data weaving system has good scalability and realizes the efficiency and intelligence of the enterprise-level data weaving process. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0028] Figure 1 A system diagram of a data weaving system based on dialogue and multi-agent collaboration provided by an embodiment of this specification is shown;

[0029] Figure 2 A schematic diagram of a master agent in a data weaving system based on dialogue and multi-agent collaboration provided by an embodiment of this specification is shown;

[0030] Figure 3 A schematic diagram of a data retrieval agent in a data weaving system based on dialogue and multi-agent collaboration provided by an embodiment of this specification is shown;

[0031] Figure 4 A schematic diagram of a data computing agent in a data weaving system based on dialogue and multi-agent collaboration provided by an embodiment of this specification is shown;

[0032] Figure 5 A schematic diagram of a data governance agent in a data weaving system based on dialogue and multi-agent collaboration provided by an embodiment of this specification is shown;

[0033] Figure 6 A schematic diagram of the process of agent NL2SQL in a data weaving system based on dialogue and multi-agent collaboration provided by an embodiment of this specification is shown;

[0034] Figure 7 A schematic diagram showing a flow chart of a data weaving system based on dialogue and multi-agent collaboration provided by an embodiment of this specification;

[0035] Figure 8 A flow chart showing a method for implementing a data weaving system based on dialogue and multi-agent collaboration provided in an embodiment of this specification is shown;

[0036] Figure 9 A schematic diagram showing an implementation device of a data weaving system based on dialogue and multi-agent collaboration provided in an embodiment of this specification. DETAILED DESCRIPTION

[0037] The solution provided in this specification is described below in conjunction with the accompanying drawings.

[0038] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0039] As a data management architecture, the Data Weaving system provides enterprises with a way to effectively manage and utilize their data assets, enabling the connection, integration, and unified view of enterprise data. The Data Weaving system strives to provide consistent interfaces and services across diverse data sources, enabling data to flow seamlessly between systems. It leverages artificial intelligence and machine learning technologies to automatically discover data relationships, conduct data governance, and optimize data access paths. It ensures data security and privacy while meeting various regulatory compliance requirements. Furthermore, it can be easily expanded and adapted as data volumes grow or business needs change.

[0040] To achieve these goals, the construction of a data weaving system relies on a variety of underlying technologies and tools. This involves a variety of existing technologies and new development methods, making technical implementation challenging. Currently, implementing a data weaving system requires building a global knowledge graph. Building a global knowledge graph requires organizing and standardizing data from various sources, while ensuring the accuracy and consistency of this data. Data federation is also necessary to enable cross-domain data sharing and analysis without moving data, which places higher demands on data access control and privacy protection. As data volumes grow or business needs change, expansion and maintenance become cumbersome, and insufficient automation and intelligence capabilities limit its effectiveness in practical applications.

[0041] Depend on Figure 1 The system diagram of the data weaving system based on dialogue and multi-agent collaboration shows that the present invention realizes the data weaving system through a master agent and multiple domain agents. The data weaving system includes a master agent and a domain agent. Each domain agent autonomously manages a data domain or data center, and the master agent and the domain agent adopt a master-slave collaboration mode to achieve the fusion of target data. The master agent receives user requests in a dialogue mode, performs intent recognition, task decomposition and allocation, result fusion and status monitoring, etc., to ensure that the system collaborates and operates efficiently as planned. The domain agent is a collection of agents composed of multiple sub-agents, which receives task instructions assigned by the master agent and feeds back the results. The domain agent includes a data retrieval agent, a data calculation agent and a data governance agent.

[0042] Depend on Figure 2 The schematic diagram of the main agent in the data weaving system based on dialogue and multi-agent collaboration shows that the main agent receives user requests through natural language dialogue mode, uses large models and RAG retrieval enhancement generation technology to perform intention recognition, task decomposition and allocation, data fusion and status monitoring, etc., and coordinates communication and collaboration between various agents, recovers from abnormal situations, and ensures the efficient and stable operation of the entire system.

[0043] The main agent uses a large language model to perform multi-level semantic analysis on natural language data requests in conversational mode. Through contextual analysis and reasoning, it identifies user intent, such as the specific scope of data retrieval, the details of the required content, and whether data fusion calculations are involved.

[0044] The main agent determines the task to be performed based on the identified user intention and assigns the task to the corresponding agent according to the task type. For example, if the task type is data retrieval, the task is assigned to the data directory search agent; if the task type is data calculation, the task is assigned to the data calculation agent. At the same time, combined with the asset directory and data permissions maintained by each data domain, the security of task assignment and data access is ensured. The main agent needs to verify the permissions requested by the user and control data access and operations based on the permissions.

[0045] The master agent coordinates the communication between the agents in each domain, manages the communication protocols between different agents, and ensures the security and reliability of data transmission. The communication between the agents in each domain includes message passing, data sharing, and task collaboration, etc., to realize information exchange and collaborative work between the agents in each domain.

[0046] The master agent integrates a fusion data warehouse to cache the result data processed by each domain agent. Based on the user's intention and the result data received from the domain agent, it determines the data fusion algorithm, executes the fusion calculation task, and returns the fused result data to the user.

[0047] The main agent monitors the system status in real time and detects failures or abnormal behaviors of domain agents. Once a failure is detected, it reallocates tasks or starts backup agents to ensure the stable operation of the system. Through various fault-tolerant strategies such as task retry, task transfer, and backup agent startup, it optimizes resource allocation to cope with different types of failure situations and enhance system robustness.

[0048] The main agent receives the task execution results from the domain agent, which include various information such as task completion status, data processing results, and external interface calls. The main agent evaluates the task execution results to determine whether the goals have been achieved as expected. If the evaluation results show that the task has not met expectations, the main agent activates the reflection mechanism to analyze potential problems in the execution process. The main agent adjusts its strategy or optimizes the use of tools accordingly to improve the efficiency and accuracy of subsequent task execution. Specifically, the reflection mechanism is a method of self-assessment and self-improvement used to analyze potential problems in the execution process and optimize future actions. It has a wide range of applications in fields such as artificial intelligence, robotics, and software engineering, including the process of collecting data, analyzing data, evaluating results, generating reports, formulating improvement plans, implementing improvements, and iterating cycles.

[0049] A domain agent is a collection of sub-agents built for a specific data domain. Each sub-agent in the sub-agent collection receives task instructions from the main agent. An enterprise usually has multiple data domains or data centers, such as the production data domain, sales data domain, financial data domain, and human resources data domain. A corresponding domain agent is built for each data domain or data center. When cross-domain business scenarios are involved, the relevant domain agents and the main agent work together in a master-slave distributed model. Based on the user's intention, the main agent clarifies the collaboration strategy between the domain agents, disassembles and assigns orchestration tasks to each domain agent, and orchestrates or integrates the feedback results of each domain agent and outputs them to the user. Each domain agent receives and executes the assigned tasks, and feeds back the execution results of the domain agent to the main agent.

[0050] When adding a new data domain, you only need to add the corresponding domain agent and incorporate it into the main agent for collaboration, making it easy to expand and adjust.

[0051] like Figure 3 The data retrieval agent in a data weaving system based on dialogue and multi-agent collaboration demonstrates that the data retrieval agent receives query intent from the master agent, translates the intent, searches and retrieves multiple data sources within the managed data domain, and returns the results to the master agent. The data retrieval agent builds table and view structures based on the characteristics and capacity constraints of the large model. Specifically, the core, associated, and transaction tables within the table structure are designed taking into account factors such as the large model's computing resource requirements, data processing speed, and data consistency. The view structure simplifies complex query logic and provides easy-to-understand and easy-to-use data views. Simplifying the table and view structures involves removing redundant fields, using enumeration types for fields with limited options, standardizing data types, and reducing the number of associations between tables. Using these simplified table and view structures, the large model's prompt words are constructed.

[0052] Large model prompt words include:

[0053] 1. Table / view description;

[0054] 2. Field descriptions and attributes related to the user's question, including data type, granularity, sorting direction, and units;

[0055] 3. Table / view SQL statement specifications, such as filtering conditions for time or dimension information.

[0056] Sort out the business semantic knowledge system and build a business data semantic knowledge framework, where the business semantic system includes the above-mentioned large model prompt words.

[0057] The business data semantic knowledge framework is uniformly vectorized to obtain the data semantic layer knowledge base.

[0058] Based on user intent and the data semantic layer knowledge base, vector retrieval and recall operations are performed to extract relevant information from the data semantic layer knowledge base. The information is optimized and sorted using re-ranking technology to achieve fast and accurate search and retrieval in multiple data sources.

[0059] It's important to note that due to the capacity constraints of large models in multi-table joins, to improve NL2SQL's accuracy, we've built a large model prompting project. By simplifying table and view structures, we can transform complex join problems into a single task focused solely on the user question and table field extraction. Prompts include: table / view descriptions; field descriptions and attributes (data type, granularity, sorting direction, units, etc.) related to the user question; and table / view SQL statement specifications, such as filtering conditions for time and dimension information.

[0060] like Figure 6 As shown, the specific NL2SQL algorithm flow of the large model is as follows:

[0061] A. Users enter query requests using natural language.

[0062] B. Use advanced semantic understanding models (such as pre-trained language models) to parse user input and accurately identify user intent.

[0063] C. For complex questions containing multiple query elements, they should be split into several single query intents to achieve more accurate understanding and processing.

[0064] For situations where the intent depends on the context to be fully understood, coreference resolution or lexical replacement operations are required to ensure that the intent can be accurately parsed.

[0065] Standardize synonyms and terminology within the context of your specific business area to ensure consistency and accuracy.

[0066] Select the most appropriate semantic library based on context relevance.

[0067] D. Based on the business type, extract the core indicators and keywords to be queried, and load the semantic layer definition of the indicators or fields.

[0068] Through the SQL template matching mechanism, vector retrieval is used to identify the "Few-shot" instances that are closest to the existing sample SQL structure.

[0069] If a specific complex business scenario is encountered, query guidance suitable for the scenario will be retrieved from the predefined prompt word template set.

[0070] E. With the help of the code generation model, generate accurate SQL query statements based on the previous intent parsing and template matching results.

[0071] F. Ensure that the generated SQL code is valid and executable. Through test cases or database-related verification, generate the final executable SQL query and present the query results.

[0072] like Figure 4 The data computing agent in the data weaving system based on dialogue and multi-agent collaboration shows that the data computing agent receives the computing task issued by the main agent. The task information will carry the user's intention, the required data asset identification, etc. The data computing agent will generate detailed information of the task according to the user's intention, such as: data asset identification, calculation dimensions (time dimension, regional dimension, other dimensions, etc.), calculation indicators, etc.

[0073] According to the data asset identifier, data asset information is obtained from the domain's data management system, such as: data source type, data source connection method, data asset magnitude, data asset metadata information, etc.

[0074] Collect relevant data of the original data that meets the conditions from the data source, such as: data volume, maximum / minimum value of indicators, number of dimensions, etc.

[0075] Considering the order of magnitude, data availability, and data source type, determine whether downstream computation is necessary, as well as the size of the original and result data. Depending on the situation, determine whether a distributed computing engine (such as Spark) is necessary. For data sources where downstream computation is not feasible, and where the original or result data volumes are relatively large, a distributed computing engine is used. In this case, the computing agent generates Spark SQL using NL2SQL technology based on user intent and detailed task information, and schedules the Spark program for computation.

[0076] If the data source type can be downlinked (such as MySQL, DAMO, ClickHouse), the computing intelligence will use NL2SQL to generate standard SQL for that data source type based on the user's intent and detailed task information, and downlink the data calculation to the data source for calculation. If the data source type is other NoSQL (such as Redis), a dedicated natural language conversion command model similar to NL2SQL (such as NL2Redis) will be used to generate the corresponding syntax and perform the relevant data calculation.

[0077] The data agent receives the data calculation results and, based on the size of the results, decides whether to return the results directly or output them to the fusion database and return the relevant information to the main agent. If the data volume is large, it is output to the fusion database; if the data volume is small, it is directly returned to the main agent.

[0078] The data computing agent integrates a general large language model and a code large language model, building upon NL2SQL (Natural Language to SQL) technology. This agent automatically identifies data sources, analyzes data formats and structures, performs SQL joins based on the semantics of target tables and fields, and dynamically selects and adjusts data processing strategies based on demand. Through self-learning and optimization mechanisms, it efficiently executes data retrieval, processing, and computing tasks in complex and ever-changing data environments.

[0079] During processing, different strategies are adopted for heterogeneous data. For structured data, SQL queries or DataFrame operations (such as Pandas and Spark DataFrame) are used for processing. For unstructured data, natural language processing (NLP) technologies such as text classification, sentiment analysis, and entity recognition are used. For semi-structured data, appropriate parsing and transformation technologies such as JSON parsing or XML parsing are selected.

[0080] To ensure timely processing, we dynamically adjust processing strategies based on data characteristics. For large data volumes, we use distributed computing frameworks (such as Spark) for efficient processing. For real-time data streams, we use stream processing frameworks (such as Flink) for immediate processing. After processing is complete, the resulting data is aggregated and output.

[0081] like Figure 5 The data governance agent in the data weaving system based on dialogue and multi-agent collaboration shows that the data governance agent uses the task splitting and tool calling capabilities of the large model to monitor data quality and security, including using algorithms and rules to check whether there are missing values, duplicate values, format errors, and sensitive data in important data. Through data association analysis, it verifies whether the logical relationship between data is correct, and determines through security policies whether sensitive data has been desensitized or encrypted, and whether users have data access rights.

[0082] The data governance agent handles data quality and security issues, including format conversion or filling in missing values for issues that can be automatically fixed, such as data format errors or missing data; automatically desensitizing or encrypting the output sensitive data according to security policies when sensitive data is not desensitized or encrypted; providing feedback of insufficient permissions for data without access permissions; and providing feedback of potential data risks for complex quality or security issues and notifying the administrator for manual processing.

[0083] The data governance agent conducts real-time monitoring and continuous improvement, including real-time monitoring of data input and output, detecting fluctuations in data volume, and optimizing abnormal fluctuations through a combination of manual and intelligent methods to ensure the stability and regularity of data flow. It also detects the presence of sensitive data, ensures that it has been processed in accordance with security policies, and audits user access rights to ensure data security. This process enables comprehensive management and continuous maintenance of data quality and security, enhancing data reliability, availability, and security.

[0084] like Figure 7 As shown, the process of weaving data of the conversation request includes the following steps:

[0085] Step 1: The master agent receives the user request.

[0086] In step 2, the main agent splits the user request into one or more tasks through intent recognition and task decomposition, and determines the relevant tasks of each domain agent, including data retrieval tasks, data calculation tasks, and data governance tasks.

[0087] Step 3: The master agent issues the corresponding data retrieval task to the data retrieval agent in the domain agent to which the data retrieval task needs to be issued.

[0088] Step 4: The data retrieval agent performs data retrieval tasks on the managed data domain.

[0089] Step 5: If the retrieval result is obtained, the data retrieval agent returns the data retrieval result to the main agent.

[0090] Step 6: The master agent sends the corresponding data computing task to the data computing agent in the domain agent to which the data computing task needs to be sent.

[0091] Step 7: The data computing agent extracts data from the data domain to perform computing tasks.

[0092] Step 8: After the data computing agent completes the computing task, if there is no need to trigger the data governance task, it returns the data computing result to the main agent;

[0093] If a data governance task needs to be triggered, the data governance agent will be triggered to check the quality and security of the data, and perform corresponding desensitization, encryption or permission control on the data according to the existing quality rules and security policies.

[0094] Step 9: The data governance agent feeds back the processing results to the main agent.

[0095] In step 10, the master agent caches the data of the domain agents that need to be fused into the fusion data warehouse, performs fusion calculations on the received relevant data, generates the final results and feeds them back to the user.

[0096] Step 11: The main intelligent agent desensitizes, encrypts or controls the permissions of the final result again according to the overall security strategy, and then feeds it back to the user.

[0097] This invention leverages the intelligent collaboration capabilities of the master agent, combined with the autonomous and intelligent management capabilities of domain agents and the parallel processing capabilities of these agents, to achieve rapid retrieval of enterprise-wide data, efficient and intelligent data computation and governance, significantly improving the overall effectiveness of data weaving. The master agent is also responsible for monitoring the operating status of each domain agent, dynamically adjusting tasks based on actual conditions, achieving load balancing, and handling and recovering from failures, ensuring stable and efficient system operation.

[0098] By assigning tasks from the master agent to each domain agent, distributed task processing is achieved. Each domain agent, relying on its proactive and intelligent management of local data, can quickly provide feedback on processing results, effectively reducing data retrieval latency, improving the system's overall response speed and operational efficiency, and enhancing system flexibility.

[0099] Each enterprise's data domain autonomously manages its asset catalog and data permissions through its own dedicated domain agent. This eliminates the need to aggregate enterprise data assets into a centralized catalog or knowledge graph, or synchronize individual security policies, reducing complexity. This approach also enables each data domain to independently update and maintain its data assets and allows them to flexibly control and output data according to their own security policies, ensuring real-time, accurate, and secure data.

[0100] Administrators of each enterprise's data domains gain a deeper understanding of their data assets and permissions, enabling more precise allocation and management of permissions. Fine-grained permission management of data within the domain ensures that only authorized users can access their data assets. This approach improves the accuracy and security of data permissions and avoids issues such as inaccurate or untimely permission allocation that can arise from centralized management.

[0101] In this invention, different sub-agents within the domain agent are specifically designed to perform distinct tasks. They leverage technologies like large models and RAG to complete the retrieval, computation, and governance of domain data. For example, the data retrieval agent is responsible for enhanced data retrieval, the data computation agent is responsible for intelligent data extraction and processing, and the data governance agent is responsible for efficient data governance. This specialized solution approach improves the domain agent's problem-solving capabilities and work efficiency.

[0102] The data computing agent in this invention integrates advanced large models and machine learning technologies, can automatically identify data source types and data formats, and then intelligently optimize data access paths, generate the most appropriate data processing strategies, and greatly improve data processing efficiency.

[0103] like Figure 8 It is shown that, according to an embodiment of the first aspect, a method for implementing a data weaving system based on dialogue and multi-agent collaboration is provided, wherein the data weaving system includes a main agent and at least one domain agent, a domain agent is a collection of sub-agents that manage a data domain, wherein a domain agent includes a data retrieval agent and a data calculation agent.

[0104] The following steps are involved:

[0105] 110. The main intelligent agent identifies the user's intention and decomposes and allocates the data processing tasks according to the user's intention. The decomposed data processing tasks include data retrieval tasks, data calculation tasks and data fusion tasks.

[0106] 120. The main agent sends a data retrieval task to the data retrieval agent in the domain agent. The data retrieval task is used to obtain the storage information of the field to be retrieved in the data domain.

[0107] 130. The data retrieval agent performs data retrieval tasks on the data domain it manages and obtains retrieval results. The retrieval results include the table structure where the field to be retrieved is located and the storage information in the table structure, and returns the retrieval results to the main agent.

[0108] 140. The master agent issues a data computing task to the data computing agent in the domain agent. The data computing task is used to perform calculations based on the retrieval results and / or one or more fields in the data domain.

[0109] 150. The data computing agent executes the data computing task and obtains a result data set, which includes a two-dimensional table structure calculated based on the search results or fields in the data domain, and returns the result data set to the main agent.

[0110] 160. The main agent performs a data fusion task on the retrieval results and / or result data sets, and returns the fused data sets to the user.

[0111] In some embodiments, step 130 specifically includes:

[0112] 131. The data retrieval agent builds the table structure and view structure of the data domain based on the large language model and simplifies the table structure and view structure;

[0113] 132. Use the simplified table structure and view structure to build large model prompt words, which are used to describe the table structure and view structure;

[0114] 133. Based on the big language model, organize the business semantic knowledge system of the data domain. The business semantic system is used to describe the business terminology information of the data domain, including big model prompt words;

[0115] 134. Use the business semantic knowledge system to build a business data semantic knowledge framework. The business data semantic knowledge framework includes metadata of the data domain, including the definition of table structure and field definitions.

[0116] 135. Perform unified knowledge vectorization on the business data semantic knowledge framework to obtain the data semantic layer knowledge base of the data domain. The data semantic layer knowledge base is a repository for managing and organizing metadata of the data domain.

[0117] 136. Based on the field information to be retrieved, a vector search and recall operation is performed on the data semantic layer knowledge base to obtain the table structure where the field to be retrieved is located and the storage information in the table structure.

[0118] Step 150 specifically includes:

[0119] 151. The data computing agent generates computing task information according to the user intention in the data computing task, and the computing task information includes the data asset identifier;

[0120] 152. According to the data asset identifier in the computing task information, obtain data asset information from the data domain. The data asset information includes the type of data source, the connection method of the data source, and the data volume of the data asset;

[0121] 153. When the amount of data in the data source is less than the preset amount of data and all the data involved in the calculation are in the data source, and the conditions for sinking to the data source for calculation are met, based on the user's intention and the calculation task information, a data query statement and / or a grammatical statement applicable to the data source type is generated, and the data calculation task, as well as the data query statement and / or grammatical statement applicable to the data source type, are sunk to the data source for calculation to obtain a result data set from the data source;

[0122] 154. When the amount of data in the data source exceeds the preset amount or other data in the domain is required for calculation, and the conditions for sinking to the data source for calculation are not met, it is necessary to rely on the high-performance computing engine of the data computing agent to perform data calculations. According to the user's intention and computing task information, a data query statement and / or a grammatical statement suitable for the computing engine is generated, and the data that meets the conditions is extracted to the computing engine. The computing engine then executes the data computing task and obtains the result data set.

[0123] 155. Return the result data set to the main agent.

[0124] In some embodiments, the domain agent also includes a data governance agent, and the data processing task also includes a data governance task;

[0125] Also includes:

[0126] The master agent issues data governance tasks to the data governance agent in the domain agent. Data governance tasks are used to manage and maintain the quality, security, and permissions of data in the data domain or result dataset.

[0127] The data governance agent performs data governance tasks on the managed data domain and / or result data set, obtains the governed result data set, and returns the governed result data set to the main agent.

[0128] In some embodiments, further comprising:

[0129] The data governance agent determines the data to be governed based on the data governance task. The data to be governed includes the data in the data domain or the result data set returned by the data computing agent.

[0130] Check whether there are any data quality issues in the data to be managed. Automatically repair any data quality issues that can be repaired, such as missing values, duplicate values, and / or format errors. For data quality issues that cannot be repaired automatically, feedback indicates that the data has hidden problems.

[0131] If sensitive data in the data to be governed is not desensitized or encrypted, desensitize or encrypt the sensitive data according to the security policy. For sensitive data to which you do not have access rights, provide a message indicating insufficient permissions.

[0132] Monitor the input and output of data in the data domain in real time, detect abnormal fluctuations in the amount of data in the data domain, and provide feedback on abnormal fluctuations.

[0133] In some embodiments, further comprising:

[0134] The main agent communicates with the user through natural language to obtain the user's intent, which includes the user's query requirements. The query requirements include the query time range, time granularity, data objects, and output requirements.

[0135] Decompose the user's query requirements into tasks and issue data retrieval tasks to the data retrieval agent in the domain agent;

[0136] The master agent receives the storage information of the to-be-retrieved field fed back by the data retrieval agent in the domain agent;

[0137] Obtaining fields related to the user's query requirements from among the to-be-retrieved fields fed back by the data retrieval agent;

[0138] Determine the relevant data domains where the fields related to the user's query requirements are located;

[0139] Generate an execution plan based on the user's query requirements;

[0140] According to the execution plan, data computing tasks are issued to the data computing agents in the relevant data domain.

[0141] In some embodiments, further comprising:

[0142] The master agent receives the result data set fed back by the data calculation agents from multiple domain agents and caches the result data set into the fusion data warehouse;

[0143] The result data set is extracted from the fusion data warehouse for fusion calculation to generate a fused result data set.

[0144] In some more specific embodiments, it further includes:

[0145] The data computing agent determines whether the data type of the search results or fields in the data domain is structured data and processes it using general database queries or data table DataFrame operations;

[0146] The data type of the fields in the search results or data domain is unstructured data and is processed using natural language processing methods;

[0147] The data type of the fields in the search results or data domain is semi-structured data, which is parsed and converted into structured data or unstructured data before processing.

[0148] According to another embodiment, there is also provided a data weaving system based on dialogue and multi-agent collaboration, the data weaving system including a main agent and at least one domain agent, wherein a domain agent is a collection of sub-agents that manage a data domain;

[0149] The main agent is used to identify user intent and, based on the user intent, to decompose and allocate data processing tasks. The decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks.

[0150] The master agent is also used to issue data retrieval tasks to the data retrieval agent in the domain agent. The data retrieval task is used to obtain the storage information of the field to be retrieved in the data domain;

[0151] The data retrieval agent is used to perform data retrieval tasks on the managed data domain, obtain retrieval results, which include the table structure of the field to be retrieved and the storage information in the table structure, and return the retrieval results to the main agent;

[0152] The master agent is used to issue data computing tasks to the data computing agents in the domain agent. The data computing tasks are used to perform calculations based on the search results or one or more fields in the data domain.

[0153] The data computing agent is used to perform data computing tasks, obtain a result data set, which includes a two-dimensional table structure calculated based on the search results or fields in the data domain, and return the result data set to the main agent;

[0154] The main agent is used to perform data fusion tasks on the retrieval results and / or result datasets and return the fused datasets to the user.

[0155] This embodiment utilizes intelligent agent technology, and users can make data weaving requests through a dialogue mode. The system has the ability to automatically identify user intentions and intelligently disassemble and reorganize the requests, thereby greatly improving the convenience of user use.

[0156] Leveraging multi-agent collaborative technology, the data weaving process is significantly simplified and optimized, eliminating the need for centralized global knowledge graph management and reducing the risk of data inconsistencies. Through close collaboration and a rational division of labor among multiple agents, tasks can be completed efficiently. This collaborative mechanism also enhances the fault tolerance of the data weaving process and effectively integrates and allocates various resources, ensuring the continuous and stable operation of the data weaving system.

[0157] Leveraging the agent's autonomy and intelligence, the domain agent can autonomously and intelligently manage the data catalog, data quality, and data security within its domain. It provides intelligent data-enhanced retrieval capabilities, automatically generates and executes SQL statements for data processing based on optimal strategies and NL2SQL technology to complete data computation tasks, and implements intelligent security strategies for data within the domain.

[0158] like Figure 9According to a third aspect, a device for implementing a data weaving system based on dialogue and multi-agent collaboration is provided. The data weaving system includes a main agent and at least one domain agent. A domain agent is a collection of sub-agents that manage data assets of a data domain. A domain agent includes a data retrieval agent and a data calculation agent.

[0159] The first processing device is configured to identify user intent through dialogue with the user and, based on the user intent, to decompose and allocate data processing tasks, wherein the decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks; and to issue data retrieval tasks to the data retrieval agent in the domain agent, wherein the data retrieval tasks are used to obtain the storage information of the fields to be retrieved in the data domain;

[0160] The second processing device is used for the data retrieval agent to perform a data retrieval task on the managed data domain, obtain a retrieval result, which includes the table structure where the field to be retrieved is located and the storage information in the table structure, and return the retrieval result to the main agent;

[0161] The third processing device is used for the master agent to issue data computing tasks to the data computing agent in the domain agent, and the data computing tasks are used to perform calculations based on the search results or one or more fields in the data domain;

[0162] A fourth processing device is configured to cause the data computing agent to execute the data computing task, obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or fields in the data domain, and return the result data set to the main agent;

[0163] The fifth processing device is used for the main agent to perform data fusion tasks on the retrieval results and / or result data sets, and return the fused data sets to the user.

[0164] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.

Claims

1. A method for implementing a data weaving system based on dialogue and multi-agent collaboration, characterized in that: The data weaving system includes a main agent and at least one domain agent, wherein a domain agent is a collection of sub-agents that manage data assets of a data domain, wherein a domain agent includes a data retrieval agent and a data calculation agent; The method comprises: The master agent identifies user intent and, based on the user intent, decomposes and allocates data processing tasks, wherein the decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks; The master agent issues a data retrieval task to the data retrieval agent in the domain agent, wherein the data retrieval task is used to obtain storage information of the field to be retrieved in the data domain; The data retrieval agent performs a data retrieval task on the managed data domain, obtains a retrieval result, and the retrieval result includes the table structure where the field to be retrieved is located and the storage information in the table structure, and returns the retrieval result to the master agent; The master agent issues a data computing task to the data computing agent in the domain agent, wherein the data computing task is used to perform calculations based on the search results and / or one or more fields in the data domain; The data computing agent executes the data computing task to obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or the fields in the data domain, and returns the result data set to the master agent; The master agent performs a data fusion task on the search results and / or the result data set, and returns the fused data set to the user; The main agent communicates with the user through natural language to obtain the user's intention, which includes the user's query requirements, including the query time range, time granularity, data objects and output requirements; Decompose the user's query requirements into tasks and issue data retrieval tasks to the data retrieval agent in the domain agent; The master agent receives the storage information of the to-be-retrieved field fed back by the data retrieval agent in the domain agent; Obtaining fields related to the user's query requirements from the to-be-retrieved fields fed back by the data retrieval agent; Determine the relevant data domains where the fields related to the user's query requirements are located, and at the same time, ensure the allocation of data calculation tasks in combination with the asset catalog maintained by each data domain; Generate an execution plan based on the user's query requirements; According to the execution plan, the data computing task is issued to the data computing agent in the relevant data domain.

2. The method according to claim 1, characterized in that The data retrieval agent performs a data retrieval task on the managed data domain, obtains a retrieval result, and includes the table structure of the field to be retrieved and the storage information in the table structure, and returns the storage information of the field to be retrieved to the main agent, specifically including: The data retrieval agent constructs a table structure and a view structure of the data domain based on a large language model, and simplifies the table structure and the view structure; Using the simplified table structure and view structure, construct a large model prompt word, wherein the large model prompt word is used to describe the table structure and view structure; Based on the large language model, sorting out the business semantic knowledge system of the data domain, wherein the business semantic system is used to describe the business terminology information of the data domain, including the large model prompt words; Constructing a business data semantic knowledge framework using the business semantic knowledge system, wherein the business data semantic knowledge framework includes metadata of the data domain, and the metadata includes definitions of the table structure and fields; Performing unified knowledge vectorization processing on the business data semantic knowledge framework to obtain a data semantic layer knowledge base of the data domain, wherein the data semantic layer knowledge base is a repository for managing and organizing metadata of the data domain; Based on the field information to be retrieved, a vector search and recall operation is performed on the data semantic layer knowledge base to obtain the table structure where the field to be retrieved is located and the storage information in the table structure.

3. The method according to claim 1, characterized in that The data computing agent executes the data computing task to obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or the fields in the data domain, and returns the result data set to the main agent, specifically including: The data computing agent generates computing task information according to the user intention in the data computing task, wherein the computing task information includes a data asset identifier; According to the data asset identifier in the computing task information, data asset information is obtained from the data domain, wherein the data asset information includes the type of data source, the connection method of the data source, and the data volume of the data asset; When the amount of data in the data source is less than a preset amount of data and the data involved in the calculation are all in the data source, and the conditions for being sunk to the data source for calculation are met, a data query statement applicable to the data source type and / or a grammatical statement applicable to the data source type are generated according to the user intention and the calculation task information, the data calculation task, and the data query statement applicable to the data source type and / or the grammatical statement applicable to the data source type are sunk to the data source for calculation, and a result data set of the data source is obtained; When the amount of data in the data source is greater than the preset amount of data or other data in the domain is required to participate in the calculation, and the conditions for sinking to the data source for calculation are not met, it is necessary to rely on the high-performance computing engine of the data computing agent to perform data calculations. According to the user intention and the computing task information, a data query statement and / or a grammatical statement suitable for the computing engine is generated, and data that meets the conditions is extracted into the computing engine. The computing engine executes the data computing task and obtains the result data set; The result data set is returned to the master agent.

4. The method according to claim 1, wherein The domain agent also includes a data governance agent, and the data processing task also includes a data governance task; The method further comprises: The master agent issues data governance tasks to the data governance agent in the domain agent, and the data governance tasks are used to manage and maintain the quality, security and permissions of the data in the data domain or the result data set; The data governance agent performs data governance tasks on the managed data domain and / or the result data set, obtains the governed result data set, and returns the governed result data set to the main agent.

5. The method according to claim 4, characterized in that The data governance agent performs data governance tasks on the managed data domain and / or the result data set to obtain a governed result data set, and returns the governed result data set to the master agent, specifically including: The data governance agent determines the data to be governed according to the data governance task, where the data to be governed includes data in the data domain or a result data set returned by the data computing agent; Check whether there are data quality issues in the data to be managed, and automatically repair any data quality issues that can be automatically repaired, such as missing values, duplicate values, and / or format errors. For data quality issues that cannot be automatically repaired, feedback indicates that the data has hidden problems. If the sensitive data in the data to be managed has not been desensitized or encrypted, the sensitive data will be desensitized or encrypted according to the security policy, and if there is no access permission for the sensitive data, insufficient permission information will be fed back; The input and output of data in the data domain are monitored in real time, abnormal fluctuations in the amount of data in the data domain are detected, and abnormal fluctuations are fed back.

6. The method according to claim 1, characterized in that The method further comprises: The master agent receives a result data set fed back by the data calculation agents in the plurality of domain agents, and caches the result data set in a fusion data warehouse; The result data set is extracted from the fusion data warehouse for fusion calculation to generate a fused result data set.

7. The method according to claim 3, characterized in that The method further comprises: The data computing agent determines that the data type of the search result or the field in the data domain is structured data, and processes it using a general database query or a data table DataFrame operation; The data type of the search result or the field in the data domain is unstructured data and is processed using a natural language processing method; The data type of the search result or the field in the data domain is semi-structured data, which is processed after being parsed and converted into structured data or unstructured data.

8. A data weaving system based on dialogue and multi-agent collaboration, characterized by: The data weaving system includes a main agent and at least one domain agent. A domain agent is a collection of sub-agents that manage a data domain. The main agent is used to identify user intentions and, based on the user intentions, to decompose and allocate data processing tasks, wherein the decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks; The master agent is further configured to issue a data retrieval task to the data retrieval agent in the domain agent, wherein the data retrieval task is configured to obtain storage information of the field to be retrieved in the data domain; The data retrieval agent is used to perform data retrieval tasks on the managed data domain, obtain retrieval results, including the table structure where the field to be retrieved is located and the storage information in the table structure, and return the retrieval results to the main agent; The master agent is used to issue data computing tasks to the data computing agent in the domain agent, wherein the data computing tasks are used to perform calculations based on the search results and / or one or more fields in the data domain; The data computing agent is configured to execute the data computing task, obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or the fields in the data domain, and return the result data set to the master agent; The master agent is configured to perform a data fusion task on the search results and / or the result dataset, and return the fused dataset to the user; The main agent is used to communicate with the user through natural language to obtain the user's intention, which includes the user's query requirements, including the query time range, time granularity, data objects and output requirements; Decompose the user's query requirements into tasks and issue data retrieval tasks to the data retrieval agent in the domain agent; The master agent is configured to receive the storage information of the to-be-retrieved field fed back by the data retrieval agent in the domain agent; Obtaining fields related to the user's query requirements from the to-be-retrieved fields fed back by the data retrieval agent; Determine the relevant data domains where the fields related to the user's query requirements are located, and at the same time, ensure the allocation of data calculation tasks in combination with the asset catalog maintained by each data domain; Generate an execution plan based on the user's query requirements; According to the execution plan, data computing tasks are issued to data computing agents in relevant data domains.

9. A data weaving system implementation device based on dialogue and multi-agent collaboration, characterized in that: The data weaving system includes a main agent and at least one domain agent, wherein a domain agent is a collection of sub-agents that manage data assets of a data domain, wherein a domain agent includes a data retrieval agent and a data calculation agent; The first processing device is used for the master agent to identify user intent and, based on the user intent, to decompose and allocate data processing tasks, wherein the decomposed data processing tasks include data retrieval tasks, data calculation tasks, and data fusion tasks; the master agent issues data retrieval tasks to the data retrieval agents in the domain agents, wherein the data retrieval tasks are used to obtain storage information of the fields to be retrieved in the data domain; a second processing device configured to cause the data retrieval agent to perform a data retrieval task on the managed data domain, obtain a retrieval result including the table structure where the to-be-retrieved field is located and information stored in the table structure, and return the retrieval result to the master agent; a third processing device, configured for the master agent to issue a data computing task to a data computing agent in a domain agent, wherein the data computing task is configured to perform calculations based on the search results and / or one or more fields in the data domain; a fourth processing device configured to cause the data computing agent to execute the data computing task, obtain a result data set, wherein the result data set includes a two-dimensional table structure calculated based on the search results or fields in the data domain, and return the result data set to the master agent; a fifth processing device, configured for the master agent to perform a data fusion task on the search results and / or the result dataset, and return the fused dataset to the user; The first processing device is configured to communicate with the user through natural language to obtain the user's intention, wherein the user's intention includes the user's query requirements, and the query requirements include the query time range, time granularity, data objects, and output requirements; Decompose the user's query requirements into tasks and issue data retrieval tasks to the data retrieval agent in the domain agent; The master agent is configured to receive the storage information of the to-be-retrieved field fed back by the data retrieval agent in the domain agent; Obtaining fields related to the user's query requirements from the to-be-retrieved fields fed back by the data retrieval agent; Determine the relevant data domains where the fields related to the user's query requirements are located, and at the same time, ensure the allocation of data calculation tasks in combination with the asset catalog maintained by each data domain; Generate an execution plan based on the user's query requirements; According to the execution plan, data computing tasks are issued to data computing agents in relevant data domains.

Citation Information

Patent Citations

  • Heterogeneous cluster data processing method and system based on multiple data centers and electronic equipment

    CN111143057A

  • Task execution method and device, equipment and storage medium

    CN116680061A

  • Data management system for intelligent operation and maintenance service

    CN117194399A

  • Complex information retrieval system and method

    CN118779364A