Data integration method and device based on large model Agent
By using a large-model agent-based intelligent data integration method, the problems of low efficiency and poor accuracy in traditional data integration are solved, and efficient and accurate automated execution of data integration tasks is achieved.
Patent Information
- Application Number
- CN202510896785.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional data integration solutions rely on manually written configuration files, resulting in low efficiency, poor accuracy, difficulty in quickly responding to dynamic changes in requirements, and high development and debugging costs.
An intelligent data integration method based on a large model agent is adopted. The intelligent interaction agent of the large model determines the text information in the requirement description, the parameter parsing agent parses the associated parameters and generates the target configuration file, and finally the execution scheduling agent sends it to the data integration tool to execute the task.
It has enabled intelligent data integration, improved efficiency and accuracy, reduced reliance on manually written configuration files, shortened configuration generation time, and reduced error rate.
Smart Images

Figure CN120929149A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and big data processing technology, and in particular to a data integration method and apparatus based on a large model agent. Background Technology
[0002] Data integration, as a core component of enterprise digital transformation, is facing numerous challenges. Modern enterprise data sources have expanded from traditional relational databases (such as Oracle and MySQL) to 32 types of heterogeneous systems, including NoSQL (MongoDB), cloud services (Snowflake), and real-time streaming (Kafka) (IDC 2024 report). The demand for cross-system data flow has surged, leading to a complex data ecosystem. Traditional ETL tools rely on technicians manually configuring task rules or writing JSON configuration files line by line. This results in long processing times per task, high error rates, difficulty in quickly responding to dynamic changes in requirements, and high development and debugging costs. Summary of the Invention
[0003] This application provides a data integration method and apparatus based on a large model agent to solve the problem that data integration schemes in related technologies rely on manually written configuration files, resulting in low data integration efficiency and poor accuracy.
[0004] Firstly, this application provides a data integration method based on a large-scale agent model, the method comprising:
[0005] Obtain the user's input description of needs, and use a smart interaction agent based on a large model to determine the various textual information in the description of needs that are associated with data integration;
[0006] The parameter parsing agent based on the large model parses each text message to determine the data integration and association parameters; wherein, the data integration and association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0007] Based on the data integration and correlation parameters, the Agent generates the target configuration file based on the configuration file of the large model;
[0008] The execution scheduling agent based on the large model sends the target configuration file to the data integration tool, and the data integration tool executes the corresponding data integration task according to the target configuration file.
[0009] The above technical solution has the following advantages or beneficial effects:
[0010] This application addresses the problem of low efficiency and poor accuracy in data integration due to the reliance on manually written configuration files in related technologies. It proposes an intelligent data integration scheme based on a large-model agent. After obtaining the user's input requirement description, the intelligent interaction agent based on the large model first identifies the various textual information in the requirement description that are related to data integration. Then, the parameter parsing agent based on the large model parses each textual information to determine the data integration-related parameters. Next, based on these parameters, a configuration file generation agent based on the large model generates the target configuration file. Finally, the execution scheduling agent of the large model sends the target configuration file to the data integration tool, which then executes the corresponding data integration task according to the target configuration file. This application achieves intelligent generation of the target configuration file for data integration based on the intelligent interaction agent, parameter parsing agent, and configuration file generation agent of the large model. The execution scheduling agent enables the sending of the target configuration file to the data integration tool, thereby executing the data integration task. This avoids the problem of low efficiency and poor accuracy caused by the reliance on manually written configuration files, thus improving the efficiency and accuracy of data integration.
[0011] Secondly, this application provides a data integration device based on a large model agent, the device comprising:
[0012] The first determining module is used to obtain the user's input description of needs, and the intelligent interaction agent based on the large model determines the various text information related to data integration in the description of needs.
[0013] The second determining module is used to parse the various text information based on the parameter parsing agent of the large model and determine the data integration association parameters; wherein, the data integration association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0014] The configuration file generation module is used to integrate related parameters based on the data and generate the target configuration file for the Agent based on the configuration file of the large model.
[0015] The sending module is used to send the target configuration file to the data integration tool based on the large model execution scheduling agent, and the data integration tool executes the corresponding data integration task according to the target configuration file.
[0016] Thirdly, this application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0017] Memory, used to store computer programs;
[0018] A processor, used to execute a program stored in memory, implements the method described.
[0019] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described herein.
[0020] Fifthly, this application provides a computer program product comprising an executable program that is executed by a processor to implement the method described. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This application provides a schematic diagram of the first data integration process based on a large model agent.
[0023] Figure 2 A schematic diagram illustrating the process of determining the various textual information related to data integration in the requirements description provided for this application;
[0024] Figure 3 This application provides a second schematic diagram of a data integration process based on a large model agent.
[0025] Figure 4 This application provides a schematic diagram of a third data integration process based on a large-model agent.
[0026] Figure 5 This application provides a schematic diagram of the fourth data integration process based on a large model agent.
[0027] Figure 6 A schematic diagram illustrating the process of generating target configuration files for the Agent based on the large model provided in this application;
[0028] Figure 7 This application provides a schematic diagram of the sixth type of data integration process based on a large model agent.
[0029] Figure 8 Detailed flowchart of data integration based on large model agent provided for this application;
[0030] Figure 9 A schematic diagram of the data integration device based on a large model agent provided in this application;
[0031] Figure 10 A schematic diagram of the electronic device structure provided in this application. Detailed Implementation
[0032] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0033] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0034] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0035] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0036] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0038] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
[0039] Figure 1 The first data integration process based on a large model agent provided in this application includes the following steps:
[0040] S101: Obtain the user's input description of requirements, and use the intelligent interaction agent based on the large model to determine the various text information related to data integration in the description of requirements;
[0041] S102: The parameter parsing agent based on the large model parses the various text information to determine the data integration and association parameters; wherein, the data integration and association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0042] S103: Based on the data integration and correlation parameters, generate the target configuration file for the Agent based on the configuration file of the large model;
[0043] S104: The execution scheduling agent based on the large model sends the target configuration file to the data integration tool, and the data integration tool executes the corresponding data integration task according to the target configuration file.
[0044] The data integration method based on large model agents provided in this application is applied to electronic devices, such as PCs, computers, smart terminals, servers, etc.
[0045] The electronic device is equipped with a data integration system. After logging into the system, a requirement description input box is displayed on the system interface, where the user enters their requirements. The electronic device acquires the user's input requirement description and, based on a large-scale model, an intelligent interaction agent determines the various textual information in the requirement description that are associated with data integration. For example, the requirement description is "Synchronize Oracle client tables to Hive every Saturday morning." The intelligent interaction agent, based on the large-scale model, determines the various textual information in the requirement description associated with data integration as "Oracle," "Hive," and "every Saturday morning."
[0046] The parameter parsing agent based on the large model analyzes each text message to determine the data integration and association parameters. These parameters include the source database connection parameters, the field mapping rules from the source database to the target database, and the data integration triggering condition parameters. For example, given the requirement description "synchronize Oracle client tables to Hive every Saturday morning," parsing the text message "Oracle" determines the source database connection parameter as "identify data source characteristics (Oracle connection string format: jdbc:oracle:thin:@host:port:SID)." Parsing the text message "every Saturday morning" determines the data integration triggering condition parameter as "extract scheduling parameters (cron expression: 0 0 0? *SAT*)."
[0047] It should be noted that if the requirement description includes field mapping rules, the above method can be used to extract the text information of the field mapping rules related to data integration. Then, based on the parameter parsing agent of the large model, this text information can be parsed to obtain the field mapping rule parameters from the source database to the target database. If the requirement description does not include field mapping rules, for example, if the requirement description is "synchronize Oracle client tables to Hive every Saturday morning," the field mapping rules corresponding to the data integration from the Oracle client tables to Hive can be determined based on the pre-saved default field mapping rules for different source databases to target databases. The field mapping rule parameters from the source database to the target database can then be obtained through parsing.
[0048] Based on the data integration and correlation parameters, the Agent generates the target configuration file based on the large model's configuration file. Optionally, the Agent generation process first loads the template library of the data integration tool, such as DataX, Sqoop, Logstash, or Airflow. Taking DataX as an example:
[0049]
[0050]
[0051] The target configuration file can be obtained through the above process. Then, the execution scheduling agent based on the large model sends the target configuration file to the data integration tool, which then executes the corresponding data integration task according to the target configuration file.
[0052] This application addresses the problem of low efficiency and poor accuracy in data integration due to the reliance on manually written configuration files in related technologies. It proposes an intelligent data integration scheme based on a large-model agent. After obtaining the user's input requirement description, the intelligent interaction agent based on the large model first identifies the various textual information in the requirement description that are related to data integration. Then, the parameter parsing agent based on the large model parses each textual information to determine the data integration-related parameters. Next, based on these parameters, a configuration file generation agent based on the large model generates the target configuration file. Finally, the execution scheduling agent of the large model sends the target configuration file to the data integration tool, which then executes the corresponding data integration task according to the target configuration file. This application achieves intelligent generation of the target configuration file for data integration based on the intelligent interaction agent, parameter parsing agent, and configuration file generation agent of the large model. The execution scheduling agent enables the sending of the target configuration file to the data integration tool, thereby executing the data integration task. This avoids the problem of low efficiency and poor accuracy caused by the reliance on manually written configuration files, thus improving the efficiency and accuracy of data integration.
[0053] Figure 2 A schematic diagram illustrating the process of determining the various textual information related to data integration in the requirements description provided for this application includes the following steps:
[0054] The intelligent interaction agent based on the large model determines that the various textual information related to data integration in the requirement description includes:
[0055] S201: The intelligent interaction agent based on the large model determines the target intent corresponding to the demand description; according to the pre-saved correspondence between each intent and each text information type, the target text information type corresponding to the target intent is determined;
[0056] S202: Based on the target text information types corresponding to the target intent, the intelligent interaction agent based on the large model determines the text information in the requirement description that corresponds to each of the target text information types respectively; wherein, the text information includes source database type text information, target database type text information, and trigger condition text information.
[0057] A large-model-based intelligent interaction agent can perform semantic analysis on the requirement description to obtain the target intent corresponding to the requirement description. Electronic devices pre-store the correspondence between each intent and each text information type, including the "data integration" intent. After determining the target intent, based on the pre-stored correspondence between each intent and each text information type, the corresponding target text information types can be determined. Then, based on the target text information types corresponding to the target intent, the large-model-based intelligent interaction agent determines the text information in the requirement description that corresponds to each of the target text information types. For the "data integration" target intent, the corresponding target text information types are source database type, target database type, and trigger condition type. The trigger condition type includes the trigger time type. The text information corresponding to the source database type is source database type text information (e.g., Oracle); the text information corresponding to the target database type is target database type text information (e.g., Hive); and the text information corresponding to the trigger condition type is trigger condition text information (e.g., every Saturday morning).
[0058] In one optional implementation, the intelligent interaction agent based on the large model determines that the various text information related to data integration in the requirement description also includes at least one of the following: data type conversion text information, data cleaning rule text information, field operation text information, and data integration strategy text information; wherein, data type includes numeric type, date and time type, and string type; field operation includes field data desensitization operation; and data integration strategy includes full strategy and incremental strategy.
[0059] The parameter parsing agent based on the large model parses the various text information and determines that the data integration and association parameters also include at least one of the following: data type conversion parameters, data cleaning rule parameters, field operation parameters, and data integration strategy parameters.
[0060] Data type conversion text information includes phrases like "Convert the string type in the source database to a numeric type in the target database." Data cleaning rule text information includes phrases like "Delete values greater than the preset first threshold" and "Delete values less than the preset second threshold." Field operation text information includes phrases like "Desensitize the mobile phone number field." Data integration strategy text information includes phrases like "Use a full-scale strategy for data integration" and "Use an incremental strategy for data integration." A full-scale strategy means integrating all data each time. An incremental strategy means integrating only the data added compared to the previous integration.
[0061] Correspondingly, the parameter parsing agent based on the large model parses each piece of text information to determine that the data integration association parameters also include at least one of the following: data type conversion parameters, data cleaning rule parameters, field operation parameters, and data integration strategy parameters. For example, if the intelligent interaction agent based on the large model determines that the text information related to data integration in the requirement description also includes "data type conversion text information, data cleaning rule text information, and field operation text information," then the parameter parsing agent based on the large model will parse the "data type conversion text information, data cleaning rule text information, and field operation text information" to determine that the data integration association parameters also include "data type conversion parameters, data cleaning rule parameters, and field operation parameters."
[0062] Figure 3 The second data integration process based on a large model agent provided in this application includes the following steps:
[0063] S301: Obtain the user's input description of requirements, and use a large-scale intelligent interaction agent to determine the various text information related to data integration in the description of requirements;
[0064] S302: The parameter parsing agent based on the large model parses each text information to determine the data integration and association parameters; wherein, the data integration and association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0065] S303: Determine the parameter types supported by the configuration file generator Agent of the large model; convert the data integration and association parameters according to the parameter types to obtain the data integration and association parameters of the parameter types;
[0066] S304: Based on the data integration and association parameters of the parameter type, generate the target configuration file of the Agent based on the configuration file of the large model;
[0067] S305: The execution scheduling agent based on the large model sends the target configuration file to the data integration tool, and the data integration tool executes the corresponding data integration task according to the target configuration file.
[0068] The data integration and association parameters are transformed according to their parameter types; in other words, natural language parameters are converted into technical parameters. The transformed technical parameters (or, as described below, data integration and association parameters of the specified parameter types obtained by transforming the data integration and association parameters according to their parameter types) are illustrated with an example:
[0069]
[0070]
[0071] Figure 4 The third data integration process based on a large model agent provided in this application includes the following steps:
[0072] S401: Obtain the user's input description of requirements, and use the intelligent interaction agent based on the large model to determine the various text information related to data integration in the description of requirements;
[0073] S402: The parameter parsing agent based on the large model parses each text information to determine the data integration and association parameters; wherein, the data integration and association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0074] S403: The rule verification agent based on the large model performs static and dynamic verification checks on data integration; if both static and dynamic verification checks pass, the agent generates the target configuration file based on the configuration file of the large model according to the data integration association parameters.
[0075] The static verification check includes the integrity check of the data integration and association parameters and the data compatibility check between the source database and the target database.
[0076] The dynamic verification check includes the existence check of the source data table in the source database, the existence check of the target data table in the target database, and the matching check of the field mapping rule parameters from the source database to the target database.
[0077] S404: The execution scheduling agent based on the large model sends the target configuration file to the data integration tool, and the data integration tool executes the corresponding data integration task according to the target configuration file.
[0078] The data integration and association parameter integrity check refers to verifying whether the data integration and association parameters determined by the parameter parsing agent based on the large model include all the parameters necessary for data integration. For example, all necessary parameters for data integration include "source database connection parameters, field mapping rule parameters from the source database to the target database, and data integration trigger condition parameters." If the parameter parsing agent based on the large model determines that the data integration and association parameters include "source database connection parameters, field mapping rule parameters from the source database to the target database, and data integration trigger condition parameters," then the data integration and association parameter integrity check passes. If the parameter parsing agent determines that the data integration and association parameters include "source database connection parameters and field mapping rule parameters from the source database to the target database," but do not include "data integration trigger condition parameters," then the data integration and association parameter integrity check fails.
[0079] Data compatibility checking between the source and target databases refers to verifying whether the data types supported by the source and target databases match. For example, if both the source and target databases support numeric and string data types, then the data types supported by the source and target databases match, and the data compatibility check passes. Conversely, if the source database supports numeric data types while the target database supports date and time data types, then the data types supported by the source and target databases do not match, and the data compatibility check fails.
[0080] The existence check of source data tables in the source database refers to checking whether the source data tables used for data integration exist in the source database. If they exist, the existence check of source data tables in the source database passes; if they do not exist, the existence check of source data tables in the source database fails.
[0081] The existence check of the target data table in the target database refers to checking whether the target data table for data integration exists in the target database. If it exists, the existence check of the target data table in the target database passes; if it does not exist, the existence check of the target data table in the target database fails.
[0082] The matching check of field mapping rule parameters from the source database to the target database refers to checking whether the data types supported by the mapped fields in the source database and target database are consistent. For example, field A in the source database is mapped to field M in the target database. If both field A and field M in the target database support the data type "numeric", then the data types supported by the mapped fields in the source database and target database are consistent, and the matching check of the field mapping rule parameters from the source database and target database passes. If field A in the source database supports the data type "numeric", but field M in the target database supports the data type "date and time", then the data types supported by the mapped fields in the source database and target database are inconsistent, and the matching check of the field mapping rule parameters from the source database and target database fails.
[0083] Figure 5 The fourth data integration process based on a large model agent provided in this application includes the following steps:
[0084] S501: Obtain the user's input description of requirements, and use the intelligent interaction agent based on the large model to determine the various text information related to data integration in the description of requirements;
[0085] S502: The parameter parsing agent based on the large model parses each text information to determine the data integration and association parameters; wherein, the data integration and association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0086] S503: Based on the data integration and correlation parameters, generate the target configuration file for the Agent based on the configuration file of the large model;
[0087] S504: Display the target configuration file in the editable display interface. If a modification instruction for the target configuration file is received, obtain and modify the target configuration file according to the modification content carried by the modification instruction to obtain the modified target configuration file.
[0088] S505: The execution scheduling agent based on the large model sends the modified target configuration file to the data integration tool, and the data integration tool executes the corresponding data integration task according to the modified target configuration file.
[0089] Based on the data integration association parameters, after the large model-based configuration file generation agent generates the target configuration file, it is displayed on the editable display interface. The editable display interface shows a modify button and a confirm button. If the user clicks the confirm button, it is determined that the target configuration file will not be modified. At this time, the large model-based execution scheduling agent sends the target configuration file to the data integration tool, which then executes the corresponding data integration task based on the target configuration file. If the user clicks the modify button and the user's modifications to the target configuration file are received (the modifications are carried in the modification instruction), it is determined that the modification instruction for the target configuration file has been received. The modification content carried in the modification instruction is obtained, and the target configuration file is modified according to the modification content to obtain the modified target configuration file. Then, the large model-based execution scheduling agent sends the modified target configuration file to the data integration tool, which then executes the corresponding data integration task based on the modified target configuration file.
[0090] Figure 6 The process of generating the target configuration file for the Agent based on the large model provided in this application is illustrated in the following steps:
[0091] S601: Generate target configuration file based on large model configuration file; generate configuration comment information at the location corresponding to the configuration code in the target configuration file.
[0092] S602: Add data integration concurrency information to the target configuration file; wherein, the data integration concurrency information is used to characterize the number of threads that process in parallel during the data integration process.
[0093] Generate configuration comment information at the location corresponding to the configuration code in the target configuration file, for example:
[0094]
[0095]
[0096] The content following the # symbol is the configuration comment information. Generating configuration comments at the corresponding locations in the configuration code of the target configuration file helps users understand the configuration file.
[0097] Figure 7 The sixth data integration process based on a large model agent provided in this application includes the following steps:
[0098] S701: Obtain the user's input description of requirements, and use a large-scale intelligent interaction agent to determine the various text information related to data integration in the description of requirements;
[0099] S702: The parameter parsing agent based on the large model parses each text information to determine the data integration and association parameters; wherein, the data integration and association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0100] S703: Based on the data integration and correlation parameters, generate the target configuration file for the Agent based on the configuration file of the large model;
[0101] S704: The execution scheduling agent based on the large model sends the target configuration file to the data integration tool, and the data integration tool executes the corresponding data integration task according to the target configuration file;
[0102] S705: The execution scheduling agent based on the large model performs lifecycle management of data integration tasks. When monitoring data integration anomalies, it determines the type of data integration anomaly through log parsing and outputs prompt information containing anomaly repair strategies based on the type of data integration anomaly.
[0103] When monitoring data integration anomalies, a data integration anomaly message will be output. The execution scheduling agent based on the large model performs log parsing to determine the type of data integration anomaly. For example, if the execution scheduling agent based on the large model detects an anomaly in the data integration process based on the data integration concurrency information, the output message might include an anomaly remediation strategy such as "modify the data integration concurrency information in the configuration file." As another example, if the execution scheduling agent based on the large model detects a program malfunction during data integration, the output message might include an anomaly remediation strategy such as "restart the data integration function."
[0104] The data integration method based on large model agents provided in this application belongs to the interdisciplinary field of artificial intelligence and big data processing technology. Specifically, it involves an intelligent data integration method and system based on a large model (large language model). Through multi-round natural language interaction, it parses user needs, automatically generates structured data integration task configuration files, and drives the underlying data synchronization engine to achieve efficient integration and automated processing of cross-platform, multi-source heterogeneous data.
[0105] The improvements in this application are as follows: It pioneers an intelligent docking architecture between the large model agent and DataX (data synchronization tool), realizing the automatic conversion from natural language to structured configuration; it develops a multi-level parameter extraction mechanism, which automatically distinguishes between required and optional parameters and generates compliant JSON through semantic parsing; it builds a dynamic validation rule engine, integrating a constraint condition library of more than 200 data integration scenarios; and it innovates a context-aware dialogue model, supporting cross-round requirement supplementation and configuration iteration optimization.
[0106] The core advantage of this application lies in constructing a natural language-driven intelligent data integration paradigm. Through deep collaboration between large models and data engines, it achieves end-to-end automation from requirement description to task configuration. Compared to traditional manual coding configuration methods, the system breaks through by reducing configuration generation time to minutes and significantly reducing the error rate by relying on a dynamic validation rule base. It also supports real-time integration with more than 10 heterogeneous data sources and can be expanded as needed, providing a high-precision, low-barrier one-stop solution for multi-scenario data fusion.
[0107] This application leverages advanced multi-agent collaboration technology to facilitate the entire data integration process through guided natural language interaction. By having a large-scale model agent interact with the user in natural language, it understands the user's needs, automatically extracts data integration elements, generates a standardized DataX Job configuration file through multi-level parameter validation, and ultimately drives the DataX execution engine to complete scheduled / real-time data synchronization across heterogeneous data sources. This entire process realizes a next-generation data integration paradigm characterized by "dialogue-based requirements, automated configuration, and intelligent execution."
[0108] The core functional modules include:
[0109] 1. Intelligent Interaction Agent (Conversation Agent).
[0110] Functional positioning: Human-computer interaction entry point and demand understanding center;
[0111] Core capabilities: • Guided multi-turn dialogue management (context preservation / ambiguity clarification), natural language requirement parsing (entity recognition / intent classification), dynamic form generation (guided required parameters / recommendation of optional parameters).
[0112] 2. Parameter Extraction Agent.
[0113] Functional positioning: Structured parameter extraction engine.
[0114] Core capabilities: Data source feature extraction (JDBC / API / NoSQL connection parameters, etc.), transformation rule parsing (field mapping / type conversion / cleaning rules), and scheduling strategy identification (full / incremental strategies, triggering conditions).
[0115] 3. Rule Validation Agent.
[0116] Functional positioning: Guardian of configuration compliance.
[0117] Core capabilities: Parameter integrity check (required field validation) and logical consistency validation (such as source / target field type compatibility).
[0118] 4. Generate the Agent (CodeGen Agent) from the configuration file (Job).
[0119] Functionality: DataX configuration generator.
[0120] Core capabilities: Templated JSON construction (based on the DataX plugin system), dynamic parameter injection (concurrency optimization), and enhanced readability (automatic generation of configuration comments).
[0121] 5. Execute the scheduling agent (Orchestration Agent).
[0122] Functional role: Mission execution commander.
[0123] Core capabilities: Job lifecycle management (startup / monitoring / retry), exception handling (log parsing / error classification / self-healing suggestions), and execution feedback (data display feedback based on runtime data).
[0124] The data integration process will be illustrated below with a specific example.
[0125] This application provides a data integration system based on a large-scale model agent. This system achieves intelligent data integration through multi-level agent collaboration. Detailed implementation steps are as follows:
[0126] Step 1: Natural Language Interaction Layer.
[0127] User input:
[0128] Users initiate natural language requests via Web / CLI / API (e.g., "Synchronize the Oracle customer table to Hive every Saturday morning, the mobile phone number field needs to be anonymized").
[0129] Dialogue Management:
[0130] The intelligent interaction agent initiates a dialogue session and maintains the dialogue context.
[0131] Intent recognition is performed using a large model (categorized as "data synchronization + data anonymization").
[0132] Dynamic form generation:
[0133] Based on the identified intent, a parameter collection form is generated by calling a preset template.
[0134] Example output:
[0135] [Required parameter]:
[0136] 1. Source database type: Oracle
[0137] 2. Target storage type: Hive
[0138] 3. Synchronization frequency: Every Saturday at 00:00
[0139] [Optional parameters]:
[0140] Anonymization rules: Mobile phone number (MD5 encrypted / partially hidden).
[0141] Step 2: Parameter structure parsing.
[0142] Parameter extraction:
[0143] The Agent performs the following operations during parameter parsing:
[0144] Identify data source characteristics (Oracle connection string format: jdbc:oracle:thin:@host:port:SID); parse field-level operations (identifying the "phone number" field requires applying data masking rules);
[0145] Extract scheduling parameters (cron expression: 0 0 0? *SAT*).
[0146] Parameter standardization:
[0147] The natural language parameters are converted into technical parameters. An example of the converted technical parameters is shown below:
[0148]
[0149]
[0150] Step 3: Rule validation and optimization.
[0151] Static validation:
[0152] Rule verification: The agent performs basic checks.
[0153] Validation of the completeness of required parameters (form);
[0154] Data type compatibility checks (e.g., Oracle DATE matches Hive TIMESTAMP).
[0155] Dynamic verification:
[0156] Calling a large model for logical reasoning:
[0157] Check if the source table exists in the source database;
[0158] Check if the target table exists in the target database (execute DESC customers);
[0159] Verify that the fields in the source table and the target table match.
[0160] Step 4: Configure generation and injection.
[0161] Template rendering:
[0162] Job generates Agent and loads DataX template library:
[0163]
[0164] Parameter injection:
[0165] Dynamically populate template parameters and generate the final JSON:
[0166]
[0167] Step 5: Task execution and monitoring.
[0168] Assignment Submission:
[0169] The execution scheduling agent submits tasks via the DataX API:
[0170] python datax.py / jobs / oracle2hive.json.
[0171] Status monitoring:
[0172] Real-time capture and display of DataX runtime logs.
[0173] Results table shown:
[0174] The sampling results and dirty data are displayed.
[0175] 1. This application abstracts a unified configuration layer, converting the configuration generated from a large model into DSLs for different engines (such as Sqoop's MapReduce tasks). Executors are pluggable, dynamically loading different engine drivers through the adapter pattern.
[0176] 2. Code fusion solution. Alternative: Natural language to low-code module.
[0177] Areas for improvement:
[0178] Generate DataX configuration and React front-end code simultaneously based on user requirements;
[0179] It provides an editable flowchart interface and supports configuration write-back modification.
[0180] 3. Multimodal demand input mechanism.
[0181] Alternative solution: In addition to natural language interaction, support voice command recognition to generate structured requirement text. The system can integrate a voice recognition engine, allowing users to describe their needs via voice. Advantages: Lowers the user barrier and improves the flexibility of human-computer interaction.
[0182] Figure 8 This application provides a detailed flowchart of the data integration process based on a large-model agent. It includes: user interaction layer, parameter structuring, rule validation and optimization, configuration generation and injection, and task execution and monitoring.
[0183] User interaction layer execution: natural language input requirements, multi-turn dialogue to clarify details, dynamic form generation: guidance on required parameters, and recommendations for optional parameters.
[0184] Parameter-structured execution: data source parsing, transformation rule parsing, and scheduling strategy parsing.
[0185] Rule validation and optimization: parameter integrity detection and logical consistency detection.
[0186] Configuration generation and injection execution: template library template rendering and parameter injection.
[0187] Task execution and monitoring: job submission, status monitoring, and result feedback.
[0188] Figure 9 The schematic diagram of the data integration device based on large model agents provided in this application includes:
[0189] The first determining module 11 is used to obtain the user's input demand description and determine the various text information related to data integration in the demand description based on the intelligent interaction agent of the large model.
[0190] The second determining module 12 is used to parse the various text information based on the parameter parsing agent of the large model and determine the data integration association parameters; wherein, the data integration association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters;
[0191] The configuration file generation module 13 is used to integrate related parameters based on the data and generate a target configuration file for the Agent based on the configuration file of the large model.
[0192] The sending module 14 is used to send the target configuration file to the data integration tool based on the large model execution scheduling agent, and the data integration tool executes the corresponding data integration task according to the target configuration file.
[0193] The first determining module 11 is specifically used to determine the target intent corresponding to the requirement description by an intelligent interaction agent based on a large model; determine each target text information type corresponding to the target intent according to the pre-saved correspondence between each intent and each text information type; and determine the text information in the requirement description corresponding to each target text information type according to the target text information type corresponding to the target intent by the intelligent interaction agent based on the large model; wherein, the text information includes source database type text information, target database type text information, and trigger condition text information.
[0194] The second determining module 12 is further configured to determine the parameter types supported by the configuration file generation Agent of the large model; and to convert the data integration and association parameters according to the parameter types to obtain the data integration and association parameters of the parameter types.
[0195] The configuration file generation module 13 is specifically used to generate the target configuration file for the Agent based on the configuration file of the large model by integrating and associating parameters according to the data of the parameter type.
[0196] The second determining module 12 is also used to perform static and dynamic verification checks on the data integration based on the rule verification agent of the large model; if both the static and dynamic verification checks pass, the process of generating the target configuration file based on the configuration file of the large model and the data integration association parameters is carried out.
[0197] The static verification check includes the integrity check of the data integration and association parameters and the data compatibility check between the source database and the target database.
[0198] The dynamic verification check includes the existence check of the source data table in the source database, the existence check of the target data table in the target database, and the matching check of the field mapping rule parameters from the source database to the target database.
[0199] The configuration file generation module 13 is also used to display the target configuration file on the editable display interface. If a modification instruction for the target configuration file is received, the module obtains and modifies the target configuration file according to the modification content carried by the modification instruction to obtain the modified target configuration file.
[0200] The sending module 14 is specifically used to send the modified target configuration file to the data integration tool based on the large model execution scheduling agent.
[0201] The configuration file generation module 13 is specifically used to generate a target configuration file based on a large model, generate configuration comment information at the location corresponding to the configuration code in the target configuration file, and add data integration concurrency information to the target configuration file; wherein, the data integration concurrency information is used to characterize the number of threads that process in parallel during the data integration process.
[0202] The sending module 14 is also used for lifecycle management of data integration tasks based on the execution scheduling agent of the large model. When monitoring data integration anomalies, it determines the type of data integration anomaly through log parsing and outputs prompt information containing anomaly repair strategies according to the type of data integration anomaly.
[0203] This application also provides an electronic device, such as Figure 10 As shown, it includes: processor 21, communication interface 22, memory 23 and communication bus 24, wherein processor 21, communication interface 22 and memory 23 communicate with each other through communication bus 24;
[0204] The memory 23 stores a computer program, which, when executed by the processor 21, causes the processor 21 to perform any of the above method steps.
[0205] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0206] Communication interface 22 is used for communication between the above-mentioned electronic device and other devices.
[0207] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0208] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0209] This application also provides a computer-readable storage medium storing a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform any of the above method steps.
[0210] This application provides a computer program product, which includes an executable program that, when executed by a processor, implements the method described herein.
[0211] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0212] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A data integration method based on a large-scale agent model, characterized in that, The method includes: Obtain the user's input description of needs, and use a smart interaction agent based on a large model to determine the various textual information in the description of needs that are associated with data integration; The parameter parsing agent based on the large model parses each text message to determine the data integration and association parameters; wherein, the data integration and association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters; Based on the data integration and correlation parameters, the Agent generates the target configuration file based on the configuration file of the large model; The execution scheduling agent based on the large model sends the target configuration file to the data integration tool, and the data integration tool executes the corresponding data integration task according to the target configuration file.
2. The method as described in claim 1, characterized in that, The intelligent interaction agent based on the large model determines that the various textual information related to data integration in the requirement description includes: The intelligent interaction agent based on the large model determines the target intent corresponding to the requirement description; according to the pre-saved correspondence between each intent and each text information type, it determines each target text information type corresponding to the target intent; according to each target text information type corresponding to the target intent, the intelligent interaction agent based on the large model determines the text information in the requirement description that corresponds to each target text information type; wherein, the text information includes source database type text information, target database type text information, and trigger condition text information.
3. The method as described in claim 2, characterized in that, The intelligent interaction agent based on the large model determines that the various text information related to data integration in the requirement description also includes at least one of the following: data type conversion text information, data cleaning rule text information, field operation text information, and data integration strategy text information; wherein, data type includes numeric type, date and time type, and string type; field operation includes field data desensitization operation; data integration strategy includes full strategy and incremental strategy; The parameter parsing agent based on the large model parses the various text information and determines that the data integration and association parameters also include at least one of the following: data type conversion parameters, data cleaning rule parameters, field operation parameters, and data integration strategy parameters.
4. The method as described in claim 1, characterized in that, After the large model-based parameter parsing agent parses the various text information and determines the data integration and association parameters, and before the large model-based configuration file generation agent generates the target configuration file based on the data integration and association parameters, the method further includes: Determine the parameter types supported by the Agent generated from the configuration file of the large model; convert the data integration and association parameters according to the parameter types to obtain the data integration and association parameters of the specified parameter types; Based on the data integration and correlation parameters, the Agent target configuration file is generated based on the configuration file of the large model, including: Based on the data integration and association parameters of the parameter type, the Agent is generated and the target configuration file is generated based on the configuration file of the large model.
5. The method as described in claim 1, characterized in that, After the large model-based parameter parsing agent parses the various text information and determines the data integration and association parameters, and before the large model-based configuration file generation agent generates the target configuration file based on the data integration and association parameters, the method further includes: The rule-validation agent based on the large model performs static and dynamic validation checks on data integration; if both static and dynamic validation checks pass, the agent generates the target configuration file based on the configuration file of the large model and the data integration association parameters. The static verification check includes the integrity check of the data integration and association parameters and the data compatibility check between the source database and the target database. The dynamic verification check includes the existence check of the source data table in the source database, the existence check of the target data table in the target database, and the matching check of the field mapping rule parameters from the source database to the target database.
6. The method as described in claim 1, characterized in that, Based on the data integration and association parameters, after the large model-based configuration file generation agent generates the target configuration file, and before the large model-based execution scheduling agent sends the target configuration file to the data integration tool, the method further includes: The target configuration file is displayed in an editable display interface. If a modification instruction for the target configuration file is received, the target configuration file is modified according to the modification content carried by the modification instruction to obtain the modified target configuration file. The execution scheduling agent based on the large model sends the target configuration file to the data integration tool, including: The execution scheduling agent based on the large model sends the modified target configuration file to the data integration tool.
7. The method as described in claim 1, characterized in that, The agent-based target configuration file generation based on the large model includes: The agent generates a target configuration file based on the large model configuration file, and generates configuration annotation information at the location corresponding to the configuration code in the target configuration file; data integration concurrency information is added to the target configuration file; wherein, the data integration concurrency information is used to characterize the number of threads that process in parallel during the data integration process.
8. The method as described in claim 1, characterized in that, The method further includes: The execution scheduling agent based on the large model manages the lifecycle of data integration tasks. When monitoring data integration anomalies, it determines the type of data integration anomaly through log parsing and outputs a prompt message containing anomaly repair strategies based on the type of data integration anomaly.
9. A data integration device based on a large-scale agent model, characterized in that, The device includes: The first determining module is used to obtain the user's input description of needs, and the intelligent interaction agent based on the large model determines the various text information related to data integration in the description of needs. The second determining module is used to parse the various text information based on the parameter parsing agent of the large model and determine the data integration association parameters; wherein, the data integration association parameters include the connection parameters of the source database, the field mapping rule parameters from the source database to the target database, and the data integration trigger condition parameters; The configuration file generation module is used to integrate related parameters based on the data and generate the target configuration file for the Agent based on the configuration file of the large model. The sending module is used to send the target configuration file to the data integration tool based on the large model execution scheduling agent, and the data integration tool executes the corresponding data integration task according to the target configuration file.
10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-8.
Citation Information
Patent Citations
Configurable data integration method based on Flume
CN111104397A
Data integration task construction device based on artificial intelligence interaction
CN118193854A
Data migration method and device, related equipment and computer program product
CN118503227A
Cloud native container platform integration method and system
CN119248420A
A multi-source data fusion sharing platform based on large model and its sharing method
CN119782399A