Intelligent distribution method and device for multi-task data integration engine
By extracting metadata and server status information from data sources and using a task scheduling model to determine the estimated completion time of the task engine, we address the high cost of getting started with data integration tools and the configuration complexity caused by their diversity, thus enabling cross-platform synchronization of all types of data and simplified data synchronization.
Patent Information
- Application Number
- CN202510548476.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-09
AI Technical Summary
Existing data integration tools have high startup costs and configuration complexity caused by tool diversity when synchronizing across platforms. They are unable to fully cover all types of data synchronization scenarios and have problems with incompatible cross-platform data synchronization configurations.
By extracting the metadata information of the data source to be assigned, data conversion information, the server's operating status, and the task information of multiple task engines, the task scheduling model is used to extract non-time series features and time series features, determine the estimated completion time of each task engine, and assign the data conversion task to the target task engine with the shortest estimated completion time, thus achieving cross-platform synchronization.
It enables easy cross-platform synchronization of all types of data, reduces learning costs, improves the flexibility and transparency of data integration, simplifies the data synchronization process, and improves the efficiency and controllability of the data integration process.
Smart Images

Figure CN120610786A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method and device for intelligently allocating a multi-task data integration engine. Background Art
[0002] Data integration tools are software tools used to collect, integrate, transform, and load (ETL) data from different data sources. They can be used to transfer data from one or more sources to a target system, such as a data warehouse, data lake, or business intelligence tool. Data integration tools typically perform tasks such as extraction, transformation, and loading on data. Extraction involves extracting data from multiple data sources, including relational databases, files, and web APIs; transformation involves converting and cleaning the extracted data, such as merging, splitting, filtering, and converting data types; and loading involves loading the converted data into a target system, such as a data warehouse, data lake, or business intelligence tool.
[0003] By using data integration tools, businesses can integrate data from multiple sources and transform it into a consistent format for easier analysis and understanding. This helps businesses make more informed business decisions and identify new business opportunities. Common data integration tools include DataX, Kettle, and Spark.
[0004] DataX is an open-source data integration tool with the following advantages: it provides ETL (Extract, Transform, and TL) process descriptions in JSON format and supports data migration between multiple data sources, making data synchronization easy. It also uses multi-threading and streaming read / write technologies to improve data synchronization efficiency and stability. It also supports custom plug-ins, allowing users to extend DataX's functionality based on their needs. However, its disadvantages include a lack of intuitive visualization. The JSON-formatted data processing configuration can make complex ETL processes difficult to understand and debug, especially for users unfamiliar with the JSON format, who struggle to quickly grasp the structure and key aspects of the entire process.
[0005] Kettle is an open-source ETL tool, also known as Pentaho Data Integration. Its advantages include: a graphical interface and support for a wide range of data sources and target systems, allowing users to easily perform data conversion operations; powerful conversion and cleansing capabilities, allowing users to customize various complex data conversion requirements; and support for custom plug-ins, allowing users to extend Kettle's functionality according to their needs. However, its disadvantages include the lack of intuitiveness when dealing with complex logic or critical information, requiring users to carefully review the configuration details of each component to fully understand the entire process.
[0006] Spark is a popular distributed computing framework that can also be used for ETL tasks. Its advantages include support for a variety of data sources, including HDFS, Hive, and JDBC; its use of in-memory and distributed computing technologies allows for rapid processing of large amounts of data; and its ability to easily scale to multiple machines for distributed computing. However, its disadvantages include a high learning curve for users without programming skills, as writing and maintaining ETL code can lead to code errors and debugging difficulties during data processing. Furthermore, its pure code implementation can make information sharing and understanding more difficult during team collaboration.
[0007] In summary, there is a need to provide a simple and intelligent allocation method and apparatus for a multi-task data integration engine that can be used for cross-platform synchronization of all types of data. Summary of the Invention
[0008] To solve the above problems, the present application proposes a method and device for intelligent allocation of a multi-task data integration engine.
[0009] On the one hand, the present application proposes a method for intelligently allocating a multi-task data integration engine, comprising:
[0010] Extract metadata information of the data source to be assigned, data conversion information, server operation status and task information of multiple task engines;
[0011] The task scheduling model extracts non-time series features and time series features from the metadata information, the data conversion information, the running status of the server and the task information of multiple task engines;
[0012] The task scheduling model determines an estimated completion time of each task engine in the plurality of task engines according to the timing feature and the non-timing feature;
[0013] The task scheduling model allocates the data conversion task corresponding to the to-be-allocated data source to a target task engine according to the estimated completion time, and the target task engine is the task engine with the shortest estimated completion time.
[0014] Preferably, before the task scheduling model allocates the task corresponding to the data source to be allocated to the target task engine according to the expected completion time, it also includes: converting the preset configuration file of the data source to be allocated into a task configuration file for the multiple task engines.
[0015] Preferably, the task scheduling model extracts non-time series features and time series features from the metadata information, data conversion information, the running status of the server, and task information of multiple task engines, including:
[0016] Acquire timing parameter variables and non-timing parameter variables from the metadata information, the data conversion information, the running status of the server, and the task information of the plurality of task engines;
[0017] The task scheduling model extracts the timing features from the timing parameter variables and extracts the non-timing features from the non-timing parameter variables.
[0018] Preferably, before the task scheduling model determines the estimated completion time of each of the plurality of task engines according to the timing feature and the non-timing feature, the task scheduling model further includes:
[0019] determining a plurality of status types of the computing resource according to the operating status;
[0020] Determining a configuration mapping table of the server according to the multiple status types and the configuration standard of the server;
[0021] The model to be trained is trained according to the configuration standard, the configuration mapping table and historical data to obtain the trained task scheduling model.
[0022] Preferably, the task scheduling model determines the estimated completion time of each task engine in the plurality of task engines according to the timing characteristics and the non-timing characteristics, including:
[0023] The task scheduling model determines the estimated completion time of each task engine in each server processing the to-be-assigned data source according to the timing characteristics, the non-timing characteristics, the configuration standard, and the configuration mapping table.
[0024] Preferably, determining the configuration mapping table of the server according to the multiple status types and the configuration standard of the server includes:
[0025] Determine the energy consumption of each task engine corresponding to the different state types and the operating states, and obtain a state energy consumption table;
[0026] Determining a configuration standard of the server according to the status performance table;
[0027] A configuration mapping table of the server is determined according to the configuration standard.
[0028] Preferably, the training of the to-be-trained model according to the configuration standard, the configuration mapping table and the historical data to obtain the trained task scheduling model includes:
[0029] Adjust the preset loss function of the model to be trained according to the configuration standard to obtain a preset model;
[0030] Adjusting the historical data according to the configuration mapping table;
[0031] The adjusted historical data is used to train the preset model to obtain the trained task scheduling model.
[0032] Preferably, after the task scheduling model allocates the data conversion task corresponding to the to-be-allocated data source to the target task engine according to the estimated completion time, the method further includes:
[0033] The target task engine is used to collect, integrate, convert and load the data source to be allocated according to the task configuration file.
[0034] Preferably, before extracting metadata information of the data source to be allocated, data conversion information, the running status of the server and task information of multiple task engines, the method further includes:
[0035] Parsing a preset configuration file to obtain data conversion information of the data source to be processed;
[0036] Scan the data source to be processed to obtain metadata information of the data source to be processed;
[0037] Acquiring the running status of the server, wherein the running status includes: computing resource information of the server, task execution information of the server, and running environment information of the server;
[0038] The task information of the multiple task engines is acquired, where the task information includes a current task amount of each task engine and an execution status of the current task.
[0039] In a second aspect, the present application proposes an intelligent allocation device for a multi-task data integration engine, comprising:
[0040] Parameter extraction module, used to extract metadata information of the data source to be assigned, data conversion information, server operation status and task information of multiple task engines;
[0041] A task scheduling module is used to extract non-time series features and time series features from the metadata information, data conversion information, the operating status of the server and the task information of multiple task engines through a task scheduling model; the task scheduling model determines the estimated completion time of each task engine in the multiple task engines based on the time series features and the non-time series features; the task scheduling model allocates the data conversion task corresponding to the to-be-allocated data source to a target task engine based on the estimated completion time, and the target task engine is the task engine with the shortest estimated completion time.
[0042] The advantages of this application are: by extracting the metadata information, data conversion information, server operating status and task information of multiple task engines of the data source to be allocated, and using the task scheduling model to extract non-time series features and time series features therefrom, non-time series features and time series features can be extracted from the data source to be allocated of all types of data and the data conversion information corresponding to various platforms; the task scheduling model is used to determine the estimated completion time of each task engine in multiple task engines based on the time series features and non-time series features, and the task engine with the shortest estimated completion time is used as the target engine, and the data conversion tasks corresponding to the data source to be allocated are allocated to the target task engine, thereby enabling cross-platform synchronization of all types of data to be achieved in a simple and convenient manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to denote the same components. In the drawings:
[0044] Figure 1 This is a schematic diagram of the steps of a method for intelligently allocating a multi-task data integration engine provided by the present application;
[0045] Figure 2 This is a schematic diagram of the model training process of the intelligent allocation method of a multi-task data integration engine provided by the present application;
[0046] Figure 3 This is a schematic diagram of the overall architecture of an intelligent allocation method for a multi-task data integration engine provided by the present application;
[0047] Figure 4This is a schematic diagram of an intelligent allocation method for a multi-task data integration engine provided by the present application for collecting, integrating, converting and loading the data source according to a preset configuration file;
[0048] Figure 5 This is a schematic diagram of a flow chart of converting a preset configuration file into a DataX configuration file according to an intelligent allocation method of a multi-task data integration engine provided by the present application;
[0049] Figure 6 This is a schematic diagram of a flow chart of converting a preset configuration file into a kettel configuration file according to an intelligent allocation method of a multi-task data integration engine provided by the present application;
[0050] Figure 7 This is a schematic diagram of a flow chart of converting a preset configuration file into a Flink configuration file according to an intelligent allocation method of a multi-task data integration engine provided by this application;
[0051] Figure 8 This is a schematic diagram of an intelligent allocation device for a multi-task data integration engine provided by the present application. DETAILED DESCRIPTION
[0052] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0053] With the deepening of digital transformation, the popularization of the Internet, the application of Internet of Things technology and the advancement of data collection technology, because different data integration tools have their own advantages and scope of application, in order to better meet different business needs, companies usually use multiple data integration tools at the same time.
[0054] However, this process often faces two common challenges: the high cost of getting started with data integration tools and the sheer diversity and similarity of these tools. The high cost of getting started with data integration tools stems primarily from the high cost of getting started with various ETL tools, which often makes it difficult for new users to create tasks and achieve fast and efficient task configuration. The sheer diversity and similarity of data integration tools stems from the fact that a wide variety of tools exist within an enterprise, leaving users without the technical expertise to determine which one to choose.
[0055] To address the high cost of getting started with data integration tools, an innovative visual multi-data source ETL tool has emerged. Its intuitive user interface significantly simplifies the process of synchronizing business data to the target database, enabling zero-programming data migration. However, despite its powerful and convenient offline ETL capabilities, data synchronization requirements in today's complex and ever-changing business scenarios often span real-time processing, offline batch processing, and complex master-slave table synchronization logic. These diverse requirements are difficult for a single tool to fully address.
[0056] To address the diversity and similarity of data integration tools, there is an existing ETL method and system for massive multi-source heterogeneous data that supports interface adaptation. This system primarily adaptively matches different data sources with corresponding ETL tools by setting basic information about the data source and target database. However, even after matching the corresponding ETL tool, users still need to configure parameters according to the specifications of each ETL tool due to differences in design architecture, configuration syntax, and supported data sources and target formats. This incompatibility limits users' ability to flexibly switch between or combine different data integration tools, increasing the complexity and cost of data integration work.
[0057] Although the above two existing methods can solve some problems, they cannot completely cover all types of data synchronization scenarios in an ETL tool environment, nor can they solve the problem of incompatible cross-platform data synchronization configurations.
[0058] Therefore, how to easily synchronize all types of data across platforms is a problem that needs to be solved.
[0059] First, to solve the above problems, the embodiments of the present application propose a method for intelligent allocation of a multi-task data integration engine, such as Figure 1 As shown, including:
[0060] S101, extracting metadata information of a data source to be allocated, data conversion information, server operation status, and task information of multiple task engines;
[0061] S102, the task scheduling model extracts non-time series features and time series features from metadata information, data conversion information, server operation status, and task information of multiple task engines;
[0062] S103, the task scheduling model determines an estimated completion time of each task engine among the multiple task engines according to the timing characteristics and the non-timing characteristics;
[0063] S104 , the task scheduling model allocates the data conversion task corresponding to the to-be-allocated data source to a target task engine according to the estimated completion time. The target task engine is the task engine with the shortest estimated completion time.
[0064] Metadata information includes metadata such as the source data's table structure, table list, field definitions, and fields. Data conversion information includes conversion mode information (mode), conversion plugin information (plugin), conversion process data source information (connection), and plugin connection information (hop). The server's operating status includes information about the server's computing resources, task execution, and operating environment. Current server computing resource information includes information about the server's CPU and currently idle CPUs. Server operating environment information includes information about temperature. Task execution information includes information about the computing resources occupied by each task currently executing on the server. Task engine task information includes information about the engine's current task load and the execution status of each task. Conversion mode information describes whether the process is real-time or offline. Conversion process data source information indicates the data source used throughout the conversion process. Conversion plugin information stores plugin information, including the necessary parameters for running the plugin. Plugin connection information indicates the connection between plugins. Conversion plugin information also includes information about plugins used to implement data operations, such as data sorting, data deduplication and merging, and data aggregation calculations, as well as information about plugins used to implement data conversion components, such as field name mapping, data dictionary conversion, and data encryption and decryption. The engine's current task volume includes the number of data conversion tasks currently being executed by each task engine and the time required.
[0065] Task completion information includes the task execution time for multiple task engines corresponding to the assigned data source, the server resources occupied by the task, and the total time required for the server to process the assigned data using the task engines in the current state. The total time is the sum of the waiting time and the number of task engines.
[0066] Before extracting the metadata information, data conversion information, server operation status and task information of multiple task engines of the data source to be allocated, it also includes: parsing the preset configuration file to obtain the data conversion information of the data source to be processed; scanning the data source to be processed to obtain the metadata information of the data source to be processed; obtaining the operation status of the server, wherein the operation status includes: the computing resource information of the server, the task execution information of the server and the operation environment information of the server; obtaining the task information of multiple task engines, wherein the task information includes the current task amount of each task engine and the execution status of the current task.
[0067] Scanning the data source to be processed includes: determining the location of the data source to be processed based on a preset configuration file, and scanning the data source to be processed. The preset configuration file includes pre-configured data source configuration information, such as the data source type, connection information, and driver information. The connection information includes, for example, the IP address, port number, username, and password.
[0068] Before the task scheduling model allocates the task corresponding to the data source to be allocated to the target task engine according to the expected completion time, it also includes: converting the preset configuration file of the data source to be allocated into a task configuration file for multiple task engines.
[0069] Converting a preset configuration file of a data source to be allocated into a task configuration file for multiple task engines includes: determining the task configuration file according to data conversion information of the data source to be processed.
[0070] Because the task scheduling model likely already parsed the preset configuration file to obtain the data conversion information for the data source to be processed before assigning the task corresponding to the data source to be assigned to the target task engine based on the estimated completion time, the preset configuration file for the data source to be assigned can be converted into a task configuration file for multiple task engines directly using the existing data conversion information. In the absence of data conversion information, before determining the task configuration file based on the data conversion information for the data source to be processed, the process also includes parsing the preset configuration file to obtain the data conversion information for the data source to be processed.
[0071] The task scheduling model extracts non-time series features and time series features from metadata information, data conversion information, the operating status of the server and the task information of multiple task engines, including: obtaining time series parameter variables and non-time series parameter variables from metadata information, data conversion information, the operating status of the server and the task information of multiple task engines; the task scheduling model extracts time series features from time series parameter variables and extracts non-time series features from non-time series parameter variables.
[0072] Before the task scheduling model determines the estimated completion time of each of the multiple task engines based on the timing characteristics and non-timing characteristics, it also includes: determining multiple state types of computing resources based on the running state; determining the configuration mapping table of the server based on the multiple state types and the configuration standards of the server; training the model to be trained based on the configuration standards, the configuration mapping table and historical data to obtain a trained task scheduling model.
[0073] The status type includes the current availability of computing resources. For example, if the current availability of computing resources is 100%, meaning no other tasks are occupying the server's computing resources, the status type can be set to Excellent, Class A, or Level 1, etc., to describe the server's computing resource availability's ability to support the task engine's task execution. The status type can be obtained by dividing the current availability of computing resources into multiple segments, with each status type (i.e., each segment) corresponding to a different level.
[0074] The task scheduling model determines the estimated completion time of each task engine in multiple task engines based on timing characteristics and non-timing characteristics, including: the task scheduling model determines the estimated completion time of each task engine in each server processing the allocated data source based on timing characteristics, non-timing characteristics, configuration standards and configuration mapping table.
[0075] In an implementation of the present application, a configuration mapping table of the server is determined based on multiple state types and configuration standards of the server, including: determining the energy consumption of each task engine corresponding to different state types and operating states to obtain a state efficiency table; determining the configuration standard of the server based on the state efficiency table; and determining the configuration mapping table of the server based on the configuration standard.
[0076] The energy consumption of the running state indicates the occupation of the server's computing resource information, the server's task execution information, the server's running environment information, and the like on the server's energy consumption.
[0077] Determining the configuration standard of the server according to the state efficiency table includes: determining the configuration of the server corresponding to the state energy efficiency according to the state energy efficiency table, and using the server configuration as the configuration standard.
[0078] Since state energy efficiency includes the corresponding amount of computing resources, it is necessary to determine the physical configuration required by the server based on the computing resources, such as memory, storage, and CPU, to support these computing resources.
[0079] The configuration standard is used to use a server configuration as the basic configuration standard for the server. The configuration standard can also select an existing server configuration based on the state performance table. The server configuration mapping table is used to map the configuration of the server that is different from the configuration standard according to the configuration standard, and obtain a configuration mapping table of multiple offsets of the server that is different from the configuration standard compared to the configuration standard. Among them, the multiple offsets include the offset of the server's hardware configuration, the offset of the server's response time, the offset of the server's operating status, and the offset of the execution time, resource consumption and other information required for the task engine to execute tasks on this server. The offset of the server's operating status includes: the offset of the server's computing resource information, the offset of the server's task execution information and the offset of the server's operating environment information.
[0080] The model to be trained is trained according to the configuration standard, the configuration mapping table and the historical data to obtain a trained task scheduling model, including: adjusting the preset loss function of the model to be trained according to the configuration standard to obtain the preset model; adjusting the historical data according to the configuration mapping table; using the adjusted historical data to input the training preset model to obtain the trained task scheduling model.
[0081] Among them, the preset model includes: a first sub-model and a second sub-model. Among them, the first sub-model includes a model corresponding to a scheduling algorithm framework based on deep learning, and the second sub-model includes one of a long short-term memory network (LSTM) model or a gated recurrent unit (GRU) model. Among them, the model corresponding to the scheduling algorithm framework based on deep learning includes: a scheduling algorithm framework based on deep learning such as a deep Q-network (DQN) or a policy gradient (Policy Gradient) algorithm. Therefore, the trained task scheduling model also includes the first sub-model and the second sub-model. For multiple non-time series parameter variables, a convolutional network in a model corresponding to a scheduling algorithm framework based on deep learning can be used to process the parameter variable to obtain multiple non-time series features; for multiple time series parameter variables, a long short-term memory network model or a gated recurrent unit model can be used to process the time series information in the time series parameter variable to obtain multiple time series features.
[0082] like Figure 2As shown, the adjusted historical data is used to train the preset model to obtain a trained task scheduling model, including: the preset model calculates the average time consumed by each historical task corresponding to the historical data source to be assigned in the historical data in all task engines (the average expected completion time), and the completion time of each historical task corresponding to each task engine, wherein the average expected completion time is determined according to the expected completion time. The historical task with the largest average expected completion time is assigned to the task engine with the shortest completion time (such as being placed in the planned execution list of the task engine with the shortest completion time), thereby excluding the historical task with the largest average expected completion time to form a new list of tasks to be assigned. Among them, for determining the expected completion time of each task engine in multiple task engines based on the timing features and the non-timing features, the task scheduling model can call the first sub-model to process and calculate the extracted non-timing features, and call the second sub-model to process and calculate the extracted timing features; the results obtained by processing the non-timing features and the timing features are further processed by the preset loss function in the preset model, thereby training and updating the model. The average time taken by each historical task in all task engines (the average estimated completion time) and the completion time of each historical task corresponding to each task engine are obtained, which includes: each historical task (r1, ..., r m ) in each task engine (x1, ..., x n ) in the estimated completion time, i.e., the time consumed; and the average time consumed Among them, the task scheduling model will process the same historical task (r1, ..., r m ) for each task engine (x1, ..., x n ) According to the time consumption (t 11 ,…,t 1n ) are arranged from small to large, and the average time taken by the same historical task in all task engines is also calculated. With each historical task (r1, ..., r m ) corresponds to the following:
[0083]
[0084] like Figure 2As shown, the server's computing resource information, the current task volume of each task engine, the execution status of the current task, and the task list are updated according to the task list to be assigned. Among them, the task list and other information also include the historical data source to be assigned corresponding to each task in the multiple tasks. Based on the updated computing resource information, the current task volume of each task engine and the execution status of the current task, the task list and other information, as well as the list of tasks to be assigned, the preset model is used to execute the steps of extracting non-time series features and time series features from the metadata information corresponding to the historical data, the data conversion information, the operating status of the server and the task information of multiple task engines; and the step of determining the expected completion time of each task engine in the multiple task engines based on the time series features and the non-time series features; thereafter, the average expected completion time is determined again based on the expected completion time; the historical task with the largest average expected completion time is assigned to the task engine with the earliest completion time, and the step of updating the server's computing resource information, the current task volume of each task engine and the execution status of the current task, as well as the task list and other information based on the list of tasks to be assigned is executed; the above steps are repeated until the expected completion time of each task engine is close (within the preset threshold range), and if the sum of the differences between the expected completion times of each task engine is less than the preset threshold, the training is stopped. According to the algorithm results (the estimated completion time of each historical task corresponding to each task engine), the algorithm result file is returned, where the algorithm result file includes a list of each historical task and its corresponding task engine), and a trained task scheduling model is obtained.
[0085] The configuration standard also includes task execution data. By evaluating and recording the response time, execution time, resource consumption and other information of a single task corresponding to different task engines under various conditions of sufficient and insufficient computing resources; and evaluating the change in the operating state of the server when different computing engines execute a single task under various conditions of sufficient and insufficient computing resources; the corresponding task execution data of different task engines executing a single task in different computing resources is obtained. Among them, the task execution data includes: obtaining the response time, execution time, resource consumption and other information of a single task when each task engine executes a single task in a variety of computing resources, as well as the change in the operating state of the server. The configuration mapping table also includes the offset of the task execution data.
[0086] The preset loss function of the model to be trained is adjusted through the configuration standard, and the historical data is adjusted according to the configuration mapping table so that the adjusted preset loss function can be applied to the historical data including the operating status of different servers and different task engines.
[0087] In an embodiment of the present application, after the task scheduling model assigns the data conversion task corresponding to the data source to be assigned to the target task engine according to the expected completion time, it also includes: using the target task engine to collect, integrate, convert and load the data source to be assigned according to the task configuration file.
[0088] Below, the implementation methods of this application are further described.
[0089] like Figure 3 As shown, the implementation method of the present application includes three parts: data source management layer, data processing layer and process creation layer.
[0090] At the data source management level, the implementation method of the present application includes a configuration interface for determining a preset configuration file of the data source to be allocated. The user configures the information of the data source to be allocated through the management interface to obtain the preset configuration file. Among them, the information of the data source to be allocated that needs to be configured includes: data source type, connection information (such as IP address, port number, user name, password, etc.) and driver information, etc. Afterwards, according to this preset configuration file, the connected data source is automatically scanned to obtain metadata information such as data table structure and fields, and store this information in the metadata warehouse. By providing standard API interfaces in the form of RESTful and WebService, other modules or upper-level applications are allowed to obtain relevant information of the data source through HTTP requests, such as data table lists, field definitions, etc., thereby achieving the purpose of shielding the differences between different data sources through a unified interface and data model, so that upper-level applications can access and operate data in a consistent manner.
[0091] At the process creation layer, determine the data conversion information in the preset configuration file. Through the various components provided in the graphical interface of the process creation layer, such as input components including library table input, real-time stream input, etc., output components including library table output, log output, etc., data operation components including data sorting, data deduplication and merging, data summary calculation, etc., data conversion components including field name mapping, data dictionary conversion, data encryption and decryption, etc., the synchronization tasks of the data source to be assigned are determined. By arranging various components and dragging and dropping to build processes, it is used to describe data synchronization tasks covering most scenarios such as offline and real-time. Among them, a json file is used as an example, in which mode is used to describe whether it is a real-time process or an offline process; connection is used to represent the data source information used in the entire conversion process; plugin is used to store plugin information, including the necessary parameters for running the plugin; hop is used to represent the connection information between plugins.
[0092] In the data processing layer, it is used to execute steps S101 to S104. First, the obtained preset configuration file is parsed by the intelligent conversion engine to extract key information therein, including conversion mode information, conversion plug-in information, conversion process data source information and plug-in connection information, or multiple timing parameter variables and multiple non-timing parameter variables corresponding to these information, for input into the task scheduling model. The task scheduling model extracts multiple timing features and multiple non-timing features from the multiple timing parameter variables and multiple non-timing parameter variables as inputs to the first sub-model and the second sub-model, and finally obtains the estimated completion time of each task engine in the multiple task engines, and determines the target task engine (ETL engine) that is most suitable for the data synchronization process of the data source to be assigned based on the estimated completion time.
[0093] like Figure 3 As shown, first, based on the evaluated and recorded changes in the server's operating state (server information and status, resource consumption, etc.) corresponding to the execution of synchronization tasks by different computing engines (task engines) under both sufficient and insufficient computing resources, as well as the response time (wait time) and execution time (time required to complete the synchronization task) of a single task, a server configuration is selected as a standard (configuration standard). Configuration mapping tables for each server are then set based on the configuration standard. Parameters such as response time and execution time can be ranges or averages.
[0094] Secondly, analyze the input parameters, that is, determine multiple first parameter variables from the metadata information of the data source to be allocated, data conversion information, the operating status of the server and the task information of multiple task engines according to the configuration file (preset configuration file), such as data conversion mode, data source related information, plug-in and related information, the current computing resource status of the server, the current task volume and execution status of each task engine, etc., wherein the multiple parameter variables include multiple timing parameter variables and multiple non-timing parameter variables.
[0095] Then, the first sub-model in the task scheduling model extracts features from multiple non-time-series parameter variables to obtain multiple non-time-series features, and the second sub-model in the task scheduling model extracts features from multiple time-series parameter variables to obtain multiple time-series features. The convolutional network of the first sub-model in the task scheduling model processes the metadata information of the data source to be assigned, data conversion information, the operating status of the server, and the multiple non-time-series parameter variables in the task information of multiple task engines, and the LSTM or GRU model in the second sub-model is used to process the time-series information extracted from the time-series parameter variables to obtain multiple time-series features. According to the server status information, the default parameter range and mean value of different status types are set, including determining the configuration standards of the task execution data and the server. The state value and configuration mapping table in the configuration standard are introduced into the loss function of the task scheduling model, and the configuration mapping is completed. Afterwards, the task scheduling model determines the expected completion time of each of the multiple task engines based on the time-series features and non-time-series features.
[0096] According to the feedback results of the task scheduling model, the task completion information (including the estimated completion time of each task engine) is obtained.
[0097] Finally, based on the estimated completion time, the task engine with the shortest estimated completion time is selected as the target task engine. The task profile, which the target task engine recognizes, is then sent to the target task engine for configuration. Finally, the target task engine processes the pending data source according to the task profile, completing the synchronization of the pending data source and completing the entire task.
[0098] Figure 5 As shown in the figure, take the conversion from unified configuration file (preset configuration file) to datax file as an example. Both datax and unified configuration files are in json format. You only need to extract connection, plugin and relation from the unified configuration file, and fill them into the readerPlugin, writerPlugin and setting parameters required by the datax configuration file after conversion. Figure 6 As shown, the unified configuration file can also be converted into a kettle configuration file in a similar way or as shown in Figure 7 The following figure shows the Flink SQL configuration file. The conversion process from a unified configuration file (preset configuration file) to a Kettle file or Flink SQL file is the same as the conversion process to a DataX file and is not repeated here.
[0099] Similarly, configuration files for Kettle, DataX, and Flink can be converted to a unified configuration file using the same method. Once the configuration files are successfully converted, the corresponding execution engines (computing engines) can execute the corresponding configuration files to perform data synchronization tasks.
[0100] The implementation method of the present application not only proposes a unified configuration file description framework that comprehensively covers data synchronization scenarios, but also proposes a two-way conversion mechanism between mainstream ETL tool configuration files and unified configuration files, as well as an algorithm for intelligently allocating ETL engines to parsing configuration files. The implementation method of the present application is centered on a user-friendly drag-and-drop interface, comprehensively and accurately depicting the entire ETL process to ensure that a wide range of needs and complex scenarios in the field of data synchronization are covered. Through this interface, users can easily build the required ETL tasks, and the system will automatically convert this task into a unified JSON format (preset configuration file) for storage at the bottom layer. When the user completes the design of the ETL scenario in the visual interface, the system immediately starts the intelligent analysis mechanism, and uses the task scheduling model to conduct in-depth analysis and calculation of the features obtained from the preset configuration file (such as JSON configuration) to select the ETL tool engine that is most suitable for the current task as the target task engine. Subsequently, the preset configuration file is automatically converted into the task configuration file required by the target task engine, and the corresponding ETL tool in the target task engine is directly called to process the data source to be allocated, achieving seamless connection from design to execution.
[0101] Second, as Figure 8 As shown, according to the embodiment of the present application, a multi-task data integration engine intelligent allocation device is also proposed, including:
[0102] The parameter extraction module 100 is used to extract metadata information of the data source to be allocated, data conversion information, the operating status of the server and task information of multiple task engines;
[0103] The task scheduling module 200 is used to extract non-temporal features and temporal features from metadata information, data conversion information, server operating status and task information of multiple task engines through a task scheduling model; the task scheduling model determines the estimated completion time of each task engine in the multiple task engines based on the temporal features and non-temporal features; the task scheduling model assigns the data conversion task corresponding to the data source to be assigned to the target task engine according to the estimated completion time, and the target task engine is the task engine with the shortest estimated completion time.
[0104] In the implementation mode of the present application, it also includes: a preprocessing module, which scans the data source to be processed and obtains metadata information of the data source to be processed; parses the preset configuration file to obtain data conversion information of the data source to be processed; obtains the running status of the server, wherein the running status includes: the computing resource information of the server, the task execution information of the server and the running environment information of the server; obtains the task information of multiple task engines, and the task information includes the current task amount of each task engine and the execution status of the current task.
[0105] The pre-processing module is further used to convert the preset configuration file of the data source to be allocated into a task configuration file for multiple task engines.
[0106] The implementation manner of the present application also includes: a data processing module that uses a target task engine to collect, integrate, convert and load data sources according to a task configuration file.
[0107] In the implementation manner of the present application, it further includes: a configuration module, which is used to obtain a preset configuration file of the data source to be allocated.
[0108] In the implementation of the present application, by extracting metadata information, data conversion information, server operating status and task information of multiple task engines of the data source to be assigned, a task scheduling model is used to extract non-time series features and time series features, so that non-time series features and time series features can be extracted from the data source to be assigned of all types of data and the data conversion information corresponding to various platforms; the estimated completion time of each task engine in the multiple task engines is determined based on the extracted non-time series features and time series features; the data conversion task corresponding to the data source to be assigned is assigned to the task engine with the shortest estimated completion time, so that the target task engine can be conveniently determined based on the data source to be assigned. Finally, a task configuration file that can be recognized by the target task engine is determined based on the preset configuration file, and the target task engine is used to collect, integrate, convert and load the data source to be assigned based on the task configuration file, so that cross-platform synchronization of all types of data can be intelligently achieved, and it is simple and convenient. The implementation of the present application realizes automatic selection of the target task engine through the task scheduling model, thereby reducing the learning cost of using ETL tools and intelligently achieving the determination of the target task engine, which is convenient and fast. Using the target task engine to automatically collect, integrate, convert, and load data sources according to preset configuration files can further simplify the process of data synchronization for the assigned data sources, thereby further reducing labor costs and improving practicality. The implementation method of the present application also has strong compatibility and can host tasks created by mainstream ETL tools on the platform through conversion, so that these tasks can also be managed and displayed under a unified interface, further improving the transparency and controllability of the data integration process. The implementation method of the present application enhances process flexibility and maintainability. By using a unified JSON format to describe the ETL process, the system not only achieves configuration standardization and modularization, but also makes the adjustment, optimization, and expansion of the process easier. At the same time, the flexibility of using JSON format as a preset configuration file also supports more complex synchronization logic and data conversion requirements, ensuring the efficient operation and long-term maintenance of the process. Finally, before collecting, integrating, converting, and loading the data source, the preset configuration file in JSON format is converted into a task configuration file that can be recognized by the target task engine, thereby improving the degree of automation of the target task engine's processing. The implementation method of the present application can also intelligently match the optimal ETL tool (target task engine), and can automatically select the most suitable ETL tool engine according to the ETL scenario constructed by the user through the system's built-in intelligent analysis engine (task scheduling model), thereby maximizing the use of existing resources and improving data processing efficiency and performance. This intelligent tool selection mechanism avoids manual trial and error and unnecessary waste of resources. The implementation method of the present application also improves the comprehensiveness and compatibility of data synchronization. The implementation party of the present application not only supports current mainstream ETL tools, but also achieves seamless compatibility with customized preset configuration files through the configuration file conversion mechanism.This enables users to easily migrate existing ETL tasks to the platform for unified management while maintaining the ability to quickly adapt to emerging tools and technologies. The implementation methods of this application also enhance process transparency and controllability. By displaying the execution status, data flow, and processing results of the ETL process through a unified interface, users can monitor and track every aspect of data synchronization in real time. This high degree of transparency and controllability helps to discover and resolve problems in a timely manner, ensuring the accuracy and reliability of data synchronization.
[0109] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for intelligently allocating a multi-task data integration engine, characterized in that: include: Extract metadata information of the data source to be assigned, data conversion information, server operation status and task information of multiple task engines; The task scheduling model extracts non-time series features and time series features from the metadata information, the data conversion information, the running status of the server and the task information of multiple task engines; The task scheduling model determines an estimated completion time of each task engine in the plurality of task engines according to the timing feature and the non-timing feature; The task scheduling model allocates the data conversion task corresponding to the to-be-allocated data source to a target task engine according to the estimated completion time, and the target task engine is the task engine with the shortest estimated completion time.
2. The intelligent allocation method according to claim 1, characterized in that: Before the task scheduling model allocates the task corresponding to the to-be-allocated data source to the target task engine according to the estimated completion time, the method further includes: converting a preset configuration file of the to-be-allocated data source into a task configuration file for the multiple task engines.
3. The intelligent allocation method according to claim 1, characterized in that: The task scheduling model extracts non-time series features and time series features from the metadata information, data conversion information, the running status of the server, and task information of multiple task engines, including: Acquire timing parameter variables and non-timing parameter variables from the metadata information, the data conversion information, the running status of the server, and the task information of the plurality of task engines; The task scheduling model extracts the timing features from the timing parameter variables and extracts the non-timing features from the non-timing parameter variables.
4. The intelligent allocation method according to claim 1, characterized in that: Before the task scheduling model determines the estimated completion time of each task engine in the plurality of task engines according to the timing feature and the non-timing feature, the method further includes: determining a plurality of status types of the computing resource according to the operating status; Determining a configuration mapping table of the server according to the multiple status types and the configuration standard of the server; The model to be trained is trained according to the configuration standard, the configuration mapping table and historical data to obtain the trained task scheduling model.
5. The intelligent allocation method according to claim 4, characterized in that: The task scheduling model determines an estimated completion time of each task engine in the plurality of task engines according to the timing feature and the non-timing feature, including: The task scheduling model determines the estimated completion time of each task engine in each server processing the to-be-assigned data source according to the timing characteristics, the non-timing characteristics, the configuration standard, and the configuration mapping table.
6. The intelligent allocation method according to claim 4, characterized in that: The determining the configuration mapping table of the server according to the multiple status types and the configuration standard of the server includes: Determine the energy consumption of each task engine corresponding to the different state types and the operating states, and obtain a state energy consumption table; Determining a configuration standard of the server according to the status performance table; A configuration mapping table of the server is determined according to the configuration standard.
7. The intelligent allocation method according to claim 4, characterized in that: The step of training the model to be trained according to the configuration standard, the configuration mapping table, and the historical data to obtain the trained task scheduling model includes: Adjust the preset loss function of the model to be trained according to the configuration standard to obtain a preset model; Adjusting the historical data according to the configuration mapping table; The adjusted historical data is used to train the preset model to obtain the trained task scheduling model.
8. The intelligent allocation method according to claim 2, characterized in that: After the task scheduling model allocates the data conversion task corresponding to the to-be-allocated data source to the target task engine according to the estimated completion time, the method further includes: The target task engine is used to collect, integrate, convert and load the data source to be allocated according to the task configuration file.
9. The intelligent allocation method according to claim 1, characterized in that: Before extracting metadata information of the data source to be allocated, data conversion information, the running status of the server, and task information of multiple task engines, the process further includes: Parsing a preset configuration file to obtain data conversion information of the data source to be processed; Scan the data source to be processed to obtain metadata information of the data source to be processed; Acquiring the running status of the server, wherein the running status includes: computing resource information of the server, task execution information of the server, and running environment information of the server; The task information of the multiple task engines is acquired, where the task information includes a current task amount of each task engine and an execution status of the current task.
10. An intelligent allocation device for a multi-task data integration engine, characterized in that: include: Parameter extraction module, used to extract metadata information of the data source to be assigned, data conversion information, server operation status and task information of multiple task engines; A task scheduling module is configured to extract non-time series features and time series features from the metadata information, the data conversion information, the operating status of the server, and the task information of the plurality of task engines through a task scheduling model; the task scheduling model determines an estimated completion time of each of the plurality of task engines based on the time series features and the non-time series features; The task scheduling model allocates the data conversion task corresponding to the to-be-allocated data source to a target task engine according to the estimated completion time, and the target task engine is the task engine with the shortest estimated completion time.