Multi-source heterogeneous offline task processing method, system, device and medium
Patent Information
- Application Number
- CN202310952399.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2043-07-31
AI Technical Summary
[0002]随着5G技术在工业领域的广泛应用,边缘端的工业设备类型逐渐呈现多样化,其中,业务端、应用端或边缘端,与云端或大数据平台之间海量数据离线集成过程中,存在着格式不一、多源异构、数据杂乱和难以溯源等特性,尤其是在复杂流程的数据加工中,大多数集成系统则是任务复杂、流程无序、难以管理且开发低效
[0035] The aforementioned multi-source heterogeneous offline task processing methods, systems, devices, and media, through the design of the SeaTunnel offline plugin model, decompose the entire Spark data integration process to ensure compatibility with multiple heterogeneous data sources, eliminate storage format inconsistencies, and simplify the Spark data integration process. A DAG-based visual configuration reduces user complexity, while dynamic orchestration of plugin relationships increases code reusability and avoids redundant code development. A seamless DFS task construction algorithm is designed to reduce manual intervention. Using a DFS-based algorithm within a DAG (Directed Acyclic Graph), SeaTunnel task construction and task dependency relationship establishment can be completed quickly and effectively, significantly reducing the complexity of the task flow. Ultimately, SeaTunnel offline task construction achieves integrated Spark data processing across multiple stages, effectively reducing the difficulty of using Spark technology. It establishes a unified offline synchronization model, batch processing method, and efficient data development tools, forming a simple and efficient one-stop offline task construction standard solution. This reduces the technical difficulty of Spark task configuration, ensures compatibility with multiple heterogeneous data sources, eliminates data format limitations, and significantly improves the compatibility of multi-source heterogeneous offline task processing.
Smart Images

Figure CN117055970B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology and relates to a method, system, device and medium for processing multi-source heterogeneous offline tasks. Background Technology
[0002] With the widespread application of 5G technology in the industrial field, the types of industrial equipment at the edge are gradually becoming more diversified. Among them, the offline integration of massive data between the business end, application end, or edge end and the cloud or big data platform is characterized by inconsistent formats, heterogeneous sources, messy data, and difficulty in tracing the source. Especially in the data processing of complex processes, most integration systems are characterized by complex tasks, disordered processes, difficulty in management, and inefficient development.
[0003] Currently, the Spark computing engine is a key technology for offline data processing. Many scholars have conducted in-depth research on Spark offline task construction methods, such as methods to construct Spark offline tasks by configuring multiple data source plugins, methods to construct Spark offline tasks by defining task component configurations and relationships through the interface, methods to construct Spark offline tasks by configuring data source components through the interface, and methods to construct Spark offline tasks by obtaining the system resources required by Linux commands. However, the aforementioned traditional technologies still have technical problems of insufficient task processing compatibility when facing heterogeneous offline data processing tasks from a large number of cloud-edge industrial devices. Summary of the Invention
[0004] To address the problems existing in the above-mentioned traditional methods, this invention proposes a multi-source heterogeneous offline task processing method, a multi-source heterogeneous offline task processing system, a computer device, and a computer-readable storage medium, which can significantly improve task processing compatibility.
[0005] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:
[0006] On the one hand, a method for processing multi-source heterogeneous offline tasks is provided, including the following steps:
[0007] Add input and output sources on the data source management page for multi-source heterogeneous data, configure data source connection information, pass connectivity tests, and save the connection information to the data source table; input sources include MySQL data source, DB2 data source, MongoDB data source, and Doris data source, and output sources include Hive data source, Phoenix data source, and ClickHouse data source;
[0008] On the dataset management page for multi-source heterogeneous data, select a data table, obtain the metadata configuration information of the data table, and save it to the dataset table and data field table; the metadata configuration information includes the data table name, data fields, field types, and primary key fields;
[0009] Create project projects and workflow canvases corresponding to multi-source heterogeneous data, save the project projects to the project project table and the workflow canvas to the workflow table;
[0010] The Spark data integration process is decomposed using an offline integration model based on unit plugin decomposition, and four types of integration plugins are generated in the data integration process: execution environment plugins, input plugins, transformation plugins, and output plugins.
[0011] Create an offline integration pipeline based on the SeaTunnel offline integration pipeline model in the workflow canvas; the pipeline modules of the offline integration pipeline include the Spark execution environment configuration module, the Spark input module, the Spark transformation module, and the Spark output module;
[0012] In the SeaTunnel offline integration pipeline, the execution environment plugin, input plugin, transformation plugin, and output plugin are dynamically selected and their parameter information is configured based on the DAG drag-and-drop operation. The execution environment plugin is selected as the Spark execution environment plugin, the input plugin is selected as the MySQL data source plugin, the DB2 data source plugin, or the MongoDB data source plugin, the transformation plugin is selected as the SQL transformation plugin, and the output plugin is selected as the Hive data source plugin or the Phoenix data source plugin. The plugin parameter information includes execution environment parameters, input table parameters, SQL statements, and output table parameters.
[0013] Use the data integration partitioning model to select the data partitioning mode from the target table configuration of the output plugin;
[0014] Based on the offline integration plug-in relationship orchestration model, the DAG node connection method is used to dynamically orchestrate the four types of integration plug-in relationships in the offline integration pipeline, and the DFS algorithm is used to verify the non-closed loop of plug-in relationships, generating a DAG integration pipeline directed acyclic graph.
[0015] Build a SeaTunnel offline task configuration template based on Spark's task parameter configuration template;
[0016] Based on the DAG integrated pipeline directed acyclic graph, the offline integrated pipeline is converted into a Task entity using a pipeline task entity transformation model based on the DFS algorithm and entered into the task definition table. Additionally, a task configuration instance is generated based on the SeaTunnel offline task configuration template.
[0017] Based on the DAG integrated pipeline directed acyclic graph, the pipeline task dependency transformation model based on the DFS algorithm is used to convert the plug-in relationship of the offline integrated pipeline into the task dependency relationship of the Task entity and record it into the task dependency table.
[0018] Based on the SeaTunnel scheduling task instance model, the Task entity, Task configuration instance, and Task dependency relationship are combined to construct the SeaTunnel scheduling task instance, thus completing the SeaTunnel offline task scheduling.
[0019] On the other hand, a multi-source heterogeneous offline task processing system is also provided, including:
[0020] The source configuration module is used to add input and output sources in the data source management page for multi-source heterogeneous data, configure data source connection information, pass connectivity tests, and save the connection information to the data source table; input sources include MySQL data source, DB2 data source, MongoDB data source, and Doris data source, and output sources include Hive data source, Phoenix data source, and ClickHouse data source;
[0021] The metadata management module is used to select a data table in the dataset management page for multi-source heterogeneous data, obtain the metadata configuration information of the data table, and save it to the dataset table and the data field table; the metadata configuration information includes the data table name, data fields, field types, and primary key fields;
[0022] The job creation module is used to create project projects and workflow canvases corresponding to multi-source heterogeneous data, save project projects to the project project table and workflow canvases to the workflow table;
[0023] The integration decomposition module is used to decompose the Spark data integration process using an offline integration model based on unit plugin decomposition, and to generate four types of integration plugins in the data integration process: execution environment plugins, input plugins, transformation plugins, and output plugins.
[0024] The pipeline creation module is used to create offline integration pipelines based on the SeaTunnel offline integration pipeline model in the workflow canvas. The pipeline module of the offline integration pipeline includes the Spark execution environment configuration module, the Spark input module, the Spark transformation module, and the Spark output module.
[0025] The plugin configuration module is used to dynamically select execution environment plugins, input plugins, transformation plugins, and output plugins based on DAG drag-and-drop operations within the module composition of the SeaTunnel offline integration pipeline, and configure plugin parameter information. The execution environment plugin is selected as the Spark execution environment plugin, the input plugin is selected as the MySQL data source plugin, the DB2 data source plugin, or the MongoDB data source plugin, the transformation plugin is selected as the SQL transformation plugin, and the output plugin is selected as the Hive data source plugin or the Phoenix data source plugin. The plugin parameter information includes execution environment parameters, input table parameters, SQL statements, and output table parameters.
[0026] The partition selection module is used to select a data partitioning mode from the target table configuration of the output plugin using the data integration partitioning model;
[0027] The plug-in orchestration module is used to dynamically orchestrate the four types of integrated plug-in relationships in the offline integrated pipeline based on the offline integrated plug-in relationship orchestration model, using the DAG node connection method, and to perform non-closed-loop verification of plug-in relationships through the DFS algorithm, generating a DAG integrated pipeline directed acyclic graph.
[0028] The template building module is used to build SeaTunnel offline task configuration templates based on Spark task parameter configuration templates.
[0029] The entity encapsulation module is used to convert offline integrated pipelines into Task entities and enter them into the task definition table based on the DAG integrated pipeline directed acyclic graph and pipeline task entity transformation model based on DFS algorithm, and to generate task configuration instances based on SeaTunnel offline task configuration template.
[0030] The dependency conversion module is used to convert the plug-in relationships of the offline integrated pipeline into the task dependency relationships of the Task entities based on the DAG integrated pipeline directed acyclic graph and the pipeline task dependency conversion model based on the DFS algorithm, and record them into the task dependency table.
[0031] The task scheduling module is used to combine the Task entity, task configuration instance, and task dependency relationship according to the SeaTunnel scheduling task instance model to construct the SeaTunnel scheduling task instance and complete the SeaTunnel offline task scheduling.
[0032] In another aspect, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described multi-source heterogeneous offline task processing method.
[0033] Furthermore, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described multi-source heterogeneous offline task processing method.
[0034] One of the above technical solutions has the following advantages and beneficial effects:
[0035] The aforementioned multi-source heterogeneous offline task processing methods, systems, devices, and media, through the design of the SeaTunnel offline plugin model, decompose the entire Spark data integration process to ensure compatibility with multiple heterogeneous data sources, eliminate storage format inconsistencies, and simplify the Spark data integration process. A DAG-based visual configuration reduces user complexity, while dynamic orchestration of plugin relationships increases code reusability and avoids redundant code development. A seamless DFS task construction algorithm is designed to reduce manual intervention. Using a DFS-based algorithm within a DAG (Directed Acyclic Graph), SeaTunnel task construction and task dependency relationship establishment can be completed quickly and effectively, significantly reducing the complexity of the task flow. Ultimately, SeaTunnel offline task construction achieves integrated Spark data processing across multiple stages, effectively reducing the difficulty of using Spark technology. It establishes a unified offline synchronization model, batch processing method, and efficient data development tools, forming a simple and efficient one-stop offline task construction standard solution. This reduces the technical difficulty of Spark task configuration, ensures compatibility with multiple heterogeneous data sources, eliminates data format limitations, and significantly improves the compatibility of multi-source heterogeneous offline task processing. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart illustrating a multi-source heterogeneous offline task processing method in one embodiment;
[0038] Figure 2 This is a schematic diagram of the SeaTunnel offline integration pipeline in one embodiment;
[0039] Figure 3 This is a schematic diagram illustrating the classification and relationship of offline plugins in one embodiment;
[0040] Figure 4 This is a schematic diagram illustrating the dynamic arrangement of offline plugin relationships in one embodiment;
[0041] Figure 5 This is a schematic diagram of the DFS offline task construction process in one embodiment;
[0042] Figure 6 This is a schematic diagram of the module composition of a multi-source heterogeneous offline task processing system in one embodiment. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0045] It should be noted that, in this document, the reference to "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The presentation of this phrase in various locations throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments.
[0046] Those skilled in the art will understand that the embodiments described herein can be combined with other embodiments. The term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items, and all possible combinations thereof.
[0047] Spark offline tasks refer to a class of tasks running on the Apache Spark distributed computing framework. These tasks do not require real-time response or instant computation results, but rather perform batch processing operations on large-scale datasets. Offline tasks typically run in the background and can process large amounts of data, performing complex computations and analyses. In developing this invention, the inventors discovered that traditional methods for constructing Spark offline tasks by configuring multiple data source plugins suffer from a lack of task dependency orchestration, leading to chaotic execution processes; methods that define task component configurations and relationships through a user interface do not decouple task component configurations, resulting in low code reusability due to repetitive configurations; methods that configure data source components through a user interface require customized page task configurations and are cumbersome to operate; and methods that obtain system resources through Linux commands result in poor interactivity due to Linux script command coding.
[0048] Against this backdrop, most Spark parameter configuration schemes still have some practical problems, such as: some methods have a high degree of data format customization, or have few types of data source plugins and poor extensibility, or lack of plugin configuration and poor code reusability, or have customized page configuration and extremely poor user experience, or lack of task decoupling and cumbersome configuration, or lack of task dependency relationships and chaotic process, or manual intervention and easy to make mistakes, or lack of user permission control and chaotic management.
[0049] SeaTunnel is already an important offline data processing platform for multi-source heterogeneous ETL (Extract-Transform-Load, which describes the process of extracting, transforming, and loading data from the source to the destination), and it uses the Spark computing engine to achieve massive data integration. However, most methods still use traditional manual configuration, which leads to problems such as high customization of SeaTunnel task configuration, low development efficiency, and difficulty in management. Therefore, it is crucial to create a Spark offline task building solution and system based on SeaTunnel.
[0050] To address the shortcomings of traditional technologies, this invention aims to create a simple, efficient, and highly compatible one-stop offline task construction standard solution based on the SeaTunnel multi-source heterogeneous offline task construction approach. This solution reduces the technical difficulty of configuring Spark offline tasks, is compatible with multiple heterogeneous data sources, and eliminates data format limitations. It designs a SeaTunnel offline plugin model to simplify the Spark data integration process; DAG visual configuration reduces user operation difficulty; dynamic plugin orchestration increases code reusability and avoids redundant code development; a seamless DFS task construction algorithm reduces manual intervention; and DFS task dependency construction reduces process complexity. Ultimately, SeaTunnel task construction integrates Spark data processing across multiple stages, reducing the difficulty of using Spark technology, establishing a unified offline synchronization model, batch processing method, and efficient data development approach, and improving task processing compatibility and efficiency.
[0051] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0052] Please see Figure 1 In one embodiment, a multi-source heterogeneous offline task processing method is provided, including the following processing steps S1 to S12:
[0053] S1. In the data source management page for multi-source heterogeneous data, add input and output sources, configure data source connection information, pass connectivity tests, and save the connection information to the data source table. Input sources include MySQL data source, DB2 data source, MongoDB data source, and Doris data source, while output sources include Hive data source, Phoenix data source, and ClickHouse data source.
[0054] It is understandable that for the various industrial devices in the current application scenario, which serve as different heterogeneous data sources, when building an offline task system based on the SeaTunnel platform for this task scenario, input sources such as MySQL data source, DB2 data source, MongoDB data source, and Doris data source are added to the data source management page for multi-source heterogeneous data. Output sources such as Hive data source, Phoenix data source, and ClickHouse data source are added. Data source connection information is configured, connectivity tests are performed, and this connection information is saved to the system's data source table, thus completing the import of offline data source connection configuration information and connectivity testing.
[0055] S2, in the dataset management page for multi-source heterogeneous data, select a data table, obtain the metadata configuration information of the data table and save it to the dataset table and data field table; the metadata configuration information includes the data table name, data fields, field types and primary key fields.
[0056] It is understandable that in the system's multi-source heterogeneous data dataset management page, you can select data tables from input sources such as MySQL data source, DB2 data source, MongoDB data source, and Doris data source, as well as data table metadata configuration information such as data table name, data fields, field types, and primary key fields. This metadata configuration information is then saved to the system's dataset table and data field table to complete the import of data table and field metadata information from offline data sources.
[0057] S3 creates project engineering and workflow canvas corresponding to multi-source heterogeneous data, saves project engineering to project engineering table and workflow canvas to workflow table.
[0058] It is understandable that after importing the aforementioned information, you can also create project projects and workflow canvases, and save this information to the system's project project table and workflow table for later use.
[0059] S4 utilizes an offline integration model based on unit plugin decomposition to decompose the Spark data integration process and generates four types of integration plugins in the data integration process: execution environment plugins, input plugins, transformation plugins, and output plugins.
[0060] Understandably, the next step is to design SeaTunnel offline integration plugin types, specifically four plugin types (execution environment plugin E, input plugin I, transformation plugin T, and output plugin O), and decompose the entire Spark data integration process to decouple the offline task processing flow. The offline integration model based on unit plugin decomposition used can be illustrated as follows:
[0061]
[0062] In this system, PlugIn() represents the offline integration model, where n1 is the number of input plugins, n2 is the number of transformation plugins, n3 is the number of output plugins, i is any value within the specified range, E represents the execution environment plugin, I represents the input plugin, T represents the transformation plugin, and O represents the output plugin. Correspondingly, SeaTunnel offline data source plugins were also designed, including plugins for MySQL, DB2, MongoDB, Hive, Phoenix, Doris, and ClickHouse, and their basic information is entered into the system's plugin configuration table. Through the SeaTunnel offline data plugin design, the system is compatible with multiple heterogeneous data sources, eliminating limitations imposed by inconsistent storage formats.
[0063] S5 creates an offline integration pipeline based on the SeaTunnel offline integration pipeline model in the workflow canvas; the pipeline modules of the offline integration pipeline include the Spark execution environment configuration module, the Spark input module, the Spark transformation module, and the Spark output module.
[0064] Understandably, the next step is to create the SeaTunnel offline integration pipeline. Within the workflow canvas described above, create an offline integration pipeline consisting of four modules: Spark execution environment configuration, Spark input, Spark transformation, and Spark output. Figure 2 As shown in Figure 2, the defined SeaTunnel offline integration pipeline model can be represented as follows:
[0065]
[0066] Where PipeLine() is the offline integration pipeline, m1 is the number of input modules, m2 is the number of transformation modules, m3 is the number of output modules, SE is the Spark execution environment configuration module, SI is the Spark input module, ST is the Spark transformation module, SO is the Spark output module, and the arrows indicate the direction of data flow in the pipeline.
[0067] In S6, within the module composition of the SeaTunnel offline integration pipeline, the execution environment plugin, input plugin, transformation plugin, and output plugin are dynamically selected and their parameter information is configured based on the DAG drag-and-drop operation. The execution environment plugin is selected as the Spark execution environment plugin, the input plugin is selected as the MySQL data source plugin, the DB2 data source plugin, or the MongoDB data source plugin, the transformation plugin is selected as the SQL transformation plugin, and the output plugin is selected as the Hive data source plugin or the Phoenix data source plugin. The plugin parameter information includes execution environment parameters, input table parameters, SQL statements, and output table parameters.
[0068] Understandably, the next step is to configure the SeaTunnel pipeline plugins. In the offline integration pipeline, the plugins for the pipeline components are specified. Based on the DAG drag-and-drop method, the execution environment plugin for the Spark execution environment configuration module, the input plugin for the Spark input module, the transformation plugin for the Spark transformation module, and the output plugin for the Spark output module are configured, and the corresponding plugin parameters are pre-configured.
[0069] Specifically, selecting an execution environment plugin for the Spark execution environment configuration module of the SeaTunnel offline integration pipeline involves specifying the Spark execution environment configuration module in step S6 of the offline integration pipeline, dynamically selecting the Spark execution environment plugin based on a DAG drag-and-drop method, and pre-configuring its execution environment parameters. Selecting an input plugin for the Spark input module of the SeaTunnel offline integration pipeline involves specifying the Spark input module in step S6 of the offline integration pipeline, dynamically selecting a MySQL data source plugin, a DB2 data source plugin, or a MongoDB data source plugin based on a DAG drag-and-drop method, and pre-configuring its input table parameters. Selecting a transformation plugin for the Spark transformation module of the SeaTunnel offline integration pipeline involves specifying the Spark transformation module in step S6 of the offline integration pipeline, dynamically selecting the SQL converter (i.e., the SQL transformation plugin) based on a DAG drag-and-drop method, and pre-configuring the SQL statements. For example... Figure 3 As shown, selecting the output plugin for the Spark output module in the SeaTunnel offline integration pipeline involves specifying the Spark output module in step S6 of the offline integration pipeline. The Hive / Phoenix data source plugin is dynamically selected using a drag-and-drop DAG approach, and its target table parameters are pre-configured. This SeaTunnel visual plugin parameter configuration enables low-code development and simplifies the Spark offline task building process.
[0070] S7 utilizes the data integration partitioning model to select the data partitioning mode from the target table configuration of the output plugin.
[0071] Understandably, after configuring the plugin, you need to select the partitioning mode for the Hive / Phoenix output plugin. Here, a data integration partitioning model is used to select a partitioning model (e.g., full model FT) in the target table configuration. This data integration partitioning model can be shown as follows:
[0072]
[0073] In this function, TablePartition() represents the data partitioning model, n is the number of integrated pipelines, m is the number of output plugins, T is the target table, and a1 and a2 are integer parameters, with a1 + a2 = 1. Furthermore, the output plugins can be configured with either full partitioning (FT) or partitioning (PT) modes (e.g., by day, month, or year) to facilitate data partitioning and storage.
[0074] S8. Based on the offline integration plug-in relationship orchestration model, the DAG node connection method is used to dynamically orchestrate the four types of integration plug-in relationships in the offline integration pipeline, and the DFS algorithm is used to verify the non-closed loop of the plug-in relationship to generate a DAG integration pipeline directed acyclic graph.
[0075] Understandably, the next step is to design the relationships between the various integration plugins in the SeaTunnel offline integration pipeline. These relationships can include: input plugin → transformation plugin, transformation plugin → output plugin, and input plugin → output plugin. The offline integration plugin relationship orchestration model used can be shown in Figure 4 below:
[0076]
[0077] In this model, `PlugInRelation()` represents the plugin relationship orchestration model, where `n1` is the number of input plugins, `n2` is the number of transformation plugins, `n3` is the number of output plugins, `E` represents the execution environment plugin, `I` represents the input plugin, `T` represents the transformation plugin, and `O` represents the output plugin. Thus, the plugin relationships between these four types of plugins are dynamically orchestrated using node connections. A Depth-First Search (DFS) algorithm is used to verify non-loop relationships, ultimately generating a Directed Acyclic Graph (DAG) in the integration pipeline, as shown below. Figure 4 As shown, dynamic orchestration of SeaTunnel integration plugin relationships simplifies the dependency configuration for Spark offline tasks.
[0078] S9 builds the SeaTunnel offline task configuration template based on Spark's task parameter configuration template.
[0079] It can be understood that after generating the DAG (Directed Acyclic Graph) integrated pipeline, tasks can be created in the graph using output nodes. Then, based on Spark's task parameter configuration template, SeaTunnel offline task configuration parsing is performed to parse out four types of plugins: environment plugins, input plugins, processing plugins, and output plugins. The attributes of these plugins are recorded in the system's task definition table. The Spark-based task parameter configuration template can be represented as follows:
[0080] {env{spark.app.name="task-seatunnel-spark-0-0001",…},source{MySQL{…}},transform{sql{…}},sink{Phoenix{…}}}.
[0081] In this configuration, `env{…}` represents the Spark execution environment module configuration, `app` represents the application, `name` represents the name, `source{…}` represents the input source module configuration, `transform{…}` represents the data transformation module configuration, `sink{…}` represents the output source module configuration, and `MySQL`, `sql`, and `Phoenix` are plugin definition identifiers. Following the SeaTunnel task configuration specification, a SeaTunnel offline task configuration template is constructed by defining a Spark-based task parameter configuration template.
[0082] S10: Based on the DAG integrated pipeline directed acyclic graph, the offline integrated pipeline is converted into a Task entity using the pipeline task entity transformation model based on the DFS algorithm and entered into the task definition table. Additionally, a task configuration instance is generated based on the SeaTunnel offline task configuration template.
[0083] It is understandable that the DFS algorithm, which is the depth-first traversal algorithm of a binary tree, requires the SeaTunnel integration plugin relationship parsing here. That is, from the DAG integrated pipeline directed acyclic graph, the pipeline task entity transformation model based on the DFS algorithm is used to parse out all the task entities and task configurations in the integrated pipeline, and these task entities and task configurations are entered into the task definition table.
[0084] Specifically, starting with the output module plugin in the pipeline, a Depth-First Search (DFS) pre-traversal is used to find the associated input module plugins, transformation module plugins, and plugin configurations. These integrated plugins are then converted into Task entities, and task configuration instances are generated based on the SeaTunnel task configuration template. The pipeline task entity transformation model based on the DFS algorithm is defined as shown in Figure 5 below:
[0085]
[0086] Wherein, EntityShift() is the entity transformation model, k1 is the number of input module plugins, k2 is the number of transformation module plugins, k3 is the number of output module plugins, SPI is the input module plugin, SPT is the transformation module plugin, SPO is the output module plugin, Task is the task entity, Config is the task configuration instance, and the arrow indicates the task transformation direction.
[0087] S11. Based on the DAG integrated pipeline directed acyclic graph, the pipeline task dependency transformation model based on the DFS algorithm is used to convert the plug-in relationship of the offline integrated pipeline into the task dependency relationship of the Task entity and record it into the task dependency table.
[0088] It can be understood that in the offline integration pipeline, starting with the output module plugin, a Depth-First Search (DFS) pre-traversal is used to find other associated output module plugins and establish the dependency relationship (Lineage) of the task entity (Task). The pipeline task dependency transformation model based on the DFS algorithm is defined as shown in Figure 6 below:
[0089]
[0090] In this table, Lineage() represents the task dependency relationship, k3 represents the number of output module plugins, SPO represents the output module plugin, Task represents the task entity, and the variable i can be any value within the specified range. The arrow indicates the direction of dependency transformation. After completing the above processing steps, the pipeline task dependency transformation model based on the DFS algorithm is used to convert the output plugin relationships in the offline integration pipeline into task entity dependencies and record them in the system's task dependency table, such as... Figure 5 As shown.
[0091] S12, based on the SeaTunnel scheduling task instance model, combine the Task entity, Task configuration instance, and Task dependency relationship to construct the SeaTunnel scheduling task instance, and complete the SeaTunnel offline task scheduling.
[0092] It can be understood that, based on the defined SeaTunnel scheduling task instance model, the aforementioned Task entities, task configuration instances, and task dependencies are combined to construct the SeaTunnel scheduling task instance in the current application scenario, and to implement the final SeaTunnel task scheduling. The defined SeaTunnel scheduling task instance model can be shown as follows:
[0093]
[0094] Here, SeaTunnel() represents the structure of the scheduling task, l1 represents the number of task entities, and l2 represents the number of task dependencies.
[0095] The aforementioned multi-source heterogeneous offline task processing method decomposes the entire Spark data integration process by designing the SeaTunnel offline plugin model. This ensures compatibility with multiple heterogeneous data sources, eliminates storage format inconsistencies, and simplifies the Spark data integration process. A DAG-based visual configuration reduces user complexity, while dynamic plugin relationship orchestration increases code reusability and avoids redundant code development. A seamless DFS task construction algorithm minimizes manual intervention. Using a DFS (Depth-First Search) algorithm within a DAG, SeaTunnel tasks and their dependencies are quickly and efficiently constructed, significantly reducing the complexity of the task flow. Ultimately, SeaTunnel offline task construction integrates Spark data processing across multiple stages, effectively reducing the difficulty of using Spark technology. It establishes a unified offline synchronization model, batch processing method, and efficient data development tools, forming a simple and efficient one-stop offline task construction standard solution. This reduces the technical difficulty of Spark task configuration, ensures compatibility with multiple heterogeneous data sources, eliminates data format limitations, and significantly improves the compatibility of multi-source heterogeneous offline task processing.
[0096] In one embodiment, the above-described multi-source heterogeneous offline task processing method may further include the following processing steps:
[0097] Create the metadata table and task configuration table for SeaTunnel offline tasks; the metadata table includes the data source table, dataset table, and data field table, and the task configuration table includes the project table, plugin configuration table, workflow table, task definition table, and task dependency table.
[0098] It is understood that before starting the above steps S1 to S12, this embodiment pre-sets task-related system tables, such as 3 metadata tables: data source table, dataset table, and data field table; and 5 task configuration tables: project engineering table, plugin configuration table, workflow table, task definition table, and task dependency table, etc., to facilitate quick invocation during the execution of subsequent steps.
[0099] In one embodiment, the above-described multi-source heterogeneous offline task processing method may further include the following processing steps:
[0100] Configure the SeaTunnel storage mode for offline integrated pipelines; SeaTunnel storage modes include full coverage storage, historical backup storage, time interval update storage, and / or incremental update storage.
[0101] It is understandable that after selecting the partitioning mode of the output plugin, the SeaTunnel storage mode of the offline integrated pipeline can also be configured. For example, any one of the following modes, or a combination of several, can be used: full coverage storage, historical backup storage, time interval update storage, and incremental update storage, to ensure efficient and flexible storage of pipeline output data.
[0102] Furthermore, the conversion plugin also includes a data filtering plugin, a field alias plugin, and a date conversion plugin. Optionally, in this embodiment, the conversion plugin T may also include a data filtering plugin for performing data filtering operations, a field alias plugin for performing data field alias conversion processing, and a date conversion plugin for performing data date conversion processing, thereby further expanding the types of task plugins to improve the system's task processing capabilities.
[0103] In one embodiment, the above-described multi-source heterogeneous offline task processing method may further include the following processing steps:
[0104] Bind the Task entity and its dependencies to the target user role.
[0105] Specifically, since the Task entity and task dependencies have been obtained in the above steps, and different task roles in a task have different permissions, in order to efficiently and accurately deliver the tasks in the system to the target users, the Task entity and task dependencies can be bound to the target user roles. In this way, through SeaTunnel task role permission control, permission isolation between user roles and Spark offline tasks is also achieved.
[0106] It should be understood that, although the above process Figure 1 and Figure 5 The steps in the diagram are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order in which these steps are executed; they can be performed in other orders. Furthermore, the above process... Figure 1 and Figure 5 At least some of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0107] In one embodiment, such as Figure 6As shown, a multi-source heterogeneous offline task processing system 100 is provided, including a source configuration module 01, a metadata management module 02, a job creation module 03, an integration and decomposition module 04, a pipeline creation module 05, a plug-in configuration module 06, a partition selection module 07, a plug-in orchestration module 08, a template construction module 09, an entity encapsulation module 10, a dependency transformation module 11, and a task scheduling module 12. The source configuration module 01 is used to add input and output sources in the multi-source heterogeneous data source management page, configure data source connection information, pass connectivity tests, and save the connection information to the data source table. Input sources include MySQL data sources, DB2 data sources, MongoDB data sources, and Doris data sources; output sources include Hive data sources, Phoenix data sources, and ClickHouse data sources. The metadata management module 02 is used to select data tables in the multi-source heterogeneous data dataset management page, obtain the metadata configuration information of the data tables, and save it to the dataset table and data field table. The metadata configuration information includes the data table name, data fields, field types, and primary key fields. The Job Creation Module 03 is used to create project projects and workflow canvases corresponding to multi-source heterogeneous data, save project projects to the project project table and workflow canvases to the workflow table.
[0108] The Integration Decomposition Module 04 decomposes the Spark data integration process using a unit-based plugin decomposition offline integration model, generating four types of integration plugins: execution environment plugins, input plugins, transformation plugins, and output plugins. The Pipeline Creation Module 05 creates an offline integration pipeline within the workflow canvas based on the SeaTunnel offline integration pipeline model. The pipeline modules consist of a Spark execution environment configuration module, a Spark input module, a Spark transformation module, and a Spark output module. The Plugin Configuration Module 06 dynamically selects execution environment plugins, input plugins, transformation plugins, and output plugins based on a DAG drag-and-drop operation within the SeaTunnel offline integration pipeline's module composition, configuring plugin parameters. The execution environment plugin is selected as the Spark execution environment plugin; the input plugin is selected as a MySQL data source plugin, a DB2 data source plugin, or a MongoDB data source plugin; the transformation plugin is selected as an SQL transformation plugin; and the output plugin is selected as a Hive data source plugin or a Phoenix data source plugin. Plugin parameters include execution environment parameters, input table parameters, SQL statements, and output table parameters.
[0109] The partition selection module 07 selects the data partitioning mode from the target table configuration of the output plugin using the data integration partitioning model. The plugin orchestration module 08 dynamically orchestrates the four types of integration plugin relationships in the offline integration pipeline using a DAG node connection method based on the offline integration plugin relationship orchestration model, and performs non-loop verification of plugin relationships using the DFS algorithm, generating a DAG integrated pipeline directed acyclic graph. The template construction module 09 constructs a SeaTunnel offline task configuration template based on Spark's task parameter configuration template. The entity encapsulation module 10 converts the offline integration pipeline into Task entities and records them in the task definition table using a pipeline task entity transformation model based on the DFS algorithm, based on the DAG integrated pipeline directed acyclic graph, and generates task configuration instances based on the SeaTunnel offline task configuration template. The dependency transformation module 11 converts the plugin relationships of the offline integration pipeline into task dependency relationships of Task entities and records them in the task dependency table, based on the DAG integrated pipeline directed acyclic graph and using a pipeline task dependency transformation model based on the DFS algorithm. The task scheduling module 12 is used to combine the task entity, task configuration instance and task dependency relationship according to the SeaTunnel scheduling task instance model and construct the SeaTunnel scheduling task instance to complete the SeaTunnel offline task scheduling.
[0110] The aforementioned multi-source heterogeneous offline task processing system 100 decomposes the entire Spark data integration process by designing the SeaTunnel offline plugin model. This ensures compatibility with multiple heterogeneous data sources, eliminates storage format inconsistencies, and simplifies the Spark data integration process. A DAG-based visual configuration reduces user complexity, while dynamic plugin relationship orchestration increases code reusability and avoids redundant code development. A seamless DFS task construction algorithm minimizes manual intervention. Employing a DFS-based algorithm within a DAG (Directed Acyclic Graph), it quickly and effectively completes SeaTunnel task construction and task dependency relationship establishment, significantly reducing the complexity of the task flow. Ultimately, SeaTunnel offline task construction achieves integrated Spark data processing across multiple stages, effectively reducing the difficulty of using Spark technology. It establishes a unified offline synchronization model, batch processing method, and efficient data development tools, forming a simple and efficient one-stop offline task construction standard solution. This reduces the technical difficulty of Spark task configuration, ensures compatibility with multiple heterogeneous data sources, eliminates data format limitations, and significantly improves the compatibility of multi-source heterogeneous offline task processing.
[0111] In one embodiment, the aforementioned multi-source heterogeneous offline task processing system 100 may further include: a table creation module for creating a metadata table and a task configuration table for SeaTunnel offline tasks; the metadata table includes a data source table, a dataset table, and a data field table, and the task configuration table includes a project engineering table, a plugin configuration table, a workflow table, a task definition table, and a task dependency table.
[0112] For specific limitations regarding the multi-source heterogeneous offline task processing system 100, please refer to the corresponding limitations of the multi-source heterogeneous offline task processing method above, which will not be repeated here. Each module in the multi-source heterogeneous offline task processing system 100 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of a device with data processing capabilities, or stored in software within the memory of the aforementioned device, so that the processor can call and execute the operations corresponding to each module. The aforementioned device can be, but is not limited to, various types of data computing and processing devices already existing in the art.
[0113] In one embodiment, a computer device is also provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following processing steps: adding input sources and output sources on the data source management page of multi-source heterogeneous data, configuring data source connection information, passing a connectivity test, and saving the connection information to a data source table; the input sources include MySQL data sources, DB2 data sources, MongoDB data sources, and Doris data sources, and the output sources include Hive data sources, Phoenix data sources, and ClickHouse data sources; selecting a data table on the dataset management page of multi-source heterogeneous data, obtaining the metadata configuration information of the data table, and saving it to a dataset. The data integration process involves creating a collection table and data field tables; metadata configuration information includes table name, data fields, field types, and primary key fields; creating project projects and workflow canvases corresponding to the multi-source heterogeneous data, saving the project projects to the project project table and the workflow canvas to the workflow table; decomposing the Spark data integration process using an offline integration model based on unit plugin decomposition, and generating four types of integration plugins in the data integration process; integration plugins include execution environment plugins, input plugins, transformation plugins, and output plugins; creating an offline integration pipeline in the workflow canvas based on the SeaTunnel offline integration pipeline model; the pipeline modules of the offline integration pipeline include the Spark execution environment configuration module, Spark... The SeaTunnel offline integration pipeline consists of three modules: input, transformation, and output. Within this pipeline, the execution environment plugin, input plugin, transformation plugin, and output plugin are dynamically selected and their parameters configured based on a drag-and-drop DAG operation. The execution environment plugin is selected as the Spark execution environment plugin; the input plugin is selected as a MySQL data source plugin, a DB2 data source plugin, or a MongoDB data source plugin; the transformation plugin is selected as the SQL transformation plugin; and the output plugin is selected as a Hive data source plugin or a Phoenix data source plugin. Plugin parameters include execution environment parameters, input table parameters, SQL statements, and output table parameters. Data integration partitioning is utilized. The model selects the data partitioning mode from the target table configuration of the output plugin; based on the offline integration plugin relationship orchestration model, it dynamically orchestrates the four types of integration plugin relationships in the offline integration pipeline using a DAG node connection method, and performs non-loop verification of plugin relationships using the DFS algorithm to generate a DAG integration pipeline directed acyclic graph; it constructs a SeaTunnel offline task configuration template based on Spark's task parameter configuration template; based on the DAG integration pipeline directed acyclic graph, it uses a pipeline task entity transformation model based on the DFS algorithm to convert the offline integration pipeline into Task task entities and enters them into the task definition table, and generates task configuration instances based on the SeaTunnel offline task configuration template;Based on the DAG (Directed Acyclic Graph) integrated pipeline, a pipeline task dependency transformation model based on the DFS (Depth-First Search) algorithm is used to convert the plug-in relationships of the offline integrated pipeline into task dependencies of Task entities and record them in the task dependency table. Based on the SeaTunnel scheduling task instance model, Task entities, task configuration instances, and task dependencies are combined to construct SeaTunnel scheduling task instances, thus completing the SeaTunnel offline task scheduling.
[0114] It is understood that, in addition to the memory and processor mentioned above, the computer equipment described above also includes other hardware and software components not listed in this specification. The specific components can be determined according to the model of the computer equipment in different application scenarios, and will not be listed and described in detail in this specification.
[0115] In one embodiment, when the processor executes the computer program, it can also implement the steps or sub-steps added in the various embodiments of the above-described multi-source heterogeneous offline task processing method.
[0116] In one embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When executed by a processor, the computer program performs the following processing steps: adding input and output sources in the data source management page of multi-source heterogeneous data, configuring data source connection information, passing a connectivity test, and saving the connection information to the data source table; the input sources include MySQL data sources, DB2 data sources, MongoDB data sources, and Doris data sources, and the output sources include Hive data sources, Phoenix data sources, and ClickHouse data sources; selecting a data table in the dataset management page of multi-source heterogeneous data, obtaining the metadata configuration information of the data table, and saving it to the dataset table and the dataset management page. According to the field table; metadata configuration information includes table name, data fields, field types, and primary key fields; create corresponding project projects and workflow canvases for multi-source heterogeneous data, save the project project to the project project table and the workflow canvas to the workflow table; decompose the Spark data integration process using an offline integration model based on unit plugin decomposition, and generate four types of integration plugins in the data integration process; integration plugins include execution environment plugins, input plugins, transformation plugins, and output plugins; create an offline integration pipeline in the workflow canvas based on the SeaTunnel offline integration pipeline model; the pipeline modules of the offline integration pipeline include the Spark execution environment configuration module, Spark input... The SeaTunnel offline integration pipeline consists of three modules: a Spark transformation module and a Spark output module. Within this module structure, the execution environment plugin, input plugin, transformation plugin, and output plugin are dynamically selected and their parameters configured based on a drag-and-drop DAG operation. The execution environment plugin is selected as the Spark execution environment plugin, the input plugin as a MySQL data source plugin, a DB2 data source plugin, or a MongoDB data source plugin, the transformation plugin as an SQL transformation plugin, and the output plugin as a Hive data source plugin or a Phoenix data source plugin. Plugin parameters include execution environment parameters, input table parameters, SQL statements, and output table parameters. A data integration partitioning model is utilized. Select the data partitioning mode from the target table configuration of the output plugin; based on the offline integration plugin relationship orchestration model, dynamically orchestrate the four types of integration plugin relationships in the offline integration pipeline using the DAG node connection method, and perform non-closed-loop verification of plugin relationships using the DFS algorithm to generate a DAG integration pipeline directed acyclic graph; construct the SeaTunnel offline task configuration template based on the Spark task parameter configuration template; based on the DAG integration pipeline directed acyclic graph, use the pipeline task entity transformation model based on the DFS algorithm to convert the offline integration pipeline into Task task entities and enter them into the task definition table, and generate task configuration instances based on the SeaTunnel offline task configuration template;Based on the DAG (Directed Acyclic Graph) integrated pipeline, a pipeline task dependency transformation model based on the DFS (Depth-First Search) algorithm is used to convert the plug-in relationships of the offline integrated pipeline into task dependencies of Task entities and record them in the task dependency table. Based on the SeaTunnel scheduling task instance model, Task entities, task configuration instances, and task dependencies are combined to construct SeaTunnel scheduling task instances, thus completing the SeaTunnel offline task scheduling.
[0117] In one embodiment, when the computer program is executed by the processor, it can also implement the steps or sub-steps added in the various embodiments of the multi-source heterogeneous offline task processing method described above.
[0118] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus DRAM (RDRAM), and interface DRAM (DRDRAM), etc.
[0119] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0120] The above embodiments merely illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, all of which fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for processing multi-source heterogeneous offline tasks, characterized in that, Including the following steps: In the data source management page for multi-source heterogeneous data, add input and output sources, configure data source connection information, pass connectivity tests, and save the connection information to the data source table. The input sources include MySQL data sources, DB2 data sources, MongoDB data sources, and Doris data sources, and the output sources include Hive data sources, Phoenix data sources, and ClickHouse data sources. On the dataset management page for multi-source heterogeneous data, select a data table, obtain the metadata configuration information of the data table, and save it to the dataset table and the data field table; the metadata configuration information includes the data table name, data fields, field types, and primary key fields; Create a project and workflow canvas corresponding to the multi-source heterogeneous data, save the project to the project table and the workflow canvas to the workflow table; An offline integration model based on unit plugin decomposition is used to decompose the Spark data integration process, and four types of integration plugins are generated in the data integration process. The integration plugins include execution environment plugins, input plugins, transformation plugins, and output plugins. An offline integration pipeline is created in the workflow canvas based on the SeaTunnel offline integration pipeline model; the pipeline modules of the offline integration pipeline include a Spark execution environment configuration module, a Spark input module, a Spark transformation module, and a Spark output module; In the modular composition of the SeaTunnel offline integration pipeline, the execution environment plugin, input plugin, transformation plugin, and output plugin are dynamically selected and the plugin parameter information is configured based on the DAG drag-and-drop operation. The execution environment plugin is selected as the Spark execution environment plugin, the input plugin is selected as the MySQL data source plugin, the DB2 data source plugin, or the MongoDB data source plugin, the transformation plugin is selected as the SQL transformation plugin, and the output plugin is selected as the Hive data source plugin or the Phoenix data source plugin. The plugin parameter information includes execution environment parameters, input table parameters, SQL statements, and output table parameters. The data partitioning model is used to select a data partitioning mode from the target table configuration of the output plugin; Based on the offline integration plug-in relationship orchestration model, the DAG node connection method is used to dynamically orchestrate the four types of integration plug-in relationships in the offline integration pipeline, and the DFS algorithm is used to verify the non-closed loop of the plug-in relationship to generate a DAG integration pipeline directed acyclic graph. Build a SeaTunnel offline task configuration template based on Spark's task parameter configuration template; Based on the DAG integrated pipeline directed acyclic graph, the offline integrated pipeline is converted into a Task entity using a pipeline task entity transformation model based on the DFS algorithm and entered into the task definition table. Additionally, a task configuration instance is generated based on the SeaTunnel offline task configuration template. Based on the DAG integrated pipeline directed acyclic graph, the plug-in relationship of the offline integrated pipeline is converted into the task dependency relationship of the Task entity using the pipeline task dependency transformation model based on the DFS algorithm and recorded in the task dependency table. Based on the SeaTunnel scheduling task instance model, the Task entity, the Task configuration instance, and the Task dependency relationship are combined to construct a SeaTunnel scheduling task instance, thereby completing the SeaTunnel offline task scheduling.
2. The multi-source heterogeneous offline task processing method according to claim 1, characterized in that, It also includes the following steps: Create a metadata table and a task configuration table for SeaTunnel offline tasks; the metadata table includes a data source table, a dataset table, and a data field table, and the task configuration table includes a project table, a plugin configuration table, a workflow table, a task definition table, and a task dependency table.
3. The multi-source heterogeneous offline task processing method according to claim 1 or 2, characterized in that, The output plugin's data partitioning modes include full mode (FT) or partition mode (PT).
4. The multi-source heterogeneous offline task processing method according to claim 3, characterized in that, It also includes the following steps: Configure the SeaTunnel storage mode of the offline integrated pipeline; the SeaTunnel storage mode includes full coverage storage, historical backup storage, time interval update storage, or incremental update storage.
5. The multi-source heterogeneous offline task processing method according to claim 1, characterized in that, The conversion plugin also includes a data filtering plugin, a field aliasing plugin, and a date conversion plugin.
6. The multi-source heterogeneous offline task processing method according to claim 1, characterized in that, Also includes: Bind the Task entity and the Task dependency to the target user role.
7. A multi-source heterogeneous offline task processing system, characterized in that, include: The source configuration module is used to add input and output sources in the data source management page for multi-source heterogeneous data, configure data source connection information, pass connectivity tests, and save the connection information to the data source table. The input sources include MySQL data sources, DB2 data sources, MongoDB data sources, and Doris data sources, and the output sources include Hive data sources, Phoenix data sources, and ClickHouse data sources. The metadata management module is used to select a data table in the dataset management page of multi-source heterogeneous data, obtain the metadata configuration information of the data table, and save it to the dataset table and the data field table; the metadata configuration information includes the data table name, data fields, field types, and primary key fields; The job creation module is used to create project projects and workflow canvases corresponding to the multi-source heterogeneous data, save the project projects to the project project table, and save the workflow canvases to the workflow table; The integration decomposition module is used to decompose the Spark data integration process using an offline integration model based on unit plugin decomposition, and to generate four types of integration plugins in the data integration process; the integration plugins include execution environment plugins, input plugins, transformation plugins and output plugins. The pipeline creation module is used to create an offline integration pipeline based on the SeaTunnel offline integration pipeline model in the workflow canvas; the pipeline module of the offline integration pipeline includes a Spark execution environment configuration module, a Spark input module, a Spark transformation module, and a Spark output module. The plugin configuration module is used to dynamically select execution environment plugins, input plugins, transformation plugins, and output plugins based on DAG drag-and-drop operations and configure plugin parameter information in the module composition of the SeaTunnel offline integration pipeline. The execution environment plugin is selected as the Spark execution environment plugin, the input plugin is selected as the MySQL data source plugin, the DB2 data source plugin, or the MongoDB data source plugin, the transformation plugin is selected as the SQL transformation plugin, and the output plugin is selected as the Hive data source plugin or the Phoenix data source plugin. The plugin parameter information includes execution environment parameters, input table parameters, SQL statements, and output table parameters. The partition selection module is used to select a data partitioning mode from the target table configuration of the output plugin using the data integration partitioning model; The plug-in orchestration module is used to dynamically orchestrate the four types of integrated plug-in relationships in the offline integrated pipeline according to the offline integrated plug-in relationship orchestration model, using the DAG node connection method, and to perform non-closed-loop verification of plug-in relationships through the DFS algorithm, generating a DAG integrated pipeline directed acyclic graph. The template building module is used to build SeaTunnel offline task configuration templates based on Spark task parameter configuration templates. The entity encapsulation module is used to convert the offline integrated pipeline into a Task entity and enter it into the task definition table based on the DAG integrated pipeline directed acyclic graph and pipeline task entity conversion model based on DFS algorithm, and to generate a task configuration instance based on the SeaTunnel offline task configuration template. The dependency conversion module is used to convert the plug-in relationship of the offline integrated pipeline into the task dependency relationship of the Task entity based on the DAG integrated pipeline directed acyclic graph and the pipeline task dependency conversion model based on the DFS algorithm, and record it into the task dependency table. The task scheduling module is used to combine the Task entity, the task configuration instance, and the task dependency relationship according to the SeaTunnel scheduling task instance model to construct a SeaTunnel scheduling task instance and complete the SeaTunnel offline task scheduling.
8. The multi-source heterogeneous offline task processing system according to claim 7, characterized in that, Also includes: The table creation module is used to create metadata tables and task configuration tables for SeaTunnel offline tasks. The metadata tables include data source tables, dataset tables, and data field tables, while the task configuration tables include project tables, plugin configuration tables, workflow tables, task definition tables, and task dependency tables.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the multi-source heterogeneous offline task processing method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-source heterogeneous offline task processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Spark-based data processing method and system
CN107463595A
Intelligent data conversion method based on Spark
CN113641739A