A Spark-based multi-component data synchronization and processing method, system, and tool
The Spark platform realizes multi-component data synchronization and processing, solves the shortcomings of multi-component data synchronization and processing in the existing technology, supports cross-component data transmission, multi-table association and visual processing, and improves data processing efficiency.
Patent Information
- Application Number
- CN202211497067.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-11-25
AI Technical Summary
The prior art cannot effectively realize the synchronization and processing of multi-component data, especially in terms of multi-table association, cross-component synchronization and visual processing when data model updates.
Spark is used as the data management platform, and component information is configured through the data source management interface and connectivity verification is performed. Table nodes are generated and synchronization parameters are configured using drag and drop. Combined with the security authentication module and parameter processing module, Spark tasks are generated and data synchronization and function processing are performed in DolphinScheduler or Yarn.
It supports data transmission and function processing between multiple components, simplifies the complexity of task configuration, provides a visual interface to view task status, realizes data transmission and multi-table association across clusters, and improves data processing efficiency.
Smart Images

Figure CN115809299B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of big data processing, and particularly relates to a multi-component data synchronization and processing method, system and tool based on Spark. Background Art
[0002] At present, more and more enterprises begin to pay attention to the value brought by data. Traditional Oracle and MySQL databases can no longer meet the storage and computing requirements of a large amount of data. Due to the unique storage characteristics of big data technologies such as HDFS, Hive, HBase, and Kudu, they have become common solutions. In the process of only data storage and management, enterprises actually have the following requirements: (1) When the data model is updated, it is necessary to process the data or perform multi-table association; (2) It is necessary to complete data synchronization across components; (3) When migrating clusters, it is necessary to complete cross-cluster data migration.
[0003] Currently, the mainstream components for realizing big data dump are DataX and Flume. DataX itself is an offline data synchronization framework, which is built with a Framework+plugin architecture. The reading and writing of data sources are abstracted into Reader / Writer plugins and incorporated into the entire synchronization framework. Data synchronization is completed by writing json information of Reader and Writer. Flume is a highly available, highly reliable, distributed system for massive log collection, aggregation and transmission provided by Cloudera. Flume supports customizing various data senders in the log system to collect data, and the synchronization configuration is completed by writing properties files of Source and Sink.
[0004] Among the above two big data dump components, DataX supports a wide variety of components, but it is necessary to complete the writing of complex json information of Reader and Writer. The range of components supported by Flume is relatively small. If it is necessary to complete the data synchronization of new components, corresponding code must be written according to the Flume framework. In addition, neither DataX nor Flume supports the cascading or linkage relationship between multiple tasks, nor does it support function processing and multi-table association of data, nor does it have a visual interface to view the running status and running results of tasks. Therefore, the prior art has not yet proposed an effective technical solution for synchronizing and processing data between multiple components to meet the multi-table association, cross-component synchronization and visual processing of data. Summary of the Invention
[0005] In order to solve the problem that the existing solution is inconvenient in the process of multi-component data synchronization and processing, which reduces the data dump and processing efficiency of users; the present invention provides a multi-component data synchronization and processing method, system and tool based on Spark.
[0006] The present invention is implemented by the following technical solutions:
[0007] A multi-component data synchronization and processing method based on Spark, which is used to synchronize data between different components, execute various function processing tasks, and visualize the task configuration process. The multi-component data synchronization and processing method includes the following steps:
[0008] Step 1: Use Spark as the data management platform. In the data source management interface, complete the configuration of the data source information of the components and pass the connectivity verification.
[0009] Step 2: Select a table under a certain data source in the task configuration management interface, and generate a corresponding table node on the canvas in a drag-and-drop manner; generate another table node in the same way, and complete the connection between the two table nodes.
[0010] Step 3: The user double-clicks on the two table nodes to complete the synchronization parameter configuration, and the parameter generation module will generate the corresponding parameter information.
[0011] Step 4: After the user clicks to submit, the security authentication module performs security authentication on the submitted configuration information.
[0012] Step 5: After receiving the parameter information, the parameter processing module initializes SparkConf and Configuration according to the parameter information.
[0013] Step 6: After completing the initialization of SparkConf and Configuration, submit the large data packet and the large data task parameters to DolphinScheduler or Yarn to generate the corresponding Spark task. Among them, after the tasks submitted to DolphinScheduler and Yarn obtain resources, they will complete the corresponding data synchronization.
[0014] Step 7: Spark recursively parses the task parameter information, then reads the source table data, and writes it into the target table.
[0015] As a further improvement of the present invention, in Step 6, when there are multiple tasks that need to be submitted to DolphinScheduler or Yarn for configuration, select the method of submitting simultaneously or submitting sequentially according to the preset priority order for parameter configuration.
[0016] As a further improvement of the present invention, in the data source information configuration of Step 1, the selectable data sources include: HBase, Hive, Kafka, Kudu, Oracle, MySQL.
[0017] As a further improvement of the present invention, during the parameter configuration of the data table in step three, the parameter configuration items when HBase is used as the source table include: columns, column families, filtering methods, start Rowkey, end Rowkey, and values included in the Rowkey.
[0018] The parameter configuration items when Oracle or MySQL is used as the source table include: data filtering conditions.
[0019] The parameter configuration items when Hive or Kudu is used as the source table include: whether it is a partitioned table, partition fields, date partition format, synchronization method, cycle type, number of cycles, and time period.
[0020] The parameter configuration items when HBase, Hive, Kudu, Oracle, or MySQL is used as the target table include: field mapping, saving method, partition value format, and partition fields.
[0021] The parameter configuration items when Kafka is used as the target table include: broker address, field selection, key value selection, and key value concatenation symbol.
[0022] The configuration information of the connection points for association includes: join method, join fields, field mapping relationship, and function selection. Among them, the functions supported in the connection point configuration information include: fill, subString, replace, cast, and addColumn.
[0023] As a further improvement of the present invention, in step four, if security authentication is enabled, the keytab and Principal information required by Kerberos are used to authenticate through Kerberos to ensure the authenticity and security of the identities of both communication parties.
[0024] As a further improvement of the present invention, during the cross-component synchronization of data in step six, the format of the task parameters passed to the big data package is:
[0025] {"spark_param":{}, "left_child":{}, "right_child":{}, "sink":{}};
[0026] Among them, "spark_param" represents the system parameter item in Spark; "left_child" and "right_child" represent the source table nodes. If "right_child" is empty, it represents the data synchronization between single tables. If both "left_child" and "right_child" are not empty, it represents that after the two tables are associated, the data is written into the "sink" target table. Among them, "left_child" and "right_child" adopt a nested structure to meet the requirements of multi-table association.
[0027] As a further improvement of the present invention, in step seven, Spark reads the data of each table to generate a DataFrame, and based on the DataFrame, function processing and join operations are completed.
[0028] The present invention also includes a multi-component data synchronization and processing system based on Spark. This system adopts the multi-component data synchronization and processing method based on Spark as described above, performs data synchronization between components such as HBase, Hive, Kafka, Kudu, Oracle, and MySQL, executes various function processing capabilities, and realizes task visualization configuration. This type of multi-component data synchronization and processing system includes: a data source management module, a parameter generation module, a parameter processing module, a big data task generation module, and a security authentication module.
[0029] The data source management module is used to configure the data source connection information of HBase, Hive, Kafka, Kudu, Oracle, and MySQL; the data source connection information includes the connection url, username, password, parameters, etc., and supports manual addition of configuration information.
[0030] The parameter generation module is used to automatically generate the corresponding task json information according to the task configuration parameters filled in by the user on the web interface to reduce the configuration complexity.
[0031] The parameter processing module is used to parse the task json information passed by the parameter generation module into parameter information recognizable by the big data generation module;
[0032] The big data task generation module is used to initialize sparkConf and Configuration according to the parameter information parsed by the parameter processing module, and then submit the parameters and the written Spark big data program package to DolphinScheduler or Yarn, and generate corresponding calculation tasks; the Spark big data program parses the parameter information passed by the backend and completes the data synchronization and function processing tasks between the corresponding components.
[0033] The security authentication module is used to perform security authentication login on the configuration information manually added by the user in the data source management module through Kerberos.
[0034] As a further improvement of the present invention, in the data source management module, when configuring multiple tasks, the user configures the execution order between multiple tasks according to the requirements of the actual scenario to achieve simultaneous execution or sequential execution.
[0035] The present invention also includes a multi-component data synchronization and processing tool based on Spark, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the multi-component data synchronization and processing method based on Spark as described above are implemented, realizing data synchronization, function processing, and visualization operations between multiple components.
[0036] The technical solution provided by the present invention has the following beneficial effects:
[0037] The solution provided by the present invention uses Spark as a data management tool to realize cross-component data and function processing tasks. This solution supports components very comprehensively, including HBase, Hive, Kafka, Kudu, Oracle, MySQL; and it not only supports data transmission within the same component, but also supports data transmission between different components, overcoming the limitations of data storage formats between components.
[0038] The data management function of the solution of the present invention is rich, supporting data transfer across clusters. The user only needs to configure the data sources to be synchronized in the data source management module to complete data transfer between multiple clusters. It supports function processing and multi-table association. The supported functions include fill, subString, filter, and writing to the target table after multi-table association. It supports task scheduling. The user can configure the execution order between multiple tasks according to the actual scenario, supporting simultaneous and sequential execution. It supports visual task management. The user can view the execution status, running logs, and data synchronization count of tasks on the web interface. It supports visual interface configuration, allowing the user to complete task configuration in a drag-and-drop form on the web interface, greatly simplifying the complexity of task configuration. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of the steps of a multi-component data synchronization and processing method based on Spark provided in Embodiment 1 of the present invention.
[0040] Figure 2 It is a visual interface for generating different table nodes in Spark in Embodiment 1.
[0041] Figure 3It is a visual interface for the data transfer logic between the data source table node and the target table node in Embodiment 1.
[0042] Figure 4 It is a step flowchart of a multi-component data synchronization and processing method based on Spark provided in Embodiment 1 of the present invention. Detailed implementation manners
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.
[0045] Embodiment 1
[0046] This embodiment provides a multi-component data synchronization and processing method based on Spark, which is used for data synchronization between different components, performing various function processing tasks, and visualizing the task configuration process. As Figure 1 shown, the multi-component data synchronization and processing method includes the following steps:
[0047] Step 1: Use Spark as the data management platform. In the data source management interface, complete the configuration of the data source information of the component and pass the connectivity verification.
[0048] In the data source information configuration, the selectable data sources include: HBase, Hive, Kafka, Kudu, Oracle, MySQL.
[0049] Step 2: Select a table under a certain data source in the task configuration management interface, and generate a corresponding table node on the canvas by dragging and dropping; generate another table node in the same way, and complete the connection between the two table nodes. The visualization state of this step is as Figure 2 shown.
[0050] Step 3: The user double-clicks on the two table nodes to complete the synchronization parameter configuration, and the parameter generation module will generate the corresponding parameter information.
[0051] Among them, the parameter configuration items when HBase is used as the source table include: columns, column families, filtering methods, start Rowkey, end Rowkey, and values included in the Rowkey.
[0052] The parameter configuration items when Oracle and MySQL are used as the source table include: data filtering conditions.
[0053] The parameter configuration items when Hive and Kudu are used as the source table include: whether it is a partitioned table, partition fields, date partition format, synchronization method, cycle type, number of cycles, and time period.
[0054] The parameter configuration items when HBase, Hive, Kudu, Oracle, and MySQL are used as the target table include: field mapping, saving method, partition value format, and partition fields.
[0055] The parameter configuration items when Kafka is used as the target table include: broker address, field selection, key value selection, and key value concatenation symbol.
[0056] The configuration information for the connection points used for association includes: join method, join fields, field mapping relationship, and function selection. Among them, the functions supported in the connection point configuration information include: fill, subString, replace, cast, and addColumn.
[0057] Step 4: After the user clicks to submit, the security authentication module performs security authentication on the submitted configuration information. If security authentication is enabled, Kerberos authentication is performed using the keytab and Principal information required by Kerberos to ensure the authenticity and security of the identities of both communication parties.
[0058] Step 5: After the parameter processing module receives the parameter information, it initializes SparkConf and Configuration according to the parameter information.
[0059] Step 6: After completing the initialization of SparkConf and Configuration, the large data package and large data task parameters are submitted to DolphinScheduler or Yarn to generate corresponding Spark tasks. Among them, after the tasks submitted to DolphinScheduler and Yarn obtain resources, they will complete the corresponding data synchronization. In the specific processing process, the visualization result of the data transfer logic is as Figure 3 shown.
[0060] When there are multiple tasks that need to be submitted to DolphinScheduler or Yarn for configuration, select the method of submitting them simultaneously or in sequence according to the preset priority order for parameter configuration.
[0061] Specifically, during the cross-component synchronization process of data, the format of the task parameters passed to the large data packet is as follows:
[0062] {"spark_param":{}, "left_child":{}, "right_child":{}, "sink":{}};
[0063] Among them, "spark_param" represents the system parameter item in Spark; "left_child" and "right_child" represent the source table nodes. If "right_child" is empty, it represents the data synchronization between single tables. If both "left_child" and "right_child" are not empty, it represents the data is written into the "sink" target table after the association of two tables. Among them, "left_child" and "right_child" adopt a nested structure to meet the requirements of multi-table association.
[0064] Step 7: Spark recursively parses the task parameter information, then reads the source table data and writes it into the target table. Spark reads the data of each table to generate a DataFrame, and completes function processing and join operations based on the DataFrame.
[0065] Embodiment 2
[0066] Based on Embodiment 1, this embodiment further provides a multi-component data synchronization and processing system based on Spark, which is the simplest system environment capable of implementing the required data synchronization and processing tasks. In the actual application process, the required multi-component data synchronization and processing system can be rebuilt, or the existing large-scale complex system can be simplified to support the implementation of the corresponding task processing functions.
[0067] The multi-component data synchronization and processing system based on Spark provided in this embodiment adopts the multi-component data synchronization and processing method based on Spark in Embodiment 1 to perform data synchronization among HBase, Hive, Kafka, Kudu, Oracle, and MySQL components, execute various function processing capabilities, and achieve task visualization configuration.
[0068] As Figure 4 shown, this type of multi-component data synchronization and processing system includes: a data source management module, a parameter generation module, a parameter processing module, a big data task generation module, and a security authentication module.
[0069] Among them, the data source management module is used to configure the data source connection information of HBase, Hive, Kafka, Kudu, Oracle, and MySQL; the data source connection information includes the connection url, username, password, parameters, etc., and supports manual addition of configuration information. In the data source management module, when configuring multiple tasks, the user configures the execution order between multiple tasks according to the requirements of the actual scenario to achieve simultaneous execution or sequential execution.
[0070] The parameter generation module is used to automatically generate the corresponding task json information according to the task configuration parameters filled in by the user on the web interface to reduce the configuration complexity.
[0071] The parameter processing module is used to parse the task json information passed by the parameter generation module into parameter information recognizable by the big data generation module;
[0072] The big data task generation module is used to initialize sparkConf and Configuration according to the parameter information parsed by the parameter processing module, and then submit the parameters and the written Spark big data program package to DolphinScheduler or Yarn, and generate the corresponding calculation tasks; the Spark big data program parses the parameter information passed by the backend to complete the data synchronization and function processing tasks between the corresponding components.
[0073] The security authentication module is used to perform security authentication login on the configuration information manually added by the user in the data source management module through Kerberos.
[0074] Embodiment 3
[0075] This embodiment provides a multi-component data synchronization and processing tool based on Spark. The data synchronization and processing tool is a data processing device running a specific microcomputer system; it includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the multi-component data synchronization and processing method based on Spark in Embodiment 1 to achieve data synchronization, function processing, and visualization operations between multiple components.
[0076] The computer device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers) that can execute programs. The computer device of this embodiment includes at least, but is not limited to, a memory and a processor that can communicate with each other through a system bus.
[0077] In this embodiment, the memory (i.e., the readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various data that have been output or will be output.
[0078] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device.
[0079] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the various technical features in the above-described embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0080] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A multi-component data synchronization and processing method based on Spark, which is used for data synchronization between different components, performing various function processing tasks, and visualizing the task configuration process; the multi-component data synchronization and processing method includes the following steps: Step 1: Use Spark as the data management platform. In the data source management interface, complete the configuration of the data source information of the components and pass the connectivity verification; Step 2: Select a table under a certain data source in the task configuration management interface, and generate a corresponding table node on the canvas in a drag-and-drop manner; generate another table node in the same way, and complete the connection between the two table nodes; Step 3: The user double-clicks on the two table nodes to complete the synchronization parameter configuration, and the parameter generation module will generate the corresponding parameter information; Step 4: After the user clicks to submit, the security authentication module performs security authentication on the submitted configuration information; Step 5: After receiving the parameter information, the parameter processing module initializes SparkConf and Configuration according to the parameter information; Step 6: After initialization, submit the large data package and large data task parameters to DolphinScheduler or Yarn to generate corresponding Spark tasks; and complete the corresponding data synchronization; during the cross-component data synchronization process, the format of the task parameters passed to the large data package is: {"spark_param":{},"left_child":{},"right_child":{},"sink":{}}; "spark_param" represents the system parameter items in Spark; "left_child" and "right_child" represent the source table nodes. If "right_child" is empty, it represents the data synchronization between single tables. If both "left_child" and "right_child" are not empty, it represents that the data is written into the "sink" target table after the association of two tables; among them, "left_child” and "right_child” adopt a nested structure to meet the requirements of multi-table association; Step 7: Spark recursively parses the task parameter information, then reads the source table data and writes it into the target table.
2. The method for multi-component data synchronization and processing based on Spark according to claim 1, wherein: In Step 6, when there are multiple tasks that need to be submitted to DolphinScheduler or Yarn for configuration, select the method of submitting simultaneously or submitting sequentially according to the preset priority order for parameter configuration.
3. The multi-component data synchronization and processing method based on Spark according to claim 1, characterized in that: During the data source information configuration process in Step 1, the selectable data sources include: HBase, Hive, Kafka, Kudu, Oracle, MySQL.
4. The method for multi-component data synchronization and processing based on Spark according to claim 1, wherein: During the data table parameter configuration process in Step 3, the parameter configuration when HBase is the source table includes: columns, column families, filtering methods, start Rowkey, end Rowkey, Rowkey contains value; The parameter configuration items when Oracle and MySQL are the source tables include: data filtering conditions; The parameter configuration items when Hive and Kudu are the source tables include: whether it is a partitioned table, partition fields, date partition format, synchronization method, cycle type, number of cycles, time period; The parameter configuration items when HBase, Hive, Kudu, Oracle, and MySQL are the target tables include: field mapping, saving method, partition value format, partition fields; The parameter configuration items when Kafka is the target table include: broker address, field selection, key value selection, key value concatenation symbol; The configuration information of the connection points for association includes: join method, join fields, field mapping relationship, function selection; among them, the functions supported in the connection point configuration information include: fill, subString, replace, cast, addColumn.
5. The multi-component data synchronization and processing method based on Spark according to claim 1, characterized in that: In step 4, if the security authentication is enabled, the keytab and Principal information required by Kerberos are used to perform Kerberos authentication to ensure the authenticity and security of the identities of both communication parties.
6. The method for multi-component data synchronization and processing based on Spark according to claim 1, characterized in that: In step 7, Spark reads the data of each table to generate a DataFrame, and based on the DataFrame, function processing and join operations are completed.
7. A Spark-based multi-component data synchronization and processing system, characterized in that: It adopts the Spark-based multi-component data synchronization and processing method described in any one of claims 1-6, performs data synchronization among components such as HBase, Hive, Kafka, Kudu, Oracle, and MySQL, executes various function processing capabilities, and realizes task visualization configuration; the multi-component data synchronization and processing system includes: A data source management module, which is used to configure the data source connection information of HBase, Hive, Kafka, Kudu, Oracle, and MySQL; the data source connection information includes the connection url, username, password, parameters, etc., and supports manual addition of configuration information; A parameter generation module, which is used to automatically generate corresponding task json information according to the task configuration parameters filled in by the user on the web interface to reduce the configuration complexity; A parameter processing module, which is used to parse the task json information passed by the parameter generation module into parameter information recognizable by the big data generation module; A big data task generation module, which is used to initialize sparkConf and Configuration according to the parameter information parsed by the parameter processing module, and then submit the parameters and the written Spark big data program package to DolphinScheduler or Yarn, and generate corresponding computing tasks; the Spark big data program parses the parameter information passed by the backend and completes the data synchronization and function processing tasks among the corresponding components; A security authentication module, which is used to perform security authentication login on the configuration information manually added by the user in the data source management module through Kerberos.
8. The Spark-based multi-component data synchronization and processing system according to claim 7, wherein: In the data source management module, when configuring multiple tasks, the user configures the execution order among various tasks according to the actual scenario requirements to achieve simultaneous execution or sequential execution.
9. A multi-component data synchronization and processing tool based on Spark, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, it realizes the steps of the Spark-based multi-component data synchronization and processing method described in any one of claims 1-6, and realizes data synchronization, function processing, and visualization operations among multiple components.
Citation Information
Patent Citations
Multi-source heterogeneous data governance method and device
CN112579625A
Data migration method and system, storage medium and electronic equipment
CN113672591A