Method and system for guaranteeing consistency of task double-cluster data

By using the data synchronization engine in a HADOOP dual-cluster environment, the task execution results are synchronized from the primary cluster to the standby cluster, solving the problem of data inconsistency after the task is completed, reducing manual verification and processing time, and improving data consistency.

CN120067119APending Publication Date: 2025-05-30SI-TECH INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111209.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the HADOOP dual cluster environment, after the task is completed, inconsistencies are found through post-audit audit, which requires manual verification and processing, resulting in a large time consumption and a great impact on the processing of downstream tasks.

Method used

Through data synchronization, two clusters of the same type are obtained, the data synchronization engine is determined, the main cluster and the backup cluster are specified, and the data synchronization engine is used to synchronize the task execution results from the main cluster to the backup cluster to achieve data consistency.

Benefits of technology

It reduces the time for manual verification and processing, lowers the maintenance threshold, improves the consistency of data in dual clusters, and reduces the impact on downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067119A_ABST
    Figure CN120067119A_ABST
Patent Text Reader

Abstract

The invention discloses a method for guaranteeing task double-cluster data consistency, and belongs to the technical field of data processing. The method comprises the following steps: acquiring two clusters of the same type; determining a data synchronization engine according to the type of the cluster; determining a main cluster and a standby cluster from the two clusters; and performing data synchronization according to the data synchronization engine, the main cluster and the standby cluster. According to the method, a flexible data synchronization engine is adopted, the tasks only need to be configured at a time, then data synchronization is automatically carried out subsequently, the integrity, accuracy and consistency of the dual-cluster data of the tasks can be guaranteed, and the time for manual checking and problem positioning is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for ensuring data consistency in a dual-cluster of tasks. Background Art

[0002] Ensuring data consistency in the ETL task dual-cluster: After the task is completed in the HADOOP dual-cluster, the consistency auditing platform conducts post-event auditing through data extraction. In case of inconsistency, manual verification and processing are carried out, and downstream tasks are also processed.

[0003] Although the post-event auditing method can detect inconsistency problems, in an actual production system, downstream tasks of the task have already run, and the processing process requires processing the downstream tasks as well, which consumes a large amount of time and is time-consuming and laborious. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for ensuring data consistency in a dual-cluster of tasks, which solves the problem of data consistency in the dual-cluster of tasks through data synchronization, and avoids the timeliness of post-event auditing and the secondary processing of downstream tasks.

[0005] To solve the above technical problems, the present invention provides a method for ensuring data consistency in a dual-cluster of tasks, including the following steps:

[0006] Obtain two clusters of the same type;

[0007] Determine a data synchronization engine according to the type of the cluster;

[0008] Determine a primary cluster and a standby cluster from the two clusters;

[0009] According to the data synchronization engine, the primary cluster and the standby cluster perform data synchronization.

[0010] Preferably, according to the data synchronization engine, the primary cluster and the standby cluster perform data synchronization, specifically including the following steps:

[0011] Create ETL tasks in the primary cluster and the standby cluster respectively;

[0012] Execute the ETL task in the primary cluster to obtain the task execution result;

[0013] Based on the data synchronization engine, synchronize the task execution result to the standby cluster.

[0014] Preferably, the types of the clusters include HADOOP clusters and GBASE clusters.

[0015] Preferably, the HADOOP cluster adopts the Distcp data synchronization engine;

[0016] The GBASE cluster adopts the Gload data synchronization engine.

[0017] Preferably, the two clusters of the HADOOP cluster type are the HA cluster and the HB cluster respectively;

[0018] The HA cluster is the main cluster, and the HB cluster is the standby cluster.

[0019] Preferably, the two clusters of the GBASE cluster type are the GA cluster and the GB cluster respectively;

[0020] The GA cluster is the main cluster, and the GB cluster is the standby cluster.

[0021] Preferably, the task execution result includes the table name, cycle, and storage method.

[0022] The present invention also provides a system for ensuring data consistency in a dual-cluster of tasks, including:

[0023] An acquisition module, configured to acquire two clusters of the same type;

[0024] A data synchronization engine determination module, configured to determine the data synchronization engine according to the type of the cluster;

[0025] A main cluster and standby cluster setting module, configured to determine the main cluster and the standby cluster from the two clusters;

[0026] A synchronization module, configured to perform data synchronization on the main cluster and the standby cluster according to the data synchronization engine.

[0027] Compared with the prior art, the beneficial effects of the present invention are:

[0028] The present invention realizes data consistency of tasks in a dual-cluster by adopting different data synchronization engines for different clusters, reduces the time for manual verification and problem handling, lowers the maintenance threshold, and improves consistency.

[0029] The present invention adopts a flexible data synchronization engine. After the task is configured only once, subsequent data synchronization is automatically performed, which can ensure data integrity, accuracy, and consistency of the task in the dual-cluster, and reduce the time for manual verification and problem location. Description of the Drawings

[0030] The following further describes in detail the specific embodiments of the present invention with reference to the drawings.

[0031] Figure 1 It is a flowchart of a method for ensuring data consistency in a dual-cluster of tasks;

[0032] Figure 2 It is a flowchart of data synchronization between the main cluster and the standby cluster. Detailed Implementation Modes

[0033] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.

[0034] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0035] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0036] The present invention will be further described in detail below with reference to the accompanying drawings:

[0037] The present invention provides a method for ensuring data consistency in a dual-cluster for safeguard tasks, including the following steps:

[0038] Obtain two clusters of the same type;

[0039] Determine a data synchronization engine according to the type of the cluster;

[0040] Determine a primary cluster and a standby cluster from the two clusters;

[0041] Synchronize data between the primary cluster and the standby cluster according to the data synchronization engine.

[0042] Preferably, synchronizing data between the primary cluster and the standby cluster according to the data synchronization engine specifically includes the following steps:

[0043] Create ETL tasks in the primary cluster and the standby cluster respectively;

[0044] Execute the ETL task in the primary cluster to obtain the task execution result;

[0045] Based on the data synchronization engine, synchronize the task execution results to the standby cluster.

[0046] Preferably, the types of the clusters include HADOOP clusters and GBASE clusters.

[0047] Preferably, the HADOOP cluster adopts the Distcp data synchronization engine;

[0048] The GBASE cluster adopts the Gload data synchronization engine.

[0049] Preferably, the two clusters of the HADOOP cluster type are respectively the HA cluster and the HB cluster;

[0050] The HA cluster is the primary cluster, and the HB cluster is the standby cluster.

[0051] Preferably, the two clusters of the GBASE cluster type are respectively the GA cluster and the GB cluster;

[0052] The GA cluster is the primary cluster, and the GB cluster is the standby cluster.

[0053] Preferably, the task execution results include table name, period, and storage method.

[0054] The present invention also provides a system for ensuring data consistency in a dual-cluster of tasks, including:

[0055] An acquisition module, configured to acquire two clusters of the same type;

[0056] A data synchronization engine determination module, configured to determine the data synchronization engine according to the type of the cluster;

[0057] A primary cluster and standby cluster setting module, configured to determine the primary cluster and the standby cluster from the two clusters;

[0058] A synchronization module, configured to perform data synchronization on the primary cluster and the standby cluster according to the data synchronization engine.

[0059] To better illustrate the technical effects of the present invention, the present invention provides the following specific embodiments to illustrate the above technical process:

[0060] Embodiment 1, A method for ensuring data consistency in a dual-cluster of tasks, adding a data synchronization engine device to the existing system; after the task execution in the primary cluster is completed, the downstream of the task in the primary cluster starts to execute, and at the same time, the same task in the standby cluster pulls data in a data synchronization manner. After the data pulling is completed, the downstream task starts to execute and runs iteratively.

[0061] A method for ensuring data consistency in a dual-cluster of tasks, as Figure 1 shown, specifically includes the following steps:

[0062] 1. Data Synchronization Engine

[0063] Classified by cluster, different clusters adopt different data synchronization engines

[0064] Cluster Classification Data Synchronization Engine HADOOP Distcp GBASE Gload

[0065] The HADOOP cluster adopts the Distcp data synchronization engine, and the GBASE cluster adopts the Gload data synchronization engine;

[0066] 2. Cluster Configuration

[0067] HADOOP Dual Cluster: Two clusters, HA and HB, with HA designated as the primary cluster and HB as the standby cluster;

[0068] GBASE Dual Cluster: Two clusters, GA and GB, with GA designated as the primary cluster and GB as the standby cluster;

[0069] 3. Task Configuration

[0070] 1) The ETL task is created in the GA cluster to automatically create tasks in the GB cluster as well;

[0071] 2) Data Synchronization Method: Synchronize from the primary cluster to the standby cluster, including table name, cycle, and storage method;

[0072] 4. Processing Solutions, as Figure 2 shown:

[0073] 1) The GA cluster is the primary cluster, and the GB cluster is the standby cluster.

[0074] 2) Task AAA exists as a dual-cluster task in both the GA cluster and the GB cluster.

[0075] 3) After the primary cluster task AAA is completed, the downstream of the primary cluster task AAA starts to execute.

[0076] 4) The standby cluster starts the data synchronization engine and synchronizes data from the primary to the standby according to the synchronization table, cycle, and storage method configured for task AAA.

[0077] 5) After the data synchronization is completed, the standby cluster task runs to completion, and the downstream of the standby cluster task AAA starts to run.

[0078] This embodiment adopts a flexible data synchronization engine. After the task is configured only once, subsequent data synchronization is automatically performed, which can ensure the integrity, accuracy, and consistency of data in the dual cluster, and reduce the time for manual verification and problem location.

[0079] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules, units or components is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units, modules or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0080] The unit may or may not be physically separated. The components shown as units can be one physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0081] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0082] In particular, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present invention are executed. It should be noted that the above computer-readable medium of the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above.

[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0084] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for ensuring data consistency of dual clusters of tasks, characterized in that: The following steps are involved: Get two clusters of the same type; Determine the data synchronization engine based on the cluster type; Determine the primary cluster and the secondary cluster from the two clusters; Based on the data synchronization engine, the primary cluster and the standby cluster synchronize data.

2. The method for ensuring data consistency of dual clusters of tasks according to claim 1 is characterized in that: According to the data synchronization engine, the active cluster and the standby cluster perform data synchronization, which specifically includes the following steps: Create ETL tasks in the primary cluster and the standby cluster respectively; Execute the ETL task in the main cluster and obtain the task execution results; Based on the data synchronization engine, the task execution results are synchronized to the standby cluster.

3. The method for ensuring data consistency of dual clusters of tasks according to claim 2 is characterized by: The types of clusters include HADOOP clusters and GBASE clusters.

4. The method for ensuring data consistency of dual clusters of tasks according to claim 3 is characterized by: The HADOOP cluster uses the Distcp data synchronization engine; The GBASE cluster adopts the Gload data synchronization engine.

5. The method for ensuring data consistency of dual clusters of tasks according to claim 4 is characterized by: The two clusters of the HADOOP cluster type are respectively an HA cluster and an HB cluster; The HA cluster is a primary cluster, and the HB cluster is a backup cluster.

6. The method for ensuring data consistency of dual clusters of tasks according to claim 5 is characterized by: The two clusters of the GBASE cluster type are GA cluster and GB cluster; The GA cluster is a main cluster, and the GB cluster is a standby cluster.

7. The method for ensuring data consistency of dual clusters of tasks according to claim 6 is characterized by: The task execution result includes table name, cycle and storage method.

8. A system for ensuring data consistency of dual clusters of tasks, used to implement the method for ensuring data consistency of dual clusters of tasks as claimed in any one of claims 1 to 7, characterized in that: include: The acquisition module is used to obtain two clusters of the same type; A data synchronization engine determination module is used to determine a data synchronization engine according to the type of cluster; A primary cluster and a standby cluster setting module is used to determine a primary cluster and a standby cluster from two clusters; The synchronization module is used to synchronize data between the primary cluster and the standby cluster according to the data synchronization engine.