Database cluster data processing method and device, equipment, medium and product
By determining the data fork points in the database cluster and comparing the data blocks by dichotomizing the dichotomy method, the problem of large amount of calculations during the synchronization of data blocks in the database cluster is solved, and fast and effective data consistency and synchronization efficiency are achieved.
Patent Information
- Application Number
- CN202411975521.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
In a database cluster scenario, when there are a large number of data blocks that need to be synchronized between the clusters of the database, the number of data blocks that need to be compared to determine the consistency points of the data block data is large, resulting in large amounts of calculations and long time.
By determining the data fork point between the source cluster and the target cluster, the data blocks are determined by dichotomous method, and the data blocks in the source cluster and the target cluster are consistently compared until the data consistency point is determined, and the data blocks are synchronized according to the consistent point.
It reduces the comparative calculations between the source cluster and the target cluster, and can determine the data consistency points faster, thereby quickly completing data block synchronization, improving the availability of database products.
Smart Images

Figure CN119938784A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of database technology, and in particular to a data processing method, device, equipment, medium and product of a database cluster. Background Art
[0002] In a database cluster scenario, different databases can be configured to store the same data, which is stored in the form of data blocks. When a database changes its timeline configuration or encounters an exception, the two databases continue to obtain and store data blocks at the same time, but the stored data blocks cannot be guaranteed to be the same.
[0003] In the prior art, a database with a timeline configuration change or anomaly will synchronize its data blocks stored after the occurrence of the situation to other databases to ensure that the data blocks stored between databases are consistent. Each data block in different database clusters is compared separately to determine the location of the data block to start synchronizing.
[0004] With the prior art, when a large number of data blocks need to be synchronized between database clusters, a large number of data blocks need to be compared when the database determines a data consistency point for starting synchronization of data blocks, resulting in a large amount of required calculations and a long time consumption. Summary of the invention
[0005] The present application provides a data processing method, device, equipment, medium and product for a database cluster to overcome the technical problem in the prior art that when a large number of data blocks need to be synchronized between database clusters, the database needs to compare a large number of data blocks when determining a data consistency point for starting synchronization of data blocks.
[0006] A first aspect of the present application provides a data processing method for a database cluster, comprising determining a data bifurcation point between a source cluster and a target cluster when a bifurcation occurs in a timeline or a data block between the source cluster and the target cluster; based on the data bifurcation point, using a binary search method to determine the data blocks, and performing a consistency comparison on the data blocks in the source cluster and the target cluster until a data consistency point between the source cluster and the target cluster is determined; wherein, in the process of using the binary search method to determine the data blocks, determining a data block for a next sampling based on a current comparison result; and synchronizing the data blocks in the source cluster to the target cluster according to the data consistency point.
[0007] A second aspect of the present application provides a data processing device for a database cluster, comprising: a determination module, used to determine a data bifurcation point between the source cluster and the target cluster when a bifurcation occurs in the timeline or data block between the source cluster and the target cluster; a comparison module, used to determine the data blocks by binary search based on the data bifurcation point, and to perform a consistency comparison on the data blocks in the source cluster and the target cluster until a data consistency point between the source cluster and the target cluster is determined; wherein, in the process of determining the data blocks by binary search, the data blocks for the next sampling are determined based on the current comparison result; and a synchronization module, used to synchronize the data blocks in the source cluster to the target cluster according to the data consistency point.
[0008] The third aspect of the present application provides an electronic device, comprising: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the method described in the first aspect of the present application.
[0009] A fourth aspect of the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed, implement the method described in the first aspect of the present application.
[0010] A fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, implements the method described in the first aspect of the present application.
[0011] In summary, the data processing method, device, equipment, medium and product of the database cluster provided in this embodiment, in order to achieve the synchronization of data blocks between database clusters, when determining the data consistency point between the destination cluster and the source cluster, it is not necessary to compare the consistency of each data block in the source cluster and the target cluster, but to determine the data block by binary search, determine the consistency comparison of some data blocks in the source cluster and the target cluster, and finally synchronize the data blocks in the source cluster to the target cluster according to the determined data consistency point. Therefore, in the data processing method of the database cluster provided in this embodiment, the comparison of the database in the source cluster and the target cluster can be reduced, and the data consistency point between the source cluster and the target cluster can be determined at a faster speed with a smaller amount of calculation, so as to complete the synchronization of the data blocks in the source cluster to the target cluster more quickly and effectively, and improve the efficiency of data block synchronization when the cluster is restored, and also improve the availability of database products, which is more conducive to the application and promotion of databases and their products. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0013] Figure 1 A schematic diagram of the application scenario of this application;
[0014] Figure 2 A schematic diagram of a state of a data block stored in a database cluster;
[0015] Figure 3 A schematic diagram of another state of data blocks stored in a database cluster;
[0016] Figure 4 A schematic diagram of determining data consistency points in the prior art;
[0017] Figure 5 A schematic diagram of a flow chart of an embodiment of a data processing method for a database cluster provided in the present application;
[0018] Figure 6 A schematic diagram of the binary method for determining data blocks provided by this application;
[0019] Figure 7 A flowchart of another embodiment of the data processing method for a database cluster provided by the present application;
[0020] Figure 8 A schematic diagram of the structure of a data processing device for a database cluster provided in this application;
[0021] Fig. 9 A schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0023] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0024] Figure 1 This is a schematic diagram of the application scenario of this application, such as Figure 1 As shown in the figure, in a database cluster scenario, different databases can be configured to store the same data, which is stored in the form of data blocks. However, when a database changes its timeline configuration or an exception occurs, the two databases continue to obtain and store data blocks at the same time, but the stored data blocks cannot be guaranteed to be the same. Therefore, when the data blocks in the database cluster are inconsistent, the data blocks need to be synchronized. Figure 1 In the scenario shown, the destination cluster can obtain data blocks from the source cluster and update its stored data blocks based on the data blocks of the source cluster to ensure that the data blocks stored in the destination cluster are consistent with the data blocks stored in the source cluster. The database cluster that provides data to other database clusters for synchronization is recorded as the source cluster, and the database cluster that receives data provided by the source cluster for synchronization is recorded as the target cluster.
[0025] Specifically, Figure 2 A schematic diagram of a state of data blocks stored in a database cluster, such as Figure 2 As shown, before time T, the source cluster DB1 stores data blocks D1, D2...DX in sequence through timeline 1, and the destination cluster DB2 stores the same data blocks D1, D2...DX in sequence through timeline 1. At this time, the data blocks stored in the source cluster DB1 and the destination cluster DB2 are the same and synchronized.
[0026] Assume that Figure 2At time T in the diagram, when the source cluster DB1 has a timeline configuration change, etc., which causes a fork in the timeline or data block between the source cluster DB1 and the target cluster DB2, the source cluster DB1 stores data blocks DX+1, DX+2, ..., DX+n in sequence through timeline 2 after time T, and the target cluster DB2 still stores data blocks Da, Db, ..., Dn in sequence through timeline 1 after time T. Since the timeline used by the source cluster DB1 has changed, the data stored between the source database DB1 and the target database DB2 after time T are different. At this time, the data block DX+1 of the source cluster DB1 and the data block Da of the target cluster DB2 after time T are recorded as the data fork point between the source cluster DB1 and the target cluster DB2.
[0027] Subsequently, when the timeline or data block of the source cluster DB1 and the target cluster DB2 diverges, the target cluster DB2 can update the data blocks stored after the data bifurcation point through the data blocks stored by the source cluster DB1 after the data bifurcation point, so as to synchronize the data blocks stored by the target cluster DB2 and the source cluster DB1, and ensure that the data blocks stored between the databases are consistent.
[0028] Figure 3 A schematic diagram of another state of data blocks stored in a database cluster, such as Figure 3 As shown, the destination cluster DB2 can update the data blocks stored by the destination cluster DB2 through timeline 1 through the data blocks DX+1, DX+2...DX+n stored by the source cluster DB1 after the data bifurcation point and through timeline 2. Specifically, the destination cluster DB2 deletes the data blocks Da, Db...Dn stored after the data bifurcation point, and stores the data blocks DX+1, DX+2...DX+n after the data bifurcation point through timeline 2. Figure 3 and Figure 2 It can be seen from the comparison that after the synchronization of the data blocks, the data blocks stored in the destination cluster DB2 after the data bifurcation point are the same as the data blocks stored in the source cluster DB1.
[0029] In order to achieve synchronization of data blocks between database clusters, an important prerequisite is to find the same data blocks between the target cluster DB2 and the source cluster DB1 after the data bifurcation point, so as to determine the position to start synchronization when synchronizing data blocks. Before this data block, the data blocks of the target cluster DB2 and the source cluster DB1 are the same. Before this data block, the data blocks of the target cluster DB2 and the source cluster DB1 begin to be inconsistent and need to be synchronized. Therefore, this data block can be called the data consistency point between the target cluster DB2 and the source cluster DB1.
[0030] Figure 4It is a schematic diagram of determining data consistency points in the prior art, wherein, in order to determine the data consistency points between the target cluster DB2 and the source cluster DB1, the target cluster DB2 will compare the data blocks DX+1, DX+2...DX+n stored by the source cluster DB1 after the data bifurcation point through timeline 2 with the data blocks Da, Db...Dn stored by the target cluster DB2 after the data bifurcation point, one by one, so as to determine the data consistency points.
[0031] For example, in Figure 4 In the example shown, data blocks DX+1, DX+2, ... DX+e of the source cluster DB1 are the same as data blocks Da, Db, ... De of the destination cluster DB2, while data blocks DX+e, ... DX+n of the source cluster DB1 are not completely the same as data blocks De, ... Dn of the destination cluster DB2. In this case, data block De in the destination cluster DB2 and data block DX+e in the source cluster DB1 are recorded as data consistency points. When the destination cluster DB2 subsequently synchronizes data blocks, it can first store data blocks after data block DX through timeline 2 after the data bifurcation point according to the data bifurcation point, and obtain data blocks DX+e, ... DX+n from the source cluster DB1 according to the data consistency point, and update the stored data blocks De, ... Dn, to achieve data block synchronization.
[0032] However, in Figure 4 In the prior art shown, when the number of data blocks that need to be synchronized between database clusters is large, the target cluster DB2 needs to compare a large number of data blocks when determining the data consistency point for starting to synchronize the data blocks. Once the number of data blocks between the data bifurcation point and the data consistency point is large, the target cluster DB2 needs to calculate more and it will take more time to determine the data consistency point.
[0033] Therefore, based on the problems existing in the above-mentioned prior art, the present application provides a data processing method for a database cluster to reduce the amount of calculation required to determine the data consistency point and reduce the time spent on determining the data consistency point. The technical solution of the present application is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0034] Figure 5 A flowchart of an embodiment of a data processing method for a database cluster provided in the present application is shown in FIG. Figure 5 The method shown can be applied to Figure 1 In the scenario shown, it is specifically executed by the destination cluster DB2, for example, it can be executed by a program script run by the destination cluster DB2. Specifically, the data processing method of the data block cluster provided in this embodiment includes:
[0035] S101: When a timeline or data block between a source cluster DB1 and a target cluster DB2 diverges, a data divergence point between the source cluster DB1 and the target cluster DB2 is determined.
[0036] Specifically, the target cluster DB2 can appear in the source cluster DB1 as follows Figure 2 When the timeline or data block shown in the figure is bifurcated, by traversing each line in the historical file information of the two clusters, the timeline and divergence point where the data blocks between the source cluster DB1 and the target cluster DB2 go to different branches are determined, so as to determine the data bifurcation point before the source cluster DB1 and the target cluster DB2 through the timeline and divergence point, and obtain the WAL file between the source cluster DB1 and the target cluster DB2 for subsequent processing.
[0037] S102: Based on the data bifurcation point determined in S101, a binary search method is used to determine the data blocks, and a consistency comparison is performed on the data blocks in the source cluster DB1 and the target cluster DB2 until a data consistency point between the source cluster DB1 and the target cluster DB2 is determined.
[0038] Specifically, in the embodiment of the present application, the target cluster DB2 does not need to perform consistency comparison on each data block in the source cluster DB1 and the target cluster DB2, but determines the consistency comparison on some data blocks in the source cluster DB1 and the target cluster DB2 by using the binary search method to determine the data blocks. In addition, in the process of using the binary search method to determine the data blocks, the data blocks to be sampled next time are also determined based on the comparison results of the current time, so that the data blocks to be sampled next time are jumpy relative to the current data blocks, rather than continuous.
[0039] In one embodiment, when comparing the consistency of data blocks, the target cluster DB2 calculates a first hash value for the current data block in the source cluster DB1 determined at the current time, and calculates a second hash value for the current data block in the target cluster DB2 determined at the current time, and compares the first hash value and the second hash value. If the first hash value is the same as the second hash value, it is determined that the current data block in the source cluster DB1 determined at the current time is consistent with the current data block in the source cluster DB1 determined at the current time; if the first hash value is different from the second hash value, it is determined that the current data block in the source cluster DB1 determined at the current time is inconsistent with the current data block in the source cluster DB1 determined at the current time.
[0040] In another embodiment, in order to improve the comparison efficiency, before comparing the data blocks, the target cluster DB2 first calculates whether the WAL file header data of the source cluster DB1 and the target cluster DB2 are consistent, and names the file header position as the data reading point for subsequent data block reading operations. When it is determined that the file header data of the source cluster DB1 and the target cluster DB2 are inconsistent, the subsequent data block comparison process is performed, thereby entering the following Figure 5 The calculation logic for finding data consistency points is shown in the figure. If the WAL file header data of the source cluster DB1 and the target cluster DB2 are consistent, the previous WAL file is used to calculate whether the header file data is consistent.
[0041] S103: Synchronize the data blocks in the source cluster DB1 to the target cluster DB2 according to the data consistency point determined in S102.
[0042] Specifically, the destination cluster DB2 can interact with the source cluster DB1 according to the determined data consistency point, so as to obtain the data blocks after the data consistency point in the source cluster DB1. The destination cluster DB2 deletes the data blocks after the data consistency point stored in itself, and reuses the data blocks after the data consistency point in the source cluster DB1 for storage after the data consistency point, so as to achieve the following: Figure 3 The data blocks stored in the destination cluster DB2 and the source cluster DB1 are consistent.
[0043] In summary, in order to achieve synchronization of data blocks between database clusters, the data processing method for database clusters provided in this embodiment does not need to compare the consistency of each data block in the source cluster DB1 with that in the target cluster DB2 when determining the data consistency point between the target cluster DB2 and the source cluster DB1. Instead, the data blocks are determined by binary search, and some data blocks in the source cluster DB1 and the target cluster DB2 are compared for consistency, and finally the data blocks in the source cluster are synchronized to the target cluster according to the determined data consistency point. Therefore, in the data processing method for database clusters provided in this embodiment, the comparison of the databases in the source cluster DB1 and the target cluster DB2 can be reduced, and the data consistency point between the source cluster DB1 and the target cluster DB2 can be determined at a faster speed with a smaller amount of calculation, so as to more quickly and effectively complete the synchronization of the data blocks in the source cluster DB1 to the target cluster DB2 during cluster recovery, and also improve the availability of database products, which is more conducive to the application and promotion of databases and their products.
[0044] Furthermore, in an embodiment of the present application, based on the consistent size of each data block of the WAL file, and the ordered and monotonically increasing LSN of the WAL file, the data blocks that need to be compared for consistency can be determined according to the binary search method. The following, in conjunction with the accompanying drawings, describes the method of determining data blocks using the binary search method in an embodiment of the present application, as well as the processing logic of determining the data blocks for the next comparison based on the current comparison result.
[0045] Figure 6 The following is a schematic diagram of the binary method used in this application to determine the data block. Figure 6 As shown, for each determined database, the destination cluster DB2 performs a consistency comparison between the current data block in the source cluster DB1 determined at the current time and the current data block in the target cluster DB2.
[0046] If the comparison result of the current data block is consistent, the current consistent data block is recorded as a temporary data consistency point, and the next data block to be compared is determined by binary search from the data block range after the current data block in the source cluster DB1 and the target cluster DB2. Figure 6 Taking the current data block D50 as an example, if the comparison results of the current data block D50-1 of the source cluster DB1 and the current data block D50-2 of the target cluster DB2 are consistent, the binary search method is used to determine that the next data block to be compared is D75 from the data block range D51-D100 after the current data block in the source cluster DB1 and the target cluster DB2. Therefore, the next data block to be compared is the current data block D75-1 and the current data block D75-2 of the target cluster DB2.
[0047] If the comparison result of the current data block is inconsistent, the next data block to be compared is determined by binary search from the data block range after the current data block in the source cluster DB1 and the target cluster DB2. Figure 6 Taking the current data block D50 as an example, if the comparison results of the current data block D50-1 of the source cluster DB1 and the current data block D50-2 of the target cluster DB2 are inconsistent, the binary search method is used to determine that the data block to be compared next is D25 from the data block range D1-D50 before the current data block in the source cluster DB1 and the target cluster DB2. Therefore, the data block to be compared next is the current data block D25-1 and the current data block D25-2 of the target cluster DB2.
[0048] In the specific implementation process, the destination cluster DB2 repeats the following steps when performing the above comparison: Figure 6The operation of determining the next comparison data block in the example shown. In a specific implementation, the destination cluster DB2 can first set the traversal times n=1 by means of a loop traversal, and based on the LSN of the WAL file, calculate the LSN+WAL file data block / (n*2) by the formula to obtain the reading point of the data block, thereby reading the next comparison data block from the source cluster DB1 and the target cluster DB2 respectively, and modifying the traversal times by n=n+1. Subsequently, when the destination cluster DB2 is based on the current comparison result being consistent, the formula: the previous data reading point+WAL file data block / (n*2) is used to traverse backward to obtain the next comparison data block; when the current comparison result is inconsistent, the formula: the previous data reading point-WAL file data block / (n*2) is used to traverse forward to obtain the next comparison data block. Finally, the destination cluster DB2 repeats the above traversal process, and determines that the traversal is completed when n is greater than or equal to the number of data blocks / 2.
[0049] Based on the data block comparison process of the above-mentioned cyclic traversal, the destination cluster DB2 records multiple temporary data consistency points. Therefore, after the traversal is completed, the destination cluster DB2 determines the smallest temporary data consistency point from all recorded temporary data consistency points that are consistent with the comparison as the final data consistency point. This consistency point is the temporary consistency point with the smallest number of data blocks between the data bifurcation point.
[0050] Finally, after the destination cluster DB2 determines the data consistency point, it can synchronize the data blocks in the source cluster DB1 to the destination cluster DB2 according to the determined data consistency point. When synchronizing the data blocks, the destination cluster DB2 further determines whether the determined data consistency point is before the last checkpoint. If the data consistency point is before the last checkpoint, it means that the data blocks before the last checkpoint have been written from the source cluster DB1 to the destination cluster DB2 and stored on the disk, and the subsequent data block synchronization processing can be directly performed based on the determined data consistency point; if the data consistency point is after the last checkpoint, it means that the data blocks before a checkpoint have not been written from the source cluster DB1 to the destination cluster DB2 and stored on the disk, and the subsequent data block synchronization processing is performed based on the last checkpoint as the data consistency point.
[0051] Figure 7 A flowchart of another embodiment of the data processing method for the database cluster provided in the present application is shown as follows: Figure 7A specific implementation method of a data processing method for a database cluster is shown. Among them, when the destination cluster synchronizes the data block, it first determines the data bifurcation point between the data blocks of the destination cluster and the source cluster from the history file, and determines the WAL file where the bifurcation point is located. Read the first data block at the data bifurcation point from the WAL file and calculate the hash value, and calculate the hash value of the data block from the corresponding position of the source cluster. When the two hash values are inconsistent, read the previous WAL file from the bifurcation point. If the two hash values are consistent, the number of data blocks currently calculated is recorded as x, and the data block in the destination cluster initially read is the data block corresponding to the LSN obtained by adding x / 2 to the LSN of the 0th data block. Then compare the data block with the hash value of the corresponding data block in the source cluster. If the hash values are consistent, store the current temporary data consistency bit, and obtain the new data block in the destination cluster by adding the LSN of the current data block to the data block of the WAL file / (n*2). If the hash values are inconsistent, obtain the new data block in the destination cluster by subtracting the data block of the WAL file / (n*2) from the LSN of the current data block. Then, compare the updated data block with the hash value of the corresponding data block in the source cluster again until the traversal is completed. From the multiple temporary data consistency bits recorded, determine the data consistency point between the destination cluster and the source cluster with the smallest corresponding LSN. Finally, according to the determined data consistency point, the data blocks in the source cluster can be synchronized to the target cluster.
[0052] More specifically, based on Figure 7In the embodiment shown, the target cluster can determine the data consistency point by executing the following steps in sequence. 1. By traversing each line in the historical information files of the two clusters, find the timeline and divergence point where the data goes to different branches. Among them, find the corresponding WAL file through the timeline and divergence point. 2. First, determine whether the WAL file header data is consistent, and name the file header position as the data reading point to improve efficiency. 3. Obtain this data block of the target cluster according to the data reading point, and calculate its hash value. 4. Obtain the hash value corresponding to the data at this location from the source cluster. 5. Determine whether the hash value of the data location of the target cluster is consistent with the hash value of the source cluster. If not, it indicates that the data of the source cluster and the target cluster are still different, and go to step 7. If they are consistent, it indicates that the data of the source cluster and the target cluster at this location are consistent, and go to the logic of finding the consistent point, step 8. 6. Continue to use the previous WAL file to enter the calculation and go to step 3. 7. Calculate the consistent point, set the number of traversals to n = 1: During the traversal process, execute a) WAL file header LSN + WAL file data block / (n*2) as the data reading point; b) traverse The number of traversals n = n + 1; c) Calculate whether the hash values of the source cluster and the target cluster data at this point are consistent; d); Determine whether the traversal is complete: n> = number of data blocks / 2, at this time the file traversal is complete, and go to step 9 to exit the traversal; e) Inconsistency indicates that the data consistency point is close to the end of the file header, then the new data reading point is the previous data reading point-WAL file data block / (n*2), and go to step c); f) If consistent, record this data reading point lsn as a temporary consistency point, and continue to traverse backwards, the new data reading point is the previous data reading point + WAL file data block / (n*2), and go to step c). 8. Get the minimum temporary consistency bit recorded as a quasi-consistent point. 9. Determine whether the quasi-consistent point at this time is before the last checkpoint in the source cluster: a) If it is before the last checkpoint, it indicates that the data has been stored on the disk, and it can be directly used as a consistency point; b) If it is after the last checkpoint, it indicates that the data has not been stored on the disk, and the last checkpoint is used as the consistency point.
[0053] In the above embodiments of the present application, the data processing method of the database cluster provided by the embodiments of the present application is introduced. In order to realize the functions in the data processing method of the database cluster provided by the embodiments of the present application, the target cluster as the execution subject can be realized by hardware structure and / or software module, for example, in the form of hardware structure, software module, or hardware structure plus software module to realize the above functions. Whether one of the above functions is executed in the form of hardware structure, software module, or hardware structure plus software module depends on the specific application and design constraints of the technical solution.
[0054] For example, Figure 8 A schematic diagram of the structure of a data processing device for a database cluster provided in this application, such as Figure 8 The device shown can be used to execute the data processing method of the database cluster provided in any embodiment of the present application. In one embodiment, Figure 8 The data processing device 1000 of the database cluster shown includes: a determination module 1001, a comparison module 1002 and a synchronization module 1003. The determination module 1001 is used to determine the data bifurcation point between the source cluster and the target cluster when the timeline or data block between the source cluster and the target cluster bifurcates; the comparison module 1002 is used to determine the data blocks by binary search based on the data bifurcation point, and to perform consistency comparison on the data blocks in the source cluster and the target cluster until the data consistency point between the source cluster and the target cluster is determined; in the process of determining the data blocks by binary search, the data blocks to be sampled next time are determined based on the comparison result of the current time; the synchronization module 1003 is used to synchronize the data blocks in the source cluster to the target cluster according to the data consistency point.
[0055] The specific implementation method and principle of the data processing device of the above database cluster refer to the description of the data processing method of the above database cluster, which will not be repeated here.
[0056] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the processing module can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The function of the above-mentioned module is determined. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or instructions in the form of software.
[0057] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0058] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0059] For example, Fig. 9 A schematic diagram of the structure of an electronic device provided in this application, such as Fig. 9 The device shown can be used to execute the data processing method of the database cluster provided in any embodiment of the present application. In one embodiment, Fig. 9The electronic device 2000 shown includes one or more processors 2001 and a memory 2002; wherein the memory 2002 is used to store computer executable instructions, and the processor 2001 can execute the computer executable instructions stored in the memory 2002. When the computer executable instructions are executed by the processor 2001, the processor 2001 implements the data processing method of any database cluster in the aforementioned embodiments of the present application. In one embodiment, Figure 8 The electronic device 2000 shown further includes a communication interface 2003 , wherein the processor 2001 can communicate with other devices via the communication interface 2003 , for example, to obtain data blocks of a source cluster.
[0060] The present application also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed, they can be used to implement the data processing method of any database cluster in the aforementioned embodiments of the present application.
[0061] An embodiment of the present application also provides a chip for executing instructions, wherein the chip is used to execute the data processing method of any database cluster as described above in the present application.
[0062] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed, implements any of the aforementioned database cluster data processing methods of the present application.
[0063] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method for a database cluster, characterized in that: include: When a timeline or a data block between a source cluster and a target cluster diverges, determining a data divergence point between the source cluster and the target cluster; Based on the data bifurcation point, a binary search method is used to determine the data blocks, and a consistency comparison is performed on the data blocks in the source cluster and the target cluster until a data consistency point between the source cluster and the target cluster is determined; wherein, in the process of determining the data blocks by the binary search method, the data blocks to be sampled next time are determined based on the comparison result of the current time; According to the data consistency point, the data blocks in the source cluster are synchronized to the target cluster.
2. The method according to claim 1, characterized in that The method of using binary sampling method to perform consistency comparison on the data blocks in the source cluster and the target cluster until the data consistency point between the source cluster and the target cluster is determined includes: Performing consistency comparison between the current data block in the source cluster determined at the time and the current data block in the target cluster; If the current data blocks are consistent, the consistent data blocks are recorded, and a binary search method is used to determine the next data block to be compared from the data block range after the current data block in the source cluster and the target cluster; If the current data blocks are inconsistent, a binary search method is used to determine the next data block to be compared from the data block range before the current data block in the source cluster and the target cluster; Until the next data block to be compared cannot be determined by using the binary search method, the data consistency point between the source cluster and the target cluster is determined based on all consistent databases.
3. The method according to claim 2, characterized in that The comparing the consistency of the current data block in the source cluster determined at the time with the current data block in the target cluster includes: Calculate a first hash value of a current data block in the source cluster determined at a current time, and a second hash value of a current data block in the target cluster; If the first hash value is the same as the second hash value, determining that the current data block in the source cluster determined at the current time is consistent with the current data block in the target cluster; If the first hash value is different from the second hash value, it is determined that the current data block in the source cluster currently determined is inconsistent with the current data block in the target cluster.
4. The method according to claim 3, characterized in that Before calculating the first hash value of the current data block in the source cluster determined at the current time and the second hash value of the current data block in the target cluster, the method further includes: Compare whether the file header data of the source cluster and the target cluster are consistent. If they are inconsistent, calculate the first hash value of the current data block in the source cluster determined at the time and the second hash value of the current data block in the target cluster.
5. The method according to any one of claims 2 to 4, characterized in that: The determining, based on all consistent data blocks, a data consistency point between the source cluster and the target cluster includes: From all consistent data blocks, determine the data block with the least number of data blocks between it and the data bifurcation point as the data consistency point between the source cluster and the target cluster.
6. The method according to claim 5, characterized in that The synchronizing the data blocks in the source cluster to the target cluster according to the data consistency point includes: Determine whether the data consistency point is before the last checkpoint; If yes, synchronizing the data blocks after the data consistency point in the source cluster to the target cluster; If not, the data blocks after the last checkpoint in the source cluster are synchronized to the target cluster.
7. A data processing device for a database cluster, characterized in that: include: A determination module, configured to determine a data bifurcation point between the source cluster and the target cluster when a bifurcation occurs in a timeline or a data block between the source cluster and the target cluster; A comparison module is used to compare the data blocks in the source cluster and the target cluster for consistency based on the data bifurcation point by using a binary search method to determine the data blocks, until a data consistency point between the source cluster and the target cluster is determined; wherein, in the process of using the binary search method to determine the data blocks, the data blocks to be sampled next time are determined based on the comparison result of the current time; A synchronization module is used to synchronize the data blocks in the source cluster to the target cluster according to the data consistency point.
8. An electronic device, characterized in that: include: Memory and processor; The memory stores computer executable instructions; The processor executes the computer executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: Computer executable instructions are stored, and when the computer executable instructions are executed, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed.