Data synchronization optimization method and system of DRBD block device and medium

By partitioning DRBD block devices and using dynamic bitmaps and generation identifier exchanges to optimize data synchronization of DRBD block devices, the network bandwidth consumption and latency issues during split-brain recovery are resolved, enabling efficient and flexible data synchronization. This approach is suitable for storage systems with high consistency and high availability requirements.

CN120768906APending Publication Date: 2025-10-10HANGZHOU EBOYLAMP ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510640977.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing DRBD block devices suffer from high network bandwidth consumption, long synchronization delay, and low efficiency when fully synchronizing in a split-brain recovery scenario, and are unable to effectively reduce redundant data transmission.

Method used

The DRBD block device is divided into multiple partitions, and a dynamic bitmap is used to mark the partition data modification status. Through partition management and dynamic bitmap fusion, only the partition data is modified synchronously. Combined with generation identifier exchange and priority synchronization mechanism, differentiated synchronization is achieved.

Benefits of technology

It significantly shortens data synchronization time, reduces resource consumption, improves synchronization efficiency, ensures data consistency and high availability, and is suitable for large-scale storage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120768906A_ABST
    Figure CN120768906A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data synchronization, in particular to a data synchronization optimization method and system of DRBD block equipment and a medium. The method comprises the following steps: dividing DRBD block equipment into a plurality of partitions; distributing a flag bit for each partition so as to track and record the modification state of data of each partition, and obtaining a dynamic bitmap of each DRBD block device based on the flag bits of different partitions; performing initialization synchronization on DRBD block devices of the master node and the slave node; when the split brain recovers, the DRBD block devices of the master node and the slave node exchange respective dynamic bitmaps and fuse the dynamic bitmaps to obtain a modified partition set; the master node synchronizes the modified partition data to the slave node only based on the modified partition set, and resets the dynamic bitmap of the DRBD block device of each node after synchronization is completed, so that invalid data transmission during full-amount synchronization is reduced fundamentally, data synchronization time is shortened remarkably, data synchronization efficiency is improved, and resource consumption is reduced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data synchronization, and in particular to a data synchronization optimization method for a DRBD block device. Background Art

[0002] DRBD (Distributed Replicated Block Device) is a distributed replicated storage system used to build high-availability computer clusters. It consists of a kernel driver module, userspace management tools, and related shell scripts. Its operating principle is similar to RAID 1 storage. By replicating block devices between geographically distributed servers, DRBD ensures that if one node fails, data can be quickly restored from another node, thereby improving system reliability.

[0003] In traditional DRBD block device data synchronization, efficiency and performance issues are particularly prominent, especially in split-brain recovery scenarios. For example, in a dual-machine production environment, network issues can cause a split-brain phenomenon between the DRBD block device master and slave nodes. During a split-brain period, the two nodes lose communication and cannot determine each other's status. The DRBD block device data on both nodes may be updated independently. Therefore, when recovering from a split-brain, it is necessary to ensure that the DRBD block device data between the two nodes is consistent. To this end, a full data synchronization mechanism is typically used to synchronize all DRBD block device data from the master node to the slave node. Specifically, during the data synchronization process, the entire data block device content of the source node must be completely copied to the destination node (including data that was not modified during the split-brain period). Therefore, when the amount of synchronized data is large, it will inevitably lead to significant network bandwidth consumption and synchronization delays, seriously affecting the business stability and read and write performance of the production environment.

[0004] To this end, the existing technology has made some useful explorations into data synchronization methods between master and slave nodes of DRBD block devices, and is dedicated to optimizing data synchronization between DRBD block devices.

[0005] For example, Chinese patent CN114385573A, a DRBD block device initialization method, apparatus, computer device, and storage medium, requires the user to set a data synchronization range in a configuration file when initializing the DRBD block device. DRBD then performs data synchronization between the master and slave replica block devices based on the set data synchronization range. The master replica block device synchronizes all data within the set data synchronization range to the slave replica block device. Although this solution can reduce the amount of synchronized data to a certain extent by setting the synchronization range in the configuration file during initialization, it cannot dynamically identify the data blocks that actually need to be synchronized within the range. Even if some data has not been modified after initialization, it will still be fully synchronized, occupying excess link bandwidth, resulting in wasted network bandwidth and low DRBD data synchronization efficiency.

[0006] For example, Chinese patent CN115033423A, a dual-machine hot standby DRBD initialization synchronization method, device, and equipment, modifies the corresponding bit in the quick synchronization bitmap to "1" or "0" based on whether the disk has data or not, and determines whether to synchronize based on the bit value. When initializing full synchronization, it is not necessary to synchronize the entire disk to the peer node, only the disk blocks with data need to be synchronized to the peer node, thereby improving the initialization synchronization efficiency of DRBD. However, this solution only optimizes the synchronization range during initialization. In a split-brain recovery scenario, it is still necessary to synchronize the block device area of ​​all data on the source end (even if the data in some areas has not been modified). Redundant data transmission cannot be avoided, resulting in excessive synchronization time for TB-level devices and low recovery efficiency.

[0007] For example, Chinese patent CN115277606A describes a method for optimizing DRBD data synchronization. This solution detects peak disk read / write and network bandwidth, calculates the DRBD write rate and the time required for DRBD to complete a single write request, and automatically writes these values ​​to the DRBD configuration file to dynamically adjust the parameters. This prevents DRBD synchronization data timeouts, excessive resource consumption, and improves the stability of the DRBD system. Although this solution avoids resource exhaustion by adjusting the synchronization rate and timeout parameters, it does not reduce the actual amount of data synchronized. For large-scale data synchronization scenarios (such as split-brain recovery), it cannot shorten synchronization time and only alleviates timeout issues rather than optimizing the root cause. Summary of the Invention

[0008] In response to the above technical problems, the present invention proposes a data synchronization optimization method, system and medium for DRBD block devices, aiming to fundamentally reduce invalid data transmission during full synchronization, significantly shorten data synchronization time, improve data synchronization efficiency, and reduce resource consumption. It is suitable for large-scale storage scenarios with strict requirements on data consistency and high availability.

[0009] In a first aspect, the present application provides a method for optimizing data synchronization of a DRBD block device, comprising the following steps:

[0010] Step 101: Divide the DRBD block devices on the master node and the slave node into several partitions;

[0011] Step 102: assign a flag bit to each partition, where the flag bit is used to mark the modification status of the partition data, and obtain a dynamic bitmap of each DRBD block device based on the flag bits of different partitions;

[0012] Step 103: Initialize and synchronize the DRBD block devices of the master and slave nodes.

[0013] Step 104: When the split-brain is recovered, the DRBD block devices of the master and slave nodes exchange their respective dynamic bitmaps, merge the dynamic bitmaps of the DRBD block devices of the master and slave nodes, and obtain a modified partition set;

[0014] Step 105 , the master node filters the unmodified partition data, and synchronizes the modified partition data to the slave nodes based on the modified partition set only, and resets the dynamic bitmap of the DRBD block device of each node after synchronization is completed.

[0015] In some embodiments, in step 102, a flag bit is assigned to each partition, where the flag bit is used to mark the modification status of the partition data. A dynamic bitmap of each DRBD block device is obtained based on the flag bits of different partitions, including:

[0016] A binary flag is assigned to each partition. The flag is set to "0" by default. When the data in the partition is modified, the flag corresponding to the partition is set to "1";

[0017] The DRBD block device of each node forms a dynamic bitmap based on the flag bit of each partition, and the index of each flag bit in the dynamic bitmap is mapped one-to-one with the number of each partition.

[0018] In some embodiments, in step 103, initializing and synchronizing the DRBD block device includes:

[0019] Set all flags in the dynamic bitmap of the DRBD block device of the primary node to "1" to indicate that all data needs to be synchronized;

[0020] The master node reads the partition data corresponding to the flag position "1" and synchronizes the read partition data to the slave node in sequence to make the DRBD block device data of the master and slave nodes completely consistent. After each synchronization, the master and slave nodes simultaneously set the flag position of the corresponding synchronization partition to "0";

[0021] After the initial synchronization is completed, the dynamic bitmap of the DRBD block device of the master node is traversed again, and the corresponding partition data of the flag position "1" in the dynamic bitmap is synchronized to the slave node.

[0022] In some embodiments, before step 104, the method further includes:

[0023] Step 201: Constructing generation identifiers for tracking data versions and dynamic bitmap versions for the DRBD block devices of the master and slave nodes;

[0024] Step 202: When the connection between the split-brain recovery nodes is re-established, the nodes at both ends exchange their generation identification information;

[0025] In step 203, the generation identification information of the local node is compared with the generation identification information of the opposite node, and whether a split-brain occurs in the DRBD block device of the dual-end node is determined based on the comparison result. If it is determined that a split-brain does not occur, the data synchronization strategy is determined based on the comparison result. If it is determined that a split-brain occurs, step 104 is executed.

[0026] In some embodiments, the generation identifier includes a current data version identifier and a dynamic bitmap version identifier. In step 203, the generation identifier information of the local node is compared with the generation identifier information of the opposite node. Based on the comparison result, it is determined whether a split-brain occurs in the DRBD block device of the dual-end node. If it is determined that a split-brain does not occur, a data synchronization strategy is determined based on the comparison result, including:

[0027] Compare the current data version identifier of the local node with the current data version identifier of the peer node. If the current data version identifiers of the two nodes are consistent, it is determined that no split-brain has occurred. If the current data version identifiers of the two nodes are inconsistent, it is determined that a split-brain has occurred.

[0028] When it is determined that no brain split has occurred, the dynamic bitmap version identifier of the local node is compared with the dynamic bitmap version identifier of the opposite node. If the dynamic bitmap version identifiers of the two nodes are consistent, no data synchronization is required. If the dynamic bitmap version identifiers of the two nodes are inconsistent, the node with the newer dynamic bitmap version identifier is used as the source end and the other end node is used as the destination end. The corresponding modified partition data in the dynamic bitmap of the source end is synchronized to the destination end, and the dynamic bitmap of the source end is reset after the synchronization is completed.

[0029] In some embodiments, in step 104, the dynamic bitmaps of the DRBD block devices of the master node and the slave node are merged to obtain a modified partition set, including:

[0030] Perform a logical OR operation on the dynamic bitmaps of the DRBD block devices of the master node and the slave node to generate a dynamic bitmap that merges the modification status of both nodes;

[0031] The set of partitions with flag position "1" in the dynamic bitmap of the modification status of both parties is fused to form a modified partition set.

[0032] In some embodiments, the following steps are also included:

[0033] Periodically count the number of writes to each partition in each DRBD block device in each cycle;

[0034] Based on the number of writes to each partition in each cycle, the access popularity of each partition is dynamically evaluated to obtain the access popularity value of each partition;

[0035] In step 105, the master node filters the unmodified partition data and synchronizes the modified partition data to the slave node based on the modified partition set, including:

[0036] Determine the synchronization priority of each modified partition in the modified partition set based on the current access heat value of each partition, wherein the synchronization priority is proportional to the access heat value;

[0037] The master node synchronizes the corresponding modified partition data to the slave node in descending order according to the synchronization priority of each modified partition.

[0038] In some embodiments, based on the number of writes to each partition in each cycle, the access popularity of each partition is dynamically evaluated to obtain the access popularity value of each partition, including:

[0039] Get the timestamp of each write to each partition;

[0040] The timeliness attenuation factor of each partition is calculated based on the timestamp of the last write in the current cycle of each partition;

[0041] The total number of writes in the current cycle of each partition and the timeliness attenuation factor are weighted and summed to obtain the access heat value of each partition in the current cycle.

[0042] In a second aspect, the present application provides a data synchronization optimization system for a DRBD block device, comprising:

[0043] A dynamic bitmap construction module is used to divide the DRBD block devices on the master node and the slave node into several partitions, assign a flag bit to each partition, and use the flag bit to mark the modification status of the partition data. Based on the flag bits of different partitions, a dynamic bitmap of each DRBD block device is obtained;

[0044] Initialization synchronization module, used to initialize and synchronize the DRBD block devices of the master and slave nodes;

[0045] The split-brain recovery module is used to exchange the dynamic bitmaps of the DRBD block devices of the master and slave nodes during split-brain recovery, and merge the dynamic bitmaps of the DRBD block devices of the master and slave nodes to obtain the modified partition set;

[0046] The data synchronization module is used by the master node to filter the unmodified partition data and synchronize the modified partition data to the slave node based on the modified partition set only. After the synchronization is completed, the dynamic bitmap of the DRBD block device of each node is reset.

[0047] In a third aspect, a computer-readable storage medium stores a computer program thereon, wherein when the computer program is executed by a processor, the method for optimizing data synchronization of a DRBD block device as described above is implemented.

[0048] The beneficial technical effects of the present invention include at least:

[0049] 1. A data synchronization optimization method, system, and media for DRBD block devices are used. Different from the static range optimization of existing technologies, this system builds an efficient and reliable partition-level modification tracking system through partition management and dynamic bitmap marking. By decoupling state in the time dimension, the time-consuming full synchronization is decomposed into the superimposed transmission of "baseline data initialization synchronization + incremental dirty data". The design of coordinated state fusion and differential synchronization overcomes the contradiction of "full and incremental synchronization cannot be achieved at the same time" in traditional synchronization solutions. While ensuring strong consistency, it reduces service latency and improves network resource utilization. It also implements dynamic differential data filtering, fundamentally avoiding the transmission of invalid data (i.e., unmodified redundant data), and achieves efficient, flexible, and highly available distributed storage management.

[0050] 2. Dividing the DRBD block device into multiple partitions, the "dirty page marking" concept from the storage system is creatively introduced into distributed block device synchronization. Through a minimalist metadata design, an efficient and reliable partition-level modification tracking system is constructed. The determination of data modifications is simplified from "block-by-block comparison" to "flag bit query", providing accurate dynamic recording of data modification status for DRBD split-brain recovery, thereby providing data support for subsequent differentiated synchronization and avoiding full-data transmission. This overcomes the static configuration limitations and full-data synchronization redundancy of existing technologies and dynamically identifies data areas that need to be synchronized. Moreover, the space complexity of the dynamic bitmap is O(n), where n is the number of partitions, which is much lower than traditional logging solutions and can significantly reduce resource consumption.

[0051] 3. Generation identifier exchange is used as the cornerstone of DRBD data consistency management. The modification tracking mechanism based on the dynamic bitmap is an optimization layer based on this. The two form a layered verification mechanism: First, generation identifiers handle data version control and policy decisions at the macro level. Specifically, this embodiment introduces the version control concept of distributed systems into the storage layer. Through the global uniqueness of generation identifiers, the complex split-brain recovery problem is transformed into a traceable metadata comparison problem, achieving accurate synchronization direction decisions, and quickly judging full or differential synchronization. The synchronization granularity is adaptive, and the split time point and impact range can be located through historical data version identifiers, and data conflicts can be traced. Then, after the generation identifier establishes the synchronization strategy, the modification tracking mechanism based on the dynamic bitmap handles the data partition optimization at the micro level, fundamentally reducing invalid data transmission during full synchronization. This layered collaborative design realizes a split-brain recovery paradigm with traceable strategies, quantifiable execution, and locatable conflicts. When facing TB-level devices, it can greatly improve the data synchronization efficiency of split-brain recovery, while further ensuring data consistency and reliability in split-brain scenarios.

[0052] 4. By introducing a priority synchronization mechanism, the technical defects of the traditional DRBD "blind synchronization" solution are compensated. First, the modification popularity of each partition is counted. When synchronizing data, the synchronization priority is sorted according to the access popularity of the partition. The partitions with high frequency modifications are synchronized first to ensure that key data is restored first, shorten the recovery time of the business critical path, and significantly reduce the business interruption time. It can further optimize the overall synchronization performance based on the parameter optimization of existing technologies (such as Chinese patent CN115277606A).

[0053] Other features and advantages of the present invention will be disclosed in detail in the following specific embodiments and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The present invention will be further described below with reference to the accompanying drawings:

[0055] Figure 1 This is a flow chart of a data synchronization optimization method for a DRBD block device according to an embodiment of the present invention.

[0056] Figure 2 This is a structural diagram of a data synchronization optimization system for a DRBD block device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The following is an explanation and description of the technical solutions of the embodiments of the present invention in conjunction with the drawings of the embodiments of the present invention. However, the following embodiments are only preferred embodiments of the present invention and are not exhaustive. Based on the embodiments in the implementation manner, other embodiments obtained by those skilled in the art without creative work are all within the scope of protection of the present invention.

[0058] In the following description, the appearance of terms such as "inner", "outer", "upper", "lower", "left", "right", etc. is merely for the convenience of describing the embodiments and simplifying the description, and does not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0059] Embodiment one:

[0060] Please refer to the accompanying Figure 1 , Figure 1 A flowchart of a data synchronization optimization method of a DRBD block device provided by an embodiment of the present specification is shown.

[0061] As Figure 1 shown, the data synchronization optimization method of the DRBD block device can at least include the following steps:

[0062] Step 101, divide the DRBD block device on the master node and the slave node into several partitions.

[0063] Among them, the DRBD block device works in a master-slave (Primar / Secondany) mode, the master node is responsible for processing write requests, and the slave node is a backup. The DRBD block device on the master node is promoted to Primany and is responsible for receiving write data. When the data reaches the DRBD block device, one continues to write to the local disk to realize the persistence of data, and at the same time, one copy of the received data to be written is sent to the DRBD block device on the opposite end (i.e. the slave node). The corresponding DRBD block device on the other host again stores the received data into its own disk, realizing shared storage.

[0064] Among them, the partition size of the DRBD block device can be pre-set in the DRBD resource configuration file according to the business needs of the production environment, and this embodiment does not limit it.

[0065] It can be understood that the partitioning process of the DRBD block device is crucial, which will directly affect the data synchronization time. By reasonably configuring the partition size, it can ensure that the system is more efficient in utilizing network resources during data synchronization, reduces invalid data synchronization, shortens the time required for synchronization, and thus improves the overall system performance.

[0066] Step 102, assign a flag bit to each partition, the flag bit is used to mark the modification state of the partition data, and the dynamic bitmap of each DRBD block device is obtained based on the flag bits of different partitions.

[0067] It is understandable that when the communication between the master and slave nodes is disconnected and DRBD split-brain occurs, the slave node may be upgraded to the master node and allow business data to be written. At this time, the DRBD block devices of both nodes allow data modification. The modification tracking of traditional DRBD block devices relies on logs or full comparisons, but the logs need to record detailed operations and the metadata overhead is very high. For this reason, this embodiment divides the DRBD block device into multiple partitions, and each partition records the data modification status through a dynamic bitmap, so that the master and slave nodes can independently record the modification status of their respective partitions during the split-brain period, providing a data basis for subsequent differentiated synchronization, thereby overcoming the static configuration limitations and full synchronization redundancy in the existing technology, and dynamically identifying the data areas that need to be synchronized.

[0068] The data on a DRBD block device includes service data and metadata. Service data includes actual stored user data, while metadata includes information for managing data synchronization. This embodiment uses a dynamic bitmap as metadata, binding it to other DRBD metadata to ensure that modified states are not lost after a node restart.

[0069] Specifically, in this embodiment, in step 102, a flag bit is assigned to each partition. The flag bit is used to mark the modification status of the partition data. A dynamic bitmap of each DRBD block device is obtained based on the flag bits of different partitions, including:

[0070] Each partition is assigned a binary flag bit, which is set to "0" by default. When the data in the partition is modified (that is, data is written to a partition), the flag bit corresponding to the partition is set to "1";

[0071] The DRBD block device of each node forms a dynamic bitmap based on the flag bit of each partition. The index of each flag bit in the dynamic bitmap is mapped one-to-one with the number of each partition.

[0072] It can be understood that the implementation principle of the modification tracking mechanism based on the dynamic bitmap of this embodiment is as follows: First, a strict mapping logic between partitions and binary bits is adopted, and each physical partition uniquely corresponds to a flag bit (Bit) in the dynamic bitmap, ensuring that any data modification can be accurately located to the corresponding partition; secondly, the flag bit status is updated (set to 1 or cleared to 0) through kernel-level atomic operations, effectively avoiding state conflicts caused by concurrent writing of multiple threads. In the data writing process, the DRBD kernel adopts a write-time marking strategy - before the application's I / O request is actually written to the disk, the flag position 1 operation of the corresponding partition is triggered first, regardless of whether the modification range is a single byte or the entire partition. This marking process will be triggered. Finally, the dynamic bitmap, as key metadata, maintains persistent synchronization with the DRBD standard metadata area (usually stored at the end of the disk or a designated external storage device), ensuring that the historical modification status can still be accurately loaded after the node is restarted or the fault is recovered, and ensuring the consistency maintenance of the distributed storage system.

[0073] It is understandable that this embodiment creatively introduces the concept of "dirty page marking" in storage systems into distributed block device synchronization. Through a minimalist metadata design, it builds an efficient and reliable partition-level modification tracking system, simplifying the determination of data modifications from "block-by-block comparison" to "flag bit query." This provides DRBD split-brain recovery with the ability to accurately and dynamically record data modification status, thereby providing data support for subsequent differentiated synchronization and avoiding full-data transmission. This overcomes the static configuration limitations and full-data synchronization redundancy of existing technologies and dynamically identifies data areas that require synchronization. Moreover, the spatial complexity of the dynamic bitmap is O(n), where n is the number of partitions, which is far lower than traditional logging solutions and can significantly reduce resource consumption.

[0074] Step 103: Initialize and synchronize the DRBD block devices of the master and slave nodes.

[0075] It can be understood that, in this embodiment, after the DRBD block device is created, a working thread drbd_work is created for sending and receiving data. After the DRBD resource is created, both nodes are initially in the slave node state. After one end node is upgraded to the master node, the drbd_work thread of the master node begins to initialize the full synchronization of data.

[0076] Specifically, in this embodiment, in step 103, initializing and synchronizing the DRBD block device includes:

[0077] Step 1031: Set all flags in the dynamic bitmap of the DRBD block device of the master node to "1" to indicate that all data needs to be synchronized.

[0078] In step 1032, the master node reads the partition data corresponding to the flag position "1", and synchronizes the read partition data to the slave node in sequence to make the DRBD block device data of the master and slave nodes completely consistent, and after each synchronization, the master and slave nodes simultaneously set the flag position of the corresponding synchronization partition to "0".

[0079] Specifically, the drbd_work thread of the master node traverses the flag content in the dynamic bitmap, reads the corresponding partition data content with the flag set to "1", and sends the read partition data to the DRBD block device of the slave node in sequence, so that the data of the DRBD block devices of the master and slave nodes are completely consistent.

[0080] Step 1033: After the initial synchronization is completed, the dynamic bitmap of the DRBD block device of the master node is traversed again, and the partition data corresponding to the flag position "1" in the dynamic bitmap is synchronized to the slave node.

[0081] It can be understood that this embodiment takes into account that during the initialization of full synchronization, the DRBD block device of the master node can still receive business I / O, that is, when business data is written, the corresponding flag bit of the modification will still be recorded in the dynamic bitmap metadata area. Therefore, after the initialization synchronization is completed, this embodiment traverses the dynamic bitmap of the DRBD block device of the master node again, and synchronizes the corresponding partition data with the flag position "1" in the dynamic bitmap to the slave node, ensuring that all modifications that occur during the initialization synchronization are captured and transmitted, thereby ensuring that the master and slave data are accurately consistent, and avoiding data deviation caused by traditional solutions due to ignoring writes during synchronization.

[0082] Next, the existing brain split detection method is used to monitor the link connection status between the DRBD master and slave nodes. For example, the heartbeat mechanism is used to monitor whether the connection between the master and slave nodes is normal. When it is detected that the heartbeat interruption time between the master and slave nodes exceeds the set threshold, the peer node is considered to be faulty and the faulty node is automatically isolated to ensure data consistency. This embodiment does not go into details about this.

[0083] Step 104 : When the split brain is recovered, the DRBD block devices of the master and slave nodes exchange their respective dynamic bitmaps, and merge the dynamic bitmaps of the DRBD block devices of the master and slave nodes to obtain a modified partition set.

[0084] Specifically, in this embodiment, in step 104, the dynamic bitmaps of the DRBD block devices of the master node and the slave node are merged to obtain a modified partition set, including:

[0085] Perform a logical OR operation on the dynamic bitmaps of the DRBD block devices of the master node and the slave node to generate a dynamic bitmap that merges the modification status of both nodes;

[0086] The set of partitions with flag position "1" in the dynamic bitmap of the modification status of both parties is fused to form a modified partition set.

[0087] For example, a logical "OR" operation is performed on the dynamic bitmap 1010 of the DRBD block device of the master node and the dynamic bitmap 1001 of the DRBD block device of the slave node to obtain 1011, ie, 3 partitions of the 4 partitions that need to be synchronously set to "1".

[0088] Step 105 , the master node filters the unmodified partition data, and synchronizes the modified partition data to the slave nodes based on the modified partition set only, and resets the dynamic bitmap of the DRBD block device of each node after synchronization is completed.

[0089] The reset update is implemented as follows: after the master node synchronizes the modified partition data to the slave node, the master node and the slave node simultaneously set the flag position of the corresponding modified partition to "0".

[0090] It can be understood that in this embodiment, when recovering from a split-brain, that is, when re-establishing the connection between nodes, the master and slave nodes exchange dynamic bitmaps, merge the modification states of both parties through a logical "OR" operation, and quickly locate the difference partitions without comparing data block by block, reflecting the advantage of fast recovery decision-making, and avoiding the adverse effects of the prior art that only relies on one-way synchronization of source data, ensuring that the data modified by both parties is overwritten. Combined with partition management, only the partition data with differences between the two parties (that is, the data in the DRBD block device partition where the data is actually updated during the split-brain) is synchronized, and the consistent area (the area with the flag bit set to "0") is directly skipped, which fundamentally reduces the amount of synchronized data, significantly shortens the data synchronization time, and improves the data synchronization efficiency. It is suitable for large-scale storage scenarios with strict requirements on data consistency and high availability.

[0091] To sum up, different from the static range optimization of the existing technology, this embodiment builds an efficient and reliable partition-level modification tracking system through partition management and dynamic bitmap marking. Through state decoupling in the time dimension, the long-lasting full synchronization is decomposed into the superimposed transmission of "baseline data initialization synchronization + incremental dirty data". The design of collaborative state fusion and differential synchronization breaks through the contradiction of "full and incremental cannot be achieved at the same time" in traditional synchronization solutions. Under the premise of ensuring strong consistency, it achieves reduced business latency and improved network resource utilization, and realizes dynamic differential data filtering, which avoids the transmission of invalid data (that is, unmodified redundant data) from the root, and realizes efficient, flexible and highly available distributed storage management.

[0092] Example 2:

[0093] This embodiment only compares Figure 1 The differences between the first embodiment and the second embodiment are described below. The technical concepts of the remaining designs are similar to those of the first embodiment and will not be described in detail in this embodiment.

[0094] In this embodiment, before step 104, the following steps are further included:

[0095] Step 201: construct generation identifiers for tracking data versions and dynamic bitmap versions for the DRBD block devices of the master and slave nodes.

[0096] It can be understood that the generation identifier is used to identify the inter-generation relationship of data, so as to accurately determine the data consistency between nodes.

[0097] Step 202: When the connection between the split-brain recovery nodes is re-established, the nodes at both ends exchange their generation identification information;

[0098] In step 203, the generation identification information of the local node is compared with the generation identification information of the opposite node, and whether the DRBD block device of the double-end node has split brain is judged according to the comparison result. If it is judged that the split brain does not occur, the data synchronization strategy is determined according to the comparison result, and if it is judged that the split brain occurs, step 104 is executed.

[0099] Specifically, in the embodiment, the generation identification includes a current data version identification (denoted as current uuid, and the update time is when the master node is activated for the first time, and a new current uuid is generated at the same time after synchronization is completed) and a dynamic bitmap version identification (denoted as bitmap uuid, and the bitmap is updated each time the bitmap is changed due to writing, and the bitmap uuid of the master and the slave remains consistent after synchronization is completed). In step 203, the generation identification information of the local node is compared with the generation identification information of the opposite node, and whether the DRBD block device of the double-end node has split brain is judged according to the comparison result. If it is judged that the split brain does not occur, the data synchronization strategy is determined according to the comparison result, including:

[0100] The current data version identification of the local node is compared with the current data version identification of the opposite node. If the current data version identifications of the two nodes are consistent, the two parties belong to the same data generation, and it is judged that the split brain does not occur. If the current data version identifications of the two nodes are inconsistent, it is judged that the split brain occurs.

[0101] When it is judged that the split brain does not occur, the dynamic bitmap version identification of the local node is compared with the dynamic bitmap version identification of the opposite node. If the dynamic bitmap version identifications of the two nodes are consistent, data synchronization is not needed. If the dynamic bitmap version identifications of the two nodes are inconsistent, the node with the newer dynamic bitmap version identification is taken as the source end, and the other node is taken as the destination end. The corresponding modified partition data (that is, the partition data corresponding to the flag position “1” in the dynamic bitmap) in the dynamic bitmap of the source end is synchronized to the destination end, and the dynamic bitmap of the source end is reset after synchronization is completed (that is, the source end and the destination end simultaneously set the flag position corresponding to the modified partition to “0”).

[0102] Among them, the source end and the destination end usually refer to the direction of data synchronization, and the roles change with the switching of the master and the slave. The master node is the source end (the data sender), and the slave node is the destination end (the data receiver). When the split brain is recovered, the master and the slave roles may be switched.

[0103] It can be understood that, in the embodiment, whether the split brain occurs is quickly judged by the current data version identification. When it is judged that the split brain does not occur, the different partition is quickly located by the dynamic bitmap version identification, and only the incremental difference data is transmitted, so that the synchronization cost is minimized.

[0104] Further, the generation identifier can also include the last two historical data version identifiers (denoted as historical1 / historical2 uuid), and when a brain split is determined, the following two cases are included to determine the master and slave nodes:

[0105] 1. The current uuid is different, but there is a historical inheritance relationship, and the new and old of the data version is determined. Specifically, if the current uuid of node A is equal to the historical1 uuid or historical2 uuid of node B, it means that the data of node A is older, and therefore node B is the master node and node A is the slave node.

[0106] 2. The current uuid is different, and there is no historical inheritance relationship, that is, the current uuid of both nodes is not in the historical1 / historical2 uuid of the other party, which means that both sides have independently made irreconcilable data changes during the split (complete brain split), so human intervention is needed to specify the master node.

[0107] It can be understood that the exchange of generation identifiers is the cornerstone of DRBD data consistency management, and the modification tracking mechanism based on dynamic bitmap is an optimization layer based on this, and the two form a layered verification mechanism:

[0108] 1. The generation identifier handles the macro-level data version control and strategy decision, specifically, the version control idea of the distributed system is introduced into the storage layer in this embodiment, through the global uniqueness of the generation identifier, the complex brain split recovery problem is converted into a traceable metadata comparison problem, the precise synchronization direction decision is realized, and the full or differential synchronization is quickly judged, the synchronization granularity is self-adaptive, and through the historical data version identifier, the split time point and the influence range can be located, and the data conflict can be traced back;

[0109] 2. After the generation identifier establishes the synchronization strategy, the modification tracking mechanism based on dynamic bitmap handles the micro-level data partition optimization, which fundamentally reduces the invalid data transmission during full synchronization;

[0110] This layered collaborative design realizes the brain split recovery paradigm of traceable strategy, quantifiable execution, and locatable conflict, which can greatly improve the data synchronization efficiency of brain split recovery when facing TB-level devices, and further guarantees the data consistency and reliability in the brain split scenario.

[0111] Embodiment three:

[0112] This embodiment only describes the part that is different from the previous embodiment, and the technical concept of the remaining design is similar to the previous embodiment, which will not be described here.

[0113] The embodiment also includes the following steps:

[0114] Step 301, periodically counting the number of writes in each period for each partition in each DRBD block device;

[0115] Step 302, dynamically evaluating the access heat of each partition based on the number of writes in each period, to obtain the access heat value of each partition.

[0116] It can be understood that the number of writes in each period for each partition may change, and therefore the access heat value of each partition obtained by evaluation is dynamically adjusted.

[0117] On the one hand, the access heat value in the embodiment can be directly determined according to the direct proportion of the counting frequency of the number of writes, and on the other hand, in the embodiment, the access heat of each partition is dynamically evaluated based on the number of writes in each period, to obtain the access heat value of each partition, including:

[0118] Step 3021, obtaining the timestamp of each write of each partition;

[0119] Step 3022, calculating the timeliness decay factor of each partition based on the timestamp of the last write in the current period, which can be expressed as:

[0120]

[0121] Wherein, R represents the timeliness decay factor of the partition, T current represents the current system timestamp, T last represents the timestamp of the last write operation of the partition, and w represents the total period length, and the constraint condition is: if T current -T last > w, then R = 0.

[0122] Step 3023, weighted sum of the total number of writes in the current period and the timeliness decay factor of each partition, to obtain the access heat value of each partition in the current period.

[0123] It can be understood that the embodiment designs a composite index considering the number of writes and the interval time of writes by deeply deconstructing the space-time characteristics of write behavior, wherein the design of the timeliness decay factor formula makes the recent write partition obtain higher weight, avoids the interference of historical writes on real-time decision, and realizes the qualitative change breakthrough from "quantity statistics" to "value perception" at the calculation logic level of the access heat value, thereby accurately identifying high-value data partitions and providing an engineering feasible technical path for the priority synchronization of high-value data.

[0124] In step 105, the master node filters the unmodified partition data and synchronizes the modified partition data to the slave node based on the modified partition set, including:

[0125] Step 1051: Based on the current access heat value of each partition, determine the synchronization priority of each modified partition in the modified partition set. The synchronization priority is proportional to the access heat value. That is, the larger the access heat value, the higher the synchronization priority.

[0126] In step 1052, the master node synchronizes the corresponding modified partition data to the slave node in descending order of synchronization priority of each modified partition.

[0127] This embodiment introduces a priority synchronization mechanism to compensate for the technical defects of the traditional DRBD "blind synchronization" solution. First, the modification popularity of each partition is counted. When synchronizing data, the synchronization priority is sorted according to the access popularity of the partition, and the partitions with high frequency modifications are synchronized first to ensure that critical data is restored first, shorten the recovery time of the business critical path, and significantly reduce the business interruption time. It can further optimize the overall synchronization performance based on the parameter optimization of existing technologies (such as Chinese patent CN115277606A).

[0128] To summarize, this embodiment combines the three seemingly independent technical units of "data status recording (partition management)", "synchronization strategy decision (generation identification)" and "business priority scheduling" into an end-to-end efficient synchronization system through the through-coordination of dynamic bitmaps and generation identification. It takes partition management as the cornerstone and adopts dynamic bitmaps to provide dynamic recording capabilities of data modification status, providing a basis for subsequent optimization. It solves the problem of "what data to synchronize" through state fusion and generation identification decision-making. It reduces the amount of synchronized data from the root through dynamic filtering, significantly shortens the data synchronization time, and improves the data synchronization efficiency. The collaborative priority synchronization mechanism supplements the business logic, solves the problem of "how to synchronize efficiently", further shortens the recovery time of the critical path, and achieves the technical effect of "1+1>2". It provides an efficient and reliable DRBD block device data synchronization solution, which is particularly suitable for scenarios such as data centers that have strict requirements on data consistency and high availability.

[0129] Please see the attached Figure 2 , Figure 2 A schematic diagram of the data synchronization optimization system structure of a DRBD block device provided in one embodiment of this specification.

[0130] like Figure 2 As shown, the data synchronization optimization system of the DRBD block device may include at least a dynamic bitmap construction module 1, an initialization synchronization module 2, a split-brain recovery module 3, and a data synchronization module 4, wherein:

[0131] Dynamic bitmap construction module 1 is used to divide the DRBD block devices on the master node and the slave node into several partitions, assign a flag bit to each partition, and use the flag bit to mark the modification status of the partition data. Based on the flag bits of different partitions, a dynamic bitmap of each DRBD block device is obtained;

[0132] Initialization synchronization module 2 is used to initialize and synchronize the DRBD block devices of the master and slave nodes;

[0133] Split-brain recovery module 3 is used to, when split-brain recovery occurs, exchange the dynamic bitmaps of the DRBD block devices of the master and slave nodes, merge the dynamic bitmaps of the DRBD block devices of the master and slave nodes, and obtain a modified partition set;

[0134] The data synchronization module 4 is used for the master node to filter the unmodified partition data, synchronize the modified partition data to the slave node based on the modified partition set only, and reset the dynamic bitmap of the DRBD block device of each node after the synchronization is completed.

[0135] It is understandable that the technical concept of the data synchronization optimization system for a DRBD block device provided in this embodiment is similar to the technical concept of the data synchronization optimization method for a DRBD block device described above, and this embodiment will not be repeated here.

[0136] Another embodiment of the present disclosure provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of the aforementioned embodiments. If the components of the aforementioned electronic device are implemented as software functional units and used as independent downstream task predictions or tasks, they can be stored in the computer-readable storage medium.

[0137] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of this specification is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. Available media may be magnetic media (eg, floppy disks, hard disks, magnetic tapes), optical media (eg, digital versatile discs (DVDs)), or semiconductor media (eg, solid state disks (SSDs)).

[0138] The above description is merely an illustration of the preferred embodiments disclosed in this application and the technical principles employed. Those skilled in the art should understand that the scope of protection provided by this disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0139] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

Claims

1. A data synchronization optimization method for a DRBD block device, characterized in that: The following steps are involved: Step 101: Divide the DRBD block devices on the master node and the slave node into several partitions; Step 102: assign a flag bit to each partition, where the flag bit is used to mark the modification status of the partition data, and obtain a dynamic bitmap of each DRBD block device based on the flag bits of different partitions; Step 103: Initialize and synchronize the DRBD block devices of the master and slave nodes. Step 104: When the split-brain is recovered, the DRBD block devices of the master and slave nodes exchange their respective dynamic bitmaps, merge the dynamic bitmaps of the DRBD block devices of the master and slave nodes, and obtain a modified partition set; Step 105: The master node filters the unmodified partition data, and synchronizes the modified partition data to the slave nodes based on the modified partition set only, and resets the dynamic bitmap of each node after synchronization is completed.

2. A data synchronization optimization method for a DRBD block device according to claim 1, characterized in that: In step 102, a flag bit is allocated to each partition, where the flag bit is used to mark the modification status of the partition data. A dynamic bitmap of each DRBD block device is obtained based on the flag bits of different partitions, including: Each partition is assigned a binary flag. The flag is set to "0" by default. When the data in the partition is modified, the flag corresponding to the partition is set to "1". The DRBD block device of each node forms a dynamic bitmap based on the flag bit of each partition, and the index of each flag bit in the dynamic bitmap is mapped one-to-one with the number of each partition.

3. The data synchronization optimization method of a DRBD block device according to claim 2, characterized in that: In step 103, the DRBD block device is initialized and synchronized, including: Set all flags in the dynamic bitmap of the primary node's DRBD block device to "1" to indicate that all data needs to be synchronized; The master node reads the data of the partition corresponding to the flag position "1" and synchronizes the read partition data to the slave node in sequence to make the DRBD block device data of the master and slave nodes completely consistent. After each synchronization, the master and slave nodes simultaneously set the flag position of the corresponding synchronization partition to "0"; After the initial synchronization is completed, the dynamic bitmap of the DRBD block device of the master node is traversed again, and the corresponding partition data with the flag position "1" in the dynamic bitmap is synchronized to the slave node.

4. The data synchronization optimization method of a DRBD block device according to claim 1, characterized in that: Before step 104, the method further includes: Step 201: Constructing generation identifiers for tracking data versions and dynamic bitmap versions for the DRBD block devices of the master and slave nodes; Step 202: When the connection between the split-brain recovery nodes is re-established, the nodes at both ends exchange their generation identification information; In step 203, the generation identification information of the local node is compared with the generation identification information of the opposite node, and whether a split-brain occurs in the DRBD block device of the dual-end node is determined based on the comparison result. If it is determined that a split-brain does not occur, the data synchronization strategy is determined based on the comparison result. If it is determined that a split-brain occurs, step 104 is executed.

5. The data synchronization optimization method of a DRBD block device according to claim 4, characterized in that: The generation identifier includes the current data version identifier and the dynamic bitmap version identifier. In step 203, the generation identifier information of the local node is compared with the generation identifier information of the opposite node. Based on the comparison result, it is determined whether a split-brain occurs in the DRBD block device of the dual-end node. If it is determined that a split-brain does not occur, a data synchronization strategy is determined based on the comparison result, including: Compare the current data version identifier of the local node with the current data version identifier of the peer node. If the current data version identifiers of the two nodes are consistent, it is determined that no split-brain has occurred. If the current data version identifiers of the two nodes are inconsistent, it is determined that a split-brain has occurred. When it is determined that no brain split has occurred, the dynamic bitmap version identifier of the local node is compared with the dynamic bitmap version identifier of the opposite node. If the dynamic bitmap version identifiers of the two nodes are consistent, there is no need to synchronize data. If the dynamic bitmap version identifiers of the two nodes are inconsistent, the node with the newer dynamic bitmap version identifier is used as the source end and the other end node is used as the destination end. The corresponding modified partition data in the dynamic bitmap of the source end is synchronized to the destination end, and the dynamic bitmap of the source end is reset after the synchronization is completed.

6. The data synchronization optimization method of a DRBD block device according to claim 2, characterized in that: In step 104, the dynamic bitmaps of the DRBD block devices of the master node and the slave node are merged to obtain a modified partition set, including: Perform a logical OR operation on the dynamic bitmaps of the DRBD block devices of the master node and the slave node to generate a dynamic bitmap that merges the modification status of both parties; The set of partitions with flag position "1" in the dynamic bitmap of the modification status of both parties is combined to form the modified partition set.

7. The data synchronization optimization method of a DRBD block device according to claim 1, characterized in that: The following steps are also included: Periodically count the number of writes to each partition in each DRBD block device in each cycle; Based on the number of writes to each partition in each cycle, the access popularity of each partition is dynamically evaluated to obtain the access popularity value of each partition; In step 105, the master node filters the unmodified partition data and synchronizes the modified partition data to the slave node based on the modified partition set, including: Determine the synchronization priority of each modified partition in the modified partition set based on the current access heat value of each partition, wherein the synchronization priority is proportional to the access heat value; The master node synchronizes the corresponding modified partition data to the slave node in descending order according to the synchronization priority of each modified partition.

8. The data synchronization optimization method of a DRBD block device according to claim 7, characterized in that: Based on the number of writes to each partition in each cycle, the access popularity of each partition is dynamically evaluated to obtain the access popularity value of each partition, including: Get the timestamp of each write to each partition; The timeliness attenuation factor of each partition is calculated based on the timestamp of the last write in the current cycle of each partition; The total number of writes in the current cycle of each partition and the timeliness attenuation factor are weighted and summed to obtain the access heat value of each partition in the current cycle.

9. A data synchronization optimization system for DRBD block devices, characterized in that: include: A dynamic bitmap construction module is used to divide the DRBD block devices on the master node and the slave node into several partitions, assign a flag bit to each partition, and use the flag bit to mark the modification status of the partition data. Based on the flag bits of different partitions, a dynamic bitmap of each DRBD block device is obtained; Initialization synchronization module, used to initialize and synchronize the DRBD block devices of the master and slave nodes; The split-brain recovery module is used to exchange the dynamic bitmaps of the DRBD block devices of the master and slave nodes during split-brain recovery, merge the dynamic bitmaps of the DRBD block devices of the master and slave nodes, and obtain the modified partition set; The data synchronization module is used by the master node to filter the unmodified partition data, synchronize the modified partition data to the slave node based on the modified partition set only, and reset the dynamic bitmap of each node after the synchronization is completed.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for optimizing data synchronization of a DRBD block device according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • DRBD block device initialization method and device, computer device and storage medium

    CN114385573A

  • Method for optimizing DRBD data synchronization

    CN115277606A