A high-reliability continuous data protection method and system based on local disk
By configuring three disk architectures and the kernel module synchronization mechanism, the reliability problem of backup storage media in the CDP solution is solved, efficient and reliable data recovery and security guarantee are achieved, and are suitable for core business scenarios.
Patent Information
- Application Number
- CN202510811461.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The reliability of the backup storage medium in the existing CDP solution is insufficient, resulting in a break in the backup chain in a single point of failure, and data recovery at any point in time cannot be achieved. Random read and write operations affect disk I/O efficiency, limiting data security guarantee capabilities.
Three physical disk architectures are adopted, the main disk processes read and write requests, and the two backup disks record data in parallel. They are synchronized in real time through the kernel module to divide the continuous and disconnected data areas, record additional information to ensure data consistency and recovery capabilities, and combine time stamps and disconnected marks to achieve data recovery.
It realizes data recovery capabilities at any point in time, improves data security guarantees, reduces the impact of single point failures, and optimizes data recovery efficiency and system performance.
Smart Images

Figure CN120353648B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data storage technology, and in particular relates to a high-reliability continuous data protection method and system based on a local disk. Background Art
[0002] Current CDP solutions suffer from a core flaw: insufficient reliability of backup storage media. When the primary storage media is functioning normally but the backup storage media is damaged, the system can still respond to real-time read and write requests, but data recovery at any point in time is impossible due to a broken backup chain. This single point of failure seriously restricts data security in core business scenarios. Although existing solutions theoretically achieve any-time recovery capabilities by recording write operation data sequences in real time on the backup storage media, their architectural design using a single backup storage media poses the risk of irreversible loss of backup data. Furthermore, disk anomalies (such as partial sector damage or sudden power outages) can easily cause discontinuous backup records, leading to interrupted recovery operations. Furthermore, the traditional solution's random read and write operation mode for the backup storage media significantly impacts disk I / O efficiency, further limiting system performance. These deficiencies collectively constitute a key bottleneck in the practical deployment of existing CDP technology. Summary of the Invention
[0003] (1) Purpose of the invention
[0004] In order to overcome the above shortcomings, the purpose of the present invention is to provide a highly reliable continuous data protection method and system based on local disks to solve the hidden dangers of irreversible loss due to single point failure of a single spare disk and the technical problems of low integrity and efficiency of data recovery.
[0005] (2) Technical solution
[0006] To achieve the above objectives, the present application provides the following technical solutions:
[0007] A highly reliable continuous data protection method and system based on a local disk, characterized by comprising the following steps:
[0008] (1) Configure three physical disks, including one primary disk and two backup disks; the primary disk for write operation data is used to process normal read and write requests of the operating system, and the two backup disks for write operation data are used to record write operation data to achieve continuous data protection;
[0009] (2) The kernel module simulates the logical block device at the operating system layer, and the write operation data kernel module manages the read and write operations of the primary disk and synchronously controls the write operation data of the two backup disks;
[0010] (3) Run the data synchronization kernel thread in the write operation data kernel module to detect and synchronize the data differences between the two backup disks in real time to ensure that the write operation data of the two backup disks are complete and consistent;
[0011] (4) Divide each spare disk into a continuous data area and a disconnected data area; wherein the continuous data area for write operation data records normal write operation data in sequence, and the disconnected data area for write operation data records write operation data that cannot be written into the continuous data area due to anomalies;
[0012] (5) For each write operation data from the operating system, the write operation data on both backup disks contains the following fields:
[0013] A disconnect flag, used to indicate whether the next write operation data is located in the disconnect data area;
[0014] Offset address, indicating the relative position of the next write operation data in the disconnect data area;
[0015] The number of disconnect write requests indicates the number of subsequent write operation data in the disconnect data area;
[0016] Timestamp, used to uniquely identify write operation data and determine the recovery order;
[0017] Write request data length and corresponding write request data;
[0018] (6) When the primary disk performs data recovery, based on the timestamp sequence and disconnection mark information, combined with the write operation data of the continuous data area and the disconnected data area, the data at the target time is completely restored from any available backup disk;
[0019] By configuring a primary disk and dual backup disks, a dual-backup mechanism for data redundancy is implemented. When the primary disk is operating normally, both backup disks simultaneously record write operations. Failure of either backup disk will not cause a break in the backup chain, effectively eliminating the single point of failure of a single backup storage medium in traditional CDP solutions. Furthermore, the synchronization mechanism of the dual backup disks ensures the feasibility of data recovery at any point in time, significantly improving data security in core business scenarios.
[0020] In some implementations, the specific operations of the data synchronization kernel thread for writing the operation data in step (3) are:
[0021] This embodiment ensures real-time data consistency between the two backup disks by periodically detecting and synchronizing data differences between them. If data is lost on one backup disk due to an anomaly, the system automatically extracts the missing write data from the other backup disk and restores it to the disconnected data area. This mechanism not only reduces the risk of data inconsistency caused by hardware failures or unexpected anomalies, but also enhances the backup system's fault tolerance, ensuring that backup data integrity is maintained even in extreme situations.
[0022] In some implementations, performing a data recovery operation on a write operation data master disk includes the following steps:
[0023] (a) Check whether the primary disk needs data recovery;
[0024] (b) If the primary disk needs to recover data, a piece of write operation data is extracted from the continuous data area of any available backup disk in timestamp order;
[0025] (c) determining whether the next write operation data is located in the disconnect data area according to the disconnect mark;
[0026] (d) when the disconnect flag is true, traverse the disconnect data area according to the offset address and the disconnect write request number to obtain a complete write operation data sequence;
[0027] (e) Replay the write operation data in timestamp order and completely restore the data at the target time to the primary disk.
[0028] This embodiment achieves accurate data recovery in complex scenarios through a recovery process based on timestamp sequence and dynamic determination of disconnection markers. During the recovery process, the system intelligently identifies the write operation data sequences between continuous and disconnected data areas and ensures the timing accuracy of data recovery through timestamp replay. This design effectively resolves the recovery interruption problem caused by discontinuous backup records in traditional solutions, while also improving the efficiency and reliability of recovery operations.
[0029] In some implementations, when extracting write operation data from the spare disk in step (b), a spare disk with higher data integrity is preferentially selected as the data source.
[0030] By prioritizing backup disks with higher integrity as data sources, this embodiment optimizes data recovery efficiency and quality. This strategy minimizes the impact of partial data corruption on the recovery process, ensuring that recovery operations are performed based on the most complete backup records, thereby reducing the probability of recovery failure and improving system robustness.
[0031] In some implementations, the detection period of the write operation data synchronization kernel thread is dynamically adjusted according to the write operation data load of the primary disk: when the write operation data frequency is higher than a preset threshold, the detection period is shortened to improve the real-time synchronization; when the write operation data frequency is lower than the preset threshold, the detection period is extended to reduce system resource usage.
[0032] By dynamically adjusting the data synchronization cycle, a balance is achieved between system resource utilization efficiency and data consistency. This allows for rapid response to data discrepancies in high-load scenarios, avoiding backup inconsistencies caused by synchronization delays. It also reduces resource consumption under low load, making it particularly suitable for real-time business systems sensitive to I / O performance, effectively optimizing overall system efficiency.
[0033] In some implementations, the triggering conditions for writing the operation data in step (a) to detect whether the primary disk needs to perform data recovery include:
[0034] It is detected that the number of bad sectors on the primary disk exceeds the preset threshold;
[0035] Receive the target recovery time point specified by the user;
[0036] The system log records abnormal power outages or incomplete write operations.
[0037] Through intelligent judgment based on multi-dimensional trigger conditions, this embodiment proactively identifies potential risks to primary disks (such as physical damage or logical errors) and supports user-defined recovery requirements. Combined with analysis of abnormal event logs, the system automatically initiates recovery operations, eliminating delays caused by manual intervention and enhancing data protection proactivity and recovery scenario coverage.
[0038] In some implementations, the following steps are further included:
[0039] (7) When both the primary disk and the two backup disks are unavailable, the snapshot data of the most recent full backup is pulled from the remote cloud storage and completed to the latest state based on the incremental logs after the snapshot time point.
[0040] By integrating a cloud storage disaster recovery linkage mechanism, a three-tiered local-cloud protection system has been established. Even in extreme disaster scenarios (such as a complete loss of local storage), business data can be quickly rebuilt through the coordinated recovery of cloud snapshots and incremental logs. This design overcomes the geographical limitations of traditional local CDP solutions and provides a viable technical path for cross-regional disaster recovery.
[0041] On the other hand, the present application provides a highly reliable continuous data protection system based on a local disk, characterized by comprising the following modules:
[0042] (1) Disk configuration module, used to configure three physical disks, including one primary disk and two backup disks; the primary disk for write operation data is used to process normal read and write requests of the operating system, and the two backup disks for write operation data are used to record write operation data in parallel to achieve continuous data protection;
[0043] (2) Data synchronization module, including a logic block device simulation unit embedded in the operating system kernel layer, which is used to manage the read and write operations of the primary disk and control the synchronization of write operation data of the two backup disks;
[0044] (3) an area division module, used to divide each spare disk into a continuous data area and a disconnected data area; wherein the continuous data area for write operation data records normal write operation data in sequence, and the disconnected data area for write operation data records write operation data that cannot be written into the continuous data area due to anomalies;
[0045] (4) Write operation data module, used for each write operation data from the operating system. The write operation data in the two standby disks contains the following fields:
[0046] A disconnect flag, used to indicate whether the next write operation data is located in the disconnect data area;
[0047] Offset address, indicating the relative position of the next write operation data in the disconnect data area;
[0048] The number of disconnect write requests indicates the number of subsequent write operation data in the disconnect data area;
[0049] Timestamp, used to uniquely identify write operation data and determine the recovery order;
[0050] Write request data length and corresponding write request data;
[0051] (5) A data recovery module is used to completely recover the data at the target time from any available backup disk based on the timestamp sequence and disconnection mark information, combined with the write operation data of the continuous data area and the disconnection data area when the primary disk performs data recovery.
[0052] In some implementations, the write operation data synchronization module also includes: a data synchronization kernel thread unit, which periodically detects the differences between the two spare disks. If data is found to be missing on a spare disk, the missing write operation data is extracted from the other spare disk and synchronized to the disconnected data area of the spare disk.
[0053] In some implementations, the specific recovery process of the data recovery module includes:
[0054] (a) a trigger detection unit for detecting whether data recovery is required on the primary disk;
[0055] (b) a write operation data extraction unit, extracting write operation data from a continuous data area of any available spare disk based on a timestamp sequence;
[0056] (c) a disconnection determination unit, which determines whether the next write operation data is located in the disconnection data area according to the disconnection mark;
[0057] (d) a disconnect traversal unit, which, when the disconnect flag is true, traverses the disconnect data area according to the offset address and the disconnect write request number to obtain a complete write operation data sequence;
[0058] (e) Data replay unit, which replays the write operation data in the order of timestamps and completely restores the data at the target time to the main disk. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a framework diagram of the local disk-based high-reliability continuous data protection system of the present invention;
[0060] Figure 2 This is the flowchart of the kernel module processing write operations;
[0061] Figure 3 This is a flowchart of the kernel module's data synchronization kernel thread processing the loss of read and write data on a backup disk;
[0062] Figure 4 This is the operation flow chart when the master disk performs recovery;
[0063] Figure 5 This is a diagram of the internal framework of the spare disk. DETAILED DESCRIPTION
[0064] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present invention.
[0065] One aspect of the present invention provides a highly reliable continuous data protection method based on a local disk, comprising the following steps:
[0066] (1) Configure three physical disks, including one primary disk and two backup disks; the primary disk for write operation data is used to process normal read and write requests of the operating system, and the two backup disks for write operation data are used to record write operation data to achieve continuous data protection;
[0067] (2) The kernel module simulates the logical block device at the operating system layer, and the write operation data kernel module manages the read and write operations of the primary disk and synchronously controls the write operation data of the two backup disks;
[0068] (3) Run the data synchronization kernel thread in the write operation data kernel module to detect and synchronize the data differences between the two backup disks in real time to ensure that the write operation data of the two backup disks are complete and consistent;
[0069] (4) Divide each spare disk into a continuous data area and a disconnected data area; wherein the continuous data area for write operation data records normal write operation data in sequence, and the disconnected data area for write operation data records write operation data that cannot be written into the continuous data area due to anomalies;
[0070] (5) For each write operation data from the operating system, the write operation data in both standby disks contains the following fields:
[0071] A disconnect flag, used to indicate whether the next write operation data is located in the disconnect data area;
[0072] Offset address, indicating the relative position of the next write operation data in the disconnect data area;
[0073] The number of disconnect write requests indicates the number of subsequent write operation data in the disconnect data area;
[0074] Timestamp, used to uniquely identify write operation data and determine the recovery order;
[0075] Write request data length and corresponding write request data;
[0076] (6) When the primary disk performs data recovery, based on the timestamp sequence and disconnection mark information, combined with the write operation data of the continuous data area and the disconnected data area, the data at the target time is completely restored from any available spare disk.
[0077] The core basis of this application is to configure three physical disks: one primary disk and two spare disks. The primary disk directly handles the read and write requests of the operating system, while the two spare disks record all write operation data in parallel. In specific implementation, the primary disk and the spare disk need to be connected through independent SATA or NVMe interfaces to ensure that the data transmission channels do not interfere with each other. During the system initialization phase, the spare disk is divided into a continuous data area and a disconnected data area through a formatting tool: the continuous data area adopts a sequential write mode to store the write operation data of the normal process; the disconnected data area adopts a dynamic allocation strategy to specifically capture unfinished write operations caused by sudden anomalies (such as power outages and disk bad sectors). This partition design not only optimizes storage efficiency, but also provides a structured log basis for subsequent data recovery.
[0078] Specifically, the data synchronization kernel thread is a key component for ensuring the consistency of dual spare disks. This thread is embedded in the operating system kernel layer and scans the write operation data of the two spare disks at a millisecond period (the default setting is 100ms). When it is detected that there is missing write operation data in the continuous data area of a spare disk, the system will immediately extract the missing data from the corresponding position of the other spare disk and append it to the disconnected data area of the current spare disk. For example, if spare disk A loses write operation data No. 100-150 due to a transient power supply fluctuation, the kernel thread will copy these write operation data from the continuous data area of spare disk B and insert them into the disconnected data area of spare disk A in the order of the original timestamp. This process is fully automated and transparent to upper-layer applications, so that the integrity of the backup chain can be maintained in the event of hardware failure or sudden abnormality.
[0079] The following describes the timestamp-based intelligent data recovery process:
[0080] When data recovery is needed on the primary disk, the system first initiates the recovery process based on trigger conditions (such as excessive bad sectors, a user-specified time point, or an exception log). The core of the recovery operation is timestamp-driven replay of write operation data:
[0081] Step 1: Extract the earliest unrecovered write operation data from the contiguous data area of any available spare disk in timestamp order. For example, if the target recovery time is T, the latest complete write operation data before T is preferentially extracted.
[0082] Step 2: Determine whether the subsequent operation will jump to the disconnect data area based on the "disconnect flag" of the current write operation data. If the flag is "true", the associated write operation data is batch read from the disconnect data area based on the offset address and the number of disconnect write requests.
[0083] Step 3: The write operation data from the continuous data area and the disconnected data area are merged according to the timestamp and replayed to the main disk one by one. During this process, the system strictly verifies the length and content of the write request data to ensure that it is completely consistent with the original operation.
[0084] This process solves the recovery interruption problem caused by backup failure in traditional solutions through dynamic path selection (flexible jump between continuous areas and disconnected areas). It is especially suitable for scenarios where abnormalities frequently occur during long-term operation.
[0085] Combine Figure 5 The data distribution example of the spare disk is shown. The specific implementation of the continuous data area and the disconnected data area is further explained as follows:
[0086] On the backup disk, the continuous data area records normal write operations in timestamp order (e.g., 4999, 5000). If a system exception (such as a momentary power outage or a bad disk sector) prevents a write operation (e.g., timestamps 5001 and 5002) from being written to the continuous data area, the operation is redirected to the disconnected data area. In this case, the system records the write operation data offset (e.g., 0x00), the disconnect flag (set to "true"), the number of disconnected write requests (two), and the corresponding data length and content in write operation data number 5000 in the continuous data area.
[0087] Data synchronization of disconnected data area:
[0088] If write data is missing from the contiguous data area of a spare disk (such as spare disk A) at timestamps 5001 and 5002, the data synchronization kernel thread detects the discrepancy and immediately extracts the complete write data with timestamps 5001 and 5002 from the contiguous data area of the other spare disk (such as spare disk B). The system then appends this write data to the disconnected data area of spare disk A in the order of its original timestamps. This process ensures data consistency between the two spare disks. Even if a single spare disk experiences a local failure, it can be quickly repaired using the redundant records of the other spare disk.
[0089] Operations when performing data recovery on the primary disk:
[0090] When the primary disk needs to be restored to the target time point (for example, timestamp 5003), the system performs the following steps:
[0091] Extract write operation data from the continuous data area: Extract write operation data numbered 5000 from the continuous data area of any available spare disk (such as spare disk A) in timestamp order.
[0092] Determine the disconnect flag: Check whether the disconnect flag of the write operation data No. 5000 is "true". If it is "true", locate the starting position of the disconnect data area according to the offset address (0x00).
[0093] Traversing the disconnected data area: Based on the number of disconnected write requests (for example, 2, corresponding to timestamps 5001 and 5002), the complete data and timestamps of the two write operation data are read in batches from the disconnected data area.
[0094] Merge and replay: The write operation data numbered 5000 in the continuous data area is merged with the write operation data numbers 5001 and 5002 in the disconnected data area in timestamp order and replayed back to the primary disk one by one. During this process, the system strictly verifies the length, content, and timestamp of each write operation data to ensure that the restored data is completely consistent with the target time point.
[0095] Through the above process, the present invention not only solves the recovery interruption problem caused by backup chain breakage in traditional solutions, but also realizes efficient recovery in abnormal scenarios through a dynamic jump mechanism.
[0096] Preferably, the present application optimizes the synchronization period to obtain a dynamic optimization scheme for the adaptive synchronization period:
[0097] To balance system performance and data consistency, the detection period of the data synchronization kernel thread can be dynamically adjusted according to the write operation data load of the primary disk. The specific implementation method is as follows:
[0098] When the write operation frequency exceeds 1000 times per second (customizable through preset thresholds), the synchronization period is shortened to 50ms to quickly respond to data differences;
[0099] When the data write operation frequency is less than 200 times per second, the synchronization period is extended to 500ms to reduce the use of CPU and I / O resources.
[0100] This strategy is similar to adjusting the rhythm of traffic lights according to traffic flow. It can not only ensure real-time performance under high load, but also reduce system overhead under low load. It is especially suitable for resource-sensitive environments such as cloud computing.
[0101] Preferably, the present application also optimizes the recovery mechanism to obtain an active recovery mechanism with multi-dimensional trigger conditions:
[0102] The system supports three data recovery trigger conditions, forming a multi-layer protection:
[0103] Physical layer detection: Use the SMART tool to monitor the health status of the main disk in real time, and automatically trigger recovery when the number of bad sectors exceeds 0.1% of the total capacity;
[0104] User-level commands: Allow administrators to specify any historical time point through the command line or graphical interface to achieve on-demand recovery;
[0105] Abnormal event backtracking: Analyze abnormal events in system logs (such as incomplete write operation data IDs), automatically locate the associated time window, and initiate recovery.
[0106] This multi-condition trigger mechanism significantly reduces the frequency of manual intervention while covering a wide range of scenarios from hardware failures to logical errors.
[0107] In extreme cases (such as when both the primary disk and the backup disk are damaged at the same time), the system can switch to cloud-based disaster recovery mode:
[0108] Download the most recent full snapshot from a pre-configured remote cloud storage (e.g., the baseline backup generated every morning);
[0109] Based on the incremental logs after the snapshot time point (stored in an independent cloud database), replay the write operation data one by one to the latest state.
[0110] This solution breaks through the geographical limitations of traditional CDP solutions through a local-cloud hybrid architecture. Even in the event of physical disasters such as fire and floods, it can still achieve cross-regional data reconstruction, providing ultimate disaster recovery protection for key industries such as finance and healthcare.
[0111] This implementation utilizes a layered design (hardware architecture, kernel synchronization, and intelligent recovery) and multi-level optimization (dynamic periodization, interpolation algorithms, and cloud-based integration) to build a highly reliable, adaptive, and scalable continuous data protection system. These technical features work together to address the single point of failure inherent in traditional solutions while also providing innovative mechanisms to address complex anomaly scenarios, ultimately achieving data recoverability at any point in time and under any failure scenario.
[0112] Another aspect of the present invention provides a highly reliable continuous data protection system based on a local disk, comprising the following modules:
[0113] (1) Disk configuration module, used to configure three physical disks, including one primary disk and two backup disks; the primary disk for write operation data is used to process normal read and write requests of the operating system, and the two backup disks for write operation data are used to record write operation data in parallel to achieve continuous data protection;
[0114] (2) Data synchronization module, including a logic block device simulation unit embedded in the operating system kernel layer, which is used to manage the read and write operations of the primary disk and control the synchronization of write operation data of the two backup disks;
[0115] (3) an area division module, used to divide each spare disk into a continuous data area and a disconnected data area; wherein the continuous data area for write operation data records normal write operation data in sequence, and the disconnected data area for write operation data records write operation data that cannot be written into the continuous data area due to anomalies;
[0116] (4) Write operation data module, used for each write operation data from the operating system. The write operation data in the two standby disks contains the following fields:
[0117] A disconnect flag, used to indicate whether the next write operation data is located in the disconnect data area;
[0118] Offset address, indicating the relative position of the next write operation data in the disconnect data area;
[0119] The number of disconnect write requests indicates the number of subsequent write operation data in the disconnect data area;
[0120] Timestamp, used to uniquely identify write operation data and determine the recovery order;
[0121] Write request data length and corresponding write request data;
[0122] (5) A data recovery module is used to completely recover the write operation data at the target time from any available spare disk based on the timestamp sequence and disconnection mark information, combined with the write operation data of the continuous data area and the disconnection data area when the primary disk performs data recovery.
[0123] In some implementations, the data synchronization module also includes: a data synchronization kernel thread unit, which periodically detects the differences between the two spare disks. If data is found to be missing on a spare disk, the missing write operation data is extracted from the other spare disk and synchronized to the disconnected data area of the spare disk.
[0124] In some implementations, the specific recovery process of the data recovery module includes:
[0125] (a) a trigger detection unit for detecting whether data recovery is required on the primary disk;
[0126] (b) a write operation data extraction unit, extracting write operation data from a continuous data area of any available spare disk based on a timestamp sequence;
[0127] (c) a disconnection determination unit, which determines whether the next write operation data is located in the disconnection data area according to the disconnection mark;
[0128] (d) a disconnect traversal unit, which, when the disconnect flag is true, traverses the disconnect data area according to the offset address and the disconnect write request number to obtain a complete write operation data sequence;
[0129] (e) Data replay unit, which replays the write operation data in the order of timestamps and completely restores the data at the target time to the main disk.
[0130] It should be understood that the above-described specific embodiments of the present invention are merely illustrative or illustrative of the principles of the present invention and do not constitute limitations of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention should be included within the scope of protection of the present invention. In addition, the appended claims are intended to cover all variations and modifications that fall within the scope and metes and bounds of the appended claims, or equivalents thereof.
Claims
1. A highly reliable continuous data protection method based on a local disk, characterized in that: The following steps are involved: (1) Three physical disks are configured, including one primary disk and two backup disks; the primary disk is used to process normal read and write requests of the operating system, and the two backup disks are used to record write operation data to achieve continuous data protection; (2) Simulating a logical block device at the operating system layer through a kernel module, wherein the kernel module manages the read and write operations of the primary disk and synchronously controls the write operation data of the two backup disks; (3) running a data synchronization kernel thread in the kernel module to detect and synchronize data differences between the two backup disks in real time, ensuring that the write operation data of the two backup disks are complete and consistent; (4) Dividing each spare disk into a continuous data area and a disconnected data area; wherein the continuous data area records normal write operation data in sequence, and the disconnected data area records write operation data that cannot be written into the continuous data area due to an abnormality; (5) For each write operation data from the operating system, the write operation data in both standby disks contains the following fields: A disconnect flag, used to indicate whether the next write operation data is located in the disconnect data area; Offset address, indicating the relative position of the next write operation data in the disconnect data area; The number of disconnect write requests indicates the number of subsequent write operation data in the disconnect data area; Timestamp, used to uniquely identify write operation data and determine the recovery order; Write request data length and corresponding write request data; (6) When the primary disk performs data recovery, based on the timestamp sequence and disconnection mark information, combined with the write operation data of the continuous data area and the disconnected data area, the data at the target time is completely restored from any available spare disk.
2. The method according to claim 1, characterized in that The specific operations of the data synchronization kernel thread in step (3) are: The difference between the two spare disks is periodically detected. If data is missing on one spare disk, the missing write operation data is extracted from the other spare disk and synchronized to the disconnected data area of the spare disk.
3. The method according to claim 1 or 2, characterized in that The data recovery operation performed by the master disk includes the following steps: (a) Check whether the primary disk needs data recovery; (b) If the primary disk needs to recover data, a piece of write operation data is extracted from the continuous data area of any available backup disk in timestamp order; (c) determining whether the next write operation data is located in the disconnect data area according to the disconnect mark; (d) when the disconnect flag is true, traverse the disconnect data area according to the offset address and the disconnect write request number to obtain a complete write operation data sequence; (e) Replay the write operation data in timestamp order and completely restore the data at the target time to the primary disk.
4. The method according to claim 3, characterized in that When extracting write operation data from the backup disk in step (b), the backup disk with higher data integrity is preferentially selected as the data source.
5. The method according to claim 2, characterized in that The detection period of the data synchronization kernel thread is dynamically adjusted according to the write operation data load of the main disk: when the write operation data frequency is higher than the preset threshold, the detection period is shortened to improve the real-time synchronization; when the write operation data frequency is lower than the preset threshold, the detection period is extended to reduce system resource usage.
6. The method according to claim 3, characterized in that The triggering conditions for detecting whether the primary disk needs data recovery in step (a) include: It is detected that the number of bad sectors on the primary disk exceeds the preset threshold; Receive the target recovery time point specified by the user; The system log records abnormal power outages or incomplete write operations.
7. The method according to any one of claims 1-2, 4-6, characterized in that: The following steps are also included: (7) When both the primary disk and the two backup disks are unavailable, the snapshot data of the most recent full backup is pulled from the remote cloud storage and completed to the latest state based on the incremental logs after the snapshot time point.
8. A high-reliability continuous data protection system based on a local disk, characterized in that: Includes the following modules: (1) A disk configuration module configured to configure three physical disks, including one primary disk and two backup disks; the primary disk is used to process normal read and write requests of the operating system, and the two backup disks are used to record write operation data in parallel to achieve continuous data protection; (2) Data synchronization module, including a logic block device simulation unit embedded in the operating system kernel layer, which is used to manage the read and write operations of the primary disk and control the synchronization of write operation data of the two backup disks; (3) a region division module, configured to divide each spare disk into a continuous data region and a disconnected data region; wherein the continuous data region records normal write operation data in sequence, and the disconnected data region records write operation data that cannot be written into the continuous data region due to an abnormality; (4) Write operation data module, used for each write operation data from the operating system. The write operation data in the two spare disks contains the following fields: A disconnect flag, used to indicate whether the next write operation data is located in the disconnect data area; Offset address, indicating the relative position of the next write operation data in the disconnect data area; The number of disconnect write requests indicates the number of subsequent write operation data in the disconnect data area; Timestamp, used to uniquely identify write operation data and determine the recovery order; Write request data length and corresponding write request data; (5) A data recovery module is used to completely recover the data at the target time from any available backup disk based on the timestamp sequence and disconnection mark information, combined with the write operation data of the continuous data area and the disconnection data area when the primary disk performs data recovery.
9. The system according to claim 8, characterized in that The data synchronization module also includes: a data synchronization kernel thread unit, which periodically detects the differences between the two spare disks. If data is found to be missing on one spare disk, the missing write operation data is extracted from the other spare disk and synchronized to the disconnected data area of the spare disk.
10. The system according to claim 8, wherein: The specific recovery process of the data recovery module includes: (a) a trigger detection unit for detecting whether data recovery is required on the primary disk; (b) a write operation data extraction unit, extracting write operation data from a continuous data area of any available spare disk based on a timestamp sequence; (c) a disconnection determination unit, which determines whether the next write operation data is located in the disconnection data area according to the disconnection mark; (d) a disconnect traversal unit, which, when the disconnect flag is true, traverses the disconnect data area according to the offset address and the disconnect write request number to obtain a complete write operation data sequence; (e) Data replay unit, which replays the write operation data in the order of timestamps and completely restores the data at the target time to the main disk.
Citation Information
Patent Citations
Block-level disk data protection system and method thereof
CN103019890A
Hard disk data backup and restore method
CN1419196A