Efficient storage synchronization system and method for active-active data center

By partitioning and labeling, synchronous optimization, conflict detection and processing and status monitoring in the dual-active data center scenario, network bandwidth waste and data inconsistency are solved in traditional synchronization technology, and efficient and automatic data synchronization and recovery are achieved.

CN120295558APending Publication Date: 2025-07-11INFORMATION CENT OF YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510163972.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional data synchronization technology has network bandwidth waste and synchronization delay in the dual-active data center scenario, and lacks efficient conflict detection and processing mechanisms, resulting in data inconsistency.

Method used

The data partitioning and labeling module are used to divide the stored data into multiple logical partitions, track updates and optimize transmission through the data synchronization module, conflict detection and processing module are used to resolve conflicts, and data consistency is ensured through the synchronization status monitoring module.

Benefits of technology

It realizes refined data management, reduces the amount of synchronized data, improves synchronization efficiency, ensures data integrity and consistency, reduces manual intervention, and improves the system's self-healing ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295558A_ABST
    Figure CN120295558A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient storage synchronization system and method oriented to an active-active data center, and relates to the technical field of data synchronization. The efficient storage synchronization system comprises a data partitioning and labeling module, a data storage module and a data synchronization module, the data synchronization module is used for tracking all data updating in the logic partition and synchronously updating the data; the synchronous data transmission optimization module is used for preprocessing the data in the synchronization process and reducing the data volume needing to be transmitted; the conflict detection and processing module is used for detecting whether the synchronized data conflicts or not in the synchronization process, and if yes, processing the conflicts according to a preset priority processing strategy; and the synchronization state monitoring module is used for checking whether the synchronization state data of the two data centers are consistent or not after data synchronization is completed, and if the synchronization state data of the two data centers are inconsistent, triggering data recovery operation. By automatically triggering the data recovery operation, it is ensured that the data is rapidly recovered to the consistent state, and the self-healing capacity and reliability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data synchronization, and particularly to an efficient storage synchronization system and method for a dual-active data center. Background Art

[0002] With the increasing requirements of enterprises for business continuity and data availability, the dual-active data center architecture has gradually become the mainstream. A dual-active data center means that two data centers provide services to the outside world simultaneously and serve as backups for each other. When any one of the data centers fails, the other data center can immediately take over the business to ensure that the business is not interrupted.

[0003] Traditional data synchronization technologies such as log-based synchronization and timestamp-based synchronization have certain limitations in the dual-active data center scenario. Full-volume synchronization or coarse-grained synchronization will cause waste of network bandwidth and synchronization delay, and lack an efficient conflict detection and handling mechanism, which is likely to lead to data inconsistency. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed.

[0005] Therefore, the problem to be solved by the present invention is: how to solve the problems that traditional methods have certain limitations in the dual-active data center scenario, full-volume synchronization or coarse-grained synchronization will cause waste of network bandwidth and synchronization delay, and lack an efficient conflict detection and handling mechanism, which is likely to lead to data inconsistency.

[0006] To solve the above technical problems, the present invention provides the following technical solution: An efficient storage synchronization system for a dual-active data center, including a data partitioning and tagging module, a data synchronization module, a synchronization data transmission optimization module, a conflict detection and handling module, and a synchronization status monitoring module; the data partitioning and tagging module is used to divide the stored data into multiple logical partitions, and each logical partition is assigned a unique tag; the data synchronization module is used to track all data updates within the logical partition and synchronize the updated data; the synchronization data transmission optimization module preprocesses the data during the synchronization process to reduce the amount of data to be transmitted; the conflict detection and handling module detects whether the synchronized data conflicts during the synchronization process. If a conflict occurs, the conflict is handled according to a preset priority processing strategy; the synchronization status monitoring module checks whether the synchronization status data of the two data centers is consistent after the data synchronization is completed. If the synchronization status data of the two data centers is inconsistent, a data recovery operation is triggered.

[0007] As a preferred solution of the efficient storage synchronization system for a dual-active data center according to the present invention, the data partitioning and tagging module includes a data partitioning unit, a tag assignment unit, a partition mapping unit, and a partition optimization unit; the data partitioning unit is used to divide the stored data into multiple logical partitions according to a preset partitioning strategy; the tag assignment unit assigns a unique tag to each logical partition, and the tag includes a partition identifier, partition metadata, and identification information of the data center to which the partition belongs; the partition mapping unit is used to establish a mapping relationship between the logical partition and the physical storage location, generate a logical partition mapping table, and update the logical partition mapping table in real time; the partition optimization unit dynamically adjusts the partition size and distribution according to the data access pattern and synchronization requirements.

[0008] As a preferred solution of the efficient storage synchronization system for a dual-active data center according to the present invention, the data synchronization module includes a log capture unit, an incremental data extraction unit, a log compression unit, a log tagging unit, an incremental data cache unit, and a log cleaning unit; the log capture unit is used to monitor the update operations of the data partition in real time, record all data changes including insert, update, and delete operations through a log capture mechanism, and generate corresponding incremental logs; the incremental data extraction unit extracts the incremental data to be synchronized according to the incremental logs and filters out irrelevant operations; the log compression unit compresses the incremental logs; the log tagging unit assigns a unique tag to each incremental log; the incremental data cache unit temporarily caches the extracted incremental data, waits for preprocessing by the transmission optimization module, and supports breakpoint resumption and retransmission mechanisms; the log cleaning unit clears the processed incremental logs after the incremental data is successfully synchronized, releases the storage space, and archives them to historical logs regularly.

[0009] As a preferred solution of the efficient storage synchronization system for a dual-active data center according to the present invention, the synchronization data transmission optimization module includes a data compression unit, a data chunking unit, a deduplication unit, a metadata management unit, a transmission protocol optimization unit, and a preprocessing cache unit; the data compression unit is used to compress the synchronization data, adopting a lossless compression algorithm to reduce the data transmission volume; the data chunking unit divides the data to be synchronized into data blocks of a fixed size or a dynamic size; the deduplication unit identifies and deletes duplicate data blocks through a hash algorithm, only transmits unique data blocks, and reorganizes the data at the target data center; the metadata management unit generates metadata for each data block, including data block identification, fingerprint information, compression status, and location information; the transmission protocol optimization unit adopts the QUIC protocol, combines with the data compression unit and the deduplication unit to improve the transmission performance; the preprocessing cache unit is used to temporarily store the data blocks that have been compressed and deduplicated.

[0010] As a preferred solution of the efficient storage synchronization system for a dual-active data center according to the present invention, wherein: the conflict detection and handling module includes a version number management unit, a conflict detection unit, a conflict log recording unit, a conflict handling strategy unit, a conflict notification unit, and a conflict recovery unit; the version number management unit assigns a unique version number to each data partition or data block, and increments the version number when the data is updated to identify the latest state of the data; the conflict detection unit compares the version numbers of the source data center and the target data center during the data synchronization process, and determines a data conflict if the version numbers are inconsistent; the conflict log recording unit is used to record all detected data conflict information, including the identification, version number, and timestamp of the conflicting data; the conflict handling strategy unit resolves conflicts according to the preset conflict handling strategy; the conflict notification unit sends a conflict notification to the administrator or relevant business system when a conflict is detected and cannot be automatically processed, and provides conflict details and handling suggestions; the conflict recovery unit ensures that the conflicting data is consistent in the dual-active data center after the conflict is handled, and records the conflict handling result.

[0011] As a preferred solution of the efficient storage synchronization system for a dual-active data center according to the present invention, wherein: the synchronization status monitoring module includes a synchronization status collection unit, a consistency verification unit, a difference analysis unit, an automatic repair unit, a status log recording unit, an alarm notification unit, and a visualization unit; the synchronization status collection unit collects the status information of the data partitions from the source data center and the target data center respectively after the data synchronization is completed, including the data version number, timestamp, data fingerprint, and metadata; the consistency verification unit verifies whether the data is consistent by comparing the status information of the source data center and the target data center, and generates a difference report if inconsistencies are found; the difference analysis unit analyzes the difference report to determine the type of difference, and the types of difference include data loss, data conflict, or data corruption; the automatic repair unit is used to automatically trigger a data recovery operation when it is detected that the status data of the two data centers is inconsistent; the status log recording unit is used to record the status information, verification results, and repair operations of each synchronization, and generate a synchronization log; the alarm notification unit is used to send an alarm notification to the system administrator or relevant business system when data inconsistency or synchronization failure is detected; the visualization unit is used to provide a graphical interface to display the status data of the data synchronization and the verification status data results in real time.

[0012] As a preferred solution of the efficient storage synchronization system for a dual-active data center according to the present invention, wherein: the data recovery operation includes data resynchronization operation, data reconstruction operation, and backup recovery operation.

[0013] Another object of the present invention is to provide a method for an efficient storage synchronization system for a dual-active data center, which can solve the problem of efficient storage synchronization for a dual-active data center through the method of efficient storage synchronization for a dual-active data center.

[0014] To solve the above technical problems, the present invention provides the following technical solutions: An efficient storage synchronization method for a dual-active data center, including dividing storage data into multiple logical partitions according to a preset policy, and assigning a unique label to each partition, including a partition identifier, metadata, and data center identification information. At the same time, establish a mapping relationship between the partition and the physical storage location, and dynamically adjust the partition size and distribution to optimize storage and synchronization efficiency; monitor the update operations of the data partition in real time, record data changes through a log capture mechanism, generate incremental logs, extract the incremental data to be synchronized, and filter out irrelevant operations; compress the synchronized data, use a lossless compression algorithm to reduce the data transmission volume, divide the data into fixed or dynamic-sized data blocks, identify and delete duplicate data blocks through a hash algorithm, and only transmit unique data blocks, and generate metadata; during the synchronization process, compare the data version numbers of the source data center and the target data center. If the version numbers are inconsistent, it is determined as a data conflict, record the conflict information, automatically or semi-automatically resolve the conflict according to a preset policy, and ensure data consistency after the processing is completed; after the data synchronization is completed, collect the status information of the data partition from the source data center and the target data center, including the version number, timestamp, data fingerprint, and metadata, and check whether the data is consistent by comparing the status information. If it is inconsistent, generate a difference report and automatically trigger a data recovery operation to ensure data consistency between the two data centers.

[0015] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned efficient storage synchronization method for a dual-active data center.

[0016] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of the above-mentioned efficient storage synchronization method for a dual-active data center.

[0017] The beneficial effects of the present invention are as follows: An efficient storage synchronization method for a dual-active data center provided by the present invention realizes refined management of data, reduces the amount of data to be synchronized, and improves synchronization efficiency by dividing storage data into multiple logical partitions and assigning unique tags; reduces network bandwidth occupancy by the amount of transmitted data, further improving synchronization efficiency; ensures that all changes can be captured and synchronized by tracking all data updates within the logical partition, avoiding data omission. After synchronization is completed, it ensures data integrity and consistency by verifying whether the synchronization status data of the two data centers is consistent; detects and processes data conflicts during synchronization to avoid data inconsistency problems caused by conflicts; can quickly and automatically resolve conflicts through a preset priority processing strategy, reducing the need for manual intervention; when detecting inconsistent synchronization status, automatically triggers a data recovery operation to ensure that the data can be quickly restored to a consistent state, improving the self-healing ability and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 FIG. 9 is an architecture diagram of an efficient storage synchronization system for a dual-active data center provided by the first embodiment of the present invention.

[0020] Figure 2 FIG. 13 is a unit structure diagram of a data partitioning and tagging module of an efficient storage synchronization system for a dual-active data center provided by the first embodiment of the present invention.

[0021] Figure 3 FIG. 17 is a unit structure diagram of a data synchronization module of an efficient storage synchronization system for a dual-active data center provided by the first embodiment of the present invention.

[0022] Figure 4 FIG. 21 is a unit structure diagram of a synchronization data transmission optimization module of an efficient storage synchronization system for a dual-active data center provided by the first embodiment of the present invention.

[0023] Figure 5 FIG. 25 is a unit structure diagram of a conflict detection and handling module of an efficient storage synchronization system for a dual-active data center provided by the first embodiment of the present invention.

[0024] Figure 6 FIG. 29 is a unit structure diagram of a synchronization status monitoring module of an efficient storage synchronization system for a dual-active data center provided by the first embodiment of the present invention.

[0025] Figure 7 This is a flowchart of an efficient storage synchronization method for a dual-active data center provided by the second embodiment of the present invention. Detailed implementation manners

[0026] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification.

[0027] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0028] Example 1, referring to Figures 1 to 6 , which is the first embodiment of the present invention. This embodiment provides an efficient storage synchronization system for a dual-active data center, including: a data partitioning and tagging module 11, a data synchronization module 12, a synchronized data transmission optimization module 13, a conflict detection and handling module 14, and a synchronization status monitoring module 15.

[0029] The data partitioning and tagging module 11: is used to divide the stored data into multiple logical partitions, and each logical partition is assigned a unique tag.

[0030] The data synchronization module 12: is used to track all data updates within the logical partition and synchronize these updated data.

[0031] The synchronized data transmission optimization module 13: preprocesses the data during the synchronization process to reduce the amount of data to be transmitted.

[0032] The conflict detection and handling module 14: during the synchronization process, detects whether the synchronized data conflicts. If a conflict occurs, it processes the conflict according to a preset priority processing strategy.

[0033] The synchronization status monitoring module 15: after the data synchronization is completed, checks whether the synchronization status data of the two data centers are consistent. If the synchronization status data of the two data centers are inconsistent, it triggers a data recovery operation.

[0034] The present invention discloses an efficient storage synchronization system for a dual-active data center. By dividing storage data into multiple logical partitions and assigning unique tags, it realizes refined management of data, reduces the amount of data to be synchronized, and improves synchronization efficiency; by reducing the amount of data transmitted, it reduces network bandwidth occupancy and further enhances synchronization efficiency; by tracking all data updates within a logical partition, it ensures that all changes can be captured and synchronized, avoiding data omission. After synchronization is completed, by checking whether the synchronization status data of the two data centers is consistent, it ensures data integrity and consistency; during the synchronization process, it detects and handles data conflicts, avoiding data inconsistency problems caused by conflicts; through a preset priority processing strategy, it can quickly and automatically resolve conflicts, reducing the need for manual intervention; when detecting inconsistent synchronization status, it automatically triggers a data recovery operation to ensure that the data can be quickly restored to a consistent state, improving the self-healing ability and reliability of the system.

[0035] Referring to Figure 2 , in a preferred embodiment, the data partitioning and tagging module 11 includes: a data partitioning unit 111, a tag assignment unit 112, a partition mapping unit 113, and a partition optimization unit 114.

[0036] The data partitioning unit 111: is used to divide the storage data into multiple logical partitions according to a preset partitioning strategy, where the preset partitioning strategy includes but is not limited to partitioning based on data access frequency, data type, data size, or business requirements.

[0037] The tag assignment unit 112: assigns a unique tag to each logical partition, and the tag includes a partition identifier, partition metadata, and identification information of the data center to which the partition belongs.

[0038] The partition mapping unit 113: is used to establish a mapping relationship between the logical partition and the physical storage location, generate a logical partition mapping table, and update the logical partition mapping table in real time.

[0039] The partition optimization unit 114: dynamically adjusts the partition size and distribution according to the data access pattern and synchronization requirements, and optimizes data storage and synchronization efficiency.

[0040] Referring to Figure 3 , in a preferred embodiment, the data synchronization module 12 includes: a log capture unit 121, an incremental data extraction unit 122, a log compression unit 123, a log marking unit 124, an incremental data cache unit 125, and a log cleaning unit 126.

[0041] The log capture unit 121: is used to monitor the update operations of the data partition in real time, record all data changes, including insert, update, and delete operations, through a log capture mechanism, and generate corresponding incremental logs.

[0042] Incremental Data Extraction Unit 122: Extracts incremental data to be synchronized according to the incremental log and filters out irrelevant operations.

[0043] Log Compression Unit 123: Performs compression processing on the incremental log to reduce the log storage space and transmission overhead, while preserving the integrity and traceability of the log.

[0044] Log Tagging Unit 124: Assigns a unique tag to each incremental log. The tag can be a timestamp or a serial number, used to identify the order of data changes and ensure the orderliness and consistency of data synchronization.

[0045] Incremental Data Caching Unit 125: Temporarily caches the extracted incremental data, waits for the Transmission Optimization Module 13 to perform preprocessing, and supports the breakpoint resumption and retransmission mechanisms.

[0046] Log Cleaning Unit 126: After the incremental data is successfully synchronized, clears the processed incremental log, releases the storage space, and archives it to the historical log regularly.

[0047] Refer to Figure 4 , in the preferred embodiment, the Synchronous Data Transmission Optimization Module 13 includes a Data Compression Unit 131, a Data Chunking Unit 132, a Duplicate Data Deletion Unit 133, a Metadata Management Unit (134), a Transmission Protocol Optimization Unit 135, and a Preprocessing Caching Unit 136.

[0048] Data Compression Unit 131: Used to perform compression processing on the synchronous data, adopting a lossless compression algorithm to reduce the data transmission volume.

[0049] Data Chunking Unit 132: Divides the data to be synchronized into data chunks of a fixed size or a dynamic size;

[0050] Duplicate Data Deletion Unit 133: Identifies and deletes duplicate data chunks through a hash algorithm, only transmits unique data chunks, and reorganizes the data at the target data center.

[0051] Metadata Management Unit 134: Generates metadata for each data chunk, including data chunk identification, fingerprint information, compression status, and location information.

[0052] Transmission Protocol Optimization Unit 135: Adopts the QUIC protocol, combines with the data compression unit and the duplicate data deletion unit to improve the transmission performance.

[0053] Specifically: The QUIC protocol is adopted as the basic protocol for data transmission. By leveraging its characteristics of low latency, multiplexing, and fast connection establishment based on UDP, the transmission efficiency is improved. The data compression technology is deeply integrated with the QUIC protocol to perform real-time compression on the transmitted data at the protocol layer, reducing the amount of transmitted data and occupying less network bandwidth. A duplicate data deletion mechanism is embedded in the QUIC protocol. Through data fingerprint technologies such as SHA-256 and MD5, duplicate data blocks are identified and deleted, and only unique data blocks are transmitted, further reducing redundant data transmission. Combining with the adaptive congestion control algorithm of the QUIC protocol, the transmission rate is dynamically adjusted to avoid network congestion while maximizing network utilization. By using the built-in error recovery mechanism of the QUIC protocol, packets are quickly retransmitted when lost or damaged to ensure the reliability and integrity of data transmission. The built-in TLS encryption function of the QUIC protocol is adopted to ensure the security of data transmission and prevent data leakage or tampering. Support for enabling the multi-path transmission function in the QUIC protocol, using multiple network paths to transmit data in parallel to further improve the transmission speed and fault tolerance. The transmission performance of the QUIC protocol, including latency, throughput, packet loss rate, and retransmission rate, is monitored in real time, and the protocol parameters are dynamically optimized according to the monitoring results to ensure the maximization of transmission efficiency.

[0054] The preprocessing cache unit 136: It is used to temporarily store the data blocks that have been compressed and deduplicated.

[0055] Refer to Figure 5 , in the preferred embodiment, the conflict detection and handling module 14 includes: a version number management unit 141, a conflict detection unit 142, a conflict log recording unit 143, a conflict handling strategy unit 144, a conflict notification unit 145, and a conflict recovery unit 146.

[0056] The version number management unit 141: Assigns a unique version number to each data partition or data block and increments the version number when the data is updated to identify the latest state of the data.

[0057] The conflict detection unit 142: During the data synchronization process, compares the version numbers of the source data center and the target data center. If the version numbers are inconsistent, it is determined that a data conflict occurs.

[0058] The conflict log recording unit 143: It is used to record all detected data conflict information, including the identification, version number, and timestamp of the conflict data.

[0059] The conflict handling strategy unit 144: Automatically or semi-automatically resolves conflicts according to the preset conflict handling strategy.

[0060] Specifically, the preset conflict handling strategy can be formulated according to business requirements and technical characteristics.

[0061] Among them, the preset conflict handling strategy can be timestamp priority or business priority. For example, in the order system of a dual-active data center, two data centers simultaneously update the same order: Data center A changes the order status from "pending payment" to "paid". Data center B changes the order status from "pending payment" to "cancelled". If the conflict is handled according to the timestamp priority strategy, the update with the more recent timestamp shall prevail; if according to the business priority strategy, the "paid" status takes precedence over the "cancelled" status and payment operations are generally more important than cancellation operations, then the order status shall be set to "paid"; if the system does not automatically handle the conflict after detecting it, it shall notify the administrator or relevant business personnel to resolve it manually.

[0062] Conflict notification unit 145: When a conflict is detected and cannot be automatically handled, it sends a conflict notification to the administrator or relevant business system, and provides conflict details and handling suggestions.

[0063] Conflict recovery unit 146: After the conflict handling is completed, it ensures that the conflict data is consistent in the dual-active data center and records the conflict handling result.

[0064] Refer to Figure 6 As shown in

[0065] Synchronization status acquisition unit 151: After data synchronization is completed, it acquires the status data of the data partitions from the source data center and the target data center respectively, including data version number, timestamp, data fingerprint and metadata.

[0066] Consistency verification unit 152: By comparing the status data of the source data center and the target data center, it verifies whether the status data of the two data centers is consistent. If it is detected that the status data of the two data centers is inconsistent, it generates a difference report.

[0067] Difference analysis unit 153: Analyzes the difference report to determine the difference type, and the difference type includes data loss, data conflict or data corruption.

[0068] Automatic repair unit 154: Used to automatically trigger a data recovery operation when it is detected that the status data of the two data centers is inconsistent, so as to ensure data consistency.

[0069] Among them, the data recovery operations include data resynchronization operations, data reconstruction operations, and backup recovery operations. For example, when data loss or partial data inconsistency is detected, the missing or incorrect data is retrieved from the source data center or backup and synchronized to the target data center to restore consistency; when the data is corrupted or cannot be directly used, the corrupted data blocks or files are reconstructed using check codes such as CRC, redundant data such as RAID, or backup data; when the data is lost or corrupted but there is a complete operation log, the data changes are re-executed by playing back the operation records in the log to restore to the latest correct state.

[0070] Status log recording unit 155: Used to record the status data, verification results, and repair operations of each synchronization, and generate a synchronization log.

[0071] Alarm notification unit 156: Used to send alarm notifications to the system administrator or relevant business systems when data inconsistency or synchronization failure is detected.

[0072] Visualization unit 157: Used to provide a graphical interface to display the status data of data synchronization and the results of verification status data in real time.

[0073] Embodiment 2, referring to Figure 7 , which is the second embodiment of the present invention. Different from the previous embodiment, it provides an efficient storage synchronization method for a dual-active data center, including:

[0074] S1. Divide the storage data into multiple logical partitions according to a preset policy, and assign a unique label to each partition, including a partition identifier, metadata, and data center identification information. At the same time, establish a mapping relationship between the partition and the physical storage location, and dynamically adjust the partition size and distribution to optimize storage and synchronization efficiency.

[0075] S2. Real-time monitor the update operations of the data partitions, record the data changes through a log capture mechanism, and generate an incremental log. Then extract the incremental data to be synchronized and filter out irrelevant operations.

[0076] S3. Compress the synchronization data, use a lossless compression algorithm to reduce the data transmission volume; divide the data into data blocks of fixed or dynamic size, identify and delete duplicate data blocks through a hash algorithm, and only transmit unique data blocks, and generate metadata.

[0077] S4. During the synchronization process, compare the data version numbers of the source data center and the target data center. If the version numbers are inconsistent, it is determined as a data conflict. Record the conflict information, automatically or semi-automatically resolve the conflict according to a preset policy, and ensure data consistency after the processing is completed.

[0078] S5. After the data synchronization is completed, collect the status information of the data partitions from the source data center and the target data center, including version numbers, timestamps, data fingerprints, and metadata. Verify whether the data is consistent by comparing the status information. If it is inconsistent, generate a difference report and automatically trigger data recovery operations, such as re-synchronization, data reconstruction, or backup restoration, to ensure data consistency between the two data centers.

[0079] If the above-described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, and other various media that can store program codes.

[0080] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0081] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other suitable processing, and then stored in a computer memory.

[0082] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0083] Embodiment 3, the third embodiment of the present invention, which is different from the previous two embodiments is: to verify and illustrate the technical effects adopted in the present invention to verify the actual effects of the method.

[0084] Data partitioning unit 111: used to divide the stored data into multiple logical partitions according to a preset partitioning strategy, where the preset partitioning strategy includes but is not limited to partitioning based on data access frequency, data type, data size, or business requirements.

[0085] Label assignment unit 112: assigns a unique label to each logical partition, where the label includes a partition identifier, partition metadata, and identification information of the data center to which the partition belongs; according to business requirements, partitions the order data according to order type, and the order types include ordinary orders, group purchase orders, pre-sale orders, etc. The label for the ordinary order partition is "ORD001", and the label for the group purchase order partition is "ORD002". The label contains a partition identifier such as "ORD001", partition metadata such as the order type being an ordinary order, and the identification information of the data center to which it belongs is data center A. Partition mapping unit 113: used to establish a mapping relationship between the logical partition and the physical storage location, generate a logical partition mapping table, and update the logical partition mapping table in real time; the ordinary order partition is stored on the 1st to 10th disks of disk array 1 in data center A, and the group purchase order partition is stored on the 11th to 20th disks of disk array 1. Generate a logical partition mapping table and update it in real time to reflect changes in the physical storage location of the partition.

[0086] Partition optimization unit 114: dynamically adjusts the partition size and distribution according to the data access pattern and synchronization requirements to optimize the data storage and synchronization efficiency. If the access frequency of group purchase orders increases significantly, the system will automatically adjust the size of the group purchase order partition, allocate more storage space, and optimize its distribution in physical storage to improve storage and synchronization efficiency.

[0087] Log Capture Unit 121: It is used to monitor the update operations of data partitions in real time, record all data changes through a log capture mechanism, including insert, update, and delete operations, and generate corresponding incremental logs.

[0088] Incremental Data Extraction Unit 122: Extracts the incremental data to be synchronized according to the incremental logs and filters out irrelevant operations.

[0089] Log Compression Unit 123: Performs compression processing on the incremental logs, reduces the log storage space and transmission overhead, while preserving the integrity and traceability of the logs. For example, the LZ77 algorithm is used to compress the logs to reduce the log storage space and transmission overhead.

[0090] Log Tagging Unit 124: Assigns a unique tag to each incremental log. The tag can be a timestamp or a serial number. For example, a timestamp "2025-02-12 10:00:00" is assigned to a log of order status update to identify the order of data changes and ensure the orderliness and consistency of data synchronization; Incremental Data Cache Unit 125: Temporarily caches the extracted incremental data, waits for the Transmission Optimization Module 13 to perform preprocessing, and supports breakpoint resumption and retransmission mechanisms. For example, if a network interruption occurs during transmission, the system can recover the data from the cache queue and retransmit it.

[0091] Log Cleaning Unit 126: After the incremental data is successfully synchronized, it cleans the processed incremental logs, releases the storage space, and archives them to historical logs regularly.

[0092] Data Compression Unit 131: Used to perform compression processing on the synchronized data, adopting a lossless compression algorithm such as Gzip to reduce the data transmission volume.

[0093] Data Chunking Unit 132: Divides the data to be synchronized into data chunks of a fixed size or a dynamic size. For example, order data is divided into data chunks of 1MB each.

[0094] Duplicate Data Deletion Unit 133: Identifies and deletes duplicate data chunks through a hash algorithm, only transmits unique data chunks, and reorganizes the data at the target data center.

[0095] Metadata Management Unit 134: Generates metadata for each data chunk, including data chunk identification, fingerprint information, compression status, and location information.

[0096] Transmission Protocol Optimization Unit 135: Adopts the QUIC protocol, combines with the Data Compression Unit and the Duplicate Data Deletion Unit to improve the transmission performance.

[0097] Preprocessing Cache Unit 136: Used to temporarily store the data chunks that have been compressed and deduplicated.

[0098] Version Number Management Unit 141: Assigns a unique version number to each data partition or data block, and increments the version number when the data is updated to identify the latest state of the data. For example, the initial version number of a data partition is 1, and when the data is updated, the version number is incremented to 2.

[0099] Conflict Detection Unit 142: During the data synchronization process, compares the version numbers of the source data center and the target data center. If the version numbers are inconsistent, it is determined that a data conflict has occurred. For example, when Data Center A changes the order status from "Pending Payment" to "Paid", the version number is incremented to 2, while Data Center B changes the order status from "Pending Payment" to "Cancelled" at the same time, and the version number is also incremented to 2. During synchronization, when comparing the version numbers of the two data centers, it is found that the version numbers are inconsistent, and a data conflict is determined.

[0100] Conflict Log Recording Unit 143: Used to record all detected data conflict information, including the identification, version number, and timestamp of the conflicting data.

[0101] Conflict Handling Strategy Unit 144: Automatically or semi-automatically resolves conflicts according to the preset conflict handling strategy.

[0102] Conflict Notification Unit 145: When a conflict is detected and cannot be automatically processed, sends a conflict notification to the administrator or the relevant business system, and provides conflict details and handling suggestions.

[0103] Conflict Recovery Unit 146: After the conflict is resolved, ensures that the conflicting data is consistent in the dual-active data centers and records the conflict resolution result.

[0104] Synchronization Status Collection Unit 151: After the data synchronization is completed, collects the status data of the data partitions from the source data center and the target data center respectively, including the data version number (e.g., 2 for Data Center A and 2 for Data Center B), timestamp (e.g., 2025-02-12 10:00:00 for Data Center A and 2025-02-12 10:00:01 for Data Center B), data fingerprint, and metadata; Consistency Verification Unit 152: Verifies whether the status data of the two data centers is consistent by comparing the status data of the source data center and the target data center. If it is detected that the status data of the two data centers is inconsistent, a difference report is generated.

[0105] Difference Analysis Unit 153: Analyzes the difference report to determine the type of difference. The types of differences include data loss, data conflict, or data corruption.

[0106] Automatic Repair Unit 154: It is used to automatically trigger data recovery operations when it detects that the status data of two data centers is inconsistent. For example, if the Synchronization Status Acquisition Unit 151 discovers that a certain piece of data in the target data center is missing, the Automatic Repair Unit 154 will retrieve the data from the source data center again and synchronize it to the target data center to ensure data consistency between the two data centers.

[0107] Status Logging Unit 155: It is used to record the status data, verification results, and repair operations of each synchronization, and generate synchronization logs. For example, it records that the synchronization time is 2025-02-12 10:00:05, the verification result is data inconsistency, and the repair operation is to resynchronize order data.

[0108] Alarm Notification Unit 156: It is used to send alarm notifications to system administrators or relevant business systems when data inconsistency or synchronization failure is detected. Specifically, the alarm notification includes an alarm title and alarm content. For example, the alarm title is "Data Synchronization Alarm - Inconsistent Data in the General Order Partition", and the alarm content is "Inconsistent data in the general order partition is detected:

[0109] Data Center A: Version number is 2, timestamp is 2025-02-12 10:00:00; data fingerprint is SHA-256: abcdef1234567890;

[0110] Data Center B: Version number is 2, timestamp is 2025-02-12 10:00:01; data fingerprint is SHA-256: ghijkl9876543210

[0111] Recovery Operation: The automatic repair operation has been triggered, and it is attempting to synchronize data from Data Center A to Data Center B.

[0112] Recipient: System administrator (such as administrator email: admin@example.com) and relevant business systems.

[0113] Visualization Unit 157: It is used to provide a graphical interface to display the status data of data synchronization and the results of verification status data in real time. Specifically, the data synchronization status includes a synchronization progress bar, synchronization status, and synchronization time. Synchronization Progress Bar: Displays the progress of the current synchronization operation, for example, "80% completed".

[0114] Synchronization Status: Displays the current synchronization status, such as "In Progress", "Success", or "Failure".

[0115] Synchronization Time: Displays the start time and estimated completion time of the synchronization operation.

[0116] The results of verification status data include data consistency status, difference report, and repair operation status.

[0117] Consistency status: display the "consistent" or "inconsistent" status; Difference report: if the data is inconsistent, display a detailed difference report, including information such as the identification of the different data blocks, data fingerprints, version numbers, etc.; Repair operation status: if an automatic repair operation is triggered, display the progress and results of the repair operation.

[0118] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An efficient storage synchronization system for a dual-active data center, characterized in that: including, a data partitioning and tagging module (11), a data synchronization module (12), a synchronized data transmission optimization module (13), a conflict detection and handling module (14), and a synchronization status monitoring module (15); the data partitioning and tagging module (11) is used to divide stored data into multiple logical partitions, and each logical partition is assigned a unique tag; the data synchronization module (12) is used to track all data updates within the logical partitions and synchronize the updated data; the synchronized data transmission optimization module (13) preprocesses the data during the synchronization process to reduce the amount of data to be transmitted; the conflict detection and handling module (14) detects whether the synchronized data conflicts during the synchronization process. If a conflict occurs, it processes the conflict according to a preset priority processing strategy; the synchronization status monitoring module (15) checks whether the synchronization status data of the two data centers is consistent after the data synchronization is completed. If the synchronization status data of the two data centers is inconsistent, it triggers a data recovery operation.

2. The high-efficiency storage synchronization system for a dual-active data center according to claim 1, characterized in that: the data partitioning and tagging module (11) includes a data partitioning unit (111), a tag assignment unit (112), a partition mapping unit (113), and a partition optimization unit (114); the data partitioning unit (111) is used to divide stored data into multiple logical partitions according to a preset partitioning strategy; the tag assignment unit (112) assigns a unique tag to each logical partition. The tag includes a partition identifier, partition metadata, and identification information of the data center to which the partition belongs; the partition mapping unit (113) is used to establish a mapping relationship between the logical partition and the physical storage location, generate a logical partition mapping table, and update the logical partition mapping table in real time; the partition optimization unit (114) dynamically adjusts the partition size and distribution according to the data access pattern and synchronization requirements.

3. The highly efficient storage synchronization system for a dual-active data center according to claim 2, characterized in that: the data synchronization module (12) includes a log capture unit (121), an incremental data extraction unit (122), a log compression unit (123), a log marking unit (124), an incremental data cache unit (125), and a log cleaning unit (126); the log capture unit (121) is used to monitor the update operations of the data partitions in real time, record all data changes, including insert, update, and delete operations, through a log capture mechanism, and generate corresponding incremental logs; the incremental data extraction unit (122) extracts the incremental data to be synchronized according to the incremental logs and filters out irrelevant operations; the log compression unit (123) performs compression processing on the incremental logs; the log marking unit (124) assigns a unique mark to each incremental log; the incremental data cache unit (125) temporarily caches the extracted incremental data, waits for the transmission optimization module (13) to perform preprocessing, and supports breakpoint resumption and retransmission mechanisms; the log cleaning unit (126) cleans the processed incremental logs, releases the storage space, and archives them to the historical logs regularly after the incremental data is successfully synchronized.

4. The efficient storage synchronization system for a dual-active data center according to claim 3, wherein: the synchronized data transmission optimization module (13) includes Data compression unit (131), data chunking unit (132), deduplication unit (133), metadata management unit (134), transport protocol optimization unit (135), and preprocessing cache unit (136); The data compression unit (131) is used to compress the synchronization data, adopting a lossless compression algorithm to reduce the data transmission volume; The data chunking unit (132) divides the data to be synchronized into data chunks of fixed size or dynamic size; The deduplication unit (133) identifies and deletes duplicate data chunks through a hash algorithm, transmits only unique data chunks, and reorganizes the data at the target data center; The metadata management unit (134) generates metadata for each data chunk, including data chunk identifier, fingerprint information, compression status, and location information; The transport protocol optimization unit (135) adopts the QUIC protocol, combines the data compression unit (131) and the deduplication unit (133) to improve the transmission performance; The preprocessing cache unit (136) is used to temporarily store the data chunks that have been compressed and deduplicated.

5. The high-efficiency storage synchronization system for a dual-active data center according to claim 4, characterized in that: The conflict detection and handling module (14), including, version number management unit (141), conflict detection unit (142), conflict log recording unit (143), conflict handling strategy unit (144), conflict notification unit (145), and conflict recovery unit (146); The version number management unit (141) assigns a unique version number to each data partition or data chunk, and increments the version number when the data is updated to identify the latest state of the data; The conflict detection unit (142) compares the version numbers of the source data center and the target data center during the data synchronization process. If the version numbers are inconsistent, it is determined that a data conflict occurs; The conflict log recording unit (143) is used to record all detected data conflict information, including the identifier, version number, and timestamp of the conflict data; The conflict handling strategy unit (144) resolves conflicts according to the preset conflict handling strategy; The conflict notification unit (145) sends a conflict notification to the administrator or relevant business systems and provides conflict details and handling suggestions when a conflict is detected and cannot be automatically processed; The conflict recovery unit (146) ensures the consistency of the conflict data in the active-active data centers after the conflict is handled and records the conflict handling result.

6. The high-efficiency storage synchronization system for a dual-active data center according to claim 5, wherein: The synchronization status monitoring module (15), including, synchronization status collection unit (151), consistency verification unit (152), difference analysis unit (153), automatic repair unit (154), status log recording unit (155), alarm notification unit (156), and visualization unit (157); The synchronization status collection unit (151) collects the status information of the data partitions from the source data center and the target data center respectively after the data synchronization is completed, including data version number, timestamp, data fingerprint, and metadata; The consistency check unit (152) checks whether the data is consistent by comparing the status information of the source data center and the target data center. If inconsistency is found, a difference report is generated; The difference analysis unit (153) analyzes the difference report to determine the type of difference, and the types of difference include data loss, data conflict or data corruption; The automatic repair unit (154) is used to automatically trigger a data recovery operation when the status data of the two data centers is detected to be inconsistent; The status log recording unit (155) is used to record the status information, check results and repair operations of each synchronization, and generate a synchronization log; The alarm notification unit (156) is used to send an alarm notification to the system administrator or relevant business systems when data inconsistency or synchronization failure is detected; The visualization unit (157) is used to provide a graphical interface to display the status data of data synchronization and the results of check status data in real time.

7. The highly efficient storage synchronization system for a dual-active data center according to claim 6, wherein: The data recovery operation includes data re-synchronization operation, data reconstruction operation and backup recovery operation.

8. A method of using an efficient storage synchronization system for a dual-active data center as described in any one of claims 1 to 7, characterized in that: Include, Divide the stored data into multiple logical partitions according to a preset policy, and assign a unique label to each partition, including a partition identifier, metadata and data center identification information. At the same time, establish a mapping relationship between the partition and the physical storage location, and dynamically adjust the partition size and distribution to optimize storage and synchronization efficiency; Monitor the update operations of data partitions in real time, record data changes through a log capture mechanism, generate incremental logs, extract the incremental data to be synchronized, and filter out irrelevant operations; Compress the synchronized data, use a lossless compression algorithm to reduce the data transmission volume, divide the data into data blocks of fixed or dynamic size, identify and delete duplicate data blocks through a hash algorithm, only transmit unique data blocks, and generate metadata; During the synchronization process, compare the data version numbers of the source data center and the target data center. If the version numbers are inconsistent, it is determined as a data conflict, record the conflict information, automatically or semi-automatically resolve the conflict according to a preset policy, and ensure data consistency after the processing is completed; After the data synchronization is completed, collect the status information of the data partitions from the source data center and the target data center, including version number, timestamp, data fingerprint and metadata. Check whether the data is consistent by comparing the status information. If it is inconsistent, generate a difference report and automatically trigger a data recovery operation to ensure data consistency between the two data centers.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the efficient storage synchronization method for a dual-active data center.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the efficient storage synchronization method for a dual-active data center.