Data synchronization method, apparatus, device, and medium

By creating a mapping relationship between the primary and secondary volumes using high-density snapshot technology, physical volume snapshots are generated. This solves the problem of multiple snapshot starts and stops at the second-level cycle in existing technologies, improving storage space utilization and the number of replication relationships, while reducing storage costs and system performance impact.

CN119988489BActive Publication Date: 2026-03-31INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing remote replication technologies cannot achieve short-cycle, multiple start-stop snapshots with second-level intervals and asynchronous replication. Furthermore, ROW snapshots consume a large amount of storage space, limiting the number of replication relationships.

Method used

High-density snapshot technology is used to periodically create high-density snapshots on both the primary and secondary volumes, establish mapping relationships, generate physical volume snapshots, reduce discard and flush processes, configure only one physical volume snapshot during the synchronization process, and achieve data synchronization by temporarily changing the volume.

Benefits of technology

It improves storage space utilization, reduces the time consumed in starting and stopping snapshots, supports more replication relationships, and reduces storage costs and system performance impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988489B_ABST
    Figure CN119988489B_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field and discloses a data synchronization method, device and equipment and medium, the method comprises the following steps: periodically creating high-density snapshots on the one side of a main volume and the one side of an auxiliary volume respectively, establishing a first mapping relationship between the high-density snapshot on the one side of the main volume and the high-density snapshot on the one side of the auxiliary volume, generating identification information corresponding to I / O requests issued between a t-1th period and a tth period according to the t-1th period on the one side of the main volume, creating a change volume corresponding to the main volume after creating the high-density snapshot on the tth period on the one side of the main volume, extracting preset data after performing an operation corresponding to the I / O request from a preset storage location in the change volume according to the identification information, generating an entity volume snapshot, establishing a second mapping relationship between the high-density snapshot created on the tth period on the one side of the main volume and the entity volume snapshot, and synchronizing data of the auxiliary volume according to the entity volume snapshot and the first and second mapping relationships.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a data synchronization method, device, equipment and medium. BACKGROUND

[0002] With the rapid development of information technology, the demand for data protection is increasing. Traditional snapshot technology, such as Copy-On-Write (COW) and Redirect-On-Write (ROW).

[0003] The current remote replication change volume snapshot method uses ROW snapshot. ROW snapshot needs to perform discard (check and remove those data rows without changes) and flush (process of writing data in memory to disk, ensuring that all data changes are persisted to storage device) processes on the data on the target volume every time it starts. The above processes bring additional overhead, which leads to the inability to support second-level periodic asynchronous replication, a complex scenario of short-period multiple start-stop snapshots.

[0004] In addition, the storage space occupied by the ROW snapshot is large. A periodic asynchronous remote replication consists of four volumes, which requires four ROW snapshots. If a cluster can support a maximum of 10,000 volumes, it can only support a maximum of 2,500 periodic asynchronous remote replications, which limits the number of replication relationships. SUMMARY

[0005] Therefore, the present application provides a data synchronization method, device, equipment and medium to solve the problem that the existing remote replication technology cannot realize second-level periodic, asynchronous replication, a short-period, multiple start-stop snapshot.

[0006] In a first aspect, the present application provides a data synchronization method. The method is applied to an asynchronous remote replication system, which includes a primary volume and a secondary volume. The method includes:

[0007] Periodically creating high-density snapshots on the primary volume side and the secondary volume side, respectively, and establishing a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side;

[0008] According to the (t-1)th period on the primary volume side, generating identification information corresponding to the I / O request issued between the (t-1)th period and the (t)th period, t being a positive integer greater than or equal to 2;

[0009] After creating a high-density snapshot on the primary volume side in the (t)th period, creating a change volume corresponding to the primary volume;

[0010] In the change volume, according to the identification information, preset data after operation corresponding to the I / O request is extracted from a preset storage location, and an entity volume snapshot is generated, wherein the preset data carries the identification information;

[0011] A second mapping relationship between the high-density snapshot created on the main volume side in the tth period and the entity volume snapshot is established;

[0012] According to the entity volume snapshot, the first mapping relationship, and the second mapping relationship, data synchronization is performed on the auxiliary volume.

[0013] The data synchronization method provided by the application has the following advantages: periodically creating high-density snapshots on the main volume side and the auxiliary volume side, rather than manually or event-triggered high-density snapshot creation actions, providing continuous data protection and reducing the risk of data loss. In write-intensive applications, periodic creation does not require creation operations at each data change, which can avoid affecting system performance due to frequent snapshot operations. Establishing a first mapping relationship between the high-density snapshot on the main volume side and the high-density snapshot on the auxiliary volume side helps quickly recover the data of the main volume from the auxiliary volume when needed. According to the t-1th period on the main volume side, identification information corresponding to the I / O request issued between the t-1th period and the tth period is generated. The high-density snapshot does not occupy the entity volume space, but only identifies the I / O request issued and the data after the operation. Subsequently, the data after the operation corresponding to each batch of I / O requests is read from the change volume corresponding to the main volume to generate an entity volume snapshot. Each period of asynchronous remote replication only configures one entity volume snapshot in the synchronization process, reducing the space requirement of periodic asynchronous replication and the time consumption of starting and stopping the auxiliary volume change snapshot, while improving the storage space utilization and configuration quantity limit of periodic asynchronous replication due to the limitation of traditional ROW snapshots. The high-density snapshot discards the discard and flush processes, improves the resource consumption and performance impact in each remote replication process, and further reduces the minimum period supported by periodic asynchronous replication.

[0014] The data synchronization process using the high-density snapshot realizes periodic asynchronous remote replication, which can reduce the minimum period supported by periodic asynchronous remote replication and the time consumption of starting and stopping the snapshot in the synchronization process. The high-density snapshot is converted to an entity volume snapshot by creating a temporary change volume on the main volume side, which only occupies the storage space of three volumes instead of the snapshot mechanism of the original ROW main and auxiliary two-side change volumes occupying the storage space of four volumes, greatly improving the storage space utilization and greatly improving the upper limit of the configuration quantity of periodic asynchronous replication.

[0015] In an optional embodiment, after the data synchronization of the auxiliary volume according to the entity volume snapshot, the first mapping relationship, and the second mapping relationship, the method further comprises:

[0016] deletion change volume.

[0017] Specifically, after creating the high-density snapshot of the main volume side in the tth period, a temporary change volume corresponding to the main volume is created, which only exists in the period synchronization process, and its function is to store the data after the operation corresponding to the I / O request is performed in the tth period, and synchronize all data to the secondary end. Since the data also carries the same identification information as the I / O request, the change volume does not need to exist all the time to ensure data security. After synchronization is completed, the change volume can be automatically deleted, that is, during the period interval at the end of the synchronization process, the period asynchronous replication only occupies the space of two volumes, improving the utilization rate of storage resources and reducing storage costs.

[0018] In an optional embodiment, the secondary volume is synchronized according to the entity volume snapshot, the first mapping relationship, and the second mapping relationship, and specifically includes:

[0019] According to the entity volume snapshot and the second mapping relationship, the high-density snapshot of the main volume side in the tth period corresponding to the entity volume snapshot is determined;

[0020] According to the high-density snapshot of the main volume side in the tth period and the first mapping relationship, the high-density snapshot of the secondary volume side in the tth period is determined;

[0021] According to the preset data in the entity volume snapshot, the data synchronization of the secondary volume is completed.

[0022] Specifically, according to the preset data in the entity volume snapshot, the data that needs to be synchronized in the tth period can be accurately determined, so as to realize accurate data synchronization. According to the high-density snapshot of the main volume side in the tth period and the first mapping relationship, the high-density snapshot of the secondary volume side in the tth period is determined, the data synchronization of the secondary volume is completed, and the data consistency between the main volume and the secondary volume is ensured. Through the entity volume snapshot and the mapping relationship, the data is flexibly managed, ensuring that the correct data state can be accurately and accurately restored in the data recovery process, improving the accuracy of data recovery, and minimizing the risk of data loss.

[0023] In an optional embodiment, according to the t-1th period of the main volume side, identification information corresponding to the I / O request issued between the t-1th period and the tth period is generated, specifically including:

[0024] According to the t-1th period of the main volume side, a generation relationship is generated;

[0025] According to the generation relationship, identification information corresponding to the I / O request issued between the t-1th period and the tth period is generated.

[0026] Specifically, based on the (t-1)th cycle on the main volume side, a generation relationship is generated. Based on this generation relationship, identification information corresponding to the I / O requests issued between the (t-1)th cycle and the tth cycle is generated. By accurately capturing and identifying I / O requests under each generation, the system can more effectively manage and schedule each batch of data, ensuring data integrity and accuracy. Furthermore, adding identification information to I / O requests based on the generation relationship helps to accurately roll back to the required generation. Subsequently, based on the identification information corresponding to the I / O requests, the data for which corresponding operations have been performed is also appended with the same identification information. When periodic replication is initiated, the changed data for each batch is stored in the changed volume for physical volume snapshot generation, instead of operating directly on the main volume, which reduces the impact on the main volume's I / O performance.

[0027] In one alternative implementation, the method further includes:

[0028] Obtain the network transmission rate, packet loss rate, link bandwidth, number of remote replications per link, data transmission volume in each period, and the time period corresponding to the period in the cluster to which the asynchronous remote replication system belongs;

[0029] Based on the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications per link in the cluster to which the asynchronous remote replication system belongs, the average remote replication load of each remote replication link is determined. The cluster to which the remote replication system belongs includes multiple remote replication links, and each remote replication link includes multiple remote replication operations.

[0030] The remote replication load is determined based on the data transfer volume within each cycle and the corresponding time period of the cycle;

[0031] Adjust the time period corresponding to the cycle based on the remote replication load and the average remote replication load.

[0032] Specifically, by acquiring key performance indicators such as network transmission rate, packet loss rate, link bandwidth, number of remote copies per link, data transmission volume in each period, and the time period corresponding to the period, the performance of the asynchronous remote replication system can be monitored and optimized in real time, and problems that may affect data replication efficiency and reliability can be identified and resolved in a timely manner.

[0033] By calculating the average remote replication load and the actual remote replication load of each link, and adjusting the corresponding time period based on the remote replication load and the average remote replication load, resources can be allocated more rationally, load balancing can be achieved, the efficiency and stability of the entire system can be improved, and performance bottlenecks caused by uneven load can be avoided.

[0034] In one optional implementation, the average remote replication load of each remote replication link is determined based on the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications per link within the cluster to which the asynchronous remote replication system belongs. Specifically, this includes:

[0035] The actual quality of each remote replication link is determined based on the network transmission rate and packet loss rate.

[0036] Determine the optimal quality for each remote replication link based on the link bandwidth in the cluster;

[0037] Determine the bandwidth load for each remote replication link based on actual and optimal quality.

[0038] The average remote replication load for each remote replication link is determined based on the bandwidth load and the number of remote replications for each link.

[0039] Specifically, based on network transmission rate and packet loss rate, the actual quality of each remote replication link is determined, and the optimal quality of each remote replication link is determined based on the link bandwidth in the cluster. Adjusting the actual quality through the optimal quality allows for more rational allocation of network resources and improves overall network efficiency. By determining the bandwidth load of each remote replication link, the load of each link can be balanced, preventing some links from being overloaded and affecting replication efficiency. Combining bandwidth load and the number of remote replications, the average remote replication load of each link is determined, which helps to adjust the cycle according to the business load. The cycle is shortened when the business load is high, reducing the amount of synchronization data in each synchronization process and improving the synchronization rate; the asynchronous remote replication cycle is lengthened when the business load is low or even zero, reducing the time occupied by temporary volume changes during synchronization and improving storage space utilization.

[0040] In one optional implementation, the period corresponds to a preset time period;

[0041] The step of adjusting the time period corresponding to the cycle based on the remote replication load and the remote replication average load specifically includes:

[0042] When the remote replication load is greater than the remote replication average load, the time period corresponding to the cycle is reduced to a first preset multiple of the preset time period;

[0043] or,

[0044] When the remote replication load is less than the remote replication average load, the time period corresponding to the cycle is amplified to a second preset multiple of the preset time period, wherein the first preset multiple is less than the second preset multiple.

[0045] Specifically, this method can adjust the periodic asynchronous remote replication cycle in the system according to the business load. When the business load is high, the cycle is shortened, the amount of data synchronized in each synchronization process is reduced, and the synchronization rate is improved. When the business load is low or there is no business, the periodic asynchronous remote replication cycle is lengthened, which reduces the space occupation time of temporary volume changes during the synchronization process and improves storage space utilization.

[0046] In a second aspect, the present invention provides a data synchronization device, the device comprising:

[0047] Create a module for periodically creating high-density snapshots on both the main volume and the secondary volume side;

[0048] The processing module is used to establish the first mapping relationship between the high-density snapshots on the main volume side and the high-density snapshots on the auxiliary volume side; based on the (t-1)th cycle on the main volume side, it generates identification information corresponding to the I / O requests issued between the (t-1)th cycle and the tth cycle, where t is a positive integer greater than or equal to 2;

[0049] The creation module is also used to create a change volume corresponding to the main volume after a high-density snapshot is created on the main volume side in the t-th period;

[0050] The extraction module is used to extract preset data from a preset storage location within the modified volume, based on the identification information, after performing the operation corresponding to the I / O request. The preset data carries the identification information.

[0051] The processing module is also used to generate physical volume snapshots based on preset data; and to establish a second mapping relationship between the high-density snapshot created on the main volume side in the t-th period and the physical volume snapshot.

[0052] The synchronization module is used to synchronize data on the secondary volume based on the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0053] This invention provides a data synchronization device with the following advantages: It achieves asynchronous remote replication based on high-density snapshots through the collaborative work of various modules. High-density snapshots are created periodically on both the primary and secondary volumes, rather than being created manually or by events, providing continuous data protection and reducing the risk of data loss. Establishing a primary mapping relationship between the high-density snapshots on the primary and secondary volumes facilitates rapid recovery of the primary volume's data from the secondary volume when needed. High-density snapshots do not occupy physical volume space; they merely batch-identify the issued I / O requests and the data after the executed operations. Subsequent creation of change volumes corresponding to the primary volume reads the data after the executed operations corresponding to each batch of I / O requests and converts it into physical volume snapshots. Each periodic asynchronous remote replication configures only one physical volume snapshot during the synchronization process, reducing the space requirements of periodic asynchronous replication and the time consumption of starting and stopping secondary volume change volume snapshots, while improving the storage space utilization and configuration quantity limitations of periodic asynchronous replication due to traditional ROW snapshots. High-density snapshots eliminate the discard and flush processes, increasing resource consumption and performance impact during each remote replication process, and further reducing the minimum cycle supported by periodic asynchronous replication.

[0054] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the data synchronization method described in the first aspect or any corresponding embodiment thereof.

[0055] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the data synchronization method described in the first aspect or any corresponding embodiment thereof.

[0056] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the data synchronization method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0057] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0058] Figure 1 This is a flowchart illustrating a data synchronization method provided in an embodiment of the present invention;

[0059] Figure 2 This is a flowchart illustrating another data synchronization method provided in an embodiment of the present invention;

[0060] Figure 3 This is a flowchart illustrating another data synchronization method provided in an embodiment of the present invention;

[0061] Figure 4 This is a flowchart illustrating another data synchronization method provided in an embodiment of the present invention;

[0062] Figure 5 This is a structural block diagram of a data synchronization device provided in an embodiment of the present invention;

[0063] Figure 6 This is a schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0065] With the rapid development of information technology, the demand for data protection is increasing. Traditional snapshot technologies, such as Copy-On-Write (COW) and Redirect-On-Write (ROW), while achieving data backup and recovery to a certain extent, suffer from problems such as large storage space consumption and low recovery efficiency. Especially in critical applications, users have increasingly higher demands for Continuous Data Protection (CDP), which traditional snapshot technologies can hardly meet. To address this, the industry has developed high-density snapshot technology, which provides users with continuous data protection capabilities through intensive data protection cycles and efficient storage space utilization. However, the configuration of existing high-density snapshot technologies is relatively complex, and their flexibility and configurability in different application scenarios need improvement. For example, in disaster recovery scenarios such as periodic asynchronous remote replication, dual-active remote replication, and 3DC, the differences in bitmaps between high-density snapshots and traditional snapshots hinder their flexible application.

[0066] The scenarios in which the current snapshot is applied for remote replication include:

[0067] Periodic asynchronous remote replication and dual-active change volume snapshots enable remote replication by periodically starting and stopping change volume snapshots to achieve information synchronization between the primary and secondary ends. The key to this functionality lies in the snapshot's bitmap recording the data differences between the source and change volumes. Every period, these differences are synchronized to the secondary end. Furthermore, the primary and secondary change volumes store the data from the previous period, allowing for disaster recovery by initiating reverse change volume snapshots. High-density snapshots, however, have limited application in remote replication scenarios due to their overly frequent data protection cycles and lack of memory space support for configuring bitmaps.

[0068] Specifically, the current remote replication change volume snapshot method uses ROW snapshots. In each cycle, remote replication performs snapshot-related operations according to the following steps.

[0069] 1) Start the primary client to a snapshot of the primary change volume.

[0070] 2) Snapshot from the start of the terminal to the secondary change volume.

[0071] 3) Transfer the changed volume data from the primary end to the secondary end.

[0072] 4) Stop the primary side from changing volume snapshots.

[0073] 5) Stop changing volume snapshots from the slave end.

[0074] In the above process, the periodic start and stop of ROW snapshots are used to achieve periodic data synchronization and protection. When a failure occurs on the primary and secondary ends and data is lost, a reverse snapshot of the changed volume to the source volume can be started to restore the source volume data to the most recent period and protect data security.

[0075] The aforementioned ROW snapshots require discarding and flushing the data on the target volume each time they are started, which cannot support complex scenarios such as short-cycle asynchronous replication with multiple snapshot starts and stops. In addition, ROW snapshots occupy a large amount of storage space. One cycle of asynchronous remote replication consists of four volumes and requires four ROW snapshots. If the maximum number of volumes that a cluster can support is 10,000, it can only support a maximum of 2,500 cycle of asynchronous remote replication, which limits the number of replication relationships.

[0076] To address the aforementioned problems, this invention provides a data synchronization embodiment. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system (computer device) including a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0077] This embodiment provides a data synchronization method that can be used in the aforementioned terminal devices, such as mobile phones and tablet computers.Figure 1 This is a flowchart illustrating a data synchronization method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, this method is applied to an asynchronous remote replication system, which includes a primary volume and a secondary volume. The method flow includes the following steps:

[0078] Step S101: Periodically create high-density snapshots on both the main volume and the auxiliary volume side, and establish the first mapping relationship between the high-density snapshots on the main volume side and the high-density snapshots on the auxiliary volume side.

[0079] Specifically, the client presets a high-density snapshot creation cycle based on actual scenario requirements. The preset cycle can be a specific time interval on the order of seconds. The main volume side and the auxiliary volume side create corresponding high-density snapshots according to the preset cycle. The high-density snapshots created on the main volume side correspond one-to-one with the high-density snapshots created on the auxiliary volume side, thereby establishing the first mapping relationship between the high-density snapshots on the main volume side and the high-density snapshots on the auxiliary volume side.

[0080] Creating a high-density snapshot does not require storage space on the primary and secondary volumes; it serves as a time boundary. It triggers the synchronization of data from the primary volume to the secondary volume within the same period prior to the creation of the high-density snapshot.

[0081] Step S102: Based on the (t-1)th cycle on the main volume side, generate identification information corresponding to the I / O requests issued between the (t-1)th cycle and the tth cycle.

[0082] Where t is a positive integer greater than or equal to 2.

[0083] Specifically, as described above, creating a high-density snapshot does not require the storage space of the primary and secondary volumes; instead, it serves as a time boundary. This boundary can be used to mark batches of I / O requests. For example, if the t-th cycle is the second cycle, then I / O requests issued between the first and second cycles can generate corresponding batch markers based on the cycle number (i.e., 2), indicating that the batch of I / O requests was issued before the moment the high-density snapshot is created in the second cycle.

[0084] Step S103: After the high-density snapshot created on the main volume side in the t-th cycle, create a change volume corresponding to the main volume.

[0085] Step S104: Within the modified volume, based on the identification information, extract the preset data after performing the operation corresponding to the I / O request from the preset storage location, and generate a physical volume snapshot.

[0086] The preset data contains identification information.

[0087] Step S105: Establish a second mapping relationship between the high-density snapshot created on the main volume side in the t-th period and the physical volume snapshot.

[0088] Specifically, after the high-density snapshot created in the t-th period, in order to synchronize the changed data in the primary volume from the (t-1)-th period to the secondary volume, a temporary change volume can be created on the primary volume side. A physical volume snapshot is then created within this temporary change volume. This physical volume snapshot stores preset data, which is the data stored in a preset storage location after performing operations corresponding to I / O requests. This data also carries identification information.

[0089] Therefore, the asynchronous remote replication system can retrieve the data from the preset storage location based on the identification information corresponding to the I / O requests issued between the (t-1)th cycle and the tth cycle, add it to the physical volume snapshot, and establish a second mapping relationship between the high-density snapshot created on the primary volume side in the tth cycle and the physical volume snapshot. This is used to subsequently synchronize the data in the secondary volume based on the physical volume snapshot.

[0090] Step S106: Synchronize the data on the secondary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0091] Specifically, based on the entity volume snapshot, the first mapping relationship, and the second mapping relationship, the preset data in the entity volume snapshot is synchronized to the auxiliary volume.

[0092] The data synchronization method provided in this embodiment periodically creates high-density snapshots on both the primary and secondary volumes, providing continuous data protection and reducing the risk of data loss. Furthermore, in write-intensive applications, periodic creation eliminates the need for creation operations with every data change, avoiding performance impacts caused by frequent snapshot operations. Establishing a primary mapping relationship between the high-density snapshots on the primary and secondary volumes facilitates rapid recovery of primary volume data from the secondary volume when needed. Based on the (t-1)th cycle on the primary volume, identification information corresponding to the I / O requests issued between the (t-1)th and tth cycles is generated. High-density snapshots do not occupy physical volume space; they merely batch-identify the issued I / O requests and the data after the executed operations. Subsequent creation of a modified volume corresponding to the primary volume reads the data after the executed operations corresponding to each batch of I / O requests and converts it into a physical volume snapshot. Each periodic asynchronous remote replication configures only one physical volume snapshot during the synchronization process, reducing the space requirements of periodic asynchronous replication and the time consumption for starting, stopping, and changing volume snapshots. It also improves storage space utilization and reduces configuration limitations imposed by traditional ROW snapshots. High-density snapshots eliminate the discard and flush processes, increasing resource consumption and performance impact during each remote replication process, further reducing the minimum cycle time supported by periodic asynchronous replication.

[0093] Using a high-density snapshot data synchronization process to implement periodic asynchronous remote replication can reduce the minimum period supported by periodic asynchronous remote replication and the time spent starting and stopping snapshots during the synchronization process. By creating a temporary change volume on the primary volume side, high-density snapshots are converted into physical volume snapshots, occupying only three volumes of storage space instead of the original snapshot mechanism that occupied four volumes of storage space on both the primary and secondary ROW change volumes. This greatly improves storage space utilization and significantly increases the upper limit of the number of periodic asynchronous replication configurations.

[0094] In an optional embodiment, the temporary modified volume is equivalent to the modified volume of the primary volume in the ROW (Remotely Owned) system, in order to further save storage space and reduce resource consumption in the asynchronous remote replication system. Therefore, based on the foregoing embodiments, after synchronizing data on the secondary volume according to the physical volume snapshot, the method further includes:

[0095] Delete the modified volume.

[0096] Specifically, after creating a high-density snapshot on the primary volume side during the t-th cycle, a temporary change volume corresponding to the primary volume is created. This change volume only exists during the cycle synchronization process. Its purpose is to store the data after the operations corresponding to the I / O requests are performed within the t-th cycle and to synchronize all data to the secondary volume. Since this data also carries the same identification information as the I / O requests, the change volume does not need to exist indefinitely to ensure data security. The change volume can be automatically deleted after synchronization is completed. That is, during the cycle interval at the end of the synchronization process, the cycle asynchronous replication only occupies the space of two volumes, greatly improving the utilization of storage resources and reducing storage costs.

[0097] Based on any of the foregoing embodiments, this embodiment provides a data synchronization method that can be used in the aforementioned mobile terminals, such as mobile phones and tablet computers. Figure 2 This is a flowchart illustrating the data synchronization method provided in this embodiment of the invention. Specifically, data synchronization is performed on the secondary volume based on the entity volume snapshot, the first mapping relationship, and the second mapping relationship, as follows: Figure 2 As shown, the process includes the following steps:

[0098] Step S201: Based on the physical volume snapshot and the second mapping relationship, determine the high-density snapshot of the main volume side corresponding to the physical volume snapshot in the t-th period.

[0099] Specifically, the second mapping relationship is the mapping relationship between the high-density snapshot created on the main volume side in the t-th period and the physical volume snapshot. Therefore, based on the physical volume snapshot and the second mapping relationship, the high-density snapshot on the main volume side in the t-th period corresponding to the physical volume snapshot can be determined.

[0100] Step S202: Based on the high-density snapshot of the main volume side in the t-th period and the first mapping relationship, determine the high-density snapshot of the auxiliary volume side in the t-th period.

[0101] Specifically, the first mapping relationship is the first mapping relationship between the high-density snapshot on the main volume side and the high-density snapshot on the auxiliary volume side. Therefore, based on the high-density snapshot on the main volume side in the tth period and the first mapping relationship, the high-density snapshot on the auxiliary volume side in the tth period can be determined.

[0102] Step S203: Based on the preset data in the physical volume snapshot, complete the data synchronization of the auxiliary volume.

[0103] Specifically, the preset data stored in the physical volume snapshot is the data resulting from the operation corresponding to the I / O request performed in the preset storage location. The I / O request includes the location information and operation type for the operation on the main volume data; the operation type can be either a read operation or a write operation.

[0104] For example, if the I / O request identified by the method above contains a write operation and includes a write location, the data is written to the corresponding write location in the primary volume based on the write location and write operation type. Generation identification information consistent with the I / O request is added to this data. The physical volume snapshot retrieves all data containing this generation identification information from the write location and transfers this data to the secondary volume, completing the data synchronization of the secondary volume.

[0105] Specifically, this method, based on preset data in the physical volume snapshot, can accurately determine the data that needs to be synchronized in the t-th period, thereby achieving precise data synchronization. Based on the high-density snapshot of the primary volume in the t-th period and the first mapping relationship, a high-density snapshot of the secondary volume in the t-th period is determined, completing the data synchronization of the secondary volume and ensuring data consistency between the primary and secondary volumes. Through physical volume snapshots and mapping relationships, data is flexibly managed, ensuring accurate and error-free restoration to the correct data state during data recovery, improving the accuracy of data recovery, and minimizing the risk of data loss.

[0106] Based on any of the foregoing embodiments, in an optional embodiment, this embodiment provides a data synchronization method that can be used in the aforementioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 3 This is a flowchart illustrating a data synchronization method provided in an embodiment of the present invention.

[0107] Specifically, as mentioned above, based on the (t-1)th cycle on the main volume side, identification information corresponding to the I / O requests issued between the (t-1)th cycle and the tth cycle is generated. This process includes the following steps:

[0108] Step S301: Generate the generation relationship based on the (t-1)th cycle on one side of the main volume.

[0109] Step S302: Based on the generational relationship, generate identification information corresponding to the I / O requests issued between the (t-1)th cycle and the tth cycle.

[0110] Specifically, in a concrete example, Continuous Data Protect I / O (CDP) is used to monitor in real time the I / O requests issued by the client between the (t-1)th cycle and the tth cycle. Based on the (t-1)th cycle on the main volume side, after generating the (t-1)th generation, a (t-1)th generation identifier is added to the I / O requests issued between the (t-1)th cycle and the tth cycle. The generation identifier is marked using CDP, such as CDP1, CDP2, etc.

[0111] Subsequently, based on the identified I / O requests, the location information and operation type of the data to be executed are determined. Then, according to the location information and operation type, the data corresponding to the identified I / O request is read to generate the physical volume snapshot for that period. By accurately reading and storing data based on the location information and operation type of the I / O requests, data integrity and accuracy are ensured. Storing the changed data for each period in the changed volume for physical volume snapshot generation, rather than operating directly on the primary volume, reduces the impact on primary volume I / O performance.

[0112] The overall operation flow of the method described above in this application will be illustrated below with a specific example. See below for details. Figure 4 As shown:

[0113] The default interval for creating high-density snapshots on both the primary and secondary volumes is 300 seconds. The replication relationship between primary volume A and secondary volume B is configured. Every 300 seconds, high-density snapshots cdp-ax and cdp-bx are created on primary volume A and secondary volume B, where x represents the number of high-density snapshots created on primary volume A and secondary volume B. I / O requests issued between the (t-1)th period and the tth period are marked with cdp-at-1. After the corresponding operation is performed on the I / O request marked with cdp-at-1, the data in primary volume A that has performed the corresponding operation is marked as cdp-at-1.

[0114] Create a high-density snapshot cdp-at for the primary volume in period t, and simultaneously create a high-density snapshot cdp-bt for the secondary volume in period t. Create a temporary change volume A' for volume A, and convert the high-density snapshot cdp-at into a physical volume snapshot. The conversion process involves transferring data identified as cdp-at-1 from the primary volume A to the change volume A' based on the I / O request identified as cdp-at-1. Then, synchronize all data from the physical volume snapshot to the secondary volume B corresponding to period t on the secondary volume side. After the transfer is complete, delete the change volume A', and the synchronization for this period is finished.

[0115] Based on the foregoing embodiments, in an optional embodiment, for a special scenario where multiple remote replications simultaneously enter the data synchronization process and interval period in a cluster, a method for balancing the configuration period based on remote replication service load is proposed, specifically including:

[0116] Step a1: Obtain the network transmission rate, packet loss rate, link bandwidth, number of remote replications per link, data transmission volume in each period, and the time period corresponding to the period in the cluster to which the asynchronous remote replication system belongs.

[0117] Step a2: Determine the average remote replication load of each remote replication link based on the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications per link in the cluster to which the asynchronous remote replication system belongs.

[0118] The cluster to which the remote replication system belongs includes multiple remote replication links, and each remote replication link includes multiple remote replication operations.

[0119] Step a3: Determine the remote replication load based on the data transmission volume in each cycle and the corresponding time period of the cycle.

[0120] Specifically, based on the data transmission volume in each cycle and the cycle time period T, the formula for determining the remote replication load is load_rc = data transmission volume / T.

[0121] Step a4: Adjust the time period corresponding to the cycle based on the remote replication load and the average remote replication load.

[0122] Based on the foregoing embodiments, in an optional embodiment, the remote replication load can be determined according to the data transmission volume within each cycle and the corresponding time period. However, to adjust the time period corresponding to the cycle based on the remote replication load and the average remote replication load, it is also necessary to determine the average remote replication load of each remote replication link based on the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications per link in the cluster to which the asynchronous remote replication system belongs. Specifically, this includes:

[0123] Step b1: Determine the actual quality of each remote replication link based on the network transmission rate and packet loss rate;

[0124] Specifically, based on the network transmission rate and packet loss rate, the formula for determining the actual quality of each remote replication link is LQ (Link Quality) = lb (Network Transmission Rate / 75) - Link Data Packet Loss Rate.

[0125] Step b2: Determine the optimal quality for each remote replication link based on the link bandwidth in the cluster;

[0126] Specifically, the formula for determining the optimal quality of each remote replication link based on the link bandwidth B in the cluster is LQ_max = lb(B / 75).

[0127] Step b3: Determine the bandwidth load for each remote replication link based on the actual quality and the optimal quality;

[0128] Specifically, based on the actual quality LQ and the optimal quality LQ_max, the formula for determining the bandwidth load of each remote replication link is load = (LQ / LQ_max) * B.

[0129] Step b4: Determine the average remote replication load for each remote replication link based on the bandwidth load and the number of remote replications for each remote replication link.

[0130] Specifically, based on the bandwidth load of each remote replication link (load) and the number of remote replications (N), the formula for determining the average remote replication load of each remote replication link is load_rc_average = load / N.

[0131] Based on the foregoing embodiments, it is known that the period corresponds to a preset time period. In an optional embodiment, the time period corresponding to the period is adjusted according to the remote replication load and the average remote replication load, specifically including:

[0132] When the remote replication load is greater than the remote replication average load, the time period corresponding to the cycle will be reduced to the first preset multiple of the preset time period;

[0133] Specifically, when the remote replication load (load_rc) is greater than the remote replication average load (load_rc_average), the time period corresponding to the cycle is reduced to a preset multiple of the preset time period. For example, it could be T / 2.

[0134] Alternatively, when the remote replication load is less than the average remote replication load, the time period corresponding to the cycle is amplified to a second preset multiple of the preset time period.

[0135] The first preset multiple is less than the second preset multiple.

[0136] Specifically, when the remote replication load (load_rc) is less than the remote replication average load (load_rc_average), the time period corresponding to the cycle is increased to a preset multiple of the preset time period. For example, it could be 2T.

[0137] It should be noted that when adjusting the actual period of remote replication in the balanced period mode, the adjustment needs to be made within the remote replication period interval, and the period needs to be recalculated after the adjustment. In addition, if the adjusted period is not within the range of the system's supported periods, no adjustment operation should be performed.

[0138] Specifically, this method can adjust the periodic asynchronous remote replication cycle in the system according to the business load. When the business load is high, the cycle is shortened, the amount of data synchronized in each synchronization process is reduced, and the synchronization rate is improved. When the business load is low or there is no business, the periodic asynchronous remote replication cycle is lengthened, which reduces the space occupation time of temporary volume changes during the synchronization process and improves storage space utilization.

[0139] Building upon the aforementioned embodiments, CDP technology is used to monitor I / O requests in real time, identifying I / O requests across consecutive cycles. Based on the identified I / O requests in each cycle, the corresponding data for each I / O request is executed, identifying and storing only the actual changed portion of the data, rather than the entire dataset. In an optional embodiment, before generating a physical volume snapshot from the identified data, the changed data can be compressed in real time to reduce storage requirements and network bandwidth. Specifically, this includes:

[0140] Step c1: Compress the data corresponding to the I / O request.

[0141] Step c2: During the data compression process, generational relationships and corresponding identification information are generated for each compression interval.

[0142] Specifically, the generational relationship and corresponding identification information for each compression interval can be consistent with the I / O request identifier, such as CDP1, CDP2, etc.

[0143] Step c3: Generate a physical volume snapshot based on the compressed data and identification information.

[0144] Because the data has been compressed, physical volume snapshots take up less storage space, reducing storage requirements and network bandwidth.

[0145] This embodiment also provides a data synchronization device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0146] This embodiment provides a data synchronization device, such as... Figure 5 The module includes: a creation module 501, a processing module 502, an extraction module 503, and a synchronization module 504.

[0147] Create module 501 to periodically create high-density snapshots on both the main volume and the secondary volume side;

[0148] Processing module 502 is used to establish a first mapping relationship between high-density snapshots on the main volume side and high-density snapshots on the auxiliary volume side; and to generate identification information corresponding to the I / O requests issued between the (t-1)th and tth cycles on the main volume side, where t is a positive integer greater than or equal to 2.

[0149] The creation module 501 is also used to create a change volume corresponding to the main volume after a high-density snapshot is created on the main volume side in the t-th period;

[0150] The extraction module 503 is used to extract preset data from a preset storage location within the modified volume after performing the operation corresponding to the I / O request, based on the identification information. The preset data carries the identification information.

[0151] The processing module 502 is also used to generate a physical volume snapshot based on preset data; and to establish a second mapping relationship between the high-density snapshot created on the main volume side in the t-th period and the physical volume snapshot.

[0152] Synchronization module 504 is used to synchronize data on the secondary volume based on the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0153] In an optional embodiment, the processing module 502 is further configured to delete the changed volume after synchronizing the data of the secondary volume according to the entity volume snapshot, the first mapping relationship, and the second mapping relationship.

[0154] In an optional embodiment, the processing module 502 is specifically configured to determine, based on the physical volume snapshot and the second mapping relationship, the high-density snapshot of the main volume side corresponding to the physical volume snapshot in the t-th period.

[0155] Based on the high-density snapshot of the main volume in the t-th period and the first mapping relationship, determine the high-density snapshot of the auxiliary volume in the t-th period;

[0156] Based on the preset data in the physical volume snapshot, complete the data synchronization of the auxiliary volume.

[0157] In an optional embodiment, the processing module 502 is specifically used to generate a generation relationship based on the (t-1)th cycle on the main volume side;

[0158] Based on the generational relationship, generate identification information corresponding to the I / O requests issued between the (t-1)th cycle and the tth cycle.

[0159] In an optional embodiment, the extraction module 503 is specifically used to obtain the network transmission rate, packet loss rate, link bandwidth, number of remote replications per link, data transmission volume in each period, and the time period corresponding to the period in the cluster to which the asynchronous remote replication system belongs.

[0160] The processing module 502 is specifically used to determine the average remote replication load of each remote replication link based on the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications of a single link in the cluster to which the asynchronous remote replication system belongs. The cluster to which the remote replication system belongs includes multiple remote replication links, and each remote replication link includes multiple remote replication operations.

[0161] The remote replication load is determined based on the data transfer volume within each cycle and the corresponding time period of the cycle;

[0162] Adjust the time period corresponding to the cycle based on the remote replication load and the average remote replication load.

[0163] In an optional embodiment, the processing module 502 is specifically configured to determine the actual quality of each remote replication link based on the network transmission rate and packet loss rate.

[0164] Determine the optimal quality for each remote replication link based on the link bandwidth in the cluster;

[0165] Determine the bandwidth load for each remote replication link based on actual and optimal quality.

[0166] The average remote replication load for each remote replication link is determined based on the bandwidth load and the number of remote replications for each link.

[0167] In an optional embodiment, the period corresponds to a preset time period. The processing module 502 is specifically used to reduce the time period corresponding to the period to a first preset multiple of the preset time period when the remote replication load is greater than the remote replication average load.

[0168] or,

[0169] When the remote replication load is less than the remote replication average load, the time period corresponding to the cycle is amplified to a second preset multiple of the preset time period, wherein the first preset multiple is less than the second preset multiple.

[0170] In this embodiment, the data synchronization device is presented in the form of a functional module. Here, a module refers to an application-specific integrated circuit (ASIC), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0171] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0172] This invention provides a data synchronization device that achieves data synchronization from a primary volume to a secondary volume through the collaborative work of modules. High-density snapshots are periodically created on both the primary and secondary volumes, rather than being created manually or by events, providing continuous data protection and reducing the risk of data loss. Furthermore, in write-intensive applications, periodic creation avoids the need for creation operations with every data change, preventing performance impact from frequent snapshot operations. Establishing a first mapping relationship between high-density snapshots on the primary and secondary volumes facilitates rapid recovery of primary volume data from the secondary volume when needed. Based on the (t-1)th cycle on the primary volume, identification information corresponding to I / O requests issued between the (t-1)th and tth cycles is generated. High-density snapshots do not occupy physical volume space; they only batch-identify the issued I / O requests and the data after the executed operations. Subsequent creation of a modified volume corresponding to the primary volume reads the data after the executed operations corresponding to each batch of I / O requests and converts it into a physical volume snapshot. Each periodic asynchronous remote replication configures only one physical volume snapshot during the synchronization process, reducing the space requirements of periodic asynchronous replication and the time consumption for starting, stopping, and changing volume snapshots. It also improves storage space utilization and reduces configuration limitations imposed by traditional ROW snapshots. High-density snapshots eliminate the discard and flush processes, increasing resource consumption and performance impact during each remote replication process, further reducing the minimum cycle time supported by periodic asynchronous replication.

[0173] This invention also provides a computer device. Figure 6 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 6 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 6 Take a processor 10 as an example.

[0174] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include an integrated circuit. The integrated circuit may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPRS), or any combination thereof.

[0175] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0176] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0177] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0178] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0179] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0180] This invention also provides a computer-readable storage medium. The methods provided in the above embodiments can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0181] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0182] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method of data synchronization, the method comprising: The method is applied to an asynchronous remote replication system comprising a primary volume and a secondary volume, and the method comprises: periodically creating high-density snapshots on the primary volume side and the secondary volume side respectively, and establishing a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side; generating, according to the (t-1)th period of the primary volume side, identification information corresponding to the I / O request issued between the (t-1)th period and the tth period, t being a positive integer greater than or equal to 2; after the high-density snapshot created on the tth period of the primary volume side, creating a change volume corresponding to the primary volume; in the change volume, extracting, according to the identification information, preset data after performing an operation corresponding to the I / O request from a preset storage location, and generating an entity volume snapshot, wherein the preset data carries the identification information; establishing a second mapping relationship between the high-density snapshot created on the tth period of the primary volume side and the entity volume snapshot; determining, according to the entity volume snapshot and the second mapping relationship, the high-density snapshot on the tth period of the primary volume side corresponding to the entity volume snapshot; determining, according to the high-density snapshot on the tth period of the primary volume side and the first mapping relationship, the high-density snapshot on the tth period of the secondary volume side; completing data synchronization of the secondary volume according to the preset data in the entity volume snapshot.

2. The method of claim 1, wherein, After the data synchronization of the secondary volume according to the entity volume snapshot, the first mapping relationship and the second mapping relationship, the method further comprises: deleting the change volume.

3. The method according to claim 1 or 2, characterized in that, The generation of the identification information corresponding to the I / O request issued between the (t-1)th period and the tth period according to the (t-1)th period of the primary volume side specifically comprises: generating a generation relationship according to the (t-1)th period of the primary volume side; generating the identification information corresponding to the I / O request issued between the (t-1)th period and the tth period according to the generation relationship.

4. The method of claim 1, wherein, The method further comprises: obtaining network transmission rate, packet loss rate, link bandwidth, number of remote replications of a single link, data transmission amount in each period and time period corresponding to the period in a cluster to which the asynchronous remote replication system belongs; determining remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth and number of remote replications of a single link in the cluster to which the asynchronous remote replication system belongs, wherein the cluster to which the remote replication system belongs comprises a plurality of remote replication links, and each remote replication link comprises a plurality of remote replication operations; determining remote replication load according to the data transmission amount in each period and the time period corresponding to the period; adjusting the time period corresponding to the period according to the remote replication load and the remote replication average load.

5. The method of claim 4, wherein, The determination of the remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth and number of remote replications of a single link in the cluster to which the asynchronous remote replication system belongs specifically comprises: determining actual quality of each remote replication link according to the network transmission rate and the packet loss rate; determining optimal quality of each remote replication link according to link bandwidth in the cluster; determining bandwidth load of each remote replication link according to the actual quality and the optimal quality; determining remote replication average load of each remote replication link according to the bandwidth load of each remote replication link and remote replication quantity.

6. The method according to claim 4 or 5, characterized in that, The period corresponds to a preset time period. The adjusting the time period corresponding to the period according to the remote replication load and the remote replication average load specifically comprises: when the remote replication load is greater than the remote replication average load, reducing the time period corresponding to the period to a first preset multiple of the preset time period; or, when the remote replication load is less than the remote replication average load, enlarging the time period corresponding to the period to a second preset multiple of the preset time period, wherein the first preset multiple is less than the second preset multiple.

7. A data synchronization apparatus, characterized by comprising: The device comprises: a creating module configured to periodically create high-density snapshots on a primary volume side and a secondary volume side respectively; a processing module configured to establish a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side, and generate identification information corresponding to I / O requests issued between a t-1th period and a tth period on the primary volume side, t being a positive integer greater than or equal to 2; the creating module is further configured to create a change volume corresponding to the primary volume after the high-density snapshot created on the primary volume side in the tth period; an extracting module configured to extract preset data after performing operations corresponding to the I / O requests from a preset storage location in the change volume according to the identification information, wherein the preset data carries the identification information; the processing module is further configured to generate an entity volume snapshot according to the preset data, and establish a second mapping relationship between the high-density snapshot created on the primary volume side in the tth period and the entity volume snapshot; a synchronizing module configured to determine the high-density snapshot of the primary volume side in the tth period corresponding to the entity volume snapshot according to the entity volume snapshot and the second mapping relationship, determine the high-density snapshot of the secondary volume side in the tth period according to the high-density snapshot of the primary volume side in the tth period and the first mapping relationship, and complete data synchronization of the secondary volume according to the preset data in the entity volume snapshot.

8. A computer device, comprising: comprise: a memory and a processor, which are communicatively connected with each other, and the memory stores computer instructions, and the processor executes the computer instructions to perform the data synchronization method in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to make a computer execute the data synchronization method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • User snapshot synchronization method and device based on periodic asynchronous remote replication relationship

    CN118170716A

  • Data volume remote copying method and device, computer equipment and storage medium

    CN118708130A