Data synchronization method and device, equipment and medium

By periodically creating high-density snapshots and establishing mapping relationships, the problem that the existing technology cannot achieve asynchronous replication of seconds is solved, and more efficient storage space utilization and shorter replication cycles are achieved.

CN119988489AActive Publication Date: 2025-05-13INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510062653.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Existing remote replication technologies cannot implement short-period, multiple start-stop snapshots such as second-level cycles, asynchronous replication, and ROW snapshots occupy a large amount of storage space, limiting the number of replication relationships.

Method used

The method of periodically creating high-density snapshots is adopted to establish a mapping relationship between the main volume and the secondary volume. By generating identification information and creating a change volume, the data corresponding to the I/O request is converted into a physical volume snapshot to achieve data synchronization.

Benefits of technology

It reduces the space requirement and start-stop time consumption of periodic asynchronous replication, improves storage space utilization and configuration limits, supports shorter replication cycles, and reduces the risk of data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988489A_ABST
    Figure CN119988489A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a data synchronization method, device, equipment and medium, and the method comprises the steps: periodically and respectively creating high-density snapshots on one side of a main volume and one side of an auxiliary volume, establishing a first mapping relation between the high-density snapshots on one side of the main volume and the high-density snapshots on one side of the auxiliary volume, and according to the (t-1) th period of one side of the main volume, establishing a second mapping relation between the high-density snapshots on one side of the main volume and the high-density snapshots on one side of the auxiliary volume; the method comprises the steps of generating identification information corresponding to an I / O request issued between a (t-1) th period and a tth period, creating a change volume corresponding to a main volume after a high-density snapshot created on one side of the main volume in the tth period, extracting preset data after executing operation corresponding to the I / O request from a preset storage position according to the identification information in the change volume, and sending the extracted preset data to the tth period. And establishing a first mapping relationship between the high-density snapshot created on one side of the main volume in the t-th period and the entity volume snapshot, generating an entity volume snapshot, establishing a second mapping relationship between the high-density snapshot created on one side of the main volume in the t-th period and the entity volume snapshot, and performing data synchronization on the auxiliary volume according to the entity volume snapshot, the first mapping relationship and the second mapping relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data synchronization method, device, equipment and medium. Background Art

[0002] With the rapid development of information technology, the demand for data protection is increasing. Traditional snapshot technologies, such as Copy-On-Write (COW) and Redirect-On-Write (ROW), are used to protect data.

[0003] The current remote replication change volume snapshot method uses ROW snapshots. Each time a ROW snapshot is started, the data on the target volume needs to be discarded (checked and removed from data rows that have not changed) and flushed (the process of writing data in memory to disk to ensure that all data changes are persisted to the storage device). The above process brings additional overhead and cannot support complex scenarios such as asynchronous replication with a short period of seconds and multiple snapshot starts and stops.

[0004] In addition, ROW snapshots occupy a large amount of storage space. One cycle of asynchronous remote replication consists of four volumes and requires four ROW snapshots. If the upper limit of the number of volumes that a cluster can support is 10,000, it can only support a maximum of 2,500 cycles of asynchronous remote replication, which limits the number of replication relationships. Summary of the invention

[0005] In view of this, the present invention provides a data synchronization method, device, equipment and medium to solve the problem that the existing remote replication technology cannot achieve short-cycle, multiple start-stop snapshots such as second-level cycle and asynchronous replication.

[0006] In a first aspect, the present invention provides a data synchronization method, which is applied to an asynchronous remote replication system, the asynchronous remote replication system including a primary volume and a secondary volume, and the method includes:

[0007] Periodically creating high-density snapshots on the primary volume side and the secondary volume side, respectively, and establishing a first mapping relationship between the high-density snapshots on the primary volume side and the high-density snapshots on the secondary volume side;

[0008] Generate identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle according to the t-1th cycle on the primary volume side, where t is a positive integer greater than or equal to 2;

[0009] After the high-density snapshot created in the t-th period on the master volume side, creating a change volume corresponding to the master volume;

[0010] In the changed volume, according to the identification information, preset data after executing the operation corresponding to the I / O request is extracted from a preset storage location, and a physical volume snapshot is generated, wherein the preset data carries the identification information;

[0011] Establishing a second mapping relationship between the high-density snapshot created on the master volume side in the t-th period and the physical volume snapshot;

[0012] Data of the auxiliary volume is synchronized according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0013] A data synchronization method provided by the present invention has the following advantages: high-density snapshots are periodically created on one side of the primary volume and one side of the secondary volume, rather than being triggered by humans or events to create high-density snapshots, providing continuous data protection and reducing the risk of data loss. In addition, in write-intensive applications, periodic creation does not require a creation operation every time the data changes, which can avoid affecting system performance due to frequent snapshot operations. Establishing a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side helps to quickly restore the data of the primary volume from the secondary volume when needed. According to the t-1th cycle on the primary volume side, identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle is generated. The high-density snapshot does not occupy the physical volume space, but only batches the issued I / O requests and the data after the operation has been executed. The subsequent creation of the change volume corresponding to the primary volume reads the data after the operation corresponding to each batch of I / O requests and converts it into a physical volume snapshot. Each periodic asynchronous remote replication only configures one physical volume snapshot during the synchronization process, reducing the space requirements of periodic asynchronous replication and the time consumption of starting and stopping auxiliary volumes to change volume snapshots. At the same time, it improves the storage space utilization and configuration quantity restrictions of periodic asynchronous replication due to traditional ROW snapshot restrictions. High-density snapshots abandon the discard and flush processes, improve resource consumption and performance impact during each remote replication process, and further reduce the minimum period supported by periodic asynchronous replication.

[0014] Using the data synchronization process of high-density snapshots to implement periodic asynchronous remote replication can reduce the minimum period supported by periodic asynchronous remote replication and the time spent on starting and stopping snapshots during the synchronization process. By creating a temporary change volume on the primary volume side, high-density snapshots can be converted to physical volume snapshots, which only occupies the storage space of three volumes instead of the original snapshot mechanism in which the change volumes on both the primary and secondary sides of ROW occupy the storage space of four volumes. This greatly improves storage space utilization and significantly increases the upper limit of the configuration quantity of periodic asynchronous replication.

[0015] In an optional implementation, after synchronizing data on the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship, the method further includes:

[0016] Delete the change volume.

[0017] Specifically, after creating a high-density snapshot in the tth cycle on the primary volume side, a temporary change volume corresponding to the primary volume is created. This change volume only exists during the periodic synchronization process, and its function is to store the data after the operation corresponding to the I / O request is executed in the tth cycle, and synchronize all the data to the secondary side. Since the data also carries the identification information consistent with the I / O request, there is no need for the change volume to exist all the time to ensure data security. After the synchronization is completed, the change volume can be automatically deleted, that is, in the periodic interval at the end of the synchronization process, the periodic asynchronous replication only occupies the space of two volumes, which improves the utilization of storage resources and reduces storage costs.

[0018] In an optional implementation, synchronizing data on the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship specifically includes:

[0019] According to the physical volume snapshot and the second mapping relationship, determine the high-density snapshot of the primary volume corresponding to the physical volume snapshot in the tth period;

[0020] Determine the high-density snapshot of the auxiliary volume side in the tth period according to the high-density snapshot of the primary volume side in the tth period and the first mapping relationship;

[0021] Complete data synchronization of the auxiliary volume based on the preset data in the physical volume snapshot.

[0022] Specifically, based on the preset data in the physical volume snapshot, the data that needs to be synchronized in the tth cycle can be accurately determined, thereby achieving accurate data synchronization. Based on the high-density snapshot of the primary volume in the tth cycle and the first mapping relationship, the high-density snapshot of the auxiliary volume in the tth cycle is determined to complete the data synchronization of the auxiliary volume, ensuring the data consistency between the primary and auxiliary volumes. Through the physical volume snapshot and mapping relationship, data can be flexibly managed to ensure that the correct data state can be accurately restored during the data recovery process, improve the accuracy of data recovery, and minimize the risk of data loss.

[0023] In an optional implementation, according to the t-1th cycle on the primary volume side, generating identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle specifically includes:

[0024] Generate a generation relationship based on the t-1th cycle on the main volume side;

[0025] According to the generation relationship, identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle is generated.

[0026] Specifically, according to the t-1th cycle on the master volume side, a generation relationship is generated, and according to the generation relationship, identification information corresponding to the I / O requests issued between the t-1th cycle and the tth cycle is generated. By accurately capturing and identifying the I / O requests under each generation, the system can more effectively manage and schedule each batch of data, ensuring the integrity and accuracy of the data. And adding identification information to the I / O request according to the generation relationship helps to accurately roll back to the data of the required generation according to the generation relationship. Subsequently, according to the identification information corresponding to the I / O request, the data on which the corresponding operation has been performed is also attached with the same identification information. When the periodic replication is started, the change data of each batch is stored in the change volume to perform the physical volume snapshot generation operation, rather than directly operating on the master volume, which can reduce the impact on the master volume I / O performance.

[0027] In an optional embodiment, the method further includes:

[0028] Obtain the network transmission rate, packet loss rate, link bandwidth, number of remote replications on a single link, data transmission volume in each cycle, and the time period corresponding to the cycle in the cluster to which the asynchronous remote replication system belongs;

[0029] Determine the remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth, and remote replication quantity of a single link in the cluster to which the asynchronous remote replication system belongs, wherein the cluster to which the remote replication system belongs includes multiple remote replication links, and each remote replication link includes multiple remote replication operations;

[0030] Determine the remote replication load based on the data transmission volume in each cycle and the time period corresponding to the cycle;

[0031] Adjust the time period corresponding to the cycle according to the remote replication load and the remote replication average load.

[0032] Specifically, by obtaining key performance indicators such as network transmission rate, packet loss rate, link bandwidth, number of remote replications on a single link, data transmission volume in each cycle, and time period corresponding to the cycle, the performance of the asynchronous remote replication system can be monitored and optimized in real time, and problems that may affect data replication efficiency and reliability can be discovered and resolved in a timely manner.

[0033] By calculating the average remote replication load of each remote replication link and the actual remote replication load of each link, and adjusting the time period corresponding to the cycle according to the remote replication load and the average remote replication load, resources can be allocated more reasonably, load balancing can be achieved, the efficiency and stability of the entire system can be improved, and performance bottlenecks caused by uneven load can be avoided.

[0034] In an optional implementation, the remote replication average load of each remote replication link is determined according to the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications of a single link in the cluster to which the asynchronous remote replication system belongs, specifically including:

[0035] Determine the actual quality of each remote replication link based on the network transmission rate and packet loss rate;

[0036] Determine the optimal quality of each remote replication link based on the link bandwidth in the cluster;

[0037] Determine the bandwidth load of each remote replication link based on the actual quality and the optimal quality;

[0038] The average load of remote replication on each remote replication link is determined based on the bandwidth load of each remote replication link and the number of remote replications.

[0039] Specifically, the actual quality of each remote replication link is determined based on the network transmission rate and packet loss rate, and the optimal quality of each remote replication link is determined based on the link bandwidth in the cluster. Adjusting the actual quality by the optimal quality can more reasonably allocate network resources and improve overall network efficiency. By determining the bandwidth load of each remote replication link, the load of each link can be balanced to avoid overloading of some links and affecting replication efficiency. Combining the bandwidth load and the number of remote replications, the average load of remote replication for each link is determined, which helps to adjust the cycle according to the business load. The cycle with large business load is reduced, the amount of synchronized data in each synchronization process is reduced, and the synchronization rate is improved; the asynchronous remote replication cycle is increased in the cycle with little or no business, which reduces the space occupied by temporary changes to the volume during the synchronization process and improves the storage space utilization.

[0040] In an optional embodiment, the period corresponds to a preset time period;

[0041] The adjusting the time period corresponding to the cycle according to the remote replication load and the remote replication average load specifically includes:

[0042] When the remote replication load is greater than the remote replication average load, shortening the time period corresponding to the cycle to a first preset multiple of the preset time period;

[0043] or,

[0044] When the remote replication load is less than the remote replication average load, the time period corresponding to the cycle is enlarged to a second preset multiple of the preset time period, wherein the first preset multiple is less than the second preset multiple.

[0045] Specifically, the method can adjust the cycle of periodic asynchronous remote replication in the system according to the business load. The cycle when the business load is large can be shortened, the amount of synchronized data in each synchronization process can be reduced, and the synchronization rate can be improved. The cycle of periodic asynchronous remote replication when the business is small or even no business can be increased to reduce the space occupied by temporary changes to the volume during the synchronization process and improve the storage space utilization rate.

[0046] In a second aspect, the present invention provides a data synchronization device, the device comprising:

[0047] The creation module is used to periodically create high-density snapshots on the primary volume side and the secondary volume side;

[0048] The processing module is used to establish a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side; and generate identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle according to the t-1th cycle on the primary volume side, where t is a positive integer greater than or equal to 2;

[0049] The creation module is also used to create a change volume corresponding to the primary volume after the high-density snapshot created in the tth period on the primary volume side;

[0050] An extraction module, used to extract, in the changed volume, preset data after executing an operation corresponding to the I / O request from a preset storage location according to the identification information, wherein the preset data carries the identification information;

[0051] The processing module is further used to generate a physical volume snapshot according to preset data; establish a second mapping relationship between the high-density snapshot created in the tth period on the master volume side and the physical volume snapshot;

[0052] The synchronization module is used to synchronize data of the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0053] A data synchronization device provided by the present invention has the following advantages: asynchronous remote replication based on high-density snapshots is realized through the collaborative work of various modules. High-density snapshots are periodically created on the primary volume side and the auxiliary volume side, rather than being manually or event-triggered to create high-density snapshots, to provide continuous data protection and reduce the risk of data loss. A first mapping relationship is established between the high-density snapshot on the primary volume side and the high-density snapshot on the auxiliary volume side, which helps to quickly restore the data of the primary volume from the auxiliary volume when needed. The high-density snapshot does not occupy the physical volume space, but only batches the issued I / O requests and the data after the operation has been executed. The subsequent creation of the change volume corresponding to the primary volume reads the above-mentioned data after the operation corresponding to each batch of I / O requests and converts it into a physical volume snapshot. Each periodic asynchronous remote replication only configures one physical volume snapshot during the synchronization process, which reduces the space requirement of the periodic asynchronous replication and the time consumption of starting and stopping the auxiliary volume change volume snapshot, while improving the storage space utilization and configuration quantity limit of the periodic asynchronous replication due to the traditional ROW snapshot limit. High-density snapshots eliminate the discard and flush processes, improve resource consumption and performance impact during each remote replication process, and further reduce the minimum period supported by periodic asynchronous replication.

[0054] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data synchronization method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0055] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data synchronization method of the first aspect or any corresponding embodiment thereof.

[0056] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the data synchronization method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0058] Figure 1 It is a flowchart of a data synchronization method provided by an embodiment of the present invention;

[0059] Figure 2 is a flow chart of another data synchronization method provided by an embodiment of the present invention;

[0060] Figure 3 is a flow chart of another data synchronization method provided by an embodiment of the present invention;

[0061] Figure 4 is a flowchart of another data synchronization method provided by an embodiment of the present invention;

[0062] Figure 5 is a structural block diagram of a data synchronization device provided by an embodiment of the present invention;

[0063] Figure 6 It is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0065] With the rapid development of information technology, the demand for data protection is increasing. Although traditional snapshot technologies, such as Copy-On-Write (COW) and Redirect-On-Write (ROW), have achieved data backup and recovery to a certain extent, they have problems such as large storage space occupation and low recovery efficiency. Especially in critical applications, users have an increasing demand for continuous data protection (CDP), and traditional snapshot technologies can no longer meet this demand. To this end, the industry has developed high-density snapshot technology to provide users with continuous data protection functions through intensive data protection cycles and efficient storage space utilization. However, the configuration method of existing high-density snapshot technology is relatively complex, and its flexibility and configurability in different application scenarios need to be improved. For example, in disaster recovery scenarios such as periodic asynchronous remote replication, active-active remote replication, and 3DC, the difference between high-density snapshots and traditional snapshots in bitmaps makes it impossible to use them flexibly.

[0066] The current snapshot is applied to remote replication in the following scenarios:

[0067] Periodic asynchronous remote replication and active-active change volume snapshots. Remote replication will periodically start and stop change volume snapshots to achieve direct information synchronization between the primary and secondary ends of remote replication. The key to achieving the above functions is that the snapshot bitmap records the data difference between the source volume and the change volume, and synchronizes the difference data to the secondary end every cycle. In addition, the primary and secondary change volumes store the primary and secondary end data of the previous cycle respectively, and the reverse change volume snapshot can be started to restore data to achieve disaster recovery. High-density snapshots are limited in application in remote replication scenarios because the data protection cycle is too intensive and the memory space does not support the configuration bitmap.

[0068] Specifically, the current remote replication change volume snapshot method uses ROW snapshots. In each cycle, remote replication performs snapshot-related operations according to the following steps.

[0069] 1) The master starts the snapshot of the master change volume.

[0070] 2) Start from the client to the snapshot of the secondary change volume.

[0071] 3) Transmit the changed volume data on the primary side to the secondary side.

[0072] 4) Stop the primary side from changing the volume snapshot.

[0073] 5) Stop changing the volume snapshot from the slave.

[0074] In the above process, the periodic start and stop of ROW snapshots are used to achieve periodic synchronization and protection of data. When a failure occurs on the primary or secondary side, resulting in data loss, a reverse snapshot of the change volume to the source volume can be started to restore the source volume data to the most recent period to protect data security.

[0075] The above ROW snapshot needs to discard and flush the data on the target volume every time it is started. It cannot support complex scenarios such as asynchronous replication with a short cycle of seconds and multiple snapshot starts and stops. In addition, ROW snapshots occupy a large storage space. One cycle of asynchronous remote replication consists of four volumes and requires four ROW snapshots. If the upper limit of the number of volumes that a cluster can support is 10,000, it can only support a maximum of 2,500 cycles of asynchronous remote replication, which limits the number of replication relationships.

[0076] To solve the above problems, an embodiment of the present invention provides a data synchronization embodiment. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system (computer device) including a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0077] In this embodiment, a data synchronization method is provided, which can be used for the above-mentioned terminal equipment, such as a mobile phone, a tablet computer, etc. Figure 1 is a flow chart of a data synchronization method provided by an embodiment of the present invention, such as Figure 1 As shown, the method is applied to an asynchronous remote replication system, the asynchronous remote replication system includes a primary volume and a secondary volume, and the method flow includes the following steps:

[0078] Step S101 : periodically create high-density snapshots on the primary volume side and the secondary volume side, respectively, and establish a first mapping relationship between the high-density snapshots on the primary volume side and the high-density snapshots on the secondary volume side.

[0079] Specifically, the client presets a period for creating high-density snapshots according to actual scenario requirements. The preset period can be a specific time interval in seconds. The primary volume side and the auxiliary volume side respectively create corresponding high-density snapshots according to the preset period, wherein the high-density snapshot created on the primary volume side corresponds to the high-density snapshot created on the auxiliary volume side one-to-one, thereby establishing a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the auxiliary volume side.

[0080] Creating a high-density snapshot does not occupy the storage space of the primary and secondary volumes, and is used as a time boundary. It is used to trigger synchronization of data in the same cycle on the primary volume to the secondary volume before the high-density snapshot is created.

[0081] Step S102: Generate identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle according to the t-1th cycle on the primary volume side.

[0082] Wherein, t is a positive integer greater than or equal to 2.

[0083] Specifically, as described above, creating a high-density snapshot does not require the storage space of the primary and secondary volumes, and is used as a time boundary. This boundary can be used to mark the batch of I / O requests issued. For example, if the tth cycle is the second cycle, then the I / O requests issued between the first cycle and the second cycle can generate a corresponding batch mark based on the number of cycles (i.e., 2), which is used to indicate that the batch of I / O requests is the I / O request issued before the moment when the high-density snapshot is created in the second cycle.

[0084] Step S103: After the high-density snapshot is created in the tth period on the primary volume side, a change volume corresponding to the primary volume is created.

[0085] Step S104: in the changed volume, according to the identification information, extracting preset data after executing the operation corresponding to the I / O request from the preset storage location, and generating a physical volume snapshot.

[0086] The preset data carries identification information.

[0087] Step S105: establishing a second mapping relationship between the high-density snapshot created at the primary volume side in the tth period and the physical volume snapshot.

[0088] Specifically, after the high-density snapshot is created in the tth cycle, in order to synchronize the changed data in the primary volume from the t-1th cycle to the tth cycle to the secondary volume, a temporary change volume can be created on the primary volume side. A physical volume snapshot is created in the temporary change volume. The physical volume snapshot stores preset data, which is data stored in a preset storage location after the operation corresponding to the I / O request is executed. The data also carries identification information.

[0089] Therefore, the asynchronous remote replication system can extract the data from the preset storage location according to the identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle, add it to the physical volume snapshot, and establish a second mapping relationship between the high-density snapshot created in the tth cycle on the primary volume side and the physical volume snapshot. This is used to complete the data synchronization in the auxiliary volume according to the physical volume snapshot later.

[0090] Step S106: Synchronize data on the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0091] Specifically, according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship, the preset data in the physical volume snapshot is synchronized to the auxiliary volume.

[0092] The data synchronization method provided in this embodiment periodically creates high-density snapshots on the primary volume side and the secondary volume side, providing continuous data protection and reducing the risk of data loss. In write-intensive applications, periodic creation does not require a creation operation every time the data changes, which can avoid affecting system performance due to frequent snapshot operations. Establishing a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side helps to quickly restore the data of the primary volume from the secondary volume when needed. According to the t-1th cycle on the primary volume side, identification information corresponding to the I / O requests issued between the t-1th cycle and the tth cycle is generated. The high-density snapshot does not occupy the physical volume space, but only batches the issued I / O requests and the data after the operation has been executed. The subsequent creation of the change volume corresponding to the primary volume reads the above-mentioned data after the operation corresponding to each batch of I / O requests and converts it into a physical volume snapshot. Each periodic asynchronous remote replication only configures one physical volume snapshot during the synchronization process, reducing the space requirements of periodic asynchronous replication and the time consumption of starting and stopping auxiliary volumes to change volume snapshots. At the same time, it improves the storage space utilization and configuration quantity restrictions of periodic asynchronous replication due to traditional ROW snapshot restrictions. High-density snapshots abandon the discard and flush processes, improve resource consumption and performance impact during each remote replication process, and further reduce the minimum period supported by periodic asynchronous replication.

[0093] Using the data synchronization process of high-density snapshots to implement periodic asynchronous remote replication can reduce the minimum period supported by periodic asynchronous remote replication and the time spent on starting and stopping snapshots during the synchronization process. By creating a temporary change volume on the primary volume side, high-density snapshots can be converted to physical volume snapshots, which only occupies the storage space of three volumes instead of the original snapshot mechanism in which the change volumes on both the primary and secondary sides of ROW occupy the storage space of four volumes. This greatly improves storage space utilization and significantly increases the upper limit of the configuration quantity of periodic asynchronous replication.

[0094] In an optional embodiment, the temporary change volume is equivalent to the change volume of the primary volume in the ROW, in order to further save storage space in the asynchronous remote replication system and reduce resource consumption. Therefore, based on the above embodiment, after synchronizing the data of the auxiliary volume according to the physical volume snapshot, the method further includes:

[0095] Delete the change volume.

[0096] Specifically, after creating a high-density snapshot in the tth cycle on the primary volume side, a temporary change volume corresponding to the primary volume is created. This change volume only exists during the periodic synchronization process, and its function is to store the data after the operation corresponding to the I / O request is executed in the tth cycle, and synchronize all the data to the secondary side. Since the data also carries the identification information consistent with the I / O request, there is no need for the change volume to exist all the time to ensure data security. After the synchronization is completed, the change volume can be automatically deleted, that is, in the periodic interval at the end of the synchronization process, the periodic asynchronous replication only occupies the space of two volumes, which greatly improves the utilization rate of storage resources and reduces storage costs.

[0097] Based on any of the above embodiments, this embodiment provides a data synchronization method, which can be used in the above mobile terminals, such as mobile phones, tablet computers, etc. Figure 2 is a flow chart of a data synchronization method provided by an embodiment of the present invention. Specifically, according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship, data synchronization is performed on the auxiliary volume, such as Figure 2 As shown, the process includes the following steps:

[0098] Step S201: Determine, according to the physical volume snapshot and the second mapping relationship, a high-density snapshot of the primary volume corresponding to the physical volume snapshot in the tth period.

[0099] Specifically, the second mapping relationship is a mapping relationship between the high-density snapshot created on the primary volume side in the tth period and the physical volume snapshot, so according to the physical volume snapshot and the second mapping relationship, the high-density snapshot on the primary volume side in the tth period corresponding to the physical volume snapshot can be determined.

[0100] Step S202: determining the high-density snapshot of the auxiliary volume in the tth period according to the high-density snapshot of the primary volume in the tth period and the first mapping relationship.

[0101] Specifically, the first mapping relationship is a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the auxiliary volume side. Therefore, the high-density snapshot on the auxiliary volume side in the tth period can be determined according to the high-density snapshot on the primary volume side in the tth period and the first mapping relationship.

[0102] Step S203: completing data synchronization of the auxiliary volume according to the preset data in the physical volume snapshot.

[0103] Specifically, the preset data stored in the physical volume snapshot is the data after the operation corresponding to the I / O request is executed in the preset storage location. The I / O request includes the location information and operation type of the operation on the primary volume data, and the operation type includes a read operation type or a write operation type.

[0104] For example, if there is a write operation in the I / O request identified by the method above, and the I / O request includes a write location, the data is written to the corresponding write location in the primary volume according to the write location and the write operation type, and the generation identification information consistent with the I / O request is added to the data. The physical volume snapshot obtains all data containing the generation identification information from the write location, and transmits the data to the secondary volume to complete the data synchronization of the secondary volume.

[0105] Specifically, the method can accurately determine the data that needs to be synchronized in the tth cycle based on the preset data in the physical volume snapshot, thereby achieving accurate data synchronization. According to the high-density snapshot of the primary volume side in the tth cycle and the first mapping relationship, the high-density snapshot of the auxiliary volume side in the tth cycle is determined to complete the data synchronization of the auxiliary volume, ensuring the data consistency between the primary volume and the auxiliary volume. Through the physical volume snapshot and mapping relationship, data can be flexibly managed to ensure that the correct data state can be accurately restored during the data recovery process, improve the accuracy of data recovery, and minimize the risk of data loss.

[0106] Based on any of the foregoing embodiments, in an optional embodiment, this embodiment provides a data synchronization method, which can be used in the above-mentioned mobile terminal, such as a mobile phone, a tablet computer, etc. Figure 3 It is a flowchart of a data synchronization method provided by an embodiment of the present invention.

[0107] Specifically, as described above, according to the t-1th cycle on the primary volume side, identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle is generated. The process includes the following steps:

[0108] Step S301, generating a generation relationship according to the t-1th cycle on the main volume side.

[0109] Step S302: Generate identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle according to the generation relationship.

[0110] Specifically, in a specific example, continuous data protection (Continuous Data Protect I / On, CDP for short) is used to monitor the I / O requests issued by the client between the t-1th cycle and the tth cycle in real time, and after generating the t-1th generation according to the t-1th cycle on the primary volume side, the t-1th generation identifier is added to the I / O requests issued between the t-1th cycle and the tth cycle. The generation identifier is marked by CDP, such as CDP1, CDP2, etc.

[0111] Subsequently, the location information and operation type of the operation data are determined based on the identified I / O request. Based on the location information and operation type, the data corresponding to the identified I / O request is read to generate the physical volume snapshot corresponding to the cycle. By accurately reading and storing data based on the location information and operation type of the I / O request, the integrity and accuracy of the data are ensured. By storing the changed data of each cycle in the change volume for physical volume snapshot generation operations, rather than directly operating on the primary volume, the impact on the primary volume I / O performance can be reduced.

[0112] The following is a specific example to illustrate the overall operation process of the above method of the present application. Figure 4 As shown:

[0113] The interval for creating high-density snapshots on the primary and secondary volumes is preset to 300 seconds. The replication relationship between primary volume A and secondary volume B is configured. High-density snapshots cdp-ax and cdp-bx are created for primary volume A and secondary volume B every 300 seconds, where x represents the number of high-density snapshots of primary volume A and secondary volume B. The I / O requests issued between the t-1th cycle and the tth cycle are marked as cdp-at-1. After the I / O requests marked with cdp-at-1 perform the corresponding operations, the data after the corresponding operations in primary volume A are marked as cdp-at-1.

[0114] Create a high-density snapshot cdp-at of the tth period on the primary volume side, and create a high-density snapshot cdp-bt of the tth period on the secondary volume side. Create a temporary change volume A' of volume A, and convert the high-density snapshot cdp-at to a physical volume snapshot. The process of converting the physical volume snapshot is to transfer the data marked as cdp-at-1 in the primary volume A to the change volume A' according to the I / O request marked as cdp-at-1. Then, all the data in the physical volume snapshot is synchronously transferred to the secondary volume B corresponding to the tth period on the secondary volume side. After the transfer is completed, the change volume A' is deleted, and the synchronization of this cycle is completed.

[0115] Based on the above embodiments, in an optional embodiment, for a special scenario where multiple remote replications in a cluster simultaneously enter the data synchronization process and interval period, a method for balancing the configuration period based on the remote replication service load is proposed, which specifically includes:

[0116] Step a1, obtaining the network transmission rate, packet loss rate, link bandwidth, number of remote replications on a single link, data transmission volume in each cycle, and time period corresponding to the cycle in the cluster to which the asynchronous remote replication system belongs.

[0117] Step a2: determining the remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications on a single link in the cluster to which the asynchronous remote replication system belongs.

[0118] The cluster to which the remote replication system belongs includes multiple remote replication links, and each remote replication link includes multiple remote replication operations.

[0119] Step a3: determining the remote replication load according to the data transmission volume in each cycle and the time period corresponding to the cycle.

[0120] Specifically, according to the data transmission volume in each cycle and the cycle time period T, the formula for determining the remote replication load is load_rc=data transmission volume / T.

[0121] Step a4: adjusting the time period corresponding to the cycle according to the remote replication load and the remote replication average load.

[0122] Based on the above-mentioned embodiment, in an optional embodiment, the remote replication load can be determined according to the data transmission volume in each cycle and the time period corresponding to the cycle. However, in order to adjust the time period corresponding to the cycle according to the remote replication load and the remote replication average load, it is also necessary to determine the remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth in the cluster to which the asynchronous remote replication system belongs, and the number of remote replications of a single link, specifically including:

[0123] Step b1, determining the actual quality of each remote replication link according to the network transmission rate and packet loss rate;

[0124] Specifically, according to the network transmission rate and the packet loss rate, the formula for determining the actual quality of each remote replication link is LQ (Link Quality) = lb (network transmission rate / 75) - link data packet loss rate.

[0125] Step b2, determining the optimal quality of each remote replication link according to the link bandwidth in the cluster;

[0126] Specifically, the formula for determining the optimal quality of each remote replication link according to the link bandwidth B in the cluster is LQ_max=lb(B / 75).

[0127] Step b3, determining the bandwidth load of each remote replication link according to the actual quality and the optimal quality;

[0128] Specifically, according to the actual quality LQ and the optimal quality LQ_max, the formula for determining the bandwidth load of each remote replication link is load=(LQ / LQ_max)*B.

[0129] Step b4: determining the remote replication average load of each remote replication link according to the bandwidth load of each remote replication link and the remote replication quantity.

[0130] Specifically, according to the bandwidth load of each remote replication link and the number of remote replication links N, a formula for determining the remote replication average load of each remote replication link is load_rc_average=load / N.

[0131] Based on the above embodiments, it can be known that the cycle corresponds to a preset time period. In an optional embodiment, the time period corresponding to the cycle is adjusted according to the remote replication load and the remote replication average load, specifically including:

[0132] When the remote replication load is greater than the remote replication average load, the time period corresponding to the cycle is reduced to a first preset multiple of the preset time period;

[0133] Specifically, when the remote replication load load_rc is greater than the remote replication average load load_rc_average, the time period corresponding to the cycle is reduced to a preset multiple of the preset time period, for example, T / 2.

[0134] Alternatively, when the remote replication load is less than the remote replication average load, the time period corresponding to the cycle is enlarged to a second preset multiple of the preset time period.

[0135] The first preset multiple is smaller than the second preset multiple.

[0136] Specifically, when the remote replication load load_rc is less than the remote replication average load load_rc_average, the time period corresponding to the cycle is enlarged to a preset multiple of the preset time period, for example, 2T.

[0137] It should be noted that when adjusting the actual period of remote replication in balanced period mode, it is necessary to adjust it in the remote replication period interval and recalculate the period after adjustment. In addition, if the adjusted period is not within the system supported period range, the adjustment operation is not performed.

[0138] Specifically, the method can adjust the cycle of periodic asynchronous remote replication in the system according to the business load. The cycle when the business load is large can be shortened, the amount of synchronized data in each synchronization process can be reduced, and the synchronization rate can be improved. The cycle of periodic asynchronous remote replication when the business is small or even no business can be increased to reduce the space occupied by temporary changes to the volume during the synchronization process and improve the storage space utilization rate.

[0139] On the basis of the above-mentioned embodiments, CDP technology is used to monitor I / O requests in real time, identify I / O requests between consecutive cycles, identify the data corresponding to the I / O request execution according to the I / O request identified in each cycle, and only identify and store the actual changed part of the data, rather than the full amount of data. In an optional embodiment, before reading the identified data to generate a physical volume snapshot, the changed data can be compressed in real time to reduce storage requirements and network transmission bandwidth. Specifically, it includes:

[0140] Step c1, compressing the data corresponding to the I / O request.

[0141] Step c2: during the data compression process, a generation relationship and corresponding identification information are generated for each compression interval.

[0142] Specifically, the generation relationship and corresponding identification information of each compression interval may be consistent with the I / O request identification, such as CDP1, CDP2, and so on.

[0143] Step c3: Generate a physical volume snapshot based on the compressed data and identification information.

[0144] Because the data has been compressed, the physical volume snapshot occupies less storage space, reducing storage requirements and network transmission bandwidth.

[0145] In this embodiment, a data synchronization device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0146] This embodiment provides a data synchronization device, such as Figure 5 As shown, it includes: a creation module 501, a processing module 502, an extraction module 503, and a synchronization module 504.

[0147] A creation module 501 is used to periodically create high-density snapshots on the primary volume side and the secondary volume side respectively;

[0148] The processing module 502 is used to establish a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side; and generate identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle according to the t-1th cycle on the primary volume side, where t is a positive integer greater than or equal to 2.

[0149] The creation module 501 is further used to create a change volume corresponding to the primary volume after the high-density snapshot created in the tth period on the primary volume side;

[0150] The extraction module 503 is used to extract preset data after the operation corresponding to the I / O request is performed from a preset storage location in the changed volume according to the identification information, wherein the preset data carries the identification information;

[0151] The processing module 502 is further used to generate a physical volume snapshot according to preset data; establish a second mapping relationship between the high-density snapshot created in the tth period on the master volume side and the physical volume snapshot;

[0152] The synchronization module 504 is used to synchronize the data of the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0153] In an optional embodiment, the processing module 502 is further configured to delete the changed volume after synchronizing the data of the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

[0154] In an optional embodiment, the processing module 502 is specifically configured to determine, according to the physical volume snapshot and the second mapping relationship, a high-density snapshot of the primary volume corresponding to the physical volume snapshot in the tth period;

[0155] Determine the high-density snapshot of the auxiliary volume side in the tth period according to the high-density snapshot of the primary volume side in the tth period and the first mapping relationship;

[0156] Complete data synchronization of the auxiliary volume based on the preset data in the physical volume snapshot.

[0157] In an optional embodiment, the processing module 502 is specifically configured to generate a generation relationship according to the t-1th cycle on the main volume side;

[0158] According to the generation relationship, identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle is generated.

[0159] In an optional embodiment, the extraction module 503 is specifically used to obtain the network transmission rate, packet loss rate, link bandwidth, number of remote replications of a single link, data transmission volume in each cycle, and time period corresponding to the cycle in the cluster to which the asynchronous remote replication system belongs;

[0160] The processing module 502 is specifically used to determine the remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications of a single link in the cluster to which the asynchronous remote replication system belongs, wherein the cluster to which the remote replication system belongs includes multiple remote replication links, and each remote replication link includes multiple remote replication operations;

[0161] Determine the remote replication load based on the data transmission volume in each cycle and the time period corresponding to the cycle;

[0162] Adjust the time period corresponding to the cycle according to the remote replication load and the remote replication average load.

[0163] In an optional embodiment, the processing module 502 is specifically configured to determine the actual quality of each remote replication link according to the network transmission rate and the packet loss rate;

[0164] Determine the optimal quality of each remote replication link based on the link bandwidth in the cluster;

[0165] Determine the bandwidth load of each remote replication link based on the actual quality and the optimal quality;

[0166] The average load of remote replication on each remote replication link is determined based on the bandwidth load of each remote replication link and the number of remote replications.

[0167] In an optional embodiment, the cycle corresponds to a preset time period, and the processing module 502 is specifically configured to reduce the time period corresponding to the cycle to a first preset multiple of the preset time period when the remote replication load is greater than the remote replication average load;

[0168] or,

[0169] When the remote replication load is less than the remote replication average load, the time period corresponding to the cycle is enlarged to a second preset multiple of the preset time period, wherein the first preset multiple is less than the second preset multiple.

[0170] The data synchronization device in this embodiment is presented in the form of a functional module, where the module refers to an application-specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0171] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0172] A data synchronization device provided by an embodiment of the present invention realizes data synchronization from a primary volume to a secondary volume through the cooperation between modules. High-density snapshots are periodically created on the primary volume side and the secondary volume side, rather than being triggered by humans or events to create high-density snapshots, providing continuous data protection and reducing the risk of data loss. In addition, in write-intensive applications, periodic creation does not require a creation operation every time data is changed, which can avoid affecting system performance due to frequent snapshot operations. A first mapping relationship between a high-density snapshot on the primary volume side and a high-density snapshot on the secondary volume side is established, which helps to quickly restore the data of the primary volume from the secondary volume when needed. According to the t-1th cycle on the primary volume side, identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle is generated. The high-density snapshot does not occupy the physical volume space, but only batches the issued I / O requests and the data after the operation has been executed. The change volume corresponding to the primary volume is subsequently created to read the data after the operation corresponding to each batch of I / O requests and convert them into physical volume snapshots. Each periodic asynchronous remote replication only configures one physical volume snapshot during the synchronization process, reducing the space requirements of periodic asynchronous replication and the time consumption of starting and stopping auxiliary volumes to change volume snapshots. At the same time, it improves the storage space utilization and configuration quantity restrictions of periodic asynchronous replication due to traditional ROW snapshot restrictions. High-density snapshots abandon the discard and flush processes, improve resource consumption and performance impact during each remote replication process, and further reduce the minimum period supported by periodic asynchronous replication.

[0173] An embodiment of the present invention further provides a computer device, Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.

[0174] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include an integrated circuit. The integrated circuit may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0175] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0176] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0177] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0178] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.

[0179] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0180] The embodiment of the present invention also provides a computer-readable storage medium. The method provided in the above embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or is implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0181] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0182] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A data synchronization method, characterized in that: The method is applied to an asynchronous remote replication system, the asynchronous remote replication system includes a primary volume and a secondary volume, and the method includes: Periodically creating high-density snapshots on the primary volume side and the secondary volume side, respectively, and establishing a first mapping relationship between the high-density snapshots on the primary volume side and the high-density snapshots on the secondary volume side; Generate identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle according to the t-1th cycle on the primary volume side, where t is a positive integer greater than or equal to 2; After the high-density snapshot created in the t-th period on the master volume side, creating a change volume corresponding to the master volume; In the changed volume, according to the identification information, preset data after executing the operation corresponding to the I / O request is extracted from a preset storage location, and a physical volume snapshot is generated, wherein the preset data carries the identification information; Establishing a second mapping relationship between the high-density snapshot created on the master volume side in the t-th period and the physical volume snapshot; Data of the auxiliary volume is synchronized according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

2. The method according to claim 1, characterized in that After synchronizing the data of the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship, the method further includes: Delete the change volume.

3. The method according to claim 1 or 2, characterized in that: The step of synchronizing data of the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship specifically includes: Determine, according to the physical volume snapshot and the second mapping relationship, a high-density snapshot of the primary volume corresponding to the physical volume snapshot in the tth period; Determining the high-density snapshot of the auxiliary volume in the tth period according to the high-density snapshot of the primary volume in the tth period and the first mapping relationship; The data synchronization of the auxiliary volume is completed according to the preset data in the physical volume snapshot.

4. The method according to claim 1 or 2, characterized in that: The generating, according to the t-1th cycle on the primary volume side, identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle specifically includes: Generate a generation relationship according to the t-1th period on one side of the main volume; According to the generation relationship, identification information corresponding to the I / O request issued between the t-1th cycle and the tth cycle is generated.

5. The method according to claim 1, characterized in that The method further comprises: Obtaining the network transmission rate, packet loss rate, link bandwidth, number of remote replications on a single link, data transmission volume in each cycle, and time period corresponding to the cycle in the cluster to which the asynchronous remote replication system belongs; Determine the remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth, and remote replication quantity of a single link in the cluster to which the asynchronous remote replication system belongs, wherein the cluster to which the remote replication system belongs includes multiple remote replication links, and each remote replication link includes multiple remote replication operations; Determining the remote replication load according to the data transmission volume in each cycle and the time period corresponding to the cycle; The time period corresponding to the cycle is adjusted according to the remote replication load and the remote replication average load.

6. The method according to claim 5, characterized in that Determining the remote replication average load of each remote replication link according to the network transmission rate, packet loss rate, link bandwidth, and the number of remote replications of a single link in the cluster to which the asynchronous remote replication system belongs specifically includes: determining the actual quality of each remote replication link according to the network transmission rate and the packet loss rate; Determining the optimal quality of each of the remote replication links according to the link bandwidth in the cluster; Determining a bandwidth load of each of the remote replication links according to the actual quality and the optimal quality; The average load of remote replication on each remote replication link is determined based on the bandwidth load of each remote replication link and the number of remote replications.

7. The method according to claim 5 or 6, characterized in that: The cycle corresponds to a preset time period; The adjusting the time period corresponding to the cycle according to the remote replication load and the remote replication average load specifically includes: When the remote replication load is greater than the remote replication average load, shortening the time period corresponding to the cycle to a first preset multiple of the preset time period; or, When the remote replication load is less than the remote replication average load, the time period corresponding to the cycle is enlarged to a second preset multiple of the preset time period, wherein the first preset multiple is less than the second preset multiple.

8. A data synchronization device, characterized in that: The device comprises: The creation module is used to periodically create high-density snapshots on the primary volume side and the secondary volume side; A processing module, configured to establish a first mapping relationship between the high-density snapshot on the primary volume side and the high-density snapshot on the secondary volume side; and generate identification information corresponding to an I / O request issued between the t-1th cycle and the tth cycle according to the t-1th cycle on the primary volume side, where t is a positive integer greater than or equal to 2; The creation module is further configured to create a change volume corresponding to the master volume after the high-density snapshot created in the t-th period on the master volume side; an extraction module, configured to extract, in the change volume, preset data after executing an operation corresponding to the I / O request from a preset storage location according to the identification information, wherein the preset data carries the identification information; The processing module is further used to generate a physical volume snapshot according to the preset data; establish a second mapping relationship between the high-density snapshot created by the master volume side in the tth period and the physical volume snapshot; A synchronization module is used to synchronize data of the auxiliary volume according to the physical volume snapshot, the first mapping relationship, and the second mapping relationship.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data synchronization method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data synchronization method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data processing method and related device

    CN108762988A

  • User snapshot synchronization method and device based on periodic asynchronous remote replication relationship

    CN118170716A

  • Data volume remote copying method and device, computer equipment and storage medium

    CN118708130A

  • Snapshot processing method and device based on dual-live volume and computer related equipment

    CN119271132A

  • Non-disruptive baseline and resynchronization of a synchronous replication relationship

    US20170154093A1

Cited By

  • Enterprise trusted data access method and system based on block chain

    CN120934850A