Mirror volume data synchronization method, system, electronic device and storage medium
By obtaining application scenario information and read and write frequency in the mirror volume, determining the target synchronization strategy and algebraic values, and synchronizing the recent snapshot data to the auxiliary space, it solves the problem of the mirror volume losing historical data after the auxiliary space failure recovery, and improves the high availability of the mirror volume.
Patent Information
- Application Number
- CN202510824425.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-19
AI Technical Summary
In the prior art, when the mirror volume is re-on after the auxiliary space failure recovery, the multi-generation snapshot data before the latest data may be lost, resulting in a reduced high availability of the mirror volume.
By obtaining the application scenario information of the mirror volume, determining the target synchronization strategy, and determining the target synchronization algebra value based on the read and write frequency information, positioning the most recent snapshot data of the corresponding algebra in the main space, synchronizing it to the auxiliary space, ensuring that both the main space and the auxiliary space retain complete historical snapshot data.
Improves the high availability of mirrored volumes, and even if the main space fails, the secondary space can provide historical snapshot data to ensure data integrity.
Smart Images

Figure CN120353768B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a mirror volume data synchronization method, system, electronic device, and storage medium. Background Art
[0002] Currently, in order to achieve high data availability of data volumes, ordinary data volumes are often converted into mirrored volumes. The mirrored volume is mapped to the host as a volume. In the underlying storage pool, there are two volumes of storage space, namely the primary space and the secondary space. The data between the primary space and the secondary space is synchronized. When the primary space fails, the secondary space serves as the primary space to provide services to the outside world.
[0003] In related technologies, taking the case of a secondary space being put back online after failure recovery as an example, usually only the latest data of the primary space during the secondary space failure recovery period is synchronized. During this period, the primary space usually saves multiple generations of snapshot data based on continuous data protection (CDP), resulting in the secondary space that is put back online may lose multiple generations of snapshot data before the latest data. If the primary space fails subsequently, the secondary space will serve as the primary space to provide services to the outside world. The secondary space will not be able to provide these generations of snapshot data due to data loss, thereby reducing the high availability of the mirrored volume. Summary of the Invention
[0004] The present application provides a mirror volume data synchronization method, system, electronic device and storage medium to at least solve the problem of reducing the high availability of mirror volumes in related technologies.
[0005] This application provides a mirror volume data synchronization method, including:
[0006] Obtain application scenario information of the mirror volume; wherein the application scenario information includes at least business information;
[0007] Determine the target synchronization strategy for the mirrored volume based on its application scenario information.
[0008] When the target synchronization strategy is partial synchronization, the read and write frequency information of the mirror volume is determined based on the business information;
[0009] Determine the target synchronization algebra value based on the read and write frequency information of the mirror volume;
[0010] According to the target synchronization generation value, locate the latest snapshot data of the corresponding generation in the main space, and use the latest snapshot data of the corresponding generation as the data to be synchronized;
[0011] Synchronize the data to be synchronized to the auxiliary space;
[0012] The mirror volume includes primary space and secondary space.
[0013] The present application also provides a mirror volume data synchronization device, comprising:
[0014] An acquisition module, configured to acquire application scenario information of the mirror volume; wherein the application scenario information includes at least business information;
[0015] A first determining module is used to determine a target synchronization strategy for the mirror volume according to application scenario information of the mirror volume;
[0016] The second determination module is used to determine the read and write frequency information of the mirror volume according to the business information when the target synchronization strategy is partial synchronization;
[0017] A third determining module is used to determine a target synchronization algebra value according to the read and write frequency information of the mirror volume;
[0018] A first synchronization module is configured to locate the latest snapshot data of the corresponding generation in the main space according to the target synchronization generation value, and use the latest snapshot data of the corresponding generation as the data to be synchronized;
[0019] The second synchronization module is used to synchronize the data to be synchronized to the auxiliary space;
[0020] The mirror volume includes primary space and secondary space.
[0021] The present application also provides a mirror volume data synchronization system, comprising: a primary space, a secondary space, and a mirror volume data synchronization device;
[0022] The primary space and secondary space belong to different storage pools;
[0023] The main space is used to store snapshot data;
[0024] The auxiliary space is used to synchronize the snapshot data stored in the primary space;
[0025] The mirror volume data synchronization device synchronizes the data to be synchronized in the primary space to the secondary space based on any of the above mirror volume data synchronization methods.
[0026] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned mirror volume data synchronization methods when executing the computer program.
[0027] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned mirror volume data synchronization methods are implemented.
[0028] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned mirror volume data synchronization methods when executed by a processor.
[0029] Through this application, the target synchronization generation value that matches the application scenario information of the mirror volume is determined, and according to the target synchronization generation value, the latest snapshot data of the corresponding generation is located in the primary space, so that the latest snapshot data of the corresponding generation is used as the data to be synchronized, and finally the data to be synchronized is synchronized to the secondary space, so that the primary space and secondary space of the mirror volume both retain complete historical snapshot data. Even if the primary space fails, the secondary space can provide historical snapshot data, thereby improving the high availability of the mirror volume. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0031] Figure 1 This is a schematic diagram of the network structure based on the embodiments of the present application;
[0032] Figure 2 A schematic diagram of a process for synchronizing mirrored volume data provided by an embodiment of the present application;
[0033] Figure 3 A schematic diagram of the data structure of a mirror volume provided in an embodiment of the present application;
[0034] Figure 4 A schematic diagram of the data structure of another mirror volume provided in an embodiment of the present application;
[0035] Figure 5 A schematic diagram of the data structure of another mirror volume provided in an embodiment of the present application;
[0036] Figure 6 A schematic diagram of the structure of a mirror volume data synchronization device provided in an embodiment of the present application;
[0037] Figure 7 A schematic diagram of the structure of a mirror volume data synchronization system provided in an embodiment of the present application;
[0038] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0040] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0041] CDP is a data protection technology. It adds a timestamp-like record (generation id, gen_id) to each I / O transaction on the volume, recording the snapshot generation to which each I / O transaction belongs. This ensures that the data corresponding to CDP snapshots taken at different times belongs to the volume itself. In the volume's underlying storage space, the gen_id is used to distinguish data at different times at the same location, ensuring that older data is not deemed invalid and recycled. Compared to traditional non-CDP snapshots, CDP snapshots do not require the creation of a separate target volume. Data from a specific generation of snapshots can be quickly found simply by gen_id, eliminating the traditional snapshot's copy-on-write (COW) data replication process. This means that CDP snapshots offer virtually lossless performance and stronger data protection, allowing for recovery to almost any point in time.
[0042] For example, for a volume, before a CDP snapshot is created, the gen_id value of all IO is 0. After a CDP snapshot is created, the snapshot becomes the 0th generation snapshot, and its corresponding data is all the data in the source volume with gen_id 0. After the snapshot is created, the gen_id of the source volume will increase by 1, that is, the gen_id of subsequent IO will become 1. If another CDP snapshot is taken, this new CDP snapshot will become the 1st generation snapshot, and the gen_id will increase by 1 again, and the gen_id of subsequent IO will become 2, and so on. Because the host IO is not uniform, some locations of the source volume may often receive IO, while some locations may have less IO. Therefore, the gen_id of each location is not equal. After multiple rounds of CDP snapshots, the gen_id of hot areas will often be very large or directly equal to the number of snapshots, while the gen_id of unpopular areas may hardly change. When reading the IO of the source volume, you only need to read the data with the largest gen_id at each location, which is the latest data. When reading data with a CDP snapshot, you need to specify the gen_id. If the id exists in the area, the corresponding data is read directly. If not, the gen_id-1 is checked to see if it exists. The search continues until the data is found. The data found is the data corresponding to the specified gen_id.
[0043] Mirrored volume technology is a high-availability technology. For a standard volume, a single volume is mapped to the host. Similarly, the underlying storage pool has only one storage space for that volume. All data written to that volume has only one copy. If a storage pool failure occurs (for example, a software failure causes the RAID in the storage pool to cease service, or a hardware failure causes the number of hard drives in the RAID to exceed the RAID tolerance limit), that volume, or even all volumes in the entire storage pool, becomes unavailable. If the failure cannot be corrected, the business data in those volumes cannot be recovered. Mirrored volume technology provides an additional layer of data security. For a mirrored volume, a single volume is mapped to the host, but two storage spaces exist in the underlying storage pool. When a write I / O arrives at the mirrored volume module, it copies the data identically into both storage spaces. This is known as mirrored volume dual-write mode. When reading data, only one of the storage spaces is needed, as the two copies are identical. When creating a mirrored volume, it's generally recommended to create two storage spaces in two different storage pools. This way, if one storage space fails, the original data can be retrieved from the other space. Even if the failed space cannot be recovered, the business data of that volume will not be lost. Furthermore, if the failed space is recovered, the thin volume module also provides a data resynchronization function, copying the differential data written during the failure from the healthy space to the newly recovered space. Once the differential data synchronization is complete, high data availability can be maintained.
[0044] The mirrored volume feature in related technologies requires initial data synchronization when converting a regular volume to a mirrored volume. Specifically, after executing the command to convert a mirrored volume, an additional storage space of the same size is allocated in the specified storage pool (the original storage space is called the primary space, and the newly created space is called the secondary space). A bitmap is then allocated in memory for the volume. Each bit in this bitmap manages a fixed-size location within the volume. Each fixed-size data block is called a grain and is typically 256KB in size. Initially, all bits in the bitmap are set to 1. Then, starting from the beginning of the volume, data is copied from the primary space to the newly created secondary space at a grain granularity. After each location is copied, the corresponding bit in the bitmap is reset to 0, indicating that the data at that location has been replicated twice. When the entire bitmap is set to 0, the mirrored volume is considered highly available.
[0045] During the initial synchronization of a mirrored volume, if there is an IO write to the mirrored volume, when the IO reaches the mirrored volume module, it will first read the value of the bit corresponding to that location. If the value is 0, it indicates that the location has completed data synchronization, and the IO needs to be double-written. Here, the mirrored volume module will copy the same data and then write two identical copies of the data to two storage spaces respectively; if the corresponding bit is 1, it indicates that the location has not completed data synchronization. At this time, the data to be written only needs to be written to the original storage space (i.e., the primary space). Then, when the initial synchronization task processes the location, the latest data is copied to the newly created storage space (i.e., the secondary space). This ensures that the original volume can also receive business IO normally during the initial synchronization, and after the initial synchronization is completed, the data in the two storage spaces is also up-to-date and identical.
[0046] If a storage space fails, the remaining storage space is switched to the primary space, and the failed space is switched to the secondary space. Reading and writing operations now access only the primary space. Furthermore, if data at a location in the primary space changes, the corresponding bit in the bitmap is modified from 0 to 1 to indicate that the location is different between the two storage spaces.
[0047] When the other storage space is restored, the mirrored volume module resynchronizes the data based on the bitmap values, resynchronizing the latest data at the location marked as 1 in the primary to the secondary. Similarly, the bit is reset to 0 after each location is synchronized. Host I / O continues normally during the resynchronization process, using the same logic as the initial synchronization. When the bitmap returns to all 0s, the mirrored volume regains high availability.
[0048] When converting an existing volume to a mirrored volume, the mirrored volume technology used in related technologies only synchronizes the latest data from the original volume to the newly added secondary storage space. If the original volume has multiple CDPs created, the secondary space will only receive the latest generation of data from the primary space; data from any CDPs will not be synchronized to the secondary space. In this case, if the primary space fails and the secondary space is switched to primary, the mirrored volume will no longer be able to provide data from the pre-existing CDPs. For example, when a secondary space recovers from a failure and comes online again, typically only the latest data from the primary space during the recovery period is synchronized. During this period, the primary space typically preserves multiple generations of snapshot data using Continuous Data Protection (CDP). As a result, the secondary space that comes online again may lose multiple generations of snapshot data prior to the latest data. If the primary space subsequently fails and the secondary space becomes the primary space, it will be unable to provide these generations of snapshot data due to data loss, reducing the high availability of the mirrored volume.
[0049] In order to solve the above-mentioned technical problems, the embodiments of the present application provide a mirror volume data synchronization method, system, electronic device and storage medium, the method comprising: obtaining application scenario information of the mirror volume; wherein the application scenario information includes at least business information; determining the target synchronization strategy of the mirror volume based on the application scenario information of the mirror volume; when the target synchronization strategy is partial synchronization, determining the read and write frequency information of the mirror volume based on the business information; determining the target synchronization generation value based on the read and write frequency information of the mirror volume; locating the latest snapshot data of the corresponding generation in the primary space based on the target synchronization generation value, and using the latest snapshot data of the corresponding generation as the data to be synchronized; synchronizing the data to be synchronized to the auxiliary space; wherein the mirror volume includes the primary space and the auxiliary space. The method provided by the above scheme determines the target synchronization generation value that matches the application scenario information of the mirror volume, and locates the latest snapshot data of the corresponding generation in the primary space according to the target synchronization generation value, so as to use the latest snapshot data of the corresponding generation as the data to be synchronized, and finally synchronizes the data to be synchronized to the secondary space, so that the primary space and the secondary space of the mirror volume both retain complete historical snapshot data. Even if the primary space fails, the secondary space can provide historical snapshot data, thereby improving the high availability of the mirror volume.
[0050] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0051] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the mirror volume data synchronization method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0052] First, the network structure on which this application is based is described:
[0053] The mirror volume data synchronization method, system, electronic device and storage medium provided in the embodiments of the present application are suitable for synchronizing the primary and secondary spaces of the snapshot data of the mirror volume. Figure 1 The figure shows a schematic diagram of the network structure based on the embodiment of the present application, which mainly includes a user end and a host end. The host end is used to send a service request to the host end to perform data read and write operations on the storage device in the host end. When responding to the data read and write operations initiated by the user end, the host end performs CDP snapshot processing on the storage device and obtains the CDP snapshot based on the mirror volume management. The host end synchronizes the primary and secondary spaces of the mirror volume snapshot data based on the mirror volume data synchronization method provided in the embodiment of the present application.
[0054] The embodiment of the present application provides a mirror volume data synchronization method for synchronizing primary and secondary spaces of snapshot data of a mirror volume. The execution subject of the embodiment of the present application is an electronic device, such as a server, desktop computer, laptop computer, tablet computer, or other electronic device that can be used for mirror volume data synchronization.
[0055] like Figure 2 FIG. 1 is a flow chart of a mirror volume data synchronization method provided by an embodiment of the present application, the method comprising:
[0056] Step 201: Obtain application scenario information of the mirror volume.
[0057] The application scenario information includes at least business information, which includes at least business type, snapshot data creation frequency, and user read / write request initiation status.
[0058] Step 202: Determine a target synchronization strategy for the mirror volume based on the application scenario information of the mirror volume.
[0059] The target synchronization strategy is divided into at least three types: partial synchronization, full synchronization, and latest synchronization.
[0060] Specifically, in one embodiment, the user's protection requirements for the historical snapshot data of the mirror volume can be determined based on the application scenario information of the mirror volume; and the target synchronization strategy of the mirror volume can be determined based on the user's protection requirements for the historical snapshot data of the mirror volume.
[0061] Specifically, if the mirrored volume application scenario information indicates that the requirements for high availability and historical data protection are extremely high, and if one storage space fails, the other storage space still needs to be able to recover data to any historical CDP time, the user's requirements for protecting the mirrored volume's historical snapshot data are determined to be high, and the mirrored volume's target synchronization policy is determined to be full synchronization. If the mirrored volume application scenario information indicates that the user has set a scheduled CDP policy, for example, automatically taking a CDP snapshot of the volume every day and regularly checking for data tampering every week, snapshots older than a week are generally considered less important, and only the most recent week's snapshots are needed to meet the historical data protection requirements. Therefore, the user's requirements for protecting the mirrored volume's historical snapshot data are determined to be moderate, and the mirrored volume's target synchronization policy is determined to be partial synchronization. If the mirrored volume application scenario information indicates that the user is more concerned with high data availability and snapshot data loss is not a major concern, the user's requirements for protecting the mirrored volume's historical snapshot data are determined to be low, and the mirrored volume's target synchronization policy is determined to be latest synchronization.
[0062] Step 203: When the target synchronization strategy is partial synchronization, the read and write frequency information of the mirror volume is determined according to the business information.
[0063] Specifically, the read / write frequency information of the mirror volume can be determined based on the initiation of the user read / write request represented by the service information, wherein the read / write frequency information at least represents whether the mirror volume has frequent write operations or frequent read operations.
[0064] Step 204: Determine the target synchronization algebra value based on the read and write frequency information of the mirror volume.
[0065] Specifically, when the read and write frequency information of the mirror volume indicates that the mirror volume write operation is frequent, a larger algebra value can be selected as the target synchronization algebra value; when the read and write frequency information of the mirror volume indicates that the mirror volume read operation is frequent, a smaller algebra value can be selected as the target synchronization algebra value.
[0066] Step 205 : Locate the latest snapshot data of the corresponding generation in the main space according to the target synchronization generation value, and use the latest snapshot data of the corresponding generation as the data to be synchronized.
[0067] For example, taking the target synchronization generation value as 5, the latest snapshot data (7th generation snapshot data) and the snapshot data of the last 5 generations (6th generation to 2nd generation) after the latest snapshot data are used as the data to be synchronized, that is, the first 6 generations of snapshot data are used as the data to be synchronized.
[0068] Step 206: Synchronize the data to be synchronized to the auxiliary space.
[0069] The mirror volume includes primary space and secondary space.
[0070] Specifically, the data to be synchronized may be synchronized to the auxiliary space in the same data synchronization manner as that used in the initial synchronization phase of the mirror volume and the resynchronization phase after failure recovery.
[0071] For example, Figure 3 As shown, it is a data structure diagram of a mirror volume provided by an embodiment of the present application. In actual application, the user can independently select the target synchronization generation value, that is, the user can choose to synchronize only the latest generations of data. In scenarios such as initial synchronization of the mirror volume and resynchronization after failure recovery, only the latest n generations of data in the primary space will be synchronized to the secondary space. This function is mainly suitable for scenarios where data is regularly protected to prevent tampering, but only recent data is paid attention to and not much attention is paid to data from a long time ago. For example, if the user sets a scheduled CDP policy, a CDP snapshot will be automatically taken for the volume every day, but the data will be checked regularly every week to see if it has been tampered with. At this time, it will be considered that snapshots that are more than a week old are not very important, and only the snapshots of the most recent week need to be retained to meet the data protection requirements. Then, when the same demand for high data availability is required, the volume can be converted to a mirror volume and only the latest 7 generations of CDP data (snapshot data) can be selected for synchronization. When a storage space fails, another storage space can be restored to the latest 7 generations of CDP data. When the target synchronization generation value n=3, Figure 3 There are N CDP snapshots, corresponding to gen_ids: 0 to N-1. The Nth generation data represents the latest data, which has no corresponding CDP in the auxiliary space. Therefore, when n=3, the data of four generations N-3 to N will be copied as the data to be synchronized (N-3 to N-1 are the data of three CDPs (historical snapshot data), and n is the latest data).
[0072] It should be noted that the method provided in the embodiment of the present application is mainly used in scenarios such as initial synchronization of mirror volumes and resynchronization after failure recovery. In such scenarios, the secondary space may lose multiple generations of snapshot data before the latest data. Based on the method provided in the embodiment of the present application, the primary space and secondary space of the mirror volume can retain complete historical snapshot data. Even if the primary space fails, the secondary space can provide historical snapshot data, thereby improving the high availability of the mirror volume.
[0073] Based on the above embodiment, as an implementable manner, in one embodiment, when the target synchronization strategy is partial synchronization, determining the read and write frequency information of the mirror volume according to the business information includes:
[0074] Step 2031: When the target synchronization strategy is partial synchronization, determine the number of write requests and read requests responded to by the mirror volume within a fixed time period based on the business information;
[0075] Step 2032: Determine the read frequency and write frequency of the mirror volume according to the number of write requests and read requests responded to by the mirror volume within a fixed time period.
[0076] The read / write frequency information includes at least a read frequency and a write frequency. The read frequency is the number of read requests per unit time, and the write frequency is the number of write requests per unit time.
[0077] It's important to note that the number of write requests reflects the activity of data changes, while the number of read requests reflects the need to access historical data. Frequent write scenarios require more frequent snapshot protection, while frequent read scenarios may rely more on the latest data.
[0078] Accordingly, in one embodiment, the read frequency and write frequency of the mirror volume may be determined based on the read and write frequency information of the mirror volume; and the target synchronization generation value may be determined based on the read frequency and write frequency of the mirror volume.
[0079] Among them, the target synchronization generation value is positively correlated with the write frequency, and the target synchronization generation value is negatively correlated with the read frequency.
[0080] Specifically, for business scenarios with low write frequency and high read frequency, a smaller n value (such as n=3) can be set to synchronize only the three most recent generations of snapshot data, reducing redundant storage usage. For business scenarios with high write frequency and low read frequency, a larger n value (such as n=10) can be set to retain more historical snapshot data to accommodate frequent rollbacks. By determining the target synchronization generation value based on read and write frequency, the high availability of mirrored volumes is improved while also boosting storage resource utilization.
[0081] Specifically, in one embodiment, the read and write frequencies are mapped to the value range of n through normalization processing, and the service weight coefficient adjustment strategy bias is introduced. The specific calculation formula of the target synchronization algebra value is as follows:
[0082]
[0083] in, represents the target synchronization algebra value, Represents the preset minimum target synchronization algebra value, Indicates the preset maximum target synchronization algebra value, Indicates the read frequency of the mirror volume. Indicates the minimum write frequency of the mirror volume in the current application scenario. Indicates the maximum write frequency of the mirror volume in the current application scenario. Indicates the read frequency of the mirror volume. Indicates the minimum read frequency of the mirror volume in the current application scenario. Indicates the maximum read frequency of the mirror volume in the current application scenario. Indicates the influence coefficient of the read frequency on the target synchronization algebra value.
[0084] Specifically, the target synchronization algebra value is calculated based on the above formula, and the calculation deviation caused by the difference in frequency unit or magnitude is avoided by converting the read and write frequencies into normalized values. The attenuation of the reading frequency to the n value is controlled so that the method provided in the embodiment of the present application can be adapted to the needs of various business scenarios.
[0085] Based on the above embodiment, as an implementable manner, in one embodiment, synchronizing the data to be synchronized to the auxiliary space includes:
[0086] Step 2061: Divide the data to be synchronized according to the generations of the data to be synchronized to obtain data to be synchronized of each generation;
[0087] Step 2062: Synchronize the data to be synchronized of each generation to the auxiliary space in the order of the generations.
[0088] It's important to clarify how to copy (synchronize) multiple generations of data from the same location in the primary storage space to another storage space. There are two possible copying schemes. The first is to copy based on the spatial location of the grain. This involves traversing the grains and, for each grain, reading the data from generations 0, 1, ..., and n, and then writing it to the target space (auxiliary space). Since the newly added storage space is unused and has no existing data, the generations of the data written should be consistent with the generations read from the original space. That is, generation 0 data from the original space is written to generation 0 of the target space, and generation n data from the original space is written to generation n of the target space. Once all generations of data for a grain have been copied, the next grain is processed. The other scheme is to copy based on generation. This involves traversing the grains first, reading generation 0 data from each grain and writing it to generation 0 of the target space. Generation 0 of the next grain is then processed. Once all generation 0 data for all grains has been copied, generation 1 data is processed again, and so on.
[0089] Specifically, the embodiment of the present application chooses to use solution 2 for the following reasons:
[0090] (1) From the perspective of implementation logic, all generations of data in the same volume are ultimately managed by the storage pool module. In this module, data of different generations are managed in a tree-like manner. Data of each generation is in the same tree, and data of different generations are in different trees. Therefore, when using Solution 2, when read IO reaches the storage pool module, it only needs to read the information of the tree corresponding to the generation to obtain the relevant data of different grains. This is because snapshot data belonging to the same generation are generally stored in a centralized manner in the storage pool module, that is, snapshot data storage trees of multiple grains belonging to the same generation are stored in a centralized manner. In contrast, when using Solution 1, it is necessary to read the trees corresponding to different generations to obtain data of different generations at the same location. This increases the number of data reads and is less efficient. On the other hand, when writing data to the target space, the target space is also managed by the storage pool module. Therefore, data of the same generation is written in a centralized manner, and nodes are inserted into the same tree multiple times in the storage pool module. In contrast, Solution 1 requires processing all generations of trees for each grain, which is also very inefficient.
[0091] (2) From a business perspective, if solution 1 is used, all generations of data for all grains must be synchronized before the other storage space can have complete data of a certain generation. If solution 2 is used, after synchronizing the 0th generation data, even if the data of later generations is still being synchronized, the target space actually already has a usable generation of data. Although it is not the latest, it is a complete and usable data. Even if the primary space fails at this time, the historical data can actually be restored through the secondary space, so that all data will not be lost.
[0092] Compared to Option 2, Option 1 has its advantages. Specifically, it requires fewer bitmap operations. This is because all generations of data for a grain are copied to auxiliary space before the next grain is processed. Therefore, after processing a grain, the corresponding bitmap position can be reset to 0. The number of bitmap operations is the same as the number of grains. Option 2, on the other hand, copies all grains of a given generation before copying all grains of the next generation. Therefore, the number of bitmap operations is the number of grains multiplied by the number of generations. However, since bitmap data is completely stored in memory, the overhead of modifying values in memory is completely negligible. Therefore, Option 2 still has advantages over Option 1.
[0093] Among them, the specific implementation steps of solution (2) to synchronize the data to be synchronized of each generation to the auxiliary space in the order of generations are as follows:
[0094] 1) During the initial synchronization of a mirrored volume, if the user clicks the Convert volume to mirror function button for an existing normal volume with several CDPs already created, the user will be prompted to select a target synchronization policy. If the user selects partial synchronization, the target synchronization generation value n will be determined. The resulting new storage space of the mirrored volume will only contain data from the most recent n generations of CDPs in the original space.
[0095] 2) After the user completes the previous step and clicks the OK button, the system will start the task of converting the normal volume to a mirror volume;
[0096] 3) First, the system creates a new storage space (auxiliary space) in the new storage pool. The size of the new storage space is the same as the original storage space (primary space). The system associates the new storage space with the existing storage space and creates a bitmap for the volume based on the storage space size.
[0097] 4) The system then determines the starting generation of the original storage space based on the user's selection in step 1, i.e., the value of start_gen. If the user selected full synchronization, the system will start copying from generation 0 of the CDP data. If the user selected partial synchronization and specified n, the system will start copying from the generation between (latest generation and n). If the user selected latest synchronization, the system will start copying from the latest generation.
[0098] 5) Record the value of start_gen;
[0099] 6) Reset the bitmap to all 1s;
[0100] 7) Starting from the first grain in the primary space, read the data of the start_gen generation in each grain in turn, then write the data to the corresponding generation of the corresponding grain position in the auxiliary space, and change the bit in the bitmap corresponding to the grain to 0. This continues until the data at the last grain position is copied;
[0101] 8) Increase the value of start_gen by 1;
[0102] 9) If start_gen does not exceed the latest generation value of the volume, that is, there is still data to be synchronized, repeat steps 6 to 8; if it has exceeded it, the data synchronization task is completed.
[0103] Specifically, in one embodiment, it can be determined whether the earliest generation of data to be synchronized is incremental data; when it is determined that the earliest generation of data to be synchronized is incremental data, the snapshot data of the previous generation of the earliest generation is used as the snapshot data to be analyzed; multiple spatial positions to be supplemented that have not been modified are screened in the earliest generation of data to be synchronized; for any spatial position to be supplemented, it is determined whether the snapshot data to be analyzed modifies the spatial position to be supplemented; when it is determined that the snapshot data to be analyzed modifies the spatial position to be supplemented, the earliest generation of data to be synchronized is supplemented according to the modification content of the snapshot data to be analyzed on the spatial position to be supplemented, so that the supplemented earliest generation of data to be synchronized becomes the full data.
[0104] It should be noted that since the auxiliary space initially has no data, a full copy is required when copying the first generation of data into it. Once the auxiliary space has a generation of data, only incremental data copies are required for subsequent generations. In order to achieve this function, when the mirror volume module reads the main space data, it must first determine whether the generation being read is the first generation in the auxiliary space. If so, when reading IO from the storage pool module, the data of this position in this generation is directly read. That is, after the IO request reaches the storage pool module, it will first search for any spatial position in this generation based on the recorded metadata information to see if there is any data modification. If so, the data is directly returned to the mirror volume module, that is, the earliest generation of data to be synchronized is full data. If not, it indicates that the earliest generation of data to be synchronized is incremental data. Therefore, it is searched again to see if this position (the multiple spatial positions to be supplemented that have not been modified in the earliest generation of data to be synchronized) has any data modification in the previous generation. If so, it is returned to supplement the earliest generation of data to be synchronized based on the modification content of the spatial position to be supplemented by the snapshot data to be analyzed, so that the supplemented earliest generation of data to be synchronized becomes full data.
[0105] Accordingly, in one embodiment, if the snapshot data to be analyzed has not modified the space location to be supplemented, the previous generation snapshot data of the snapshot data to be analyzed is used as the new snapshot data to be analyzed; the process returns to the step of determining whether the snapshot data to be analyzed has modified the space location to be supplemented, until the snapshot data to be analyzed has modified the space location to be supplemented or the snapshot data to be analyzed is the first generation snapshot data of the mirror volume.
[0106] Specifically, if there is no data modification in the previous generation, that is, the snapshot data to be analyzed has not been modified in the space location to be supplemented, then the previous generation will be searched again, and the search will continue until the 0th generation (the first generation snapshot data). In order to realize the above-mentioned incremental data judgment and incremental data supplement process, if the mirror volume module determines that the read generation is not the first generation in the auxiliary space, then when reading IO from the storage pool module, an additional control field READ_INCREMENT_DATA will be sent down. After the IO request arrives at the storage pool module with this control field, the storage pool module will still first search for whether there has been any data modification at this location in this generation based on the recorded metadata information. If so, the data will be returned. If not, the data will not be searched for in the previous generation again, but a special return value without difference data will be directly returned to the mirror volume. After receiving this special value, the mirror volume module directly considers that the location has completed synchronization, and only needs to modify the bitmap information, and then process the next location.
[0107] Based on the above embodiment, as an implementable manner, in one embodiment, the method further includes:
[0108] Step 301: When the target synchronization strategy is full synchronization, the snapshot data of all generations in the main space are used as data to be synchronized.
[0109] For example, Figure 4 The figure shows a schematic diagram of the data structure of another mirrored volume provided by an embodiment of the present application. When the target synchronization policy is full synchronization, all generations of data in the primary space are synchronized to the secondary space. After synchronization is complete, the corresponding snapshot data in both the primary and secondary spaces are completely consistent. This embodiment is mainly suitable for scenarios with extremely high requirements for high availability and historical data protection. When a storage space fails, the other storage space can still recover data to any historical CDP time.
[0110] Accordingly, in one embodiment, the data to be synchronized can be divided according to the spatial position of the data to be synchronized in the main space to obtain the data to be synchronized at each spatial position; and the data to be synchronized at each spatial position can be synchronized to the auxiliary space according to the order of the spatial positions.
[0111] Specifically, since the target synchronization strategy of full synchronization is generally applied to the initial synchronization of mirror volumes, in order to achieve fast synchronization, the data synchronization method of the above scheme (1) can be selected, that is, copying based on the spatial location grain, that is, traversing the grain, for each grain, reading out the 0th generation, 1st generation, ..., nth generation data in the grain respectively, and then writing them to the target space (auxiliary space). Since the newly added storage space is an unused space, there will be no data on it, so the generation of the written data can also be consistent with the generation read from the original space, that is, the 0th generation data of the original space is written to the 0th generation of the target space, and the nth generation data of the original space is written to the nth generation of the target space. When all the generation data of the grain are copied, the next grain is processed. The number of bitmap operations in this scheme is small, because all the generation data of the same grain are copied to the auxiliary space before the next grain is processed. Therefore, after a grain is processed, the corresponding position of the bitmap can be changed to 0. The bitmap is operated as many times as there are grains, which improves the data synchronization efficiency to a certain extent.
[0112] Based on the above embodiment, as an implementable manner, in one embodiment, the method further includes:
[0113] Step 401: When the target synchronization strategy is the latest synchronization, the snapshot data of the latest generation in the main space is used as the data to be synchronized.
[0114] For example, Figure 5 The figure shows another data structure diagram of a mirrored volume provided by an embodiment of the present application. When the target synchronization policy is the latest synchronization, only the latest data in the primary space is synchronized to the secondary space. This implementation is mainly suitable for scenarios that prioritize high data availability and where the loss of snapshot data is not a major concern. If a storage space fails, the other storage space can provide the latest data to ensure normal business operations, but all previously created CDP snapshots will be lost.
[0115] Based on the above embodiment, as an implementable manner, in one embodiment, the method further includes:
[0116] Step 501: If a failure occurs in the primary space while the primary space is synchronizing the latest snapshot data to the secondary space, a service rollback is performed on the mirror volume.
[0117] It should be noted that in related technologies, if a primary volume fails while synchronizing data from the primary to the secondary volume, the mirrored volume cannot be used until the primary volume is restored, as the secondary volume does not yet have a complete copy of the data. If an irreparable failure occurs on the primary volume, the data in the volume will be permanently unusable. To address this technical issue, the present embodiment of the application rolls back the mirrored volume if a failure occurs on the primary volume while the primary volume is synchronizing the latest snapshot data to the secondary volume, thereby ensuring the data integrity of the mirrored volume.
[0118] Specifically, when the primary space of a mirrored volume fails and becomes inaccessible, differentiated processing will be performed based on the status of the current data synchronization task. The processing logic is as follows:
[0119] If the primary and secondary spaces are not synchronizing data when a failure occurs, it indicates that the data in the primary and secondary spaces are the same at this time. Before the failure occurs, the host write IO will be double-written in the primary and secondary spaces, and the data in the two spaces will be consistent after the write is completed. After the failure occurs, the mirror volume module will mark the primary space as unavailable and downgrade it to the secondary space, and promote the original secondary space to the new primary space. When the host write IO reaches the mirror volume module, when the double write is about to be performed, it is found that the new secondary space (that is, the original failed primary space) is unavailable, then the IO will only be written to the new primary space (that is, the original secondary space). Since the location is only written in one space, there is a difference in the data of the primary and secondary spaces. Therefore, after the IO is written, the data at the corresponding position in the bitmap will be changed from 0 to 1 to record the difference at that location.
[0120] And because the new primary space retains the historical generation data of the original primary space (i.e., the current auxiliary space), that is, there are still n generations of CDP on the mirror volume at this time, the mirror volume can also use the relevant functions provided by the CDP module, such as: (1) If the volume data is tampered with, the CDP rollback function can be used to roll back the volume data to the data at a certain time before the tampering; (2) If the data at a certain historical moment of the volume is accessed, such as making an analysis report on the data at a certain time of a day, or doing corresponding tests, the physical volume conversion function provided by the CDP module can be used to map the historical data at a certain moment in the mirror volume to another volume, and the historical data of the mirror volume can be directly obtained by accessing the volume without directly rolling back the mirror volume data. In this way, the latest data of the mirror volume is retained and its historical data can be accessed.
[0121] When the failed space (formerly the primary space, now the secondary space) is restored, the mirror volume module will synchronize data again. During data synchronization, the data that has changed during the failure will be synchronized to the other space based on the information recorded in the bitmap.
[0122] Among them, if the data synchronization task from the primary space to the secondary space is in progress when the failure occurs, that is, the primary space fails while synchronizing the latest snapshot data from the primary space to the secondary space, then the secondary space does not have all the primary space data. It may only have the data of the previous generations that have been copied, and the first half of the data of a generation that is currently being copied. At this time, when the primary space fails, the secondary space cannot directly process the host IO. Here, it cannot directly process the host IO. It means that from the perspective of the business application on the server, it cannot directly process IO, because at this time the business application reads and writes the latest data on the volume, and the secondary space does not have the complete latest data at this time. Some may only have complete data from a certain historical moment. This data cannot be directly provided to the business application for use, and the business application needs to make corresponding adjustments before it can be used. Therefore, when this scenario occurs, the system will report an alarm message to inform the user that the primary space has failed during the synchronization of the mirror volume data, and the user needs to manually intervene to select the next processing strategy.
[0123] There are two possible handling strategies: The first is to maintain the status quo without intervention. This strategy can be chosen if the user believes a temporary suspension of service is acceptable, or if the outage occurred during a low-traffic period (such as the early morning hours). This strategy allows engineers to restore the primary space before resuming service. With this strategy, after the outage, the data in the primary and secondary spaces remains unchanged, and bitmap information is no longer updated. Therefore, the role transition between the primary and secondary spaces is not triggered (that is, the faulty primary space is not demoted to a secondary space, nor is a healthy secondary space promoted to the primary space). Once the primary space is restored, subsequent data synchronization tasks can be resumed from the point of interruption based on the bitmap information until all data synchronization is complete.
[0124] The second processing strategy is to intervene and roll back the business applications on the server and the data on the auxiliary storage space to the data at a certain available time (data that has been fully synchronized to a certain generation of the auxiliary space). That is, the business is rolled back to the mirror volume and the business is restarted using the data at that time. When the data on the auxiliary storage space is rolled back, the business application is developed and advanced based on the data at a certain historical moment.
[0125] This policy is designed to require intervention rather than automatic configuration because these operations may change the data on the volume (rolling back to a certain historical point) and require the cooperation of the business applications on the server (rolling back or maintaining the status quo). These operations have a significant impact on business data and must be judged based on the actual business scenario at the time of the failure. All of these require comprehensive judgment by the user.
[0126] Rolling back and restarting business applications is a function provided by the business application software, while data rollback on the secondary storage space is provided by the storage software's mirrored volume module. When this scenario occurs, after the user clicks the function button in the storage software, the mirrored volume module first marks the failed primary space as unavailable and demotes it to a secondary space. It then promotes the previously functioning secondary space to primary. The mirrored volume module then triggers a rollback of the latest and most complete generation of data on the secondary space (the generation immediately preceding the generation being synchronized at the time of the failure) and resets the bitmap data to all zeros.
[0127] After the service is activated, all write IOs issued by the host are written to the new primary space, and the bitmap at the corresponding position is changed to 1 to record the data difference. When the failed space is restored, manual intervention is still required to select whether the failed space will continue to serve as the secondary space after recovery or be promoted to the primary space again.
[0128] If the user believes that the data before the failure is more important, he or she can choose to promote the recovered space to the primary space. At this time, the mirror volume module will mark the failed space as normal and promote it to the primary space, downgrade the existing primary space to the secondary space, and use the previous generation of the generation that the secondary space was synchronizing when the failure occurred to roll back the data, rolling back to the same data as a historical generation of the primary space. Then, the bitmap will be reset to 1, and data synchronization from the primary space to the secondary space will be re-performed starting from this historical generation.
[0129] If the user believes that the data before the failure is no longer needed and the restored data is more important, they may choose to keep the restored space in the secondary space. The mirrored volume module will then mark the failed space as normal and roll back the data to the generation immediately preceding the generation being synchronized at the time of the failure. However, the bitmap information will not be reset. After the rollback is complete, the only difference between the primary and failed spaces is the newly written data after the service is restored. This data can then be incrementally synchronized to the failed space based on the bitmap information.
[0130] The embodiment of the present application provides a mirror volume data synchronization method, which obtains application scenario information of the mirror volume; wherein the application scenario information includes at least business information; determines the target synchronization strategy of the mirror volume according to the application scenario information of the mirror volume; in the case where the target synchronization strategy is partial synchronization, determines the read and write frequency information of the mirror volume according to the business information; determines the target synchronization generation value according to the read and write frequency information of the mirror volume; locates the latest snapshot data of the corresponding generation in the primary space according to the target synchronization generation value, and uses the latest snapshot data of the corresponding generation as the data to be synchronized; synchronizes the data to be synchronized to the secondary space; wherein the mirror volume includes the primary space and the secondary space. The method provided by the above scheme determines the target synchronization generation value that matches the application scenario information of the mirror volume, locates the latest snapshot data of the corresponding generation in the primary space according to the target synchronization generation value, uses the latest snapshot data of the corresponding generation as the data to be synchronized, and finally synchronizes the data to be synchronized to the secondary space, so that the primary space and the secondary space of the mirror volume both retain complete historical snapshot data. Even if the primary space fails, the secondary space can still provide historical snapshot data, thereby improving the high availability of the mirror volume. In addition, the CDP data synchronization policy for the mirrored volume in this application is optional. Users can flexibly select the target synchronization policy according to the actual application scenario, which is more user-friendly.
[0131] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0132] An embodiment of the present application further provides a mirror volume data synchronization device for executing the mirror volume data synchronization method provided in the above embodiment.
[0133] like Figure 6 FIG2 is a schematic diagram of the structure of the mirror volume data synchronization device provided by an embodiment of the present application. The mirror volume data synchronization device 60 includes: an acquisition module 601, a first determination module 602, a second determination module 603, a third determination module 604, a first synchronization module 605 and a second synchronization module 606.
[0134] Among them, the acquisition module is used to obtain the application scenario information of the mirror volume; wherein the application scenario information includes at least business information; the first determination module is used to determine the target synchronization strategy of the mirror volume according to the application scenario information of the mirror volume; the second determination module is used to determine the read and write frequency information of the mirror volume according to the business information when the target synchronization strategy is partial synchronization; the third determination module is used to determine the target synchronization algebra value according to the read and write frequency information of the mirror volume; the first synchronization module is used to locate the latest snapshot data of the corresponding algebra in the primary space according to the target synchronization algebra value, and use the latest snapshot data of the corresponding algebra as the data to be synchronized; the second synchronization module is used to synchronize the data to be synchronized to the auxiliary space.
[0135] The mirror volume includes primary space and secondary space.
[0136] For the description of the features in the embodiment corresponding to the mirror volume data synchronization device, reference can be made to the relevant description of the embodiment corresponding to the mirror volume data synchronization method, which will not be repeated here.
[0137] An embodiment of the present application further provides a mirror volume data synchronization system for executing the mirror volume data synchronization method provided in the above embodiment.
[0138] like Figure 7 FIG. 1 is a schematic diagram of the structure of a mirror volume data synchronization system provided by an embodiment of the present application. The mirror volume data synchronization system includes: a primary space, a secondary space, and a mirror volume data synchronization device.
[0139] The primary space and the secondary space belong to different storage pools; the primary space is used to store snapshot data; and the secondary space is used to synchronize the snapshot data stored in the primary space.
[0140] The mirror volume data synchronization device synchronizes the data to be synchronized in the primary space to the secondary space based on the mirror volume data synchronization method provided by the above embodiment.
[0141] For descriptions of features in the embodiments corresponding to the mirrored volume data synchronization system, reference may be made to the descriptions of the embodiments corresponding to the mirrored volume data synchronization method, which will not be described in detail here.
[0142] The embodiment of the present application also provides an electronic device, such as Figure 8 As shown, it is a structural diagram of an electronic device provided in an embodiment of the present application, including a processor 10 and a memory 20, wherein a computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any of the above-mentioned mirror volume data synchronization method embodiments.
[0143] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned mirror volume data synchronization method embodiments when running.
[0144] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0145] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned mirror volume data synchronization method embodiments are implemented.
[0146] An embodiment of the present application further provides another computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned mirror volume data synchronization method embodiments are implemented.
[0147] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0148] The above describes in detail the mirror volume data synchronization method, system, electronic device, and storage medium provided by this application. This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is intended only to facilitate understanding of the method and core concepts of this application. It should be noted that those skilled in the art may make various improvements and modifications to this application without departing from the principles of this application, and such improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A mirror volume data synchronization method, characterized in that: include: Obtaining application scenario information of the mirror volume; wherein the application scenario information at least includes business information; Determining a target synchronization strategy for the mirror volume according to application scenario information of the mirror volume; When the target synchronization strategy is partial synchronization, determining the read and write frequency information of the mirror volume according to the business information; Determining a target synchronization algebra value according to the read and write frequency information of the mirror volume; According to the target synchronization generation value, locate the latest snapshot data of the corresponding generation in the main space, and use the latest snapshot data of the corresponding generation as the data to be synchronized; Synchronize the data to be synchronized to the auxiliary space; The mirror volume includes the primary space and the secondary space.
2. The mirror volume data synchronization method according to claim 1, wherein: The determining, according to the application scenario information of the mirror volume, a target synchronization strategy of the mirror volume includes: Determining the user's protection requirements for historical snapshot data of the mirror volume based on the application scenario information of the mirror volume; A target synchronization strategy for the mirror volume is determined according to the user's requirement for protecting the historical snapshot data of the mirror volume.
3. The mirror volume data synchronization method according to claim 1, wherein: When the target synchronization strategy is partial synchronization, determining the read and write frequency information of the mirror volume according to the service information includes: When the target synchronization strategy is partial synchronization, determining, based on the service information, the number of write requests and the number of read requests responded to by the mirror volume within a fixed time period; Determining a read frequency and a write frequency of the mirror volume according to a number of write requests and a number of read requests responded to by the mirror volume within a fixed time period; The read and write frequency information at least includes the read frequency and the write frequency.
4. The mirror volume data synchronization method according to claim 1, wherein: The step of determining a target synchronization algebra value according to the read and write frequency information of the mirror volume includes: Determining the read frequency and write frequency of the mirror volume according to the read and write frequency information of the mirror volume; Determining a target synchronization algebra value according to a read frequency and a write frequency of the mirror volume; The target synchronization algebra value is positively correlated with the write frequency, and the target synchronization algebra value is negatively correlated with the read frequency.
5. The mirror volume data synchronization method according to claim 1, wherein: Synchronizing the data to be synchronized to the auxiliary space includes: Dividing the data to be synchronized according to the generations of the data to be synchronized to obtain data to be synchronized of each generation; Synchronize the data to be synchronized of each generation to the auxiliary space in the order of generations.
6. The mirror volume data synchronization method according to claim 5, characterized in that: Synchronizing the data to be synchronized of each generation to the auxiliary space in the order of generations includes: Determine whether the earliest generation of data to be synchronized is incremental data; When it is determined that the earliest generation of data to be synchronized is incremental data, the previous generation of snapshot data of the earliest generation is used as the snapshot data to be analyzed; Screening a plurality of unmodified to-be-supplemented spatial locations from the earliest generation of data to be synchronized; For any of the to-be-supplemented spatial locations, determining whether the to-be-supplemented spatial location is modified by the to-be-supplemented snapshot data; When it is determined that the snapshot data to be analyzed modifies the spatial position to be supplemented, the earliest generation of data to be synchronized is supplemented according to the modification content of the spatial position to be supplemented by the snapshot data to be analyzed, so that the supplemented earliest generation of data to be synchronized becomes the full data.
7. The mirror volume data synchronization method according to claim 6, wherein: The method further comprises: If the snapshot data to be analyzed does not modify the spatial position to be supplemented, taking the previous generation snapshot data of the snapshot data to be analyzed as the new snapshot data to be analyzed; Return to the step of determining whether the snapshot data to be analyzed modifies the location of the space to be supplemented, until the snapshot data to be analyzed modifies the location of the space to be supplemented or the snapshot data to be analyzed is the first-generation snapshot data of the mirror volume.
8. The mirror volume data synchronization method according to claim 1, wherein: The method further comprises: When the target synchronization strategy is full synchronization, the snapshot data of all generations in the main space are used as data to be synchronized.
9. The mirror volume data synchronization method according to claim 8, wherein: Synchronizing the data to be synchronized to the auxiliary space includes: Dividing the data to be synchronized according to the spatial positions of the data to be synchronized in the main space to obtain data to be synchronized at each spatial position; According to the order of the spatial positions, the data to be synchronized at each spatial position is synchronized to the auxiliary space.
10. The mirror volume data synchronization method according to claim 1, wherein: The method further comprises: When the target synchronization strategy is the latest synchronization, the snapshot data of the latest generation in the main space is used as the data to be synchronized.
11. The mirror volume data synchronization method according to claim 1, wherein: The method further comprises: If a failure occurs in the primary space while the primary space is synchronizing the latest snapshot data to the secondary space, the service of the mirror volume is rolled back.
12. A mirror volume data synchronization system, characterized in that: include: Data synchronization device for primary space, auxiliary space and mirror volume; The primary space and the secondary space belong to different storage pools; The main space is used to store snapshot data; The auxiliary space is used to synchronously store the snapshot data stored in the primary space; The mirror volume data synchronization device synchronizes the to-be-synchronized data in the primary space to the secondary space based on the mirror volume data synchronization method according to any one of claims 1 to 11.
13. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the mirror volume data synchronization method according to any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the mirror volume data synchronization method according to any one of claims 1 to 11 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the mirror volume data synchronization method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Continuous data protection system and method combining with snapshot technology
CN105389230A
Storage device, method and system for realizing disaster recovery backup
CN117950915A