Snapshot generation method, storage system and storage device
By generating snapshots when different hosts write operation instructions in the storage system and managing their generation and deletion, the problem of incomplete periodic snapshots is solved, and data recovery efficiency and system reliability are improved.
Patent Information
- Application Number
- CN202311866858.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, periodic snapshots may record incomplete or incorrect data, resulting in the inability to fully recover the data of the storage system before abnormal write operations, affecting the data recovery efficiency.
The storage system generates snapshots when receiving write instructions from different hosts, records a copy of the data before the write operation, and manages the generation and deletion of snapshots through timers and alarm mechanisms to ensure the effectiveness of the generated snapshots and resource utilization efficiency.
The complete rate of LUN data recovery based on snapshots is improved, the efficiency of data recovery is improved, the generation of invalid snapshots is reduced, the storage resources are saved, and the system reliability and fault location efficiency is improved.
Smart Images

Figure CN120233940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the storage field, and in particular, to a method for generating a snapshot, a storage system, and a storage device. Background Art
[0002] A logical unit number (LUN) is a logically divided storage unit, and the physical capacity of a storage system can be divided into multiple virtual logical volumes. By generating LUNs, the storage system can allocate storage resources to different hosts, which is conducive to flexible management and utilization of storage space. To ensure the reliability of services, a host usually adopts a cluster for redundancy protection. For example, in the primary / standby mode, the LUNs of the storage system are mapped to both the primary and standby hosts simultaneously. When one of the hosts is abnormal, the other host can read the data in the LUN and then quickly resume the service. To ensure data consistency, only one host (i.e., the primary host) issues a write operation instruction to the LUN of the storage system among the primary and standby hosts. However, when the host cluster software is abnormal (for example, the primary / standby heartbeat is abnormal, resulting in arbitration errors, etc.), a dual-primary state is generated, that is, both hosts issue write operation instructions to the LUN of the storage system simultaneously, which may lead to abnormal storage data. To prevent such abnormalities from damaging the useful data in the LUN, the industry uses the method of periodic snapshots to protect the data in the LUN.
[0003] Specifically, the storage system generates a snapshot of the LUN based on a certain snapshot generation period, and then the storage system can periodically save a copy of the data of the LUN, so that after an abnormality occurs, the storage system can recover the data of the LUN based on the periodic snapshot at a certain moment.
[0004] However, the periodic snapshots generated only at a certain period may record incomplete data or may record some incorrect data, resulting in the storage system may not be able to fully recover the data of the storage system before the abnormal write operation based on the periodic snapshot at a certain moment. Therefore, how the storage system generates snapshots to ensure the complete recovery of the data of the LUN is an urgent problem to be solved in the industry. Summary of the Invention
[0005] The present application provides a method for generating a snapshot, a storage system, and a storage device, which are used to improve the integrity rate of recovering the data of the LUN based on the snapshot, and further improve the efficiency of data recovery.
[0006] In a first aspect, the present application provides a method for generating a snapshot, which is applied to a storage system. The storage system may include only a single storage node, or may include multiple storage nodes. When the storage system includes multiple storage nodes, the storage system may be a distributed storage system or a centralized storage system. The method for generating a snapshot may be executed by the storage system, or by components of the storage system (such as components like a processor, a chip, or a chip system, etc.), or by a certain storage node in the storage system. Hereinafter, taking the storage system as an example, the storage system divides a physical storage array into at least one LUN, and maps one of the LUNs to at least two hosts, where the at least two hosts include a first host and a second host. The storage system sequentially receives a second write operation instruction from the second host and a first write operation instruction from the first host. Then, in the case where the first write operation instruction and the second write operation instruction come from different hosts respectively, the storage system generates a first snapshot of the LUN, and the first snapshot is used to record a data copy of the LUN before the execution of the first write operation instruction.
[0007] It should be understood that generally, the storage system receives write operation instructions from the same host (i.e., the primary host). When the two adjacent write operation instructions received by the storage system come from different hosts, it indicates that a normal switch has occurred between the primary and standby hosts, or a cluster software failure causes the primary and standby hosts to issue write operation instructions simultaneously. Therefore, the storage system protects the data of the LUN before the execution of the first write operation instruction by means of a snapshot, that is, generates a first snapshot to retain a data copy of the LUN for recovery before the execution of the first write operation instruction. This is beneficial for the storage system to perform data recovery using the first snapshot when the write operation instruction is an abnormal write operation instruction, which is beneficial for ensuring the complete recovery of the LUN data and thus improving the data recovery efficiency.
[0008] In a possible implementation manner, in the case where the first write operation instruction and the second write operation instruction come from different hosts respectively, the storage system generates a first snapshot of the LUN, including: when the first write operation instruction and the second write operation instruction come from different hosts in a host group, and the storage system has not generated a snapshot of the LUN, the storage system generates a first snapshot of the LUN.
[0009] In this embodiment, when the storage system discovers that two write operation instructions come from different hosts, the storage system first determines whether a snapshot has been generated for the LUN. When there is no snapshot for the LUN, the storage system triggers the generation of a first snapshot. This helps to avoid the storage system frequently generating snapshots when the primary and standby simultaneously issue write operations. On the one hand, only the snapshot generated during the first abnormal write operation is a valid snapshot (i.e., a snapshot that can completely restore the data before the abnormal write operation), while the snapshots generated later may be invalid snapshots (i.e., snapshots that record incorrect data); on the other hand, the number of snapshots that the storage system can store is limited, and multiple invalid snapshots generated later may overwrite the valid snapshots. Therefore, when the storage system has not generated a snapshot for the LUN, the storage system triggers the generation of a single snapshot for the LUN, which helps to save the processing resources and storage resources of the storage system and improve the validity of the generated snapshots.
[0010] In a possible implementation, the method further includes: after generating the first snapshot, if no write operation instruction from the second host is received again within a set time, the storage system deletes the first snapshot.
[0011] Exemplarily, when the storage system generates the first snapshot, the storage system also starts a first timer. During the operation of the first timer, the storage system may or may not receive a write operation instruction. If the storage system does not receive a write operation instruction from the second host before the first timer times out, the storage deletes the first snapshot. Here, the second host is a host different from the first host in the host group.
[0012] In this embodiment, the storage system does not receive a write operation instruction from a host other than the first host during the operation of the first timer. For example, within a set time (e.g., during the operation of the first timer), the storage system only receives a write operation instruction from the first host, or the storage system does not receive a write operation instruction from any host, indicating that the previous host switch (i.e., switching from the second host issuing the second write operation instruction to the first host issuing the first write operation instruction) is a normal primary-standby switch, and the storage system does not need to use the first snapshot for data recovery, so the storage system deletes the first snapshot. This helps to timely release the storage resources used for storing snapshots and improve the utilization efficiency of the storage resources.
[0013] In a possible implementation, when the storage system starts the first timer, the storage system also generates first identification information, which is used to indicate that snapshot protection has been enabled for the LUN. It can also be understood that the first identification information is used to indicate that the storage system has generated a snapshot of the LUN. If no write operation instruction from the second host is received before the first timer times out, the storage system deletes the first identification information.
[0014] In this embodiment, the first identification information generated by the storage system can be displayed to the user through a mapping view or other user interfaces, so as to prompt the user that snapshot protection has been enabled for the LUN, and improve the user experience of using the storage system.
[0015] In a possible implementation, the method further includes: after generating the first snapshot, if a write operation instruction from a second host is received within a set time, the storage system sends an alarm message.
[0016] Exemplarily, if the storage system receives a write operation instruction from a second host before the first timer times out, the storage system sends an alarm message. Among them, the alarm message is used to indicate that there is an abnormal write operation. It can be understood that the alarm message is used to indicate that the storage system receives an abnormal write operation instruction. For example, within a set time (for example, during the operation of the first timer), the storage system receives a third write operation instruction, and the storage system determines that the third write operation instruction and the first write operation instruction come from different hosts in the host group, then the storage system sends an alarm message.
[0017] In this embodiment, the storage system receives a write operation instruction from a host other than the first host during the operation of the first timer. For example, during the operation of the first timer, the storage system receives a write operation instruction from a second host (that is, a host other than the first host in the host group), indicating that the previous host switch (that is, switching from the second host issuing the second write operation instruction to the first host issuing the first write operation instruction) is an abnormal switch, and the storage system needs to trigger an alarm to avoid further damage to the data of the LUN. This is conducive to the storage system sending an alarm message in a timely manner, reducing the further damage of the data of the LUN by abnormal write operations, and thus conducive to improving the reliability of the storage system.
[0018] Optionally, the alarm message includes information about the LUN, indicating which LUN in the storage system the abnormal write operation points to. Exemplarily, the information about the LUN can be the LUN ID.
[0019] Optionally, the alarm message further includes information about the host group. The storage system may correspond to multiple host groups. The alarm message containing the information about the host group is conducive to the operation and maintenance personnel quickly locating the abnormal host group and improving the efficiency of fault location.
[0020] Optionally, the alarm information further includes information about the first host and the second host, indicating between which two hosts the abnormal write operation occurred. In addition to the first host and the second host, a host group may also include other hosts, and the abnormal write operation may only occur between the first host and the second host, that is, only the first host and the second host in the host group simultaneously send write operation instructions to the LUN, and the other hosts in the host group do not send write operation instructions to the LUN. Therefore, the alarm information includes information about the first host and the second host, which can indicate the hosts where the abnormal write operation occurred, facilitating the operation and maintenance personnel to quickly locate the abnormal hosts and improving the efficiency of fault location.
[0021] Optionally, the alarm information further includes information about the first snapshot, and the first snapshot is used to restore the data of the LUN before the first write operation instruction is executed. The alarm information including the first snapshot is beneficial for the operation and maintenance personnel to quickly determine the snapshot for data recovery and improve the efficiency of data recovery.
[0022] In a possible implementation manner, after the storage system sends the alarm information, the method further includes: the storage system receives a first indication information, where the first indication information is used to instruct the storage system to restore the data of the LUN before the first write operation instruction is executed based on the first snapshot; then, the storage system restores the data of the LUN before the first write operation instruction is executed based on the first snapshot.
[0023] Optionally, for the storage system to restore the data of the LUN, the storage system may roll back based on the first snapshot, or copy the data of the first snapshot to a newly created LUN in the storage system, which is not limited in this application.
[0024] In a possible implementation manner, after the storage system restores the data of the LUN before the first write operation instruction is executed based on the first snapshot, the method further includes: the storage system deletes the first snapshot. Optionally, if the storage system previously generated the first identification information, the storage system will also delete the first identification information. This is beneficial for timely releasing the storage resources used to store the snapshot and the first identification information and improving the utilization efficiency of the storage resources.
[0025] In a possible implementation, at least two hosts include a first host and a second host. The first write operation instruction includes identification information of the first host, and the second write operation instruction includes identification information of the second host. The method further includes: when the identification information of the first host is different from the identification information of the second host, the storage system determines that the first write operation instruction and the second write operation instruction respectively come from different hosts in the host group. Exemplarily, the identification information of the host may be the identity document (ID) of the host's initiator, may also be the ID of the host bus adapter (HBA) card, or may also be the ID of the world wide number (WWN) of the host's hardware, etc., which is not limited in this application.
[0026] In this embodiment, the storage system can identify which host a write operation instruction comes from based on the identification information of the host in the write operation instruction, which is beneficial for the storage system to accurately identify whether two adjacent received write operation instructions come from the same host, and further beneficial for the storage system to timely trigger the generation of the first snapshot or execute the first write operation instruction. This is further beneficial for improving the processing efficiency of the storage system.
[0027] In a possible implementation, before the storage system receives the first write operation instruction from the first host, the method further includes: the storage system receives first configuration information, and the first configuration information is used to indicate that the host group associated with the LUN enables the primary-secondary consistency protection function.
[0028] In this implementation, the primary-secondary consistency protection function means that whenever a write operation instruction from a host is received, the write operation instruction is not directly executed, but rather it is decided whether a snapshot needs to be generated to protect the data of the LUN. That is to say, the method for generating a snapshot provided in this application can be configured to be enabled on demand through the first configuration information, which is not only beneficial for improving the flexibility of snapshot generation, but also beneficial for saving the storage resources for storing snapshots.
[0029] Second aspect, the present application provides an implementation of a storage device. The storage device may be the storage system involved in the foregoing first aspect, or a storage node within the storage system, or a functional module or chip within the storage system. The storage device may include a processing module and a transceiver module. When the storage device is a storage system or a storage node, the processing module may be a processor, and the transceiver module may be a transceiver; the storage device may further include a storage module, and the storage module may be a memory; the storage module is used to store instructions, and the processing module executes the instructions stored in the storage module, so that the storage device executes the method in the first aspect or any implementation manner of the first aspect. When the storage device is a functional module or chip within the storage system, the processing module may be a processor, and the transceiver module may be an input / output interface, a pin, a circuit, etc.; the processing module executes the instructions stored in the storage module, so that the storage device executes the method in the first aspect or any implementation manner of the first aspect. The storage module may be a storage module within the chip (for example, a register, a cache, etc.), or a storage module outside the chip within the storage device (for example, a read-only memory, a random access memory, etc.).
[0030] Third aspect, the present application provides a storage system. The storage system may be a centralized storage system or a distributed storage system. If the storage system is a centralized storage system, the storage system may be a disk control separation architecture or a disk control integrated architecture. The storage system includes a processor and a memory. The processor is coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the device is caused to execute the method described in any implementation manner of the foregoing aspects.
[0031] Fourth aspect, the present application provides an implementation of a computer program product containing instructions. When it runs on a computer, the computer is caused to execute the method described in any implementation manner of the foregoing aspects.
[0032] Fifth aspect, the present application provides an implementation of a computer-readable storage medium, including instructions. When the instructions run on a computer, the computer is caused to execute the method described in any implementation manner of the foregoing aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1A It is an example diagram of an application scenario of the method for generating a snapshot provided by the present application;
[0034] Figure 1B It is an example diagram of generating a periodic snapshot by a storage system in the prior art;
[0035] Figure 2A flowchart of the method for generating a snapshot provided by this application;
[0036] Figure 3 An example diagram of the snapshot generated by the method for generating a snapshot provided by this application;
[0037] Figure 4 Another flowchart of the method for generating a snapshot provided by this application;
[0038] Figure 5A An example diagram of starting the first timer and triggering an alarm in the method for generating a snapshot provided by this application;
[0039] Figure 5B An example diagram of deleting the first snapshot in the method for generating a snapshot provided by this application;
[0040] Figure 6 A schematic diagram of an embodiment of the storage system provided by this application;
[0041] Figure 7 A schematic diagram of an embodiment of the storage device provided by this application. Detailed implementation manners
[0042] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments.
[0043] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0044] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be single or multiple. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, the expression "at least one of the following" or its similar expressions in this text are used to represent any combination of the listed items. For example, at least one of A, B, and (or) C can represent the following situations: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, B and C exist simultaneously, A and C exist simultaneously, and A, B, and C exist simultaneously. Here, A, B, and C can be single or multiple.
[0045] First, the application scenarios and system architectures applicable to the method for generating snapshots provided in this application will be introduced as follows:
[0046] The method for generating snapshots provided in this application is mainly applied to the scenario where a storage system allocates storage resources to different hosts through LUNs.
[0047] As Figure 1A shown, this scenario mainly involves a storage system and a host group.
[0048] Among them, the storage system can include only a single storage node or multiple storage nodes. When the storage system includes multiple storage nodes, the storage system can be a distributed storage system or a centralized storage system. The storage node includes a physical storage array, and the storage system can virtualize the physical storage arrays from one storage node or multiple storage nodes into at least one logical unit (LU) for use by the host. Each LU has a unique logical unit number (LUN). Since the host can directly perceive the logical unit number LUN, those skilled in the art usually directly use LUN to refer to the logical unit LU. Each LUN has a LUN ID, and the LUN ID is used to identify the LUN.
[0049] In addition, the host group includes at least two hosts. The user accesses the data in the storage device through the applications running on the hosts, and the host running the aforementioned applications is called an "application server". It should be understood that the aforementioned hosts can be physical machines, virtual machines, or operating systems. For example, a host group containing two hosts can be two independent physical machines, two virtual hosts running on the same physical machine, or two virtual hosts running on different physical machines. This application does not impose any restrictions. It should be understood that the aforementioned physical machines include, but are not limited to, desktop computers, servers, laptops, and mobile devices, etc. The aforementioned hosts access the storage system through a switch (e.g., a fiber optic switch or a network switch) to store and retrieve data; or, the aforementioned hosts directly access the storage system to store and retrieve data.
[0050] The storage system can map one of the LUNs to at least two hosts in the host group, so that the at least two hosts can control the LUN through the upper-layer cluster software. For example, the hosts in the host group can send write operation instructions and / or read operation instructions to the storage system. Generally, in order to ensure the reliability of the service, at least two hosts in the host group adopt the primary and standby mode for redundancy protection. For example, one of the hosts in the host group serves as the primary host to access the data of the LUN in the storage system to run the service, and the remaining hosts in the host group serve as standby hosts, which can read the data in the LUN when the primary host is abnormal to quickly resume the service. Exemplarily, as Figure 1A shown, taking the host group including Host 1 and Host 2 as an example. Host 1 and Host 2 are respectively connected to the storage system through Switch 1 and Switch 2. The storage system maps one LUN to both Host 1 and Host 2 at the same time, so that Host 1 and Host 2 can control the LUN through the upper-layer cluster software. In order to ensure the reliability of the service, Host 1 and Host 2 adopt the primary and standby mode for redundancy protection. For example, Host 1 is the primary host and Host 2 is the standby host. When working normally, Host 1 accesses the data of the LUN in the storage system to run the service. When Host 1 is abnormal, it switches to Host 2 to read the data in the LUN to quickly resume the service.
[0051] In order to ensure data consistency, among the primary and standby hosts, only one host (i.e., the primary host) sends write operation instructions to the LUN of the storage system. For example, Figure 1A in the example shown, when the cluster software in Host 1 and Host 2 is working normally, usually only Host 1, which is the primary host, sends write operation instructions to the storage system. If Host 1 is abnormal or there is a need to switch hosts for other reasons, then Host 2 is switched to the primary host. At this time, Host 2 is allowed to send write operation instructions.
[0052] However, when the host cluster software is abnormal (for example, the arbitration error is caused by the abnormal heartbeat between the primary and standby hosts), a dual-primary state is generated, that is, two hosts simultaneously issue write operation instructions to the LUN of the storage system. For the storage system in the traditional technology, the storage system does not care which host the write operation instruction comes from. Whenever the storage system receives a write operation instruction, the storage system executes the write operation instruction, which may lead to abnormal storage data.
[0053] To prevent the useful data in the LUN from being damaged due to such abnormalities, the industry uses the method of periodic snapshots to protect the data in the LUN. Among them, periodic snapshots refer to the storage system generating snapshots of the LUN at regular intervals to record the data backup of the LUN. However, the periodic snapshots generated only at a certain period may record incomplete data or may record some incorrect data.
[0054] Exemplarily, as Figure 1B shown, the host group includes Host 1 and Host 2, and the storage system maps the LUN to Host 1 and Host 2. At time T1, Host 1, as the primary host, issues write operation instruction 1 to the storage system. Before time T2, Host 1 issues write operation instruction 2 to the storage system again. Subsequently, the cluster software is abnormal, and Host 2, which was originally the standby host, successively issues write operation instruction 3 and write operation instruction 4 to the storage system, and Host 1 issues write operation instruction 5 to the storage system, etc. During the process of Host 1 and Host 2 issuing write operation instructions, the storage system treats each received write operation instruction as a normal write operation instruction, and, according to the pre-configured snapshot period, generates snapshot - T1, snapshot - T2, snapshot - T3, and snapshot - T4 at times T1, T2, T3, and T4 respectively. At this time, the operation and maintenance personnel need to use one of the snapshots to restore the data of the LUN after the execution of write operation instruction 2 and before the execution of write operation instruction 3. However, snapshot - T1 only records the data of write operation instruction 1 and lacks the data of write operation instruction 2, that is, the data recorded by snapshot - T1 is incomplete; snapshot - T2 not only records the data of write operation instruction 1 and write operation instruction 2, but also records the abnormal write operation instruction 3 data, resulting in snapshot - T2 recording some incorrect data. It can be seen that the periodic snapshots in the traditional technology may not be able to completely restore the data of the storage system before the execution of abnormal write operations based on the periodic snapshots at a certain moment.
[0055] In response to this, the present application provides a method for generating snapshots, a storage system, and a storage device, which are used to generate snapshots of the LUN when abnormal write operations occur, and are beneficial to improving the integrity rate of restoring the data of the LUN based on the snapshots, thereby improving the efficiency of data recovery.
[0056] Next, in combination with Figure 2 the main process of the method for generating snapshots provided by the present application will be introduced. As Figure 2As shown in the figure, the storage system mainly performs the following steps:
[0057] Step 201, the storage system receives a first write operation instruction from the first host.
[0058] Among them, a certain LUN in the storage system is mapped to at least two hosts, and the at least two hosts can be called a host group. The at least two hosts include a primary host and at least one standby host. Among them, the primary host is the host that runs services for accessing the data in the storage system currently, and the standby host can also access the data in the storage system, but only runs services based on the data in the storage system when the primary host is abnormal. In this step, the storage system receives the first write operation instruction from the first host, which can be understood as that the storage system receives the write operation instruction issued by a certain host in the host group.
[0059] In this embodiment, the storage system stores the source information of the previous write operation instruction. For example, the storage system stores the identification information of the host that issued the previous write operation instruction. Among them, the previous write operation instruction refers to the previous write operation instruction received by the storage device before receiving the first write operation instruction, which is called the second write operation instruction hereinafter.
[0060] The storage system stores the source information of the previous write operation instruction so that when receiving a new write operation instruction (for example, the first write operation instruction), the storage system can compare the source of the newly received write operation instruction with the source of the previous write operation instruction, and then determine whether the two write operation instructions come from different hosts. For example, after the storage system receives the first write operation instruction, the storage system will compare the information of the host that issued the first write operation instruction and the information of the host that issued the second write operation instruction, that is, the storage system determines whether the first write operation instruction and the second write operation instruction come from the same host in the host group or from different hosts in the host group respectively.
[0061] When the first write operation instruction and the second write operation instruction come from the same host in the host group, the storage system executes the first write operation instruction and does not generate a snapshot of the LUN. When the first write operation instruction and the second write operation instruction come from different hosts in the host group, the storage system executes step 202.
[0062] Step 202, when the first write operation instruction and the second write operation instruction come from different hosts respectively, the storage system generates a first snapshot of the LUN, and the second write operation instruction is the previous write operation instruction received by the storage system before receiving the first write operation instruction.
[0063] Among them, the first snapshot is used to record the data copy of the LUN before executing the first write operation instruction. For example, the first snapshot is used to record the latest data copy of the LUN before executing the first write operation instruction.
[0064] For example, when the storage system determines that two consecutively received write operation instructions come from different hosts in the host group, the storage system generates a first snapshot of the LUN before executing the currently received first write operation instruction to save the data of the storage system before executing the first write operation instruction. Since the storage system maps the LUN to a primary host and at least one standby host, generally the write operation instruction is issued by the primary host. Correspondingly, the write operation instructions received by the storage system should also come from the same host (i.e., the primary host). When two adjacent write operation instructions received by the storage system come from different hosts, it indicates that a normal switch has occurred between the primary and standby hosts, or the cluster software failure causes the primary and standby hosts to issue write operation instructions simultaneously. Therefore, the storage system protects the data of the LUN before executing the first write operation instruction by taking a snapshot, that is, generating a first snapshot to retain a data copy for restoring the LUN data before executing the first write operation instruction. Therefore, when the write operation instruction is an abnormal write operation instruction, it is beneficial for the storage system to use the first snapshot for data recovery, which is beneficial to ensuring the complete recovery of the LUN data and thus improving the data recovery efficiency.
[0065] Exemplarily, as Figure 3 shown, the storage system receives the write operation instruction 2 from host 1 at time t0 and receives the write operation instruction 3 from host 2 at time t1. Since the write operation instruction 2 and the write operation instruction 3 come from different hosts, the storage system generates snapshot - t1 before executing the write operation instruction 3, and this snapshot - t1 is used to record the data copy of the LUN of the storage system before executing the write operation instruction 3. In this example, the write operation instruction 3 is an example of the first write operation instruction, the write operation instruction 2 is an example of the second write operation instruction, and the snapshot - t1 is an example of the first snapshot. In this example, since the snapshot - t1 records the data of the write operation instructions 1 and 2 from host 1 and does not record the data of the abnormal write operation instruction 3 from host 2, the data copy in the snapshot - t1 is not affected by the write operation instruction 3. It is beneficial for the storage system to completely restore the data of the LUN when the write operation instruction 3 is not executed based on the snapshot - t1.
[0066] It should be noted that the storage system may or may not execute the first write operation instruction. For example, in Figure 3 the shown example, after the storage system generates the snapshot - t1, the storage system may or may not execute the write operation instruction 3. This embodiment does not limit whether the storage system executes the first write operation instruction, and more specific embodiments will be introduced later.
[0067] In this embodiment, since the storage system can identify the host that issues the write operation instruction, when the first write operation instruction received is from a different host in the host group than the previous received write operation instruction, the storage system generates a first snapshot of the LUN to store a data copy of the LUN before executing the first write operation instruction. That is to say, when the host that issues the write operation instruction changes, the storage system generates a snapshot of the LUN before executing the current write operation instruction, which is beneficial for the storage system to use this first snapshot for data recovery when the write operation instruction is an abnormal write operation instruction, beneficial for ensuring the complete recovery of the LUN data, and thus improving the efficiency of data recovery.
[0068] Next, in combination with Figure 4 the method for generating snapshots provided by this application will be further introduced. As Figure 4 shown, the storage system mainly performs the following steps:
[0069] Step 401, the storage system receives a second write operation instruction from the second host.
[0070] Step 402, the storage system receives a first write operation instruction from the first host.
[0071] Among them, a LUN of the storage system is mapped to a host group, and the host group includes at least two hosts. The at least two hosts include the first host and the second host. For the introduction of the storage system and the host group, please refer to the previous step 201, which will not be elaborated here.
[0072] In this embodiment, the second write operation instruction is the previous write operation instruction received by the storage system before receiving the first write operation instruction. For the introduction of the first write operation instruction and the second write operation instruction, please refer to the previous step 201, which will not be elaborated here.
[0073] In a possible implementation manner, before step 401, the storage system receives first configuration information, and the first configuration information is used to indicate that the host group associated with the LUN enables the master-slave consistency protection function. The master-slave consistency protection function means that whenever a write operation instruction from a host is received, the storage system does not directly execute the write operation instruction, but decides whether to generate a snapshot to protect the data of the LUN before executing the write operation instruction. For example, the storage system configured with the master-slave consistency protection function will execute step 403 after receiving the first write operation instruction in step 402, while the storage system not configured with the master-slave consistency protection function directly executes the first write operation instruction after receiving the first write operation instruction, without executing Figure 4 the subsequent steps of the corresponding embodiment. It should be understood that in practical applications, the master-slave consistency protection function may also have other names, which are not limited in this application.
[0074] Optionally, the first configuration information can be host group configuration or mapping view configuration. In one example, the first configuration information is host group configuration, indicating that when the host group accesses all LUNs in the storage system, the host group enables the active-standby consistency protection function. For example, the physical storage array of the storage system is divided into LUN1 and LUN2. If the storage system maps LUN1 to host group 1, and the storage system also maps LUN2 to host group 1, then at least two hosts in host group 1 will enable the active-standby consistency protection function when accessing LUN1, and at least two hosts in host group 1 will also enable the active-standby consistency protection function when accessing LUN2. In another example, the first configuration information is mapping view configuration, indicating that the host group associated with a specified LUN in the storage system enables the active-standby consistency protection function. For example, the physical storage array of the storage system is divided into LUN1 and LUN2, and the first configuration information is set only in the mapping view of LUN1. Then, even if the storage system maps LUN1 to host group 1 and the storage system also maps LUN2 to host group 1, only at least two hosts in host group 1 will enable the active-standby consistency protection function when accessing LUN1, while at least two hosts in host group 1 will not enable the active-standby consistency protection function when accessing LUN2.
[0075] After the storage system receives the first write operation instruction, the storage system will execute step 403.
[0076] Step 403, the storage system determines whether the first write operation instruction and the second write operation instruction come from different hosts respectively.
[0077] Specifically, after the storage system receives the first write operation instruction, the storage system will compare the source of the first write operation instruction with the source of the previous write operation instruction (hereinafter referred to as the second write operation instruction), that is, the storage system compares the information of the host carried by the first write operation instruction with the information of the host carried by the second write operation instruction to determine whether the two write operation instructions come from the same host in the host group or from different hosts in the host group respectively.
[0078] Optionally, the storage system determines whether the two write operation instructions come from different hosts based on the identification information of the hosts in the two write operation instructions (i.e., the first write operation instruction and the second write operation instruction).
[0079] Optionally, the write operation instruction includes the identification information of the host that issues the write operation instruction. The storage system can determine which host in the host group the write operation instruction comes from based on the identification information of the host included in the write operation instruction, and further determine whether the first write operation instruction and the second write operation instruction come from the same host. In one example, the first write operation instruction includes the identification information of the first host, and the storage system determines that the first write operation instruction comes from the first host based on the identification information of the first host in the first write operation instruction. The second write operation instruction includes the identification information of the second host, and the storage system determines that the second write operation instruction comes from the second host based on the identification information of the second host in the second write operation instruction. Since the identification information of the first host is different from the identification information of the second host, the storage system determines that the first write operation instruction and the second write operation instruction come from different hosts in the host group.
[0080] Optionally, the identification information of the host included in the write operation instruction refers to the information that can uniquely identify the host in the host group. Exemplarily, the identification information of the host can be the identity document (ID) of the host's initiator, or the ID of the host bus adapter (HBA) card, or the ID of the world wide number (WWN) of the host's hardware, etc., which is not limited in this application.
[0081] In this embodiment, when the first write operation instruction and the second write operation instruction come from the same host in the host group, the storage system executes the first write operation instruction, and the storage system does not perform the subsequent steps. When the first write operation instruction and the second write operation instruction come from different hosts in the host group, the storage system executes step 404a and step 404b.
[0082] Optionally, the storage system also executes step 404c. For example, while the storage system executes step 404a and step 404b, the storage system also executes step 404c.
[0083] Step 404a, the storage system generates a first snapshot of the LUN.
[0084] The first snapshot is used to record the data copy of the LUN before the first write operation instruction is executed.
[0085] When the storage system determines that the first write operation instruction and the second write operation instruction come from different hosts in the host group, the storage system generates a first snapshot of the LUN. The LUN is the LUN corresponding to the first write operation instruction, that is, the LUN for which the first write operation instruction is to be executed. For example, the storage system includes LUN1 and LUN2, and both LUN1 and LUN2 are mapped to the host group. If the first write operation instruction is used to perform a write operation on LUN1, when the storage system determines that the first write operation instruction and the second write operation instruction come from different hosts in the host group, the storage system generates a snapshot of LUN1, rather than generating a snapshot of LUN2.
[0086] Optionally, the storage system generates the first snapshot of the LUN only when the first write operation instruction and the second write operation instruction come from different hosts in the host group respectively and no snapshot of the LUN has been generated. It can be understood that the storage system is configured to trigger the generation of a snapshot of the LUN only when the currently received write operation instruction and the previous write operation instruction come from different hosts and the storage system has not generated a snapshot of the LUN. Since when the cluster software is abnormal, the primary host and the standby host almost simultaneously issue write operation instructions. If the storage system generates a snapshot every time it finds that two adjacent write operation instructions come from different hosts, it may cause the storage system to generate multiple snapshots in a short period of time. However, on the one hand, only the snapshot generated during the first abnormal write operation among the multiple snapshots generated by the storage system may be a valid snapshot (that is, a snapshot that can completely restore the data before the abnormal write operation), while the snapshots generated later may be invalid snapshots (that is, snapshots that record incorrect data); on the other hand, the number of snapshots that the storage system can store is limited, and multiple invalid snapshots generated later may overwrite the valid snapshots. Therefore, the storage system triggers the generation of a snapshot of the LUN only when no snapshot of the LUN has been generated, which is beneficial to saving the processing resources and storage resources of the storage system and improving the validity of the generated snapshots.
[0087] Optionally, when the first write operation instruction and the second write operation instruction come from different hosts in the host group, the storage system not only generates the first snapshot of the LUN, but also generates first identification information. The first identification information is used to indicate that snapshot protection has been started for the LUN. It can also be understood that the first identification information is used to indicate that the storage system has generated a snapshot of the LUN. Optionally, the first identification information generated by the storage system can be displayed to the user through a mapping view or other user interfaces to prompt the user that snapshot protection has been started for the LUN.
[0088] Step 404b, the storage system starts a first timer.
[0089] Among them, the duration of the first timer is a pre-configured duration.
[0090] When the storage system determines that the first write operation instruction and the second write operation instruction are from different hosts in the host group, the storage system starts the first timer. Exemplarily, as Figure 5A shown, the storage system receives the write operation instruction 2 from host 1 at time t0 and the write operation instruction 3 from host 2 at time t1. The write operation instruction 2 and the write operation instruction 3 are from different hosts, and the storage system starts the first timer.
[0091] It should be understood that when the storage system determines that the first write operation instruction and the second write operation instruction are from different hosts, the storage system is not yet sufficient to determine the reason why the first write operation instruction and the second write operation instruction are from different hosts. Exemplarily, a normal switch between the primary and standby hosts will cause the storage system to receive two write operation instructions from different hosts; or, a cluster software failure that causes the primary and standby hosts to issue write operation instructions simultaneously will also cause the storage system to receive two write operation instructions from different hosts. In this embodiment, the storage system starts the first timer to determine the reason why the first write operation instruction and the second write operation instruction are from different hosts within a set time. Specifically, the storage system will detect whether it can still receive the write operation instruction from the first host within the set time (i.e., during the operation of the first timer), or only receive the write operation instruction from other hosts (e.g., the second host), and then determine the reason for the storage system to receive two write operation instructions from different hosts based on the situation of the write operation instructions received within the set time.
[0092] Step 404c, the storage system determines that the first host is the primary host.
[0093] In this embodiment, step 404c is an optional step. For example, when the first write operation instruction and the second write operation instruction are from different hosts in the host group, the storage system not only generates the first snapshot of the LUN and starts the first timer, but also determines that the first host is the primary host.
[0094] For example, the storage system records the identification information of the first host as the identification information of the primary host. When the storage system receives the next write operation instruction, the storage system compares the identification information of the host included in the newly received write operation instruction with the identification information of the primary host (i.e., the identification information of the first host) to determine whether the newly received write operation instruction still comes from the primary host (i.e., the first host).
[0095] For example, if the storage system stores the identification information of the second host that sends the second write operation instruction, when the first write operation instruction and the second write operation instruction come from different hosts respectively, the storage system not only stores the identification information of the first host that sends the first write operation instruction, but also deletes the identification information of the second host. When the storage system receives the next write operation instruction, the storage system compares the identification information of the host included in the newly received write operation instruction with the identification information of the first host to determine whether the newly received write operation instruction still comes from the first host.
[0096] Step 405, before the first timer times out, the storage system determines whether it has received a write operation instruction from the second host.
[0097] Wherein, the second host is a host different from the first host in the host group.
[0098] After the storage system starts the first timer, the storage system may also receive write operation instructions from one or more hosts in the host group. The storage system will determine whether the received write operation instruction still comes from the first host or from other hosts (such as the second host) in the host group.
[0099] If the storage system receives a write operation instruction from the second host before the first timer times out, it means that the storage system receives write operation instructions from two different hosts because of a cluster software failure that causes the primary and standby hosts to issue write operation instructions simultaneously. In this case, the storage system sends an alarm message. For details, please refer to step 406a; if the storage system does not receive a write operation instruction from the second host before the first timer times out, it means that the storage system receives write operation instructions from two different hosts due to a normal switch between the primary and standby hosts. Then the storage system does not need to use the first snapshot to restore the data of the LUN, and the storage system deletes the first snapshot. For details, please refer to step 406b.
[0100] Step 406a, the storage system sends an alarm message.
[0101] If the storage system receives a write operation instruction from the second host before the first timer times out, it means that the storage system receives write operation instructions from two different hosts because of a cluster software failure that causes the primary and standby hosts to issue write operation instructions simultaneously. In this case, the storage system sends an alarm message. For example, during the operation of the first timer, the storage system receives a third write operation instruction, and the storage system determines that the third write operation instruction and the first write operation instruction come from different hosts in the host group respectively. Then the storage system sends an alarm message.
[0102] Exemplarily, such as Figure 5AAs shown, the storage system generates snapshot - t1 at time t1 and starts the first timer. During the operation of the first timer, the storage system receives a write operation instruction 4 from host 2 at time t2. Both write operation instruction 4 and write operation instruction 3 are from host 2, and the storage system does not send an alarm. Subsequently, the storage system receives a write operation instruction 5 from host 1 at time t3. Write operation instruction 5 and write operation instruction 4 are from different hosts, and the storage system sends an alarm message at time t3.
[0103] Optionally, the storage system sends the alarm message to the management device, and the management device displays the alarm message to the operation and maintenance personnel or users through output devices such as a display and a microphone; or, the storage system is connected to output devices such as a display and a microphone, and displays the alarm message to the operation and maintenance personnel or users through the aforementioned output devices.
[0104] Among them, the alarm message is used to indicate that there is an abnormal write operation. It can be understood that the alarm message is used to indicate that the storage system has received an abnormal write operation instruction.
[0105] Among them, the alarm message includes information about the LUN, indicating which LUN in the storage system the abnormal write operation points to. Exemplarily, the information of the LUN can be the LUN ID.
[0106] Optionally, the alarm message further includes information about the host group. The storage system may correspond to multiple host groups. The alarm message containing the information about the host group is beneficial for the operation and maintenance personnel to quickly locate the abnormal host group and improve the efficiency of fault location.
[0107] Optionally, the alarm message further includes information about the first host and the second host, indicating between which two hosts the abnormal write operation occurs. In addition to the first host and the second host, a host group may also include other hosts, and the abnormal write operation may only occur between the first host and the second host, that is, only the first host and the second host in the host group simultaneously send write operation instructions to the LUN, and the other hosts in the host group do not send write operation instructions to the LUN. Therefore, the alarm message containing the information about the first host and the second host can indicate the hosts where the abnormal write operation occurs, which is beneficial for the operation and maintenance personnel to quickly locate the abnormal hosts and improve the efficiency of fault location.
[0108] Optionally, the alarm message further includes information about the first snapshot. The first snapshot is used to restore the data of the LUN before the execution of the first write operation instruction. The alarm message containing the first snapshot is beneficial for the operation and maintenance personnel to quickly determine the snapshot for data recovery and improve the efficiency of data recovery.
[0109] In addition, after the storage system sends the alarm message, the storage system may also execute step 407.
[0110] Step 406b, the storage system deletes the first snapshot.
[0111] If the storage system does not receive a write operation instruction from the second host before the first timer times out, it indicates that the storage system receiving write operation instructions from two different hosts is due to a normal switch between the primary and standby hosts. In this case, the storage system does not need to use the first snapshot to restore the data of the LUN. Therefore, the storage system deletes the first snapshot. For example, during the operation of the first timer, the storage system does not receive a write operation instruction, or the write operation instruction received by the storage system is from the first host, indicating that the aforementioned switch from the second write operation instruction to the first write operation instruction is a normal primary-standby switch of the storage system, and the data of the LUN has not been damaged by abnormal write operations. The storage system does not need to use the first snapshot for data recovery. At this time, the storage system deletes the first snapshot. This helps to release the storage resources used for storing snapshots in a timely manner and improve the utilization efficiency of storage resources.
[0112] Exemplarily, as Figure 5B shown, the storage system generates snapshot - t1 at time t1 and starts the first timer. During the operation of the first timer, the storage system receives write operation instruction 4 from host 2 at time t2. Until the first timer times out at time t4, the storage system does not receive a write operation instruction from host 1, and the storage system deletes snapshot - t1 at time t4.
[0113] Optionally, if the storage system also generates the first identification information when generating the first snapshot, when the storage system deletes the first snapshot, the storage system also deletes the first identification information. This helps to release the storage resources in the storage system in a timely manner and improve the utilization efficiency of storage resources.
[0114] In this embodiment, steps 407 to 409 are optional steps.
[0115] Step 407, the storage system receives the first indication information.
[0116] Among them, the first indication information is used to instruct the storage system to restore the data of the LUN before executing the first write operation instruction based on the first snapshot.
[0117] Step 408, the storage system restores the data of the LUN before executing the first write operation instruction based on the first snapshot.
[0118] Optionally, for the storage system to restore the data of the LUN, it can be that the storage system performs a rollback based on the first snapshot, or it can be that the storage system copies the data of the first snapshot to a newly created LUN in the storage system. This application does not limit.
[0119] Step 409, the storage system deletes the first snapshot.
[0120] After the storage system restores the data of the LUN based on the first snapshot, the storage system may delete the first snapshot to release storage resources of the storage system. Optionally, the storage system may also delete the first identification information while deleting the first snapshot.
[0121] In this embodiment, the storage system can identify the host that issues the write operation instruction. When the first write operation instruction received is from a different host in the host group from the last received write operation instruction, the storage system generates a first snapshot of the LUN before executing the first write operation instruction and starts the first timer. If the storage system receives a write operation instruction from another host (i.e., a host other than the first host) again before the first timer times out, the storage system sends an alarm message and restores the data of the LUN based on the first snapshot under the action of the first indication information. It can be seen that when the host that issues the write operation instruction changes, the storage system generates a snapshot of the LUN before executing the current write operation instruction. When the host that issues the write operation instruction changes again, the storage system triggers an alarm and restores the data of the LUN based on the first snapshot. It is not only beneficial for the storage system to use the first snapshot for data recovery, but also beneficial for ensuring the complete recovery of the LUN data and improving the efficiency of data recovery, and it is also beneficial for avoiding the triggering of an alarm during the normal switching of the primary and standby systems and affecting the business operation.
[0122] Corresponding to the scheme given in the above method embodiment, the embodiment of the present application also provides a corresponding storage system, which includes a module or unit for executing each part of the above embodiment. The module or unit can be software, hardware, or a combination of software and hardware. The following is only a brief description of the storage system. For the implementation details of the scheme, reference can be made to the description of the above method embodiment, which will not be repeated below.
[0123] like Figure 6 FIG. 6 is a schematic diagram of a storage system 60 provided in an embodiment of the present application. Figure 2 or Figure 4 The storage system in the corresponding method embodiment can be based on the Figure 6 The structure of the storage system 60 shown in FIG. Figure 6 As shown, the storage system 60 may include a processor 601. In addition, the storage system 60 may also include a memory 603 and a communication interface 602. The processor 601 is coupled to the memory 603, and the processor 601 is coupled to the communication interface 602.
[0124] Among them, the aforementioned communication interface 602 is connected to other devices through a communication link. For example, the communication interface 602 may include an interface between the storage system 60 and a host. The storage system 60 establishes connections with at least two hosts in the host group through the communication interface 602. As another example, the communication interface 602 may include an interface between the storage system 60 and a switch. The storage system 60 is connected to the switch through the communication interface 602, and the switch establishes connections with at least two hosts in the host group. For example, the communication interface 602 may be an interface between the storage system 60 and an optical fiber switch or a network switch, which is not limited in this application.
[0125] Among them, the aforementioned processor 601 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The processor 601 may refer to a single processor or may include multiple processors, which is not specifically limited herein.
[0126] In addition, the aforementioned memory 603 is mainly used to store software programs and data. The memory 603 can exist independently and be connected to the processor 601. Optionally, the memory 603 can be integrated with the processor 601, for example, integrated within one or more chips. Among them, the memory 603 can store the program code for implementing the technical solution of the embodiments of the present application, and is controlled by the processor 601 for execution. Various computer program codes to be executed can also be regarded as the driver programs of the processor 601. The memory 603 can include volatile memory, such as random-access memory (RAM); the memory can also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 603 can also include a combination of the above types of memories. The memory 603 can refer to a single memory or can include multiple memories. Exemplarily, the memory 603 is used to store various data. For example, the aforementioned first snapshot, first identification information, etc.
[0127] It should be understood that the storage system 60 can be a centralized storage system or a distributed storage system. If the storage system 60 is a centralized storage system, the storage system 60 can be a disk-control separation architecture or a disk-control integrated architecture, which is not limited in the present application. If the storage system 60 is a centralized storage system, the storage system 60 further includes at least one engine. One engine includes at least one controller. The controller includes a processor, a memory, a front-end interface, and a back-end interface. The aforementioned processor 601 can be the processor in the engine, which is used to process data access requests from outside the storage system (servers or other storage systems), and is also used to process requests generated inside the storage system. In addition, the aforementioned communication interface 602 includes a front-end interface and a back-end interface. Among them, the front-end interface is used to communicate with the host or application server, so as to provide storage services for the host or application server. For example, the front-end interface is used to receive write operation instructions sent by the host or application server. The front-end interface is connected to the memory, and the front-end interface can temporarily store the data in the write operation instructions in the memory. The back-end interface is used to communicate with the storage medium of the memory 603 (for example, a hard disk) to expand the capacity of the storage system. For example, when the total amount of data in the memory reaches a certain threshold, the processor sends the data stored in the memory to the storage medium of the memory 603 (for example, a hard disk) through the back-end interface for persistent storage.
[0128] The processor 601 invokes the program in the memory 603 to enable the storage system 60 to implement the following functions:
[0129] In one design, the storage system 60 is used to execute the method of the storage system in the foregoing Figure 2 or Figure 4 corresponding embodiments. Wherein, the communication interface 602 is used to sequentially receive the second write operation instruction from the second host and the first write operation instruction from the first host; the processor 601 is used to generate the first snapshot of the LUN when it is determined that the first write operation instruction and the second write operation instruction come from different hosts in the host group respectively. Wherein, the first snapshot is used to record the data copy of the LUN before the execution of the first write operation instruction, and the second write operation instruction is the previous write operation instruction received by the storage system before receiving the first write operation instruction.
[0130] In a possible implementation manner, when the write operation instruction from the second host is not received again through the communication interface 602 within the set time, the processor 601 deletes the first snapshot. Wherein, the second host is a host different from the first host in the host group.
[0131] In a possible implementation manner, when the write operation instruction from the second host is received through the communication interface 602 within the set time, the processor 601 generates an alarm message and sends the alarm message through the communication interface 602. The alarm message is used to indicate that there is an abnormal write operation. Optionally, the alarm message includes the information of the LUN. Optionally, the alarm message further includes: the information of the first host and the information of the second host. Optionally, the alarm message further includes the information of the first snapshot, and the first snapshot is used to restore the data of the LUN before the execution of the first write operation instruction.
[0132] In a possible implementation manner, the communication interface 602 receives the first indication information, and the first indication information is used to indicate that the storage system restores the data of the LUN before the execution of the first write operation instruction based on the first snapshot; the processor 601 restores the data of the LUN before the execution of the first write operation instruction based on the first snapshot.
[0133] In a possible implementation manner, after the processor 601 restores the data of the LUN before the execution of the first write operation instruction based on the first snapshot, the processor 601 deletes the first snapshot.
[0134] In a possible implementation manner, the communication interface 602 is further used to receive the first configuration information, and the first configuration information is used to indicate that the master-slave consistency protection function is enabled for the host group associated with the LUN.
[0135] It should be noted that the specific implementation manners and beneficial effects of this embodiment can refer to the method of the storage system in the foregoing embodiments, and will not be elaborated here.
[0136] As shown Figure 7 in the figure, the present application also provides a storage device 70. The storage device 70 may be a storage node constituting a centralized storage system, or a storage node constituting a distributed storage system, or a component (such as an integrated circuit, a chip, etc.) in a storage node. The storage device 70 may also be other devices or modules for implementing the methods in the method embodiments of the present application.
[0137] The storage device 70 may include a processing module 701 (or referred to as a processing unit). Optionally, it may further include an interface module 702 (or referred to as a transceiver unit or transceiver module) and a storage module 703 (or referred to as a storage unit). The interface module 702 is used to communicate with other devices. The interface module 702 may be, for example, a transceiver module or an input / output module.
[0138] In a possible design, as Figure 7 described in one or more of the modules may be implemented by one or more processors, or by one or more processors and a memory; or by one or more processors and a transceiver; or by one or more processors, a memory, and a transceiver. The present application embodiments do not limit this. The processor, memory, and transceiver may be provided separately or integrated together.
[0139] The storage device 70 has the function of implementing the storage system described in the embodiments of the present application. For example, the storage device 70 includes modules or units or means corresponding to the steps involved in the storage system described in the embodiments of the present application. The function or unit or means may be implemented by software, or by hardware, or by hardware executing corresponding software, or by a combination of software and hardware. For specific details, please refer to the corresponding descriptions in the foregoing Figure 2 or Figure 4 corresponding method embodiments, which will not be elaborated here.
[0140] In addition, the present application provides a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are fully or partially generated. For example, it implements the methods related to the storage system as described in the foregoing Figure 2 or Figure 4 described. For another example, it implements the methods related to the storage system as described in the foregoing Figure 2 or Figure 4Methods related to the storage system therein. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0141] In addition, the present application also provides a computer-readable storage medium storing a computer program, which is executed by a processor to implement the storage system-related method as described above Figure 2 or Figure 4 the method related to the storage system therein.
[0142] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0143] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
Claims
1. A method for generating a snapshot, which is applied to a storage system, where a logical unit number (LUN) of the storage system is mapped to at least two hosts, and the at least two hosts include a first host and a second host, and is characterized in that, Comprising: Successively receiving a second write operation command from the second host and a first write operation instruction from the first host; When the first write operation instruction and the second write operation instruction are from different hosts respectively, generating a first snapshot of the LUN, where the first snapshot is used to record a data copy of the LUN before executing the first write operation instruction.
2. The method according to claim 1, wherein When the first write operation instruction and the second write operation instruction are from different hosts respectively, generating a first snapshot of the LUN, including: When the first write operation instruction and the second write operation instruction are from different hosts respectively and no snapshot of the LUN has been generated, generating a first snapshot of the LUN.
3. The method according to claim 1 or 2, characterized in that, The first snapshot is used to record the latest data copy of the LUN before executing the first write operation instruction.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: After generating the first snapshot, if no write operation instruction from the second host is received again within a set time, deleting the first snapshot.
5. The method according to any one of claims 1 to 4, characterized in that The method further includes: After generating the first snapshot, if a write operation instruction from the second host is received within a set time, sending an alarm message.
6. The method according to claim 5, characterized in that, The alarm message includes information of the LUN.
7. The method according to claim 6, characterized in that, The alarm message further includes: information of the first host and information of the second host.
8. The method according to claim 6 or 7, characterized in that The alarm message further includes identification information of the first snapshot.
9. The method according to any one of claims 1 to 8, characterized in that, Before receiving the first write operation instruction from the first host, the method further includes: Receiving first configuration information, where the first configuration information is used to indicate that the at least two hosts associated with the LUN enable the primary-backup consistency protection function.
10. A storage system, characterized in that, Comprising a processor and a memory, where the memory is used to store data and a program that can run on the processor, and the processor calls the program in the memory to execute the method according to any one of claims 1 to 9.
11. A storage device, characterized in that, Comprising a module for executing the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, Stored with instructions, when the instructions run on a computer, enabling the computer to execute the method according to any one of claims 1 to 9.