Data recovery method and apparatus, and computing device

By using the data recovery method with semantic-level data as granularity in the storage area network (SAN), the data of the target object is restored using the snapshot mapping relationship set, the problems of slow recovery speed and low efficiency in the prior art are solved, and fast and efficient data recovery is achieved.

WO2025161401A1PCT designated stage Publication Date: 2025-08-07HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/117834
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2024-09-09
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

The prior art data recovery speed is slow and inefficient when data stored through a storage area network (SAN) is attacked by ransomware, because recovery methods with logical unit number (LUN) or storage volume as granularity lead to a large recovery granularity.

Method used

A data recovery method is provided, by determining the semantic level data of the target object to be restored, determining the storage address of the data block using the mapping relationship set in the first snapshot, and restoring the data of the target object, realizing data recovery with semantic level data as granularity.

Benefits of technology

Data recovery with semantic-level data as granularity is realized, with fast recovery speed and high efficiency, avoiding the problems of slow speed and low efficiency when recovering with LUN or storage volume as granularity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024117834_07082025_PF_FP_ABST
    Figure CN2024117834_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data security, and discloses a data recovery method and apparatus, and a computing device. The method is applied to a storage system that employs a block-level storage technique for data storage. The method comprises: determining a target object for which data must be recovered; on the basis of a first mapping set, determining, in a first snapshot, a first storage address at which a data block of the target object is stored; and, on the basis of data stored at the first storage address in the first snapshot, recovering data of the target object. The target object is any kind of semantic-level data. The storage system stores data blocks of the target object which are stored in a block-level manner. The first snapshot is a snapshot obtained after a snapshot operation has been performed on data of the storage system at a first time point. The first mapping set comprises correspondences between objects maintained by the storage system and storage addresses of data blocks of the objects as of the first time point. The method enables data recovery at the granularity of semantic-level data.
Need to check novelty before this filing date? Find Prior Art

Description

Data recovery method, device and computing equipment

[0001] This application claims priority to Chinese patent application number 202410161112.9, filed on February 2, 2024, and entitled “Data recovery method, device and computing equipment”, and claims priority to Chinese patent application number 202410911375.7, filed on July 8, 2024, and entitled “Data recovery method, device and computing equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of data security technology, and in particular to a data recovery method, apparatus, and computing device. Background Art

[0003] A storage area network (SAN) is a dedicated high-speed network that provides access to block-level storage. It typically consists of interconnected client devices, switching equipment, and storage devices. SAN is one of the most widely used mainstream storage products today, providing block storage services for critical application data in enterprises, finance, data centers, and other sectors.

[0004] When data stored on a SAN is attacked by malicious software (e.g., infected with a ransomware virus) and needs to be recovered, the related technology generally restores the data by taking a snapshot of the storage area corresponding to the logical unit number (LUN) or storage volume. This method has a large granularity when restoring data, resulting in slow data recovery.

[0005] Summary of the Invention

[0006] The present application provides a data recovery method, apparatus, and computing device, which can realize data recovery with semantic-level data as the granularity.

[0007] The technical solutions provided in this application are as follows:

[0008] In a first aspect, the present application provides a data recovery method, which is applied to a storage system that uses block-level storage for data storage, and the method includes: determining a target object for which data needs to be recovered; determining a first storage address of a data block storing the target object in the first snapshot based on a first mapping relationship set associated with a first snapshot; and recovering the data of the target object based on the data stored at the first storage address in the first snapshot. The target object is any kind of semantic-level data, and the data block of the target object is stored in the storage system in a block-level manner. The first snapshot is a snapshot obtained after a snapshot operation is performed on the data of the storage system at a first time, and the first mapping relationship set associated with the first snapshot includes the correspondence between the storage address of the object maintained by the storage system and the data block of the object as of the first time.

[0009] For storage systems that use block-level storage for data storage, the method provided in this application can realize data recovery at the granularity of semantic-level data. Semantic-level data refers to data with semantics that can be understood / recognized by users. As an example, semantic-level data includes but is not limited to files (such as document files, worksheet files, etc.), databases, data tables, etc. Compared with recovering data at the granularity of logical unit number (LUN) or storage volume, the granularity when recovering data using the method provided in this application is small, so the data recovery speed is fast and the efficiency is high.

[0010] In one possible design, the method further includes: receiving an input / output (IO) request, the IO request being used to instruct a write operation to be performed on pending data, the IO request including an identifier of an object to which the pending data belongs, where the object to which the pending data belongs is any type of semantic-level data; and updating the first mapping relationship set based on the identifier of the object to which the pending data belongs and a storage address of the pending data. This possible design enables maintaining and updating the first mapping relationship set based on received IO requests.

[0011] In another possible design, determining the target object for data recovery includes: receiving an IO request carrying a first identifier and first data, where the first data is data of a first object represented by the first identifier; detecting whether the first data is infected with a virus or has been subjected to a malicious attack; and if it is determined that the first data is infected with a virus or has been subjected to a malicious attack, determining the first object as the target object. This possible design achieves the purpose of determining the target object for data recovery based on the data carried in the received IO request.

[0012] In another possible design, determining the target object for data recovery includes: detecting data blocks stored in the storage system to determine target data blocks in the storage system that are infected with a virus or have been maliciously attacked; querying a first mapping relationship set based on the storage address of the target data block to determine a first object corresponding to the storage address of the target data block; and determining the first object as the target object. This possible design allows the storage system to determine the target object for data recovery by detecting data blocks stored in the storage system in the event that a missed detection (virus missed detection or attack missed detection) occurs when determining the object for data recovery based on data carried in a received IO request, or in the event that the storage system is at risk of being infected with a virus or being attacked.

[0013] In another possible design, before determining the first object as the target object, the method further includes: sending a first recovery request to the user-end device, the first recovery request including an identifier of the first object, the first recovery request being used to confirm whether to recover the data of the first object; and receiving a confirmation message returned by the user-end device, the confirmation message being used to indicate that the data of the first object is to be recovered. Through this possible design, the user-end device participates in determining the target object for which data recovery is required, thereby increasing user participation in the implementation of the present application solution. Furthermore, for data that the user considers unimportant, the user-end device can indicate that it does not need to be recovered, thereby saving computing resources of the storage system while ensuring user satisfaction.

[0014] In another possible design, sending the first restore request to the user-end device includes sending the first restore request to the user-end device via a preset protocol interface. Receiving a confirmation message returned by the user-end device includes receiving the confirmation message returned by the user-end device via the preset protocol interface. This possible design enables transmission of restore requests and confirmation messages between the storage system and the user-end device without affecting normal data transmission via an existing standard protocol interface.

[0015] In another possible design, determining the target object for which data recovery is required includes: upon receiving a second recovery request sent by the user-end device, determining the first object as the target object. The second recovery request includes an identifier of the first object, and the second recovery request is used to request recovery of the data of the first object. This possible design enables the user-end device to specify the target object for which data recovery is required. Thus, when the user-end device determines that data recovery is required for a particular object, the method provided in this application can be used to perform data recovery. This possible design can thus improve the user experience.

[0016] In another possible design, restoring the target object's data based on the data stored at the first storage address in the first snapshot includes: detecting whether the data stored at the first storage address in the first snapshot is infected by a virus or subjected to a malicious attack; and if it is determined that the data stored at the first storage address in the first snapshot is not infected by a virus or subjected to a malicious attack, determining the data stored at the first storage address in the first snapshot as the restored data of the target object. This possible design ensures that the data used to restore the target object's data is secure.

[0017] In another possible design, the method further includes: updating a first mapping relationship set currently maintained by the storage system according to the identifier of the target object and the first storage address.

[0018] In another possible design, the method further includes: upon determining that the data stored at the first storage address in the first snapshot is infected by a virus or maliciously attacked, determining, based on the first mapping relationship set associated with the second snapshot, a second storage address storing the data block of the target object in the second snapshot; and restoring the data of the target object based on the data stored at the second storage address in the second snapshot. The second snapshot is a snapshot obtained by performing a snapshot operation on the storage system at a second time prior to the first time, and the first mapping relationship set associated with the second snapshot includes a correspondence between the storage addresses of the objects maintained by the storage system and the data blocks of the objects as of the second time. This possible design allows for further recovery of the target object data using the previous snapshot when the data used to restore the target object data is unsafe.

[0019] In another possible design, before determining the first storage address of the data block storing the target object in the first snapshot based on the first mapping relationship set associated with the first snapshot, the method further includes: performing at least one snapshot operation on the data of the storage system to obtain at least one snapshot, wherein the first snapshot is one of the at least one snapshots.

[0020] In another possible design, performing at least one snapshot operation on the data of the storage system to obtain at least one snapshot includes: periodically performing a snapshot operation on the data of the storage system to obtain at least one snapshot.

[0021] In another possible design, the first snapshot is a snapshot obtained after a snapshot operation is performed on the data in the storage system before the current moment and most recently from the current moment.

[0022] In another possible design, the target object is a file or a database / table.

[0023] In another possible design, the above-mentioned storage address is represented by a logical block address (LBA) and an identifier (ID) of the LUN or storage volume to which the LBA belongs, or the storage address is represented by a physical block address (PBA) and the ID of the LUN or storage volume to which the PBA belongs.

[0024] In another possible design, the snapshot is obtained by performing a snapshot operation on the data of the storage system at a granularity of LUN or storage volume.

[0025] In another possible design, the storage system is a data storage system based on a storage area network (SAN), or a direct-attached storage (DAS) system.

[0026] In a second aspect, the present application provides a data recovery method, which is applied to a user-end device accessing a storage system, and the storage system uses block-level storage for data storage. The method includes: generating an IO request for indicating a write operation on the data of a first object, the IO request includes an identifier of the first object, and the data of the first object is any kind of semantic-level data; sending the IO request to the storage system, and the identifier of the first object carried in the IO request is used to update a first mapping relationship set. The first mapping relationship set includes a correspondence between an object and the storage address of the object's data, and the first mapping relationship set is used to determine the storage address of the restored data of the first object in a snapshot when restoring the data of the first object, wherein the snapshot refers to a snapshot obtained after performing a snapshot operation on the data of the storage system. By carrying the identifier of the object in the IO request, it is possible to update the first mapping relationship set maintained by the storage system and used for data recovery at the granularity of semantic-level data.

[0027] In one possible design, the above method also includes: receiving a first recovery request sent by the storage system, the first recovery request includes an identifier of the first object, and the first recovery request is used to confirm whether to restore the data of the first object; when it is determined that the data of the first object is to be restored, returning a confirmation message to the storage system, and the confirmation message is used to indicate that the data of the first object is to be restored.

[0028] In another possible design, the receiving of the first recovery request sent by the storage system includes: receiving the first recovery request sent by the storage system through a preset protocol interface. The returning of the confirmation message to the storage system includes: sending the confirmation message to the storage system through the preset protocol interface.

[0029] In another possible design, when determining to restore the data of the first object, returning a confirmation message to the storage system includes: querying a third mapping relationship set based on the identifier of the first object carried in the first recovery request to determine the first object, wherein the third mapping relationship set is used to record the identifier of the object and the correspondence between the objects; outputting an alarm message, wherein the alarm message is used to indicate that the data of the first object is infected with a virus or is subjected to a malicious attack; and returning a confirmation message to the storage system in response to a first recovery indication input by the user based on the alarm information, wherein the first recovery indication is used to indicate the restoration of the data of the first object.

[0030] In another possible design, before querying the third mapping relationship set based on the identifier of the first object carried in the first recovery request, the method further includes: determining the identifier of the first object; and updating the third mapping relationship set based on the identifier of the first object.

[0031] In another possible design, the method further includes: sending a second recovery request to the storage system, where the second recovery request includes an identifier of the first object, and the second recovery request is used to request recovery of data of the first object.

[0032] In another possible design, the storage address is represented by an LBA and an ID of a LUN or storage volume to which the LBA belongs, or the storage address is represented by a PBA and an ID of a LUN or storage volume to which the PBA belongs.

[0033] In another possible design, the first object is a file or a database / table.

[0034] In another possible design, the snapshot is obtained by performing a snapshot operation on the data of the storage system at a granularity of LUN or storage volume.

[0035] In another possible design, the storage system is a SAN-based data storage system, or the storage system is a DAS storage system.

[0036] It can be understood that the description of the beneficial effects of the solution provided by the second aspect and any possible design method in the second aspect can refer to the description of the beneficial effects of the solution provided by the first aspect and any possible design method in the first aspect, and no further details will be given.

[0037] In a third aspect, the present application provides a data recovery device. The data recovery device is used to execute any one of the methods provided in the first aspect above. The present application can divide the data recovery device into functional modules according to any one of the methods provided in the first aspect above. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. Exemplarily, the present application can divide the data recovery device into a determination unit and a recovery unit, etc. according to the function. The description of the possible technical solutions and beneficial effects executed by each of the divided functional modules can refer to the solutions provided by the first aspect above and any possible design method in the first aspect, and will not be repeated here.

[0038] In a fourth aspect, the present application provides a data recovery device. The data recovery device is used to execute any of the methods provided in the second aspect above. The present application can divide the data recovery device into functional modules according to any of the methods provided in the second aspect above. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. Exemplarily, the present application can divide the data recovery device into a generation unit and a sending unit, etc. according to the function. The description of the possible technical solutions and beneficial effects executed by each of the divided functional modules can refer to the solutions provided in the second aspect above and any possible design method in the second aspect, and will not be repeated here.

[0039] In a fifth aspect, the present application provides a data recovery device. The data recovery device is used to execute any one of the methods provided in the first or second aspects above. The present application can divide the data recovery device into functional modules according to any one of the methods provided in the first or second aspects above. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. Exemplarily, the present application can divide the data recovery device into a transceiver unit and a processing unit, etc. according to the function. The description of the possible technical solutions and beneficial effects executed by the above-mentioned divided functional modules can refer to the solutions provided by any possible design method in the first or second aspects above, and will not be repeated here.

[0040] In a sixth aspect, the present application provides a data recovery system comprising a user-end device and a storage system. The storage system is configured to execute the method provided in the first aspect and any possible design of the first aspect, and the user-end device is configured to execute the method provided in the second aspect and any possible design of the second aspect.

[0041] In a seventh aspect, the present application provides a computing device. The computing device includes: a memory, a network interface, and one or more processors. The one or more processors receive or send data via the network interface, and the one or more processors are configured to read program instructions stored in the memory to execute the method provided in the first aspect and any possible design of the first aspect, or to execute the method provided in the second aspect and any possible design of the second aspect.

[0042] In an eighth aspect, the present application provides a data recovery method, which includes: a user-end device determines a target object for which data needs to be recovered, and queries a second mapping relationship set based on an object flag of the target object to determine the storage address of a data block of the target object, i.e., a target storage address, and sends a data recovery instruction including the target storage address to a storage system, wherein the data recovery instruction is used to instruct the recovery of the data of the target object based on the target storage address. Wherein, the target object is any semantic-level data, the data block of the target object is stored in a block-level manner in the storage system, and the storage system stores data in a block-level storage manner. The second mapping relationship set includes the correspondence between the object maintained by the user-end device and the storage address of the data block of the object in the storage system.

[0043] In a storage system that stores data in a block-level storage manner, by executing the method provided by the present application on a user-end device of the storage system, it is possible to instruct the storage system to recover data in the granularity of semantic-level data (i.e., object). Compared with recovering data in the granularity of LUN or volume, the granularity when recovering data using the method provided by the present application is small, so the data recovery speed is fast and the efficiency is high. Moreover, for SAN storage systems that store data in the form of data blocks, current academic research can achieve ransomware detection and recovery at the data block level. However, since the SAN storage system does not have a file system structure, and what is presented on the user side of the SAN storage system is not a data block, but a file with semantic information, and one file usually corresponds to multiple data blocks. Therefore, data recovery based on data block granularity may cause the file information displayed on the user side to be lost. For example, when the file being ransomed on the user side corresponds to 10 data blocks in the storage system, if the storage system detects and recovers 9 data blocks and misses 1 data block, the ransomed file may still be unavailable to the user side. The method provided in this application recovers data at the granularity of semantic-level data (i.e., objects), which can solve the problem of loss of user-side displayed file information that may occur when data recovery is performed at the granularity of data blocks.

[0044] In one possible design, the method further includes: the storage system receiving a data recovery instruction including a target storage address from the user-end device, and recovering the data of the target object based on the data stored at the target storage address in the first snapshot. The first snapshot is a snapshot obtained after performing a snapshot operation on the data in the storage system at the first time.

[0045] Through this possible design, the storage system responds to the data recovery instruction of the user-end device and can implement data recovery with semantic-level data (ie, object) as the granularity.

[0046] In another possible design, the method further includes: the storage system detecting whether the data stored in the storage system is infected by a virus or maliciously attacked, and sending the storage address of an abnormal data block to the user-end device. The abnormal data block is a data block containing data infected by a virus or maliciously attacked. The user-end device determines a target object for which data recovery is required, including: the user-end device receiving the storage address of the abnormal data block sent by the storage system, querying a second mapping relationship set based on the storage address of the abnormal data block to determine the abnormal object, and determining the target object within the abnormal object.

[0047] Through this possible design, it is possible to achieve data recovery for objects to which data in the storage system that has been infected by a virus or has been maliciously attacked belongs. For example, it is possible to achieve data recovery for objects to which all data in the storage system that has been infected by a virus or has been maliciously attacked belongs. For another example, it is possible to achieve data recovery for objects to which data with higher popularity in the data in the storage system that has been infected by a virus or has been maliciously attacked belongs. For another example, by increasing the interaction between the user and the user-end device, it is possible to achieve data recovery for objects to which data specified by the user in the data in the storage system that has been infected by a virus or has been maliciously attacked belongs. In this case, since the user participates in determining the target object for which data needs to be recovered, the user participation in the implementation of the solution can be improved. For data that the user considers unimportant in the data that has been infected by a virus or has been maliciously attacked, the user indicates that it does not need to be recovered, which can save the use of computing resources of the control module while ensuring user satisfaction.

[0048] In another possible design, the user terminal device determines the target object for which data needs to be restored, including: the user terminal device determines the target object in response to a received object restoration instruction, wherein the object restoration instruction includes an object identifier of the target object.

[0049] Through this possible design, data of an object specified by a user can be restored.

[0050] In another possible design, the storage system recovers the data of the target object based on the data stored at the target storage address in the first snapshot, including: the storage system detects whether the data stored at the target storage address in the first snapshot is infected by a virus or maliciously attacked, and if it is determined that the data stored at the target storage address in the first snapshot is not infected by a virus or maliciously attacked, determines the data stored at the target storage address in the first snapshot as the recovered data of the target object.

[0051] In another possible design, the method further includes: if the storage system determines that the data stored at the target storage address in the first snapshot is infected by a virus or is under a malicious attack, restoring the data of the target object based on the data stored at the target storage address in a second snapshot. The second snapshot is a snapshot obtained by performing a snapshot operation on the storage system at a second time before the first time.

[0052] Through the above two possible design methods, when the data used to restore the target object data is not secure, the previous snapshot can be further used to restore the target object data, thereby ensuring the security of the restored data.

[0053] In yet another possible design manner, the target object is a file or a database / table.

[0054] In another possible design, the storage address is represented by an LBA and an ID of a LUN or storage volume to which the LBA belongs, or the storage address is represented by a PBA and an ID of a LUN or storage volume to which the PBA belongs.

[0055] In another possible design, the snapshot is obtained by performing a snapshot operation on the data of the storage system at a granularity of LUN or storage volume.

[0056] In another possible design, the storage system is a SAN-based data storage system, or the storage system is a DAS storage system.

[0057] In a ninth aspect, the present application provides a data recovery system, which includes a user-end device and a storage system. The user-end device is used to determine the target object for which data needs to be recovered, query the second mapping relationship set according to the object flag of the target object to determine the storage address of the data block of the target object, that is, the target storage address, and send a data recovery instruction including the target storage address to the storage system, where the data recovery instruction is used to instruct the recovery of the data of the target object according to the target storage address. The target object is any semantic-level data, the data block of the target object is stored in the storage system in a block-level manner, and the storage system uses a block-level storage method to store data. The second mapping relationship set includes the correspondence between the object maintained by the user-end device and the storage address of the data block of the object in the storage system.

[0058] In one possible design, a storage system is configured to receive a data recovery instruction including a target storage address from a user terminal device, and to recover data of a target object based on data stored at the target storage address in a first snapshot. The first snapshot is a snapshot obtained after a snapshot operation is first performed on data in the storage system.

[0059] In another possible design, the storage system is further configured to detect whether data stored in the storage system is infected by a virus or maliciously attacked, and to send the storage address of an abnormal data block to a user-end device. The abnormal data block is a data block containing data infected by a virus or maliciously attacked. The user-end device is specifically configured to receive the storage address of the abnormal data block sent by the storage system, query the second mapping relationship set based on the storage address of the abnormal data block to determine an abnormal object, and determine the target object within the abnormal object.

[0060] In another possible design, the user terminal device is specifically configured to determine the target object in response to a received object recovery instruction, wherein the object recovery instruction includes an object identifier of the target object.

[0061] In another possible design, the storage system is specifically used to detect whether the data stored at the target storage address in the first snapshot is infected by a virus or maliciously attacked, and if it is determined that the data stored at the target storage address in the first snapshot is not infected by a virus or maliciously attacked, the storage system determines the data stored at the target storage address in the first snapshot as the data after the target object is recovered.

[0062] In another possible design, the storage system is further configured to restore the data of the target object based on the data stored at the target storage address in a second snapshot, if it is determined that the data stored at the target storage address in the first snapshot is infected by a virus or is under a malicious attack. The second snapshot is a snapshot obtained by performing a snapshot operation on the storage system at a second time before the first time.

[0063] In yet another possible design manner, the target object is a file or a database / table.

[0064] In another possible design, the storage address is represented by an LBA and an ID of a LUN or storage volume to which the LBA belongs, or the storage address is represented by a PBA and an ID of a LUN or storage volume to which the PBA belongs.

[0065] In another possible design, the snapshot is obtained by performing a snapshot operation on the data of the storage system at a granularity of LUN or storage volume.

[0066] In another possible design, the storage system is a SAN-based data storage system, or the storage system is a DAS storage system.

[0067] It can be understood that the description of the beneficial effects of the solution provided by the ninth aspect and any possible design method in the ninth aspect can refer to the description of the beneficial effects of the solution provided by the eighth aspect and any possible design method in the eighth aspect, and no further details will be given.

[0068] In the tenth aspect, the present application provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium, and the computer-readable storage medium includes computer program instructions. When the computer program instructions are executed by a computing device or a processor, the computing device or the processor executes the method provided in the first aspect and any possible design method of the first aspect, or executes the method provided in the second aspect and any possible design method of the second aspect, or executes the method provided in the eighth aspect and any possible design method of the eighth aspect.

[0069] In the eleventh aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device or a processor, causes the computing device or the processor to execute the method provided in the first aspect and any possible design method of the first aspect, or to execute the method provided in the second aspect and any possible design method of the second aspect, or to execute the method provided in the eighth aspect and any possible design method of the eighth aspect.

[0070] In the twelfth aspect, the present application provides a chip, which includes a processor for running program instructions or codes. The chip or a device including the chip can be used to execute the method provided in the first aspect and any possible design method in the first aspect, or the chip or a device including the chip is used to execute the method provided in the second aspect and any possible design method in the second aspect, or the chip or a device including the chip is used to execute the method provided in the eighth aspect and any possible design method in the eighth aspect. Exemplarily, the chip also includes: an input interface, an output interface, and a memory. The input interface, output interface, processor, and memory of the chip are connected through the internal connection path of the chip, the memory in the chip is used to store the program instructions or codes run by the processor, and the input interface and output interface of the chip are used for connection and communication between the chip and other chips or devices.

[0071] It can be understood that any of the data recovery devices, data recovery systems, computing devices, computer-readable storage media, computer program products or chips provided above can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods and will not be repeated here.

[0072] In this application, the names of the above-mentioned data recovery device, data recovery system, etc. do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear with other names. As long as the functions of each device or functional module are similar to those of this application, they are all within the scope of protection of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] FIG1 is a schematic diagram of an implementation environment of the method provided in an embodiment of the present application;

[0074] FIG2 is a schematic diagram of another implementation environment of the method provided in the embodiment of the present application;

[0075] FIG3 is a schematic diagram of a method for maintaining a first mapping relationship set provided in an embodiment of the present application;

[0076] FIG4 is a schematic diagram of updating a first mapping relationship set provided by an embodiment of the present application;

[0077] FIG5 is a flow chart of a data recovery method provided in an embodiment of the present application;

[0078] FIG6 is a flow chart of another data recovery method provided in an embodiment of the present application;

[0079] FIG7 is another flow chart of a data recovery method according to an embodiment of the present application;

[0080] FIG8 is a flow chart of another data recovery method provided in an embodiment of the present application;

[0081] 9 is a schematic diagram of a process in which a user terminal device determines a target object requiring data recovery based on an abnormal data block in a storage system according to an embodiment of the present application;

[0082] FIG10 is another flow chart of a data recovery method according to an embodiment of the present application;

[0083] FIG11 is a schematic structural diagram of a data recovery device provided in an embodiment of the present application;

[0084] FIG12 is a schematic structural diagram of another data recovery device provided in an embodiment of the present application;

[0085] FIG13 is a schematic structural diagram of another data recovery device provided in an embodiment of the present application;

[0086] FIG14 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0087] FIG15 is a schematic structural diagram of a data recovery system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0088] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0089] To facilitate understanding, the technology and background involved in the embodiments of this application are explained below.

[0090] 1) Ransomware

[0091] Ransomware is a type of malware that uses various powerful encryption algorithms. It encrypts the victim's critical data, rendering it unusable, or locks the victim's computing device, rendering it unusable. Furthermore, ransomware attackers threaten to keep the victim's data encrypted or their device locked. Only after the victim pays a ransom will the attackers decrypt the data or restore the device. In recent years, ransomware, which restricts user access to files and demands a ransom, has become increasingly prevalent, becoming one of the fastest-growing cyberthreats.

[0092] Ransomware is primarily spread through emails, Trojans, and web exploits. Ransomware detection technology identifies ransomware behavior patterns as it encrypts data, allowing it to detect it promptly and prevent further data from being maliciously encrypted.

[0093] 2) Logical unit number (LUN)

[0094] In a storage area network (SAN), a LUN is a number used to identify a logical unit (LUN). This LUN is a device addressable via the Small Computer System Interface (SCSI). In other words, the storage system partitions the physical hard drive into logically addressed sections, allowing user hosts to access each section. Each such partition is called a LUN. LUNs are also commonly referred to as logical disks created within a SAN storage system.

[0095] 3) Storage Volume

[0096] The directory on the container host that is bound to the container is called a storage volume.

[0097] A storage volume is essentially a file or directory that exists directly on the container's host machine as a file or directory, bypassing the default union file system. Specifically, a storage volume establishes a binding relationship between a directory in the container's host machine's local file system and a directory in the container's internal file system. Therefore, when data is written to this directory in the container, the container writes the data directly to the directory on the host machine that has been bound to the container.

[0098] 4) Logical block address (LBA) and physical block address (PBA)

[0099] LBA is a common mechanism used on data storage devices to indicate data location. Hard drives are the most common device to use this mechanism. LBA can refer to the address of a data block or the data block pointed to by an address. PBA, on the other hand, refers to a physical address on a hard drive or the physical block pointed to by a physical address.

[0100] On a hard disk drive (HDD), since data on the HDD can be directly overwritten, the relationship between LBA and PBA is one-to-one. However, on a solid state drive (SSD), the relationship between LBA and PBA is not fixed because the storage medium used by the SSD must be erased before it can be written, and reading and writing are done in pages, while erasing is done in data blocks (composed of multiple pages). Therefore, the SSD requires a flash translation layer (FTL) to convert LBA and PBA to adapt to the existing file system. The FTL is used to complete the translation (or mapping) from the host logical address space to the flash physical address space in the flash translation layer.

[0101] 5) Snapshot

[0102] A snapshot is a backup of the data stored on a disk / volume / book.

[0103] When a snapshot operation is performed on a disk / volume / book to obtain a snapshot of the disk / volume / book, no copying is involved. Therefore, even if the amount of data is huge, the snapshot operation can be completed in a very short time. Taking the example of performing a snapshot operation on the disk at time 1 to obtain snapshot 1 of the disk data, the disk system does not need to copy the data itself on the disk, but only needs to notify the disk to retain all disk blocks storing data, and does not allow overwriting of the retained disk blocks. Therefore, when any subsequent modification, addition, or deletion operations are performed on the data on the disk, the disk blocks where the original data is located will not be overwritten, but the modified parts will be written to other free and available disk blocks on the disk. In this way, the data stored in the aforementioned retained disk blocks can be called snapshot 1 of the disk system at time 1.

[0104] Currently, for SAN storage systems, after the storage system detects ransomware, one technical solution is to recover data by rolling back the LUN or volume where the data infected with the ransomware is located. Due to the large granularity of the LUN or volume, the data recovery speed is slow and the efficiency is low. In addition, during the rollback of the LUN or volume, the normal data in the LUN or volume that is not infected by the virus will be rolled back as well, which will result in the loss of normal data. In another technical solution, by backing up the data of the SAN storage system, the storage system can recover the data based on the backup copy after detecting the ransomware. However, this method will bring additional storage costs.

[0105] Based on this, embodiments of the present application provide a data recovery method for storage systems that use block-level storage. This method enables data recovery at the semantic level. Compared to recovering data at the LUN or volume level, the method provided in embodiments of the present application achieves smaller granularity, resulting in faster and more efficient data recovery.

[0106] Semantic-level data refers to data with semantic meanings that can be understood / recognized by users. As an example, semantic-level data includes but is not limited to files (such as document files, worksheet files, etc.), databases, data tables, etc. In the embodiments of this application, a type of "semantic-level data" is referred to as an "object."

[0107] Referring to Figure 1, Figure 1 shows a schematic diagram of an implementation environment of the method provided in an embodiment of the present application. As shown in Figure 1, the storage system includes a control module and multiple storage devices. The control module communicates with the user-end device that accesses the storage system and processes input / output (IO) requests from the user-end device. As an example, the control module responds to the IO request from the user-end device and performs read and write operations on the data indicated by the IO request. For example, the control module writes the data carried by the IO request to the storage device in the storage system based on the indication of the IO request. For another example, the control module modifies the data stored in the storage device based on the indication of the IO request. For another example, the control module reads data from the storage device in the storage system based on the indication of the IO request.

[0108] When the data carried in the IO request from the user-end device or the user data stored in the above-mentioned storage device is infected with a virus (such as a ransomware virus) or tampered with by a malicious attack and needs to be recovered, or when the data stored in the storage device needs to be recovered due to user misoperation, the data recovery method provided in the embodiment of the present application can be used to achieve data recovery with semantic-level data as the granularity. Compared with recovering data with LUN or volume as the granularity, the granularity of data recovery using the method provided in the embodiment of the present application is small, so the data recovery speed is fast and the efficiency is high.

[0109] The user-end device shown in FIG1 is any computing device (or user host) capable of accessing the storage system. For example, the computing device includes, but is not limited to, a general-purpose computer, laptop, tablet computer, vehicle-mounted device, mobile phone, artificial intelligence (AI) device, or other terminal device capable of accessing the storage system. Another example includes, but is not limited to, a server, virtual machine, or other network-capable device capable of accessing the storage system.

[0110] The storage device in the storage system shown in FIG1 can be implemented as a disk array, or as a computing device including a disk or a disk array, without limitation. Exemplary disk arrays include, but are not limited to, HDDs, solid state drives (SSDs), and the like.

[0111] The control module of the storage system shown in Figure 1 can be any computing device with computing processing capabilities or a functional module within a computing device. Alternatively, the control module can be implemented by another device independent of the storage devices in the storage system, or the functions implemented by the control module can be implemented by at least one storage device in the storage system, without limitation.

[0112] In one example, when the storage system shown in FIG1 is a direct-attached storage (DAS) system, the control module in the storage system is a storage server that is directly connected to the storage devices in the storage system (e.g., directly connected via a cable). In another example, when the storage system shown in FIG1 is a SAN storage system, the control module in the storage system is a SAN storage controller, and the user-end device communicates with the SAN storage controller via the SAN network shown in FIG2 . Optionally, the SAN network is implemented as a fiber optic network, or as an Ethernet network (e.g., a Gigabit Ethernet network or a 10 Gigabit Ethernet network), without limitation.

[0113] In the embodiments of the present application, the user-end device accessing the storage system and the control module of the storage system are configured with, in addition to a standard protocol interface for transmitting I / O requests and I / O request response messages, a communication protocol interface for transmitting the identifier of the object to which the data to be recovered belongs. This communication protocol interface is hereinafter referred to as the preset protocol interface. The detailed usage of the preset protocol interface is described in the method below and is not further elaborated here.

[0114] It should be understood that the above content is an illustrative description of the implementation environment of the data recovery method provided in the embodiment of the present application, and does not constitute a limitation on the implementation environment of the data recovery method. A person skilled in the art will know that as business needs change, the implementation environment of the method can be adjusted according to application requirements, and the embodiments of the present application do not list them one by one.

[0115] The present application also provides a data recovery system comprising the aforementioned storage system and a user-end device accessing the storage system. The method provided in the present application is applied to the data recovery system, with the storage system and the user-end device executing the corresponding method steps. For detailed descriptions, please refer to the method description below and will not be repeated here.

[0116] The embodiment of the present application also provides a data recovery device, which can be applied to the control module of the storage system described above, and is used to execute the corresponding steps executed by the storage system control module in the method of the embodiment of the present application. The device can also be applied to the user-end device described above, and is used to execute the corresponding steps executed by the user-end device in the method of the embodiment of the present application. Optionally, the device can be any computing device with computing processing capabilities, or a functional module in the computing device. For example, the computing device includes but is not limited to terminal devices such as general-purpose computers, laptops, tablet computers, vehicles, mobile phones, AI devices, etc. For another example, the computing device includes but is not limited to devices with network functions such as servers.

[0117] The following describes the implementation process of the data recovery method provided in the embodiment of the present application.

[0118] First, when implementing the data recovery method provided in the embodiment of the present application, the control module of the storage system needs to record and update the correspondence between the object to which the data to be written indicated by the IO request and the storage address of the data in the storage system when receiving the IO request from the user-end device. Alternatively, when the user-end device needs to modify or delete the data stored in the storage system according to the data to be written to the storage system, or to record and update the correspondence between the object to which the data to be processed belongs and the storage address of the data to be processed in the storage system. The data to be processed may be data to be written to the storage system, or data used to modify the data stored in the storage system, or data stored in the storage system to be deleted. In the embodiment of the present application, the object to which the data belongs refers to any kind of semantic-level data, for example, the object to which the data belongs is a document file or a table file, etc., but is not limited to this.

[0119] In an embodiment of the present application, the set recorded and updated by the storage system and including the at least one "correspondence between the object to which the data indicated by the IO request for writing belongs and the storage address of the data in the storage system" is referred to as a first mapping relationship set. In this case, the process by which the storage system control module, upon receiving an IO request from a user-end device, records and updates the "correspondence between the object to which the data indicated by the IO request for writing belongs and the storage address of the data in the storage system" can also be understood as the process by which the storage system control module maintains / updates the first mapping relationship set based on the IO request from the user-end device. Furthermore, in an embodiment of the present application, the set recorded and updated by the user-end device and including the at least one "correspondence between the object to which the data to be processed belongs and the storage address of the data to be processed in the storage system" is referred to as a second mapping relationship set. In this case, the process by which the user-end device, upon needing to write data to the storage system or delete / modify data stored in the storage system, records and updates the "correspondence between the object to which the data to be processed belongs and the storage address of the data to be processed in the storage system" can also be understood as the process by which the user-end device maintains / updates the second mapping relationship set based on the data to be processed.

[0120] Optionally, the first mapping relationship set and the second mapping relationship may be implemented as a mapping relationship table, or as a relationship graph, etc., which is not limited.

[0121] Referring to Figure 3, Figure 3 shows a schematic diagram of a method for updating and maintaining a first mapping relationship set provided by an embodiment of the present application. Optionally, the method is applied to the implementation environment shown in Figure 1 or Figure 2, and the corresponding steps are performed by the control module of the user terminal device and storage system in the implementation environment shown in Figure 1 or Figure 2. As shown in Figure 3, the method includes the following steps.

[0122] Step 101: The user terminal device generates an IO request for instructing to perform a write operation on data to be processed. The IO request includes an identifier of an object to which the data to be processed belongs.

[0123] The write operation includes at least one of a data write operation, a data delete operation, or a data modification operation. Specifically, the data write operation refers to the operation of writing data to the storage system, the data delete operation refers to the operation of deleting data from the storage system, and the data modification operation refers to the operation of modifying data stored in the storage system.

[0124] The identifier of the object to which the data to be processed belongs included in the IO request is used to uniquely identify the object and to update the first mapping relationship set. The object to which the data to be processed belongs is any kind of semantic-level data. For example, the object is a document file or a table file, etc., but is not limited thereto. The first mapping relationship set includes the correspondence between the object and the storage address of the data block of the object in the storage system. The object refers to the object to which the data on which the IO request indicates to perform a write operation belongs, and the data of the object is the data on which the IO request indicates to perform a write operation. In addition, the first mapping relationship set can be used to determine the storage address of the restored data block of the object in the snapshot when it is necessary to restore the data of a certain object (such as the first storage address, the second storage address, etc. described below). The snapshot is a snapshot obtained after performing a snapshot operation on the data of the storage system.

[0125] Taking the case where the data to be processed is the first data and the object to which the first data belongs is the first object as an example, when the user-end device needs to generate an IO request to perform a write operation on the first data, it first determines the identifier of the first object. For example, the user-end device uses a preset algorithm to encode the metadata of the first object, thereby obtaining the identifier of the first object. The metadata of the first object includes but is not limited to the name, size, attributes, etc. of the first object, and the preset algorithm includes but is not limited to a hash algorithm. For another example, the user-end device uses a preset algorithm to encode the file directory of the first object in the user-end device, thereby obtaining the identifier of the first object. For another example, the identifier of the first object is represented by the metadata of the first object. This is not limited to this.

[0126] Furthermore, the user terminal device generates a corresponding IO request according to the write operation to be performed on the first data and the identifier of the first object.

[0127] In one example, when the write operation performed on the first data is a data writing operation, the user terminal device obtains the first data that needs to be written to the storage system, determines the identifier of the first object, and generates a corresponding IO write request based on the identifiers of the first data and the first object.

[0128] In some embodiments, the user terminal device records a correspondence between objects and object identifiers. In this embodiment of the present application, the set of at least one "correspondence between objects and object identifiers" recorded by the user terminal device is referred to as a third mapping relationship set. Optionally, the third mapping relationship set can be implemented as a mapping relationship table or a relationship graph, etc., without limitation. It is understood that the objects in the correspondence relationship can be represented as object names, directories, or metadata, etc., without limitation.

[0129] In this case, each time the user-end device determines the identifier of an object, it updates the third mapping relationship set according to the identifier of the object. For example, after determining the identifier of the first object, the user-end device updates the third mapping relationship set according to the identifier of the first object. Specifically, the user-end device may first traverse the third mapping relationship set according to the identifier of the first object. When the third mapping relationship set includes a corresponding relationship containing the identifier of the first object, the user-end device does not process it. When the third mapping relationship set does not include a corresponding relationship containing the identifier of the first object, the user-end device adds a corresponding relationship to the third mapping relationship set, and the newly added corresponding relationship is used to record the corresponding relationship between the identifier of the first object and the first object. The first object in the newly added corresponding relationship can be represented as the name, directory or metadata of the first object, etc., but is not limited to this.

[0130] Step 102: The user terminal device sends an IO request to the storage system.

[0131] Exemplarily, after generating an IO request, the user-end device sends the IO request to the control module of the storage system through a communication protocol interface between the user-end device and the control module of the storage system.

[0132] In one example, when the storage system is a SAN storage system, after generating an IO request, the user-end device sends the IO request to the SAN storage controller through the communication protocol interface between itself and the SAN storage controller and via the SAN network between the user-end device and the storage system.

[0133] Step 103: The control module of the storage system receives an IO request from the user-end device.

[0134] Exemplarily, the control module of the storage system receives the IO request from the user-end device through the communication protocol interface between the control module of the storage system and the user-end device.

[0135] In one example, when the storage system is a SAN storage system, the SAN storage controller receives an IO request from the user-end device transmitted via the SAN network between the user-end device and the storage system through a communication protocol interface between the SAN storage controller and the user-end device.

[0136] Step 104: The control module of the storage system updates the first mapping relationship set according to the identifier of the object to which the data to be processed belongs and the storage address of the data to be processed carried in the IO request.

[0137] Specifically, after receiving the IO request, the control module first extracts the processing operation type indicated by the IO request from the IO request, and when it is determined that the type of processing operation is a write operation, updates the first mapping relationship set based on the storage address of the data to be processed on which the write operation is performed as indicated by the IO request and the identifier of the object to which the data to be processed belongs carried by the IO request.

[0138] When the control module determines that it is necessary to update the first mapping relationship set based on the identifier of the object to which the data to be processed carried by the received IO request belongs and the storage address of the data to be processed, the control module collects the IO data of the IO request. Taking the example that the data to be processed carried by the IO request is the first data, and the object to which the data to be processed carried by the IO request belongs is the first object, the IO data of the IO request collected by the control module include: the identifier of the first object carried by the IO request, the storage address of the first data on which the IO request indicates to perform a write operation, the type of write operation indicated by the IO request (i.e., an operation to write data, an operation to modify data, or an operation to delete data), and the operation length of the write operation indicated by the IO request. Optionally, the IO data also includes the time when the IO request is received.

[0139] Among them, the operation length of the write operation performed by the IO request refers to the length of the first data that the IO request indicates to perform the write operation. The storage address of the first data that the IO request indicates to perform the write operation refers to the address used to store the first data in the storage device of the storage system, and the storage address can be represented by the PBA of the physical block storing the first data and the identifier (identifier, ID) of the LUN to which the PBA belongs or the ID of the storage volume to which it belongs. Alternatively, the storage address can be represented by the LBA of the logical block storing the first data and the ID of the LUN to which the LBA belongs or the ID of the storage volume to which it belongs, and there is no limitation on this. Generally, the PBA is represented by the first address of the physical block, and the LBA is represented by the first address of the logical block, and the sizes of the physical block and the logical block are generally preset sizes.

[0140] Because the IO request sent by the user-end device to the control module carries the identifier of the first object, the type of write operation indicated by the IO request, and the length of the write operation indicated by the IO request, and when the IO request is used to instruct the execution of an operation to delete data, the IO request sent by the user-end device to the control module optionally carries the storage address of the first data indicated by the IO request to execute the operation to delete data, the control module can extract the identifier of the first object, the type of write operation indicated by the IO request, the length of the write operation indicated by the IO request, and the storage address of the first data indicated by the IO request to execute the write operation from the received IO request.

[0141] In addition, when the write operation indicated by the IO request is an operation to write data, optionally, when the storage address (e.g., LBA) of the first data on which the IO request indicates to perform the write operation is specified by the user-end device, the IO request sent by the user-end device to the control module carries the storage address assigned by the user-end device to the first data. In response, the control module can extract the storage address of the first data from the received IO request. Optionally, when the storage address (e.g., LBA or PBA) of the first data on which the IO request indicates to perform the write operation is specified by the storage system, the IO request sent by the user-end device to the control module will not carry the storage address of the first data on which the IO request indicates to perform the write operation. In this case, after receiving the IO request, the control module assigns a storage address to the first data on which the IO request indicates to perform the write operation, thereby obtaining the storage address of the first data on which the IO request indicates to perform the write operation.

[0142] Furthermore, the control module updates the first mapping relationship set according to the IO data collected based on the IO request. Specifically, the control module updates the first mapping relationship set according to the identifier of the first object and the storage address of the first data in the collected IO data.

[0143] In one case, when the control module indicates that the type of write operation to be performed according to the IO request included in the IO data is an operation of writing data or an operation of modifying data, the control module adds a correspondence between the identifier of the first object and the storage address of the first data in the first mapping relationship set. In addition, the control module uses the operation length of the write operation indicated by the IO request included in the IO data as the attribute information of the correspondence between the identifier of the newly added first object and the storage address of the first data. Optionally, the control module uses the time of receiving the IO request included in the IO data as the attribute information of the correspondence between the identifier of the newly added first object and the storage address of the first data. It can be understood that when the storage system supports in-situ overwriting to modify data, and the length of the modified data is less than the length of the data before modification, the control module does not need to update the first mapping relationship set according to the IO request indicating the execution of the operation of modifying data.

[0144] In another case, when the type of write operation performed by the control module according to the IO request indication included in the IO data is a data deletion operation, the control module searches for the correspondence between the identifier of the first object and the storage address of the first data in the first mapping relationship set based on the identifier of the first object and the storage address of the first data included in the IO data, deletes the correspondence from the first mapping relationship set, and deletes the attribute information of the correspondence.

[0145] Taking the first mapping relationship set as a mapping relationship table, and the IO data collected by the control module according to the received IO request including: the identifier of the first object, the storage address of the first data indicated by the IO request to perform a write operation, the type of write operation indicated by the IO request, the operation length L of the write operation indicated by the IO request, and the time of receiving the IO request as an example, referring to Figure 4, Figure 4 shows a schematic diagram of updating the first mapping relationship set provided in an embodiment of the present application.

[0146] As shown in Table (1) in FIG4 , the first mapping relationship set includes the correspondence between the identifier of the first object and the storage address of the second data, and the correspondence between the identifier of the second object and the storage address of the third data. When the control module indicates that the type of write operation performed according to the IO request included in the IO data is an operation of writing data, for example, the control module determines according to the IO data that the IO request is used to indicate writing the first data of the first object to the storage system, then the control module adds the correspondence between the identifier of the first object and the storage address of the first data to the first mapping relationship set, and uses L and T as the attribute information of the newly added correspondence, thereby obtaining the first mapping relationship set shown in Table (2) in FIG4 (the attribute information of the correspondence is not shown in FIG4 ). Alternatively, as shown in Table (3) in FIG4 , the first mapping relationship set includes the correspondence between the identifier of the second object and the storage address of the third data. When the type of write operation performed by the control module according to the IO request indication included in the IO data is an operation of writing data, for example, the control module determines according to the IO data that the IO request is used to instruct writing the first data of the first object to the storage system, then the control module adds a correspondence between the identifier of the first object and the storage address of the first data to the first mapping relationship set, and uses L and T as attribute information of the newly added correspondence, thereby obtaining the first mapping relationship set shown in Table (4) in Figure 4.

[0147] As shown in Table (5) in FIG4 , the first mapping relationship set includes a correspondence between the identifier of the first object and the storage address of the first data, a correspondence between the identifier of the first object and the storage address of the second data, and a correspondence between the identifier of the second object and the storage address of the third data. When the control module indicates that the write operation type to be performed according to the IO request included in the IO data is an operation of deleting data, for example, the control module determines according to the IO data that the IO request is used to indicate the deletion of the first data of the first object from the storage system, then the control module traverses the first mapping relationship set according to the identifier of the first object and the storage address of the first data included in the data, finds the correspondence between the identifier of the first object and the storage address of the first data in the first mapping relationship set, deletes the correspondence, and deletes the attribute information of the correspondence, thereby obtaining the first mapping relationship set shown in Table (6) in FIG4 .

[0148] Thus, through steps 101 to 104, the storage system achieves the purpose of maintaining and updating the first mapping relationship set according to the received IO request. Furthermore, based on the first mapping relationship set maintained by the storage system, the data recovery method provided in the embodiment of the present application can be implemented.

[0149] The following describes the process of implementing the data recovery method provided in the embodiment of the present application based on the first mapping relationship set maintained by the control module.

[0150] Referring to Figure 5 , a schematic flow chart illustrating a data recovery method provided in an embodiment of the present application is shown. Optionally, the method is applied to the implementation environment shown in Figure 1 or Figure 2 , with the corresponding steps being performed by the user-end device and the control module of the storage system in the implementation environment shown in Figure 1 or Figure 2 . As shown in Figure 5 , the method includes the following steps.

[0151] Step 201: The control module of the storage system determines a target object whose data needs to be restored.

[0152] The data of the target object is any kind of semantic-level data, and the storage system stores data blocks of the target object in a block-level manner.

[0153] The process of the control module determining the target object for which data needs to be restored can be implemented in the following ways.

[0154] Method 1: The control module determines the target object whose data needs to be restored according to the IO request from the user terminal device.

[0155] After the control module receives any IO request from a user-end device, if the control module determines, based on the IO request, that the type of write operation to be performed on the data indicated by the IO request is a data write operation or a data modification operation, the control module performs a security check on the data to be written carried by the IO request, and determines, based on the check result, whether the object to which the data to be written belongs is identified as a target object requiring data recovery. In this embodiment of the present application, because the IO request received by the control module of the storage system carries the object identifier of the object to which the data to be written belongs, the object to which the data to be written belongs is the object represented by the object identifier carried by the IO request.

[0156] The following description uses as an example an IO request received by a control module instructing to write first data (i.e., data to be written), where the first data is data of a first object represented by a first identifier. In this case, the IO request carries the first identifier and the first data, and instructs the IO request to write the first data carried by it to the storage system.

[0157] First, the control module performs a security check on the first data carried by the received IO request, for example, the control module detects whether the first data is infected with a virus (such as a ransomware virus) or is subject to a malicious attack. Furthermore, after the control module performs a security check on the first data carried by the IO request, the control module determines whether the first data is infected with a virus or is subject to a malicious attack based on the obtained detection result. However, the embodiment of the present application does not describe in detail the process by which the control module detects whether the data to be written carried by the IO request is infected with a virus or is subject to a malicious attack.

[0158] It is understandable that the embodiments of the present application do not limit the order of writing the first data to the storage device of the storage system and performing a security check on the first data. For example, the control module performs a security check on the first data while writing the first data to the storage device of the storage system. For another example, the control module first performs a security check on the first data, and writes the first data to the storage device of the storage system when it is determined based on the test result of the first data that the first data is not infected with a virus or is not subject to a malicious attack, and discards the first data when it is determined based on the test result of the first data that the first data is infected with a virus or is subject to a malicious attack. At this time, the control module does not need to update the first mapping relationship set based on the IO request carrying the first data. For another example, the control module first writes the first data to the storage device of the storage system, and then performs a security check on the first data.

[0159] In a first possible implementation, when the control module determines, based on the detection result, that the first data carried in the IO request is infected with a virus or is under a malicious attack, the control module determines the first object represented by the first identifier carried in the IO request as the target object.

[0160] In a second possible implementation, when the control module determines based on the detection result that the first data carried by the IO request is infected with a virus or is under a malicious attack, the control module sends a first recovery request to the user-end device that sent the IO request. The first recovery request includes the object identifier carried by the IO request (i.e., the identifier of the first object). The first recovery request is used to determine whether the data of the first object needs to be recovered. Exemplarily, the control module sends the first recovery request to the user-end device through a preset protocol interface configured by itself. In the case where the first data carried by the IO request is infected with a virus or is under a malicious attack, the object identifier carried by the IO request is also called an abnormal object identifier.

[0161] In response, the user-end device receives the first restoration request. For example, the user-end device receives the first restoration request sent by the control module via a preset protocol interface configured within the user-end device. Next, the user-end device queries the third mapping relationship set based on the identifier of the first object carried in the first restoration request to determine the first object. The third mapping relationship set is used to record the correspondence between objects and their identifiers. For detailed information on configuring the third mapping relationship set on the user-end device, please refer to the relevant description of step 101.

[0162] After the user-end device determines that the object represented by the identifier carried in the received first recovery request is the first object by querying the third mapping relationship set, it outputs an alarm message, which is used to indicate that the data of the first object is infected with a virus or is under malicious attack, and is used to determine whether the data of the first object needs to be restored. Exemplarily, the user-end device can output the alarm message through an output interface (such as a display screen, a monitor, a speaker, or other output interface). Based on the alarm message output by the user-end device, the user can choose whether to restore the data of the first object based on his or her own needs.

[0163] Furthermore, the user-end device can receive a first recovery indication input by the user in response to the alarm information and input to the user-end device through the input interface of the user-end device (such as a mouse, keyboard, touch screen, microphone, etc.). When the first recovery indication received by the user-end device is used to indicate that there is no need to restore the data of the first object, the user-end device ends the data recovery process in response. When the first recovery indication received by the user-end device is used to indicate the restoration of the data of the first object, the user-end device returns a confirmation message to the control module in response to the first recovery indication input by the user based on the alarm information. For example, the user-end device returns a confirmation message to the control module through a preset protocol interface configured by itself. The confirmation message is used to indicate the restoration of the data of the first object represented by the object identifier carried by the first recovery request. Exemplarily, the above-mentioned confirmation message may be a response message to the first recovery request. Optionally, the confirmation message carries the identifier of the first object.

[0164] In response, the control module receives a confirmation message returned by the user terminal device. For example, the control module receives the confirmation message returned by the user terminal device through a preset protocol interface configured by the control module. Further, in response to the confirmation message, the control module determines the first object as the target object.

[0165] It should be understood that in the second possible implementation, the user participates in determining the target objects for which data needs to be restored, thereby increasing user participation in the solution implementation process. For data deemed unimportant by the user, the user indicates that it does not need to be restored. This can save computing resources on the control module while ensuring user satisfaction.

[0166] Method 2: The control module detects the data blocks in the storage system and determines the target object for which data needs to be restored based on the detection results.

[0167] Exemplarily, a data block may be a data block (block) represented by the smallest storage unit when a storage system stores data, or a physical block composed of at least one block, or a logical block corresponding to a physical block, without limitation.

[0168] Optionally, the control module may periodically perform security checks on data blocks in the storage system, and determine target objects for which data needs to be restored based on the detection results of each period.

[0169] Optionally, the control module may periodically perform security checks on all or some of the data blocks in the storage system and determine the target objects for data recovery based on the test results of each period. The data stored in some of the data blocks may be hot data stored in the storage system, which is not limited to this.

[0170] After the control module detects data blocks stored in the storage system and determines, based on the detection results, a target data block in the storage system that is infected with a virus or under a malicious attack, the control module queries the most recently updated first mapping relationship set based on the storage address of the target data block, thereby determining the object identifier corresponding to the storage address of the target data block. A detailed description of the first mapping relationship set can be found in the description of the method shown in FIG. 3 and is not further elaborated here. Furthermore, the present embodiment does not describe in detail the process by which the control module detects whether a data block in the storage system is infected with a virus or under a malicious attack.

[0171] Taking the example of a case where the control module determines that the object identifier corresponding to the storage address of the target data block is the identifier of the first object, in one possible implementation, the control module determines the first object represented by the identifier as the target object. In another possible implementation, the control module sends a first recovery request carrying the identifier of the first object to the user-end device, and upon receiving a confirmation message returned by the user-end device indicating recovery of the data of the first object, determines the first object as the target object. The detailed process can be referred to the relevant description in Method 1 and is not repeated here.

[0172] Mode 3: The control module determines the target object whose data needs to be restored according to the second restoration request sent by the user terminal device.

[0173] In this method, when a user determines that data stored in a storage system needs to be restored based on their needs, the user inputs a second restore instruction to the user-side device via an input interface of the user-side device. In response to the second restore instruction input by the user, the user-side device generates a second restore request and sends the second restore request to the control module of the storage system. The second restore request includes the object identifier of the object to which the data to be restored belongs, and requests the restoration of the data of that object.

[0174] As an example, when a user needs to restore first data that was incorrectly operated on a first object, the user inputs a second restoration instruction including the identifier of the first object to the user-end device via an input interface of the user-end device. In response to the second restoration instruction, the user-end device generates a second restoration request including the identifier of the first object. The user-end device then sends the second restoration request to the control module via a preset protocol interface configured within the user-end device. In response, the control module receives the second restoration request sent by the user-end device via its preset protocol interface and identifies the first object represented by the object identifier carried in the second restoration request as the target object.

[0175] It should be noted that the above methods for determining the target object for which data recovery is required are merely exemplary and do not constitute a limitation to the embodiments of the present application.

[0176] Step 202: The control module of the storage system determines a first storage address of a data block storing a target object in the first snapshot according to a first mapping relationship set associated with the first snapshot.

[0177] The first snapshot is a snapshot obtained after a snapshot operation is performed on the data in the storage system at the first time. The embodiment of the present application does not describe the snapshot process in detail.

[0178] It should be understood that the control module of the storage system performs at least one snapshot operation on the data of the storage system during operation, thereby obtaining at least one snapshot, and the first snapshot is one of the at least one snapshot. In one example, the control module of the storage system periodically performs a snapshot operation on the data of the storage system during operation, thereby obtaining at least one snapshot. The embodiment of the present application does not limit the specific value of the periodic duration of the snapshot operation performed by the control module. In another example, the control module of the storage system performs at least one snapshot operation on the data of the storage system during operation by triggering a condition, thereby obtaining at least one snapshot. The embodiment of the present application does not specifically limit the implementation form of the trigger condition. In another example, the control module of the storage system combines the periodic snapshot operation on the data of the storage system and the snapshot operation on the data of the storage system triggered by a triggering condition during operation, thereby obtaining at least one snapshot, which is not limited.

[0179] Optionally, the first snapshot is a snapshot obtained after the control module performs a snapshot operation on the data of the storage system before the current moment and most recently before the current moment.

[0180] After the control module determines the target object for which data recovery is required, the control module selects any snapshot obtained from historical snapshot operations on the storage system's data as the snapshot used to recover the target object's data. For example, the control module selects the first snapshot obtained from the most recent snapshot operation on the storage system's data before the current moment as the snapshot used to recover the target object's data.

[0181] After the control module determines that the first snapshot is a snapshot for restoring the target object data, the control module determines a first mapping relationship set associated with the first snapshot based on the first time at which the first snapshot was obtained during the snapshot operation. The first mapping relationship set associated with the first snapshot is the first mapping relationship set most recently updated, maintained, and persistently stored by the storage system's control module as of the first time. That is, the first mapping relationship set associated with the first snapshot includes the correspondences between objects and storage addresses of their data blocks maintained by the storage system's controller as of the first time.

[0182] Exemplarily, the control module determines the first mapping relationship set associated with the first snapshot based on the time of the attribute information record of each corresponding relationship in the currently updated first mapping relationship set. The time of the attribute information record of each corresponding relationship in the first mapping relationship set associated with the first snapshot includes the first time and the time before the first time.

[0183] As another example, the control module determines the first mapping relationship set persistently stored at the first time, or the first mapping relationship set with the same ID as the first snapshot, as the first mapping relationship set associated with the first snapshot. It should be understood that in an embodiment of the present application, optionally, each time the control module performs a snapshot operation on the data of the storage system to obtain a snapshot, it will also persistently store the first mapping relationship set that was most recently updated in its own memory at the time of performing the snapshot operation. At this time, the snapshot corresponds to the persistently stored first mapping relationship set. Exemplarily, the control module performs a snapshot operation on the data of the storage system at the first time to obtain snapshot 1, and at the same time, the control module persistently stores a copy of the first mapping relationship set that was most recently updated in its own memory based on the received IO request at the first time. At this time, the copy of the persistently stored first mapping relationship set corresponds to snapshot 1. Furthermore, after obtaining each snapshot, the control module will associate the persistently stored first mapping relationship set that has a corresponding relationship with the snapshot. Taking snapshot 1 and the first mapping relationship set of persistent storage corresponding to snapshot 1 (denoted as mapping relationship set 1) as an example, in one example, the control module associates snapshot 1 with mapping relationship set 1 by obtaining the time of snapshot 1 and the time recorded in the attribute information of the most recently updated corresponding relationship in mapping relationship set 1 (i.e., the time when the IO request was received in the above IO data). In another example, the control module associates snapshot 1 with mapping relationship set 1 by marking an ID with a binding relationship for each of snapshot 1 and mapping relationship set 1. The IDs with a binding relationship can be the same ID or different IDs, and there is no limitation on this.

[0184] Then, the control module searches for a corresponding relationship containing the target object identifier in the first mapping relationship set associated with the first snapshot according to the identifier of the target object, and determines the storage address in the corresponding relationship as the first storage address for storing the target object data block in the first snapshot.

[0185] In one example, in conjunction with Table (1) in FIG4 , when the target object is the first object, the control module searches, based on the identifier of the first object, for a corresponding relationship containing the identifier of the first object in the first mapping relationship set shown in Table (1) associated with the first snapshot, and determines the storage address of the second data in the corresponding relationship as the first storage address for storing the target object data block in the first snapshot. In other words, the data stored in the storage system for the target object in the first snapshot is the second data of the target object.

[0186] In another example, in conjunction with Table (5) in FIG4 , when the target object is the first object, the control module searches for a corresponding relationship containing the first object identifier in the first mapping relationship set shown in Table (5) associated with the first snapshot based on the identifier of the first object, and determines the storage address of the first data and the storage address of the second data in the corresponding relationship as the first storage address for storing the target object data block in the first snapshot. In other words, the data stored in the storage system for the target object in the first snapshot includes the first data and the second data of the target object.

[0187] Step 203: The control module of the storage system restores the data of the target object according to the data stored at the first storage address in the first snapshot.

[0188] In a possible implementation, the process of the control module restoring the data of the target object according to the data stored at the first storage address in the first snapshot can be implemented through steps 2031 to 2034 shown in FIG. 6 .

[0189] Step 2031: The control module of the storage system performs a security check on the data stored at the first storage address in the first snapshot to determine whether the data stored at the first storage address in the first snapshot is infected with a virus or is attacked by a malicious attack.

[0190] The embodiments of the present application do not limit the specific implementation method of the security detection performed by the control module.

[0191] When the control module determines that the data stored at the first storage address in the first snapshot is not infected by a virus or has not been attacked maliciously, the control module executes steps 2032 to 2033. When the control module determines that the data stored at the first storage address in the first snapshot is infected by a virus or has been attacked maliciously, the control module executes step 2034.

[0192] Step 2032: The control module of the storage system determines the data stored at the first storage address in the first snapshot as the restored data of the target object, so as to restore the target object data.

[0193] The control module determines the target data stored at the first storage address in the first snapshot as the latest data after the target object is restored, and deletes the target object data stored in the storage system after the first time (the time when the snapshot operation to obtain the first snapshot was executed) to restore the target object data. The length of the target data is the operation length recorded in the attribute information of the target correspondence (i.e., the operation length described in the IO data above), and the target correspondence is the correspondence that includes the first storage address in the first mapping relationship set associated with the first snapshot.

[0194] Taking the case where the target object is the first object and the snapshot used to restore the data of the first object is the first snapshot obtained by executing the snapshot operation at the first time, as an example, the first mapping relationship set associated with the first snapshot is the first mapping relationship set persistently stored at the first time, which is recorded as mapping relationship set 1. In this case, when the control module determines, based on the mapping relationship set 1 shown in Table (6) in FIG4 , that the first storage address used to store the target object data in the first snapshot only includes the storage address of the second data (recorded as address 2), and the control module finds, based on the first mapping relationship set (recorded as mapping relationship set 2) that is most recently updated at the current moment after the first time, that the corresponding relationship containing the first object identifier includes: the identifier of the first object and the storage address of the second data (i.e., address 2), and the identifier of the first object and the storage address of the fourth data (recorded as address 4), this indicates that the control module wrote the fourth data of the first object to address 4 of the storage system after the first time. At this time, when the control module restores the data of the first object based on the second data stored at address 2 in the first snapshot, the control module determines the second data stored at address 2 as the data after the recovery of the first object, and deletes the fourth data stored at address 4 in the storage system after the first time, thereby realizing the recovery of the first object data.

[0195] Step 2033: The control module of the storage system updates the most recently updated first mapping relationship set according to the identifier of the target object and the first storage address.

[0196] Exemplarily, the control module searches for a corresponding relationship including the target object identifier in the latest updated first mapping relationship set according to the identifier of the target object, and deletes the corresponding relationships including the target object identifier except the corresponding relationship between the target object identifier and the first storage address.

[0197] Taking the case where the target object is the first object and the snapshot used to restore the data of the first object is the first snapshot obtained by executing the snapshot operation at the first time, the first mapping relationship set associated with the first snapshot is the first mapping relationship set persistently stored at the first time, which is recorded as mapping relationship set 1. In this case, when the control module determines, based on the mapping relationship set 1 shown in Table (6) in FIG4 , that the first storage address storing the target object data in the first snapshot only includes the storage address of the second data (recorded as address 2), and the control module finds the corresponding relationship containing the first object identifier in the first mapping relationship set (recorded as mapping relationship set 2) that includes: the identifier of the first object and the storage address of the second data (i.e., address 2), and the identifier of the first object and the storage address of the fourth data (recorded as address 4), this means that after the first time, the control module also wrote the fourth data of the first object to address 4 of the storage system, thereby adding the corresponding relationship between the identifier of the first object and the storage address of the fourth data to the first mapping relationship set in the memory. At this time, while or after the control module restores the first object data based on the second data stored at address 2 in the first snapshot, the control module deletes the correspondence between the identifier of the first object recorded in the mapping relationship set 2 and the storage address of the fourth data to achieve an update of the mapping relationship set 2.

[0198] Step 2034: The control module of the storage system determines the second storage address of the data block storing the target object in the second snapshot based on the first mapping relationship set associated with the second snapshot, and restores the data of the target object based on the data stored at the second storage address in the second snapshot.

[0199] The second snapshot is a snapshot obtained by performing a snapshot operation on the data in the storage system at a second time before the first time. For example, if the control module periodically performs a snapshot operation on the data in the storage system, and the control module performs a snapshot operation on the data in the storage system at the sth period to obtain the first snapshot, then the second snapshot is a snapshot obtained by performing a snapshot operation on the data in the storage system at the s-1th period. Where s is a positive integer greater than 1.

[0200] The first mapping relationship set associated with the second snapshot refers to the first mapping relationship set that is most recently updated and persistently stored by the control module of the storage system as of the second time, that is, the first mapping relationship set associated with the second snapshot includes the correspondence between the objects maintained by the control module of the storage system as of the second time and the storage addresses of the data blocks of the objects.

[0201] Specifically, the process of the control module determining the second storage address of the data block storing the target object in the second snapshot based on the first mapping relationship set associated with the second snapshot can be referred to the description of "the control module determines the first storage address of the data block storing the target object in the first snapshot based on the first mapping relationship set" in step 202, and will not be repeated here.

[0202] Furthermore, the control module performs a security check on the data stored at the second storage address in the second snapshot. If the control module determines that the data stored at the second storage address in the second snapshot is not infected by a virus or has been subjected to a malicious attack, the control module executes steps 2032 and 2033 for the second snapshot and the second storage address to update the target object data. Furthermore, if the control module determines that the data stored at the second storage address in the second snapshot is infected by a virus or has been subjected to a malicious attack, the control module executes step 2034 for the first mapping relationship set associated with the third snapshot and the third snapshot. The control module determines the third storage address of the data block of the target object stored in the third snapshot based on the first mapping relationship set associated with the third snapshot, and restores the target object data based on the data stored at the third storage address in the third snapshot. The control module repeats this process until the control module completes the restoration of the target object data.

[0203] The third snapshot is a snapshot obtained by performing a snapshot operation on the data of the storage system at a third time before the second time. For example, when the control module periodically performs a snapshot operation on the data of the storage system, and the control module performs a snapshot operation on the data of the storage system in the s-1th period to obtain the second snapshot, the third snapshot is a snapshot obtained by performing a snapshot operation on the data of the storage system in the s-2th period.

[0204] The first mapping relationship set associated with the third snapshot is the first mapping relationship set most recently updated and persistently stored by the storage system control module as of the third time. That is, the first mapping relationship set associated with the third snapshot includes the correspondences between objects and storage addresses of their data blocks maintained by the storage system control module as of the third time.

[0205] In another possible implementation, the control module directly determines the data stored at the first storage address in the first snapshot as the data after the target object is restored, and deletes the data of the target object stored in the storage system after the snapshot operation of obtaining the first snapshot is executed (i.e., after the first time), so as to realize the recovery of the target object data. The detailed process can refer to the description of step 2032 and will not be repeated here. For example, when the control module determines the target object that needs to restore data in method 3 in step 201, this possible implementation method can be used to restore the data of the target object. In addition, the control module also updates the most recently updated first mapping relationship set in the current memory based on the identifier of the target object and the first storage address. The detailed process can refer to the description of step 2033 and will not be repeated here.

[0206] At this point, through the method described in steps 201 to 203, data recovery is achieved at the granularity of semantic-level data (i.e., objects) in a storage system that stores data in a block-level storage manner. Compared with recovering data at the granularity of LUN or volume, the granularity of data recovery using the method provided in the embodiment of the present application is small, so the data recovery speed is fast and the efficiency is high.

[0207] In addition, for SAN storage systems that store data in the form of data blocks, academic research can currently achieve ransomware detection and recovery at the data block level. However, since the SAN storage system does not have a file system structure, and what is presented on the user side of the SAN storage system is not a data block, but a file with semantic information, and one file usually corresponds to multiple data blocks. Therefore, data recovery based on data block granularity may cause the loss of file information displayed on the user side. For example, when the file ransomed on the user side corresponds to 10 data blocks in the storage system, if the storage system detects and recovers 9 data blocks and misses 1 data block, the ransomed file may still be unavailable to the user side. The method provided in the embodiment of the present application performs data recovery at the granularity of semantic-level data (i.e., object), which can solve the problem of loss of file information displayed on the user side that may occur when data recovery is performed at the granularity of data blocks.

[0208] To deepen understanding of the data recovery method described in steps 201 to 203, the method provided in an embodiment of the present application is further described below with reference to an example. For example, in the case where the user-end device is a user host, the storage system is a SAN storage system, and the controller of the storage system is a SAN storage controller, reference is made to FIG7 , which shows another flow chart of the data recovery method provided in an embodiment of the present application.

[0209] As shown in Figure 7, on the user host side, the user host includes a third mapping relationship set management module, which is used to calculate the identifier of the object to which the user data to be written to the SAN storage system belongs, and to update the third mapping relationship set based on the calculated identifier. For detailed descriptions, please refer to the relevant description of step 101 above and will not be repeated here. On the SAN storage system side, the SAN storage controller includes a first mapping relationship set management module, which is used to collect IO data based on received IO requests and update the first mapping relationship set based on the collected IO data. For detailed descriptions, please refer to the descriptions of steps 101 to 104 above and will not be repeated here.

[0210] On the SAN storage system side, the SAN storage controller also includes a security detection module and a data recovery module. The security detection module is used to detect data to determine abnormal object identifiers. The object represented by the abnormal object identifier is the object whose detection results indicate that the data is infected with a virus or has been maliciously attacked. The data recovery module is used to obtain the abnormal object identifier and send a recovery request including the abnormal object identifier (such as a first recovery request including the first object identifier) ​​to the user host. A detailed description of this process can be found in the relevant description of step 201 and will not be repeated here.

[0211] On the user host side, the user host also includes an alarm recovery module, which is used to determine the target object's identifier from the abnormal object identifier received from the SAN storage controller and return the target object's identifier to the SAN storage controller via a confirmation message. Furthermore, the SAN storage controller's data recovery module restores the target object's data based on a snapshot obtained by historically performing a snapshot operation on the SAN storage system's data and the first mapping relationship set associated with the snapshot. The detailed process can be found in steps 202 to 203 and will not be further described.

[0212] The following describes the process of the user terminal device updating and maintaining the second mapping relationship set.

[0213] Specifically, the user terminal device updates the second mapping relationship set according to the object to which the data to be processed belongs and the storage address of the data to be processed.

[0214] The data to be processed can be data to be written to the storage system, data to be modified in the storage system, or data to be deleted in the storage system. The "object to which the data to be processed belongs" refers to any semantic-level data, such as, but not limited to, a document file or a spreadsheet file.

[0215] The storage address of the data to be processed may be a logical address of the data to be processed, which may be represented, for example, by the LBA of the logical block storing the data to be processed, the ID of the LUN or the ID of the storage volume to which the LBA belongs, and the length of the data to be processed. Alternatively, the storage address of the data to be processed may be a physical address where the data to be processed is stored in the storage system, which may be represented, for example, by the PBA of the physical block storing the data to be processed in the storage system, the ID of the LUN or the ID of the storage volume to which the PBA belongs, and the length of the data to be processed. It should be understood that the LBA is represented by the first address of the logical block, and the PBA is represented by the first address of the physical block, and the sizes of the physical blocks and logical blocks are generally both preset sizes.

[0216] The object to which the data to be processed belongs can be represented by the object flag of the object to which the data to be processed belongs. The object flag of the object to which the data to be processed belongs can be the object identifier of the object to which the data to be processed belongs (i.e., the object identifier described above), or can be the name, directory, or metadata of the object to which the data to be processed belongs, etc., without limitation. Among them, the process of obtaining the object identifier of the object to which the data to be processed belongs can refer to the relevant description in step 101 and will not be repeated here. It can be understood that when the object flag of the object to which the data to be processed belongs is the object identifier of the object to which the data to be processed belongs, the user terminal device also maintains the third mapping relationship set described above to record the correspondence between the object identifier of the object to which the data to be processed belongs and the object to which the data to be processed belongs.

[0217] In the first possible scenario, the data to be processed is data to be written to the storage system (recorded as data to be written). In this case, after the user-end device determines that the data to be written needs to be written to the storage system, the user-end device determines the object identifier of the object to which the data to be written belongs (recorded as the object identifier to be written). In one example, when the object identifier is the name, directory, or metadata of the object, the user-end device determines the name, directory, or metadata of the object to which the data to be written belongs as the object identifier to be written. In another example, when the object identifier is the object identifier described above, the user-end device uses the relevant process described in step 101 to determine the object identifier of the object to which the data to be written belongs, for use as the object identifier to be written, which will not be repeated.

[0218] Optionally, when the storage address of the data to be written is specified by the user-end device, the user-end device allocates a storage address for the data to be written. In this case, the storage address is typically a logical address. Furthermore, the user-end device updates the second mapping relationship set based on the determined object flag to be written and the storage address allocated for the data to be written. The detailed process of "the user-end device updates the second mapping relationship set based on the determined object flag to be written and the storage address allocated for the data to be written" can be referred to the description of "the user-end device updates the second mapping relationship set based on the object flag to be written and the storage address of the data to be written" below, and will not be repeated here.

[0219] Optionally, when the storage address of the data to be written is specified by the storage system, the user-end device generates an IO request for indicating that the data to be written is to be written to the storage system based on the data to be written, and sends the IO request to the storage system, wherein the IO request carries the data to be written. In response, the storage system receives the IO request, allocates a storage address for the data to be written carried by the IO request, and returns the storage address to the user-end device. In this case, the storage address can be a logical address or a physical address, which is not limited. Furthermore, the user-end device updates the second mapping relationship set based on the determined object flag to be written and the storage address allocated to the data to be written received from the storage system. Among them, the detailed process of "the user-end device updates the second mapping relationship set based on the determined object flag to be written and the storage address allocated to the data to be written received from the storage system" can be referred to the description of "the user-end device updates the second mapping relationship set based on the object flag to be written and the storage address allocated to the data to be written" below, which will not be repeated here.

[0220] In one example, the user-end device updates the second mapping relationship set based on the object flag to be written and the storage address of the data to be written, including: the user-end device first traverses the second mapping relationship set based on the object flag to be written, and then, when the user-end device determines that the second mapping relationship set contains the object flag to be written, continues to determine whether the storage address in the second mapping relationship set that has a corresponding relationship with the object flag to be written contains the storage address of the aforementioned data to be written. If the user-end device determines that the storage address in the second mapping relationship set that has a corresponding relationship with the object flag to be written contains the storage address of the aforementioned data to be written, the process of updating the second mapping relationship set is terminated. If the user-end device determines that the storage address in the second mapping relationship set that has a corresponding relationship with the object flag to be written does not contain the storage address of the aforementioned data to be written, and if the user-end device does not contain the object flag to be written in the second mapping relationship set, a correspondence between the object flag to be written and the storage address of the data to be written is added in the second mapping relationship set to achieve the update of the second mapping relationship set. Taking the object flag as an example, an example of the correspondence between the flag of the object to be written and the storage address of the data to be written added in the second mapping relationship set can refer to the correspondence between the identifier of the first object and the storage address of the first data added in the first mapping relationship set shown in Table (1) in Figure 4, thereby obtaining an example of the updated first mapping relationship set shown in Table (2) in Figure 4.

[0221] In the second possible scenario, the data to be processed is data to be deleted and stored in the storage system (recorded as data to be deleted). In this case, the user-end device already knows the storage address of the data to be deleted. In addition, when the user-end device determines that the data to be deleted needs to be deleted from the storage system, the user-end device determines the object flag of the object to which the data to be deleted belongs (recorded as the object flag to be deleted). Here, for a detailed description of the user-end device determining the object flag to be deleted, please refer to the description of the user-end device determining the object flag to be written in the first possible scenario, and will not be repeated here. Subsequently, the user-end device traverses the second mapping relationship set according to the object flag to be deleted to find the object flag to be deleted contained in the second mapping relationship set. Next, the user-end device deletes the storage address of the data to be deleted from the storage address contained in the second mapping relationship set that has a corresponding relationship with the object flag to be deleted, so as to achieve the correspondence between the deletion of the object flag to be deleted and the storage address of the data to be deleted. Optionally, when the user terminal device deletes the storage address of the data to be deleted from the storage address corresponding to the flag of the object to be deleted contained in the second mapping relationship set, the second mapping relationship set no longer contains the storage address corresponding to the flag of the object to be deleted. In this case, the user terminal device also deletes the flag of the object to be deleted contained in the second mapping relationship set. Taking the object flag as an example, an example of deleting the correspondence between the flag of the object to be deleted and the storage address of the data to be deleted in the second mapping relationship set can refer to the example of deleting the correspondence between the flag of the object to be deleted and the storage address of the data to be deleted in the first mapping relationship set shown in Table (5) of FIG. 4, thereby obtaining the example of the updated first mapping relationship set shown in Table (6) of FIG. 4.

[0222] In the third possible scenario, the data to be processed is the data stored in the storage system (denoted as data before modification) that is modified (denoted as data after modification). In this case, the user-end device already knows the storage address of the data before modification. Moreover, when the user-end device determines that the "data before modification" stored in the storage system needs to be modified to "data after modification", the user-end device determines the object flag of the object to which the modified data belongs (denoted as the modified object flag), and traverses the second mapping relationship set according to the modified object flag to find the modified object flag contained in the second mapping relationship set. It should be understood that the modified object flag contained in the second mapping relationship set is an object flag determined based on the data before modification. It should also be understood that the objects to which the data before modification and the data after modification belong are the same, so the object flag determined based on the data before modification and the object flag determined based on the data after modification are the same. Among them, the detailed description of the user-end device determining the modified object flag can refer to the description of the user-end device determining the object flag of the object to be written in the first possible scenario, and will not be repeated here.

[0223] Optionally, when the length of the modified data is less than the length of the data before the modification, and the storage system used to store the data before the modification supports in-place overwriting of the modified data, the user terminal device modifies the data length of the storage address of the data before the modification, in the storage address included in the second mapping relationship set and having a corresponding relationship with the modification object flag. Specifically, the user terminal device modifies the data length of the storage address of the data before the modification to the length of the data after the modification, in the storage address included in the second mapping relationship set and having a corresponding relationship with the modification object flag.

[0224] Optionally, when the length of the modified data is equal to the length of the data before modification, and the storage system used to store the data before modification supports in-situ overwriting of the modified data, the user terminal device ends the update process of the second mapping relationship set.

[0225] Optionally, when the storage system used to store the data before modification does not support in-situ overwriting of the modified data, the user-end device may first delete the data before modification stored in the storage system, and then write the modified data in the storage system. Alternatively, the user-end device may first write the modified data in the storage system, and then delete the data before modification stored in the storage system. In this case, the user-end device uses the data before modification as the data to be deleted in the second possible case, and updates the second mapping relationship set based on the data to be deleted. For detailed descriptions, please refer to the description in the second possible case, which will not be repeated here. The user-end device also uses the modified data as the data to be written in the first possible case, and updates the second mapping relationship set based on the data to be written. For detailed descriptions, please refer to the description in the first possible case, which will not be repeated here.

[0226] Based on the above process, the purpose of updating and maintaining the second mapping relationship set by the user terminal device can be achieved. Furthermore, based on the second mapping relationship set updated and maintained by the user terminal device, the data recovery method provided in the embodiment of the present application can be implemented.

[0227] The following describes the process of implementing the data recovery method provided by the embodiment of the present application based on the second mapping relationship set maintained by the user terminal device. Referring to Figure 8, Figure 8 shows a flow chart of another data recovery method provided by the embodiment of the present application. Optionally, the method is applied to the implementation environment shown in Figure 1 or Figure 2, and the corresponding steps are performed by the user terminal device and storage system in the implementation environment shown in Figure 1 or Figure 2. As shown in Figure 8, the method includes the following steps 301 to 305.

[0228] Step 301: The user terminal device determines a target object whose data needs to be restored. The target object is any semantic-level data, and the data blocks of the target object are stored in a storage system in a block-level manner.

[0229] The data of the target object is any kind of semantic-level data, and the storage system stores data blocks of the target object in a block-level manner.

[0230] The process of the user terminal device determining the target object for which data recovery is required can be implemented in the following ways. It should be noted that the following ways for determining the target object for which data recovery is required are only exemplary and do not constitute a limitation to the embodiments of the present application.

[0231] Mode 1: The user terminal device determines a target object whose data needs to be restored in response to an object restoration instruction input by a user.

[0232] In this method, when a user determines that data stored in a storage system needs to be restored based on their own needs, the user inputs an object recovery instruction to the user-end device through the input interface of the user-end device, and the object recovery instruction carries the object identifier of the object for which data needs to be restored (i.e., the target object). In response to the object recovery instruction input by the user, the user-end device determines the object to which the data to be restored indicated by the object recovery instruction belongs as the target object. In this case, the number of target objects can be one or more, and there is no limitation on this. It can be seen that the object recovery instruction here is the second recovery instruction in step 201 above.

[0233] For example, when a user needs to restore first data in a first object that was incorrectly operated, the user inputs a second restoration instruction including the object identifier of the first object to the user terminal device through an input interface of the user terminal device. The user terminal device then determines the first object as a target object in response to the second restoration instruction.

[0234] Method 2: The user-end device determines the target object for which data needs to be restored based on the abnormal data blocks in the storage system.

[0235] Referring to Figure 9 , a schematic diagram illustrating a process for a user-end device, provided in an embodiment of the present application, to determine a target object requiring data recovery based on abnormal data blocks in a storage system is shown. Optionally, this process is applied to the implementation environment shown in Figure 1 or Figure 2 , with the user-end device and storage system in the implementation environment shown in Figure 1 or Figure 2 performing the corresponding steps. As shown in Figure 9 , this process includes the following steps 3011 through 3015.

[0236] Step 3011: The storage system detects whether the data stored in the storage system and / or the received IO requests are infected by viruses or attacked by malicious means.

[0237] Specifically, the storage system detects whether the data stored in the storage system is infected by viruses or attacked maliciously, including: the storage system performs security detection on the data blocks stored in the storage system to determine whether the data is infected by viruses or attacked maliciously.

[0238] Optionally, the storage system may periodically perform security checks on the data blocks stored therein, so as to periodically determine whether the data blocks stored in the storage system are infected by viruses or attacked by malicious means.

[0239] Optionally, the storage system may periodically perform security checks on all or part of the data blocks stored in the storage system to periodically determine whether the data blocks are infected by viruses or attacked maliciously. The data stored in some of the data blocks may be hot data stored in the storage system, which is not limited to this.

[0240] When the storage system periodically performs security checks on all or part of the data blocks stored in the storage system, the embodiment of the present application executes the following steps 3012 to 3015 based on the detection results in each cycle.

[0241] Furthermore, the storage system detects whether a received IO request is infected by a virus or subjected to a malicious attack, including: the storage system performs a security check on data carried in the received IO request to determine whether the data is infected by a virus or subjected to a malicious attack. Optionally, the storage system detects whether a received IO request is infected by a virus or subjected to a malicious attack, further including: the storage system detects features of the received IO request to determine whether the IO request is a malicious IO request.

[0242] It should be understood that the embodiments of the present application do not describe in detail the process of detecting whether data is infected with a virus or is subjected to a malicious attack, and do not describe in detail the process of detecting whether an IO request is a malicious IO request.

[0243] Step 3012: The storage system sends the storage address of the abnormal data block to the user terminal device. The abnormal data block is a data block where data infected by a virus or maliciously attacked is located.

[0244] The abnormal data block is used by the user-end device to determine the target object. The storage address of the abnormal data block can be the logical address or physical address of the abnormal data block. The following examples 1 to 3 describe the process of the storage system sending the storage address of the abnormal data block to the user-end device.

[0245] In Example 1, when a storage system detects that one or more data blocks stored within it are infected by a virus or are maliciously attacked, the storage system identifies the one or more data blocks infected by the virus or are maliciously attacked as abnormal data blocks. That is, the number of abnormal data blocks is one or more. The storage system then reads the storage address of each abnormal data block and sends the storage address of each abnormal data block to the user-end device via the communication protocol interface between the storage system and the user-end device.

[0246] Example 2: When the storage system detects that the first data carried by the received IO request is infected by a virus or is under malicious attack, if the storage address of the first data carried in the IO request is allocated by the user-end device, the IO request includes the ID of the LUN or storage volume to which the first data belongs, the length of the first data, and the LBA allocated by the user-end device for the first data, then the storage system sends the ID of the LUN or storage volume to which the first data carried in the IO request belongs, the length of the first data, and the LBA for storing the first data as the storage address of the abnormal data block to the user-end device. If the storage address of the first data carried in the IO request is allocated by the storage system, then the storage system sends the ID of the LUN or storage volume to which the first data carried in the IO request belongs, the length of the first data, and the LBA or PBA allocated by the storage system for the first data as the storage address of the abnormal data block to the user-end device.

[0247] Example 3, when the storage system detects that the received IO request is a malicious IO request, the storage system determines that the data requested to be processed by the IO request is abnormal data. If the IO request carries the storage address of the data requested to be processed by the IO request, the storage system sends the storage address carried by the IO request as the storage address of the abnormal data block to the user-end device. If the IO request does not carry the storage address of the data requested to be processed by the IO request, then the IO request is an IO request that carries the first data and is used to instruct the first data to be written to the storage system. In this case, the storage system can send the storage address of the abnormal data block to the user-end device according to the process described in Example 2.

[0248] Optionally, when there are multiple abnormal data blocks, the storage system may send the storage addresses of the multiple abnormal data blocks as one piece of information to the user terminal device, or may send them as multiple pieces of information to the user terminal device, which is not limited.

[0249] Step 3013: The user terminal device receives the storage address of the abnormal data block sent by the storage system.

[0250] Exemplarily, the user terminal device receives the storage address of the abnormal data block sent by the storage system through the communication protocol interface between the user terminal device and the storage system.

[0251] Step 3014: The user terminal device queries the second mapping relationship set according to the storage address of the abnormal data block to determine the abnormal object.

[0252] The abnormal object refers to the object to which the data recorded in the abnormal data block belongs, that is, the abnormal object is an object infected by a virus or maliciously attacked. The second mapping relationship set includes the correspondence between the objects maintained by the user terminal device and the storage addresses of the data blocks of the objects in the storage system. For detailed description, please refer to the description of the second mapping relationship set above and will not be repeated here.

[0253] In one possible implementation, when the type of the storage address of the abnormal data block sent by the storage system to the user-end device is the same as the type of the storage address in the second mapping relationship set updated and maintained by the user-end device, the user-end device directly queries the second mapping relationship set based on the storage address of the abnormal data block, and determines the object represented by the object flag in the second mapping relationship set that has a corresponding relationship with the storage address of the abnormal data block as an abnormal object.

[0254] The types of storage addresses include logical addresses and physical addresses.

[0255] In another possible implementation, when the type of the storage address of the abnormal data block sent by the storage system to the user-end device is different from the type of the storage address recorded in the second mapping relationship set, the user-end device must first query its own configured address mapping relationship set based on the storage address of the abnormal data block received from the storage system (recorded as the fourth storage address) to determine the storage address (recorded as the fifth storage address) that has a mapping relationship with the fourth storage address. Then, the user-end device queries the second mapping relationship set based on the fifth storage address and determines the object represented by the object flag in the second mapping relationship set that has a corresponding relationship with the fifth storage address as an abnormal object.

[0256] The address mapping relationship set includes the mapping relationship between the logical address and the physical address of the data stored in the storage system. The present embodiment of the application does not describe in detail the process of configuring the address mapping relationship set on the user-end device. For example, the storage system generally maintains the address mapping relationship set during the data storage process. Therefore, the user-end device can obtain the mapping relationship in the address mapping relationship set from the storage system in real time and store it to obtain the address mapping relationship set maintained by itself.

[0257] As an example, when the storage address recorded in the second mapping relationship set is a logical address, and the storage address of the abnormal data block sent by the storage system to the user-end device is a physical address (denoted as physical address 1), the user-end device queries the address mapping relationship set based on the received physical address 1 to determine the logical address 1 that has a mapping relationship with physical address 1. Next, the user-end device queries the second mapping relationship set based on logical address 1 and determines the object represented by the object flag that has a corresponding relationship with logical address 1 in the second mapping relationship set as an abnormal object.

[0258] Optionally, when the object identifier recorded in the second mapping relationship set is the object identifier described in step 101 above, the user terminal device also queries the third mapping relationship set described above based on the object identifier of the abnormal object queried based on the second mapping relationship set to determine the abnormal object corresponding to the object identifier of the abnormal object.

[0259] Step 3015: The user terminal device determines a target object from the determined abnormal objects.

[0260] In a possible implementation, the user terminal device directly determines the abnormal object as the target object.

[0261] In another possible implementation, the user-end device identifies abnormal objects whose data is hot data as target objects. In one example, the user-end device first determines the heat of the data included in each abnormal object, and then identifies abnormal objects whose data heat value is greater than a threshold as target objects, where data with a heat value greater than the threshold is considered hot data.

[0262] In another possible implementation, the user terminal device determines a target object from among the abnormal objects in response to a first restoration instruction input by the user.

[0263] Specifically, after the user terminal device determines the abnormal object, it outputs an alarm message to the user. The alarm message is used to indicate that the data of the abnormal object is infected with a virus or is under malicious attack, and is used to determine whether the data of the abnormal object needs to be restored. Exemplarily, the user terminal device can output the alarm message through its own output interface (such as a display screen, a monitor, a speaker, etc.). Based on the alarm message output by the user terminal device, the user can choose whether to restore the data of the abnormal object based on their own needs, and choose which objects among the abnormal objects indicated by the alarm message to restore based on their own needs.

[0264] Furthermore, the user terminal device receives a first recovery instruction input by the user to the user terminal device through an input interface of the user terminal device (such as a mouse, keyboard, touch screen, microphone, or other input interface). When the first recovery instruction received by the user terminal device is used to indicate that it is not necessary to recover the data of all abnormal objects, the user terminal device ends the data recovery process in response. When the first recovery instruction received by the user terminal device is used to indicate that the data of all or part of the abnormal objects in the abnormal objects need to be recovered, the user terminal device determines the abnormal object selected by the user for which data recovery is required as the target object in response to the first recovery instruction. In this case, the number of target objects is one or more.

[0265] It can be seen that the alarm information and the first restoration indication here are the same as the alarm information set first restoration indication described in step 201 above.

[0266] Step 302: The user terminal device queries the second mapping relationship set according to the object identifier of the target object to determine the storage address of the data block of the target object.

[0267] The second mapping relationship set includes the correspondence between objects maintained by the user terminal device and storage addresses of data blocks of the objects in the storage system. For detailed description, please refer to the above description of the second mapping relationship set and will not be repeated here.

[0268] Specifically, the user terminal device queries the second mapping relationship set based on the object identifier of the target object and determines all storage addresses in the second mapping relationship set that have a corresponding relationship with the object identifier of the target object as the storage address of the data block of the target object. For simplicity, the "storage address of the data block of the target object" will be referred to as the target storage address below.

[0269] Step 303: The user terminal device sends a data recovery instruction to the storage system, where the data recovery instruction includes a target storage address.

[0270] The data recovery instruction is used to instruct the storage system to recover the data of the target object.

[0271] Exemplarily, the user-end device sends the storage address of the data block of the target object to the storage system through the communication protocol interface between the user-end device and the storage system.

[0272] Step 304: The storage system receives a data recovery instruction including a target storage address sent by the user terminal device.

[0273] The data recovery indication includes the storage address of the data block of the target object, the target object is any kind of semantic-level data, and the data block of the target object is stored in the storage system in a block-level manner.

[0274] Exemplarily, the storage system receives the data recovery instruction sent by the user terminal device through the communication protocol interface between the storage system and the user terminal device.

[0275] Step 305: The storage system restores the data of the target object based on the data stored at the target storage address in the first snapshot.

[0276] The first snapshot is a snapshot obtained after a snapshot operation is performed on the data in the storage system at the first time. Detailed descriptions of the first snapshot can be found in the above descriptions of step 202 and step 203, which will not be repeated here.

[0277] In one possible implementation, the storage system directly determines the data stored at the target storage address in the first snapshot as the data after the target object is restored. That is, the storage system rolls back the data of the target object to the data stored at the target storage address in the first snapshot, thereby completing the recovery of the target object data.

[0278] In another possible implementation, when the storage system restores the data of the target object based on the data stored at the target storage address in the first snapshot, the storage system first detects whether the data stored at the target storage address in the first snapshot is infected by a virus or attacked maliciously.

[0279] Then, when the storage system determines that the data stored at the target storage address in the first snapshot is not infected by a virus or maliciously attacked, the storage system determines the data stored at the target storage address in the first snapshot as the data after the target object is recovered. That is, the storage system rolls back the data of the target object to the data stored at the target storage address in the first snapshot, thereby completing the recovery of the target object data.

[0280] If the storage system determines that the data stored at the target storage address in the first snapshot is infected by a virus or is under a malicious attack, the storage system restores the data of the target object based on the data stored at the target storage address in the second snapshot. For a detailed description of the second snapshot, refer to the description of the second snapshot in step 203. For a detailed description of the storage system restoring the data of the target object based on the data stored at the target storage address in the second snapshot, refer to the description of the storage system restoring the data of the target object based on the data stored at the target storage address in the first snapshot, and are not further described.

[0281] At this point, through the method described in steps 301 to 305, data recovery is achieved in a storage system that stores data in a block-level storage manner, with semantic-level data (i.e., object) as the granularity. Compared with recovering data at the granularity of LUN or volume, the granularity of data recovery using the method provided in the embodiment of the present application is small, so the data recovery speed is fast and the efficiency is high. In addition, for SAN storage systems that store data in the form of data blocks, current academic research can achieve ransomware detection and recovery at the data block level. However, since the SAN storage system does not have a file system structure, and what is presented on the user side of the SAN storage system is not a data block, but a file with semantic information, and one file usually corresponds to multiple data blocks. Therefore, data recovery based on data block granularity may cause the file information displayed on the user side to be lost. For example, when the file that is ransomed on the user side corresponds to 10 data blocks in the storage system, if the storage system detects and recovers 9 data blocks, but misses 1 data block, the ransomed file may still be unavailable to the user side. The method provided in the embodiment of the present application performs data recovery at the granularity of semantic-level data (i.e., objects), which can solve the problem of loss of user-side displayed file information that may occur when data recovery is performed at the granularity of data blocks.

[0282] To deepen the understanding of the data recovery method described in steps 301 to 305, the method provided in the embodiment of the present application is further described below with reference to an example. Assuming that the user-end device is a user host, the storage system is a SAN storage system, and the steps performed by the storage system in the method described in Figures 8 or 9 are specifically performed by a controller of the SAN storage system, i.e., a SAN storage controller, reference is made to Figure 10, which shows another flow chart of the data recovery method provided in the embodiment of the present application.

[0283] As shown in Figure 10, on the user host side, the user host includes a second mapping relationship set management module, which is used to update the second mapping relationship set based on the object flag and storage address of the object to which the data to be processed belongs. For detailed description, please refer to the above description of the process of "user-end device updating and maintaining the second mapping relationship set", which will not be repeated here.

[0284] On the SAN storage system side, the SAN storage controller includes a security detection module and a data recovery module. The security detection module is used to detect data and determine the storage address of abnormal data blocks. A detailed description of this process can be found in the description of step 3011 and will not be repeated here. The data recovery module is used to obtain the storage address of the abnormal data block and send the obtained storage address of the abnormal data block to the user host. A detailed description of this process can be found in the description of step 3012 and will not be repeated here.

[0285] On the user host side, the user host also includes an alarm recovery module, which is used to receive the storage address of the abnormal data block from the SAN storage controller, and determine the abnormal object based on the storage address of the abnormal data block and the second mapping relationship set, and then determine the target object in the abnormal object through the alarm, and query the storage address of the data block of the target object through the second mapping relationship set. A detailed description of this process can be found in steps 3013 to 3015, as well as the relevant description of step 302, which will not be repeated here. The user host then sends the storage address of the data block of the target object to the SAN storage controller (see the description of step 303).

[0286] Then, the data recovery module of the SAN storage controller recovers the target object's data based on the snapshots previously taken of the SAN storage system's data and the storage addresses of the target object's data blocks received from the user host. The detailed process can be found in the description of step 305 and will not be repeated here.

[0287] The above mainly introduces the solution provided in the embodiment of the present application from the perspective of method.

[0288] To achieve the above functions, as shown in Figure 11, a schematic diagram of the structure of a data recovery device provided in an embodiment of the present application is shown. Data recovery device 1100 is applied to a storage system that uses block-level storage for data storage and is used to execute the data recovery method described above, for example, the portion of the method shown in Figures 3, 5, or 6 that is executed by the storage system. Data recovery device 1100 may include a determination unit 801 and a recovery unit 802.

[0289] A determination unit 801 is configured to determine a target object for which data needs to be restored, and to determine a first storage address of a data block storing the target object in the first snapshot based on a first mapping relationship set associated with the first snapshot. A recovery unit 802 is configured to restore the data of the target object based on the data stored at the first storage address in the first snapshot. The target object is any type of semantic-level data, and the storage system stores data blocks of the target object stored in a block-level manner. The first snapshot is a snapshot obtained after a snapshot operation is performed on the data of the storage system at a first time. The first mapping relationship set associated with the first snapshot includes a correspondence between objects maintained by the storage system as of the first time and the storage addresses of the data blocks of the objects.

[0290] As an example, in conjunction with FIG. 5 , the determining unit 801 may be configured to execute step 201 and step 202 , and the recovering unit 802 may be configured to execute step 203 .

[0291] Optionally, the data recovery device 1100 also includes: a receiving unit 803, used to receive an IO request, the IO request is used to indicate a write operation to the data to be processed, the IO request includes an identifier of the object to which the data to be processed belongs, and the object to which the data to be processed belongs is any semantic-level data; an updating unit 804, used to update the first mapping relationship set according to the identifier of the object to which the data to be processed belongs and the storage address of the data to be processed.

[0292] As an example, in conjunction with FIG3 , the receiving unit 803 may be used to execute step 103 , and the updating unit 804 may be used to execute step 104 .

[0293] Optionally, receiving unit 803 is configured to receive an IO request carrying a first identifier and first data, where the first data is data of a first object represented by the first identifier. Data recovery apparatus 1100 further includes a detection unit 805 configured to detect whether the first data is infected with a virus or subjected to a malicious attack. Determining unit 801 is specifically configured to determine the first object as a target object if it is determined that the first data is infected with a virus or subjected to a malicious attack.

[0294] Optionally, the detection unit 805 is further configured to detect data blocks stored in the storage system to determine target data blocks in the storage system that are infected with a virus or are under malicious attack. The determination unit 801 is specifically configured to query a first mapping relationship set based on the storage address of the target data block to determine a first object corresponding to the storage address of the target data block, and to determine the first object as the target object.

[0295] Optionally, data recovery apparatus 1100 further includes a sending unit 806 configured to send a first recovery request to a user-end device before determining the first object as the target object, the first recovery request including an identifier of the first object and used to confirm whether to recover the data of the first object. A receiving unit 803 further configured to receive a confirmation message returned by the user-end device, the confirmation message indicating that the data of the first object should be recovered.

[0296] Optionally, the sending unit 806 is further configured to send a first recovery request to the user terminal device before determining the first object as the target object, the first recovery request including an identifier of the first object and used to confirm whether to recover the data of the first object. The receiving unit 803 is further configured to receive a confirmation message returned by the user terminal device, the confirmation message being used to indicate that the data of the first object should be recovered.

[0297] Optionally, the determining unit 801 is specifically configured to determine the first object as the target object when receiving a second restoration request sent by the user terminal device. The second restoration request includes an identifier of the first object, and the second restoration request is used to request restoration of data of the first object.

[0298] Optionally, the recovery unit 802 is specifically used to detect whether the data stored at the first storage address in the first snapshot is infected by a virus or maliciously attacked; when it is determined that the data stored at the first storage address in the first snapshot is not infected by a virus or maliciously attacked, the data stored at the first storage address in the first snapshot is determined as the data after the target object is recovered.

[0299] As an example, in conjunction with FIG6 , the recovery unit 802 may be configured to execute steps 2031 to 2032 .

[0300] Optionally, the updating unit 804 is further configured to update the first mapping relationship set currently maintained by the storage system according to the identifier of the target object and the first storage address.

[0301] As an example, in conjunction with FIG. 6 , the updating unit 804 may be configured to perform step 2033 .

[0302] Optionally, the determination unit 801 is further configured to, upon determining that the data stored at the first storage address in the first snapshot is infected by a virus or maliciously attacked, determine, based on the first mapping relationship set associated with the second snapshot, a second storage address for storing the data block of the target object in the second snapshot. The recovery unit 802 is further configured to recover the data of the target object based on the data stored at the second storage address in the second snapshot. The second snapshot is a snapshot obtained by performing a snapshot operation on the storage system at a second time prior to the first time, and the first mapping relationship set associated with the second snapshot includes a correspondence between the storage addresses of the objects and the data blocks of the objects maintained by the storage system as of the second time.

[0303] As an example, in conjunction with FIG6 , the determining unit 801 and the restoring unit 802 may be configured to execute step 2034 .

[0304] Optionally, the data recovery device 1100 also includes: a snapshot unit 807, which is used to perform at least one snapshot operation on the data of the storage system before determining the first storage address of the data block storing the target object in the first snapshot based on the first mapping relationship set associated with the first snapshot, to obtain at least one snapshot, and the first snapshot is one of the at least one snapshot.

[0305] Optionally, the snapshot unit 807 is specifically configured to periodically perform a snapshot operation on the data in the storage system to obtain at least one snapshot.

[0306] Optionally, the first snapshot is a snapshot obtained after a snapshot operation is performed on data in the storage system before the current moment and most recently before the current moment.

[0307] Optionally, the target object is a file or a database / table.

[0308] Optionally, the storage address is represented by an LBA and an ID of a LUN or storage volume to which the LBA belongs, or the storage address is represented by a PBA and an ID of a LUN or storage volume to which the PBA belongs.

[0309] Optionally, the snapshot is obtained by performing a snapshot operation on the data of the storage system at a granularity of LUN or storage volume.

[0310] Optionally, the storage system is a SAN-based data storage system. Alternatively, the storage system is a DAS storage system.

[0311] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation and description of the beneficial effects of any of the above data recovery devices 1100 can refer to the above corresponding method embodiments, which will not be repeated here.

[0312] As an example, with reference to FIG14 described below, the functions implemented by the determination unit 801, the recovery unit 802, the update unit 804, the detection unit 805, and the snapshot unit 807 in the data recovery device 1100 can be implemented by the processor 1101 in FIG14 executing the program code in the memory 1102 in FIG14. The functions implemented by the receiving unit 803 and the sending unit 806 in the data recovery device 1100 can be implemented by the network interface 1103 shown in FIG14.

[0313] As shown in Figure 12, Figure 12 shows a schematic diagram of the structure of another data recovery device provided by an embodiment of the present application. Data recovery device 1200 is applied to a user-end device accessing a storage system that uses block-level storage for data storage. Data recovery device 1200 is used to execute the data recovery method described above, for example, to execute the portion of the method shown in Figures 3, 5, or 6 that is executed by the user-end device. Data recovery device 1200 may include a generation unit 901 and a sending unit 902.

[0314] A generating unit 901 is configured to generate an IO request for instructing a write operation to be performed on the data of a first object, wherein the IO request includes an identifier of the first object, and the data of the first object is any type of semantic-level data. A sending unit 902 is configured to send the IO request to a storage system, wherein the identifier of the first object carried in the IO request is used to update a first mapping relationship set. The first mapping relationship set includes a correspondence between an object and a storage address of the object's data, and the first mapping relationship set is used to determine the storage address of the restored data of the first object in a snapshot when restoring the data of the first object, wherein the snapshot refers to a snapshot obtained after performing a snapshot operation on the data of the storage system.

[0315] As an example, in conjunction with FIG3 , the generating unit 901 may be used to perform step 101 , and the sending unit 902 may be used to perform step 102 .

[0316] Optionally, data recovery apparatus 1200 further includes a receiving unit 903 configured to receive a first recovery request sent by the storage system, the first recovery request including an identifier of the first object, the first recovery request being used to confirm whether to recover the data of the first object. A sending unit 902 configured to return a confirmation message to the storage system upon determining that the data of the first object should be recovered, the confirmation message being used to indicate that the data of the first object should be recovered.

[0317] Optionally, the receiving unit 903 is specifically configured to receive the first recovery request sent by the storage system through a preset protocol interface. The sending unit 902 is specifically configured to send the confirmation message to the storage system through the preset protocol interface.

[0318] Optionally, the data recovery device 1200 further includes: a query unit 904 configured to query a third mapping relationship set based on the identifier of the first object carried in the first recovery request to determine the first object, wherein the third mapping relationship set is configured to record the correspondence between the object identifier and the object; an output unit 905 configured to output an alarm message indicating that the data of the first object is infected with a virus or is under a malicious attack; and a sending unit 902 configured to return a confirmation message to the storage system in response to a first recovery instruction input by the user based on the alarm information, wherein the first recovery instruction is configured to instruct the recovery of the data of the first object.

[0319] Optionally, the data recovery device 1200 also includes: a determination unit 906, used to determine the identifier of the first object before querying the third mapping relationship set based on the identifier of the first object carried in the first recovery request; an update unit 907, used to update the third mapping relationship set based on the identifier of the first object.

[0320] Optionally, the sending unit 902 is further configured to send a second recovery request to the storage system, where the second recovery request includes an identifier of the first object, and the second recovery request is used to request recovery of data of the first object.

[0321] Optionally, the storage address is represented by an LBA and an ID of a LUN or storage volume to which the LBA belongs, or the storage address is represented by a PBA and an ID of a LUN or storage volume to which the PBA belongs.

[0322] Optionally, the first object is a file or a database / table.

[0323] Optionally, the snapshot is obtained by performing a snapshot operation on the data of the storage system at a granularity of LUN or storage volume.

[0324] Optionally, the storage system is a SAN-based data storage system, or the storage system is a DAS storage system.

[0325] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation and description of the beneficial effects of any of the above data recovery devices 1200 can refer to the above corresponding method embodiments, which will not be repeated here.

[0326] As an example, with reference to FIG14 described below, the functions implemented by the generation unit 901, query unit 904, determination unit 906, and update unit 907 in the data recovery device 1200 can be implemented by the processor 1101 in FIG14 executing the program code in the memory 1102 in FIG14. The functions implemented by the sending unit 902 and receiving unit 903 in the data recovery device 1200 can be implemented by the network interface 1103 in FIG14. The functions implemented by the output unit 905 in the data recovery device 1200 can be implemented by the input / output interface 1105 in FIG14.

[0327] As shown in Figure 13, Figure 13 shows a structural schematic diagram of another data recovery device provided by an embodiment of the present application. The data recovery device 1300 can be applied to a storage system that uses block-level storage for data storage, and is used to execute the above-mentioned data recovery method, for example, to execute the portion of the method shown in Figure 3, Figure 5 or Figure 6 that is executed by the storage system. The data recovery device 1300 can also be applied to a user-end device that accesses the aforementioned storage system, and is used to execute the data recovery method provided by the embodiment of the present application, for example, to execute the portion of the method shown in Figure 3, Figure 5 or Figure 6 that is executed by the user-end device. The data recovery device 1300 may include a transceiver unit 1001 and a processing unit 1002. The transceiver unit 1001 is used to execute operations related to receiving and / or sending in the method provided by the embodiment of the present application, and the processing unit 1002 is used to execute other operations other than operations related to receiving and / or sending in the method provided by the embodiment of the present application. The explanation of the data recovery device 1300 and the description of its beneficial effects can be referred to the corresponding method embodiments above and will not be repeated here.

[0328] As an example, in conjunction with FIG14 described below, the functions implemented by the transceiver unit 1001 in the data recovery device 1300 can be implemented by the network interface 1103 in FIG14. The functions implemented by the processing unit 1002 in the data recovery device 1300 can be implemented by the processor 1101 in FIG14 executing the program code in the memory 1102 in FIG14.

[0329] It should be readily apparent to those skilled in the art that, in combination with the units and algorithmic steps of the various examples described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is performed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0330] It should be noted that the division of modules / units in Figures 11, 12, or 13 is schematic and merely represents a logical functional division. In actual implementation, other divisions may be employed. For example, two or more functions may be integrated into a single processing module. Such integrated modules may be implemented in either hardware or software functional modules.

[0331] The present invention provides a computing device that is used to implement some or all of the functions of the data recovery method provided in the present invention. For example, the computing device can be a control module in the storage system described above, or a user terminal device for accessing the storage system described above.

[0332] Figure 14 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in Figure 14, the computing device 1400 includes a processor 1101, a memory 1102, a network interface 1103, a bus 1104, and an input / output interface 1105. The processor 1101, the memory 1102, the network interface 1103, and the input / output interface 1105 are connected to each other via the bus 1104.

[0333] The processor 1101 may include a general-purpose processor and / or a dedicated hardware chip. A general-purpose processor may include: a central processing unit (CPU), a microprocessor or a graphics processing unit (GPU). The CPU is, for example, a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A dedicated hardware chip is a hardware module for high-performance processing. Dedicated hardware chips include at least one of a digital signal processor (DSP), a data processing unit (DPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, a neural processing unit (NPU), a tensor processing unit (TPU), an artificial intelligence (artificial intelligent) chip or a network processor (NP). The processor 1101 may also be an integrated circuit chip having signal processing capabilities. During the implementation process, part or all of the functions of the method provided in the embodiment of the present application can be completed through the hardware integrated logic circuit in the processor 1101 or instructions in software form.

[0334] The memory 1102 is used to store computer programs, including an operating system 1102a and executable code (i.e., program instructions) 1102b. The memory 1102 is, for example, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, or other types of static storage devices that can store static information and instructions, and is also, for example, a static RAM (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), or a synchronous link dynamic random access memory (synchlink). DRAM, SLDRAM) or other types of dynamic storage devices that can store information and instructions, such as read-only optical discs or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired executable code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. For example, the memory 1102 is used to store the first mapping relationship set, the second mapping relationship set, etc. The memory 1102 is, for example, independent and connected to the processor 1101 via the bus 1104. Or the memory 1102 and the processor 1101 are integrated together. The memory 1102 can store executable code, and when the executable code stored in the memory 1102 is executed by the processor 1101, the processor 1101 is used to execute part or all of the functions of the data recovery method provided in the embodiment of the present application. For the implementation method of the processor 1101 executing the process, please refer to the relevant description in the aforementioned embodiment. The memory 1102 may also include software modules and data required for other running processes such as an operating system.

[0335] The network interface 1103 uses a transceiver module, such as, but not limited to, a transceiver, to communicate with other devices or communication networks. For example, the network interface 1103 can be any one or a combination of the following devices: a network interface (e.g., an Ethernet interface), a wireless network card, or other devices with network access capabilities. The network interface 1103 includes a receiving unit for receiving data / messages and a transmitting unit for sending data / messages.

[0336] The bus 1104 is any type of communication bus used to interconnect the internal components of the computing device 1400 (e.g., the memory 1102, the processor 1101, the network interface 1103, etc.). For example, a system bus is provided. The embodiments of the present application illustrate the example of interconnecting the aforementioned components within the computing device 1400 via the bus 1104. Alternatively, the aforementioned components within the computing device 1400 may be communicatively connected to each other using other connection methods besides the bus 1104, for example, the aforementioned components within the computing device 1400 may be interconnected via an internal logical interface.

[0337] Input / output interface 1105 is used to enable human-computer interaction between a user and computing device 1400. For example, text or voice interaction between the user and computing device 1400 can be implemented. Input / output interface 1105 includes an input interface for enabling a user to input information into computing device 1400, and an output interface for enabling computing device 1400 to output information to the user. By way of example, input interfaces include, but are not limited to, a touch screen, keyboard, mouse, or microphone, while output interfaces include, but are not limited to, a display screen and speakers. A touch screen, keyboard, or mouse is used to input text / image information, a microphone is used to input voice information, a display screen is used to output text / image information, and a speaker is used to output voice information.

[0338] It should be noted that the above-mentioned multiple devices can be respectively arranged on independent chips, or at least partially or completely arranged on the same chip. Whether each device is independently arranged on different chips or integrated on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation form of the above-mentioned devices. The descriptions of the processes corresponding to the above-mentioned figures have different focuses. For parts that are not described in detail in a certain process, please refer to the relevant descriptions of other processes.

[0339] In the above embodiments, all or part of the embodiments may be implemented through software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product providing a program development platform includes one or more computer instructions. When these computer program instructions are loaded and executed on the computing device 1400, all or part of the functions of the data recovery method provided in the embodiments of the present application are implemented in whole or in part.

[0340] Furthermore, computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium stores computer program instructions that provide a program development platform.

[0341] An embodiment of the present application also provides a data recovery system, which includes a user terminal device and a storage system. The data recovery method described above is used in the data recovery system, for example, the method described in Figures 3, 5, 6 or 7 above is applied to the data recovery system.

[0342] The present application also provides another data recovery system, as shown in FIG15 . FIG15 shows a schematic diagram of the structure of a data recovery system provided by the present application. The data recovery method described above is applied to data recovery system 1500 shown in FIG15 . For example, the method described in FIG8 , FIG9 , or FIG10 above is applied to data recovery system 1500 . Data recovery system 1500 includes a user terminal device 1501 and a storage system 1502 .

[0343] User-end device 1501 is configured to determine a target object for which data recovery is required, query a second mapping relationship set based on an object identifier of the target object to determine the storage address of a data block of the target object, i.e., a target storage address, and send a data recovery instruction including the target storage address to storage system 1502. The data recovery instruction is used to instruct the recovery of data of the target object based on the target storage address. The target object is any type of semantic-level data, and the data blocks of the target object are stored in storage system 1502 at the block level. Storage system 1502 uses block-level storage to store data. The second mapping relationship set includes a correspondence between objects maintained by user-end device 1501 and the storage addresses of the data blocks of the objects in storage system 1502.

[0344] As an example, in conjunction with FIG8 , the user terminal device 1501 may be configured to execute steps 301 to 303 .

[0345] Optionally, storage system 1502 is configured to receive a data recovery instruction including a target storage address sent by user terminal device 1501, and to recover the data of the target object based on the data stored at the target storage address in the first snapshot. The first snapshot is a snapshot obtained by performing a snapshot operation on the data in storage system 1502 at the first time.

[0346] As an example, in conjunction with FIG8 , the storage system 1502 may be used to perform steps 304 to 305 .

[0347] Optionally, storage system 1502 is further configured to detect whether data stored in storage system 1502 is infected by a virus or maliciously attacked, and to send the storage address of an abnormal data block to user-end device 1501. The abnormal data block is a data block containing data infected by a virus or maliciously attacked. User-end device 1501 is specifically configured to receive the storage address of the abnormal data block sent by storage system 1502, query the second mapping relationship set based on the storage address of the abnormal data block to determine an abnormal object, and determine the target object within the abnormal object.

[0348] As an example, in conjunction with FIG. 9 , the storage system 1502 may be used to execute steps 3011 to 3012 , and the user terminal device 1501 may be used to execute steps 3013 to 3015 .

[0349] Optionally, the user terminal device 1501 is specifically configured to determine a target object in response to a received object restoration instruction, wherein the object restoration instruction includes an object identifier of the target object.

[0350] Optionally, the storage system 1502 is specifically used to detect whether the data stored at the target storage address in the first snapshot is infected by a virus or maliciously attacked, and when it is determined that the data stored at the target storage address in the first snapshot is not infected by a virus or maliciously attacked, the data stored at the target storage address in the first snapshot is determined as the data after the target object is recovered.

[0351] Optionally, storage system 1502 is further configured to restore the target object's data based on the data stored at the target storage address in a second snapshot, if it is determined that the data stored at the target storage address in the first snapshot is infected by a virus or is under a malicious attack. The second snapshot is a snapshot obtained by performing a snapshot operation on storage system 1502 at a second time prior to the first time.

[0352] Optionally, the target object is a file or a database / table.

[0353] Optionally, the storage address is represented by an LBA and an ID of a LUN or storage volume to which the LBA belongs, or the storage address is represented by a PBA and an ID of a LUN or storage volume to which the PBA belongs.

[0354] Optionally, the above-mentioned snapshot is a snapshot obtained by performing a snapshot operation on the data of the storage system 1502 at a granularity of LUN or storage volume.

[0355] Optionally, the storage system 1502 is a SAN-based data storage system, or the storage system 1502 is a DAS storage system.

[0356] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation and description of the beneficial effects of any of the above data recovery systems 1500 can refer to the above corresponding method embodiments, which will not be repeated here.

[0357] It should be noted that the hardware structure of the user terminal device 1501 in the data recovery system 1500 can refer to the description of the hardware structure of the computing device shown in Figure 14. The hardware structure of the device used to implement the storage system 1502 in the data recovery system 1500 can also refer to the description of the hardware structure of the computing device shown in Figure 14, and will not be repeated here.

[0358] An embodiment of the present application also provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium. The computer-readable storage medium includes computer program instructions. When the computer program instructions are executed by a computing device, a computer system or a processor, the computing device, the computer system or the processor executes the data recovery method provided in the embodiment of the present application.

[0359] An embodiment of the present application also provides a computer program product containing instructions. When the instructions are executed by a computing device, a computer system or a processor, the data recovery method provided by the embodiment of the present application is implemented on the computing device, the computer system or the processor.

[0360] A computer system is a system with computing processing capabilities. A computer system generally includes a processor and memory. The processor is configured to retrieve and execute instructions stored in the memory, thereby enabling the computer system to implement the data recovery method described above. Optionally, the computer system may also include at least one of an input interface and an output interface. The processor, memory, input interface, and output interface of the computer system are connected via internal connection paths.

[0361] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.

[0362] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0363] The embodiment of the present application also provides a chip, which includes a processor for running program instructions or codes. The chip or a device including the chip can be used to perform the data recovery method provided in the embodiment of the present application. Exemplarily, the chip also includes: an input interface, an output interface, and a memory. Among them, the input interface, output interface, processor, and memory of the chip are connected through the internal connection path of the chip, the memory in the chip is used to store the program instructions or codes run by the processor, and the input interface and output interface of the chip are used for connection and communication between the chip and other chips or devices.

[0364] In the embodiments of the present application, the terms "first", "second" and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "plurality" refers to a plurality, unless otherwise expressly limited.

[0365] In this application, the term "and / or" simply describes an association between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0366] It should be understood that the terminology used in the description of the various examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0367] It should be understood that determining B based on A does not mean determining B based solely on A. B can also be determined based on A and / or other information.

[0368] It will be understood that the term “comprise” (also known as “includes,” “including,” “comprises,” and / or “comprising”) when used in this specification specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0369] It should also be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0370] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the concepts and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A data recovery method, characterized in that: Applied to a storage system, the storage system adopts a block-level storage method for data storage; the method includes: Determine a target object for which data needs to be restored, where the target object is any semantic-level data, and data blocks of the target object are stored in the storage system in a block-level manner; determining, based on a first mapping relationship set associated with a first snapshot, a first storage address of a data block storing the target object in the first snapshot; the first snapshot is a snapshot obtained after a snapshot operation is performed on data in the storage system at a first time, and the first mapping relationship set associated with the first snapshot includes a correspondence between objects maintained by the storage system and storage addresses of the data blocks of the objects as of the first time; The data of the target object is restored according to the data stored at the first storage address in the first snapshot.

2. The method according to claim 1, wherein The method further comprises: Receive an input / output (IO) request, where the IO request is used to instruct a write operation to be performed on data to be processed, and the IO request includes an identifier of an object to which the data to be processed belongs, where the object to which the data to be processed belongs is any type of semantic-level data; The first mapping relationship set is updated according to the identifier of the object to which the data to be processed belongs and the storage address of the data to be processed.

3. The method according to claim 1 or 2, wherein: Determining the target object for which data needs to be restored includes: receiving an IO request carrying a first identifier and first data, where the first data is data of a first object represented by the first identifier; detecting whether the first data is infected with a virus or subjected to a malicious attack; In the case where it is determined that the first data is infected with a virus or is attacked by a malicious attack, the first object is determined as the target object.

4. The method according to claim 1 or 2, wherein: Determining the target object for which data needs to be restored includes: Detecting data blocks stored in the storage system to determine target data blocks in the storage system that are infected with a virus or are under malicious attack; querying the first mapping relationship set according to the storage address of the target data block to determine a first object corresponding to the storage address of the target data block; The first object is determined as the target object.

5. The method according to claim 3 or 4, wherein: Before determining the first object as the target object, the method further includes: Sending a first recovery request to a user terminal device, where the first recovery request includes an identifier of the first object, and the first recovery request is used to confirm whether to recover data of the first object; A confirmation message returned by the user terminal device is received, where the confirmation message is used to instruct to restore the data of the first object.

6. The method according to claim 1 or 2, wherein: Determining the target object for which data needs to be restored includes: When a second recovery request sent by the user terminal device is received, the first object is determined as the target object; the second recovery request includes an identifier of the first object, and the second recovery request is used to request recovery of data of the first object.

7. The method according to any one of claims 1 to 6, characterized in that Restoring the data of the target object according to the data stored at the first storage address in the first snapshot includes: detecting whether the data stored at the first storage address in the first snapshot is infected by a virus or attacked by a malicious agent; When it is determined that the data stored at the first storage address in the first snapshot is not infected by a virus or attacked by a malicious program, the data stored at the first storage address in the first snapshot is determined as the restored data of the target object.

8. The method according to claim 7, wherein The method further comprises: The first mapping relationship set currently maintained by the storage system is updated according to the identifier of the target object and the first storage address.

9. The method according to claim 7 or 8, wherein The method further comprises: If it is determined that the data stored at the first storage address in the first snapshot is infected by a virus or is maliciously attacked, determining, based on the first mapping relationship set associated with the second snapshot, a second storage address for storing the data block of the target object in the second snapshot; the second snapshot is a snapshot obtained by performing a snapshot operation on the storage system at a second time before the first time, and the first mapping relationship set associated with the second snapshot includes a correspondence between objects maintained by the storage system and storage addresses of the data blocks of the objects as of the second time; The data of the target object is restored according to the data stored at the second storage address in the second snapshot.

10. The method according to any one of claims 1 to 9, characterized in that The target object is a file or a database / table.

11. A data recovery device, characterized in that: Applicable to a storage system, the storage system adopts block-level storage for data storage; the device comprises: a determining unit configured to determine a target object for which data recovery is required, the target object being any type of semantic-level data, data blocks of the target object being stored in a block-level manner on the storage system, and determining, based on a first mapping relationship set associated with a first snapshot, a first storage address of the data blocks storing the target object in the first snapshot; the first snapshot being a snapshot obtained after a snapshot operation is performed on data in the storage system at a first time, the first mapping relationship set associated with the first snapshot including a correspondence between objects maintained by the storage system as of the first time and storage addresses of the data blocks of the objects; A recovery unit is configured to recover the data of the target object based on the data stored at the first storage address in the first snapshot.

12. The device according to claim 11, wherein The device further comprises: A receiving unit, configured to receive an input / output (IO) request, wherein the IO request is used to instruct a write operation to be performed on data to be processed, and the IO request includes an identifier of an object to which the data to be processed belongs, and the object to which the data to be processed belongs is any type of semantic-level data; An updating unit is configured to update the first mapping relationship set according to an identifier of the object to which the data to be processed belongs and a storage address of the data to be processed.

13. The device according to claim 11 or 12, characterized in that The device further comprises: a receiving unit, configured to receive an IO request carrying a first identifier and first data, where the first data is data of a first object represented by the first identifier; a detection unit, configured to detect whether the first data is infected with a virus or subjected to a malicious attack; The determining unit is specifically configured to determine the first object as the target object when it is determined that the first data is infected with a virus or is attacked by a malicious attack.

14. The device according to claim 11 or 12, characterized in that The device further comprises: A detection unit, configured to detect data blocks stored in the storage system to determine target data blocks in the storage system that are infected with a virus or are under malicious attack; The determining unit is specifically configured to query the first mapping relationship set according to the storage address of the target data block to determine a first object corresponding to the storage address of the target data block, and determine the first object as the target object.

15. The device according to claim 13 or 14, characterized in that The device further comprises: a sending unit, configured to send a first recovery request to a user terminal device before determining the first object as the target object, wherein the first recovery request includes an identifier of the first object and is used to confirm whether to recover data of the first object; A receiving unit is used to receive a confirmation message returned by the user terminal device, where the confirmation message is used to instruct to restore the data of the first object.

16. The device according to claim 11 or 12, characterized in that The determining unit is specifically configured to determine the first object as the target object upon receiving a second recovery request sent by a user terminal device; the second recovery request includes an identifier of the first object, and the second recovery request is used to request recovery of data of the first object.

17. The device according to any one of claims 11 to 16, characterized in that The recovery unit is specifically used to: detecting whether the data stored at the first storage address in the first snapshot is infected by a virus or attacked by a malicious agent; When it is determined that the data stored at the first storage address in the first snapshot is not infected by a virus or attacked by a malicious program, the data stored at the first storage address in the first snapshot is determined as the restored data of the target object.

18. The device according to claim 17, wherein The device further comprises: An updating unit is configured to update the first mapping relationship set currently maintained by the storage system according to the identifier of the target object and the first storage address.

19. The device according to claim 17 or 18, characterized in that The determining unit is further configured to, if it is determined that the data stored at the first storage address in the first snapshot is infected by a virus or is maliciously attacked, determine, based on the first mapping relationship set associated with the second snapshot, a second storage address for storing the data block of the target object in the second snapshot; the second snapshot is a snapshot obtained by performing a snapshot operation on the storage system at a second time before the first time, and the first mapping relationship set associated with the second snapshot includes a correspondence between objects maintained by the storage system and storage addresses of the data blocks of the objects as of the second time; The restoration unit is further configured to restore the data of the target object based on the data stored at the second storage address in the second snapshot.

20. The device according to any one of claims 11 to 19, characterized in that The target object is a file or a database / table.

21. A data recovery method, characterized in that: include: The user terminal device determines a target object for which data recovery is required, wherein the target object is any semantic-level data, and data blocks of the target object are stored in a storage system in a block-level manner. The storage system stores data in a block-level storage manner; The user terminal device queries a second mapping relationship set according to the object identifier of the target object to determine a target storage address, where the target storage address is a storage address of a data block of the target object, and the second mapping relationship set includes a correspondence between the object maintained by the user terminal device and the storage address of the data block of the object in the storage system; The user terminal device sends a data recovery instruction to the storage system, where the data recovery instruction includes the target storage address, and the data recovery instruction is used to instruct to recover the data of the target object according to the target storage address.

22. The method according to claim 21, wherein The method further comprises: The storage system receives the data recovery instruction sent by the user terminal device; The storage system restores the data of the target object based on the data stored at the target storage address in the first snapshot, where the first snapshot is a snapshot obtained after a snapshot operation is performed on the data of the storage system at a first time.

23. The method according to claim 21 or 22, wherein: The method further comprises: The storage system detects whether the data stored in the storage system is infected by a virus or attacked by a malicious agent; The storage system sends a storage address of an abnormal data block to the user terminal device, where the abnormal data block is a data block where data infected by a virus or maliciously attacked is located; The user terminal device determines a target object for which data needs to be restored, including: The user terminal device receives the storage address of the abnormal data block sent by the storage system; The user terminal device queries the second mapping relationship set according to the storage address of the abnormal data block to determine the abnormal object; The user terminal device determines the target object among the abnormal objects.

24. The method according to claim 21 or 22, wherein: The user terminal device determines a target object for which data needs to be restored, including: The user terminal device determines the target object in response to the received object recovery indication, where the object recovery indication includes the object identifier of the target object.

25. The method of claim 22, wherein: The storage system restores the data of the target object according to the data stored at the target storage address in the first snapshot, including: The storage system detects whether the data stored at the target storage address in the first snapshot is infected by a virus or attacked by a malicious agent; When determining that the data stored in the first snapshot at the target storage address is not infected by a virus or attacked by malicious software, the storage system determines the data stored in the first snapshot at the target storage address as the restored data of the target object.

26. The method of claim 25, wherein: The method further comprises: When the storage system determines that the data stored at the target storage address in the first snapshot is infected by a virus or is maliciously attacked, the storage system restores the data of the target object based on the data stored at the target storage address in a second snapshot, where the second snapshot is a snapshot obtained by performing a snapshot operation on the storage system at a second time before the first time.

27. The method according to any one of claims 21 to 26, characterized in that The target object is a file or a database / table.

28. A computing device, characterized in that include: A memory, a network interface, and one or more processors, wherein the one or more processors receive or send data through the network interface, and the one or more processors are configured to read program instructions stored in the memory to execute the method according to any one of claims 1 to 10.

29. A data recovery system, characterized in that: It includes a user-end device and a storage system, the user-end device is used to execute the method executed by the user-end device as claimed in any one of claims 21 to 27, the storage system uses block-level storage to store data, and is used to execute the method executed by the storage system as claimed in any one of claims 21 to 27.

30. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device or a processor, the computing device or the processor is caused to perform the method according to any one of claims 1 to 10 or any one of claims 21 to 27.

31. A computer-readable storage medium, characterized in that The method comprises computer program instructions which, when executed by a computing device or a processor, cause the computing device or the processor to perform the method according to any one of claims 1 to 10 or any one of claims 21 to 27.

Citation Information

Patent Citations

  • Data recovery method and device and computing equipment

    CN120429163A

  • Data recovery method and storage device

    CN108351821A

  • Data backup method, device and system

    CN111078464A

  • Data processing method and device

    CN113297007A

  • Data recovery method and related apparatus

    WO2023207280A1