Data recovery method, device and equipment
By analyzing the valid data from multiple snapshots using a fusion device, a fused snapshot or candidate data set is generated, which solves the problem of losing the latest data in the prior art, improves the flexibility and efficiency of data recovery, and ensures that the data set contains the latest uninfected data.
Patent Information
- Application Number
- CN202411118364.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies, when recovering data, often use earlier snapshots to restore the dataset, which can lead to the loss of the most recently updated data and fail to effectively preserve the latest, virus-free data.
The system acquires valid data from multiple snapshots using a fusion device. Based on audit information and availability verification, it determines the latest versions of related and unrelated data, generates a fusion snapshot or candidate data set, and restores the data set to retain the latest valid data.
It enables the preservation of the latest uninfected data during the data recovery process, ensuring that the recovered data set supports the operation of applications and improving the flexibility and efficiency of data recovery.
Smart Images

Figure CN121597481A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a data recovery method, apparatus, and device. Background Technology
[0002] A snapshot is a data protection method. A snapshot is a static image created for a collection of data. A snapshot records the state of the collection of data, not a complete copy of the collection of data. Therefore, snapshots occupy less storage space. Due to this characteristic, snapshots are widely used.
[0003] By periodically creating snapshots of the dataset, data can be recovered in case of loss or corruption. However, if the most recently created snapshot contains encrypted or virus-infected data, it will not be used for recovery. Instead, an earlier snapshot will be used, resulting in the loss of recently updated data in the recovered dataset. Summary of the Invention
[0004] This application provides a data recovery method, apparatus, and device to ensure that the recovered data retains a significant amount of the most recently updated data.
[0005] Firstly, embodiments of this application also provide a data recovery method, executed by a fusion device. In this method: the fusion device acquires multiple snapshots to be fused, wherein each snapshot is created for a specific data set, and the creation times of different snapshots are different. After acquiring the multiple snapshots, the fusion device acquires valid data from each snapshot, where the valid data in each snapshot represents data from the data set recorded in that snapshot that has not been infected by a virus. Based on the valid data from the multiple snapshots, the fusion device performs data recovery on the data set.
[0006] Using the method described above, the fusion device can extract valid data from multiple snapshots to restore the dataset. This valid data includes uninfected and updated data. By restoring the dataset in this way, the restored dataset will include the most recently updated and uninfected data, ensuring that the restored dataset not only contains valid data but also retains as much of the most recent data as possible.
[0007] In one possible implementation, when the fusion device restores a dataset based on valid data from multiple snapshots, it can directly fuse the valid data from the multiple snapshots to obtain a candidate data set, and then use the candidate data set to restore the dataset. Alternatively, it can first obtain a fused snapshot based on the valid data from multiple snapshots; when data restoration is needed, it can acquire this fused snapshot and use it to restore the dataset.
[0008] The above method provides a flexible way to restore a dataset based on valid data from multiple snapshots, making it suitable for different application scenarios. The existence of fused snapshots facilitates subsequent viewing of the dataset at historical moments based on the fused snapshots.
[0009] In one possible implementation, the dataset includes at least one set of related data, and each set of related data includes multiple related data.
[0010] When obtaining a fused snapshot based on valid data from multiple snapshots, for any set of related data, such as the first set of related data, the fusion device can determine at least one first candidate snapshot. The least one first candidate snapshot is the snapshot from which the first set of related data is valid. The fusion device determines a first snapshot from these at least one first candidate snapshots; the first snapshot is the snapshot whose creation time is closest to the current time among the at least one first candidate snapshots. The fusion device determines the latest version of the first set of related data from the first block; the latest version of the first set of related data is the first set of related data recorded in the first snapshot. The fusion device can obtain a fused snapshot based on the latest version of the first set of related data, and can also restore the dataset based on the latest version of the first set of related data.
[0011] By using the above method, when restoring the dataset, the latest version of at least one set of related data is used, ensuring that the restored dataset retains the latest and valid related data.
[0012] In one possible implementation, when determining the latest version of the first set of associated data, the fusion device may also determine at least one first candidate snapshot among multiple snapshots that meets a first availability condition. The first availability condition is that the data set recorded in the first candidate snapshot supports the operation of the first application. The fusion device determines the latest version of the first set of associated data from at least one first candidate snapshot. That is, it determines the first snapshot from at least one first candidate snapshot and obtains the latest version of the first set of associated data from the first snapshot.
[0013] Using the above method, the latest version of the first set of associated data is not only valid data, but also usable data, which means it supports the operation of the first application and ensures that the restored data set can also support the operation of the first application.
[0014] In one possible implementation, the dataset includes non-associative data, which is data in the dataset other than at least one set of associated data.
[0015] When the fusion device obtains a fused snapshot based on valid data from multiple snapshots, for any non-related data in the non-related data, such as the first non-related data, the fusion device determines at least one second candidate snapshot. This second candidate snapshot is a snapshot from the multiple snapshots where the first non-related data is valid. The fusion device then determines a second snapshot from the at least one second candidate snapshot, which is the snapshot whose creation time is closest to the current time. The fusion device obtains the latest version of the first non-related data from the second snapshot. The latest version of the first non-related data is the first non-related data recorded in the second snapshot. The fusion device can obtain the fused snapshot based on the latest version of the first non-related data and can also restore the data set based on the latest version of the first non-related data.
[0016] By using the above method, when restoring the dataset, the latest version of the non-related data is used, ensuring that the restored dataset retains the latest and valid non-related data.
[0017] In one possible implementation, when determining the latest version of the first unrelated data, the fusion device determines at least one second candidate snapshot among multiple snapshots that meets a second availability condition. The second availability condition is that the data set recorded in at least one second candidate snapshot supports the operation of a second application. The fusion device determines the latest version of the first unrelated data from the at least one second candidate snapshot, that is, determines a second snapshot from the at least one second candidate snapshot, and then obtains the latest version of the first unrelated data from the second snapshot.
[0018] Using the above method, the latest version of the first non-related data is not only valid data, but also usable data, which means it supports the operation of the second application and ensures that the restored data set can also support the operation of the second application.
[0019] In one possible implementation, the fusion device acquires audit information that describes operational information on data in the dataset; based on the audit information, the fusion device determines at least one set of related data.
[0020] Using the above method, the fusion device can analyze the changing patterns of various data from the audit information, and thus accurately obtain at least one set of related data.
[0021] In one possible implementation, the fusion device identifies at least one set of associated data based on audit information and a list of associated data. The list of associated data records one or more sets of associated data.
[0022] Using the above method, the fusion device can accurately determine all the related data that may be included in the dataset.
[0023] In one possible implementation, the fusion device, based on valid data from multiple snapshots, detects the data set recorded in any one of the multiple snapshots before obtaining the fused snapshot, to determine the valid data in the snapshot.
[0024] Using the above method, the fusion device can perform virus detection on the data set recorded in the snapshot to extract valid data. The data set that is recovered contains valid data.
[0025] In one possible implementation, when the fusion device restores the data set based on valid data from multiple snapshots, it obtains the latest version of non-related data and the latest version of at least one set of related data from the valid data in the multiple snapshots to construct a candidate data set; and uses the candidate data set to restore the data set.
[0026] Using the above method, the candidate dataset contains the latest version of non-related data and the latest version of at least one set of related data. The dataset recovered using this candidate dataset also contains the latest non-related data and related data, thus avoiding the loss of a large amount of recently updated data during data recovery.
[0027] In one possible implementation, the fusion device creates a snapshot of the candidate data set to obtain the fusion snapshot when obtaining the fusion snapshot based on valid data from multiple snapshots.
[0028] By using the above method, the existence of the fused snapshot allows for timely recovery of the data set when it is needed in the future, improving the efficiency of data recovery. It also allows for convenient querying of the status of the data set at historical moments.
[0029] Secondly, embodiments of this application also provide a fusion device that has the function of implementing the behavior in the method example of the first aspect described above. The beneficial effects can be found in the description of the first aspect and will not be repeated here. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the fusion device includes an acquisition module, a recovery module, and optionally, a detection module. These modules can perform the corresponding functions in the method example of the first aspect described above; see the detailed description in the method example for details, which will not be repeated here.
[0030] Thirdly, this application also provides a computing device, which includes a processor and a memory, and may further include a communication interface. The processor executes computer program instructions stored in the memory to perform the method provided in the first aspect or any possible implementation thereof. The memory is coupled to the processor and stores the computer program instructions and data necessary for determining the data deduplication process. The communication interface is used for communicating with other devices, such as transmitting multiple snapshots, restored data sets, fused snapshots, etc.
[0031] Fourthly, this application also provides a computing device, which includes an external device and a processor. Optionally, it may also include a memory and a communication interface. The processor cooperates with the external device to execute the method provided in the first aspect or any possible implementation of the first aspect. Alternatively, the external device executes the method provided in the first aspect or any possible implementation of the first aspect. The memory is coupled to the processor and stores computer program instructions and data required during the data recovery process. The communication interface is used for communication with other devices.
[0032] Fifthly, this application provides a computing device system including at least one computing device. Each computing device includes a memory and a processor. The processor of the at least one computing device is used to access computer program instructions in the memory to execute the methods provided in the first aspect or any possible implementation thereof.
[0033] Sixthly, this application provides a computer-readable storage medium that, when executed by a computing device, allows the computing device to perform the methods provided in the first aspect or any possible implementation thereof. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, hard disk drive (HDD), and solid-state drive (SSD).
[0034] In a seventh aspect, this application provides a computer device program product, which includes computer program instructions. When executed by a computing device, the computing device performs the methods provided in the first aspect or any possible implementation thereof. The computer program product can be a software installation package, and when it is necessary to use the methods provided in the first aspect or any possible implementation thereof, the computer program product can be downloaded and executed on the computing device.
[0035] Eighthly, this application also provides a computer chip connected to a memory, the chip being used to read and execute computer program instructions stored in the memory, and to execute the methods described in the first aspect and various possible implementations of the first aspect.
[0036] For the technical effects that can be achieved in aspects two through eight above, please refer to the description of the technical effects that can be achieved in the corresponding design schemes in aspect one above. This application will not repeat them here. Attached Figure Description
[0038] Figure 1 A schematic diagram of the structure of a data recovery system provided in this application;
[0039] Figure 2 This is a schematic diagram of a data recovery method provided in an embodiment of this application;
[0040] Figure 3 A schematic diagram of a snapshot query interface provided in this application;
[0041] Figure 4 A schematic diagram of a data recovery interface provided in this application;
[0042] Figure 5 A schematic diagram of the structure of a server provided in this application;
[0043] Figure 6 A schematic diagram of the structure of a computing device provided in this application;
[0044] Figure 7 A schematic diagram of the structure of a computing device provided in this application;
[0045] Figure 8 This is a schematic diagram of a fusion device provided in this application. Detailed Implementation
[0046] Before introducing the data recovery method, apparatus, and device provided in the embodiments of this application, some concepts involved in the embodiments of this application will be explained first:
[0047] (1) File system.
[0048] A file system is a structured form of data storage and organization. It organizes data using the concept of "files," grouping data for the same purpose into different types of files according to the structural requirements of different applications. Different file extensions are typically used to indicate different file types, and each file is assigned an easy-to-remember name, or "filename." When there are many files, they are grouped according to a certain method, with each group of files placed in the same directory (or folder). Furthermore, directories can contain subdirectories (or subfolders) in addition to files, forming a tree-like structure. This tree-like structure has a specific name: file system. There are many types of file systems, common ones being Windows' FAT / FAT32 / NTFS, and Linux's EXT2 / EXT3 / EXT4 / XFS / BtrFS, etc. To facilitate searching, directories are listed level by level from the root node down to the file itself. The names of these directories, subdirectories, and files are concatenated using special characters (e.g., "\" in Windows / DOS, " / " in Unix-like systems). This string of characters is called a file path, such as " / etc / systemd / system.conf" in Linux or "C:\Windows\System32\taskmgr.exe" in Windows. A path is a unique identifier for accessing a specific file. For example, D:\data\file.exe in Windows is a file path, representing the file.exe file located in the data directory on the D partition.
[0049] The file system is built on block devices. The file system not only records the file path, but also records which blocks make up a file and which blocks record directory / subdirectory information.
[0050] (2) Object storage.
[0051] Object storage uses a flat address space to store data, without a hierarchical structure of directories and files. An object can include user data, related metadata (such as size, date, owner, etc.), and other data attributes (such as access characteristics). Each object has a unique identifier (ID), called the object ID. The object ID is generated using a specialized algorithm (such as a hash value of the data) to ensure that each object's ID is unique. In object storage, objects are stored in buckets, and new buckets can be created when the storage space in a bucket is insufficient.
[0052] Object storage primarily involves two types of management services: metadata service and storage service. The storage service manages multiple hard drives storing data and can be deployed on one or more nodes. The metadata service manages the data's metadata. The metadata service can be deployed on the same node as the storage service or on different nodes.
[0053] (3) Snapshot, Fusion Snapshot.
[0054] A snapshot is a fully usable copy of a specified dataset (dataset pages can be called data volumes). This copy is a static image of the source data at a specific point in time. A snapshot is generally a "virtual" copy of the source data. A snapshot only saves the state of the dataset at a particular point in time; based on the snapshot, a version of the dataset can be created, which is the dataset at that specific point in time. Snapshots are time-sensitive, and the dataset created based on a snapshot is also the dataset at that specific point in time. In other words, the dataset created based on a snapshot is consistent with the dataset at a past point in time; that point in time is the snapshot's time.
[0055] A snapshot records data as a "static image," and the data recorded in the snapshot is the data in the dataset formed based on the snapshot.
[0056] In this embodiment, the creation of a snapshot of a file system is used as an example for explanation. To easily distinguish between the file system targeted by the snapshot and the file system formed based on the snapshot, the file system targeted by the snapshot is called the original file system, and the file system formed based on the snapshot is called the cloned file system. Therefore, restoring the original file system using a snapshot means replacing files in the original file system with files in the cloned file system, or replacing the original file system with the cloned file system.
[0057] In this embodiment of the application, the concept of "fusion snapshot" is introduced. A fusion snapshot is a new snapshot formed based on multiple snapshots, and the fusion snapshot merges the valid data from multiple snapshots.
[0058] (4) Audit information.
[0059] In this embodiment, data within the storage device is accessed (e.g., by user-triggered operations such as adding, deleting, querying, and modifying data), and the storage device also spontaneously operates on the data stored therein. Audit information describes the data operation information within the storage device. This audit information includes some or all of the following:
[0060] Information on data access, the user to whom the data belongs, and the application to which the data belongs.
[0061] The data access information describes the operations performed on the data, such as adding, deleting, querying, and modifying. This data access information includes the time, type, and content of the operations performed on the data.
[0062] The data within the storage device may include some user-specific data, such as data created by the user through operating the computing devices deployed on the user side, or data generated by applications supporting the user's business during operation. For this type of data, the user to whom the data belongs can be recorded.
[0063] The storage device may also contain data generated by applications running on the device, or data transferred to the device by applications deployed remotely. For this type of data, the application to which the data belongs can be recorded.
[0064] The above is merely an example of the information content that the audit information may include. This application embodiment does not limit the information content included in the audit information. Any information that describes the operation information of data in the storage device can be used as audit information.
[0065] If the data in the storage device is organized as files, then the audit information for the storage device includes: file access information, the user to whom the file belongs, and the application to which the file belongs.
[0066] (5) Related data, unrelated data, related files, and unrelated files.
[0067] First, let's explain the concept of data versions. After data is created, it changes through operations such as adding (e.g., adding data to the data set), deleting (e.g., deleting data from the data set or deleting data altogether), and modifying (e.g., renaming the data, the file or object to which the data belongs, or moving the data's location). Each time data changes (e.g., each time an operation is performed on the data), the changed data can be referred to as a version of the data or a version of the data itself. The version of data changes with each operation performed on it.
[0068] To understand data versions in conjunction with snapshots, a snapshot essentially records one version of the data. If snapshots of the same data are created at different points in time, each snapshot represents a version of the data, and different snapshots represent different versions of that data. Snapshots can be used to label data versions.
[0069] Related data refers to multiple pieces of data that are interconnected. The "relationship" between these data points manifests as a dependency, where a change in one piece of data will cause a change in another. Because of this dependency, the versions of these related data points must remain consistent; otherwise, the data will be invalid or unable to support business operations.
[0070] It should be noted that the number of data items is determined by the way the data is organized in the storage device. For example, if the data in the storage device is organized as files, then one data item is one file. Or, if the data in the storage device is organized as objects, then one data item is one object.
[0071] Here are two examples of related data:
[0072] Example 1: In databases, there are often multiple related tables, where one table points to another. When one table changes, the other tables also need to be adjusted accordingly. In other words, the versions of these multiple tables need to be consistent; otherwise, errors may occur.
[0073] Example 2: In an application's configuration data, some configuration data are related; if one configuration data changes, another will also change. The versions of these multiple configuration data must remain consistent; otherwise, the application will fail to generate functionalities and support business operations.
[0074] For ease of explanation, multiple related data are referred to as a set of related data. This includes, for example, the first set of related data and the first set of related files mentioned in this application. It should be noted that the first set of related data represents any one of at least one set of related data; the term "first set" in the first set of related data does not indicate the order of the related data. Similarly, the first set of related files only represents any one set of related files. In the embodiments of this application, whether multiple data sets belong to related data can be determined based on some or all of the following information.
[0075] Information 1: List of Associated Data (e.g., list of associated files, list of associated object files). This list of associated data is constructed based on the experience of operations and maintenance personnel, and records at least one set of associated data.
[0076] Information 2: Audit Information. By analyzing audit information, it is determined whether there are correlations between multiple data points, thereby identifying related data.
[0077] In contrast to associated data, data that is not related to other data is considered unrelated data, such as the first unrelated data and the first unrelated file mentioned in this application. The first unrelated data represents any unrelated data. Unrelated data is relatively "independent," and changes to this unrelated data will not affect other data.
[0078] If the data in the storage device is organized as files, the files within the storage device include associated files and / or non-associated files. If the data in the storage device is organized as objects, the objects within the storage device include associated objects and / or non-associated objects.
[0079] like Figure 1 The diagram shown is an architectural schematic of a data recovery system provided in an embodiment of this application. The data system includes a storage device 110 and a fusion device 120.
[0080] Storage device 110 has storage functionality. Data on storage device 110 is organized in a certain structure. This application embodiment does not limit the way data is organized on storage device 110. For example, the data on storage device 110 may be organized in the form of files, and storage device 110 may have a file system deployed. Another example is that the data on storage device 110 may be organized in the form of objects, and storage device 110 may have an object system deployed.
[0081] Regardless of how the data in storage device 110 is organized, the data stored on storage device 110 can form a data set, and the file system and object system are specific forms of representation of this data set.
[0082] In addition to storing data sets, the storage device 110 can also create and store snapshots of the stored data sets. In this embodiment, the storage space allocated by the storage device 110 for storing snapshots is called a snapshot resource pool. The storage device 110 can pre-allocate a segment of storage space as a snapshot storage pool. Each time the storage device 110 creates a snapshot, it stores the snapshot in the snapshot resource pool. If the storage space in the snapshot resource pool is full or the unoccupied storage space (also called free storage space) in the snapshot resource pool is less than a set value, the storage device 110 can expand the snapshot resource pool, that is, increase the storage space in the snapshot resource pool. The storage device 110 can also delete some snapshots in the snapshot resource pool to release the storage space in the snapshot resource pool, such as deleting the earliest stored snapshot in the snapshot resource pool.
[0083] The storage device 110 can restore stored data based on snapshots within the snapshot resource pool. Taking data organized as files within the storage device 110 as an example, the storage device 110 can create a cloned file system based on a snapshot within the snapshot resource pool. A snapshot represents the state of the original file system at a specific point in time; this point in time is the snapshot's time. The cloned file system is the original file system at the snapshot's time, meaning it is identical to the original file system at that snapshot's time. The storage device 110 uses the cloned file system to restore the original file system.
[0084] When data is organized by objects within storage device 110, the method of data recovery based on data in storage device 110 is similar, the difference being the different organization of data within storage device 110.
[0085] This application does not limit the specific form of the storage device 110. The storage device 110 can be a computing device with storage function, which includes multiple hard disks. The storage device 110 can be a computing device cluster, with a file system or object system distributed across multiple computing devices in the cluster. The storage device 110 can also be a storage system, which can be a centralized storage system or a distributed storage system.
[0086] The fusion device 120 can acquire multiple snapshots from the storage device 110 and merge them into a single fused snapshot. The fusion device 120 can also store the fused snapshot in a snapshot resource pool. The fusion device 120 can analyze and process the multiple snapshots from the storage device 110, extract valid data from the snapshots, and use the valid data from the snapshots to restore the dataset. The fusion device 120 has some or all of the following functions:
[0087] Function 1: Audit Information Analysis.
[0088] The fusion device 120 can collect operation information of data within the storage device 110, thereby obtaining audit information of the storage device 110. The fusion device 120 can also obtain audit information recorded by the storage device 110 from the storage device 110.
[0089] After acquiring audit information, the fusion device 120 analyzes the audit information to extract at least one set of related data from the data in the storage device 110.
[0090] Since audit information accumulates and changes as data within storage device 110 is manipulated, its analysis can be performed cyclically. For example, fusion device 120 can analyze audit information in real-time, or it can analyze it periodically. By analyzing the audit information, related data within storage device 110 can be identified, preparing for subsequent snapshot fusion.
[0091] Function 2: Snapshot detection.
[0092] The fusion device 120 can detect whether there is virus data (i.e., data infected by a virus) in the snapshot. After performing snapshot detection, the fusion device 120 will mark the snapshot as valid data, which is data that has not been infected by a virus.
[0093] In this embodiment, the data infected by a virus can be encrypted data, meaning the virus program encrypts the data. It can also be garbled data, meaning the virus program edits the data, making it garbled and different from the original data. This embodiment does not limit the specific method of data infection; any method that makes it impossible to obtain the original data after data manipulation can be considered as virus infection. Snapshot detection is a pre-processing operation before snapshot fusion. Snapshot detection ensures the security of multiple snapshots participating in the fusion, thereby ensuring that the subsequently obtained fused snapshot is safe, i.e., free of encrypted and virus-infected data.
[0094] Function 3: Data Recovery.
[0095] After acquiring valid data from multiple snapshots, the fusion device 120 can use this valid data to restore the dataset. There are many ways for the fusion device 120 to restore the dataset using the valid data from the multiple snapshots; two are listed here:
[0096] Recovery Method 1: The fusion device 120 merges the valid data from multiple snapshots into a candidate data set, and uses the candidate data set to replace the data set.
[0097] For the valid data in the multiple snapshots, the analysis results of the audit information analysis (i.e., whether there is related data and the identified related data) and / or the list of related data are used to determine whether there is related data in the valid data of the multiple snapshots and to determine which data is related data.
[0098] For any set of associated data, taking the first set of associated data as an example, the fusion device 120 determines the latest version of the first set of associated data from the valid data of multiple snapshots. The latest version of the first set of associated data is the first set of associated data recorded in the first snapshot. The first snapshot is the first candidate snapshot whose creation time is closest to the current time among the first candidate snapshots. The first candidate snapshot is the snapshot in which the first set of associated data is valid data among multiple snapshots.
[0099] Furthermore, the first candidate snapshot determined by the fusion device 120 can also satisfy a first availability condition: the data set recorded by the first candidate snapshot supports the operation of the first application. That is, the first candidate snapshot passes availability verification.
[0100] For any non-associative data, taking the first non-associative data as an example, the fusion device 120 determines the latest version of the first non-associative data from the valid data. The latest version of the first non-associative data is the second non-associative data recorded in the second snapshot. The second snapshot is the second candidate snapshot with the fastest creation time closest to the current time among the second candidate snapshots. The second candidate snapshot is a snapshot in which the first non-associative data is valid data among multiple snapshots.
[0101] Furthermore, the second candidate snapshot determined by the fusion device 120 can also satisfy a second availability condition, which is the first availability condition: the data set recorded by the second candidate snapshot supports the operation of the second application. That is, the second candidate snapshot passes availability verification.
[0102] The fusion device 120 acquires the latest versions of each set of related data and the latest versions of the unrelated data, merges the latest versions of the related data and the unrelated data into a candidate data set, and uses this candidate data set to restore the data set. When using the candidate data set to restore the data set, the fusion device 120 can replace the original data set with the candidate data set. Alternatively, the fusion device 120 can restore only the target data that needs to be restored from the data set; that is, the fusion device 120 only restores the target data and does not process any other data in the data set. In this case, the fusion device 120 can extract the latest version of the target data from the valid data in the multiple snapshots and replace the target data in the data set with the latest version of the target data.
[0103] Recovery Method 2: The fusion device 120 generates a fusion snapshot based on valid data from multiple snapshots, and recovers the data set based on the fusion snapshot.
[0104] When the fusion device 120 fuses multiple snapshots to obtain a fused snapshot, it extracts the valid data from each of the multiple snapshots. The method by which the fusion device 120 obtains the latest versions of each group of related data and the latest versions of the unrelated data from the valid data in the multiple snapshots is described above and will not be repeated here.
[0105] After acquiring the latest versions of each group of associated data and the latest versions of the unassociated data, the fusion device 120 generates a fusion snapshot based on these latest versions. When generating the fusion snapshot, the latest versions of each group of associated data and the latest versions of the unassociated data can be merged into a candidate data set. Creating a snapshot of this candidate data set is then called the fusion snapshot. When the fusion device 120 determines that a data set needs to be restored, it can restore the data set based on the fusion snapshot. The fusion device 120 generates a candidate data set based on the fusion snapshot and uses this candidate data set to restore the data set. The method of restoring the data set using the candidate data set is described above and will not be repeated here.
[0106] It is worth noting that, as can be seen from the descriptions of Recovery Method 1 and Recovery Method 2, Recovery Method 2 adds the operation of generating a fused snapshot on top of Recovery Method 1. The existence of the fused snapshot provides the fusion device 120 with an opportunity to perform data recovery. For example, when data recovery is not required, the fusion device 120 generates a fused snapshot and stores it. If data recovery is needed subsequently, the fusion device 120 can retrieve the fused snapshot and perform data recovery based on it. That is, Recovery Method 2 can improve the flexibility and efficiency of data recovery to a certain extent. In addition, the existence of the fused snapshot can also serve as a kind of "record" of the data set, making it easy to query the state of the data set at a certain time period or moment in the past through the fused snapshot. Moreover, the fused snapshot can be used as a snapshot to continue merging with subsequently created snapshots.
[0107] During the fusion process of multiple snapshots, the fusion device 120 performs different operations on different data. This ensures that the fused snapshot records the latest versions of each group of related data, guaranteeing that the correlation between the groups of related data is not broken and that the latest versions are valid. At the same time, it also ensures that the fused snapshot records the latest versions of non-related data. Overall, this guarantees the validity of the fused snapshot.
[0108] It should be noted that, regarding the "data recovery" function, in practical applications, the fusion device 120 may not possess a complete data recovery function, but only a portion of it. For example, the fusion device 120 may only have a snapshot fusion function, meaning it can generate a fused snapshot based on valid data from multiple snapshots. After generating the fused snapshot, it can perform data set recovery based on the fused snapshot, or it can store the fused snapshot in the snapshot resource pool of the storage device 110. The storage device 110 can then retrieve the fused snapshot from the snapshot resource pool and recover the data set based on the fused snapshot.
[0109] Function 4: Snapshot Availability Verification.
[0110] Snapshot availability verification is mainly for verifying merged snapshots or multiple snapshots to be merged. Availability verification is used to verify whether the data formed by the snapshot (such as merged snapshots or snapshots to be merged) can support business or whether it can support the operation of the application.
[0111] After obtaining the fusion snapshot, the fusion device 120 constructs a data application environment, which refers to an application environment that needs to use the data restored based on the fusion snapshot. This data application environment can be understood as an application that uses the data restored based on the fusion snapshot (such as the first application and the second application mentioned in the embodiments of this application).
[0112] This application does not limit the specific form of the fusion device 120. The fusion device 120 can be a hardware device, such as a computing device, a computing device cluster, or a chip or processor in a computing device. Alternatively, the fusion device 120 can be an external device of a computing device (such as an accelerator card or offloading card), capable of performing snapshot fusion on snapshots in the connected computing devices. The fusion device 120 can also be a software device, deployed on one or more computing devices as containers or virtual machines, etc. For example, the computing instance where the fusion device 120 resides carries a snapshot fusion service, and the storage device 110 can request snapshot fusion services from the fusion device 120. The fusion device 120 can also be an application deployed on a computing device. For example, the fusion device 120 can be software used to implement snapshot fusion, which can periodically fuse snapshots stored within the computing device where the software resides.
[0113] In this embodiment, the fusion device 120 acquires multiple snapshots to be fused, acquires valid data from the multiple snapshots, and restores the data set based on the valid data from the multiple snapshots. The valid data of each snapshot is the data recorded in the snapshot that has not been infected by the virus. The restored data set obtained by restoring the data set using the valid data from the multiple snapshots will contain more recently updated and valid data.
[0114] The following is in conjunction with the appendix Figure 2 This application provides a snapshot fusion method for illustrating an embodiment of the present application. Here, we take data organized as files in storage device 110 as an example. The snapshot fusion method when data is organized as objects or other data structures in storage device 110 is similar to the snapshot fusion method when data is organized as files in storage device 110, and will not be described again in this embodiment of the present application.
[0115] Step 200: The fusion device 120 analyzes the audit information to determine at least one set of associated files within the storage device 110.
[0116] This application does not limit the method of obtaining the audit information. For example, the fusion device 120 has monitoring permissions for the data in the storage device 110, and the fusion device 120 can record and store audit information during the process of data being operated on in the storage device 110. Alternatively, if audit software is deployed in the storage device 110, and this audit software monitors the data operations in the storage device 110, records and stores the audit information, the fusion device 120 has access permissions to this audit information and can obtain it.
[0117] After acquiring the audit information, the fusion device 120 analyzes the audit information. The analysis of the audit information by the fusion device 120 involves analyzing the correlation between files within the storage device 110, thereby identifying at least one set of related files within the storage device.
[0118] For example, the fusion device 120 may focus on situations in the audit information where operations on one file cause operations on other files. That is, when an operation such as adding, deleting, or modifying a file is performed, other files also need to be adjusted accordingly, indicating a relationship between that file and other files. The fusion device 120 may also focus on situations in the audit information where an error in one file causes operations on other files to fail. That is, if data in one file is corrupted or erroneous, operations such as adding, deleting, or modifying other files will fail, indicating a relationship between that file and other files. Furthermore, the fusion device 120 may also focus on file changes in the audit information, such as changes in the file system's location (e.g., moving a file from one directory to another) or changes in the file name. By monitoring these changes, the fusion device 120 can identify changes to other files within related files, facilitating the subsequent location of files within related files across different cloned file systems, especially for files that have undergone changes.
[0119] In reality, the process by which the fusion device 120 analyzes audit information is more complex. The above are just some examples of the analytical operations that may occur during the analysis of audit information.
[0120] Furthermore, during the analysis of audit information, the fusion device 120 first performs preprocessing such as filtering and aggregation. For example, the fusion device 120 can delete invalid information in the audit information, such as information about file revocation operations, or some duplicate information in the audit information. Another example is that the fusion device 120 can cluster the audit information at the file level, aggregating information related to the same file together.
[0121] The audit information describes file operation information within storage device 110. As file operations occur within storage device 110, the audit information changes dynamically, and the fusion device 120 can perform cyclical analysis of the audit information. The fusion device 120 can analyze the audit information periodically, or it can analyze the audit information when new information is detected or when the amount of added information reaches a set threshold.
[0122] Step 201: The fusion device 120 receives a snapshot fusion instruction, which instructs to fuse multiple snapshots.
[0123] There are many sources for snapshot fusion commands. The following are some examples of situations in which the fusion device 120 obtains snapshot fusion commands:
[0124] Scenario 1: The user triggers the snapshot fusion command.
[0125] Example 1: For users, the fusion device 120 can provide a snapshot fusion interface through which users send snapshot fusion commands to the fusion device 120.
[0126] The snapshot fusion interface refers to the function provided by the fusion device 120 to the user, that is, the fusion device 120 has the ability to provide snapshot fusion to the user. This application embodiment does not limit the specific form of the snapshot fusion interface. The snapshot fusion interface can be embodied in a preset command format; commands conforming to this command format are the snapshot fusion commands. Users can send the snapshot fusion commands to the fusion device 120 through computing devices deployed on the user side. The snapshot fusion interface can also be embodied in a visual user interface, where users can trigger the snapshot fusion commands through operations.
[0127] like Figure 3 The snapshot fusion interface described herein is an embodiment of the present application. In this interface, users can view created snapshots and the time point of each snapshot. The snapshot fusion interface provides two snapshot fusion options: a first, default snapshot fusion option, which indicates that multiple recently created snapshots with a number equal to the default value will be fused; and a second, user-defined snapshot fusion option, which indicates that multiple snapshots selected by the user will be fused.
[0128] When a user clicks the default snapshot fusion option, a snapshot fusion command is triggered. This command instructs that multiple snapshots be merged, which are the most recently created snapshots with the same number as the default value.
[0129] When a user clicks the user-defined snapshot fusion option, a snapshot fusion command is triggered, which instructs to merge multiple snapshots, namely the multiple snapshots selected by the user.
[0130] Example 2: For users, the fusion device 120 can provide a data recovery interface through which users send data recovery instructions to the fusion device 120. These data recovery instructions are used to instruct multiple snapshots to restore the file system.
[0131] The fusion device 120 splits the data recovery instruction into two instructions: a snapshot fusion instruction, which instructs the fusion of multiple snapshots, which can be multiple recently created snapshots with a number equal to a set value; and a snapshot recovery instruction, which instructs the recovery of data based on the fused snapshots.
[0132] Similar to the snapshot fusion interface, the embodiments of this application do not limit the specific form of the data recovery interface. The specific form of the data recovery interface is similar to that of the snapshot fusion interface, and can be found in the foregoing description, which will not be repeated here.
[0133] like Figure 4 The snapshot viewing interface described herein allows users to view created snapshots and the time point of each snapshot. The snapshot viewing interface provides data recovery options. Furthermore, the data recovery options also allow setting the number of snapshots, that is, allowing configuration of specific values for the set values.
[0134] When a user clicks the data recovery option, a data recovery command is triggered. This command instructs multiple snapshots to recover the data, and the number of snapshots is equal to the set value.
[0135] Scenario 2: Storage device 110 sends a snapshot fusion command to fusion device 120.
[0136] Storage device 110 can periodically create snapshots of the deployed file system, or it can create snapshots of the deployed file system when triggered by a user.
[0137] Storage device 110 can organize the snapshots in the snapshot resource pool, such as deleting multiple snapshots and storing a fused snapshot formed based on these multiple snapshots. In this case, storage device 110 can send a snapshot fusion command to fusion device 120. The snapshot fusion command may carry the multiple snapshots or their identifiers. The snapshot identifier is information configured by storage device 110 to uniquely identify the snapshot, and the snapshot identifier is recognizable by both storage device 110 and fusion device 120.
[0138] Storage device 110 can also restore the file system (i.e., the original file system) upon user triggering. For example, upon receiving a user-triggered instruction to restore the file system, storage device 110 can send a snapshot fusion instruction to fusion device 120, instructing fusion device 120 to merge multiple snapshots. The snapshot fusion instruction may carry the multiple snapshots or their identifiers. Storage device 110 obtains the fused snapshot generated by fusion device 120 by sending the snapshot fusion instruction and restores the file system based on the fused snapshot.
[0139] Storage device 110 can also determine on its own that file system recovery is needed. For example, during the operation of storage device 110, file access operations consistently fail, such as when the failure frequency exceeds a frequency threshold or the number of failures exceeds a number threshold. When storage device 110 determines that file system recovery is needed, it can send a snapshot fusion command to fusion device 120, instructing fusion device 120 to merge multiple snapshots to obtain a fused snapshot generated by fusion device 120, and then restore the file system based on the fused snapshot.
[0140] Case 3: Storage device 110 sends a data recovery command to fusion device 120.
[0141] When storage device 110 determines that file system recovery is necessary, it may send a data recovery command to fusion device 120. This data recovery command instructs the storage device 110 to recover the file system. For a scenario where storage device 110 determines that file system recovery is necessary, please refer to the relevant description in Case 2; it will not be repeated here.
[0142] The data recovery instruction of the fusion device 120 is split into two instructions: a snapshot fusion instruction, which instructs the fusion of multiple snapshots, which can be multiple recently created snapshots with a number equal to a set value; and a snapshot recovery instruction, which instructs the recovery of data based on the fused snapshots.
[0143] Scenario 4: The fusion device 120 generates a snapshot fusion command on its own.
[0144] The fusion device 120 can actively fuse snapshots in the snapshot resource pool. For example, the fusion device 120 can periodically access the snapshot resource pool, and when it determines that the number of newly created snapshots in the snapshot resource pool has reached a set value, the fusion device 120 automatically generates a snapshot fusion command to fuse multiple snapshots, which are the newly created snapshots. In other words, the fusion device 120 can fuse multiple snapshots that are close in time and have a number equal to the set value.
[0145] The above are just examples of several scenarios in which the fusion device 120 receives a snapshot fusion instruction. If the snapshot fusion instruction received by the fusion device 120 does not carry multiple snapshots to be fused (such as only carrying the identifiers of multiple snapshots), the fusion device 120 can execute step 202. If the snapshot fusion instruction received by the fusion device 120 carries multiple snapshots to be fused, the fusion device 120 can directly execute step 203.
[0146] Step 202: The fusion device 120 accesses the snapshot resource pool and obtains multiple snapshots to be fused from the snapshot resource pool.
[0147] The fusion device 120 has access to the snapshot resource pool. After receiving the snapshot fusion instruction, the fusion device 120 obtains multiple snapshots to be fused from the snapshot resource pool.
[0148] For example, storage device 110 and fusion device 120 are deployed in different computing devices. Storage device 110 sets the snapshot resource pool in memory, and fusion device 120 can access the snapshot resource pool based on RDMA. Alternatively, storage device 110 and fusion device 120 are deployed in the same computing device. The snapshot resource pool is set in the memory of that computing device, and fusion device 120 can retrieve multiple snapshots to be fused from memory based on DMA.
[0149] Step 203: The fusion device 120 detects the multiple snapshots to be fused and determines the valid files among the multiple snapshots.
[0150] After acquiring multiple snapshots to be merged, the fusion device 120 detects any snapshot to determine whether there are any virus-infected files in the snapshot. Files in the snapshot other than virus-infected files are considered valid files in the snapshot. Of course, the snapshot may also contain virus files, which are files that carry viruses. Valid files in the snapshot are files other than virus-infected files and virus files.
[0151] The process of the fusion device 120 detecting this snapshot is described below.
[0152] Step 1: Create a cloned file system based on the snapshot. This application embodiment does not limit the execution entity of Step 1. Step 1 can be executed by the storage device 110; that is, the fusion device 120 can also request the storage device 110 to first create a cloned file system based on the snapshot and obtain the cloned file system from the storage device 110. Step 1 can also be executed by the fusion device 120; that is, the fusion device 120 directly creates a cloned file system based on the snapshot.
[0153] Step 2: The fusion device 120 checks each file in the cloned file system to determine whether there are any files infected by the virus, and to determine whether there are any virus files.
[0154] Step 3: The fusion device 120 marks the valid files in the cloned file system. In step 3, the fusion device 120 can mark files other than virus files and virus-infected files as valid files. The fusion device 120 can also mark virus files and virus-infected files as invalid files, in which case the files other than invalid files are valid files.
[0155] Step 204: For any snapshot to be merged, the fusion device 120 performs availability verification on the snapshot. This step is optional. The fusion device 120 can perform this step or skip this step and directly execute step 205.
[0156] For any snapshot to be merged, after obtaining the cloned file system formed based on the snapshot, the merging device 120 can construct an application environment for the cloned file system, simulate various operations that need to be performed on the original file system in the application environment, and detect whether the cloned file system can perform these various operations normally in the application environment.
[0157] As described above regarding the data application environment, in step 204, the fusion device 120 can run applications that require the use of the cloned file system or its files. This application embodiment does not limit the number or type of such applications; any application that needs to use the cloned file system or its files during operation is applicable to this application embodiment. These applications will use the cloned file system or its files to perform various operations during operation. If such applications can use the cloned file system or its files normally and perform various operations without failure, then the cloned file system formed based on the snapshot is usable and can support the operation of the application or be used in the application environment. If such applications cannot use the cloned file system or its files normally and experience operation failures, then the cloned file system formed based on the snapshot is unusable, cannot support the operation of the application, or cannot be used in the application environment.
[0158] Using an application that requires the cloned file system and its files as database management software, in step 204, the fusion device 120 can mount the cloned file system on a host and run the database management software on that host. The database management software operates on the data in the cloned file system, such as performing operations like adding, deleting, querying, and modifying files, as well as merging and compressing files. If the database management software runs normally and all operations are performed correctly, then the cloned file system is usable; otherwise, the cloned file system is unusable.
[0159] If a first application is running during the creation of the cloned file system's application environment, then the available files in that snapshot are those that support running the first application. Similarly, if a second application is running during the creation of the cloned file system's application environment, then the available files in that snapshot are those that support running the second application.
[0160] Step 203 verifies the files in the snapshot from a data perspective to determine the valid files in the snapshot. Step 204 verifies the files in the snapshot from an application perspective to determine the usability of the files in the snapshot. These two preliminary operations can effectively ensure the validity and usability of the subsequent merged snapshot, that is, ensure that the files in the merged snapshot are valid and usable.
[0161] Step 205: The fusion device 120 merges the multiple snapshots to obtain a merged snapshot. If steps 203 and 204 have been executed, the multiple snapshots in step 205 are snapshots containing valid files and that have passed availability verification. In this case, the valid files in the snapshot are also usable files because they have passed availability verification, meaning they can support the operation of the application (such as the first application or the second application). If step 203 has been executed, the multiple snapshots in step 205 are snapshots containing valid files.
[0162] The process of the fusion device 120 performing step 205 is described below:
[0163] Step 2051: The fusion device 120 determines the valid files in the cloned file system formed based on each snapshot. That is, the fusion device 120 can determine the files in each cloned file system that are not infected by the virus.
[0164] From the perspective of file association, files are divided into associated files and non-associated files. For associated files, the fusion device 120 can determine at least one set of associated files existing in the storage system based on the analysis and / or the list of associated files in step 200.
[0165] Method 1: The fusion device 120 analyzes the audit information to determine at least one set of associated files within the storage device 110. For details on the implementation, please refer to the relevant description in step 200; it will not be repeated here.
[0166] Method 2: The fusion device 120 determines at least one set of associated files within the storage device 110 based on the associated file list. If the analysis based on audit information fails to obtain the at least one set of associated files, or if step 200 is not executed, the fusion device 120 can determine at least one set of associated files within the storage device 110 based on the associated file list. That is, the at least one set of associated files is a group of files recorded in the associated file list and included in the file system.
[0167] Method 3: The fusion device 120 determines at least one set of associated files within the storage device 110 based on audit information and a list of associated files.
[0168] For ease of explanation, in this method, the at least one set of associated files determined by the fusion device 120 based on audit information is referred to as at least one first candidate associated file group, and a first candidate associated file group is equivalent to a set of associated files. The at least one set of associated files determined by the fusion device 120 based on the list of associated files is referred to as at least one second candidate associated file group, and a second candidate associated file group is equivalent to a set of associated files. The relationship between any first candidate associated file group and any second candidate associated file group may have the following two states:
[0169] (1) There are identical files in the first candidate associated file group and the second candidate key file group.
[0170] This situation indicates that both the first candidate associated file group and the second candidate key file group contain the same one or more files. Based on this first candidate associated file group and the second candidate key file group, a set of associated files can be determined. This set of associated files includes all files from both the first candidate associated file group and the second candidate key file group.
[0171] For example, if the first candidate associated file group includes file A and file B, and the second candidate key file group includes file B and file C, then file B is present in both the first and second candidate key file groups. Therefore, it can be determined that there is an association relationship between files A, B, and C, meaning that files A, B, and C constitute a sufficiently long associated file group.
[0172] (2) The first candidate associated file group and the second candidate key file group do not have the same files.
[0173] This situation indicates that the files included in the first candidate associated file group and the second candidate key file group are different. Therefore, the first candidate associated group can be considered as a set of associated files, and the second candidate associated group can also be considered as a set of associated files.
[0174] Files other than associated files are considered non-associated files. For example, data in associated files is structured data, such as data tables; while data in non-associated files is unstructured data, such as images, audio / video files, PDF documents, and Word documents.
[0175] For any set of associated files, for ease of explanation, this set of associated files will be referred to as the first set of associated files. The fusion device 120 may execute steps 2052 to 2053. For any non-associated file, for ease of explanation, this non-associated file will be referred to as the first non-associated file. The fusion device 120 executes step 2054.
[0176] Step 2052: The fusion device 120 determines a first candidate snapshot. This first candidate snapshot is the snapshot in which the first group of associated files among multiple snapshots is a valid file. That is, the cloned file system formed based on the first candidate snapshot contains a first group of associated files, and the first group of associated files is a valid file. This application embodiment does not limit the number of first candidate snapshots; it can be one or more.
[0177] For the same set of associated files, the first set of associated files may exist in multiple cloned file systems, and the first set of associated files in each cloned file system is a version of the first set of associated files.
[0178] For example, the four snapshots to be merged are snapshot A, snapshot B, snapshot C, and snapshot D; correspondingly, the clone file systems formed based on these four snapshots are clone file system A, clone file system B, clone file system C, and clone file system D, respectively. A set of associated files includes three files, namely file 1, file 2, and file 3, and the first set of associated files is identified as file 1-2-3.
[0179] The file 1-2-3 exists in cloned file system A, and it is a valid file. The file 1-2-3 in this cloned file system is a version of file 1-2-3, which is here identified as file 1-2-3 [Ver-A]. Snapshot A, on which cloned file system A is based, is the first candidate snapshot.
[0180] The file 1-2-3 exists in cloned file system B, and it is a valid file. The file 1-2-3 in this cloned file system is a version of file 1-2-3, which is here identified as file 1-2-3 [Ver-B]. Snapshot B, on which cloned file system B is based, is the first candidate snapshot.
[0181] The file 1-2-3 exists in the cloned file system C, but it is not a valid file (e.g., some or all of the files in 1-2-3 are encrypted or infected by a virus). Therefore, the file 1-2-3 in the cloned file system C is not located. The snapshot C on which the cloned file system C is based is not the first candidate snapshot.
[0182] The file 1-2-3 exists in the cloned file system D, and the file 1-2-3 is a valid file. The file 1-2-3 in this cloned file system is a version of the file 1-2-3, which is identified here as file 1-2-3 [Ver-D]. The snapshot D on which the cloned file system D is based is the first candidate snapshot.
[0183] Step 2053: The fusion device 120 determines the latest version of the first set of associated files. The latest version of the first set of associated files is the first set of associated files recorded in the first snapshot, which is the first candidate snapshot among the first candidate blocks determined in step 2501 whose snapshot time point (the snapshot time point is also the creation time) is closest to the current time. The first set of associated files recorded in the first candidate snapshot essentially refers to the first set of associated files contained in the clone file system formed based on the first candidate snapshot.
[0184] Taking snapshots A, B, C, and D, and file 1-2-3 as an example, three first candidate snapshots were determined in step 2052: snapshot A, snapshot B, and snapshot D. Among these three snapshots, snapshot D's time point is closest to the current time point, so file 1-2-3 [Ver-D] is the latest version of file 1-2-3.
[0185] Steps 2052 to 2053 are executed for each group of associated files, and the latest version of each group of associated files is determined through steps 2052 to 2053.
[0186] It should be noted that if the fusion device 120 has performed step 204 and the first application has been run when constructing the application environment of the cloned file system, when determining the first candidate snapshot, it is necessary to determine the first candidate snapshot that meets the first availability condition. The first availability condition is: the cloned file system recorded by the first candidate snapshot supports the running of the first application.
[0187] Step 2054: The fusion device 120 determines the latest version of the non-associated file.
[0188] Step 2054 can also be broken down into two steps. First, the fusion device 120 determines a second candidate snapshot. This second candidate snapshot is a snapshot in which the first unrelated file is a valid file among multiple snapshots. The fusion device 120 can determine one or multiple second candidate snapshots. Then, the fusion device 120 determines the latest version of the first unrelated file from the second candidate snapshots. The latest version of the first unrelated file is the first unrelated file recorded in the second snapshot, which is the second candidate snapshot whose snapshot time point is closest to the current time among the determined second candidate snapshots. The first unrelated file recorded in the second candidate snapshot essentially refers to the first unrelated file contained in the cloned file system formed based on this second candidate snapshot.
[0189] Taking snapshots A, B, C, and D as an example, the non-associated file is identified as file 4.
[0190] File 4 exists in cloned file system A, and file 4 is a valid file. File 4 in this cloned file system is a version of file 4, which is identified here as file 4 [Ver-A]. Snapshot A, on which cloned file system A is based, is the second candidate snapshot.
[0191] File 4 exists in cloned file system B, and file 4 is a valid file. File 4 in this cloned file system is a version of file 4, which is identified here as file 4 [Ver-B]. Snapshot B, on which cloned file system B is based, is the second candidate snapshot.
[0192] The file 4 exists in the cloned file system C. File 4 is a valid file, and the file 4 in this cloned file system is a version of file 4, which is identified here as file 4 [Ver-C]. The snapshot C on which the cloned file system C is based is the second candidate snapshot.
[0193] The file 1-2-3 exists in cloned file system D, but it is not a valid file. File 4 in cloned file system D is not located. Snapshot D, on which cloned file system D is based, is not the second candidate snapshot.
[0194] Snapshots A, B, and C are the second candidate snapshots. Among these three snapshots, snapshot C's time point is closest to the current time point, so file 4 [Ver-C] is the latest version of file 4.
[0195] Step 2054 is performed for each non-associated file, and the latest version of each non-associated file is determined through step 2054.
[0196] It should be noted that if the fusion device 120 performs step 204 and runs the second application when constructing the application environment for the cloned file system, the second application may be the same as or different from the first application. When determining the second candidate snapshot, it is necessary to determine the first candidate snapshot that meets the second availability condition. This second availability condition is: the cloned file system recorded by the second candidate snapshot supports the operation of the second application.
[0197] Step 2055: The fusion device 120 constructs a file system based on the latest version of each group of associated files and the latest version of each non-associated file. For ease of explanation, this file system is referred to as the candidate file system.
[0198] The fusion device 120 organizes the latest versions of each group of associated files and each group of non-associated files to form a candidate file system. Specifically, for any group of associated files, the position of the latest version of that group of associated files in the candidate file system is consistent with the position of the latest version of that group of associated files in its respective clone file system; similarly, for any non-associated file, the position of the latest version of that non-associated file in the candidate file system is consistent with the position of the latest version of that non-associated file in its respective clone file system.
[0199] Continuing with snapshots A, B, C, D, and files 1-2-3 and 4 as examples, in cloned file system A, file 1-2-3 is located in directory 1-1; in cloned file system B, it is located in directory 1-2; and in cloned file system D, it is also located in directory 1-2. This means that the location of file 1-2-3 changed within the original file system between the creation of snapshot A and the creation of snapshot D. In the candidate file system, the location of the latest version of file 1-2-3 must match its location in cloned file system D, i.e., it must be located in directory 1-2.
[0200] In cloned file system A, file 4 is located in directory 1-2; in cloned file system B, file 4 is located in directory 1-2; in cloned file system C, file 4 is located in directory 1-1; that is, in cloned file system D, file 4's location in the original file system changed between the creation of snapshot A and the creation of snapshot D. The latest version of file 4 in the candidate file system must be located in the same directory as file 4 in cloned file system C, i.e., it must be located in directory 1-2.
[0201] Step 2056: The fusion device 120 creates a snapshot for the candidate file system, which is the fusion snapshot, and the candidate file system is the cloned file system formed based on the fusion snapshot.
[0202] Step 206: The fusion device 120 performs availability verification on the fused snapshot. The method by which the fusion device 120 performs availability verification on the fused snapshot is similar to the method by which the fusion device 120 performs availability verification on the snapshot itself. For details, please refer to the relevant description in step 204, which will not be repeated here.
[0203] Once the availability verification of the fused snapshot is passed, the fused snapshot can be obtained. The way the fusion device 120 processes the fused snapshot is related to the way the fusion device 120 receives the snapshot fusion instruction.
[0204] Scenario 1: The user triggers the snapshot fusion command.
[0205] After acquiring the merged snapshot, if a user directly triggers the snapshot fusion command, the fusion device 120 can display the merged snapshot to the user. For example, the fusion device 120 can prompt the user that the snapshot fusion is complete and display the merged snapshot in the snapshot fusion interface. Alternatively, the fusion device 120 can notify the user that the snapshot fusion is complete via SMS, email, or in-app notification and inform the user of the URL to view the merged snapshot.
[0206] If a user triggers a data recovery command, since the fusion device 120 has generated a candidate file system during the acquisition of the fusion snapshot, the fusion device 120 can use the candidate file system to recover the file system (i.e., the original file system). For example, it can replace the original file system with the candidate file system, or it can replace the target files that need to be recovered in the original file system with the target files in the candidate file system. After acquiring the fusion snapshot, the fusion device 120 displays the fusion snapshot and / or the candidate file system to the user. For example, the fusion device 120 can prompt the user that data recovery is complete in the snapshot viewing interface and display the fusion snapshot and the candidate file system. Alternatively, the fusion device 120 can notify the user that data recovery is complete via SMS, email, or in-application notifications, and inform the user of the URL to view the fusion snapshot and / or the candidate file system.
[0207] If the fusion device 120 fails to obtain a fusion snapshot, the fusion device 120 may notify the user that the snapshot fusion has failed or the data recovery has failed, and inform the user of the reason for the failure, such as that multiple snapshots to be fused are invalid, or that the fusion snapshot has failed the availability verification.
[0208] Scenario 2: Storage device 110 sends a snapshot fusion command to fusion device 120.
[0209] After acquiring the fusion snapshot, the fusion device 120 can send the fusion snapshot to the storage device 110; or it can store the fusion snapshot in the snapshot resource pool and notify the storage device 110 that the snapshot fusion is complete.
[0210] If the fusion device 120 fails to obtain a fused snapshot, it can notify the storage device 110 of the snapshot fusion failure and inform the user of the reason for the failure, such as multiple snapshots to be fused being invalid, or the fused snapshot failing availability verification. The storage device 110 can then re-initiate the snapshot fusion command to instruct the fusion device 120 to fuse other selected snapshots.
[0211] Case 3: Storage device 110 sends a data recovery command to fusion device 120.
[0212] Since the fusion device 120 has generated a candidate file system during the process of acquiring the fusion snapshot, the fusion device 120 can use the candidate file system to restore the file system (i.e., the original file system). For example, it can use the candidate file system to replace the file system, or it can replace the target file that needs to be restored in the original file system with the target file in the candidate file system.
[0213] After acquiring the fusion snapshot, the fusion device 120 can store the fusion snapshot in the snapshot resource pool of the storage device 110. Optionally, the fusion device 120 can also notify the storage device 110 that the fusion snapshot has been stored in the storage resource pool.
[0214] Scenario 4: The fusion device 120 generates a snapshot fusion command on its own.
[0215] After obtaining a fused snapshot, the fusion device 120 stores the fused snapshot in the snapshot resource pool, and can also delete multiple snapshots in the snapshot resource pool on which the fused snapshot is based.
[0216] If the fusion device 120 does not obtain a fusion snapshot, the fusion device 120 can re-obtain multiple snapshots from the snapshot resource pool (the multiple snapshots obtained this time may be different from or partially different from the multiple snapshots obtained previously), and obtain a fusion snapshot based on the multiple snapshots obtained.
[0217] It should be noted that the above description uses the restoration of the file system after the fusion device 120 generates a fusion snapshot as an example. In practical applications, the file system can be restored after step 2055, that is, after step 2025, a candidate file system is obtained and used to restore the file system. For example, when the fusion device 120 determines that the file system needs to be restored, such as when the received fingerprint is a data recovery command, the fusion device 120 can obtain a candidate file system after executing step 2055 and use the candidate file system to restore the file system. The fusion device 120 can execute steps 2056 and 206, or it can choose not to execute steps 2056 and 206. If the fusion device 120 generates a fusion snapshot (that is, executes step 2056), the fusion snapshot can be saved as a record of the candidate file system. When the file system needs to be restored again in the future, the fusion snapshot can be used to restore the file system. In addition, it is also ensured that the fusion snapshot facilitates the viewing of historical file systems (that is, candidate file systems). Furthermore, after generating a merged snapshot, a new snapshot is created. This merged snapshot can also be used as a snapshot to be merged with the new snapshot to generate a new merged snapshot.
[0218] The following describes the scenarios in which the snapshot fusion method provided in this application is applicable:
[0219] Scenario 1: Data recovery from storage devices.
[0220] like Figure 5 The above is a schematic diagram of the structure of a storage device, in which both the storage device 110 and the fusion device 120 are deployed. Figure 5 As shown, the storage device 130 includes at least a processor 132, memory 133, network card 134, and hard disk 105. The processor 132, memory 133, network card 134, and hard disk 135 are connected via a bus.
[0221] Processor 132 can be a central processing unit (CPU) or other specific integrated circuits. Processor 201 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 132 is the main processing core inside the storage device, handling data access requests from the outside and accessing the data within the storage device 130, such as writing data to or reading data from the storage device 130 (e.g., the hard disk 135 or memory 133). Processor 132 can also maintain the data stored in the storage device 130, such as updating data metadata, performing persistent data storage, and backing up data.
[0222] Memory 133 refers to the internal memory that directly exchanges data with the processor. Memory 133 can be dynamic random access memory (DRAM). Besides DRAM, memory 202 can also be other types of random access memory, such as static random access memory (SRAM). Additionally, memory 202 can also be read-only memory (ROM). For example, read-only memory can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). Memory 202 can also be flash memory, hard disk drive (HDD), or solid-state drive (SSD).
[0223] The network interface card 134 can be used to communicate with devices other than the storage device 130 and receive data access requests from the outside.
[0224] Hard disk 135 is used to provide storage resources, such as storing data. It can be a disk or other type of storage media, such as a solid-state drive or a shingled magnetic recording hard disk. Multiple hard disks 135 can be used to form multiple storage pools.
[0225] Taking a storage device 130 with a file system deployed on it as an example, the files in the file system within the storage device 130 are stored on the hard disk 135 within the storage device 130. The processor 132 within the storage device 130 can create snapshots for the file system and store them in a snapshot resource pool, which can be located in memory 133 or hard disk 135. The processor 133 within the storage device 130 also functions as a fusion device 120. After receiving a snapshot fusion command or data recovery command triggered by the user via the network interface card 134, the processor 132 can access the snapshot storage pool, retrieve multiple snapshots from the snapshot storage pool, and execute... Figure 2In the illustrated embodiment, steps 202-206 are executed, or steps 202-2055 are performed (i.e., steps 2025 and 206 are not executed). Upon receiving a user-triggered data recovery command, the processor 132, after executing step 2025, recovers the file system deployed on the storage device 130 using a candidate file system (upon receiving a user-triggered data recovery command). The processor 132 may execute step 2025 (optionally, step 206 may also be executed), or may not execute steps 2025 and 206.
[0226] Upon receiving a user-triggered snapshot fusion command, the processor 132 acquires the fused snapshot, stores it in the snapshot resource pool, and can transmit the fused snapshot to the user (upon receiving a user-triggered snapshot fusion command or data recovery command). When the file system needs to be restored subsequently, the processor 132 can also utilize candidate file systems to restore the file system deployed on the storage device 130 (upon receiving a user-triggered data recovery command).
[0227] The processor 132 can also determine independently when to perform data recovery on the file system deployed on the storage device 130, such as when it detects that the number of failed file operations in the file system exceeds a threshold. When the processor 132 determines that data recovery of the file system is necessary, it executes... Figure 2 The execution steps 202 to 2055 are shown in the embodiment. The processor 132 may execute step 2025 (optionally, it may also execute step 206), or it may not execute steps 2025 and 206.
[0228] Scenario 2: Support the snapshot fusion function of computing devices in the form of external devices.
[0229] like Figure 6 The diagram shown is a structural schematic of a computing device 140 provided in an embodiment of this application. The computing device 140 includes an I / O interface 141, a processor 142, a memory 143, and an external device 144. The I / O interface 141, the processor 142, the memory 143, and the external device 144 can be connected via a system bus. This system bus can be a peripheral component interconnect express (PCIe) bus, or a compute express link (CXL), universal serial bus (USB) protocol, or a bus using other protocols.
[0230] Figure 6 One example of the connection method is shown below. Figure 6 In this configuration, the external device 144 can be directly plugged into the slot on the motherboard of the computing device 140 and exchange data with the processor 142 via the PCIe bus 340.
[0231] I / O interface 141 is used to communicate with devices located outside computing device 140. For example, it can receive data access requests, snapshot fusion commands, or data recovery commands from devices outside computing device 140, or send back fused snapshots, files in the original file system, or files in the candidate file system to devices outside computing device 140 via I / O interface 141.
[0232] Processor 142 is the computing core and control core of computing device 140. The specific type of processor 142 is the same as that of processor 132. For details, please refer to the foregoing description. It will not be repeated here.
[0233] Memory 143 is typically used to store computer program instructions. Memory 143 can also be used to temporarily store data. The specific type of memory 143 is the same as that of main memory 133, as detailed above, and will not be repeated here.
[0234] The processor 142 is connected to the memory 143 via a double data rate (DDR) bus or other type of bus. The memory 143 is understood to be the RAM of the computing device 140.
[0235] Although not shown, the computing device 140 also includes persistent memory, or memory that can be remotely accessed by the computing device 140. Both persistent and remotely accessed memory can expand the storage space of the computing device 140. The persistent memory and remotely accessed memory are used to store data, such as files in a file system. The remotely accessed memory can be memory located outside the computing device 140 and connected to it via a network. This memory can be volatile memory, such as RAM, DRAM, SCM, or SRAM. It can also be non-volatile memory, such as ROM, flash memory, HDD, SSD, or SCM.
[0236] The persistent memory included in the computing device 140 can be connected to the computing device 140 via a system bus. This persistent memory can be a non-volatile memory such as ROM, flash memory, HDD, or SSD.
[0237] Within the computing device 140, the processor 142 can access the persistent memory and store the accessed data in the memory 143. In some cases, such as when the external device 144 offloads some of the functions of the processor 142, the external device 144 can access the persistent memory, retrieve the data stored in the persistent memory, and store the accessed data in the memory 143 or in the external device 144; the external device 144 can also access the memory 143 and retrieve the data in the memory 143.
[0238] In this embodiment, the computing device 140 simultaneously possesses the functions of a storage device 110 and a fusion device 120, and is capable of performing tasks such as... Figure 2 The steps performed by the storage device 110 and the fusion device 120 in the illustrated embodiment.
[0239] Inside the computing device 140, the processor 142 is able to execute, by calling computer program instructions stored in the memory 143, such as... Figure 2 The illustrated embodiment includes all steps performed by storage device 110. External device 144 performs the following steps: Figure 2 All steps performed by the fusion device 120 in the illustrated embodiment.
[0240] exist Figure 6 In this device, external device 144 is connected to computing device 140. External device 144 can be used as an external device of computing device 140. External device 144 can also be deployed inside computing device 140, such as external device 144 being located on the motherboard or backplane of computing device 140. Figure 6 A schematic diagram of an external device 144 deployed inside a computing device 140.
[0241] The external device 144 can serve as a data processing module attached to the computing device 140, undertaking some of the functions of the computing device 140 (processor 142). In other words, some functions of the computing device 140 are offloaded to the external device 144, which processes data and performs some tasks in place of the computing device 140 (such as the processor 142 in the computing device 140), thereby reducing the pressure on the processor 142 in the computing device 140 and freeing up its computing power.
[0242] In this embodiment, the external device 144 can perform snapshot fusion functionality. The external device 144 includes a processing module 1441 and a memory 1442, although not shown. The external device 144 may also include a power supply circuit. The processing module can be a data processing unit (DPU), and the processing module 1441 can also be other general-purpose processors, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processing module 1441 and the memory 1442 are connected via a system bus, which can be a PCIe-based line, or a bus using CXL, USB, or other protocols. The memory 1442 provides storage space for the processing module 1441 for storing data. The specific type and function of the memory 1442 are similar to those of the memory 143, as detailed in the foregoing description, and will not be repeated here.
[0243] The processing module 1441 is the main computing unit of the external device 144, and it undertakes the main functions of the external device 144. For example, the processing module 1441 can be used to perform functions that are offloaded from the processor 142 to the external device 144.
[0244] In this embodiment, the specific form of the external device 144 in the computing device 140 is not limited. The external device 144 can be deployed in the computing device 140 as an offloading card or an accelerator card. The external device 144 can also be a smart network card, which, in addition to offloading the processor 142, also has the functions of a network card. That is, the external device 144 can process data packets based on network protocols, such as encapsulating and transmitting data, and perform data transmission with devices outside the computing device 140 (that is, the external device 144 specifically performs the functions of the aforementioned I / O interface).
[0245] Taking a computing device 140 with a file system deployed on it as an example, the files in the file system within the computing device 140 are stored in the persistent storage of the computing device 140. The processor 142 within the computing device 140 can create snapshots for the file system and store these snapshots in a snapshot resource pool, which can be located in storage 143 or persistent storage. The processor 142 within the computing device 140 has the functionality of the storage device 110. After receiving a snapshot fusion command or data recovery command triggered by a user through the I / O interface 141, the processor 142 can transmit a snapshot fusion command, formed by splitting the received snapshot fusion command or data recovery command, to an external device 144. Upon receiving the snapshot fusion command, the external device 144 accesses the snapshot storage pool, retrieves multiple snapshots from the snapshot storage pool, and executes... Figure 2In steps 202 to 206 of the illustrated embodiment, after acquiring the fused snapshot, the fused snapshot is stored in the snapshot resource pool. The processor 142 acquires the fused snapshot from the snapshot resource pool, forms a target clone system based on the fused snapshot, and restores the file system deployed on the computing device 140 using a candidate file system (in the event of receiving a user-triggered data recovery command).
[0246] The processor 142 can also spontaneously decide whether to obtain a merged snapshot or initiate a data recovery process. In this case, the processor 142 generates a snapshot fusion command and transmits it to the external device 144. Upon receiving the snapshot fusion command, the external device 144 accesses the snapshot storage pool, retrieves multiple snapshots from the pool, and executes the fusion command. Figure 2 In steps 202 to 206 of the illustrated embodiment, after obtaining the fused snapshot, the fused snapshot is stored in the snapshot resource pool. The processor 142 obtains the fused snapshot from the snapshot resource pool, forms a target clone file system based on the fused snapshot, and restores the file system deployed on the computing device 140 using the candidate file system.
[0247] The foregoing description only illustrates the use of an external device with snapshot fusion functionality as an example. In practical applications, the external device 144 can also have data recovery functionality. After receiving a data recovery command from the processor 142, the external device 144 can execute... Figure 2 In the illustrated embodiment, steps 202 to 2055 are executed (i.e., steps 2025 and 206 may be omitted). After obtaining the candidate file system, the file system is restored. The external device 144 may execute step 2025 (optionally, step 206 may also be executed) to store the merged snapshot in the snapshot resource pool after obtaining it. The external device 144 may also omit steps 2025 and 206.
[0248] External device 144 can also determine when to perform data recovery on the file system deployed on computing device 140, such as when it detects that the number of failed file operations in the file system exceeds a threshold. When external device 144 determines that data recovery of the file system is necessary, it executes... Figure 2 The execution steps 202 to 2055 are shown in the embodiment. The processor 142 may execute step 2025 (optionally, it may also execute step 206), or it may not execute steps 2025 and 206.
[0249] Scenario 3: Systems containing multiple computing devices support data recovery functionality.
[0250] This application also provides a computing device system, the computing device system including at least one such as Figure 7 The computing device 700 shown includes a bus 701, a processor 702, a communication interface 703, and a memory 704. The processor 702, memory 704, and communication interface 703 communicate with each other via the bus 701. At least one computing device 700 in the computing device system communicates with each other via a communication path.
[0251] The processor 702 can be a CPU, or it can be other general-purpose processors, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0252] Memory 704 can be DRAM. Besides DRAM, memory 1604 can also be other random access memory (such as SRAM). Additionally, memory 1602 can also be ROM. For read-only memory, for example, it could be PROM, EPROM, etc. Memory 1604 can also be flash memory, HDD, or SSD.
[0253] Processor 702 executes the computer program instructions stored in memory 704 to perform the aforementioned tasks. Figure 2 The described method includes some or all of the steps performed by the fusion device 120. The memory may also include other software modules required for running processes, such as an operating system. The operating system may be Linux. TM UNIX TM WINDOWS TM wait.
[0254] At least one computing device 700 in the computing device system establishes communication with each other through a communication network, and each computing device 700 runs as follows: Figure 8 Any one or more modules in the fusion device 800 shown.
[0255] At least one computing device 700 in the computing device system establishes communication with each other through a communication network, and each computing device 700 runs on such a network. Figure 8 The shown can be any one or more modules in the fusion device 800.
[0256] Based on the same inventive concept as the method embodiments, this application also provides a fusion apparatus for executing the method performed by the fusion apparatus 120 in the above method embodiments. For example... Figure 8 As shown, the fusion device 800 includes an acquisition module 801, a recovery module 802, and optionally, a detection module 803. Specifically, in the fusion device 800, the modules are connected through a communication path. The fusion device 800 includes:
[0257] The acquisition module 801 is used to acquire multiple snapshots to be merged, wherein each snapshot is created for a data set and the creation time of different snapshots is different; and to acquire the valid data of multiple snapshots, wherein the valid data in each snapshot is the data in the data set recorded in each snapshot that has not been infected by the virus.
[0258] Recovery module 802 is used to recover a dataset based on valid data from multiple snapshots.
[0259] As one possible implementation, when the recovery module 802 recovers the data set, it obtains a fused snapshot based on the valid data in multiple snapshots; the recovery module 802 then recovers the data set based on the fused snapshot.
[0260] As one possible implementation, the data set includes at least one set of related data, and each set of related data includes multiple related data. When the recovery module 802 recovers the data set, for the first set of related data in the at least one set of related data, the latest version of the first set of related data is determined. The latest version of the first set of related data is the first set of related data recorded in the first snapshot. The first snapshot is the snapshot whose creation time is closest to the current time among at least one first candidate snapshot. At least one first candidate snapshot is the snapshot in which the first set of related data is valid data among multiple snapshots.
[0261] The recovery module 802 recovers the data set based on the latest version of the first set of associated data.
[0262] In one possible implementation, when the recovery module 802 determines the latest version of the first set of associated data, it determines at least one first candidate snapshot among multiple snapshots that meets a first availability condition. The first availability condition is that the data set recorded by the first candidate snapshot supports the operation of the first application. The recovery module 802 determines the latest version of the first set of associated data from at least one first candidate snapshot.
[0263] In one possible implementation, the dataset includes unrelated data, which refers to data in the dataset other than at least one set of related data. When the recovery module 802 recovers the dataset, for the first unrelated data, it determines the latest version of the first unrelated data. The latest version of the first unrelated data is the first unrelated data recorded in a second snapshot. The second snapshot is the snapshot with the creation time closest to the current time among at least one second candidate snapshot. The at least one second candidate snapshot is a snapshot among multiple snapshots where the first unrelated data is valid. The recovery module 802 obtains a merged snapshot based on the latest version of the first unrelated data.
[0264] In one possible implementation, when the recovery module 802 determines the latest version of the first non-associated data, it determines at least one second candidate snapshot among multiple snapshots that meets a second availability condition. The second availability condition is that the data set recorded by at least one second candidate snapshot supports the operation of the second application. The recovery module 802 determines the latest version of the first non-associated data from at least one second candidate snapshot.
[0265] As one possible implementation, the acquisition module 801 acquires audit information, which describes the operation information on the data in the data set; the recovery module 802 determines at least one set of related data based on the audit information.
[0266] As one possible implementation, the recovery module 802 determines at least one set of associated data based on audit information and a list of associated data.
[0267] As one possible implementation, the detection module 803 detects the data set recorded in any one of the multiple snapshots to determine the valid data in the snapshot.
[0268] As one possible implementation, the recovery module 802 obtains the latest version of non-related data and at least one set of the latest version of related data from the valid data in multiple snapshots, constructs a candidate data set, and uses the candidate data set to recover the data set.
[0269] As one possible implementation, the recovery module 802 creates a snapshot of the candidate data set to obtain a fused snapshot.
[0270] The module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0271] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device (which may be a personal computer, mobile phone, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0272] The descriptions of the processes corresponding to the above-mentioned figures each have their own emphasis. For parts of a process that are not described in detail, please refer to the relevant descriptions of other processes.
[0273] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, in the form of a computer program product. A computer program product includes computer program instructions, which, when loaded and executed on a computer, generate, in whole or in part, the product according to the embodiments of the present invention. Figure 2 The process or function described.
[0274] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., SSD).
[0275] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A data recovery method, characterized in that, include: Obtain multiple snapshots to be merged, where each snapshot is created for a specific dataset and the creation time of different snapshots is different; Obtain valid data from multiple snapshots, wherein the valid data in each snapshot is the data in the data set recorded in each snapshot that has not been infected by the virus; The data set is restored based on the valid data in the multiple snapshots.
2. The method as described in claim 1, characterized in that, The data recovery process based on valid data from the multiple snapshots includes: Based on the valid data in the multiple snapshots, a fused snapshot is obtained; The dataset is restored based on the fused snapshot.
3. The method as described in claim 1 or 2, characterized in that, The data set includes at least one set of related data, and each set of related data includes multiple related data. The recovery of the data set based on the valid data from the multiple snapshots includes: For the first set of associated data (without sequence number) in the at least one set of associated data, determine the latest version of the first set of associated data. The latest version of the first set of associated data is the first set of associated data recorded in the first snapshot. The first snapshot is the snapshot whose creation time is closest to the current time among at least one first candidate snapshot. The at least one first candidate snapshot is the snapshot in the plurality of snapshots in which the first set of associated data is valid data. The dataset is restored based on the latest version of the first set of associated data.
4. The method as described in claim 3, characterized in that, Determining the latest version of the first set of associated data includes: Determine at least one first candidate snapshot from the plurality of snapshots that satisfies a first availability condition, wherein the first availability condition is: the data set recorded in the first candidate snapshot supports the operation of a first application; The latest version of the first set of associated data is determined from the at least one first candidate snapshot.
5. The method according to any one of claims 1 to 4, characterized in that, The data set includes unrelated data, which refers to data in the data set other than the at least one set of related data. The recovery of the data set based on valid data from the multiple snapshots includes: For the first unrelated data in the unrelated data, determine the latest version of the first unrelated data. The latest version of the first unrelated data is the first unrelated data recorded in the second snapshot. The second snapshot is the snapshot whose creation time is closest to the current time among at least one second candidate snapshot. The at least one second candidate snapshot is the snapshot in the plurality of snapshots in which the first unrelated data is valid data. The fusion snapshot is obtained based on the latest version of the first non-associated data.
6. The method as described in claim 5, characterized in that, Determining the latest version of the first unrelated data includes: Determine at least one second candidate snapshot from the plurality of snapshots that meets a second availability condition, wherein the second availability condition is that the data set recorded by the at least one second candidate snapshot supports the operation of a second application; The latest version of the first non-associated data is determined from the at least one second candidate snapshot.
7. The method as described in claim 3 or 4, characterized in that, The method further includes: Obtain audit information, which describes the operation information on the data in the dataset; Based on the audit information, the at least one set of related data is determined.
8. The method as described in claim 7, characterized in that, The determination of the at least one set of related data based on the audit information includes: Based on the audit information and the list of related data, the at least one set of related data is determined.
9. The method according to any one of claims 1 to 8, characterized in that, Before obtaining the fused snapshot based on the valid data from the plurality of snapshots, the process further includes: For any one of the multiple snapshots, the data set recorded in the snapshot is examined to determine the valid data in the snapshot.
10. The method according to any one of claims 1 to 9, characterized in that, The process of restoring the data set based on valid data from the multiple snapshots includes: Obtain the latest version of the non-related data and the latest version of the at least one set of related data from the valid data in the plurality of snapshots to construct a candidate data set; The candidate data set is used to recover the data set.
11. The method as described in claim 10, characterized in that, The process of obtaining a fused snapshot based on valid data from the plurality of snapshots includes: A snapshot is created on the candidate data set to obtain the fused snapshot.
12. A fusion device, characterized in that, include: The acquisition module is used to acquire multiple snapshots to be merged, wherein each snapshot is created for a dataset and the creation time of different snapshots is different; and to acquire valid data from multiple snapshots, wherein the valid data in each snapshot is the data in the dataset recorded in each snapshot that has not been infected by the virus. The recovery module is used to recover the data set based on the valid data in the plurality of snapshots.
13. The apparatus as claimed in claim 12, characterized in that, The recovery module is used for: Based on the valid data in the multiple snapshots, a fused snapshot is obtained; The dataset is restored based on the fused snapshot.
14. The apparatus as claimed in claim 12 or 13, characterized in that, The dataset includes at least one set of related data, and each set of related data includes multiple related data. The recovery module is used for: For the first set of associated data in the at least one set of associated data, determine the latest version of the first set of associated data. The latest version of the first set of associated data is the first set of associated data recorded in the first snapshot. The first snapshot is the snapshot whose creation time is closest to the current time among at least one first candidate snapshot. The at least one first candidate snapshot is the snapshot in the plurality of snapshots in which the first set of associated data is valid data. The dataset is restored based on the latest version of the first set of associated data.
15. The apparatus as claimed in claim 14, characterized in that, The recovery module is used for: Determine at least one first candidate snapshot from the plurality of snapshots that satisfies a first availability condition, wherein the first availability condition is: the data set recorded in the first candidate snapshot supports the operation of a first application; The latest version of the first set of associated data is determined from the at least one first candidate snapshot.
16. The apparatus according to any one of claims 12 to 15, characterized in that, The data set includes unrelated data, which is data in the data set other than the at least one set of related data. The recovery module is used for: For the first unrelated data in the unrelated data, determine the latest version of the first unrelated data. The latest version of the first unrelated data is the first unrelated data recorded in the second snapshot. The second snapshot is the snapshot whose creation time is closest to the current time among at least one second candidate snapshot. The at least one second candidate snapshot is the snapshot in the plurality of snapshots in which the first unrelated data is valid data. The fusion snapshot is obtained based on the latest version of the first non-associated data.
17. The apparatus as claimed in claim 16, characterized in that, The recovery module is used for: Determine at least one second candidate snapshot from the plurality of snapshots that meets a second availability condition, wherein the second availability condition is that the data set recorded by the at least one second candidate snapshot supports the operation of a second application; The latest version of the first non-associated data is determined from the at least one second candidate snapshot.
18. The apparatus as claimed in claim 14 or 15, characterized in that, The acquisition module is further configured to: acquire audit information, wherein the audit information describes operation information on the data in the data set; The recovery module is used to determine the at least one set of associated data based on the audit information.
19. The apparatus as claimed in claim 18, characterized in that, The recovery module is used for: Based on the audit information and the list of related data, the at least one set of related data is determined.
20. The apparatus according to any one of claims 12 to 19, characterized in that, The fusion device further includes a detection module, which is used for: For any one of the multiple snapshots, the data set recorded in the snapshot is examined to determine the valid data in the snapshot.
21. The apparatus according to any one of claims 12 to 20, characterized in that, The recovery module is used for: Obtain the latest version of the non-related data and the latest version of the at least one set of related data from the valid data in the plurality of snapshots to construct a candidate data set; The candidate data set is used to recover the data set.
22. The apparatus as claimed in claim 21, characterized in that, The recovery module is used for: A snapshot is created on the candidate data set to obtain the fused snapshot.
23. A computing device, characterized in that, The computing device includes a processor and memory; The memory is used to store computer program instructions; The processor executes computer program instructions in the memory to perform the method as described in any one of claims 1 to 11.
24. A computer-readable storage medium, characterized in that, When the computer-readable storage medium is executed by a computing device, the computing device performs the method according to any one of claims 1 to 11.