Data recovery method based on cloud management platform, cloud management platform and other equipment
By selecting a target storage pool with high IO capabilities through a cloud management platform for data recovery, the problem of inefficient recovery when logical volumes fail in block storage services is solved, achieving faster data recovery speed and higher recovery efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, block storage services have low data recovery efficiency and long recovery time when logical volume fails, and cannot effectively utilize the differences in IO capabilities among multiple storage pools.
Based on the storage pool's IO capability information, the cloud management platform selects a target storage pool with higher average IO capability from multiple storage pools, restores the backup data of the faulty logical volume to these target storage pools, and utilizes the high IO capability of multiple storage pools for concurrent data writing.
It significantly improves data recovery efficiency, shortens data recovery time, and solves the problem of excessively long recovery time caused by poor IO performance of a single storage pool.
Smart Images

Figure CN121807618A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing, and more particularly to a data recovery method based on a cloud management platform, the cloud management platform, and other devices. Background Technology
[0002] With the development of cloud technology, the cybersecurity situation has become increasingly severe, and the demand for data protection in the cloud has become more and more intense. Currently, the cloud offers massive storage for tenants, among which block storage services are a type of cloud storage service. Block storage services can include multiple storage pools, each containing one or more logical volumes. Tenants' virtual machines can access specific logical volumes through block storage services to achieve secure access to massive storage.
[0003] In the event of a failure in a logical volume accessed by a tenant through a virtual machine, rapid recovery of the failed logical volume is required in a disaster recovery scenario. Block storage services primarily achieve data recovery by restoring backup data from the failed logical volume to that volume. However, this method suffers from low data recovery efficiency and a long recovery time. Summary of the Invention
[0004] This application provides a data recovery method, cloud management platform and other devices based on a cloud management platform, which can determine the target storage pool for data recovery based on the storage pool's IO capability information, which has an average IO capability higher than that of the source storage pool, so as to restore the data of the faulty logical volume to the logical volume in the target storage pool, thereby improving the data recovery efficiency of the source volume in the source storage pool and shortening the data recovery time.
[0005] Firstly, this application provides a data recovery method based on a cloud management platform. The cloud management platform manages infrastructure providing cloud services, including multiple storage nodes, each providing multiple storage pools. The method includes: the cloud management platform receiving a data recovery request, the data recovery request including information indicating a source logical volume of data to be recovered; the cloud management platform determining N target storage pools from the multiple storage pools based on the data recovery request and the IO capability information of the multiple storage pools, where N is a positive integer greater than or equal to 1, the average IO capability of the N target storage pools is higher than the IO capability of the source storage pool, the IO capability of each storage pool in the multiple storage pools is used to indicate the data recovery capability of each storage pool, and the source storage pool is the storage pool to which the source logical volume belongs among the multiple storage pools; the cloud management platform, based on the data recovery request, recovers the backup data of the source logical volume to the logical volume included in the N target storage pools.
[0006] The aforementioned multiple storage nodes can be a storage cluster, multiple storage clusters, or a distributed storage system; there are no restrictions here.
[0007] The multiple storage pools provided by the aforementioned multiple storage nodes can be all or part of the storage pools in the multiple storage nodes; there are no restrictions here.
[0008] For ease of explanation, the logical volume containing the data to be recovered will be referred to as the "source volume," and the storage pool to which the source volume belongs will be referred to as the "source storage pool."
[0009] In related technologies, data recovery of the source volume is performed directly within the source volume, specifically by restoring all backup data of the source volume to the source volume. However, if the IO performance (e.g., IO throughput) of the source storage pool is poor, this results in excessively long data recovery time and low data recovery efficiency.
[0010] However, in this embodiment, the cloud management platform can determine N target storage pools from multiple storage pools provided by multiple storage nodes based on the data recovery request and the IO capability information of multiple storage pools. Furthermore, the average IO capability of the N target storage pools is higher than the IO capability of the source storage pool containing the source volume, and the IO capability of each storage pool is used to indicate the data recovery capability of each storage pool. Thus, the data write rate of the storage pool with higher IO capability is higher than that of the storage pool with lower IO capability. Therefore, compared to restoring the backup data of the source volume to the source volume, the method of this embodiment restores the backup data of the source volume to the logical volume included in the N target storage pools with higher average IO capability, which significantly shortens the data recovery time of the logical volume and improves the data recovery efficiency of the logical volume.
[0011] In one possible implementation, the cloud management platform determines N target storage pools from the multiple storage pools based on the data recovery request and the IO capability information of the multiple storage pools, including: the cloud management platform obtaining the IO capability information of the multiple storage pools based on the data recovery request; and the cloud management platform determining N target storage pools from the multiple storage pools based on the IO capability information of the multiple storage pools.
[0012] Unlike data recovery solutions in related technologies, in this embodiment, the cloud management platform can obtain the IO capability information of multiple storage pools based on the data recovery request. However, solutions in related technologies do not support the cloud management platform obtaining the IO capability information of storage pools. Therefore, the cloud management platform in this embodiment can determine N target storage pools from multiple storage pools based on the IO capability information of each storage pool, and restore the backup data of the source volume to the logical volumes included in the N target storage pools. Because the computing power of the cloud management platform is stronger than that of storage nodes, it can filter out N target storage pools from multiple storage pools based on the IO capability information, thereby determining the target storage pools more quickly, achieving rapid recovery of the source volume, shortening the data recovery time of the source volume, and improving data recovery efficiency.
[0013] In one possible implementation, the cloud management platform obtains the IO capability information of the multiple storage pools based on the data recovery request, including: the cloud management platform receiving IO metrics from the multiple storage nodes of the multiple storage pools based on the data recovery request, the IO metrics including at least one of the following: IO throughput and IO utilization; the cloud management platform determining the IO capability information of the multiple storage pools based on the IO metrics of the multiple storage pools.
[0014] Unlike data recovery solutions in related technologies, the cloud management platform in this application embodiment can receive IO metrics from multiple storage pools across multiple storage nodes. These IO metrics include at least one of the following: IO throughput and IO utilization. Thus, the cloud management platform can determine the IO capacity information of multiple storage pools based on their IO metrics, and then use this information to determine N target storage pools. This makes the target storage pools in this application selectable, whereas in related technologies where data recovery is performed directly on the source volume, the target storage pool cannot be selected.
[0015] Furthermore, in this embodiment, the cloud management platform can receive IO metrics of multiple storage pools reported by multiple storage nodes, so that the cloud management platform can determine the IO capability information of the storage pool based on the IO metrics of the storage pool, thereby enabling the cloud management platform to share a portion of the computing load of multiple storage nodes, thereby improving the overall speed of data recovery.
[0016] In one possible implementation, the cloud management platform obtains the IO capability information of the multiple storage pools based on the data recovery request, including: the cloud management platform receiving the IO capability information of the multiple storage pools from the multiple storage nodes based on the data recovery request.
[0017] In this way, the cloud management platform can directly receive the IO capability information of each storage pool in multiple storage pools from multiple storage nodes, thereby reducing the computational overhead of the cloud management platform for IO capability information.
[0018] In one possible implementation, the cloud management platform receives IO capability information of the multiple storage pools from the multiple storage nodes based on the data recovery request, including: the cloud management platform sending a first request to the multiple storage nodes based on the data recovery request, the first request being used to obtain IO capability information of the multiple storage pools; the cloud management platform receiving the IO capability information of the multiple storage pools from the multiple storage nodes, the IO capability information of the multiple storage pools being determined based on IO metrics of the multiple storage pools, the IO metrics including at least one of the following: IO throughput and IO utilization.
[0019] The first request is used to obtain IO capability information of multiple storage pools. The multiple storage pools indicated in the first request can be all the storage pools provided by the multiple storage nodes, or they can be some of the storage pools provided by the multiple storage nodes.
[0020] The cloud management platform can specify which storage pool(s) to obtain IO capability information through the first request, or the cloud management platform can not specify the storage pool in the first request, but multiple storage nodes can select one or more or all storage pools based on the first request, and obtain the IO indicators of the storage pools selected by the multiple storage nodes to determine the IO capability information of the selected storage pools.
[0021] Unlike data recovery solutions in related technologies, the cloud management platform in this application embodiment can receive IO capability information of multiple storage pools sent by multiple storage nodes. Thus, the cloud management platform can utilize the IO capability information received from multiple storage nodes to select N target storage pools from these pools. This makes the target storage pools in this application selectable, whereas in related technologies where data recovery is performed directly on the source volume, the target storage pool cannot be selected.
[0022] In one possible implementation, where N ≥ 2, the cloud management platform restores the backup data of the source logical volume to the logical volumes included in the N target storage pools based on the data recovery request. This includes: the cloud management platform sending a second request to the plurality of storage nodes based on the data recovery request, the second request indicating the creation of a new target logical volume belonging to the N target storage pools, the target logical volume including N logical spaces located within the N target storage pools; and the cloud management platform restoring the backup data of the source logical volume to the N logical spaces within the N target storage pools based on the data recovery request.
[0023] For example, the logical space of the target logical volume is contiguous, making the addresses of the N logical spaces contiguous.
[0024] In block storage services of related technologies, it is not supported to restore backup data to logical volumes spanning multiple storage pools. However, in the embodiments of this application, the cloud management platform can send requests to multiple storage nodes providing multiple storage pools based on data recovery requests, instructing the creation of logical volumes belonging to N target storage pools (here referred to as target logical volumes); for example, the target logical volume may include N logical spaces located within the N target storage pools, with each of the N logical spaces corresponding one-to-one with the N target storage pools. Therefore, when performing data recovery on the source volume, the cloud management platform can restore the backup data of the source volume to the N logical spaces within the N target logical volumes based on the data recovery request. This allows the use of multiple storage pools with strong IO capabilities to write the backup data of the source volume, thereby accelerating the data recovery speed.
[0025] Furthermore, a logical volume within a storage pool of multiple storage nodes can include multiple data blocks, each of which is the same size. When the source volume has a large amount of data, even with strong IO capabilities in the source storage pool, writing massive amounts of data block copies to the same storage pool (here, the source storage pool) can lead to prolonged data recovery times. Therefore, a large source volume also results in a long data recovery time. To overcome the limitation of long data recovery times caused by a large source volume, in this embodiment, because data copies of each data block in the source volume can be written to multiple target storage pools with strong IO capabilities, the impact of the source volume size on data recovery time can be reduced. Even if the source volume is large, because the N logical spaces of the target volume span multiple storage pools, the writing speed of backup data to each logical space can be accelerated, thereby improving the overall recovery speed of the source volume and shortening the data recovery time.
[0026] In one possible implementation, the N target storage pools include a first storage pool and a second storage pool, the N logical spaces include a first logical space located within the first storage pool and a second logical space located within the second storage pool, and the backup data of the source logical volume includes first backup data and second backup data; the cloud management platform restores the backup data of the source logical volume to the N logical spaces within the N target storage pools based on the data recovery request, including: the cloud management platform receiving metadata of the target logical volume sent by the plurality of storage nodes, the metadata including a first matching relationship between a first access address of the first logical space and the first storage pool, and a second matching relationship between a second access address of the second logical space and the second storage pool; the cloud management platform, based on the metadata and the data recovery request, restores the first backup data of the source logical volume to the first logical space within the first storage pool according to the first access address in the first matching relationship, and restores the second backup data of the source logical volume to the second logical space within the second storage pool according to the second access address in the second matching relationship.
[0027] In this embodiment, the metadata of a logical volume created across multiple target storage pools differs from that of a logical volume created within a single storage pool. The metadata of a logical volume created across multiple target storage pools includes the matching relationship between the access addresses of the logical spaces of the target volumes distributed across the target storage pools and the target storage pools to which those logical spaces belong. Thus, during the process of restoring data copies of the source volume to a target volume across multiple target storage pools, the cloud management platform can restore each data copy of the source volume to its corresponding logical space within the corresponding target storage pool according to the matching relationship in the metadata. This enables data writing to logical volumes across multiple target storage pools and achieves cross-storage pool data recovery.
[0028] In one possible implementation, the cloud management platform, based on the metadata and the data recovery request, restores the first backup data of the source logical volume to the first logical space within the first storage pool according to the first access address in the first matching relationship, and restores the second backup data of the source logical volume to the second logical space within the second storage pool according to the second access address in the second matching relationship. This includes: the cloud management platform creating multiple threads, including a first thread and a second thread, for writing backup data of the source logical volume to the N target storage pools based on the metadata and the data recovery request; the cloud management platform writing the first backup data of the source logical volume to the first logical space within the first storage pool based on the first thread and the first access address in the first matching relationship; and during the process of the cloud management platform writing the first backup data to the first logical space through the first thread, the cloud management platform writing the second backup data of the source logical volume to the second logical space within the second storage pool based on the second thread and the second access address in the second matching relationship.
[0029] In this process, the cloud management platform can control the writing of the first backup data to the first logical space within the first storage pool using either a single first thread or multiple first threads concurrently. Similarly, the cloud management platform can control the writing of the second backup data to the second logical space within the second storage pool using either a single second thread or multiple second threads concurrently. Thus, the aforementioned multiple threads can include at least one first thread and at least one second thread. Having multiple first and second threads can also improve the data writing speed when writing backup data to a single storage pool.
[0030] In data recovery methods in related technologies, since the source storage pool where the source volume is located is unique, when writing a copy of the source volume to the source storage pool, it can only be done through a single task corresponding to a single source storage pool. However, the processing capacity of a single task is limited, which reduces the efficiency of data recovery.
[0031] However, in this embodiment, target volumes can be created across multiple target storage pools for writing data copies of the source volume. This allows each target storage pool to have at least one thread to write a partial data copy of the source volume to a target storage pool, enabling concurrent writing of source volume data copies to target volumes within different target storage pools using multiple threads corresponding to N target storage pools. Related technologies, however, do not support concurrent writing of a source volume's data copy to multiple storage pools.
[0032] In contrast to related technologies where backup data of the source volume can only be written to a single storage pool, resulting in data recovery efficiency being limited by the single-threaded task processing capability of a single storage pool, the embodiments of this application allow for separate threads to write data copies to each of the multiple target storage pools. This leverages the combined IO processing capabilities of multiple target storage pools, enabling multi-tasking data recovery of the source volume without being limited by the processing capability of a single task, thus improving data recovery efficiency.
[0033] Secondly, this application provides a cloud management platform for managing infrastructure providing cloud services. The infrastructure includes multiple storage nodes, which provide multiple storage pools. The cloud management platform includes: a receiving module for receiving a data recovery request, the data recovery request including information indicating a source logical volume of data to be recovered; a determining module for determining N target storage pools from the multiple storage pools based on the data recovery request and the input / output I / O capability information of the multiple storage pools, where N is a positive integer greater than or equal to 1, the average I / O capability of the N target storage pools is higher than the I / O capability of the source storage pool, the I / O capability of each storage pool in the multiple storage pools is used to indicate the data recovery capability of each storage pool, and the source storage pool is the storage pool to which the source logical volume belongs among the multiple storage pools; and a processing module for restoring the backup data of the source logical volume to the logical volume included in the N target storage pools based on the data recovery request.
[0034] In one possible implementation, the determining module is specifically configured to: obtain IO capability information of the plurality of storage pools based on the data recovery request; and determine N target storage pools from the plurality of storage pools based on the IO capability information of the plurality of storage pools.
[0035] In one possible implementation, the determining module is specifically configured to: receive IO metrics from the plurality of storage pools of the plurality of storage nodes based on the data recovery request, wherein the IO metrics include at least one of the following: IO throughput and IO utilization; and determine the IO capability information of the plurality of storage pools based on the IO metrics of the plurality of storage pools.
[0036] In one possible implementation, the determining module is specifically configured to: receive IO capability information of the multiple storage pools from the multiple storage nodes based on the data recovery request.
[0037] In one possible implementation, the determining module is specifically configured to: send a first request to the plurality of storage nodes based on the data recovery request, the first request being used to obtain IO capability information of the plurality of storage pools; receive the IO capability information of the plurality of storage pools from the plurality of storage nodes, the IO capability information of the plurality of storage pools being determined based on the IO metrics of the plurality of storage pools, the IO metrics including at least one of the following: IO throughput and IO utilization.
[0038] Thirdly, this application provides a cloud service system. The cloud service system includes a cloud management platform and multiple storage nodes. The cloud management platform manages the infrastructure providing the cloud service, and the infrastructure includes the multiple storage nodes, which provide multiple storage pools. The cloud management platform receives data recovery requests, which include information indicating the source logical volume of the data to be recovered. Based on the data recovery request and the IO capability information of the multiple storage pools, the cloud management platform determines N target storage pools from the multiple storage pools, where N is a positive integer greater than or equal to 1, the average IO capability of the N target storage pools is higher than the IO capability of the source storage pool, the IO capability of each storage pool in the multiple storage pools is used to indicate the data recovery capability of each storage pool, and the source storage pool is the storage pool to which the source logical volume belongs among the multiple storage pools. The cloud management platform restores the backup data of the source logical volume to the logical volume included in the N target storage pools based on the data recovery request.
[0039] In one possible implementation, the cloud management platform is further configured to send a first request to the plurality of storage nodes based on the data recovery request, the first request being used to obtain IO capability information of the plurality of storage pools; the plurality of storage nodes are configured to obtain IO metrics of the plurality of storage pools based on the first request, the IO metrics including at least one of the following: IO throughput capacity, IO utilization rate; and determine the IO capability information of the plurality of storage pools based on the IO metrics of the plurality of storage pools; the cloud management platform is further configured to receive the IO capability information of the plurality of storage pools from the plurality of storage nodes.
[0040] In this embodiment, multiple storage nodes can respond to a first request from the cloud management platform to obtain IO metrics of multiple storage pools, such as IO throughput and / or IO utilization. This allows them to determine the IO capability information of each storage pool. Finally, the multiple storage nodes can send the IO capability information of each storage pool to the cloud management platform, enabling the cloud management platform to receive the IO capability information of multiple storage pools from the multiple storage nodes. The multiple storage nodes directly calculate the IO capability information of the storage pools using the obtained IO metrics, which improves the calculation speed of IO capability information. Furthermore, since the multiple storage nodes themselves manage the storage pools, they can obtain the IO metrics of the managed storage pools faster and process the IO metrics faster than the cloud management platform, facilitating accurate calculation of the storage pool's IO capability. In this way, the cloud management platform does not need to calculate the IO capability of the storage pools but directly obtains the IO capability information of multiple storage pools provided by the multiple storage nodes to determine N target storage pools, reducing the data processing steps of the cloud management platform and improving its operating speed.
[0041] In one possible implementation, N≥2, the cloud management platform is specifically configured to send a second request to the plurality of storage nodes based on the data recovery request, the second request being used to instruct the creation of a new target logical volume belonging to the N target storage pools; the plurality of storage nodes are further configured to create a target logical volume across the N target storage pools based on the second request, the target logical volume including N logical spaces located within the N target storage pools; the cloud management platform is specifically configured to restore the backup data of the source logical volume to the N logical spaces within the N target storage pools based on the data recovery request.
[0042] In one possible implementation, the plurality of storage nodes are specifically configured to: obtain IO capability information of the N target storage pools based on the second request; determine the size of the i-th logical space to be created in the i-th target storage pool based on the IO capability information of the N target storage pools, where 1≤i≤N and i is a positive integer; and create the i-th logical space of the target logical volume in the i-th target storage pool according to the size of the i-th logical space, so as to create the target logical volume.
[0043] In this embodiment, multiple storage nodes providing multiple storage pools can obtain the IO capability information of each target storage pool to determine the size of the logical space to be created in each target storage pool within the target logical volume used for data recovery. For example, a larger proportion of the logical space about the target volume can be created in the target storage pool with stronger IO capability. For example, the larger the IO capability value of the target storage pool, the stronger its IO capability. In this way, the target storage pool with relatively stronger IO capability can be used to write a larger proportion of the backup data in the source volume, while the target storage pool with relatively weaker IO capability can be used to write a smaller proportion of the backup data in the source volume. This allows for the reasonable utilization of the IO capabilities of multiple target storage pools to write backup data occupying different proportions of the source volume to multiple target storage pools with different IO capabilities, thereby improving the overall recovery efficiency of data recovery across multiple storage pools and reducing the data recovery time.
[0044] In one possible implementation, the N target storage pools include a first storage pool and a second storage pool, the N logical spaces include a first logical space located within the first storage pool and a second logical space located within the second storage pool, and the backup data of the source logical volume includes first backup data and second backup data; the plurality of storage nodes are further configured to send metadata of the target logical volume to the cloud management platform, the metadata including a first matching relationship between a first access address of the first logical space and the first storage pool, and a second matching relationship between a second access address of the second logical space and the second storage pool; the cloud management platform is further configured to, based on the metadata and the data recovery request, restore the first backup data of the source logical volume to the first logical space within the first storage pool according to the first access address in the first matching relationship, and restore the second backup data of the source logical volume to the second logical space within the second storage pool according to the second access address in the second matching relationship.
[0045] In one possible implementation, the cloud management platform is specifically configured to: based on the metadata and the data recovery request, create multiple threads for writing backup data of the source logical volume to the N target storage pools, the multiple threads including a first thread and a second thread; based on the first thread, write the first backup data of the source logical volume to the first logical space in the first storage pool according to the first access address in the first matching relationship; during the process of writing the first backup data to the first logical space through the first thread, based on the second thread, write the second backup data of the source logical volume to the second logical space in the second storage pool according to the second access address in the second matching relationship.
[0046] The cloud management platform of each of the above embodiments can realize the functions and effects of the data recovery method of the first aspect or any possible embodiment of the first aspect, which will not be elaborated here.
[0047] Fourthly, embodiments of this application provide a computing device cluster. The computing device cluster includes at least one computing device, each of the at least one computing device including a processor and a memory. The processor of each computing device is configured to execute instructions stored in the memory of each computing device, causing the computing device cluster to perform a data recovery method executed by a cloud management platform in the first aspect or any possible implementation thereof.
[0048] Fifthly, embodiments of this application provide a computer storage medium storing one or more instructions, which, when executed by one or more computers, cause the one or more computers to implement the data recovery method of the first aspect or any possible implementation of the first aspect.
[0049] Sixthly, embodiments of this application provide a computer program product storing instructions that, when executed by a computer, cause the computer to implement the data recovery method of the first aspect or any possible implementation thereof. Attached Figure Description
[0050] Figure 1a A schematic diagram of a cloud system as an example;
[0051] Figure 1b A schematic diagram of a cloud system as an example;
[0052] Figure 2 This is a schematic diagram illustrating the data recovery process in related technologies, as an example.
[0053] Figure 3a This is a schematic diagram of the system architecture as an example.
[0054] Figure 3b A schematic diagram illustrating an exemplary data processing procedure;
[0055] Figure 4a A schematic diagram illustrating an exemplary data recovery process;
[0056] Figure 4b A schematic diagram illustrating an exemplary data recovery process;
[0057] Figure 5a This is a schematic diagram illustrating one application scenario as an example.
[0058] Figure 5bThis is a schematic diagram illustrating one application scenario as an example.
[0059] Figure 5c This is a schematic diagram illustrating one application scenario as an example.
[0060] Figure 6a This is a schematic diagram illustrating one application scenario as an example.
[0061] Figure 6b This is a schematic diagram illustrating one application scenario as an example.
[0062] Figure 7a This is a schematic diagram illustrating the structure of a cloud management platform as an example.
[0063] Figure 7b This is a schematic diagram illustrating the structure of a cloud service system as an example.
[0064] Figure 8 This is a schematic diagram of the structure of a computing device as an example.
[0065] Figure 9 This is a schematic diagram illustrating the structure of a computing device cluster as an example.
[0066] Figure 10 This is a schematic diagram illustrating the structure of a computing device cluster as an example. Detailed Implementation
[0067] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0068] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0069] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0070] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0071] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0072] Before describing the technical solutions of the embodiments of this application, a brief introduction to the technical background and technical terms involved in the embodiments of this application will be given first:
[0073] Public cloud is a cloud platform provided by a third-party public cloud provider to a wide range of individuals or businesses. In a public cloud, the hardware, software, and other infrastructure are owned and managed by the third-party public cloud provider.
[0074] A private cloud is a dedicated cloud platform provided to a single enterprise or organization. It can be operated internally by the respective enterprise or organization. Private clouds are primarily geared towards enterprise users and are also known as enterprise clouds.
[0075] Hybrid cloud refers to a cloud platform comprised of different cloud platforms. A hybrid cloud typically includes at least two cloud platforms, also known as a multi-cloud platform or multi-cloud. Optionally, a hybrid cloud integrates public and private clouds. For security reasons, some enterprise users prefer to store data in a private cloud but simultaneously desire access to the computing resources of a public cloud. In this context, hybrid clouds, which combine public and private clouds, are increasingly being adopted. Hybrid clouds combine and match public and private clouds to achieve optimal performance.
[0076] In this application's embodiments, the cloud used is a public cloud as an example. In other embodiments, it can also be a private cloud and / or a hybrid cloud, which is not limited in this application. That is to say, the data recovery method in the following embodiments of this application can also be applied to a private cloud or a hybrid cloud, which is not limited in this application.
[0077] The following is combined Figure 1a , Figure 1b The cloud system 10 of this application embodiment will be introduced.
[0078] like Figure 1a As shown, cloud system 10 may include public cloud 101 and one or more tenants (here, one tenant 102 is taken as an example).
[0079] The public cloud 101 may include a cloud management platform 103 and infrastructure 104 with communication connectivity.
[0080] The cloud management platform 103, also known as the cloud platform or simply the cloud management platform, is a software system used by cloud providers to provide cloud technology (also known as cloud computing) services and can be used to manage infrastructure 104.
[0081] Infrastructure 104 is the hardware equipment that provides cloud services. Infrastructure 104 may include multiple data centers (DCs) located in different geographical regions, with at least one data center in each region. Each data center may contain multiple physical servers, and each physical server can be used to support various cloud services. For example, physical servers can be bare metal servers; there are no restrictions here.
[0082] Cloud services may include computing services, storage services, virtual machine services, container services, network services, etc. Devices or functions accessible to tenant 102 upon logging into the cloud management platform 103 can all be considered cloud services provided by infrastructure 104.
[0083] Among them, such as Figure 1b As shown, the physical servers included in a data center can be compute nodes, storage nodes, and network nodes. For example, the infrastructure 104 shown in Figure 1 may include, for instance, compute nodes, storage nodes, and network nodes. Figure 1b The data center 200 shown may include, for example, compute nodes 11, network nodes 12, and storage nodes 13. Thus, infrastructure 104 may include compute nodes 11, network nodes 12, storage nodes 13, etc.
[0084] One or more computing nodes may have the functionality of the data recovery method provided in the embodiments of this application, and these computing nodes may also be named "management nodes" in subsequent embodiments.
[0085] In some embodiments, multiple computing nodes can form a computing cluster, multiple storage nodes can form a storage cluster, and multiple network nodes can form a network cluster. Multiple computing clusters, multiple storage clusters, and multiple network clusters can be deployed in the same or different data centers. The computing nodes can be computing servers, the storage nodes can be storage servers, and the network nodes can include, but are not limited to, network infrastructure such as routers and switches. Computing nodes and storage nodes can interact with each other through network nodes, which can perform data forwarding functions between computing nodes and storage nodes.
[0086] like Figure 1a As shown, the cloud management platform 103 provides an interface related to cloud services for tenant 102 (or a client) to remotely access cloud services. Tenant 102 can log in to the cloud management platform 103 through a pre-registered account and password on the cloud service access page, and after successful login, purchase and use the corresponding cloud services on the cloud service access page. Since the cloud management platform 103 is communicatively connected to the infrastructure 104, the cloud management platform 103 can provide various cloud services supported by the infrastructure 104 and purchased by tenant 102 to tenant 102 for use.
[0087] like Figure 1a The client shown refers to the terminal or browser on the terminal used by tenant 102 that can access cloud services. This terminal may include, but is not limited to, mobile phones, tablets, computers, personal computers (PCs), and devices in internet systems.
[0088] In some embodiments, cloud services may be deployed in a distributed manner on physical servers within a data center.
[0089] For example, storage services can be distributed across multiple storage nodes within a storage cluster, allowing multiple storage nodes to work together to support the storage service.
[0090] The following describes the storage services provided by infrastructure 104 in this application embodiment.
[0091] Storage services can be categorized into block storage services (such as Elastic Volume Service (EVS)), object storage services (OBS), and file storage services (such as Elastic File Service (SFS)).
[0092] The basic components of object storage services are buckets and objects. A bucket is a container for storing objects in OBS. Each bucket has its own storage category, access permissions, and region, and users locate buckets on the internet using their access domain names. An object is the basic unit of data storage in OBS. An object is actually a collection of a file's data and its related attribute information, specifically including three parts: key, metadata, and data.
[0093] Block storage services are storage services that use data blocks as the basic unit for data storage and access.
[0094] The following sections will introduce the block storage service from the perspectives of logic, access control, snapshots and backups, and usage.
[0095] Logical level:
[0096] Block storage services can include multiple storage pools, where the storage resources of one or more storage nodes (physical servers) can be abstracted into a single storage pool. For example, a storage pool can be understood as a logical storage unit of one or more storage nodes, and thus a storage pool can be mapped to the storage resources of one or more storage nodes. Block storage services can be deployed in a distributed manner across multiple storage nodes to manage the one or more storage pools mapped to those nodes, simplifying the management of storage resources across multiple storage nodes and improving the utilization and flexibility of storage resources.
[0097] Each storage pool can contain one or more logical volumes. A logical volume is a logical storage unit partitioned from a storage pool. Each logical storage unit within a storage pool can also be called a Logical Unit Number (LUN). Furthermore, multiple logical volumes within a storage pool can be accessible to multiple tenants; in other words, different tenants can access different logical volumes within the same storage pool to ensure data isolation.
[0098] Logical volumes are tenants (e.g.) Figure 1a The interface shown is for tenant 102 or its client to access storage resources within storage nodes in the public cloud. Each logical volume can be treated as an independent hard drive. Tenant 102 or its client can perform operations such as partitioning, formatting, installing operating systems, backup, and data recovery on logical volumes. In this way, tenant 102 can access data stored within storage nodes in the public cloud 101 (or private cloud, hybrid cloud) as if it were a local disk.
[0099] In addition, logical volumes can be configured with different performance and redundancy settings; and a single storage pool can be supported by multiple storage nodes for its data storage, allowing a single logical volume within a single storage pool to distribute data across multiple storage nodes.
[0100] A logical volume can comprise multiple data blocks of the same size. When accessing the logical volume, tenant 102 reads and writes data in units of data blocks. Each data block in the block storage service is a fixed-size unit of data; for example, each block can be 10 megabytes (MB). Therefore, the data block is the basic unit of storage in the block storage service.
[0101] Access control layer:
[0102] Block storage services offer a high degree of access control. Block storage services can configure virtual instances (such as virtual machines or containers) to access specific logical volumes to ensure the security and isolation of stored data.
[0103] Snapshot and backup level:
[0104] Block storage services typically offer snapshot functionality, allowing tenants to create data copies of data in logical volumes (e.g., creating them on a schedule, with no restrictions on the specific data copy creation strategy) to facilitate backup and recovery of data in logical volumes.
[0105] Usage process level:
[0106] The process of using block storage services typically involves the following steps: 1. Configuring the storage pool: Administrators first need to configure the storage pool in the storage system, including selecting appropriate storage nodes, setting the Redundant Array of Independent Disks (RAID) level, and other performance parameters. This allows the storage pool to be mapped to one or more selected storage nodes. 2. Creating logical volumes: Block storage services can create one or more logical volumes in the storage pool and set the size and access control of these volumes, ensuring secure access to logical volumes in multi-tenant scenarios. 3. Allocating logical volumes: Block storage services can allocate the created logical volumes to virtual instances (e.g., virtual machines or containers) on compute nodes that require storage resources. 4. Virtual instance configuration: The allocated logical volume is mounted on the tenant's virtual instance. The tenant can then format and install file systems on the mounted logical volume through the virtual instance. 5. Data operations: The tenant's client can perform data read and write operations on the logical volume mounted on the virtual instance through the virtual instance running on the compute node, accessing the logical volume as if it were a local hard drive.
[0107] Block storage services provide high-performance, highly reliable, and elastically scalable data storage solutions in cloud environments, making them well-suited for high-load and high-concurrency data access scenarios.
[0108] To ensure the security and reliability of stored data in the cloud environment, the block storage service provides the function of taking snapshots and backing up data in logical volumes. The block storage service can back up the data in the logical volume regularly (e.g., daily) so that in the event of a data disaster such as data corruption or loss in the logical volume, the data in the logical volume can be recovered using the backed-up data.
[0109] Before introducing the relevant technologies and the data recovery solutions of the embodiments of this application, the following technical names involved will first be explained and described:
[0110] The source volume is the logical volume for which data recovery is required in block storage services.
[0111] A backup set is a logical storage unit used to store a copy of the data from the source volume (it can be a logical volume in a block storage service or a bucket in OBS, there are no restrictions).
[0112] In block storage services, a target volume is a logical volume used to store recovery data from the source volume (specifically, a copy of the source volume's data read from a backup set).
[0113] Source storage pool: A storage pool containing source volumes.
[0114] The target storage pool is the storage pool containing the target volume.
[0115] The following section introduces data recovery solutions from relevant technologies.
[0116] Figure 2 An exemplary schematic diagram of a data recovery scheme in the related art is shown.
[0117] like Figure 2 As shown, the storage pool 1 within the block storage service 201 contains logical volumes, such as volume V1. Volume V1 includes P data blocks, such as block 1, block 2, block 3, etc. The bucket of the object storage service 202 (represented here as backup set 1) stores a data copy of volume V1. Specifically, object 1 in backup set 1 is a data copy of block 1 in volume V1, object 2 in backup set 1 is a data copy of block 2 in volume V1, object 3 in backup set 1 is a data copy of block 3 in volume V1, object 4 in backup set 1 is a data copy of block 4 in volume V1, and so on. Regardless of whether any data has been written to the P data blocks in volume V1, each of the P data blocks has a corresponding data copy stored in backup set 1. Even if there are empty data blocks in volume V1, the empty data blocks in the source volume must be restored during the data recovery process for volume V1.
[0118] like Figure 2As shown, block 1 and block 3 in volume V1 are corrupted. The block storage service 201 is visible to the tenant as volume V1 (e.g., drive D), but the tenant cannot know which data block in volume V1 is corrupted. They can only know that volume V1 is corrupted. In response to the tenant's client's data recovery request for volume V1, the block storage service 201 can read data copies of all data blocks in backup set 1 according to the mapping relationship between data blocks in volume V1 and objects in backup set 1 (e.g., the mapping relationship shown by the arrow), and write them to the corresponding data blocks in volume V1. Taking data recovery of block 1 as an example, the block storage service 201 can read the storage data of object 1 (specifically, an example of a data copy of block 1) from the storage node corresponding to the object storage service 202 based on the metadata of object 1 in backup set 1. Then, according to the metadata of block 1 in volume V1 (such as the starting address and offset address of block 1), the storage data of object 1 is written to the corresponding address in the storage node corresponding to block 1, thereby restoring the data copy of block 1 to volume V1. Similarly, the data copy of volume V1 in backup set 1 can be used to realize the data recovery of all data blocks (such as P data blocks) in volume V1.
[0119] exist Figure 2 In the related technologies shown, when block storage services perform data recovery on a source volume (e.g., volume V1) within a storage pool, they write a copy of the data from a backup set (e.g., backup set 1 of volume V1) to the source volume for data recovery. However, if the input / output (IO) performance (e.g., data read / write rate) of the storage pool where the source volume resides (i.e., the source storage pool) is poor, the time required to write the data copy from the source volume to the source storage pool will be long, resulting in a long data recovery time and low data recovery efficiency.
[0120] Furthermore, since the source storage pool where the source volume resides is unique, when writing a copy of the source volume's data to the source storage pool, it can only be done through a single task corresponding to that source storage pool. However, the processing capacity of a single task is limited, which reduces the efficiency of data recovery.
[0121] Furthermore, in scenarios where the source volume for data recovery comprises multiple logical volumes, while concurrent data recovery across multiple logical volumes can improve recovery efficiency, the recovery time of a single logical volume within these volumes still impacts the overall recovery time. As mentioned above, data recovery on a single source volume is performed on a block-by-block basis, and the size of each block is fixed. The larger the data volume in the source volume, the more blocks it contains. Even with a robust source storage pool's I / O capabilities, writing massive amounts of data block copies to the same storage pool (here, the source storage pool) can still lead to prolonged data recovery times. Therefore, the size of the source volume also affects the data recovery duration.
[0122] Therefore, block storage services in related technologies suffer from long recovery times and low efficiency when performing data recovery on logical volumes.
[0123] To address the aforementioned technical problems, this application provides a data recovery method and a data recovery apparatus to quickly recover data from a logical volume mounted on a virtual machine in the event of a storage data failure.
[0124] Figure 3a The diagram illustrates the system architecture of this application as an example.
[0125] Figure 3a The system structure diagram shown can be compared with... Figure 1a and Figure 1b This combination allows client cluster 800 and data center cluster 1000 to interact via cloud management platform 103. However, Figure 3a The system structure shown is not limited to the one combined with Figure 1a and Figure 1b The structure shown. (As illustrated) Figure 3a As shown, the system may include, but is not limited to, client cluster 800 and data center cluster 1000.
[0126] Client cluster 800 may include multiple clients, such as clients 102 to 10n, where n is a positive integer.
[0127] Each client in the client cluster 800 is a client of the tenant who purchased the cloud service for this application.
[0128] Continue to refer to Figure 3a Data center cluster 1000 may include, but is not limited to: management cluster 400, computing cluster 500, storage cluster 300 and storage cluster 301.
[0129] The management cluster 400 may include one or more management nodes, such as management node 1 to management node m, where m is a positive integer.
[0130] One or more management nodes can be set up in one or more data centers; there are no restrictions here.
[0131] In this embodiment, each management node may be deployed with a data recovery service 900, which can be used to recover data that has failed in the storage node. The specific functions will be described in the following embodiments.
[0132] In some embodiments, such as Figure 3a As shown, each management node in the management cluster 400 can deploy a complete data recovery service 900.
[0133] In other embodiments, the same data recovery service 900 can be distributed across multiple management nodes, allowing multiple management nodes to jointly support a single data recovery service 900.
[0134] The management node can be implemented as a physical server.
[0135] In other embodiments, the data recovery service 900 described above can also operate in, for example... Figure 1a The cloud management platform 103 shown can implement the data recovery methods described in the various embodiments of this application.
[0136] Continue to refer to Figure 3a The computing cluster 500 may include one or more computing nodes, such as computing node 1 to computing node n, where n is a positive integer.
[0137] One or more compute nodes can be located in one or more data centers; there are no restrictions here.
[0138] Each compute node may include one or more virtual instances (such as containers or virtual machines). Taking virtual machines as an example, compute node 1 may include virtual machine 1 and virtual machine 2, compute node 2 may include virtual machine 3 and virtual machine 4, and so on. The same applies to other compute nodes, which will not be elaborated here.
[0139] It should be understood that this application does not limit the number of virtual machines within a single computing node, nor does it limit whether the number of virtual machines on different computing nodes is the same.
[0140] Continue to refer to Figure 3a The storage cluster 300 may include multiple storage nodes, such as storage node 1 to storage node 20. This application does not limit the number of storage nodes in the storage cluster 300.
[0141] Multiple storage nodes within storage cluster 300 (here, storage nodes 1 to 20) collectively provide block storage service 600. In other words, as mentioned earlier, block storage service 600 can be distributed across multiple storage nodes. These multiple storage nodes can belong to one or more storage clusters; there are no restrictions here.
[0142] The storage nodes providing this block storage service 600 can be located in one or more data centers; there are no restrictions here.
[0143] The storage resources of storage nodes 1 and 2 can be abstracted as storage pool 1, the storage resources of storage nodes 3 to 5 can be abstracted as storage pool 2, the storage resources of storage nodes 6 to 10 can be abstracted as storage pool 3, and the storage resources of storage nodes 11 to 20 can be abstracted as storage pool 4. The block storage service 600 can manage storage pools 1 to 4 to manage the storage resources of storage nodes 1 to 20.
[0144] For example, the block storage service 600 has created two logical volumes in storage pool 1, namely volume V11 and volume V12, and has created volume V21 and volume V22 in storage pool 2, and has created volume V31 in storage pool 3. No logical volumes have been created in storage pool 4.
[0145] Continue to refer to Figure 3a The storage cluster 301 may include multiple storage nodes, such as storage nodes 21 to 40. This application does not limit the number of storage nodes in the storage cluster 301.
[0146] Multiple storage nodes (here, storage nodes 21 to 40) within storage cluster 300 jointly provide object storage service 700. In other words, object storage service 700 can be distributed across multiple storage nodes, which can belong to one or more storage clusters, without any restrictions.
[0147] The storage nodes providing the object storage service 700 can be located in one or more data centers, without any restrictions.
[0148] The storage resources of storage node 21 can be abstracted as storage bucket 1, while the storage resources of the other storage nodes in storage cluster 301 are abstracted as storage buckets 2 to 10, without any restrictions.
[0149] The object storage service 700 can manage buckets 1 to 10 to manage the storage resources of storage nodes 21 to 40.
[0150] In this embodiment, the data copy of volume V11 in storage pool 1 is backup set 1 in storage bucket 1. That is, in order to ensure the recovery of stored data in case of failure, the data copy of each logical volume in the storage pool can be periodically written to the storage bucket managed by the object storage service 700 so that in the event of failure of a logical volume in the storage pool, the data copy in the storage bucket can be used to recover the data of the failed logical volume.
[0151] Of course, the storage service used to manage data copies of logical volumes within the storage pool is not limited to object storage services. It can also be implemented by file storage services, which save data copies of logical volumes in the form of files. There are no restrictions here.
[0152] Continue to refer to Figure 3a Client cluster 800 can interact with computing cluster 500, computing cluster 500 can interact with storage cluster 300; client cluster 800 can interact with management cluster 400; management cluster 400 can interact with computing cluster 500, storage cluster 300, and storage cluster 301 respectively.
[0153] Let's take tenant's client 102 as an example to illustrate the interaction process between the nodes.
[0154] Tenant purchases on client 102 are as follows Figure 3a The data recovery service 900, virtual machine service, and block storage service 600 are shown. For example, the virtual machine service is as follows: Figure 3a The virtual machines 1 and 2 in compute node 1, and virtual machines 3 and 4 in compute node 2 are shown.
[0155] Specifically, the tenant purchased a certain size of storage resources (e.g., 1300G, the exact size may vary depending on needs, and is not limited here) of block storage service 600. For example, volume V11 (500G) and volume V12 (800G) in storage pool 1 correspond to the storage resources purchased by the tenant. Virtual machine 1 is then mounted with volume V11 and volume V12.
[0156] In this way, client 102 can run virtual machine 1 in compute node 1 and access volumes V11 and V12 mounted on virtual machine 1 through virtual machine 1, thereby realizing data read and write to the storage nodes mapped by volumes V11 and V12. In this way, tenants can use the storage resources provided by the storage cluster 300 of this application through block storage service 600.
[0157] For example, in the event of a data failure in volume V11, client 102 can communicate with management node 1 to send a data recovery request. Then, the data recovery service 900 in management node 1 can interact with the block storage service 600 provided by storage cluster 300 and the object storage service 700 provided by storage cluster 301 to restore and write the backup set 1 of volume V11 in storage bucket 1 to at least one logical volume in a target storage pool, such as storage pool 4, or storage pool 3 and storage pool 4.
[0158] The target volume has a higher data read / write capability than storage pool 1. Therefore, the data recovery time of the logical volume is not limited by the data read / write capability of the source storage pool (here, storage pool 1). Even if the source storage pool is currently handling a large number of IO requests, the data recovery time of the source volume (here, volume V11) can be shortened and the data recovery efficiency of the source volume can be improved by writing a copy of the data of the source volume to the target storage pool with stronger data read / write capability.
[0159] The following uses a management node to implement the data recovery method of this application as an example to illustrate the data recovery process of this application. In other embodiments, the data recovery methods of the various embodiments described below can also be implemented by, for example, a management node. Figure 1a The cloud management platform 103 shown is used to implement this. Specifically, the operations performed by the management node in the methods of the following embodiments can be replaced by the cloud management platform 103. The implementation principle and technical effect of the methods in this application embodiment performed by the cloud management platform 103 are the same as those performed by the management node, and will not be described in detail here.
[0160] When the method of this application embodiment is executed by the cloud management platform 103, this application embodiment may also provide a cloud service system, which may include, for example, Figure 1a The cloud management platform 103 shown includes multiple storage nodes (e.g., ...). Figure 3a The storage cluster 300 shown provides block storage services.
[0161] Figure 3b This is a schematic diagram illustrating the data processing process of the management node of this application as an example.
[0162] Figure 3b It can be combined with the above Figure 1a , Figure 1b , Figure 3a However, it is not limited to combining the above. Figure 1a , Figure 1b , Figure 3a The architecture.
[0163] like Figure 3b As shown, the process may include the following steps:
[0164] Step S101: Management node 1 receives a data recovery request.
[0165] The data recovery request may include, but is not limited to, information indicating the source volume to be recovered, such as the source volume identifier (ID) or the name of the source volume.
[0166] For example, please refer to Figure 3a Management node 1 can receive a data recovery request from client 102, which may include the volume ID (e.g., V11) of volume V11.
[0167] The number of source volumes can be one or more, which is not limited here. Management node 1 can perform large-scale data recovery on multiple source volumes in batches based on a data recovery request, or it can perform data recovery on the source volume corresponding to each data recovery request. The principle of the data recovery process for each source volume is the same. The following explanation uses one source volume as an example. When there are multiple source volumes, the implementation principle of the data recovery method in this application is similar, and will not be repeated here.
[0168] As mentioned above, the source volume's data can be backed up periodically (e.g., daily), allowing the source volume to have multiple backup sets, with one backup set corresponding to each day. In some embodiments, the data recovery request may also include information indicating the backup set of the source volume, such as a backup set ID. This allows tenants to specify a particular backup set for the failed source volume for data recovery.
[0169] Alternatively, in some embodiments, the management node 1 may determine the backup set for data recovery of the source volume according to a predetermined strategy, such as selecting the backup set obtained from the most recent backup of the source volume to determine the backup set ID.
[0170] Step S103: Based on the data recovery request, management node 1 determines N target storage pools.
[0171] Where N is a positive integer greater than or equal to 1, that is, the number of target storage pools determined by management node 1 can be one or more, without any restriction.
[0172] Among them, the N target storage pools are storage pools selected from multiple storage pools based on the IO capability information of multiple storage pools.
[0173] In this context, the I / O capability of each storage pool within multiple storage pools indicates its data recovery capability. In other words, a higher I / O capability means faster recovery of backup data from logical volumes to that storage pool, resulting in shorter recovery times and higher overall data recovery capability. Conversely, a lower I / O capability means slower recovery of backup data from logical volumes, resulting in slower recovery times and lower overall data recovery capability. Each of the multiple storage pools provided by multiple storage nodes can have its own I / O capability information.
[0174] As for whether the N target storage pools are determined by management node 1 from multiple storage pools, or by multiple storage nodes providing those multiple storage pools from multiple storage pools, there is no restriction here.
[0175] In one possible implementation of step S103, the management node 1 can receive information indicating N target storage pools (e.g., ID information of the N target storage pools) from multiple storage nodes based on a data recovery request. The management node 1 can then determine the N target storage pools from among the multiple storage pools based on this received information. For example, the N target storage pools could be the one with the highest IO capacity value among the multiple storage pools, or the two storage pools with the highest and second-highest IO capacity among the multiple storage pools. In other words, multiple storage nodes can use the IO capacity information of multiple storage pools to determine the N target storage pools from among the multiple storage pools.
[0176] In one possible implementation, management node 1 can determine N target storage pools from the multiple storage pools based on the data recovery request and the IO capability information of the multiple storage pools. These multiple storage pools are provided by multiple storage nodes (e.g., storage cluster 300). That is, management node 1 can determine the N target storage pools from the multiple storage pools using the data recovery request and the IO capability information of the multiple storage pools, rather than having the multiple storage nodes determine the N target storage pools.
[0177] As an implementation example, management node 1 can obtain information indicating the IO capabilities of the aforementioned multiple storage pools (e.g., the IO capability value of each storage pool) based on a data recovery request. Then, management node 1 can determine N target storage pools from the multiple storage pools based on the obtained information (e.g., the IO capability value of each storage pool in the multiple storage pools).
[0178] Of course, step S103 may include other implementation examples, which are not limited here.
[0179] Furthermore, there are no restrictions here regarding whether the multiple storage pools provided by multiple storage nodes can be all or part of the storage pools managed by the block storage service.
[0180] For example, combined with Figure 3a These multiple storage pools can provide all the storage pools for storage cluster 300, such as storage pool 1 to storage pool 4.
[0181] In addition, the average IO capability of the N target storage pools is higher than the IO capability of the source storage pool, which is the storage pool to which the source volume belongs among the above multiple storage pools.
[0182] For example, if N=1, then the IO capability of a target storage pool determined by management node 1 is higher than the IO capability of the source storage pool.
[0183] For example, if N≥2, then the average IO capability of the N target storage pools is higher than the IO capability of the source storage pool. For instance, at least one of the N target storage pools has a higher IO capability than the source storage pool, making the average IO capability of the N target storage pools higher than the IO capability of the source storage pool.
[0184] As an implementation example, please refer to Figure 3a For example, the IO capacity values of storage pool 1 to storage pool 4 are 60, 70, 80 and 90 respectively, and storage pool 1 is the source storage pool.
[0185] The average IO capacity of N target storage pools can be the average of the IO capacity of N target storage pools (for example, if storage pool 3 and storage pool 4 are target storage pools, then the average IO capacity is (80+90) / 2=85).
[0186] In one example, N=1, then the target storage pool can be any one of storage pool 2 to storage pool 4.
[0187] In one example, N ≥ 2, and the N target storage pools can be at least two of storage pools 2 to 4. Alternatively, the N target storage pools can be storage pool 1 and storage pool 4 respectively. In this case, the average IO capacity of the N target storage pools is 75, which is also higher than the IO capacity of storage pool 1. Alternatively, the storage cluster 300 may also provide storage pool 5. For example, if storage pool 5 has an IO capacity of 50, then the N target storage pools can also be storage pool 4 and storage pool 5 respectively. In this case, the average IO capacity of the two storage pools is 70, which is also higher than the IO capacity of storage pool 1.
[0188] In step S104, management node 1 restores the backup data of the source volume to the target volume included in the N target storage pools based on the data recovery request.
[0189] The target volume and the source volume are two different logical volumes; for example, the volume ID of the target volume is different from the volume ID of the source volume.
[0190] The target volume belongs to N target storage pools.
[0191] When N=1, the target volume is a logical volume within a target storage pool.
[0192] When N≥2, the N target storage pools share a single target volume, meaning that the single target volume exists across N target storage pools.
[0193] Before restoring the backup data from the source volume to the target volume, the size of the free storage resources of the target volume must be greater than or equal to the size of the source volume.
[0194] In some embodiments, the target logical volume may be a logical volume created before step S104. The created logical volume may be an empty volume with no data written to it, or it may be a volume with data written to it. However, in order to ensure data isolation between different tenants, the created logical volume is the logical volume of the tenant of client 102, and not the logical volume of other tenants.
[0195] This reduces the time overhead of creating new logical volumes during the data recovery process, thereby further improving data recovery efficiency.
[0196] In some embodiments, the target volume may also be provided by multiple storage nodes (e.g., multiple storage pools) during the implementation of step S104. Figure 3a The logical volume created by the storage cluster 300 shown.
[0197] In the aforementioned technologies, data recovery from the source volume is performed directly within the source volume, specifically by restoring all data from the source volume's backup set to the source volume. Thus, given a fixed source volume, the target storage pool is also fixed, meaning the source and target storage pools are identical. This can lead to excessively long data recovery times if the source storage pool has poor I / O throughput.
[0198] However, in this embodiment, the N target storage pools determined by the management node based on the data recovery request are at least one storage pool determined based on the IO capability information of multiple storage pools provided by multiple storage nodes. Furthermore, the average IO capability of the selected N target storage pools is higher than the IO capability of the source storage pool where the source volume resides; the data write rate of the storage pool with higher IO capability is higher than that of the storage pool with lower IO capability. Thus, the overall time for restoring the backup data of the source volume to the target volumes within the N target storage pools is significantly reduced, thereby shortening the data recovery time of the source volume and improving the data recovery efficiency of the source volume.
[0199] Example 1
[0200] In this example 1, we will combine Figure 4a Taking the above N target storage pools as one target storage pool (i.e., N=1) as an example, the implementation process of the data recovery method of this application will be described.
[0201] Figure 4a It can be combined with the above Figure 1a , Figure 1b , Figure 3a , Figure 3b However, it is not limited to combining the above. Figure 1a , Figure 1b , Figure 3a , Figure 3b The architecture shown.
[0202] In the introduction Figure 4a At that time, it will be combined with Figures 5a to 5c , Figure 6a The application scenarios shown are used to illustrate this. Figure 4a The implementation process of the data recovery method shown is as follows, however, Figure 4a The implementation process of the data recovery method shown is not limited to application. Figures 5a to 5c as well as Figure 6a The scene shown.
[0203] like Figure 4a As shown, the process may include the following steps:
[0204] Step S201: Management node 1 receives a data recovery request from client 102.
[0205] The data recovery request may include, but is not limited to, source volume ID, backup set ID, and virtual machine information.
[0206] This virtual machine information indicates the virtual machine mounted on the source volume indicated by the source volume ID. For example, the virtual machine information is the virtual machine ID.
[0207] like Figure 3aAs described in the above related embodiments, client 102 purchased virtual machines 1 to 4. Virtual machine 1 is mounted with volumes V11 and V12. Therefore, the source volume for data recovery in this failure is volume V11. The virtual machine information is the virtual machine ID of virtual machine 1, and the backup set ID is the backup set ID of volume V11. The backup set has a data copy of volume V11.
[0208] The following is combined Figure 5a , Figure 5b and Figure 6a The application scenarios are used to illustrate the methods of the embodiments of this application.
[0209] After step S201, as Figure 5a As shown, after the tenant purchased, such as Figure 3a He Ru Figure 6a After the data recovery service 900 shown, since the tenant has purchased virtual machines 1 to 4, the tenant can customize the configuration to have data recovery functionality for which virtual machines(s) they are using.
[0210] like Figure 5a As shown in (1), the tenant logs into the cloud management platform 103 through client 102, which allows client 102 to display the following: Figure 5a (1) The data recovery settings interface 201 shown includes, but is not limited to: a list of virtual machines purchased by the tenant, namely virtual machines 1 to 4, and an OK button 203; then, the tenant clicks on the option box 202 on the right side of virtual machine 1, and the cloud management platform 103 can respond to the click operation and update the data recovery settings interface 201 as shown. Figure 5a (2) The data recovery settings interface 201 is shown.
[0211] like Figure 5a (2) As shown, option box 202 is in a gray selected state. Then, the tenant clicks the OK button 203. In this way, management node 1 can respond to the tenant's click of the OK button 203 to implement data protection function for the logical volume mounted on virtual machine 1. In the event of a data failure in the logical volume mounted on virtual machine 1, the logical volume mounted on virtual machine 1 can be restored through the data recovery service 900 provided by management node 1 to restore the data of any failed logical volume mounted on virtual machine 1.
[0212] Optionally, the data recovery settings interface 201 may further provide a list of logical volumes mounted on the selected virtual machine 1, such as volume V11 and volume V12, so that the tenant can further select a specific logical volume (e.g., volume V11) mounted on the virtual machine 1 to use the data recovery service 900 for data recovery protection.
[0213] It should be understood that the service provided by management node 1 to implement the data recovery method of this application is not limited to the name of data recovery service. In addition, it can be a separate service or a sub-service embedded in a service already provided by the management node. This application does not limit this.
[0214] exist Figure 5a Following the process shown, if the logical volume mounted on virtual machine 1 fails, the tenant can trigger a data recovery request for the failed logical volume via client 102. However, for... Figure 5a In the scenario shown, even if the logical volume mounted on virtual machine 2 fails, a data recovery request for the failed logical volume will not be triggered if the virtual machine (e.g., virtual machine 2) does not have data protection enabled.
[0215] For details, please refer to Figure 5b In the event of a failure in volume V11, client 102 may display the following: Figure 5b (1) The data recovery interface 204 shown. The data recovery interface 204 may include, but is not limited to: the ID list information of the backup sets of volume V11 mounted by virtual machine 1 and the backup time information corresponding to each backup set.
[0216] Volume V11 was backed up on August 30, 2024, August 30, 2024, and September 1, 2024, respectively, to obtain Backup Set 1 (backed up on August 30, 2024), Backup Set 2 (backed up on August 31, 2024), and Backup Set 3 (backed up on September 1, 2024). Data writes to Volume V11 may occur during these three days, making the data in the three backup sets potentially different. However, if the tenant did not write any data to Volume V11 during these three days, the data in the three backup sets will be identical; this application does not impose any restrictions on this.
[0217] like Figure 5b (1) As shown, when a tenant selects radio button 205, the cloud management platform 103 can respond to the tenant's operation and restore the data recovery interface from the previous screen. Figure 5b (1) Updated to: Figure 5b (2) The data recovery interface 204 shown updates the radio button 205 to a grayed-out selected state; then, the tenant clicks the "Start Data Recovery" button 206, and the cloud management platform 103 can receive the data recovery request from the client 102, which is triggered by clicking the "Start Data Recovery" button 206. The cloud management platform 103 can send the data recovery request to management node 1. In this way, management node 1 can receive the data recovery request. The data recovery request includes the source volume ID (specifically volume V11), backup set ID (e.g., backup set 1), and virtual machine information (specifically virtual machine 1).
[0218] In addition, such as Figure 5b As shown in (3), the cloud management platform 103 can change the display interface of the client 102 from... Figure 5b (2) The data recovery interface 204 shown in the figure has been updated to be as follows: Figure 5b (3) The data recovery interface 204 is shown. Figure 5b As shown in (3), during the process of data recovery of volume V11 by management node 1 in this application embodiment, the data recovery interface 204 displayed by client 102 can display the text "Volume V11 of virtual machine 1 is being recovered" to remind tenants that data recovery of volume V11 is being performed.
[0219] In some embodiments, data recovery time for a 100GB source volume is typically in the minutes.
[0220] As a concrete example, please refer to Figure 6a After client 102 triggers a data recovery request, virtual machine 1 has mounted not only volume V12 but also volume V11. However, data blocks 1 and 2 in volume V11 have failed. The tenant, management node, storage node, or virtual machine 1 does not know which data block in volume V11 has failed. Therefore, data recovery of volume V11 is required. Thus, the data recovery service 900 in management node 1 can receive the data recovery request from client 102. This data recovery request can carry the following information: virtual machine 1, volume V11, and backup set 1.
[0221] And such as Figure 6a As shown, before client 102 sends a data recovery request, a copy of the data in volume V11 has been written to backup set 1 in bucket 1 managed by object storage service 700.
[0222] For example, volume V11 has 200 data blocks, which are as follows: Figure 6a As shown in Block 1, Block 2, Block 3, Block 4..., Object 1 in Backup Set 1 is a data copy of Block 1, Object 2 is a data copy of Block 2, Object 3 is a data copy of Block 3, Object 4 is a data copy of Block 4... and so on. All data blocks in Volume V11 have corresponding data copies stored as objects in Bucket 1.
[0223] In this fault recovery scenario, the tenant requests to restore data to volume V11 using backup set 1, that is, to use the data copy in backup set 1 as the recovery data for volume V11.
[0224] Back Figure 4a After step S201, as shown by the dashed arrow, management node 1 may optionally execute steps S301 and S302.
[0225] In step S301, management node 1 may respond to the data recovery request by instructing virtual machine 1 to mount the source volume;
[0226] Step S501: Virtual machine 1 can mount the source volume.
[0227] As an example, please refer to Figure 6a In response to a data recovery request, the data recovery service 900 in management node 1 can instruct virtual machine 1 to mount volume V11. Then, in response to the instruction of the data recovery service 900, virtual machine 1 can unmount the originally mounted volume V11 as shown by the dashed arrow, so that there is no mounting relationship between volume V11 and virtual machine 1. In this way, virtual machine 1 only mounts volume V12.
[0228] In step S302, management node 1 may respond to the data recovery request by instructing storage cluster 300 to delete the source volume.
[0229] Step S401: Storage cluster 300 can delete the source volume.
[0230] As an example, please refer to Figure 6a In response to a data recovery request, the data recovery service 900 in management node 1 can instruct the block storage service 600, which is distributed across storage cluster 300, to delete volume V11. Thus, the block storage service 600 can delete volume V11 from storage pool 1 based on the instruction from the data recovery service 900. The specific implementation of deleting volumes from storage pools by storage clusters or storage nodes can refer to existing technologies or future volume deletion logic; no restrictions are imposed here.
[0231] For example, when the block storage service 600 deletes the source volume in step S301, it can format the source volume or mark the source volume as to be deleted. It can first not delete the data in the source volume, and then delete the data in the source volume later. There is no restriction here.
[0232] In some embodiments, management node 1 may instruct storage cluster 300 to delete the source volume after confirming that virtual machine 1 has successfully unmounted the source volume, so as to prevent tenant client 102 from continuing to write data to volume V11 through virtual machine 1 during the data recovery process of volume V11.
[0233] Continue back Figure 4a After S201, such as Figure 4a As shown, the process may also include:
[0234] In step S303a, management node 1 responds to the data recovery request and obtains the IO capability information of multiple storage pools.
[0235] This application does not restrict the execution order between S303a and S302. Management node 1 can obtain the IO capability value of each storage pool after the source volume is deleted from the source storage pool. Thus, regarding... Figure 6a The IO capacity value of storage pool 1 shown is the IO capacity value of storage pool 1 after volume V11 is deleted from storage pool 1. Alternatively, management node 1 can also obtain the IO capacity value of each storage pool before the source volume is deleted from the source storage pool. Thus, regarding... Figure 6a The I / O capacity value of storage pool 1 shown is the I / O capacity value of storage pool 1 (which has volumes V11 and V12) before volume V11 is deleted from storage pool 1. In some embodiments, the change in the I / O capacity value of the source storage pool is small before and after the source volume is deleted from the source storage pool.
[0236] The multiple storage pools in step S303a can be all or some of the storage pools provided by storage cluster 300; there is no restriction here.
[0237] Using these multiple storage pools as an example Figure 3a The storage cluster 300 shown here provides all the storage pools, taking storage pool 1 to storage pool 4 as examples.
[0238] In one possible implementation 1, when the management node 1 executes S303a, the management node 1 can receive IO capability information from multiple storage pools of the storage cluster 300 based on the data recovery request.
[0239] like Figure 4a As shown, optionally, in step S402a, the storage cluster 300 can send the IO capability information of multiple storage pools to the management node 1, so that the management node 1 can obtain the IO capability information of multiple storage pools.
[0240] In this way, the management node can directly obtain the IO capability information of each storage pool from the storage cluster 300, thereby reducing the computational load on the management node.
[0241] As an example of implementing the above-described embodiment 1, such as Figure 6a As shown, the data recovery service 900 can instruct the block storage service 600 to report the IO capability information of multiple storage pools based on the data recovery request; then, the block storage service 600 can obtain the IO metrics of the multiple storage pools in response to the instruction of the data recovery service 900; then, the block storage service 600 can determine the IO capability information of the multiple storage pools based on the IO metrics of the multiple storage pools; finally, the block storage service 600 can report the IO capability information of the multiple storage pools to the data recovery service 900.
[0242] When the data recovery service 900 instructs the block storage service 600 to report the IO capability information of multiple storage pools, the data recovery service 900 can specify the storage pool, for example, instruct the storage cluster 300 to report the IO capability information of each of the specified storage pools 1 to 3. In this way, the block storage service 600 does not need to report the IO capability information of the storage pool 4 it manages.
[0243] Alternatively, when instructing the block storage service 600 to report the IO capability information of multiple storage pools, the data recovery service 900 may not specify the storage pool. In this way, the block storage service 600 may report the IO capability information of all the storage pools it manages (e.g., storage pool 1 to storage pool 4) to the data recovery service 900.
[0244] Alternatively, when instructing the block storage service 600 to report the IO capability information of multiple storage pools, the data recovery service 900 may not specify a storage pool. However, the block storage service 600 may select some storage pools from all managed storage pools (e.g., storage pool 1 to storage pool 4) according to a preset strategy and report the IO capability information of the selected storage pools to the data recovery service 900.
[0245] The IO metric may include, but is not limited to, at least one of the following: IO throughput and IO utilization.
[0246] For example, IO throughput can be the data read / write rate of a storage pool.
[0247] When determining the IO throughput capacity of a storage pool, the block storage service 600 can base its decision on multiple storage nodes mapped to that storage pool (e.g., ...). Figure 3a The storage pool 1 shown is mapped to storage node 1 and storage node 2, and its I / O throughput is determined by the I / O capacity of each storage node. The I / O throughput of a storage node is the amount of data transferred per second, typically measured in gigabytes per second (GB / s). The I / O throughput of a single storage node is determined by its hardware specifications. Generally, if a storage pool maps to multiple storage nodes (e.g., a storage cluster), the I / O throughput of the entire storage pool can be obtained based on the I / O throughput of a single storage device within that cluster.
[0248] As an example, the block storage service 600 can determine the IO throughput of the storage pool based on a first preset coefficient and the IO throughput of a single storage node among the multiple storage nodes mapped to the storage pool. This preset coefficient may be related to the number of storage nodes mapped to the storage pool, as well as other information about the storage nodes; no restrictions are imposed here.
[0249] When determining the IO utilization of a storage pool, the block storage service 600 can determine the IO utilization of the storage pool based on the IO throughput capacity (also known as the number of input / output operations per second (IOPS)) of each storage node among the multiple storage nodes mapped to the storage pool, and the utilization of one or more CPUs of each storage node.
[0250] As an example, the block storage service 600 can obtain the IO utilization of the storage pool based on a second preset coefficient, the IO throughput of a single storage node among the multiple storage nodes mapped by the storage pool, and the CPU utilization of that single storage node.
[0251] As an implementation example:
[0252] The block storage service 600 can determine the IO capability information of the storage pool based on the IO throughput capability of the storage pool. The higher the IO throughput capability of the storage pool, the higher the IO capability value.
[0253] As another implementation example:
[0254] The block storage service 600 can determine the IO capability information of a storage pool based on its IO utilization rate, where a storage pool with a lower IO utilization rate has a higher IO capability value.
[0255] As another implementation example:
[0256] The block storage service 600 can determine the IO capacity information of the storage pool based on the IO utilization and IO throughput of the storage pool.
[0257] For example, the block storage service 600 can perform weighted calculations on the IO utilization and IO throughput of a storage pool to obtain the IO capability value of the storage pool. The higher the IO throughput and the lower the IO utilization, the higher the IO capability value of the storage pool.
[0258] Please refer to Figure 3a and Figure 6a When calculating the IO capacity value of storage pool 1 in the block storage service 600, the IO capacity value of the storage pool can be determined based on the IO throughput capacity of the storage pool.
[0259] like Figure 6aAs shown, after volume V11 in storage pool 1 is deleted, the IO throughput of storage pool 1 is as follows: single-volume data read / write rate is 3GB / s; single-volume data read / write rate of storage pool 2 is 4GB / s; single-volume data read / write rate of storage pool 3 is 5GB / s; and single-volume data read / write rate of storage pool 4 is 6GB / s. Based on the obtained single-volume data read / write rate of each storage pool from storage pool 1 to storage pool 4, the block storage service 600 can determine the IO capacity value of each storage pool. For example, the IO capacity values of storage pool 1 to storage pool 4 are 60, 70, 80, and 90, respectively.
[0260] Specifically, when determining the IO capacity value of storage pool 4, volume V41 has not yet been created in storage pool 4 (for example, the state of storage pool 4 is as follows). Figure 3a The storage pool 4 shown does not contain any logical volumes, or the storage pool 4 has created volume V41, but the data copy of volume V11 has not yet been written to volume V41.
[0261] After that, as Figure 6a As shown, the block storage service 600 can report the IO capability values of each storage pool from storage pool 1 to storage pool 4 to the data recovery service 900. In this way, the data recovery service 900 can receive the IO capability values of each storage pool from storage pool 1 to storage pool 4 from the block storage service 600.
[0262] Unlike related technologies, the storage cluster 300 in this embodiment can send IO capability information of multiple storage pools to the management node. This allows the management node to select from N target storage pools using the IO capability information of multiple storage pools. This makes the target storage pool selectable, whereas in related technologies, data recovery is performed directly on the source volume, making the target storage pool unselectable.
[0263] Furthermore, in this embodiment, the storage cluster 300 can respond to a request from the management node 1 to collect IO metrics from multiple storage pools, such as IO throughput and / or IO utilization, thereby determining the IO capability information of each storage pool. Finally, the storage cluster 300 can send the IO capability information of each storage pool to the management node 1, enabling the management node 1 to receive the IO capability information from the storage cluster 300. The storage cluster 300 directly calculates the IO capability information of the storage pools using the collected IO metrics, which improves the calculation speed of IO capability information. Since the block storage service 600 in the storage cluster 300 manages the storage pools, the block storage service 600 acquires IO metrics faster and processes IO metrics faster than the management node, facilitating accurate calculation of the storage pool's IO capability. In this way, management node 1 does not need to calculate the IO capacity of the storage pool, but directly obtains the IO capacity information of multiple storage pools provided by storage cluster 300, so as to determine N target storage pools, which can reduce the data processing steps of the management node and improve the running speed of the management node.
[0264] In one possible implementation 2, when the management node 1 executes S303a, the management node 1 may also receive IO metrics from the plurality of storage pools of the storage cluster 300 based on the data recovery request. The IO metrics include at least one of IO throughput and IO utilization. Based on the IO metrics of the plurality of storage pools, the management node 1 determines the IO capability information of the plurality of storage pools.
[0265] For example, the data recovery service 900 in management node 1 can send a request to the block storage service 600 provided by storage cluster 300 based on a data recovery request. This request instructs the storage service 600 to obtain the IO metrics for each storage pool it manages. The block storage service 600 can then collect the IO metrics for each storage pool it manages and report them to the data recovery service 900. Based on the IO metrics reported by the block storage service 600, the data recovery service 900 can then determine the IO capability information (e.g., IO capability value) of each storage pool.
[0266] The specific implementation principle of the management node 1 in determining the IO capability information of the storage pool based on the IO metrics of the storage pool is the same as the implementation principle of the block storage service 600 in determining the IO capability information of the storage pool based on the IO metrics of the storage pool as described above. Please refer to the corresponding description above for details, and it will not be repeated here.
[0267] Unlike related technologies, the storage cluster 300 in this embodiment can send IO metrics of multiple storage pools to the management node. This allows the management node to determine the IO capabilities of multiple storage pools based on these metrics, which is then used to select from N target storage pools. This makes the target storage pools selectable, whereas in related technologies where data recovery is performed directly on the source volume, the target storage pool cannot be selected.
[0268] Furthermore, unlike the above-mentioned Implementation 1 where the storage cluster 300 determines the IO capability information of the storage pool based on the IO indicators of the storage pool, in this Implementation 2, the storage cluster 300 can report the collected IO indicators of the storage pool to the management node, so that the management node can determine the IO capability based on the IO indicators. This allows the management node to share some of the computing load of the storage cluster 300, thereby improving the overall speed of data recovery.
[0269] It should be understood that IO capability information is not limited to IO capability values, but can also include IO capability levels, such as... Figure 6a The I / O capability levels of storage pools 1 to 4 shown are "Level 4", "Level 3", "Level 2", and "Level 1", respectively. "Level 1" indicates a higher I / O capability than "Level 2", and so on. Of course, I / O capability levels are not limited to the four levels illustrated here.
[0270] For example, each IO capability level can be matched with an IO capability range (a range of IO capability values), and there are no restrictions on the specific matching relationship between IO capability levels and IO capability values.
[0271] Continue back Figure 4a After step S303a, management node 1 can execute step S304a.
[0272] Step S304a: Based on the IO capability information of multiple storage pools, management node 1 determines a target storage pool from among the multiple storage pools.
[0273] Among them, the IO capacity value of the target storage pool is higher than that of the source storage pool.
[0274] As an implementation example, please refer to Figure 6a Based on the IO capability values of storage pools 1 to 4 (60, 70, 80, and 90 respectively), management node 1 can select any storage pool with an IO capability value higher than that of storage pool 1 as the target storage pool.
[0275] For example, the target storage pool can be any one of storage pool 2 to storage pool 4.
[0276] exist Figure 6aIn the scenario shown, management node 1 selects storage pool 4 with the highest IO capability value based on the IO capability information of multiple storage pools. Thus, storage pool 4 is the target storage pool.
[0277] Unlike data recovery solutions in related technologies, in this embodiment, the management node can obtain IO capability information of multiple storage pools based on a data recovery request. However, solutions in related technologies do not support the management node obtaining IO capability information of storage pools. Therefore, the management node in this embodiment can determine a target storage pool from among multiple storage pools based on the IO capability information of each storage pool, and restore the backup data of the source volume to a target volume within that target storage pool. Because the management node has greater computing power than the storage nodes, it can filter target storage pools based on their IO capability information, thereby determining the target storage pool more quickly, achieving rapid recovery of the source volume, shortening the data recovery time, and improving data recovery efficiency.
[0278] Continue back Figure 4a After step S304a, management node 1 can execute step S305a.
[0279] In one possible implementation, if the target volume does not exist in the target storage pool, then in step S305a, the management node 1 sends a request to the storage cluster 300, which instructs to create a new volume (i.e., a new logical volume) in the target storage pool.
[0280] This new volume is the target volume described above.
[0281] The request may include the target storage pool's ID and the size of the logical volume to be created. This size must be the same as the source volume's size; that is, the size of the logical volume to be created must match the size of the failed source volume.
[0282] As an implementation example, please refer to Figure 6a The data recovery service 900 in management node 1 can instruct the block storage service 600 in storage cluster 300 to create a target volume in storage pool 4, and indicate the size of the target volume, which is the same as the size of volume V11, for example, 500G.
[0283] In step S403a, storage cluster 300 can create a new volume in the target storage pool based on the request in S305a.
[0284] As an implementation example, please refer to Figure 6a The block storage service 600 can create a new volume of size 500G within storage pool 4 based on the request in S305a, for example, volume ID V41.
[0285] Regarding the logical volumes managed by Block Storage Service 600, their IDs are unique. Therefore, the ID of a newly created V41 volume is different from the original volume ID.
[0286] In this example, the newly created volume V41 is an empty volume, and the block storage service 600 has divided volume V41 into multiple data blocks during the creation process. As mentioned above, the size of volume V41 is the same as the size of volume V11, and the size of each data block in the volume managed by the block storage service 600 is specified and constant. Therefore, the number of data blocks in volume V41 is the same as the number of data blocks in volume V11, for example, 100. Thus, the newly created volume V41 can include 100 data blocks, with block numbers from 1 to 100 starting from the starting address of volume V41. Similarly, starting from the starting address of the source volume (here, volume V11), the block numbers of the 100 data blocks are also from 1 to 100.
[0287] When writing data to any logical volume, the block storage service 600 writes data sequentially from the smallest block number. Only after block 1 (block number 1) is full will data be written to block 2 (block number 2), and so on.
[0288] It should be understood that the data recovery method in this application embodiment is to perform data recovery on the entire logical volume. For example, only the first 80 blocks in the source volume have data written to them, and the remaining 20 blocks have not yet been written to. However, each time the source volume is backed up, all data blocks (here, 100) are backed up. Therefore, when performing data recovery on the source volume, the data copies of the 100 backed-up data blocks can also be restored to the target volume, instead of only restoring the first 80 backed-up data blocks. Although the data copies of the last 20 data blocks are empty (e.g., 0000), the data copies of these 20 data blocks still need to be restored to the target volume.
[0289] In step S404a, storage cluster 300 can send the metadata of the newly created volume to management node 1, so that management node 1 can receive the metadata of the new volume (i.e., the target volume) from storage cluster 300.
[0290] As an implementation example, please refer to Figure 6a After creating a new volume V41 in storage pool 4, the block storage service 600 can send the metadata information of volume V41 to the data recovery service 900.
[0291] For example, this metadata information may include, but is not limited to: the volume ID of volume V41, and the ID of the storage pool to which volume V41 belongs (here, storage pool 4).
[0292] In step S306a, management node 1 can control storage cluster 300 to write backup data from the source volume to the new volume based on the metadata of the new volume.
[0293] As an implementation example, please refer to Figure 6a Data recovery service 900 can determine the target storage pool to be written backup data based on the metadata information of volume V41, which is storage pool 4 in this case. In this way, data recovery service 900 can control block storage service 600 and object storage service 700 to write the backup data of volume V11 (which can be indicated by the backup set ID in the data recovery request) into volume V41 in storage pool 4.
[0294] For example, the storage bucket 1 managed by the object storage service 700 can store the mapping relationship between backup set 1 and volume V11. This mapping relationship is the mapping relationship between the offset address of the data block in volume V11 and the object in backup set 1. Among them, between the source volume and the target volume (also called the target volume), the offset addresses of data blocks with the same block number are the same. That is, the offset address of block 1 in volume V11 is the same as the offset address of block 1 in volume V41, and the same applies to other data blocks.
[0295] In this way, the object storage service 700 can determine, based on this mapping relationship, such as Figure 6a The object 1 in the backup set 1 shown is a data copy of block 1 of volume V11 in storage pool 1, the object 2 in the backup set 1 is a data copy of block 2 of volume V11 in storage pool 1, and so on, which will not be elaborated here.
[0296] Specifically, the data recovery service 900 can create a thread to perform the data recovery operation. This thread can call the object storage service 700 to read a copy of the data of block 1 in object 1 from backup set 1, and the thread can call the block storage service 600 to write the copy of the data of block 1 into block 1 in volume V41 according to the above mapping relationship (the logical address of block 1 in volume V41 is the same as the logical address of block 1 in volume V11 in the mapping relationship); similarly, the thread can control the writing of the copy of the data of block 2 in object 2 into block 2 in volume V41; ..., and so on, the data recovery service 900 writes the copies of the data of 100 data blocks of volume V11 in 100 objects in backup set 1 into 100 data blocks of volume V41 to achieve data recovery of volume V11.
[0297] In some embodiments, such as Figure 6a As shown, in the specific implementation, the data recovery service 900 can create multiple threads, each thread is responsible for the data recovery of a certain number of data copies of volume V11. Then, multiple threads can operate in parallel to realize the concurrent writing of data copies of multiple data blocks in backup set 1 into volume V41.
[0298] like Figure 6a The dashed arrows showing different types of data between storage pool 4 and storage bucket 1 indicate that different threads control the process. For example, data recovery service 900 can use thread 1 to recover 50 data blocks with odd block numbers (e.g., block 1, block 3, ..., block 49), and use thread 2 to recover 50 data blocks with even block numbers (e.g., block 2, block 4, ..., block 50), thus writing the data copy from backup set 1 to volume V41. Data recovery service 900 can control thread 1 and thread 2 to execute the data recovery operation in parallel. "Parallel" means that while thread 1 is performing the data recovery task, thread 2 is also performing the same task. This multi-threaded concurrent writing of data copies to volume V41 further improves data recovery speed and reduces the data recovery time of the source volume.
[0299] Of course, this application does not limit whether the number of data blocks each thread is responsible for is the same, nor does it limit the specific number of multiple threads. It can be flexibly set according to the number of data blocks in the source volume.
[0300] The data recovery process described above, which involves writing a copy of the data in backup set 1 to volume V41, is not intended to limit the data recovery method of this application embodiment. In other embodiments, other specific implementation methods can be used to write a copy of the data in backup set 1 to volume V41, and the implementation details of the specific writing process are not limited.
[0301] Optionally, after using object storage service 700 and block storage service 600 to write all data copies of volume V11 in backup set 1 to volume V41, so as to restore the backup data of volume V11 (e.g., backup set 1) to volume V41, such as Figure 6a As shown, the block storage service 600 can notify the data recovery service 900 of a message indicating that the data on volume V11 has been successfully recovered.
[0302] Continue back Figure 4a Following S306a, as shown by the dashed arrow, the management node 1 may optionally also execute step S307.
[0303] In step S307, management node 1 instructs virtual machine 1 to mount a new volume.
[0304] In step S502, virtual machine 1 mounts a new volume based on the instructions in step S307.
[0305] As an implementation example, please refer to Figure 6a Data Recovery Service 900 can instruct Virtual Machine 1 to mount Volume V41, so that Virtual Machine 1 can mount Volume V41 onto Virtual Machine 1 according to the instruction.
[0306] The process of mounting volume V41 to virtual machine 1 can be implemented using any existing or future method of mounting logical volumes to virtual machines; no restrictions are imposed here.
[0307] For example, the mounting process of volume V41 to virtual machine 1 may include allocating volume V41 to virtual machine 1 and configuring the corresponding file system in the operating system of virtual machine 1 so that virtual machine 1 can access and manage the data in volume V41.
[0308] After that, as Figure 6a As shown, client 102 can initiate IO access to volume V41 from the running virtual machine 1 to access the data in volume V41.
[0309] Back Figure 5b In the tenant through such Figure 5b (2) After selecting the "Start Data Recovery" button 206, management node 1 will be able to perform the following operations: Figure 4a As shown in S201, a data recovery request is received, and thus the following is executed: Figure 4a The steps shown are to achieve data recovery of volume V11, thereby presenting the data as follows: Figure 5b (3) Data recovery decoding 204 is shown. In, as... Figure 4a After S306a (e.g., after the source volume's data is restored to the target volume), or after S502, the data recovery interface of client 102 will change from... Figure 5b (3) Updated to Figure 5b (4) shows the data recovery interface.
[0310] like Figure 5b As shown in (4), the data recovery decoding 204 may include text 208 to remind the tenant that the source volume to which it requested data recovery has been successfully recovered, and to remind the tenant of the volume ID information of the new volume to which the data in the source volume has been recovered.
[0311] Combined with Figure 3a , Figure 3b , Figure 4a , Figures 5a to 5b , Figure 5c An example is shown Figure 5a and Figure 5b The diagram shows the interface before and after data recovery of Volume V11 in the block storage service scenario.
[0312] like Figure 5c As shown in (1), before data recovery of volume V11 is performed using the method of this application embodiment, or in other words, before the data of volume V11 is successfully recovered (e.g.) Figure 4aBefore S502 shown, the block storage service interface 209 displayed by the client 102 may include information about two logical volumes mounted on the virtual machine 1, namely volume V11 (specifically described as "local disk V11") and volume V12 (specifically described as "local disk V12"). Volume V11 has a size of 500G, 400G of which has been used to store data, and 100G is available; volume V12 has a size of 800G, 700G of which has been used to store data, and 100G is available.
[0313] like Figure 5c As shown in (2), after data recovery of volume V11 using the method of this application embodiment, for example... Figure 4a Following S502, the interface 209 displayed by client 102 can include information about two logical volumes mounted on virtual machine 1: volume V41 (specifically described as "local disk V41") and volume V12 (specifically described as "local disk V12"). The information of the originally faulty volume V11 is updated to local disk V41, which is a client example of volume V41. Because data recovery was not performed within the source volume (V11), but rather within a target volume different from the source volume, the volume ID changes after data recovery. For example, before data recovery, the volume ID of a 500GB local disk was V11; after data recovery, the volume ID is updated to V41. However, according to the related technology's method of performing data recovery within the source volume, the interface displayed before and after data recovery of volume V11 is as follows... Figure 5c (1) The volume ID will not change in the interface shown.
[0314] It should be understood that Figure 5a , Figure 5b , Figure 5c This is merely an example of an application scenario for a data recovery method. It is not intended to limit the interface displayed on the client during the implementation of the data recovery method in this application, and the actions that trigger arbitrary operations are not limited to click-and-select operations.
[0315] Example 2
[0316] In this Example 2, we will combine Figure 4b Taking the above N target storage pools as multiple target storage pools (i.e., N≥2) as an example, the implementation process of the data recovery method of this application will be described.
[0317] Figure 4b It can be combined with the above Figure 1a , Figure 1b , Figure 3a , Figure 3b However, it is not limited to combining the above. Figure 1a , Figure 1b , Figure 3a , Figure 3b The architecture shown.
[0318] also, Figure 4b The data recovery process shown can also be applied to Figures 5a to 5c The application scenarios shown can be referenced in Example 1 for specific application processes. Figures 5a to 5c The details of the introduction will not be repeated here.
[0319] In addition, in the introduction Figure 4b The data recovery method shown will combine Figure 6b The application scenarios shown are used to illustrate this. Figure 4b The implementation process of the data recovery method shown is as follows, however, Figure 4b The data recovery method shown is not limited to applications in [specific fields]. Figure 6b The scene shown.
[0320] And in this example 1 Figure 4b The process shown is the same as in Example 1. Figure 4a The processes shown are largely the same; the main differences lie in creating a target volume (e.g., a new volume) across storage pools and writing backup data from the source volume across storage pools to different logical spaces within the target volume. Other similar content can be found in [reference needed]. Figure 4a The relevant embodiments can be described in detail.
[0321] akin, Figure 6b The implementation process shown is the same as Figure 6a The implementation process shown is largely the same. This example 2 mainly focuses on describing the process of creating a target volume (e.g., a new volume) across storage pools, and writing backup data from the source volume across storage pools to different logical spaces within the target volume. Other implementation processes are similar. Figure 6a The relevant introduction is sufficient, and will not be repeated here.
[0322] like Figure 4b As shown, the process may include the following steps:
[0323] Step S201: Management node 1 receives a data recovery request from client 102.
[0324] In step S301, management node 1 may respond to the data recovery request by instructing virtual machine 1 to mount the source volume;
[0325] Step S501: Virtual machine 1 can mount the source volume.
[0326] In step S302, management node 1 may respond to the data recovery request by instructing storage cluster 300 to delete the source volume.
[0327] Step S401: Storage cluster 300 can delete the source volume.
[0328] The implementation process of each of the above steps is the same as in Example 1. Figure 4a The steps with the same labels shown are implemented in the same way, and will not be described again here.
[0329] After S201, such as Figure 4b As shown, the process may also include:
[0330] In step S303b, management node 1 responds to the data recovery request and obtains the IO capability information of multiple storage pools.
[0331] This application does not impose any restrictions on the execution order between S303b and S302.
[0332] The implementation process of step S303b is the same as in Example 1. Figure 4a The implementation principle of S303a shown is mostly the same, the only difference being that when the management node 1 executes step S303b, it can receive IO metrics from the multiple storage pools of the storage cluster 300 based on the data recovery request. These IO metrics include IO throughput and / or IO utilization. Based on the IO metrics of the multiple storage pools, the management node 1 determines the IO capability information of the multiple storage pools. The specific implementation process is the same as that of "Implementation Method 2" mentioned in Example 1, and will not be repeated here.
[0333] As an implementation example, please refer to Figure 6b The data recovery service 900 can receive the IO metrics of each storage pool from storage pool 1 to storage pool 4 from the block storage service 600. Then, based on the respective IO metrics of the four storage pools, the data recovery service 900 can calculate the IO capability value (or IO capability level) of each storage pool, thereby obtaining the IO capability information of each storage pool. For example, the data recovery service 900 calculates the IO capability values of storage pool 1 to storage pool 4 to be 60, 70, 80, and 90 respectively.
[0334] After step S303b, management node 1 can execute step S304b.
[0335] In step S304b, management node 1 determines multiple target storage pools among the multiple storage pools based on the IO capability information of multiple storage pools.
[0336] The average IO capacity of these multiple target storage pools is higher than that of the source storage pool.
[0337] As an implementation example, please refer to Figure 6bBased on the IO capability values of storage pools 1 to 4 (60, 70, 80, and 90 respectively), management node 1 can select the two storage pools with the highest IO capability values as target storage pools. For example, the two target storage pools could be storage pool 3 and storage pool 4.
[0338] In one possible implementation, in step S305b, management node 1 sends a request to storage cluster 300, which instructs to create a new volume (i.e., a new logical volume) across multiple target storage pools.
[0339] The request may include the ID of each target storage pool and the size of the logical volume to be created.
[0340] As an implementation example, please refer to Figure 6b The data recovery service 900 in management node 1 can instruct the block storage service 600 in storage cluster 300 to create a target volume belonging to storage pools 3 and 4, and indicate the size of the target volume, which is the same as the size of volume V11, for example, 500G.
[0341] In step S403b, storage cluster 300 can create a new volume (as a target volume) across multiple target storage pools (e.g., N target storage pools) based on the instructions in S305b.
[0342] The target volume may include N logical spaces located within N target storage pools.
[0343] For example, the addresses between the N logical spaces are contiguous. For instance, the end address of the logical space within storage pool 3 and the start address of the logical space within storage pool 4 are contiguous to ensure that although the target volume is stored across storage pools, the logical space of the target volume remains contiguous, for example, with a start address of 0 and an offset address of 500G.
[0344] As an implementation example, please refer to Figure 6b The block storage service 600 can, based on the instructions in S305b, create a new volume of size 500G across storage pools 3 and 4, for example, volume ID V41.
[0345] For logical volumes managed by the block storage service 600, their IDs are unique. Therefore, the volume ID of a newly created V41 is different from the volume ID of volume V11.
[0346] Specifically, the block storage service 600 can create a logical space a of size A in storage pool 3, and a logical space b of size B in storage pool 4. The sum of size A and size B is the size of the new volume, for example, 500G. The block storage service 600 uses logical space a and logical space b as the logical space of the new volume, for example, the new volume is volume V41.
[0347] In some embodiments, this application does not limit the size of the logical space belonging to the new volume created by the block storage service 600 in each target storage pool, as long as the sum of the sizes of the logical spaces created within each of the multiple target storage pools is equal to the size of the new volume.
[0348] In some embodiments, storage cluster 300 (e.g., as shown in some examples) Figure 6b The block storage service 600 shown can obtain the IO capability information of each target storage pool specified by the data recovery service 900. For example, if the target storage pools specified by the data recovery service 900 are storage pool 3 and storage pool 4, the block storage service 600 can obtain the IO capability information of each of these two storage pools (for the specific acquisition process, please refer to the relevant introduction of the block storage service 600 acquiring the IO capability information of the storage pool in Example 1, which will not be repeated here). Then, based on the IO capability information of each target storage pool, the block storage service 600 determines the size of the logical space belonging to the new volume that needs to be created in each target storage pool. Finally, the block storage service 600 creates the logical space of the new volume in the corresponding target storage pool according to the determined size of the logical space to be created for each target storage pool, thereby realizing the creation of a new volume across multiple target storage pools.
[0349] by Figure 6b For example, if the target volume (i.e., the new volume to be created) is 500G, the block storage service 600 can obtain the IO capability value of storage pool 3 as 80 and the IO capability value of storage pool 4 as 90, and thus determine to create a larger logical space belonging to the new volume in the storage pool with the higher IO capability value. For example, the size B of the logical space b to be created in storage pool 4 is greater than the size A of the logical space a to be created in storage pool 3. For example, based on the respective IO capability values of storage pool 3 and storage pool 4, the block storage service 600 determines that the size A of the logical space a to be created is 125G and the size B of the logical space b to be created is 375G, totaling 500G.
[0350] Furthermore, similar to Example 1, whether creating a new volume within a single storage pool or across multiple storage pools, the logical space containing the new volume within each storage pool is divided into multiple data blocks of a specified size during the volume creation process. Therefore, in... Figure 6bIn the application scenario shown, for example, when the block storage service 600 creates a logical space a of size 125G in storage pool 3, it will also divide the logical space a into multiple data blocks, such as 50 data blocks, and assign block numbers, such as 1 to 50. Here, the 50 consecutive data blocks in the logical space a are named blocks 1 to 50.
[0351] Similarly, for example, when the block storage service 600 creates a logical space b of size 375G within storage pool 4, it will divide logical space b into multiple data blocks, such as 150 data blocks, and assign block numbers, such as 51 to 200. Here, the 150 consecutive data blocks in logical space b are named blocks 51 to 200. The block numbers between logical space a and logical space b are consecutive; in other words, the logical addresses between logical space a and logical space b are also consecutive.
[0352] In this embodiment, multiple storage nodes providing block storage services can determine the size of the logical space to be created in each target storage pool of the target volume used for data recovery based on the IO capability information of each target storage pool indicated by the management node. For example, a larger proportion of the logical space for the target volume can be created in the target storage pool with stronger IO capability. For example, the larger the IO capability value of the target storage pool, the stronger the IO capability of the target storage pool. In this way, the target storage pool with relatively stronger IO capability among the multiple target storage pools (e.g., Figure 6b Storage pool 4) is used to write a larger proportion of backup data to the source volume, while the target storage pool, which has relatively weaker IO capabilities (e.g., storage pool 4), is used to write the backup data to the source volume. Figure 6b Storage pool 3) as shown writes a smaller proportion of backup data to the source volume. This allows for the efficient use of the IO capabilities of multiple target storage pools to write different proportions of backup data to the source volume to target storage pools with varying IO capabilities, thereby improving the overall recovery efficiency across multiple storage pools and reducing data recovery time.
[0353] In some embodiments, in order to reduce the management overhead of the block storage service 600 on the target volume, when the data recovery service 900 instructs the block storage service 600 to create a new target volume belonging to N target storage pools, the block storage service 600 may further filter the N target storage pools for creating a target volume across pools as instructed by the data recovery service 900, for example, selecting one or more target storage pools from the N target storage pools as the final target storage pool for creating the target volume across pools.
[0354] by Figure 6bFor example, data recovery service 900 can select storage pools 2 to 4 as target storage pools for cross-pool target volume creation based on the IO capability information of storage pools 1 to 4, thus instructing block storage service 600 to create the target volume across these three storage pools. However, block storage service 600 can pre-set a minimum logical space (e.g., 100GB) for the logical volume created in each storage pool. Even if the logical volume is a cross-pool created logical volume, the minimum size of the logical space belonging to any storage pool within that logical volume is 100GB.
[0355] Thus, based on the size of the target volume to be created (here, 500G) and the IO capacity information of each storage pool in storage pools 2 to 4, the block storage service 600 can determine the sizes of the three logical spaces belonging to the target volume to be created in storage pools 2, 3, and 4 as 80, 120, and 300, respectively. However, the size of the logical space to be created in storage pool 2 (80G) is less than the size of the minimum logical space (here, 100G). Therefore, the block storage service 600 can further filter the multiple target storage pools indicated by the data recovery service 900. For example, it can exclude target storage pools with smaller IO capabilities from the cross-pool target volume creation process. Specifically, the block storage service 600 can select storage pools 3 and 4, which have stronger IO capabilities, as the final target storage pools, according to the following... Figure 6b The method shown is used to create volume V41 across pools.
[0356] Furthermore, as described in Example 1 above, the newly created volume V41 is an empty volume, and the block storage service 600 has divided volume V41 into multiple data blocks during the creation process. As mentioned above, the size of volume V41 is the same as the size of volume V11, and the size of each data block in the volume managed by the block storage service 600 is specified and constant. Therefore, the number of data blocks in volume V41 is the same as the number of data blocks in volume V11, for example, 100. Thus, the newly created volume V41 can include 100 data blocks, with block numbers from 1 to 100 starting from the starting address of volume V41. Similarly, starting from the starting address of the source volume (here, volume V11), the block numbers of the 100 data blocks are also from 1 to 100.
[0357] When writing data to any logical volume, the block storage service 600 writes data sequentially from the smallest block number. Only after block 1 (block number 1) is full will data be written to block 2 (block number 2), and so on.
[0358] It should be understood that the data recovery method in this application embodiment is to perform data recovery on the entire logical volume. For example, only the first 80 blocks in the source volume have data written to them, and the remaining 20 blocks have not yet been written to. However, each time the source volume is backed up, all data blocks (here, 100) are backed up. Therefore, when performing data recovery on the source volume, the data copies of the 100 backed-up data blocks can also be restored to the target volume, instead of only restoring the first 80 backed-up data blocks. Although the data copies of the last 20 data blocks are empty (e.g., 0000), the data copies of these 20 data blocks still need to be restored to the target volume.
[0359] In step S404b, storage cluster 300 can send the metadata of the newly created volume to management node 1, so that management node 1 can receive the metadata of the new volume (i.e., the target volume) from storage cluster 300.
[0360] As an implementation example, please refer to Figure 6b After creating volume V41 across storage pool 3 and storage pool 4, block storage service 600 can send the metadata information of volume V41 to data recovery service 900.
[0361] like Figure 6b As shown, unlike Example 1, the metadata information of Volume V41 may include, but is not limited to: the Volume ID of Volume V41; the first matching relationship between the logical address (also called the access address) (e.g., 0 to 125G) of logical space a of Volume V41 and the storage pool 3 (e.g., storage pool ID) to which logical space a belongs; and the second matching relationship between the logical address (e.g., 125G to 500G) of logical space b of Volume V41 and the storage pool 4 (e.g., storage pool ID) to which logical space b belongs. Similarly, if the number of target storage pools is greater than 2, the metadata information of Volume V41 may include more matching relationships, which will not be elaborated here.
[0362] In step S306b, management node 1 can control storage cluster 300 to write backup data from the source volume to a new volume across multiple target storage pools based on a data recovery request.
[0363] Specifically, based on the data recovery request, management node 1 restores the backup data of the source volume to N logical spaces within multiple target storage pools.
[0364] In related block storage services, creating a logical volume across multiple storage pools is not supported. However, in this embodiment, the management node can send a request to multiple storage nodes (e.g., storage cluster 300) providing multiple storage pools based on a data recovery request, instructing them to create a target volume belonging to N target storage pools. Multiple storage nodes can then create a target volume across the N target storage pools based on this request. This target volume can include N logical spaces located within the N target storage pools, with each logical space corresponding one-to-one with one of the N target storage pools. When performing data recovery on the source volume, the management node can restore the backup data of the source volume to the N logical spaces within the N target volumes based on the data recovery request. This allows the use of multiple storage pools with strong I / O capabilities to write the backup data of the source volume, thus accelerating data recovery. Furthermore, as mentioned above, with each data block size remaining constant, a large amount of data in the source volume can lead to prolonged data recovery times, even with strong IO capabilities in the source storage pool, when dealing with a massive number of data block copies being written to the same storage pool (here, the source storage pool). Therefore, a large source volume also results in a long data recovery time. To overcome this issue, in this embodiment, each data block copy in the source volume can be written to multiple target storage pools with strong IO capabilities, thus mitigating the impact of source volume size on data recovery time. Even with a large source volume, because the N logical spaces of the target volume span multiple storage pools, the writing speed of data copies to each logical space can be accelerated, thereby improving the overall recovery speed of the source volume and shortening the data recovery time.
[0365] In some embodiments, when performing step S306b, the management node 2 may, based on the metadata and data recovery request of the new volume, restore the first backup data (125G of backup data) of the source volume to the logical space a in the storage pool 3 according to the logical address (e.g., 0 to 125G) in the first matching relationship, and restore the second backup data (375G of backup data) of the source volume to the logical space b in the storage pool 4 according to the logical address (e.g., 125G to 500G) in the second matching relationship.
[0366] As an implementation example, as described in Example 1, the offset addresses of data blocks with the same block number are the same between the source volume and the target volume. That is, the offset address of block 1 in volume V11 is the same as the offset address of block 1 in volume V41, and the same applies to other data blocks.
[0367] Please refer to Figure 6bData recovery service 900 can determine the backup set 1 of volume V11 managed by object storage service 700 based on the data recovery request, and data recovery service 900 can determine, based on the metadata information of volume V41, the logical address (e.g., 0 to 125G) to be written to the storage pool 3, and data recovery service 900 can determine, based on the metadata information of volume V41, the logical address (e.g., 125G to 500G) to be written to the storage pool 4.
[0368] The storage bucket 1 managed by the object storage service 700 can store the mapping relationship between backup set 1 and volume V11. This mapping relationship is the mapping relationship between the offset address of the data block in volume V11 and the object in backup set 1.
[0369] In this way, the object storage service 700 can determine, based on this mapping relationship, such as Figure 6b The object 1 in the backup set 1 shown is a data copy of block 1 of volume V11 in storage pool 1, the object 2 in the backup set 1 is a data copy of block 2 of volume V11 in storage pool 1, and so on, which will not be elaborated here.
[0370] Data recovery service 900 can control block storage service 600 and object storage service 700 to write backup data of volume V11 (specifically indicated by the backup set ID in the data recovery request) into volume V41 across storage pool 3 and storage pool 4.
[0371] Specifically, the data recovery service 900 can control the recovery of the first backup data of the backup volume V11 (e.g., data copies of blocks 1 to 50 of volume V11 stored in objects 1 to 50) to logical space a within storage pool 3 according to the logical addresses (e.g., 0 to 125G) in the first matching relationship; and the data recovery service 900 can control the recovery of the second backup data of the backup volume V11 (e.g., data copies of blocks 51 to 200 of volume V11 stored in objects 51 to 200) to logical space b within storage pool 4 according to the logical addresses (e.g., 125G to 500G) in the second matching relationship.
[0372] In this embodiment, the metadata of a logical volume created across multiple target storage pools differs from the metadata of a logical volume created within a single storage pool. The metadata of a logical volume created across multiple target storage pools includes the matching relationship between the access addresses of the logical spaces of the target volumes distributed across the target storage pools and the target storage pools to which those logical spaces belong. Thus, during the process of restoring data copies of the source volume to a target volume across multiple target storage pools, the matching relationship in the metadata can be used to restore each data copy of the source volume to its corresponding logical space within the corresponding target storage pool, thereby enabling data writing to logical volumes across multiple target storage pools and achieving cross-storage pool data recovery.
[0373] In one possible implementation, when management node 1 restores partial backup data of the source volume to the corresponding logical space in the corresponding target storage pool according to each matching relationship, it can be done in the following way: management node 1 can create multiple threads for writing backup data of the source volume to the N target storage pools based on the metadata and data recovery request of volume V41. The multiple threads include a first thread and a second thread. Based on the first thread, management node 1 writes the first backup data of the source volume to the logical space a in storage pool 3 according to the logical address (e.g., 0 to 125G) in the first matching relationship. During the process of management node 1 writing the first backup data to the logical space a through the first thread, management node 1 writes the second backup data of the source volume to the logical space b in storage pool 4 according to the logical address (e.g., 125G to 500G) in the second matching relationship through the second thread.
[0374] Among them, the N target storage pools are multiple target storage pools for which the block storage service 600 has created target volumes across storage pools, such as storage pool 3 and storage pool 4.
[0375] As an implementation example, please refer to Figure 6b The data recovery service 900 can create a first thread and a second thread. The data recovery service 900 can use the first thread to write data copies of blocks 1 to 50 of volume V11 in backup set 1 to logical space a in storage pool 3. For example, it writes data copies of block 1 of volume V11 in object 1 to block 1 of volume V41 in storage pool 3; writes data copies of block 2 of volume V11 in object 2 to block 2 of volume V41 in storage pool 3, and so on, writing data copies of block 50 of volume V11 in object 50 to block 50 of volume V41 in storage pool 3, thus realizing the writing of data copies of blocks 1 to 50 of volume V11 to logical space a in volume V41 in storage pool 3.
[0376] During the data recovery process of the first 50 data blocks of volume V11, which is implemented by the data recovery service 900 through the first thread, such as Figure 6b As shown, the data recovery service 900 can also use a second thread to write data copies of blocks 51 to 200 of volume V11 in objects 51 to 200 of backup set 1 to logical space b within storage pool 4. For example, it writes data copy of block 51 of volume V11 in object 51 to block 51 of volume V41 within storage pool 4; it writes data copy of block 52 of volume V11 in object 52 to block 22 of volume V41 within storage pool 4, and so on, writing data copy of block 200 of volume V11 in object 200 to block 200 of volume V41 within storage pool 4, thus realizing the writing of data copies of blocks 51 to 200 of volume V11 to logical space b within volume V41 within storage pool 4.
[0377] In data recovery methods in related technologies, since the source storage pool where the source volume is located is unique, when writing a copy of the source volume to the source storage pool, it can only be done through a single task corresponding to a single source storage pool. However, the processing capacity of a single task is limited, which reduces the efficiency of data recovery.
[0378] However, in this embodiment, target volumes can be created across multiple target storage pools for writing data copies of the source volume. This allows each target storage pool to have at least one thread to write a partial data copy of the source volume to a target storage pool, enabling concurrent writing of source volume data copies to target volumes within different target storage pools via multiple threads corresponding to N target storage pools. Related technologies, however, do not support concurrent writing of a source volume's data copy to multiple storage pools.
[0379] In contrast to related technologies where backup data of the source volume can only be written to a single storage pool, resulting in data recovery efficiency being limited by the single-threaded task processing capability of a single storage pool, the embodiments of this application allow for separate threads to write data copies to each of the multiple target storage pools. This leverages the combined IO processing capabilities of multiple target storage pools, enabling multi-tasking data recovery of the source volume without being limited by the processing capability of a single task, thus improving data recovery efficiency.
[0380] Optionally, such as Figure 6b As shown, similar to Example 1, when the data recovery service 900 writes data copies to the target volume in the same storage pool through threads, it can use multiple threads to concurrently write backup data to the same target storage pool to further improve data recovery efficiency.
[0381] in, Figure 6a The dashed arrows with different patterns in backup set 1 to volume V41 indicate different threads used for implementation.
[0382] Figure 6b The dashed arrows with different patterns in backup set 1 to volume V41 indicate different threads used for implementation.
[0383] Continue back Figure 4a Following S306b, as shown by the dashed arrow, management node 1 may optionally also execute step S307.
[0384] In step S307, management node 1 instructs virtual machine 1 to mount a new volume.
[0385] In step S502, virtual machine 1 mounts a new volume based on the instructions in step S307.
[0386] It should be understood that Figure 4a and Figure 4b In the illustrated process, the storage cluster 300 is multiple storage nodes that provide multiple storage pools. In other embodiments, the multiple storage nodes that provide multiple storage pools may not be a storage cluster, but a distributed storage system.
[0387] In Example 2 above, taking two target storage pools as an example of N target storage pools, the implementation process of the method of this application is illustrated. When there are more than two storage pools (e.g., three storage pools) among the N target storage pools, the above volume V41 is a logical volume created across three storage pools. The principle of creating volume V41 across three storage pools is the same as the principle of creating volume V41 across two storage pools, and the implementation principle of the data recovery process is also similar, which will not be repeated here.
[0388] Furthermore, in Examples 1 and 2 above, the data copy of the source volume is stored in a bucket managed by the object storage service as an example. However, in other embodiments, the data copy of the source volume may also be stored in a storage pool provided by the block storage service or in a file provided by the file storage service. This application does not limit the storage form, storage location, or management method of the data copy of the source volume.
[0389] The above combination Figures 1a to 6b The data recovery method provided in the embodiments of this application will be introduced. Next, the management node, data center and other equipment provided in the embodiments of this application will be introduced with reference to the accompanying drawings.
[0390] This application provides a cloud management platform for managing infrastructure that provides cloud services. The infrastructure includes multiple storage nodes, which are used to provide multiple storage pools. Figure 7a For an exemplary structural diagram of the cloud management platform 700, please refer to... Figure 7aThe cloud management platform 700 may include: a receiving module 701, configured to receive a data recovery request, the data recovery request including information indicating the source logical volume of the data to be recovered; a determining module 702, configured to determine N target storage pools from the plurality of storage pools based on the data recovery request and the input / output I / O capability information of the plurality of storage pools, wherein N is a positive integer greater than or equal to 1, the average I / O capability of the N target storage pools is higher than the I / O capability of the source storage pool, the I / O capability of each storage pool in the plurality of storage pools is used to indicate the data recovery capability of each storage pool, and the source storage pool is the storage pool to which the source logical volume belongs in the plurality of storage pools; and a processing module 703, configured to restore the backup data of the source logical volume to the logical volume included in the N target storage pools based on the data recovery request.
[0391] In one possible implementation, the determining module 702 is specifically used to: obtain the IO capability information of the plurality of storage pools based on the data recovery request; and determine N target storage pools from the plurality of storage pools based on the IO capability information of the plurality of storage pools.
[0392] In one possible implementation, the determining module 702 is specifically configured to: receive IO metrics from the plurality of storage pools from the plurality of storage nodes based on the data recovery request, wherein the IO metrics include at least one of the following: IO throughput and IO utilization; and determine the IO capability information of the plurality of storage pools based on the IO metrics of the plurality of storage pools.
[0393] In one possible implementation, the determining module 702 is specifically configured to: receive IO capability information of the multiple storage pools from the multiple storage nodes based on the data recovery request.
[0394] In one possible implementation, the determining module 702 is specifically configured to: send a first request to the plurality of storage nodes based on the data recovery request, the first request being used to obtain IO capability information of the plurality of storage pools; receive the IO capability information of the plurality of storage pools from the plurality of storage nodes, the IO capability information of the plurality of storage pools being determined based on the IO metrics of the plurality of storage pools, the IO metrics including at least one of the following: IO throughput and IO utilization.
[0395] This application provides a cloud service system. Figure 7b For an exemplary structural diagram of the cloud service system 800, please refer to... Figure 7bThe cloud service system includes a cloud management platform 700 and multiple storage nodes 801. The cloud management platform 700 manages the infrastructure providing the cloud service, including the multiple storage nodes 801, which provide multiple storage pools. The cloud management platform 700 receives data recovery requests, which include information indicating the source logical volume of the data to be recovered. Based on the data recovery request and the IO capability information of the multiple storage pools, the cloud management platform 700 determines N target storage pools from the multiple storage pools, where N is a positive integer greater than or equal to 1. The average IO capability of the N target storage pools is higher than the IO capability of the source storage pool. The IO capability of each storage pool in the multiple storage pools is used to indicate the data recovery capability of each storage pool. The source storage pool is the storage pool to which the source logical volume belongs. Based on the data recovery request, the cloud management platform 700 restores the backup data of the source logical volume to the logical volumes included in the N target storage pools.
[0396] The cloud management platform 700 and multiple storage nodes 801 can communicate via a network.
[0397] The cloud management platform 700 can achieve Figure 7a The functions implemented by the cloud management platform 700 in any of the possible implementations.
[0398] In one possible implementation, the cloud management platform 700 is further configured to send a first request to the plurality of storage nodes 801 based on the data recovery request, the first request being used to obtain IO capability information of the plurality of storage pools; the plurality of storage nodes 801 are configured to obtain IO metrics of the plurality of storage pools based on the first request, the IO metrics including at least one of the following: IO throughput capacity, IO utilization rate; and determine the IO capability information of the plurality of storage pools based on the IO metrics of the plurality of storage pools; the cloud management platform 700 is further configured to receive the IO capability information of the plurality of storage pools from the plurality of storage nodes.
[0399] In one possible implementation, N≥2, the cloud management platform 700 is specifically configured to send a second request to the plurality of storage nodes based on the data recovery request, the second request being used to instruct the creation of a new target logical volume belonging to the N target storage pools; the plurality of storage nodes are further configured to create a target logical volume across the N target storage pools based on the second request, the target logical volume including N logical spaces located within the N target storage pools; the cloud management platform 700 is specifically configured to restore the backup data of the source logical volume to the N logical spaces within the N target storage pools based on the data recovery request.
[0400] In one possible implementation, the plurality of storage nodes 801 are specifically configured to: obtain IO capability information of the N target storage pools based on the second request; determine the size of the i-th logical space to be created in the i-th target storage pool based on the IO capability information of the N target storage pools, where 1≤i≤N and i is a positive integer; and create the i-th logical space of the target logical volume in the i-th target storage pool according to the size of the i-th logical space, so as to create the target logical volume.
[0401] In one possible implementation, the N target storage pools include a first storage pool and a second storage pool, the N logical spaces include a first logical space located within the first storage pool and a second logical space located within the second storage pool, and the backup data of the source logical volume includes first backup data and second backup data; the plurality of storage nodes 801 are further configured to send metadata of the target logical volume to the cloud management platform 700, the metadata including a first matching relationship between a first access address of the first logical space and the first storage pool, and a second matching relationship between a second access address of the second logical space and the second storage pool; the cloud management platform 700 is further configured to, based on the metadata and the data recovery request, restore the first backup data of the source logical volume to the first logical space within the first storage pool according to the first access address in the first matching relationship, and restore the second backup data of the source logical volume to the second logical space within the second storage pool according to the second access address in the second matching relationship.
[0402] In one possible implementation, the cloud management platform 700 is specifically configured to: based on the metadata and the data recovery request, create multiple threads for writing backup data of the source logical volume to the N target storage pools, the multiple threads including a first thread and a second thread; based on the first thread, write the first backup data of the source logical volume to the first logical space in the first storage pool according to the first access address in the first matching relationship; during the process of writing the first backup data to the first logical space through the first thread, based on the second thread, write the second backup data of the source logical volume to the second logical space in the second storage pool according to the second access address in the second matching relationship.
[0403] The effects of the cloud management platform 700 and cloud service system in the above-described embodiments are similar to the effects of the data recovery methods in the above-described embodiments, and will not be repeated here.
[0404] The modules of the cloud management platform 700 can be implemented in software or hardware. As an example of a software functional unit, a module may include code running on a computing instance. A computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Furthermore, there may be one or more computing instances. For example, module 702 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers running the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0405] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0406] As an example of a hardware functional unit, the module mentioned above may include at least one computing device, such as a server. Alternatively, the module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0407] The aforementioned computing devices can be distributed within the same region or in different regions. They can also be distributed within the same Availability Zone (AZ) or in different AZs. Similarly, they can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. These multiple computing devices can be any combination of devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0408] It should be noted that, in other embodiments, the above-described module can be used to perform corresponding steps in the data recovery method to realize all the functions of the computing device.
[0409] This application also provides a computing device 900. For example... Figure 8 As shown, the computing device 900 includes: a bus 902, a processor 904, a memory 906, and optionally, a communication interface 909. The processor 904, the memory 906, and the communication interface 909 communicate via the bus 902. The computing device 900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 900.
[0410] The 902 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 8 The bus 902 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 902 may include a path for transmitting information between various components of the computing device 900 (e.g., memory 906, processor 904, communication interface 909).
[0411] Processor 904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0412] Memory 906 may include volatile memory, such as random access memory (RAM). Memory 906 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0413] The memory 906 stores executable program code, and the processor 904 executes this executable program code to implement the functions of the aforementioned receiving module 701, determining module 702, and processing module 703, thereby realizing the data recovery method. That is, the memory 906 stores instructions for executing the data recovery method.
[0414] The communication interface 909 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 900 and other devices or communication networks.
[0415] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0416] like Figure 9 As shown, the computing device cluster includes at least one computing device 1000. The memory 1006 of one or more computing devices 1000 in the computing device cluster may store the same instructions for executing data recovery methods.
[0417] The computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other via the bus 1002.
[0418] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the data recovery method. In other words, a combination of one or more computing devices 1000 can jointly execute the instructions for executing the data recovery method.
[0419] It should be noted that the memory 1006 in different computing devices 1000 within the computing device cluster can store different instructions, which are used to execute certain functions of the cloud management platform. That is, the instructions stored in the memory 1006 of different computing devices 1000 can implement the functions of one or more modules among the receiving module 701, the determining module 702, and the processing module 703.
[0420] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 10 One possible implementation is shown. For example... Figure 10 As shown, the two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through the communication interface 1108 in each computing device.
[0421] The computing device 1100A includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other via the bus 1102.
[0422] The computing device 1100B includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other via the bus 1102.
[0423] In this type of possible implementation, the memory 1106 in computing device 1100A stores instructions for performing the functions of receiving module 701 and determining module 702. Meanwhile, the memory 1106 in computing device 1100B stores instructions for performing the functions of processing module 703.
[0424] It should be understood that Figure 10 The functions of computing device 1100A shown can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.
[0425] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 8 and Figure 10 The connection method of the computing device cluster. The difference is that the memory 1106 of one or more computing devices 1100 in the computing device cluster can store the same instructions for executing the data recovery method.
[0426] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data recovery method. In other words, a combination of one or more computing devices 1100 can jointly execute the instructions for executing the data recovery method.
[0427] It should be noted that the memory 1106 in different computing devices 1100 within the computing device cluster can store different instructions for executing some functions of the cloud service system. That is, the instructions stored in the memory 1106 of different computing devices 1100 can implement the functions of storage nodes and the cloud management platform within the cloud service system.
[0428] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the data recovery method described above.
[0429] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the data recovery method described in the above embodiments.
[0430] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A data recovery method based on a cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure providing cloud services, the infrastructure including multiple storage nodes, the multiple storage nodes providing multiple storage pools, and the method including: The cloud management platform receives a data recovery request, which includes information indicating the source logical volume of the data to be recovered. The cloud management platform determines N target storage pools from the multiple storage pools based on the data recovery request and the input / output I / O capability information of the multiple storage pools, where N is a positive integer greater than or equal to 1. The average I / O capability of the N target storage pools is higher than the I / O capability of the source storage pool. The I / O capability of each storage pool in the multiple storage pools is used to indicate the data recovery capability of each storage pool. The source storage pool is the storage pool to which the source logical volume belongs among the multiple storage pools. Based on the data recovery request, the cloud management platform restores the backup data of the source logical volume to the logical volumes included in the N target storage pools.
2. The method according to claim 1, characterized in that, Based on the data recovery request and the IO capability information of the multiple storage pools, the cloud management platform determines N target storage pools from the multiple storage pools, including: The cloud management platform obtains the IO capability information of the multiple storage pools based on the data recovery request; The cloud management platform determines N target storage pools from the multiple storage pools based on their IO capability information.
3. The method according to claim 2, characterized in that, Based on the data recovery request, the cloud management platform obtains the IO capability information of the multiple storage pools, including: Based on the data recovery request, the cloud management platform receives IO metrics from the multiple storage pools of the multiple storage nodes. The IO metrics include at least one of the following: IO throughput and IO utilization. The cloud management platform determines the IO capability information of the multiple storage pools based on their IO metrics.
4. The method according to claim 2, characterized in that, Based on the data recovery request, the cloud management platform obtains the IO capability information of the multiple storage pools, including: Based on the data recovery request, the cloud management platform receives IO capability information from the multiple storage pools of the multiple storage nodes.
5. The method according to claim 4, characterized in that, Based on the data recovery request, the cloud management platform receives IO capability information from the multiple storage pools of the multiple storage nodes, including: Based on the data recovery request, the cloud management platform sends a first request to the multiple storage nodes, the first request being used to obtain IO capability information of multiple storage pools; The cloud management platform receives IO capability information from the multiple storage pools of the multiple storage nodes. The IO capability information of the multiple storage pools is determined based on the IO metrics of the multiple storage pools. The IO metrics include at least one of the following: IO throughput and IO utilization.
6. The method according to any one of claims 1 to 5, characterized in that, If N≥2, the cloud management platform, based on the data recovery request, restores the backup data of the source logical volume to the logical volumes included in the N target storage pools, including: Based on the data recovery request, the cloud management platform sends a second request to the multiple storage nodes. The second request is used to instruct the creation of a new target logical volume belonging to the N target storage pools. The target logical volume includes N logical spaces located within the N target storage pools. Based on the data recovery request, the cloud management platform restores the backup data of the source logical volume to the N logical spaces within the N target storage pools.
7. The method according to claim 6, characterized in that, The N target storage pools include a first storage pool and a second storage pool. The N logical spaces include a first logical space located in the first storage pool and a second logical space located in the second storage pool. The backup data of the source logical volume includes first backup data and second backup data. Based on the data recovery request, the cloud management platform restores the backup data of the source logical volume to the N logical spaces within the N target storage pools, including: The cloud management platform receives metadata of the target logical volume sent by the multiple storage nodes. The metadata includes a first matching relationship between the first access address of the first logical space and the first storage pool, and a second matching relationship between the second access address of the second logical space and the second storage pool. Based on the metadata and the data recovery request, the cloud management platform restores the first backup data of the source logical volume to the first logical space in the first storage pool according to the first access address in the first matching relationship, and restores the second backup data of the source logical volume to the second logical space in the second storage pool according to the second access address in the second matching relationship.
8. The method according to claim 7, characterized in that, The cloud management platform, based on the metadata and the data recovery request, restores the first backup data of the source logical volume to the first logical space within the first storage pool according to the first access address in the first matching relationship, and restores the second backup data of the source logical volume to the second logical space within the second storage pool according to the second access address in the second matching relationship, including: Based on the metadata and the data recovery request, the cloud management platform creates multiple threads for writing backup data of the source logical volume to the N target storage pools, including a first thread and a second thread. The cloud management platform, based on the first thread, writes the first backup data of the source logical volume to the first logical space in the first storage pool according to the first access address in the first matching relationship. During the process of the cloud management platform writing the first backup data to the first logical space through the first thread, the cloud management platform, based on the second thread, writes the second backup data of the source logical volume to the second logical space in the second storage pool according to the second access address in the second matching relationship.
9. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes multiple storage nodes, which provide multiple storage pools. The cloud management platform includes: A receiving module is configured to receive a data recovery request, the data recovery request including information indicating the source logical volume of the data to be recovered; The determining module is used to determine N target storage pools from the plurality of storage pools based on the data recovery request and the input / output I / O capability information of the plurality of storage pools, where N is a positive integer greater than or equal to 1, the average I / O capability of the N target storage pools is higher than the I / O capability of the source storage pool, the I / O capability of each storage pool in the plurality of storage pools is used to indicate the data recovery capability of each storage pool, and the source storage pool is the storage pool to which the source logical volume belongs in the plurality of storage pools; The processing module is used to restore the backup data of the source logical volume to the logical volumes included in the N target storage pools based on the data recovery request.
10. The cloud management platform according to claim 9, characterized in that, The determining module is specifically used for: Based on the data recovery request, obtain the IO capability information of the multiple storage pools; Based on the IO capability information of the multiple storage pools, N target storage pools are determined from the multiple storage pools.
11. The cloud management platform according to claim 10, characterized in that, The determining module is specifically used for: Based on the data recovery request, IO metrics from the multiple storage pools of the multiple storage nodes are received, and the IO metrics include at least one of the following: IO throughput and IO utilization. Based on the IO metrics of the multiple storage pools, the IO capability information of the multiple storage pools is determined.
12. The cloud management platform according to claim 10, characterized in that, The determining module is specifically used to: receive IO capability information of the multiple storage pools from the multiple storage nodes based on the data recovery request.
13. A cloud service system, characterized in that, The cloud service system includes a cloud management platform and multiple storage nodes. The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes the multiple storage nodes, and the multiple storage nodes provide multiple storage pools. The cloud management platform is used to receive data recovery requests, the data recovery requests including information indicating the source logical volume of the data to be recovered; The cloud management platform is used to determine N target storage pools from the multiple storage pools based on the data recovery request and the input / output I / O capability information of the multiple storage pools, where N is a positive integer greater than or equal to 1, the average I / O capability of the N target storage pools is higher than the I / O capability of the source storage pool, the I / O capability of each storage pool in the multiple storage pools is used to indicate the data recovery capability of each storage pool, and the source storage pool is the storage pool to which the source logical volume belongs in the multiple storage pools; The cloud management platform is used to restore the backup data of the source logical volume to the logical volumes included in the N target storage pools based on the data recovery request.
14. The cloud service system according to claim 13, characterized in that, N≥2, The cloud management platform is specifically used to send a second request to the multiple storage nodes based on the data recovery request. The second request is used to instruct the creation of a target logical volume belonging to the N target storage pools. The plurality of storage nodes are further configured to create a target logical volume across the N target storage pools based on the second request, wherein the target logical volume includes N logical spaces located within the N target storage pools; The cloud management platform is specifically used to restore the backup data of the source logical volume to the N logical spaces within the N target storage pools based on the data recovery request.
15. The cloud service system according to claim 14, characterized in that, The plurality of storage nodes are specifically used for: Based on the second request, obtain the IO capability information of the N target storage pools; Based on the IO capability information of the N target storage pools, determine the size of the i-th logical space to be created in the i-th target storage pool, where 1≤i≤N and i is a positive integer; Based on the size of the i-th logical space, the i-th logical space of the target logical volume is created within the i-th target storage pool to create the target logical volume.
16. The cloud service system according to claim 14 or 15, characterized in that, The N target storage pools include a first storage pool and a second storage pool. The N logical spaces include a first logical space located in the first storage pool and a second logical space located in the second storage pool. The backup data of the source logical volume includes first backup data and second backup data. The plurality of storage nodes are also used to send metadata of the target logical volume to the cloud management platform. The metadata includes a first matching relationship between the first access address of the first logical space and the first storage pool, and a second matching relationship between the second access address of the second logical space and the second storage pool. The cloud management platform is further configured to restore the first backup data of the source logical volume to the first logical space in the first storage pool according to the first access address in the first matching relationship based on the metadata and the data recovery request, and restore the second backup data of the source logical volume to the second logical space in the second storage pool according to the second access address in the second matching relationship.
17. The cloud service system according to claim 16, characterized in that, The cloud management platform is specifically used for: Based on the metadata and the data recovery request, multiple threads are created to write backup data of the source logical volume to the N target storage pools, and the multiple threads include a first thread and a second thread. Based on the first thread, the first backup data of the source logical volume is written to the first logical space in the first storage pool according to the first access address in the first matching relationship. During the process of writing the first backup data to the first logical space through the first thread, based on the second thread, the second backup data of the source logical volume is written to the second logical space in the second storage pool according to the second access address in the second matching relationship.
18. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, each computing device including a processor and memory: The memory is used to store instructions; The processor is configured to, according to the instructions, cause the computing device cluster to perform the method of any one of claims 1 to 8.
19. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method according to any one of claims 1 to 8.
20. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 8.