Method, apparatus and computer program product for adjusting data placement

By adding FTT task entries to the logical block index of the storage system, the inefficiency problem caused by global scanning is solved, efficient data block placement adjustment and fast FTT queries are achieved, and data availability and maintenance efficiency are improved.

CN120276659APending Publication Date: 2025-07-08DELL PROD LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410023481.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When adjusting data placement, existing storage systems need to scan the logical block index globally to calculate the node FTT, resulting in inefficiency and the inability to quickly determine the minimum node FTT, affecting the efficiency of data availability and maintenance operations.

Method used

By adding FTT task entries to the logical block index, including the identifier of the logical block, node FTT and optimization system node FTT, adjusting the placement of data blocks instead of global scanning, improving the tracking efficiency of node FTT.

Benefits of technology

Improves the performance of the storage system, reduces the computing resource consumption of rebalancing operations, and can quickly query the minimum node FTT, ensuring data availability and security of maintenance operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276659A_ABST
    Figure CN120276659A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method, equipment and a computer program product for adjusting data placement. The method includes determining whether a node FTT of a logical block in the storage system satisfies an optimization node FTT corresponding to the logical block, and adding an FTT task entry for the logical block to a logical block index of the storage system when it is determined that the node FTT does not satisfy. The added FTT task entry includes an identifier of a logical block, its node FTT, and a corresponding optimization system node FTT. Further, instead of scanning placement information for each logical block, the storage system may adjust placement of a plurality of data blocks of the respective logical block among a plurality of nodes of the system based on FTT task entries in the logical block index. Therefore, the requirements of scanning the placement information of all the logic blocks and calculating the node FTT in the rebalance operation and the minimum node FTT query are eliminated, and the performance of the storage system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to storage technologies, and more particularly, to methods, apparatuses, and computer program products for adjusting data placement in a storage system. Background Art

[0002] In a storage system, data can be divided into logical chunks for storage. One logical chunk can include multiple physical data blocks. In a distributed storage system including multiple nodes, these data blocks can be placed on different nodes and provide a certain degree of data redundancy, such that the logical chunk can tolerate a certain number of node failures. In other words, in the case where up to this number of nodes fail, the data of the logical chunk can still be recovered.

[0003] For a distributed storage system, data availability is one of the key elements of the service level agreement (SLA) for customers, which can be reflected as the node-level failure to tolerate (FTT). The node FTT for a logical chunk indicates the number of failed nodes that the logical chunk can tolerate. Due to the specific placement of its data blocks in nodes, different logical chunks in a storage system can have different node FTTs. Summary of the Invention

[0004] Embodiments of the present disclosure provide a solution for adjusting data placement.

[0005] In a first aspect of the present disclosure, a method for adjusting data placement is provided, including: determining whether a node failure tolerance (node FTT) of a first logical chunk in a storage system meets an optimized node FTT corresponding to the first logical chunk, the storage system including multiple nodes, and the first logical chunk including multiple data blocks; in response to the node FTT not meeting the optimized node FTT, adding a first FTT task entry for the first logical chunk to a logical chunk index of the storage system, the first FTT task entry including an identifier of the first logical chunk, the node FTT, and an optimized system node FTT; and adjusting the placement of the multiple data blocks among the multiple nodes based on the first FTT task entry.

[0006] In a second aspect of the present disclosure, an electronic device is provided, including a processor and a memory coupled to the processor. The memory has instructions stored therein, and when the instructions are executed by the processor, the device performs actions, which include: determining whether the node FTT of a first logical block in a storage system meets the optimized node FTT corresponding to the first logical block, the storage system including multiple nodes, and the first logical block including multiple data blocks; in response to the node FTT not meeting the optimized node FTT, adding a first FTT task entry for the first logical block to a logical block index of the storage system, the first FTT task entry including an identifier of the first logical block, the node FTT, and an optimized system node FTT; and adjusting the placement of the multiple data blocks among the multiple nodes based on the first FTT task entry.

[0007] In a third aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a computer-readable medium and includes machine-executable instructions that, when executed, cause the machine to perform the method according to the first aspect of the present disclosure.

[0008] Note that the present invention content is provided to introduce, in a simplified form, a selection of concepts that will be further described in the detailed implementation below. The invention content section is not intended to identify the key features or main features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, where:

[0010] Figure 1 A schematic diagram of an example environment in which multiple embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 An example method for adjusting data placement according to some embodiments of the present disclosure is shown;

[0012] Figures 3A - 3F A non-limiting simplified example according to some embodiments of the present disclosure is shown, and actions according to the embodiments of the present disclosure can be performed in the scenario of this example; and

[0013] Figure 4 A schematic block diagram of a device that can be used to implement the embodiments of the present disclosure is shown.

[0014] In all the drawings, the same or similar reference numerals denote the same or similar elements. DETAILED DESCRIPTION

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure.

[0016] As used herein, the term "comprising" and variations thereof are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment". The relevant definitions of other terms will be given in the following description.

[0017] As used herein, depending on the context, the value "meeting the threshold" can, depending on the context, mean that the value is greater than the threshold, greater than or equal to the threshold, equal to the threshold, less than or equal to the threshold, less than the threshold, etc.

[0018] In a storage system, data can be divided into logical blocks for storage, and each logical block includes a plurality of physical data blocks on a storage disk. Thus, the logical block can be physically stored on different storage nodes and provide a certain degree of data redundancy. In this way, when up to a certain number of nodes fail, the data of the logical block can be recovered. The number of failed nodes that a logical block can tolerate is called the node FTT, which can reflect the availability of data in the storage system and is an important part of the SLA for customers.

[0019] Since the specific placement of each logical block among multiple nodes is different, different logical blocks in the storage system can have different node FTTs. In addition, the storage system can use multiple types of logical blocks to store data, and each type of logical block has a specific structure. For example, according to different sources, uses, and sizes of data, the storage system can store data as different types of logical blocks. Since different types of logical blocks can have different structures, the optimal node FTTs that different types of logical blocks can achieve in the same system are different.

[0020] When the node FTT of a logical block is lower than the optimal FTT it can achieve, the storage system can perform a rebalance operation on it. The rebalance operation relocates the data blocks of the logical block so that these data blocks are more evenly distributed among the nodes of the storage system. The storage system can use a logical block index to perform this operation. The logical block index can use, for example, a lookup table or other appropriate data structures. Information about the logical block can be stored in the corresponding entries of the logical block index, and this information can include the locations of the data blocks of the logical block in the storage system and the corresponding health status (e.g., whether the node where the data block is located has failed).

[0021] In a traditional storage system, to perform a rebalancing operation, it is necessary to scan the entries of all logical blocks in the logical block index, and calculate the node FTT of each logical block based on the placement of the logical block and the health status of the data block. If the calculated node FTT is lower than the corresponding optimal node FTT, the storage system will attempt to rebalance the data block of the logical block to a location that is as evenly distributed as possible among the nodes. Such a global scan needs to be repeatedly executed, and the node FTT needs to be calculated for each logical block, so the overhead is high and the efficiency is limited.

[0022] On the other hand, when performing maintenance operations on a node, such as replacing a disk, a temporary node outage, or node removal, as a safety check item before performing the operation, the storage system needs to check whether the minimum node FTT of the logical blocks in the current system meets the data availability requirements to determine whether the maintenance can be performed on the node and / or how many nodes can be maintained. However, a global scan of all logical blocks takes a lot of time and cannot quickly return the accurate minimum node FTT.

[0023] Some storage systems determine whether the data availability requirements are met by checking whether there are entries for repair tasks in the logical block index. As a non-limiting example, it is assumed that when a logical block cannot tolerate the failure of at least one node, an entry for the repair task for the logical block is added to the logical block index. In this case, if there is a repair task in the logical block index, the minimum node FTT is zero and the maintenance operation is not allowed. If there is no repair task, the maintenance operation on one node can be performed. However, this method is not perfect. For example, in the above example, if there is no repair task, the storage system cannot determine whether the minimum node FTT of the system is greater than one, so it cannot perform maintenance operations on more than one node.

[0024] To at least partially solve the above problems and other potential problems, embodiments of the present disclosure propose a scheme for adjusting data placement. When it is determined that the node FTT of a logical block in the storage system does not meet the optimized node FTT corresponding to the logical block (for example, the optimal node FTT based on the number of nodes and the logical block structure, or the optimized value considering additional constraints), an FTT task entry for the logical block is added to the logical block index of the storage system. The added FTT task entry includes the identifier of the logical block, its node FTT, and the corresponding optimized system node FTT. Furthermore, instead of scanning the placement information for each logical block, the storage system can adjust the placement of multiple data blocks of the corresponding logical block among multiple nodes in the system based on the FTT task entries in the logical block index.

[0025] In this way, the logical blocks that need to be improved by node FTT can be tracked, enabling the storage system to efficiently perform rebalancing operations on the data blocks corresponding to these logical blocks. This solution eliminates the need for rebalancing operations to repeatedly scan the placement information of all logical blocks and calculate their node FTTs, improving the performance of the storage system. Additionally, in some embodiments, based on the tracking of node FTT, the minimum node FTT for a logical block can be quickly queried.

[0026] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. First, refer to Figure 1 , which shows a schematic diagram of an example environment 100 in which multiple embodiments of the present disclosure can be implemented. The environment 100 includes an example storage system, which includes multiple storage nodes 110-1, 110-2, …, 110-K (individually or collectively referred to as nodes 110). Each storage node 110 includes one or more storage disks (not shown) and may include services and processes for storing and managing data on the disks. The data stored on the nodes 110 can include user data and system metadata. In the storage system, the data can be divided into logical blocks for storage. A logical block includes multiple physical data blocks that can be stored on different nodes. For example, data blocks 120-1, 120-2, …, 120-L (individually or collectively referred to as data blocks 120) can be multiple data blocks of the same logical block.

[0027] Figure 1 Also shown in

[0028] In a storage system, multiple types of logical blocks can be used to store data. Each type of logical block can have a specific structure and a corresponding optimal node FTT. In some embodiments, one type of logical block can be based on multiple mirror copies. This type of logical block can include N (e.g., 3) data blocks, where each data block stores one mirror copy of the same data. Each copy is stored on a different node, such that this logical block can tolerate up to N - 1 node failures, i.e., its optimal node FTT is N - 1. For ease of explanation, this type of logical block will be referred to as a normal logical block hereinafter. For example, this type of logical block can be used to store system metadata in a storage system.

[0029] In some embodiments, a storage system can utilize erasure coding (EC) technology to enhance data availability. The m + n EC mode supports using a specific algorithm to calculate n parity data blocks for m original data blocks, thereby obtaining m + n encoded data blocks. Any m of the m + n data blocks can be used to calculate all the original data. Thus, one EC data copy includes m original data blocks and n parity data blocks. For example, in a storage system that utilizes EC technology, user data can be stored based on the logical block type of the EC mode. For ease of explanation, such a logical block will be referred to as an EC logical block hereinafter.

[0030] In such some embodiments, one type of EC logical block (hereinafter referred to as the first type of EC logical block) can include one EC data copy and N - 1 mirror copies of the original data. For example, this type of EC logical block can include m original data blocks and n parity data blocks, as well as N - 1 mirror copies of the m original data blocks. For example, this type of logical block can be used for writing small - sized data multiple times. For example, during each write, the computing device 130 can write data to the m original data blocks on one or more nodes 110, and simultaneously write N - 1 copies of this data to another N - 1 groups of m data blocks respectively. In this way, the availability of these data can be enhanced using the mirror copies before the m original data blocks can be erasure - encoded.

[0031] Furthermore, when the m original data blocks are full, the computing device 130 can encode them to calculate and write n parity data blocks to one or more nodes 110, thereby obtaining an EC copy of the data. Further, after the encoding is completed, the storage system can delete the N - 1 mirror copies of the N - 1 original data of this logical block, thereby converting this logical block into another type of data block (hereinafter referred to as the second type of EC logical block). The second type of EC logical block includes one EC data copy of m + n data blocks.

[0032] In some embodiments, when writing larger data that can fill m data blocks, the storage system can also directly use the second type of EC logical blocks to store data. In this case, the storage system directly writes and encodes m raw data blocks without using the first type of EC logical blocks as a transition.

[0033] The optimal FTT of the encoded EC logical block depends on the number of available nodes in the storage system and the EC mode used. For example, for a storage cluster that includes at least m + n available nodes and uses an EC mode of m + n, when the encoded EC logical block has an optimal placement, each data block in the EC data replicas of the logical block can be placed in a different node. Therefore, in this storage cluster, the optimal node FTT of the encoded EC logical block is equal to n. In practical implementations, for example, because the storage cluster includes fewer than m + n available nodes, the optimal node FTT of the encoded EC logical block may be less than m.

[0034] To track relevant information of logical blocks, the computing device 130 can maintain a logical block index 140 in the form of one or more tables, for example. The logical block index 140 can store entries regarding the placement of logical blocks, which record the locations of the respective data blocks of the corresponding logical blocks on the nodes 110 and their health status.

[0035] For illustrative purposes, the logical block index 140 is shown as a separate entity associated with the computing device 130, but it should be understood that this is only a logical schematic representation. Depending on the implementation, the logical block index 140 can be stored centrally or distributively at multiple storage locations in the storage system. For example, as part of the storage system metadata, multiple copies of the logical block index 140 can also be distributed on multiple nodes 110. Embodiments of the present disclosure are not limited by the physical storage location of the logical block index 140.

[0036] The computing device 130 can handle the placement of logical blocks on the nodes 110 and maintain the status of their data blocks (e.g., using a logical block manager implemented thereon). For example, the computing device 130 can create various types of logical blocks and distribute their data blocks among the nodes 110. For example, the computing device 130 can encode EC logical blocks. The successful creation of a normal logical block based on N mirror replicas requires that the block can tolerate N - 1 node failures. On the other hand, when the encoded EC logical block can tolerate at least one node failure, its data does not need to be recovered.

[0037] In addition, if a certain node in the nodes 110 fails, the health status of the data blocks stored in that node will become bad. The logical blocks that own these blocks may still have an optimal placement, but some data blocks are lost, which causes the node FTT of these logical blocks to decrease. In response to this, the computing device 130 can perform a recovery action for the logical block.

[0038] For example, when a node 110 or its disk in the computing device 130 fails, the computing device 130 can change the data block status in the logical block index and create a repair task for the affected logical blocks. The repair task can also be recorded in the logical block index. For example, for ordinary logical blocks, the computing device 130 can reconstruct the lost data blocks by copying the remaining available mirror copies. For example, for encoded EC logical blocks, the computing device 130 can utilize the remaining available data blocks to perform erasure code decoding to recover the lost data blocks. When the recovered ordinary logical block can tolerate N - 1 node failures, or the encoded EC logical block can tolerate at least one node failure, the computing device 130 can remove the repair task entry.

[0039] For a specific logical block, since its data blocks may not be optimally placed in the nodes, the node FTT of the specific logical block may be lower than the optimal FTT that the logical blocks of its type can achieve. In addition, as described above, the recovered logical blocks may also not be in an optimal placement state, such that their node FTT can be increased by adjusting the placement of their data blocks. For example, an encoded or recovered EC logical block may just tolerate at most one node failure, yet according to the number of nodes in the storage system and the mode used by the EC logical block, the logical block should be able to tolerate more node failures after its placement is optimized. In this regard, the computing device 130 can also re - balance the physical placement of logical blocks among the nodes according to embodiments of the present disclosure, as will be described in more detail below.

[0040] The architecture and functions in the example environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure. There may be other devices, systems, components, etc. not shown in the example environment 100. For example, there may be multiple clients in the example environment for users to access the storage system. In addition, embodiments of the present disclosure can also be applied to other environments with different structures and / or functions.

[0041] Figure 2 A flowchart of an example method 200 for adjusting data placement according to some embodiments of the present disclosure is shown. The example method 200 can be performed, for example, by the computing device 130 as Figure 2 shown. It should be understood that the method 200 may also include additional actions not shown, and the scope of the present disclosure is not limited in this regard. The method 200 will be described in detail below in conjunction with Figure 1 the example environment 100.

[0042] At 210, determine whether the node FTT of the first logical block in the storage system satisfies the optimized node FTT corresponding to the first logical block. The storage system includes multiple nodes, and the first logical block includes multiple data blocks. For example, the computing device 130 may determine Figure 1 whether the node FTT of a certain logical block stored in the storage system in

[0043] satisfies the optimized node FTT corresponding to the logical block. In some embodiments, the storage system may store multiple types of logical blocks, and the optimized node FTT corresponding to a specific logical block corresponds to the type of the logical block. For example, the computing device 130 may use the optimal node FTT that a certain type of logical block can achieve in the storage system as the optimized node FTT for the type of logical block.

[0044] For example, as described above in connection with Figure 1 shown, the optimized node FTT for a normal logical block based on N replicas may be N - 1, and the optimized node FTT for an encoded EC logical block based on the m + n mode may be n when the number of nodes is sufficient. In some embodiments, the computing device 130 may also consider additional limiting factors on the basis of the optimal node FTT to determine the optimized node FTT for actual application. The optimized node FTT of the corresponding type may be recorded in the metadata of the storage system, for example.

[0045] The computing device 130 may perform the determination action of 210 when creating, encoding, and repairing data blocks, or when the node 110 changes, as will be described in more detail later in connection with Figures 3A - 3F shown. In some embodiments, based on the data block placement and health status of the first logical block recorded in the logical block index 110, the computing device 130 may determine the node FTT of the first logical block, and then determine whether the node FTT satisfies the corresponding optimized node FTT.

[0046] At 220, in response to the node FTT not satisfying the optimized node FTT, add a first FTT task entry for the first logical block to the logical block index of the storage system. For example, in response to determining at 210 that the current node FTT of a certain logical block does not satisfy (e.g., is lower than) its corresponding optimized node FTT, the computing device 130 may add an FTT task entry for the logical block to the logical block index 140.

[0047] The FTT task entry for the first logical block includes the identifier of the first logical block, its node FTT, and the optimized system node FTT corresponding to the first logical block. In some embodiments, the FTT task entry can be implemented as part of the entry for the logical block in the logical block index 140. In some embodiments, the FTT task entry can be stored as a separate list to facilitate scanning of the FTT task entry. Those skilled in the art will understand that in a specific implementation, various other implementations of the FTT task entry can also be adopted.

[0048] As a non-limiting example, the FTT task entry can be stored in the logical block index 140 as a key-value pair. The key can include the identifier of the corresponding logical block and the node FTT of the logical block. The value can include the optimized FTT corresponding to the logical block. Such an entry indicates that a rebalancing action should be performed for the execution of the corresponding logical block so that its node FTT is increased as much as possible to the optimized FTT recorded in the entry.

[0049] At 230, the placement of multiple data blocks of the first logical block among multiple nodes of the storage system is adjusted based on the first FTT task entry. For example, the computing device 130 can adjust the placement of multiple data blocks of the logical block among multiple nodes of the storage system based on the FTT task entry for the logical block.

[0050] In some embodiments, the computing device 130 can adjust the placement of multiple data blocks of the logical block in the nodes 110 in a way that increases the node FTT of the logical block. For example, the computing device 130 can place each data block of the logical block on as many nodes 110 as possible. In this way, the distribution of the logical block is more uniform and dispersed, thus being able to tolerate more node failures.

[0051] If the increased node FTT meets (reaches) the optimized node FTT corresponding to the logical block, it can be considered that the FTT task for the logical block is completed. In this case, the computing device 130 can remove the corresponding FTT task entry from the logical block index 140. In some embodiments, after performing the rebalancing action for the logical block, the node FTT of the logical block may still not meet the optimized node FTT. In this case, the computing device 130 can update the node FTT of the logical block in the entry and retain the entry in the logical block index 140.

[0052] In some embodiments, the computing device 130 may scan the logical block index 140 to search for FTT task entries for logical blocks in the storage system. If an FTT task entry (e.g., the first FTT task entry described above) is detected during the scan, the computing device 130 may perform a rebalancing action on the logical block targeted by the entry, i.e., adjust the placement of the multiple data blocks of the logical block among the storage nodes. In this way, the computing device 130 does not need to repeatedly scan the physical placement information of all logical blocks and the health status of each data block. Instead, the computing device 130 may implement the FTT task scanning function, perform the corresponding logical block rebalancing action based on the scanned FTT task entry, and may stop scanning when there is no FTT task entry in the logical block index 140.

[0053] Using method 200, logical blocks with a need to improve node FTT can be tracked, enabling the storage system to efficiently perform rebalancing operations on the data blocks of the corresponding logical blocks. This solution eliminates the need to scan the placement information of all logical blocks and calculate their node FTT in rebalancing operations and minimum node FTT queries, thereby saving computing resources for balancing data placement in the storage system and improving the performance of the storage system.

[0054] In some embodiments, when creating a logical block, the computing device 130 may determine whether the node FTT meets the optimized node FTT corresponding to the logical block, and add a corresponding FTT task entry to the logical block index when the node FTT does not meet the optimized node FTT. For example, if the optimized node FTT for an encoded EC logical block is n, and the node FTT of the created second type of EC logical block is less than n, the computing device 130 may add an FTT task entry for the EC logical block.

[0055] In some embodiments, when performing EC encoding on a first type of EC logical block to write computed parity data blocks, the computing device 130 may determine whether the FTT of the encoded logical block is less than the optimized node FTT for the EC logical block. If so, the computing device may add an FTT task entry for the EC logical block.

[0056] In some embodiments, changes regarding node 110 may occur in the storage system, such as addition, removal, and temporary shutdown of nodes, etc. Node changes may cause the optimized node FTT for logical blocks to change. For example, an EC logical block may be able to tolerate more or fewer node failures at most due to changes in node data. In such embodiments, the computing device 130 may implement a scanning function to check the FTT, which scans all logical blocks when needed and updates the FTT task entries accordingly. This scanning function can monitor the available nodes where data blocks can be placed and the target node FTT for the storage system.

[0057] This target node FTT reflects the number of node failures that all logical blocks in the storage system are expected to tolerate at least. The target node FTT can be determined based on the number of nodes in the storage system and the EC mode used by its EC logical blocks, and it satisfies the following rules:

[0058] Target node FTT ≥ 1;

[0059] The optimal node FTT for the encoded EC logical block ≥ the target node FTT;

[0060] The optimal node FTT for the ordinary logical block = N - 1 ≥ the target node FTT.

[0061] According to the above relationships, all logical blocks in the storage system are designed to tolerate at least a certain degree of node failures, and the optimal node FTT that each specific type of logical block can achieve may be higher than the target node FTT. Among them, the ordinary logical block depends on the number of mirror copies N in the logical block, and the optimal number of nodes for the encoded EC logical block depends on the EC mode it uses (i.e., the number of data blocks m and the number of parity blocks n) and the number of available nodes in the storage system.

[0062] In such embodiments, in response to a new node being added to the storage system, the computing device 130 may determine whether there is a change in a certain type of optimized node FTT for the logical blocks in the storage system. If such a change exists, the computing device 130 may update the FTT task entries for the logical blocks in the logical block index 140 based on this change. Otherwise, the computing device 130 may not perform related actions.

[0063] Figures 3A - 3F A non - restrictive simplified example according to some embodiments of the present disclosure is shown, which shows multiple scenarios of node changes in the storage system, and actions according to the embodiments of the present disclosure can be performed in these scenarios. In this example, the optimal node FTT that a specific type of logical block can achieve in the storage system is used as the optimized node FTT for this type. The storage system shown in FIG. 3 may be Figure 1An example implementation of the storage system will be described below in the context of Figure 1 environment 100. Figures 3A - 3F Examples will be given.

[0064] As Figure 3A shown, in scenario 300A, nodes 310-1 to 310-K (individually or collectively referred to as nodes 310) of a storage system including K available nodes are shown, and K = m + n - 1. In addition, the storage system uses EC logic blocks in m + n mode to store data. In this case, at least two of the data blocks in the EC replicas of the encoded EC logic blocks have to be located on the same node. For example, as shown by data blocks 320-1 and 320-2 on node 310-1. Thus, in this storage system, the optimized node FTT for the EC logic blocks cannot reach n supported by the EC mode.

[0065] Figure 3A The logical block index 340 of the storage system is also shown, which can be Figure 1 an example implementation of the logical block index 140. Assume that the other data blocks of the logical block where data blocks 320-1 and 320-2 are located are evenly distributed on the other nodes 310 of the storage system. This logical block is in the best placement currently achievable, so there is no FTT task entry for this logical block in the logical block 340. For the sake of clarity in illustration, other types of entries for this logical block and entries for other logical blocks in the Figures 3A - 3F are omitted.

[0066] Later, as Figure 3B shown, in scenario 300B, a new node 311 is added to the storage system. In response to the addition of this node, the computing device 130 can detect that the addition of this node makes the number of available nodes reach m + n. Thus, the storage system can support placing each data block in the EC replicas of the encoded EC logic blocks on different nodes. Therefore, the addition of node 311 raises the optimized node FTT for the encoded EC logic blocks to n.

[0067] In response to detecting that the addition of node 311 has changed the above-mentioned optimized node FTT, the computing device 130 can scan all the logical blocks to check their node FTT. If it is detected that the node FTT of a certain logical block does not meet the corresponding updated optimal node FTT, the computing device 130 can add an FTT task entry for this logical block to the logical block index 340. The FTT task entry includes the identifier of this logical block and the node FTT, as well as the updated optimized FTT corresponding to this logical block. For example, the computing device 130 can also update the information in the existing FTT task entries in the logical block index according to the check.

[0068] Furthermore, the computing device 130 can use the FTT task scanning function as described above to scan the FTT task entries in the logical block index 340 and execute the corresponding FTT tasks for the scanned FTT task entries, such as rebalancing actions for the corresponding logical blocks, etc. For example, since data blocks 320-1 and 320-2 are located on the same node, the computing device 130 will detect that the node FTT of the EC logical block it belongs to is less than n. Therefore, the FTT task entry 341 for this EC logical block is added to the logical block index 340.

[0069] The computing device 130 can then adjust the physical distribution of the logical block based on the FTT task entry 341. For example, in response to its FTT scanning function detecting the FTT task entry 341. As shown in 3C, in scenario 300C, the computing device 130 can relocate the data block 320-2 to the newly added node 311. As described above, the other data blocks in this EC logical block have already been distributed on different nodes 310. After the relocation, the node FTT of this logical block reaches n. Therefore, the computing device 130 can remove the FTT task entry 341 from the logical block index 340.

[0070] For example, due to a full storage space of a certain node or other reasons, as Figure 3D shown, in scenario 300D, another node 312 can also be added to the storage system at a later time. However, the addition of node 312 will not further improve the optimized node FTT for the logical block. Therefore, the computing device 130 does not need to scan all the logical blocks after such a node is added.

[0071] Similarly, the removal of a node may also change the optimized FTT for the logical block, for example, causing the optimized FTT to decrease. In addition, the removal of a node makes the data blocks located on that node unavailable and therefore needs to be recovered. In some embodiments, in response to a node being removed from the storage system, the computing device 130 can perform a recovery of the data in the storage system. For example, the computing device 130 can update the status of each data block to determine the logical blocks that need to be recovered, add repair tasks for them (such as repair task entries in the logical block index 340) and execute them. As described above in connection with Figure 1 what has been said, according to the type of the logical block, the computing device 130 can recover the data in a corresponding manner.

[0072] After data recovery is completed, computing device 130 may remove the repair task. Further, computing device 130 may update the FTT task entry for a logical block in the logical block index based on the placement of the recovered data across multiple nodes. For example, computing device 130 may determine whether the removal of a node changes the optimized FTT for a logical block, and scan the logical block to check whether its node FTT meets the changed optimized FTT. Based on the check, computing device 130 may add an FTT task entry accordingly, remove the corresponding FTT task entry when the node FTT of the logical block meets the changed optimized FTT, or update the information of the FTT task entry.

[0073] As Figure 3E shown, in scenario 300E, node 310-1 is removed from the storage system, causing data block 320-1 of the EC logical block to be lost. Thus, computing device 130 may calculate the value of data block 320-1 based on the remaining logical blocks in the EC logical block to recover it. In this example, the recovered data block 320-1' is first placed on node 311, such that the node FTT of the EC logical block it belongs to is lower than n.

[0074] On the other hand, due to the previous addition of node 312, the storage system still has m + n nodes, such that the optimized FTT for the EC logical block remains n. Thus, computing device 130 will detect that the logical block node FTT needs to be increased and add FTT task entry 342 to logical block index 340. After performing the corresponding FTT task, as Figure 3F shown, in scenario 300F, data block 320-1' is re-placed on node 312, such that the logical block has an optimal distribution. Then, computing device 130 may remove FTT task entry 342.

[0075] In some embodiments, a user may cause a node to be temporarily disabled (e.g., taken offline) in order to perform various maintenance actions. Such a temporary disablement may also cause the optimized FTT of the node to change. The changed optimized FTT may affect the decision-making for various operations during node failure. Thus, in response to performing an action that temporarily disables a node in the storage system, computing device 130 may update the optimized node FTT for the type of logical block in the storage system based on the status of the remaining nodes. If computing device 130 determines that the temporary disablement does not change the optimized node FTT, its value may be kept unchanged. Additionally, since the data on the temporarily disabled node will be back online later, the data on it is not actually lost. Thus, computing device 130 will not scan the logical block to update the status of its data blocks to avoid unnecessary data recovery operations.

[0076] Based on the node FTT recorded in the FTT task entry, the minimum node FTT for the logical block in the storage system can be quickly queried. This minimum node FTT can reflect how many nodes in the storage system can currently be safely maintained. Thus, the computing device 130 can quickly determine whether the maintenance operation to be performed on the node 110 is safe.

[0077] In some embodiments, the computing device 130 may receive a query for the minimum value in the node FTT for the logical block in the storage system. In response to receiving the query, the computing device determines the minimum value based on at least one of the following: various types of optimized node FTTs for the logical block in the storage system, such as the target node FTT for the storage system as described above, and the FTT task entry for the logical block in the logical block index.

[0078] In some embodiments, the computing device 130 may determine the value that ensures the safety of the maintenance action as the minimum value. For example, in one exemplary implementation, if it is determined that the scan action of the logical block node FTT has not been completed, the computing device 130 may determine the minimum value to be 0 when there is a repair task, otherwise the minimum value is 1.

[0079] Furthermore, in this example, if the aforementioned scan action has been completed, but the update process of the data block status due to node failure, etc. has not been completed, the actual FTT value of the logical block is not yet clear. In this case, the computing device 130 may check whether there is any FTT task entry in the logical block index.

[0080] If the FTT task entry exists, the computing device 130 may take 0 as the safe minimum node FTT value. That is, at this time, no maintenance operation should be performed on the node because it cannot be ensured that the storage system can tolerate node failures at this time. If the FTT task entry does not exist, the computing device 130 may take the larger value of the following two as the minimum node FTT value: 0, or the value obtained by subtracting the number of failed nodes from the target node FTT.

[0081] In addition, in this exemplary implementation, if the scan for FTT has been completed and the status of all data blocks has been updated, the computing device 130 may also check whether there is an FTT task entry in the logical block index. If the FTT task entry exists, the computing device 130 may take the smaller value of the following two as the minimum node FTT value: the target node FTT, or the minimum node FTT value recorded in the FTT task entry. If the FTT task entry does not exist, the computing device 130 may take the target node FTT as the minimum node FTT value.

[0082] Figure 4FIG. 0 shows a schematic block diagram of a device 400 that can be used to implement embodiments of the present disclosure. The device 400 can be a device or apparatus described in embodiments of the present disclosure. As Figure 4 shown, the device 400 includes a central processing unit (CPU) 401, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 402 or computer program instructions loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404. Although not shown in Figure 4 , the device 400 may further include a coprocessor.

[0083] Multiple components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0084] Each of the above-described methods or processes can be executed by the processing unit 401. For example, in some embodiments, the method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the CPU 401, one or more steps or actions in the above-described methods or processes can be executed.

[0085] In some embodiments, the above-described methods and processes can be implemented as a computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.

[0086] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0087] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0088] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages and conventional procedural programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0089] These computer - readable program instructions can be provided to the processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. The computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions that implement various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0090] The computer - readable program instructions can also be loaded onto a computer, other programmable data - processing apparatus, or other devices such that a series of operation steps are executed on the computer, other programmable data - processing apparatus, or other devices to produce a computer - implemented process, so that the instructions executed on the computer, other programmable data - processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0092] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skill in the art in the field to understand the embodiments disclosed herein.

Claims

1. A method for adjusting data placement, comprising: Determining whether a node failure tolerance number of nodes FTT of a first logical block in a storage system meets an optimized node FTT corresponding to the first logical block, the storage system including a plurality of nodes, and the first logical block including a plurality of data blocks; In response to the node FTT not meeting the optimized node FTT, adding a first FTT task entry for the first logical block to a logical block index of the storage system, the first FTT task entry including an identifier of the first logical block, the node FTT, and the optimized system node FTT; and Adjusting the placement of the plurality of data blocks among the plurality of nodes based on the first FTT task entry.

2. The method according to claim 1, wherein adjusting the placement of the plurality of data blocks among the plurality of nodes based on the FTT includes: Adjusting the placement of the plurality of data blocks in a manner that increases the node FTT; And In response to the increased node FTT meeting the optimized node FTT, removing the first FTT task entry from the logical block index.

3. The method according to claim 1, adjusting the placement of the plurality of data blocks based on the first FTT task entry includes: Scanning the logical block index to search for FTT task entries for logical blocks; And In response to detecting the first FTT task entry in the scan, adjusting the placement of the plurality of data blocks.

4. The method according to claim 1, wherein the storage system stores a plurality of types of logical blocks, and the optimized node FTT corresponding to a corresponding logical block corresponds to the type of the corresponding logical block.

5. The method according to claim 4, further comprising: In response to a new node being added to the storage system, determining whether there is a change in the optimized node FTT for the type of logical block in the storage system; And In response to the existence of the change, updating the FTT task entries for logical blocks in the logical block index based on the change.

6. The method according to claim 4, further comprising: In response to a node being removed from the storage system, performing recovery for the data in the storage system; And After the recovery is completed, updating the FTT task entries for logical blocks in the logical block index based on the placement of the recovered data on the plurality of nodes.

7. The method according to claim 4, further comprising: In response to performing an action that temporarily disables nodes of the plurality of nodes, updating the optimized node FTT for the type of logical block in the storage system based on the remaining nodes among the plurality of nodes.

8. The method according to claim 1, further comprising: Receiving a query for a minimum value of the node FTT of a logical block in the storage system; And In response to the query, determining the minimum value based on at least one of: The optimized node FTT for the type of logical block in the storage system, The target node FTT for the storage system, and The FTT task entries for logical blocks in the logical block index.

9. An electronic device, comprising: a processor; and a memory coupled to the processor, the memory having instructions stored therein, the instructions, when executed by the processor, causing the device to perform operations, the operations including: determining whether a node failure tolerance number node FTT of a first logical block in a storage system meets an optimized node FTT corresponding to the first logical block, the storage system including a plurality of nodes, and the first logical block including a plurality of data blocks; in response to the node FTT not meeting the optimized node FTT, adding a first FTT task entry for the first logical block to a logical block index of the storage system, the first FTT task entry including an identifier of the first logical block, the node FTT, and the optimized system node FTT; and adjusting placement of the plurality of data blocks among the plurality of nodes based on the first FTT task entry.

10. The device according to claim 9, wherein adjusting placement of the plurality of data blocks among the plurality of nodes based on the FTT includes: adjusting the placement of the plurality of data blocks in a manner that increases the node FTT; and in response to the increased node FTT meeting the optimized node FTT, removing the first FTT task entry from the logical block index.

11. The device according to claim 9, adjusting placement of the plurality of data blocks based on the first FTT task entry includes: scanning the logical block index to search for FTT task entries for logical blocks; and in response to detecting the first FTT task entry during the scan, adjusting the placement of the plurality of data blocks.

12. The device according to claim 9, wherein the storage system stores a plurality of types of logical blocks, and the optimized node FTT corresponding to a corresponding logical block corresponds to the type of the corresponding logical block.

13. The device according to claim 12, the operations further including: in response to a new node being added to the storage system, determining whether there is a change in the optimized node FTT for a type of logical block in the storage system; and in response to there being the change, updating FTT task entries for logical blocks in the logical block index based on the change.

14. The device according to claim 12, the operations further including: in response to a node being removed from the storage system, performing recovery for data in the storage system; and after the recovery is completed, updating FTT task entries for logical blocks in the logical block index based on placement of the recovered data on the plurality of nodes.

15. The device according to claim 12, the operations further including: in response to performing an operation that temporarily disables nodes of the plurality of nodes, updating the optimized node FTT for a type of logical block in the storage system based on the remaining nodes among the plurality of nodes.

16. The device according to claim 9, the operations further including: receiving a query for a minimum value of a node FTT of a logical block in the storage system; and In response to the query, determine the minimum value based on at least one of the following: The optimized node FTT for the type of logical block in the storage system, The target node FTT for the storage system, and The FTT task entry for the logical block in the logical block index.

17. A computer program product, tangibly stored on a computer-readable medium and including machine-executable instructions that, when executed, cause a machine to: Determine whether the node failure tolerance number node FTT of a first logical block in a storage system meets the optimized node FTT corresponding to the first logical block, the storage system including a plurality of nodes, and the first logical block including a plurality of data blocks; In response to the node FTT not meeting the optimized node FTT, add a first FTT task entry for the first logical block to the logical block index of the storage system, the first FTT task entry including an identifier of the first logical block, the node FTT, and the optimized system node FTT; and Adjust the placement of the plurality of data blocks among the plurality of nodes based on the first FTT task entry.

18. The computer program product according to claim 17, wherein adjusting the placement of the plurality of data blocks among the plurality of nodes based on the FTT includes: Adjusting the placement of the plurality of data blocks in a manner that increases the node FTT; And In response to the increased node FTT meeting the optimized node FTT, remove the first FTT task entry from the logical block index.

19. The computer program product according to claim 18, adjusting the placement of the plurality of data blocks based on the first FTT task entry includes: Scanning the logical block index to search for FTT task entries for logical blocks; And In response to detecting the first FTT task entry during the scan, adjusting the placement of the plurality of data blocks.

20. The computer program product according to claim 17, wherein the storage system stores a plurality of types of logical blocks, and the optimized node FTT corresponding to a respective logical block corresponds to the type of the respective logical block.