Method and system for recovering storage space in detachable storage system
By identifying and optimizing source and target nodes in a decomposed storage system, and utilizing QPC and REC to optimize the storage space reclamation process, the problems of uneven resource utilization and QoS impact in existing technologies are solved, achieving efficient storage space management.
Patent Information
- Application Number
- CN202411510371.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2024-10-28
- Publication Date
- 2025-10-24
AI Technical Summary
Existing decomposed storage systems cannot effectively manage cluster-level durability when compressing and reclaiming storage space, leading to uneven resource utilization and potentially affecting Quality of Service (QoS), with no predictable impact on QoS.
By retrieving multiple data levels and durability levels of storage nodes, determining the range of variation based on workload and parameters, identifying source and target nodes, and optimizing the storage space reclamation process using the Quality of Service Penalty (QPC) and Reclamation Efficiency (REC) coefficients.
It enables dynamic and effective storage space reclamation, optimizes resource utilization, reduces the impact on service quality, and improves the overall efficiency of the storage system.
Smart Images

Figure CN120832084A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present subject matter relates generally to disaggregated storage systems, and more particularly, but not exclusively, to methods and systems for dynamic reclaiming of storage space in disaggregated storage systems. BACKGROUND
[0002] Currently, flash-based storage systems are being widely adopted in data centers. The adoption of flash-based storage devices is accelerating due to multiple factors such as the need for larger storage space for emerging artificial intelligence (AI) / machine learning (ML) type workloads, availability of cheaper flash-based storage devices in the form of quad-level cell (QLC), penta-level cell (PLC), and the like, and shift from on-premise storage devices to cloud-based storage devices, and the like.
[0003] The accelerated adoption of flash in disaggregated storage systems has led to the emergence of multi-tiered storage architecture. Different tiers of the storage architecture provide different levels of endurance. The density of the storage nodes has also reached the PB (peta-byte) level. In disaggregated storage architecture, storage nodes can be added at will. These factors can require endurance management at the cluster level, which can require out of place writes at the cluster level. Out of place writes generate redundant data in the storage space. Therefore, compaction is a process of better utilizing the existing storage space by removing redundant data and cleaning up invalid objects. In existing systems, compaction starts with the identification of source storage segments and target storage segments. The source storage segment is identified based on the amount of valid data to be moved. If the amount of valid data to be moved is less, then there is a higher probability of selecting that source data segment for compaction. The selection of the target segment is based on the existing cluster level allocation policies for managing endurance and available space. Further, the data movement during compaction consumes cluster resources (e.g., CPU, network, and the like). Therefore, the amount of data to be moved is a direct factor affecting the amount of cluster resources utilized. Further, because this is a dynamic / real-time process and compaction can also impact QoS in unknown ways, the impact on quality of service (QoS) cannot be predicted. Therefore, this background operation of compaction / space reclamation cannot also impact the QoS functionality of the storage cluster.
[0004] The information disclosed in the Background section of the present disclosure is only for the purpose of strengthening the understanding of the general background of the present invention, and should not be regarded as acknowledging or implying in any form that this information constitutes prior art known to those skilled in the art. SUMMARY
[0005] In an embodiment, the present disclosure relates to a method for dynamically reclaiming storage space in a disaggregated storage system. The method includes retrieving a plurality of data levels and a plurality of durability levels of a plurality of storage nodes. The plurality of data levels correspond to an amount of data present in each storage segment of one or more storage segments associated with each of the plurality of storage nodes. The method further includes determining a first range of variation based on a workload and one or more first parameters associated with the plurality of data levels. The workload is associated with the plurality of storage nodes. The method further includes determining a second range of variation based on the workload and one or more second parameters associated with the plurality of durability levels. The method further includes identifying one or more source nodes among the plurality of storage nodes based on the first range of variation. The method further includes identifying one or more target nodes among the plurality of storage nodes based on the second range of variation. The method further includes identifying a set of storage node pairs among the one or more source nodes and the one or more target nodes based on a quality of service penalty coefficient (QPC), wherein the QPC corresponds to a number of network hops required to establish a data path between a storage node pair of the set of storage node pairs. Finally, the method includes performing a reclamation of at least one storage segment among the set of storage node pairs.
[0006] In another embodiment, a system for dynamically reclaiming storage space in a disaggregated storage system is provided. The system includes a memory and a processor. The processor is configured to retrieve a plurality of data levels and a plurality of durability levels of a plurality of storage nodes. The plurality of data levels correspond to an amount of data present in each storage segment of one or more storage segments associated with each of the plurality of storage nodes. The processor is configured to determine a first range of variation based on a workload and one or more first parameters associated with the plurality of data levels, wherein the workload is associated with the plurality of storage nodes. The processor is configured to determine a second range of variation based on the workload and one or more second parameters associated with the plurality of durability levels. The processor is configured to identify one or more source nodes among the plurality of storage nodes based on the first range of variation. The processor is configured to identify one or more target nodes among the plurality of storage nodes based on the second range of variation. The processor is configured to identify a set of storage node pairs among the one or more source nodes and the one or more target nodes based on a quality of service penalty coefficient (QPC). The QPC corresponds to a number of network hops required to establish a data path between a storage node pair of the set of storage node pairs. Finally, the processor is configured to perform a reclamation of at least one storage segment among the set of storage node pairs.
[0007] According to some embodiments, a system for dynamically reclaiming storage space in a disaggregated storage system is provided. The system can include a memory and a processor. The processor can be configured to retrieve a plurality of data levels and a plurality of durability levels of a plurality of storage nodes. The plurality of data levels can correspond to an amount of data present in each storage segment of one or more storage segments associated with each of the plurality of storage nodes. The processor can be configured to determine a first range of variation based on a workload and one or more first parameters associated with the plurality of data levels. The workload can be associated with the plurality of storage nodes. The processor can be configured to determine a second range of variation based on the workload and one or more second parameters associated with the plurality of durability levels. The processor can be configured to identify one or more source nodes among the plurality of storage nodes based on the first range of variation. The processor can be configured to identify one or more target nodes among the plurality of storage nodes based on the second range of variation. The processor can be configured to identify a set of storage node pairs among the one or more source nodes and the one or more target nodes based on a quality of service penalty coefficient (QPC). The QPC can correspond to a number of network hops required to establish a data path between a storage node pair of the set of storage node pairs. Further, the processor can be configured to perform a reclamation of at least one storage segment of a storage node pair of the set of storage node pairs based on a reclamation efficiency coefficient (REC) of the storage node pair of the set of storage node pairs. The REC can correspond to a number of storage segments that can be reclaimed from the storage node pair of the set of storage node pairs.
[0008] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent to those skilled in the art upon examination of the drawings and the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0009] The drawings are incorporated into this disclosure and constitute a part of it, illustrate example embodiments, and, together with the detailed description below, serve to explain the principles disclosed. In the drawings, the rightmost bit of the reference numbers indicates in which drawing the reference label first appears. In all the figures, like references indicate similar features and components. Some embodiments of systems and / or methods according to embodiments of the subject matter will now be described, by way of example only, and with reference to the drawings in which:
[0010] Figure 1 An environment for dynamically reclaiming storage space in a disaggregated storage system is shown in accordance with some embodiments of the present disclosure;
[0011] Figure 2 A system for dynamically reclaiming storage space in a disaggregated storage system is shown in accordance with some embodiments of the present disclosure;
[0012] Figure 3 A flow diagram depicting an example method for reclaiming storage space in a dynamic, decomposed storage system is depicted in accordance with some embodiments of the present disclosure;
[0013] Figure 4 A flow diagram depicting an example method for reclaiming storage space in a dynamic, decomposed storage system is depicted in accordance with some embodiments of the present disclosure;
[0014] Figure 5 A block diagram depicting an example scenario for reclaiming storage space in a decomposed storage system is shown in accordance with some embodiments of the present disclosure; and
[0015] Figure 6 A block diagram of an example computing system for performing embodiments consistent with the present disclosure is shown in accordance with some embodiments of the present disclosure.
[0016] Those skilled in the art will appreciate that any flowchart, flow diagram, state transition diagram, pseudocode, and the like represent various processes which can be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown. DETAILED DESCRIPTION
[0017] In this document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation of the subject matter's embodiment or implementations described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
[0018] While the disclosure is amenable to various modifications and alternative forms, specifics thereof have been shown by way of example in the drawings and will be described in detail in the following description. It should be understood, however, that the intention is not to limit the disclosure to the particular embodiments described, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the scope of the disclosure.
[0019] The terms "comprise", "comprising", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but can include other components or steps not expressly listed or inherent to such setup, device or method. In other words, without further restriction, one or more elements of a device, system or apparatus preceded by "comprising" encompasses the device, system or apparatus comprising the specific elements, as well as other elements of the device, system or apparatus.
[0020] The terms "comprise(s)," "comprising," "include(s)," "including," "has(s)," "having" or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, product, or apparatus that comprises a list of elements is not necessarily limited only those elements but can include other elements not expressly listed or inherent to such process, other steps, or other ingredients, not expressly listed or inherent to such product or apparatus. In other words, nothing in this specification should be interpreted as a limitation on the scope of such processes or disclosing only those particular combinations of elements which were specifically asserted to be essential.
[0021] The terms "a" or "an", as used herein, mean "one or more" when applied to any element. The terms "including" and / or "containing," when utilized, mean "including without limitation" and / or "containing without limitation," respectively.
[0022] The terms "comprise(s)," "comprising," "include(s)," "including," "has(s)," "having" or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, product, or apparatus that comprises a list of elements is not necessarily limited only those elements but can include other elements not expressly listed or inherent to such process, other steps, or other ingredients, not expressly listed or inherent to such product or apparatus. In other words, nothing in this specification should be interpreted as a limitation on the scope of such processes or disclosing only those particular combinations of elements which were specifically asserted to be essential.
[0023] As used herein, the terms "communication" and "in communication" can refer to reception, receipt, transmission, transfer, provision, and / or the like of information (e.g., data, signals, messages, instructions, commands, and / or the like). That one unit (e.g., a device, system, component of a device or system, combinations thereof, and / or the like) is in communication with another unit means that the one unit is capable of directly or indirectly receiving a transmission from, and / or enabling a transmission to, the other unit. This can include a direct or indirect connection (e.g., direct communication connection, indirect communication connection, and / or the like) that is wired and / or wireless in nature. Additionally, two units can be in communication with each other even though the information transmitted can be modified, processed, relayed, and / or routed between the first and second unit. For example, a first unit can be in communication with a second unit even though the first unit passively receives information and does not actively transmit information to the second unit. As another example, a first unit can be in communication with a second unit if at least one intermediary unit (e.g., a third unit located between the first and second units) processes information received from the first unit and transmits process information to the second unit. In some non-limiting embodiments, a message can refer to a network packet (e.g., data packet, and / or the like) that includes data. It will be recognized that numerous other arrangements are possible.
[0024] As used herein, the term "processor" can refer to any suitable data computation device or any suitable plurality of data computation devices. The processor can include one or more microprocessors, which can be physically located proximate to, or remote from, other microprocessors. The processor can include a CPU with at least one high-speed data processor sufficient to execute program components for executing user and / or system generated requests. The CPU can be a microprocessor such as AMD's Athlon, Duron and Opteron processors; Intel processor such as Pentium, Celeron, and Xeon; IBM and Motorola's PowerPC; Sun's SPARC; and / or the like. TMDuron and / or Opteron TM IBM and / or Motorola's
[0025] IBM and Sony's Cell processor; Intel's and / or and / or similar processors.
[0026] As used herein, the term "memory" can be any suitable device or any suitable number of devices capable of storing electronic data. A suitable memory can include a non-transitory computer-readable medium that stores instructions capable of being executed by a processor to implement a desired method. Examples of memory can include one or more memory chips, disk drives, etc. Such memory can operate using any suitable electrical, optical, and / or magnetic mode of operation.
[0027] As used herein, the term "decomposed storage" or "decomposed storage system" can refer to a form of scalable storage that can combine the performance advantages of directly attached storage and the flexibility of a storage area network (SAN). A decomposed storage system has multiple storage devices (e.g., a storage cluster) that can be configured to function as a logical storage pool that can be reconfigured as needed without modifying the physical connections between them.
[0028] In embodiments, as used herein, the term "cluster" can refer to any suitable group of storage servers that can work like a single system capable of processing in parallel. Each storage server of the group of storage servers can be depicted as a storage node. Each storage node can provide logical space for applications that can access the logical space through semantics such as files, objects, blocks, databases (DBs), etc. Each storage node can include a physical storage in the form of a flash tier for storing data. Each storage node can also include one or more storage segments.
[0029] As used herein, the term "workload" refers to the amount and type of work that a storage system is expected to handle. It can represent the demands placed on a storage system by various operations (e.g., reading, writing, updating, and querying data present in the storage system). The workload can vary based on the characteristics of the applications using the storage system. For example, an e-commerce website can have a workload that includes frequent read operations to display product information, while a financial system can have a workload that includes complex queries and intensive write operations.
[0030] As used herein, the term "read-intensive workload" may refer to a workload that is associated with frequent data retrieval but infrequent updates.
[0031] As used herein, the term "write-intensive workload" may refer to a workload that is associated with frequent data updates but infrequent data retrieval.
[0032] In an embodiment, the term "valid data" may refer to data present in a segment of a storage node of a storage system that may still be relevant / required / necessary in the current iteration.
[0033] In an embodiment, the term "invalid data" may refer to data present in a segment of a storage node of a storage system, which may be irrelevant / non-essential in a current iteration.
[0034] In the following detailed description of the embodiments of the present disclosure, reference is made to the accompanying drawings, which form a part of the present disclosure and illustrate specific embodiments in which the present disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the present disclosure, and it should be understood that other embodiments may be utilized and changes may be made without departing from the scope of the present disclosure. Therefore, the following description should not be considered restrictive.
[0035] Figure 1 An environment 100 for dynamically reclaiming storage space in a disaggregated storage system is shown. The environment 100 may include a system 101 and a disaggregated storage system 104. The system 101 may include a processor 102 and a memory 103. The disaggregated storage system 104 may include one or more storage clusters 1051, 1052, ..., 105 n .
[0036] In an embodiment, the disaggregated storage system 104 may have one or more heterogeneous flash tiers in terms of endurance level per storage node.
[0037] The system 101 may be configured to dynamically reclaim storage space in the disaggregated storage system 104. The processor 102 may be configured to dynamically reclaim storage space in the disaggregated storage system 104 for one or more storage clusters 1051, 1052, ..., 105 n Multiple storage nodes retrieve multiple data levels and multiple durability levels 106 ( Figure 2 ) and stores it in the memory 103. Multiple durability levels and one or more storage clusters 1051, 1052, ..., 105 n The current durability level of multiple storage nodes is related to Figure 5If storage node 504c has storage segment 5012 with a durability level of 20% and storage segment 5016 with a durability level of 60%, then the durability level of storage node 504c can be an aggregated value (e.g., average) of the durability levels of the individual storage segments, i.e., the durability level of storage node 504c is 40%. It is noted that the durability level is indicated for example purposes only as an aggregated value of the durability levels of the individual storage segments of a storage node and can be determined by other statistical techniques. Each data level of the plurality of data levels corresponds to an amount of data present in each of the one or more storage segments associated with each of the plurality of storage nodes of one or more storage clusters 1051, 1052, …, 105 n The processor 102 can be configured to determine a first delta range based on the workload and one or more first parameters (e.g., from the plurality of data levels). In an embodiment, the workload can be an overall workload associated with the plurality of storage nodes of the cluster at a given point in time. In an embodiment, the plurality of data levels and the plurality of durability levels can be dynamically updated.
[0038] In an embodiment, the one or more first parameters correspond to (e.g., include at least one of) a maximum invalid data level, a minimum invalid data level, an average invalid data level, a maximum valid data level, a minimum valid data level, and an average valid data level associated with the plurality of storage nodes.
[0039] In an embodiment, the processor 102 determines the first delta range based on the one or more first parameters (e.g., but not limited to, the maximum invalid data level, the minimum invalid data level, and the average invalid data level). When the plurality of data levels correspond to (e.g., indicate) an amount of invalid data present in the one or more storage segments, the processor 102 determines the first delta range as a difference between the maximum invalid data level and the average invalid data level. In an embodiment, when the workload is write intensive, the first delta range can be reduced by half.
[0040] In another embodiment, the processor 102 determines the first delta range based on the one or more first parameters (e.g., but not limited to, the maximum valid data level, the minimum valid data level, and the average valid data level). When the plurality of data levels correspond to (e.g., indicate) an amount of valid data present in the one or more storage segments, the processor 102 determines the first delta range as a difference between the average valid data level and the minimum valid data level. In an embodiment, when the workload is write intensive, the first delta range can be reduced by half.
[0041] In an embodiment, the processor 102 can be configured to determine a second range of variation based on the workload and one or more second parameters (e.g., from the plurality of durability levels) associated with the plurality of durability levels. In an embodiment, the one or more second parameters correspond to (e.g., include at least one of) a maximum durability level, a minimum durability level, and an average durability level associated with the plurality of storage nodes. The processor 102 can determine the second range of variation as a difference between the average durability level and the minimum durability level associated with the one or more storage nodes. In an embodiment, the second range of variation can be reduced by half when the workload is write intensive.
[0042] The processor 102 can be configured to identify one or more source nodes among the plurality of storage nodes based on the first range of variation. The processor 102 can be configured to identify one or more target nodes among the plurality of storage nodes based on the second range of variation.
[0043] In an embodiment, the processor 102 can be configured to identify a set of storage node pairs among the one or more source nodes and the one or more target nodes based on a quality of service penalty coefficient (QPC). The QPC corresponds to a number of network hops required to establish a data path between a storage node pair of the set of storage node pairs (i.e., between two storage nodes in a pair). In an embodiment, a maximum possible QPC can be predefined (which can be defined as “Z”). In another embodiment, the maximum possible QPC can be user-defined and can be received from the user via the I / O interface 201 (e.g., via the user interface 202). The processor 102 can be configured to set the maximum possible QPC. For example, the processor 102 can be configured to identify the set of storage node pairs for all possible (i.e., from zero to “Z”) QPCs. For example, if the QPC is 1, then a storage node pair in which a data path between a source node having a storage segment that can be reclaimed and a target node having a storage segment containing free storage space can be established with one network hop can be determined. A storage node having a source storage segment can be defined as a source node. A storage node having a target storage segment can be defined as a target node. For example, in Figure 2 FIG. 5, if the source node 504a has a source storage segment 5011 and the target node 504b has a target storage segment 5015, then a data path can be established between these two segments for transferring valid data, resulting in a QPC of 1. Figure 5
[0044] Generally, the reclaiming process is an iterative method and is performed until a predetermined number of segments are reclaimed.
[0045] The processor 102 can be configured to perform reclaiming of at least one storage segment among the set of storage node pairs.
[0046] Figure 2 Detailed block diagram of system 101 is shown. System 101 may include processor 102, input / output (I / O) interface 201, memory 103, and module 207. Memory 103 may include data 202. Data 202 may include, for example, but not limited to, the following data: data level and durability level 106, invalid data level table 203, valid data level table 204, durability level table 205, and other data 206. Module 207 may also include, for example, but not limited to, the following modules: retrieval module 208, variation range determination module 209, identification module 210, QPC determination module 211, recovery efficiency coefficient (REC) determination module 212, recovery module 213, and other modules 214.
[0047] In an embodiment, other data 206 may include various temporary data and files generated by module 207 .
[0048] As used herein, the term "module" may refer to an application-specific integrated circuit (ASIC), an electronic circuit, a hardware processor (shared, dedicated, or group) and memory, a combinational logic circuit, and / or other suitable components that execute one or more software or firmware programs, and provide the functionality. In an embodiment, each module 207 may be configured as an independent hardware computing unit. In an embodiment, other modules 214 may be used to perform various miscellaneous functions of the system 101. It will be understood that the module 207 may be represented as a single module, or as a combination of different modules.
[0049] In an embodiment, the retrieval module 208 may be configured to retrieve information from one or more storage clusters 1051, 1052, ..., 105 of the disaggregated storage system 104. n The plurality of data levels and the plurality of durability levels 106 may be retrieved from the plurality of storage nodes and may be stored in the memory 103 as part of the data 202 as the data levels and the durability levels 106. The retrieval module 208 may be configured to generate an invalid data level table 203 from the data levels 106. The invalid data level may be determined based on the invalid data levels of the storage segments in the storage nodes from the data levels 106. The invalid data level may correspond to (e.g., indicate) the amount of invalid data present in the storage segments of the storage nodes. For example, in Figure 5 In the example, storage node 504c has storage segment 5012. Storage segment 5012 may have 20% useful data / valid data and 80% useless data / invalid data. The invalid data level of storage node 504c may be 80%. When performing recycling, the invalid data level can be considered for the source node. An example invalid data level table is depicted in Table 1 below.
[0050]
[0051] Table 1
[0052] For example, in an embodiment, in Table 1 above, storage node 1 may have three segments with invalid data levels of 80-99% and two storage segments with invalid data levels of 0-19%.
[0053] The retrieval module 208 may be configured to generate a valid data level table 204 from the data level 106. The valid data level may be determined based on the data levels of the storage segments in the storage nodes from the data level 106. The valid data level may correspond to (e.g., indicate) the amount of valid data present in the storage segments of the storage nodes. For example, Figure 5 In the example, storage node 504c has storage segment 5012. Storage segment 5012 may have 20% useful data / valid data and 80% useless data / invalid data, so the valid data level of storage node 504c may be 20%. When performing recycling, the valid data level can be considered for the source node. An example valid data level table is depicted in Table 2 below.
[0054]
[0055] Table 2
[0056] For example, in an embodiment, in Table 2 above, storage node 1 may have three segments with valid data levels of 0-19% and two storage segments with valid data levels of 80-99%.
[0057] The retrieval module 208 may be configured to generate the durability level table 205 from the durability levels 106. The durability level may be determined based on the durability levels of the storage segments in the storage nodes from the durability levels 106. For example, in Figure 5 , if storage node 504c has storage segment 5012 with a durability level of 20% and storage segment 5016 with a durability level of 60%, the durability level of storage node 504c can be the aggregate value of the individual storage segments, that is, the durability level of storage node 504c is 40%. It should be noted that the durability level is indicated as the aggregate value of the durability levels of the individual storage segments of the storage node. However, the durability level can be determined in other ways. When performing recycling, the durability level can be taken into account for the target node. An example durability level table is depicted in Table 3 below.
[0058] Durability level Storage node 80-99% 60-79% 40-59% Storage node 3, storage node 1 20-39% Storage node 2 0-19% Storage node 3
[0059] Table 3
[0060] For example, in an embodiment, in Table 3 above, the durability level of storage node 2 can be 20-39%.
[0061] In an embodiment, the variation range determination module 209 is configured to determine the first variation range and the second variation range.
[0062] In an embodiment, the variation range determination module 209 determines the first variation range from the plurality of data levels based on one or more first parameters associated with the plurality of data levels, the plurality of data levels being associated with the plurality of storage nodes. In an embodiment, the one or more first parameters correspond to the maximum invalid data level, the minimum invalid data level, the average invalid data level, the maximum valid data level, the minimum valid data level, and the average valid data level associated with the plurality of storage nodes.
[0063] In an embodiment, the variation range determination module 209 determines the first variation range based on one or more first parameters, such as but not limited to, the maximum invalid data level, the minimum invalid data level, and the average invalid data level. When the plurality of data levels correspond to the amount of invalid data present in the one or more storage segments, the variation range determination module 209 determines the first variation range as the difference between the maximum invalid data level and the average invalid data level. In an embodiment, when the workload is write intensive, the first variation range can be reduced by half.
[0064] In another embodiment, the variation range determination module 209 determines the first variation range based on one or more first parameters, such as but not limited to, the maximum valid data level, the minimum valid data level, and the average valid data level. When the plurality of data levels correspond to the amount of valid data present in the one or more storage segments, the variation range determination module 209 determines the first variation range as the difference between the average valid data level and the minimum valid data level. In an embodiment, when the workload is write intensive, the first variation range can be reduced by half.
[0065] In an embodiment, the variation range determination module 209 determines the second variation range from the plurality of durability levels based on the workload and one or more second parameters associated with the plurality of durability levels. In an embodiment, the one or more second parameters correspond to the maximum durability level, the minimum durability level, and the average durability level associated with the plurality of storage nodes. The variation range determination module 209 determines the second variation range as the difference between the average durability level and the minimum durability level associated with the one or more storage nodes. In an embodiment, when the workload is write intensive, the second variation range can be reduced by half.
[0066] In an embodiment, the identification module 210 may be configured to identify one or more source nodes among the plurality of storage nodes based on a first variation range. For example, in an embodiment, when the first variation range is identified as the difference between the maximum invalid data level and the average invalid data level, all storage nodes that may fall within the range in Table 1 may be considered as one or more source nodes. For example, in another embodiment, when the first variation range is identified as the difference between the average valid data level and the minimum valid data level, all storage nodes that may fall within the range in Table 2 may be considered as one or more source nodes.
[0067] In an embodiment, the identification module 210 may be configured to identify one or more target nodes among the plurality of storage nodes based on the second variation range. For example, in an embodiment, when the second variation range is identified as the difference between the average durability level and the minimum durability level, in Table 3, all storage nodes that may fall within the range may be considered as the one or more target nodes.
[0068] In an embodiment, the QPC determination module 211 may be configured to identify a set of storage node pairs among one or more source nodes and one or more target nodes based on a QPC. The QPC corresponds to the number of network hops required to establish a data path between the source node and the target node of the storage node pairs in the set of storage node pairs. In an embodiment, a maximum possible QPC may be predefined (which may be defined as "Z"). In another embodiment, the maximum possible QPC may be user-defined and may be received from the user via the I / O interface 201. The QPC determination module 211 may be configured to set the maximum possible QPC. For example, the QPC determination module 211 may be configured to identify a set of storage node pairs for all possible (i.e., from zero to "Z") QPCs. In an embodiment, the reclaim module 213 may be configured to perform reclaim of at least one storage segment in the set of storage node pairs.
[0069] Generally speaking, the reclamation process is an iterative method and is performed until a predetermined number of segments are reclaimed.
[0070] In an embodiment, the QPC determination module 211 may be configured to initialize the QPC to zero. The QPC determination module 211 may be configured to set the maximum possible QPC. The QPC determination module 211 may be configured to determine one or more storage node pairs from the set of storage node pairs based on the QPC. That is, if the current QPC value is 1, the QPC determination module 211 may identify a storage node pair from the set of storage node pairs whose QPC value is 1.
[0071] In an example scenario, if the QPC is 1, a storage node pair can be determined in which a data path between a source node and a target node of the storage node pair can be established with one network hop, the source node having a storage segment that can be reclaimed, and the target node having a storage segment containing free storage space. For example, in Figure 5 if the source node 504a has the source storage segment 5011 and the target node 504b has the target storage segment 5015, a data path can be established between these two segments for transferring valid data, resulting in a QPC of 1.
[0072] In another example scenario, if the QPC is 0, the source and target nodes can be a storage node pair of the same storage node, for example. More specifically, there can exist two storage segments within the same storage node with a QPC of 0 (i.e., a network hop count of 0). Thus, the source storage segment and the target storage segment must be identified within the same source node, i.e., the storage node corresponding to the source storage segment and the storage node corresponding to the target storage segment can be the same. For example, in Figure 5 if the storage node 504c has the source storage segment 5012 and the target storage segment 5016, a data path can be established between these two segments (i.e., 5012 and 5016) for transferring valid data, resulting in a QPC of zero. In this scenario, the source node and the target node of the storage node pair correspond to the same storage node, i.e., the source node is 504c and the target node is 504c, and the storage node pair is 504c-504c.
[0073] The example scenarios defined above should not be considered limiting, and as in other embodiments, a zero QPC can be defined as a network hop between two storage nodes within a storage cluster.
[0074] In an embodiment, the REC determination module 212 can be configured to determine a Reclaim Efficiency Coefficient (REC) for each of the one or more storage node pairs determined by the QPC determination module 211. The REC corresponds to (e.g., indicates) the number of storage segments that can be reclaimed from the source node of each storage node pair. For example, if in Table 1, storage node 1 having three storage segments with 80-99% invalid data and two storage segments with 0-19% invalid data is identified as the source node, and in Table 3, storage node 2 with a durability level between 20-39% is identified as the target node (having an available target storage segment), the REC for the storage node pair (i.e., storage node 1 - storage node 2) can be 5 since five segments in the source node (storage node 1) can be reclaimed.
[0075] In another embodiment, the REC determination module 212 can be configured to arrange the one or more storage node pairs in a descending order based on the REC associated with each of the one or more storage node pairs.
[0076] In an embodiment, the reclaiming module 213 can be configured to perform the reclaiming of at least one storage segment of a source node of the one or more storage node pairs based on the REC.
[0077] In an embodiment, the reclaiming module 213 can be configured to perform the reclaiming of at least one storage segment of a source node of the one or more storage node pairs based on the REC.
[0078] In another embodiment, the reclaiming module 213 can be configured to perform the reclaiming of at least one storage segment of a source node of the one or more storage node pairs based on a descending order arrangement of the REC associated with each of the one or more storage node pairs.
[0079] In an embodiment, the reclaiming module 213 can be configured to update the data level and the durability level associated with the one or more storage node pairs after the reclaiming. In an embodiment, the reclaiming module 213 can be configured to determine the number of storage segments (e.g., “Y” segments) based on the reclaiming of at least one storage segment of the one or more storage node pairs. In an embodiment, if a predetermined number (e.g., “X”) of segments need to be reclaimed, and “Y” is less than “X”, the reclaiming module 213 can be configured to increase the QPC and inform the processor 102 or the identifying module 210 that the method flow can need to be repeated. Generally, the reclaiming process is an iterative method and is performed until a predetermined number of storage segments are reclaimed.
[0080] Figure 3 A flowchart of an example method 300 for dynamically reclaiming storage space in a disaggregated storage system (e.g., the disaggregated storage system 104) is depicted.
[0081] At step 301, the processor 102 of the system 101 can be configured to retrieve a plurality of data levels and a plurality of durability levels of a plurality of storage nodes. In an embodiment, the plurality of data levels correspond to an amount of data present in each storage segment of one or more storage segments associated with each of the plurality of storage nodes. In an embodiment, the plurality of data levels correspond to one of: an amount of invalid data present in the one or more storage segments and an amount of valid data present in the one or more storage segments. In an embodiment, the plurality of data levels and the plurality of durability levels are dynamically updated.
[0082] At step 303, the processor 102 of the system 101 can be configured to determine a first range of variation from the plurality of data levels based on the workload and one or more first parameters associated with the plurality of data levels. In an embodiment, the workload is associated with the plurality of storage nodes. In an embodiment, the workload is one of a read intensive workload and a write intensive workload. In an embodiment, the one or more first parameters correspond to a maximum invalid data level, a minimum invalid data level, an average invalid data level, a maximum valid data level, a minimum valid data level, and an average valid data level associated with the plurality of storage nodes.
[0083] In an embodiment, the processor 102 of the system 101 can be configured to determine the first range of variation based on the one or more first parameters, such as but not limited to, the maximum invalid data level, the minimum invalid data level, and the average invalid data level. When the plurality of data levels correspond to an amount of invalid data present in the one or more storage segments, the processor 102 determines the first range of variation as a difference between the maximum invalid data level and the average invalid data level. In an embodiment, when the workload is write intensive, the first range of variation can be reduced by half.
[0084] In another embodiment, the processor 102 of the system 101 can be configured to determine the first range of variation based on the one or more first parameters, such as but not limited to, the maximum valid data level, the minimum valid data level, and the average valid data level. When the plurality of data levels correspond to an amount of valid data present in the one or more storage segments, the processor 102 determines the first range of variation as a difference between the average valid data level and the minimum valid data level. In an embodiment, when the workload is write intensive, the first range of variation can be reduced by half.
[0085] At step 305, the processor 102 of the system 101 can be configured to determine a second range of variation from the plurality of durability levels based on the workload and one or more second parameters associated with the plurality of durability levels. The one or more second parameters correspond to a maximum durability level, a minimum durability level, and an average durability level associated with the plurality of storage nodes. The second range of variation corresponds to a difference between the average durability level and the minimum durability level.
[0086] In an embodiment, the processor 102 of the system 101 can be configured to determine a second range of variation from the plurality of durability levels based on the workload and one or more second parameters associated with the plurality of durability levels. In an embodiment, the one or more second parameters correspond to a maximum durability level, a minimum durability level, and an average durability level associated with the plurality of storage nodes. The processor 102 determines the second range of variation as a difference between the average durability level and the minimum durability level associated with the one or more storage nodes. In an embodiment, the second range of variation can be reduced by half when the workload is write intensive.
[0087] At step 307, the processor 102 of the system 101 can be configured to identify one or more source nodes among the plurality of storage nodes based on the first range of variation. For example, in an embodiment, when the first range of variation is identified as a difference between the maximum invalid data level and the average invalid data level, in Table 1, all the storage nodes that can fall within this range can be considered as the one or more source nodes. In another embodiment, when the first range of variation is identified as a difference between the average valid data level and the minimum valid data level, in Table 2, all the storage nodes that can fall within this range can be considered as the one or more source nodes.
[0088] At step 309, the processor 102 of the system 101 can be configured to identify one or more target nodes among the plurality of storage nodes based on the second range of variation. For example, in an embodiment, when the second range of variation is identified as a difference between the average durability level and the minimum durability level, in Table 3, all the storage nodes that can fall within this range can be considered as the one or more target nodes.
[0089] At step 311, the processor 102 of the system 101 can be configured to identify a set of storage node pairs among the one or more source nodes and the one or more target nodes based on the QPC. In an embodiment, a maximum possible QPC can be predefined (which can be defined as “Z”). In another embodiment, the maximum possible QPC can be user-defined and can be received from the user via the I / O interface 201. In an embodiment, the processor 102 can be configured to initialize the QPC to zero. The processor 102 can be configured to set the maximum possible QPC. The processor can then be configured to determine one or more storage node pairs among the set of storage node pairs based on the QPC. In an example scenario, if the QPC is 1, then a storage node pair in which a data path between a source node (having a segment that can be reclaimed) and a target node (having a segment containing free storage space) can be established with one network hop can be determined. For example, in Figure 5In an example scenario, if the source node 504a has the source storage segment 5011 and the target node 504b has the target storage segment 5015, a data path can be established between these two segments for transferring valid data, resulting in a QPC of 1.
[0090] In another example scenario, if the QPC is 0, it can be determined that, for example, the source node and the target node can be a pair of storage nodes of the same storage node. More specifically, there can exist two storage segments within the same storage node with a QPC of 0 (i.e., a network hop count of 0). Therefore, the source storage segment and the target storage segment must be identified within the same source node, i.e., the storage node corresponding to the source storage segment and the storage node corresponding to the target storage segment can be the same. For example, in Figure 5 In an example scenario, if the source node 504a has the source storage segment 5011 and the target node 504b has the target storage segment 5015, a data path can be established between these two segments for transferring valid data, resulting in a QPC of 1.
[0091] In another embodiment, the processor 102 can be configured to initialize the QPC to zero. The processor 102 can be configured to set a maximum possible QPC. The processor 102 can be configured to determine one or more pairs of storage nodes from among the set of pairs of storage nodes identified at step 311 based on the QPC, i.e., if the current QPC value is 1, the QPC determination module 211 can identify the pair of storage nodes from among the set of pairs of storage nodes with a QPC value of 1.
[0092] In an embodiment, the processor 102 can be configured to determine a recycling efficiency coefficient (REC) for each of the one or more pairs of storage nodes determined by the processor 102.
[0093] In another embodiment, the processor 102 can be configured to arrange the one or more pairs of storage nodes based on a descending order of the REC associated with each of the one or more pairs of storage nodes.
[0094] At step 313, the processor 102 of the system 101 can be configured to perform recycling of at least one storage segment from among the set of pairs of storage nodes.
[0095] In an embodiment, the processor 102 can be configured to perform recycling of at least one storage segment of the source node of the one or more pairs of storage nodes based at least on the REC.
[0096] In another embodiment, the processor 102 can be configured to perform the reclaiming of at least one storage segment of a source node of one or more storage node pairs based on a descending arrangement of RECs associated with each of the one or more storage node pairs.
[0097] In an embodiment, the processor 102 can be configured to update (e.g., dynamically update) the data class and the durability class associated with at least one of the set of storage node pairs. In an embodiment, the processor 102 can be configured to determine the number of storage segments (e.g., “Y” segments) based on the reclaiming of at least one storage segment of one or more storage node pairs. In an embodiment, if a predetermined number (e.g., “X”) of segments needs to be reclaimed, and “Y” is less than “X”, the processor 102 can be configured to increase the QPC, and repeat the steps starting from step 311.
[0098] Figure 4 A flowchart 400 of an example method for dynamically reclaiming storage space in a disaggregated storage system, which can be performed by the system 101, is depicted.
[0099] At step 401, the system 101 can be configured to set the number of segments that need to be reclaimed / compressed. In an embodiment, the number of segments that need to be reclaimed can be predefined. In another embodiment, the number of segments that need to be reclaimed can be defined by a user, and can be received from the user via the I / O interface 201. For example, the number of segments that need to be reclaimed can be “X”.
[0100] At step 403, the system 101 can be configured to retrieve a plurality of data classes and a plurality of durability classes 106 for a plurality of storage nodes of one or more storage clusters 1051, 1052, …, 105 n of the disaggregated storage system 104, and can store it in the memory 103 as the data classes and durability classes 106.
[0101] At step 405, the system 101 can be configured to create the durability class table 205 from the data classes and durability classes 106, and create one of the invalid data class table 203 and the valid data class table 203.
[0102] At step 407, the system 101 can be configured to determine the first range of variation and the second range of variation.
[0103] The system 101 can be configured to determine a first range of change from the plurality of data levels based on the workload and one or more first parameters associated with the plurality of data levels. In an embodiment, the workload is associated with the plurality of storage nodes. In an embodiment, the workload is one of a read intensive workload and a write intensive workload. In an embodiment, the one or more first parameters correspond to a maximum invalid data level, a minimum invalid data level, an average invalid data level, a maximum valid data level, a minimum valid data level, and an average valid data level associated with the plurality of storage nodes.
[0104] In an embodiment, the system 101 can be configured to determine the first range of change based on one or more first parameters such as, but not limited to, the maximum invalid data level, the minimum invalid data level, and the average invalid data level. When the plurality of data levels correspond to an amount of invalid data present in the one or more storage segments, the system 101 determines the first range of change as a difference between the maximum invalid data level and the average invalid data level. In an embodiment, when the workload is write intensive, the first range of change can be reduced by half.
[0105] In another embodiment, the system 101 can be configured to determine the first range of change based on one or more first parameters such as, but not limited to, the maximum valid data level, the minimum valid data level, and the average valid data level. When the plurality of data levels correspond to an amount of valid data present in the one or more storage segments, the system 101 determines the first range of change as a difference between the average valid data level and the minimum valid data level. In an embodiment, when the workload is write intensive, the first range of change can be reduced by half.
[0106] In an embodiment, the system 101 can be configured to determine a second range of change from the plurality of durability levels based on the workload and one or more second parameters associated with the plurality of durability levels. The one or more second parameters correspond to a maximum durability level, a minimum durability level, and an average durability level associated with the plurality of storage nodes. The second range of change corresponds to a difference between the average durability level and the minimum durability level.
[0107] In an embodiment, the system 101 can be configured to determine a second range of change from the plurality of durability levels based on the workload and one or more second parameters associated with the plurality of durability levels. In an embodiment, the one or more second parameters correspond to a maximum durability level, a minimum durability level, and an average durability level associated with the plurality of storage nodes. The system 101 determines the second range of change as a difference between the average durability level and the minimum durability level associated with the one or more storage nodes. In an embodiment, when the workload is write intensive, the second range of change can be reduced by half.
[0108] At step 409, the system 101 can be configured to identify one or more source nodes and one or more target nodes. The system 101 can be configured to identify one or more source nodes among the plurality of storage nodes based on the first range of variation. For example, in an embodiment, when the first range of variation is identified as the difference between the maximum level of invalid data and the average level of invalid data, in Table 1, all storage nodes that fall within this range can be considered as one or more source nodes. In another embodiment, when the first range of variation is identified as the difference between the average level of valid data and the minimum level of valid data, in Table 2, all storage nodes that fall within this range can be considered as one or more source nodes.
[0109] In an embodiment, the system 101 can be configured to identify one or more target nodes among the plurality of storage nodes based on the second range of variation. For example, in an embodiment, when the second range of variation is identified as the difference between the average level of durability and the minimum level of durability, in Table 3, all storage nodes that fall within this range can be considered as one or more target nodes.
[0110] At step 411, the system 101 can be configured to initialize the QPC to / into zero. In an embodiment, a maximum possible QPC can be predefined (which can be defined as “Z”). In another embodiment, the maximum possible QPC can be user-defined and can be received from the user via the I / O interface 201. For example, Z can be defined as five.
[0111] At step 413, the system 101 can be configured to check whether the maximum possible QPC value has been reached. If the QPC is less than or equal to “Z”, the system 101 can perform steps 415-429. If the QPC is greater than “Z”, the system 101 can be configured to terminate the method flow.
[0112] At step 415, the system 101 can be configured to identify a set of storage node pairs among the one or more source nodes and the one or more target nodes based on the QPC. Then, the system 101 can be configured to select / determine one or more storage node pairs among the set of storage node pairs based on the QPC. In an embodiment, the system 101 can be configured to initialize the QPC to / into zero. In an embodiment, a maximum possible QPC can be predefined (which can be defined as “Z”). In another embodiment, the maximum possible QPC can be user-defined and can be received from the user via the I / O interface 201. The system 101 can be configured to set the maximum possible QPC. Then the system 101 can be configured to select / determine one or more storage node pairs among the set of storage node pairs based on the QPC. In an example scenario, if the QPC is 1, then a storage node pair can be determined in which a data path between a source node (having a segment that can be reclaimed) and a target node (having a segment containing free storage space) of the storage node pair can be established with one network hop. For example, in Figure 5 Fig. 5B, if the source node 504a has the source storage segment 5011 and the target node 504b has the target storage segment 5015, then a data path can be established between these two segments for transferring valid data, resulting in a QPC of 1.
[0113] In another example scenario, if the QPC is 0, then a storage node pair can be determined in which, for example, the source node and the target node can be the same storage node. More specifically, there can exist two storage segments within the same storage node with a QPC of 0 (i.e., a network hop count of 0). Therefore, the source storage segment and the target storage segment must be identified within the same source node, i.e., the storage node corresponding to the source storage segment and the storage node corresponding to the target storage segment can be the same. For example, in Figure 5 Fig. 5C, if the storage node 504c has the source storage segment 5012 and the target storage segment 5016, then a data path can be established between these two segments for transferring valid data, resulting in a QPC of zero. In this scenario, the source node and the target node of the storage node pair correspond to the same storage node, i.e., the source node is 504c and the target node is 504c, and the storage node pair is 504c-504c.
[0114] At step 417, the system 101 can be configured to determine the REC of each storage node pair of the set of storage node pairs identified at step 415. For example, if in Table 1, storage node 1 having three storage segments with invalid data of 80-99% and two storage segments with invalid data of 0-19% is identified as the source node, and in Table 3, storage node 2 with durability level between 20-39% is identified as the target node (with available target storage segments), the REC of the storage node pair (i.e., storage node 1 - storage node 2) can be 5 since five segments in the source node (storage node 1) can be reclaimed.
[0115] At step 419, the system 101 can be configured to identify at least one storage node pair among the set of storage node pairs based on the REC associated with each storage node pair of the set of storage node pairs.
[0116] In another embodiment, the system 101 can be configured to arrange the one or more storage node pairs based on a descending order of the REC associated with each of the one or more storage node pairs.
[0117] In an embodiment, the system 101 can be configured to identify at least one storage node pair among the set of storage node pairs based on the REC associated with each storage node pair of the one or more storage node pairs. In an example, the at least one storage node pair has the highest REC among the one or more storage node pairs.
[0118] In another embodiment, the system 101 can be configured to arrange the one or more storage node pairs based on a descending order of the REC associated with each of the one or more storage node pairs.
[0119] At step 421, the system 101 can be configured to identify at least one storage node pair having the second highest REC (among the set of storage node pairs).
[0120] At step 423, the system 101 can be configured to perform reclamation of at least one storage segment of the source node of the at least one storage node pair based at least on the REC.
[0121] In another embodiment, the system 101 can be configured to perform reclamation of at least one storage segment of the source node of the one or more storage node pairs arranged based on a descending order of the REC associated with each of the one or more storage node pairs.
[0122] At step 425, the system 101 can be configured to update the invalid data level table 203, the valid data level table 204, and the durability level table 205.
[0123] At step 427, the system 101 can be configured to determine the number of storage segments (e.g., the number of recycled storage segments “Y”) based on the recycling of at least one storage segment of one or more pairs of storage nodes.
[0124] At step 429, the system 101 can be configured to check if “Y” is less than “X”, i.e., if the number of recycled storage segments is less than the number of segments to be recycled. In an embodiment, if “Y” is greater than “X”, the system 101 can terminate the method flow.
[0125] In an embodiment, if “Y” is less than “X”, at step 431, the system 101 can be configured to subtract the value of “Y” from “X” and update the value of “X”.
[0126] In an embodiment, if “Y” is less than “X”, at step 433, the system 101 can be configured to increment the QPC and repeat the method 400 from step 413 to step 429.
[0127] Figure 5 An example scenario for recycling of storage space in a disaggregated storage system is depicted in accordance with some embodiments of the present disclosure. Figure 5 A cluster 500 is disclosed. In an embodiment, the cluster can include data store 1 502a, data store 2 502b, storage node 1 504a, storage node 2 504b, storage node 3 504c, and storage node 4 504d. Storage node 1 504a and storage node 2 504b can be connected via switch 1 503a. Storage node 3 504c and storage node 4 504d can be connected via switch 2 503b.
[0128] In an embodiment, each storage node can have one or more storage segments associated with it. A storage segment can be a source storage segment (having valid data that needs to be moved to another segment so that the segment can be recycled) or a target storage segment (having free storage space so that valid data can be moved into it). Storage node 504a has source storage segment 5011. Storage node 504b has target storage segment 5015. Storage node 504c has source storage segment 5012 and target storage segment 5016. Storage node 504d has source storage segment 5013 and target storage segment 5014.
[0129] Storage node 504a can have an aggregated durability level of 50%. Storage node 504b can have an aggregated durability level of 40%. Storage node 504c can have an aggregated durability level of 40%. Storage node 504d can have an aggregated durability level of 30%.
[0130] In an embodiment, if storage node 504c has storage segment 5012 with a durability level of 20% and segment 5016 with a durability level of 60%, the durability level of storage node 504c can be the aggregate value of the individual storage segments, i.e., the durability level of storage node 504c is 40%. In an embodiment, the durability level is indicated as the aggregate value of the durability levels of the individual storage segments of a storage node. However, the durability level can be determined in other ways.
[0131] In another embodiment, if source node 504a has source storage segment 5011 and target node 504b has target storage segment 5015, a data path can be established between these two segments for transferring valid data, resulting in a QPC of 1.
[0132] In another embodiment, if the QPC is zero, it can be determined that, for example, the source node and the target node can be a storage node pair of the same storage node. More specifically, there can be two storage segments within the same storage node with a QPC of 0 (i.e., a network hop count of 0). Thus, the source storage segment and the target storage segment must be identified within the same source node, i.e., the storage node corresponding to the source storage segment and the storage node corresponding to the target storage segment can be the same. For example, if storage node 504c has source storage segment 5012 and target storage segment 5016, a data path can be established between these two segments for transferring valid data, resulting in a QPC of zero. In this scenario, the source node and the target node of the storage node pair correspond to the same storage node, i.e., the source node is 504c and the target node is 504c, and the storage node pair is 504c-504c.
[0133] Figure 6 A block diagram of an example computing system 600 for implementing embodiments consistent with the present disclosure is shown. Computing system 600 can be, but is not limited to, system 101. Computing system 600 can include a central processing unit ("CPU" or "processor") 601. Processor 601 can include at least one data processor for performing processing. Processor 601 can include specialized processing units such as, for example, an integrated system (bus) controller, memory management control unit, floating point unit, graphics processing unit, digital signal processing unit, etc.
[0134] The processor 601 can communicate with one or more input / output (I / O) devices 607 and 608 via the I / O interface 606. The I / O interface 606 can employ communication protocols / methods that are well-known in the art, such as, but not limited to: audio, analog, digital, mono, RCA, stereo, IEEE- 1394, serial bus, universal serial bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, digital visual interface (DVI), high-definition multimedia interface (HDMI), RF antenna, S-Video, VGA, IEEE 802.n / b / g / n / x, Bluetooth, cellular (e.g., code division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, etc.), etc.
[0135] Using the I / O interface 606, the computing system 600 can communicate with one or more I / O devices 607 and 608. For example, the input device 607 can be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, fax machine, dongle, biometric reader, microphone, touchscreen, touchpad, trackball, stylus, scanner, storage device, transceiver, video device / source, etc. The output device 608 can be a printer, fax machine, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, plasma display panel (PDP), organic light-emitting diode display (OLED), etc.), audio speaker, etc.
[0136] In some embodiments, the processor 601 can communicate with external elements (e.g., external computer systems, servers, network elements). The network interface 610 can employ connection protocols including, but not limited to: direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc.
[0137] In some embodiments, the processor 601 can communicate with a memory 603 (e.g., RAM, ROM, etc.) via a storage interface 602. The storage interface 602 can connect to the memory 603 (including, but not limited to, memory drives, removable disc drives, etc.) employing connection protocols (e.g., serial advanced technology attachment (SATA), integrated drive electronics (IDE), IEEE-1394, universal serial bus (USB), Fibre Channel, small computer system interface (SCSI), etc.). The memory drives can further include drum memories, magnetic disc drives, magneto-optical drives, optical drives, redundant array of independent discs (RAID), solid-state memory devices, solid state drives, etc.
[0138] Memory 603 may store a collection of program or database components (including, but not limited to, an operating system 604, data 605, etc.). In some embodiments, computing system 600 may store user / application data 605 (e.g., data, variables, records, etc.), as described in this disclosure. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases (e.g., or ).
[0139] Operating system 604 may facilitate resource management and operation of computing system 600. Examples of operating systems include, but are not limited to, APPLE OS X, UNIX-like system distributions (e.g., Berkeley Software Distribution) TM (BSD), FREEBSD TM , NETBSD TM 、OPENBSD TM etc.), LINUX DISTRIBUTION TM (For example, RED HAT TM UBUNTU TM 、KUBUNTU TM etc.), IBM TM OS / 2, MICROSOFT TM WINDOWS TM (XP TM VISTA TM / 7 / 8, 10, etc.), iOS TM 、 Android TM 、 OS, etc.
[0140] In some embodiments, the computing system 600 may be configured to communicate with the disaggregated storage system 104. The disaggregated storage system 104 may include one or more storage clusters 1051, 1052, ..., 105 n The network interface 610 can communicate with the disaggregated storage system 104 using a connection protocol including, but not limited to, direct connection, Ethernet (e.g., twisted pair 10 / 100 / 1000Base T), Transmission Control Protocol / Internet Protocol (TCP / IP), Token Ring, IEEE 802.11a / b / g / n / x, etc.
[0141] Furthermore, embodiments consistent with the present disclosure can be implemented using one or more computer-readable storage media. Computer-readable storage media refers to any type of physical memory that can store information or data that is readable by a processor. Thus, computer-readable storage media can store instructions for execution by one or more processors, including instructions for causing a processor to perform steps or stages consistent with the embodiments described herein. The term "computer-readable medium" should be taken to include tangible articles but not carrier waves and transitory signals, i.e., is non-transitory. Examples include random access memory (RAM), read only memory (ROM), volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, magnetic disks, and any other known physical storage media.
[0142] The described operations can be implemented as a method, system, or article of manufacture using standard programming and / or engineering techniques to produce software, firmware, hardware, or any combination thereof. The described operations can be implemented as code that is saved in a "non-transitory computer-readable medium" that a processor can read and execute. The processor is at least one of a microprocessor and a processor. The non-transitory computer-readable medium can include media such as magnetic storage media (e.g., hard disk drives, floppy disks, magnetic strips in identification cards, etc.), optical storage (e.g., CD ROMs, DVDs, optical disks, etc.), volatile and non-volatile memory devices (e.g., EEPROMs, ROMs, PROMs, RAMs, DRAMs, SRAMs, Flash, firmware, programmable logic, etc.). In addition, the non-transitory computer-readable medium can include all computer-readable media excluding transitory media. The code implementing the described operations can further be implemented in hardware logic (e.g., an integrated circuit chip, Programmable Gate Array (PGA), Application Specific Integrated Circuit (ASIC), etc.).
[0143] An "article of manufacture" includes a non-transitory computer-readable medium and / or hardware logic in which code implementing the described operations can be embodied. The device that the code implementing the described operations embodiments is encoded on can include a computer-readable medium or hardware logic. Of course, those skilled in the art will recognize that many modifications can be made to this configuration, and that the article of manufacture can include suitable information bearing media known in the art.
[0144] One or more embodiments described above can have the advantageous effect that the reclaiming of storage space does not affect the QoS of the cluster resources while freeing the maximum amount of source storage segments / nodes for further use. This advantageous effect is achieved because the system and method disclosed in the present disclosure takes into account (e.g., considers) various factors, such as QPC and REC and durability level and data level, when performing the reclaiming of storage segments of storage nodes. Another advantage of the present disclosure embodiments is that it reduces the dispersion of effective data due to the nature of the disaggregated storage system.
[0145] The terms "an embodiment," "one embodiment," "embodiments," "an implementation," "one implementation," "some implementations," "one or more implementations," "some embodiments,” and "one or more embodiments” mean "one or more (but not all) embodiments of the invention."
[0146] The terms "comprise,” "comprising,” "include,” "including,” and "have” and variations thereof herein are meant to be open-ended.
[0147] Unless otherwise noted, a list of items does not imply that any or all of the items are mutually exclusive.
[0148] Unless otherwise noted, the terms "a,” "an,” and "the” mean "one or more.”
[0149] The description of an embodiment of the invention, in which multiple components communicate with each other, does not imply that all of these components must be present or necessary for the practice of the invention. On the contrary, a variety of optional components are described to illustrate the wide variety of potential embodiments of the present invention.
[0150] When a single device or article is described herein, it will be readily apparent that more than one device / article (whether or not they cooperate) can be used in place of a single device / article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it will be readily apparent that a single device / article can be used in place of the more than one device or article or a different number of devices / articles can be used instead of the shown number of devices or programs. The functionality and / or the features of a device can be alternatively embodied by one or more other devices which are not explicitly described as having such
[0151] Figure 3 and Figure 4The illustrated operations show certain events occurring in a certain order. In alternative embodiments, certain operations can be performed in a different order, modified or removed. Moreover, steps can be added to the above-described logic and still conform to the described embodiments. Further, operations described herein can occur sequentially or certain operations can be processed in parallel. Moreover, operations can be performed by a single processing unit or by distributed processing units.
[0152] Finally, the language used in the specification has been principally selected for readability and instructional purposes and can not have been selected to delineate or circumscribe the subject application. Accordingly, the scope of the subject application is intended to be defined solely by the claims submitted herewith, and it is specifically intended that the claims be interpreted as such. Accordingly, the disclosure of the embodiments of the application is intended primarily for purposes of illustrating the subject application and is not intended to limit the scope of the subject application as set forth in the claims submitted herewith.
[0153] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to limit the scope of the appended claims.
Claims
1. A method for dynamically reclaiming storage space in a decomposed storage system, the method comprising: retrieving a plurality of data levels and a plurality of durability levels of a plurality of storage nodes, wherein each data level of the plurality of data levels corresponds to an amount of data present in each storage segment of one or more storage segments associated with each of the plurality of storage nodes; determining a first range of variation based on a workload and one or more first parameters associated with the plurality of data levels, wherein the workload is associated with the plurality of storage nodes; determining a second range of variation based on the workload and one or more second parameters associated with the plurality of durability levels; identifying one or more source nodes among the plurality of storage nodes based on the first range of variation; identifying one or more target nodes among the plurality of storage nodes based on the second range of variation; identifying a set of storage node pairs among the one or more source nodes and the one or more target nodes based on a quality of service penalty coefficient (QPC), wherein the QPC corresponds to a number of network hops required to establish a data path between a storage node pair of the set of storage node pairs; and performing a reclamation of at least one storage segment among the set of storage node pairs.
2. The method of claim 1, wherein, performing the reclamation of the at least one storage segment comprises: initializing the QPC to zero; and iteratively performing the following steps until a predetermined number of storage segments have been reclaimed among the set of storage node pairs: determining one or more storage node pairs among the set of storage node pairs based on the QPC; determining a reclamation efficiency coefficient (REC) for each storage node pair of the one or more storage node pairs, wherein the REC corresponds to a number of storage segments that can be reclaimed from a source node of each storage node pair; performing the reclamation of the at least one storage segment from a source node of the one or more storage node pairs based on the REC; updating data levels and durability levels associated with the one or more storage node pairs; determining a number of storage segments based on the reclamation of the at least one storage segment of the one or more storage node pairs; and incrementing the QPC.
3. The method of claim 1, wherein, the plurality of data levels correspond to one of: an amount of invalid data present in the one or more storage segments; or an amount of valid data present in the one or more storage segments.
4. The method of claim 1, wherein, the plurality of data levels and the plurality of durability levels are dynamically updated.
5. The method of claim 1, wherein, the one or more first parameters include at least one of a maximum invalid data level, a minimum invalid data level, an average invalid data level, a maximum valid data level, a minimum valid data level, and an average valid data level associated with the plurality of storage nodes.
6. The method of claim 1, wherein, when the plurality of data levels correspond to an amount of invalid data present in the one or more storage segments, the first range of variation includes a difference between a maximum invalid data level and an average invalid data level; or The first range of variation includes a difference between an average data level and a minimum data level when the plurality of data levels corresponds to an amount of valid data present in the one or more storage segments.
7. The method of claim 1, wherein, The one or more second parameters include at least one of a maximum durability level, a minimum durability level, and an average durability level associated with the plurality of storage nodes.
8. The method of claim 1, wherein, The second range of variation includes a difference between an average durability level and a minimum durability level.
9. The method of claim 1, wherein, The workload includes a read-intensive workload or a write-intensive workload.
10. The method of claim 9, wherein, The first range of variation and the second range of variation are reduced by half when the workload is the write-intensive workload.
11. A system for dynamically reclaiming storage space in a disaggregated storage system, the system for dynamically reclaiming storage space comprising: a memory; and a processor configured to: retrieve a plurality of data levels and a plurality of durability levels of a plurality of storage nodes, wherein each data level of the plurality of data levels corresponds to an amount of data present in each storage segment of one or more storage segments associated with each of the plurality of storage nodes; determine a first range of variation based on a workload and one or more first parameters associated with the plurality of data levels, wherein the workload is associated with the plurality of storage nodes; determine a second range of variation based on the workload and one or more second parameters associated with the plurality of durability levels; identify one or more source nodes among the plurality of storage nodes based on the first range of variation; identify one or more target nodes among the plurality of storage nodes based on the second range of variation; identify a set of storage node pairs among the one or more source nodes and the one or more target nodes based on a quality of service penalty coefficient (QPC), wherein the QPC corresponds to a number of network hops required to establish a data path between a storage node pair of the set of storage node pairs; and perform reclamation of at least one storage segment among the set of storage node pairs.
12. The system of claim 11, wherein, The processor is configured to perform reclamation of the at least one storage segment by: initializing the QPC to zero; and iteratively performing the following steps until a predetermined number of storage segments have been reclaimed in the set of storage node pairs: determining one or more storage node pairs among the set of storage node pairs based on the QPC; determining a reclamation efficiency coefficient (REC) for each storage node pair of the one or more storage node pairs, wherein the REC corresponds to a number of storage segments that can be reclaimed from a source node of each storage node pair; performing reclamation of the at least one storage segment from a source node of the one or more storage node pairs based on the REC; updating data levels and durability levels associated with the one or more storage node pairs; determining a number of storage segments based on reclamation of the at least one storage segment of the one or more storage node pairs; and incrementing the QPC.
13. The system of claim 11, wherein, The plurality of data levels correspond to one of: an amount of invalid data present in the one or more storage segments; or an amount of valid data present in the one or more storage segments.
14. The system of claim 11, wherein, the processor is configured to dynamically update the plurality of data levels and the plurality of durability levels.
15. The system of claim 11, wherein, the one or more first parameters include at least one of a maximum invalid data level, a minimum invalid data level, an average invalid data level, a maximum valid data level, a minimum valid data level, and an average valid data level associated with the plurality of storage nodes.
16. The system of claim 11, wherein, when the plurality of data levels correspond to an amount of invalid data present in the one or more storage segments, the first range of variation includes a difference between a maximum invalid data level and an average invalid data level; or when the plurality of data levels correspond to an amount of valid data present in the one or more storage segments, the first range of variation includes a difference between an average valid data level and a minimum valid data level.
17. The system of claim 11, wherein, the one or more second parameters include at least one of a maximum durability level, a minimum durability level, and an average durability level associated with the plurality of storage nodes.
18. The system of claim 11, wherein, the second range of variation includes a difference between an average durability level and a minimum durability level.
19. The system of claim 11, wherein, the workload includes a read-intensive workload or a write-intensive workload, and wherein, when the workload is the write-intensive workload, the first range of variation and the second range of variation are reduced by half.
20. A system for dynamically reclaiming storage space in a disaggregated storage system, the system for dynamically reclaiming storage space comprising: a memory; and a processor configured to: retrieve a plurality of data levels and a plurality of durability levels of a plurality of storage nodes, wherein each data level of the plurality of data levels corresponds to an amount of data present in each storage segment of one or more storage segments associated with each of the plurality of storage nodes; determine a first range of variation based on a workload and one or more first parameters associated with the plurality of data levels, wherein the workload is associated with the plurality of storage nodes; determine a second range of variation based on the workload and one or more second parameters associated with the plurality of durability levels; identify one or more source nodes among the plurality of storage nodes based on the first range of variation; identify one or more target nodes among the plurality of storage nodes based on the second range of variation; identify a set of storage node pairs among the one or more source nodes and the one or more target nodes based on a quality of service penalty coefficient (QPC), wherein the QPC corresponds to a number of network hops required to establish a data path between a storage node pair of the set of storage node pairs; and performing a reclaim of at least one storage segment of a storage node pair of the set of storage node pairs based on a reclaim efficiency factor REC of the storage node pair, wherein the REC corresponds to a number of storage segments that can be reclaimed from the storage node pair of the set of storage node pairs.