Data layout selection between storage devices associated with nodes of a distributed file system cluster

By monitoring and analyzing the data access frequency and available space parameters of storage devices, the optimal storage location is selected, which solves the problem of data layout imbalance in the storage system and improves storage efficiency.

CN116414796BActive Publication Date: 2026-07-03DELL PROD LP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DELL PROD LP
Filing Date
2021-12-31
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

In existing storage systems, the data layout selection among storage devices is difficult to effectively balance data access frequency and available space, resulting in reduced storage efficiency.

Method used

By monitoring the data access frequency and available space parameters of storage devices, overall performance indicators are calculated, and the optimal storage location is selected based on these indicators to reduce imbalances between storage devices.

Benefits of technology

This achieves a more optimized data layout between storage devices, improving the efficiency and performance of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116414796B_ABST
    Figure CN116414796B_ABST
Patent Text Reader

Abstract

An apparatus includes a processing device configured to receive a request to store one or more data portions at a given node of a distributed file system cluster and monitor performance parameters of each storage device associated with the given node, the performance parameters including a first performance parameter characterizing a frequency of data access and at least a second performance parameter characterizing available space. The processing device is further configured to determine an overall performance indicator for each of the storage devices associated with the given node based at least in part on the monitored performance parameters, and select at least one of the storage devices associated with the given node to store the one or more data portions based at least in part on the overall performance indicators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This field relates generally to information processing, and more specifically to storage in information processing systems. Background Technology

[0002] Storage arrays and other types of storage systems are typically shared by multiple host devices over a network. Applications running on these host devices each consist of one or more processes that perform application functionality. These processes issue input / output (I / O) operation requests to the storage system. The storage system's storage controller serves these I / O operation requests. In some information processing systems, multiple storage systems can be used to form a storage cluster. Summary of the Invention

[0003] The illustrative embodiments of this disclosure provide techniques for selecting data layout among storage devices associated with nodes in a distributed file system cluster.

[0004] In one embodiment, an apparatus includes at least one processing means, the at least one processing means including a processor coupled to memory. The at least one processing means is configured to perform the following steps: receiving a request to store one or more portions of data at a given node of two or more nodes in a distributed file system cluster; and monitoring two or more performance parameters of each of two or more storage devices associated with the given node, the two or more performance parameters including a first performance parameter characterizing data access frequency and at least a second performance parameter characterizing available space. The at least one processing means is configured to perform the following steps: determining an overall performance metric for each of the two or more storage devices associated with the given node based at least in part on the monitored two or more performance parameters; and selecting at least one storage device from the two or more storage devices associated with the given node to store the one or more portions of data, based at least in part on the determined overall performance metric. The at least one processing means is configured to perform the following step: storing the one or more portions of data on the selected at least one storage device from the two or more storage devices associated with the given node.

[0005] These and other illustrative embodiments include, but are not limited to, methods, devices, networks, systems, and processor-readable storage media. Attached Figure Description

[0006] Figure 1 This is a block diagram of an information processing system configured in an illustrative implementation for selecting data layouts among storage devices associated with nodes in a distributed file system cluster.

[0007] Figure 2 This is a flowchart of an exemplary process for selecting a data layout among storage devices associated with nodes of a distributed file system cluster, as described in an illustrative implementation.

[0008] Figure 3 A distributed file system architecture in an illustrative implementation is shown.

[0009] Figure 4 A circular selection strategy for storage volumes used in a data node is illustrated in an illustrative implementation.

[0010] Figure 5 An illustrative implementation of a storage volume availability selection strategy for data nodes is shown.

[0011] Figure 6 The illustrative implementation shows the process flow for selecting storage volumes based on real-time input-output temperature and the ratio of used space.

[0012] Figure 7A Exemplary parameter values ​​for the storage volume of a data node in an illustrative implementation are shown.

[0013] Figure 7B An illustrative implementation scheme is shown where data blocks use a cyclic selection strategy across data blocks with... Figure 7A The layout of the storage volume of the data node with parameter values.

[0014] Figure 7C An illustrative implementation scheme is shown where data blocks use an available space selection strategy across areas with Figure 7A The layout of the storage volume of the data node with parameter values.

[0015] Figure 7D The illustrative implementation shows a data block selection strategy that considers real-time input-output temperature and the ratio of used space across data blocks. Figure 7A The layout of the storage volume of the data node with parameter values.

[0016] Figure 8 The diagram shows graphs of input-output temperature and used space percentage imbalance rate using different volume selection strategies in the illustrative implementation.

[0017] Figure 9 The illustration shows graphs of input-output temperature and percentage of used space imbalance after laying out each data block using a selection strategy that takes into account real-time input-output temperature and used space ratio in an illustrative implementation.

[0018] Figure 10 and Figure 11 An example of a processing platform that can be used to implement at least a portion of an information processing system is shown in an illustrative embodiment. Detailed Implementation

[0019] This document describes illustrative embodiments with reference to exemplary information processing systems and associated computers, servers, storage devices, and other processing apparatuses. However, it should be understood that the embodiments are not limited to use with the specific illustrative system and apparatus configurations shown. Therefore, the term "information processing system" as used herein is intended to be broadly interpreted to encompass, for example, processing systems including cloud computing and storage systems, as well as other types of processing systems including various combinations of physical and virtual processing resources. Thus, an information processing system may include, for example, at least one data center or other type of cloud-based system that includes one or more clouds hosting tenants accessing cloud resources.

[0020] Figure 1 An information processing system 100 is illustrated, configured according to an illustrative embodiment to provide functionality for data layout selection among storage devices associated with nodes of a distributed file system cluster. The information processing system 100 includes one or more host devices 102-1, 102-2, ... 102-M (collectively referred to as storage array 106) communicating via a network 104 with one or more storage arrays 106-1, 106-2, ... 106-M (collectively referred to as storage array 106). The network 104 may include a storage area network (SAN).

[0021] Storage array 106-1 (e.g.) Figure 1 The storage array 106-1 (shown) includes multiple storage devices 108, each storing data utilized by one or more applications running on host device 102. The storage devices 108 are illustratively arranged in one or more storage pools. Storage array 106-1 also includes one or more storage controllers 110 that facilitate I / O processing of the storage devices 108. Storage array 106-1 and its associated storage devices 108 are examples of what is more generally referred to herein as a “storage system.” Such a storage system in this embodiment is shared by host device 102 and is therefore also referred to herein as a “shared storage system.” In embodiments where only a single host device 102 exists, host device 102 can be configured to exclusively use the storage system. In some embodiments, storage array 106 may be part of a storage cluster (e.g., where storage array 106 may be used to implement one or more storage nodes in a clustered storage system comprising multiple storage nodes interconnected by one or more networks), and it is assumed that host device 102 submits I / O operations to be processed by the storage cluster.

[0022] Host device 102 illustratively includes a corresponding computer, server, or other type of processing device capable of communicating with storage array 106 via network 104. For example, at least a subset of host device 102 may be implemented as a corresponding virtual machine of a computing service platform or other type of processing platform. In such an arrangement, host device 102 illustratively provides computing services, such as executing one or more applications on behalf of each of one or more users associated with a corresponding host device in host device 102.

[0023] The term “user” in this article is intended to be interpreted broadly as encompassing numerous arrangements of human, hardware, software, or firmware entities, and combinations thereof.

[0024] Computing and / or storage services can be provided to users based on Platform as a Service (PaaS), Infrastructure as a Service (IaaS), and / or Function as a Service (FaaS) models; however, it is understood that many other cloud infrastructure deployments can be used. Furthermore, illustrative implementations can be implemented outside the context of cloud infrastructure, such as in the case of stand-alone computing and storage systems implemented within a given enterprise.

[0025] Storage device 108 of storage array 106-1 can implement logical units (LUNs) configured to store user objects associated with host device 102. These objects may include files, blocks, or other types of objects. Host device 102 interacts with storage array 106-1 using read / write commands and other types of commands transmitted over network 104. In some embodiments, such commands more specifically include Small Computer System Interface (SCSI) commands, but other types of commands may be used in other embodiments. The terminology used extensively herein for a given I / O operation descriptively includes one or more such commands. References to terms such as “input-output” and “I / O” herein should be understood to refer to input and / or output. Therefore, an I / O operation involves at least one of input and output.

[0026] Furthermore, as used herein, the term "storage device" is intended to be broadly interpreted to encompass, for example, logical storage devices, such as LUNs or other logical storage volumes. A logical storage device can be defined in storage array 106-1 as a distinct portion comprising one or more physical storage devices. Therefore, storage device 108 can be considered to include a corresponding LUN or other logical storage volume.

[0027] The storage device 108 of the storage array 106-1 can be implemented using a solid-state drive (SSD). Such an SSD is implemented using a non-volatile memory (NVM) device such as flash memory. Other types of NVM devices that can be used to implement at least a portion of the storage device 108 include non-volatile random access memory (NVRAM), phase-change RAM (PC-RAM), and magnetic RAM (MRAM). Various types of NVM devices or other storage devices, as well as various combinations thereof, can also be used. For example, a hard disk drive (HDD) can be used in conjunction with or in place of an SSD or other type of NVM device. Therefore, at least a subset of the storage device 108 can be implemented using a wide variety of other types of electronic or magnetic media.

[0028] Assume that at least one of the storage controllers of storage array 106 (e.g., storage controller 110 of storage array 106-1) implements the functionality for selecting the data layout among the storage devices 108 of storage array 106-1 for providing storage for a distributed file system. This functionality is provided via storage performance parameter imbalance rate calculation module 112 and data layout selection module 114. Assume that storage array 106-1 includes a given node of two or more nodes in a distributed file system cluster and receives requests to store one or more portions of data in the distributed file system of the distributed file system cluster.

[0029] Storage performance parameter imbalance rate calculation module 112 is configured to monitor multiple performance parameters of at least a subset of storage devices 108 of storage array 106-1, which provide storage space for the distributed file system of a distributed file system cluster. Such performance parameters may include a first performance parameter characterizing (e.g., for data stored on each of the storage devices 108 of storage array 106-1 providing storage space for the distributed file system of the distributed file system cluster) the frequency of data access and at least a second performance parameter characterizing (e.g., for each of the storage devices 108 of storage array 106-1 providing storage space for the distributed file system of the distributed file system cluster) the available space. Storage performance parameter imbalance rate calculation module 112 is configured to determine an overall performance metric for each of two or more storage devices associated with a given node, at least in part, based on the monitored performance parameters. The overall performance metric illustratively characterizes a combination of the load and percentage of used space of the storage devices 108 of storage array 106-1 providing storage space for the distributed file system of the distributed file system cluster. Differences in the overall performance metric values ​​indicate the imbalance rate therebetween. This type of imbalance rate can be further subdivided for each parameter.

[0030] The data layout selection module 114 is configured to select, at least in part, one of the storage devices 108 of the storage array 106-1, which provides storage space for the distributed file system of the distributed file system cluster, to store different data portions based on determined overall performance metrics. This selection is illustratively performed to reduce the imbalance rate of the storage devices 108 of the storage array 106-1, which provides storage space for the distributed file system of the distributed file system cluster. The storage array 106-1 then stores the different data portions according to this selection.

[0031] In some implementation schemes, Figure 1 The storage array 106 in the embodiments provides or implements multiple different storage tiers of a tiered storage system. For example, a given tiered storage system may include a speed tier or performance tier implemented using flash storage devices or other types of SSDs and a capacity tier implemented using HDDs, wherein one or more of such tiers may be server-based. It will be apparent to those skilled in the art that a variety of other types of storage devices and tiered storage systems may be used in other embodiments. The specific storage device used in a given storage tier may vary depending on the specific requirements of a given embodiment, and a variety of different storage device types may be used in a single storage tier. As previously indicated, the term “storage device” as used herein is intended to be interpreted broadly and may therefore encompass, for example, SSDs, HDDs, flash drives, hybrid drives, or other types of storage products and devices or portions thereof, and illustratively includes logical storage devices such as LUNs.

[0032] It should be understood that a tiered storage system may include more than two storage tiers, such as one or more "performance" tiers and one or more "capacity" tiers, wherein the performance tiers illustratively provide increased IO performance characteristics relative to the capacity tiers, and the capacity tiers illustratively use storage that is relatively less expensive than the performance tiers. There may also be multiple performance tiers, each providing a different level of service or performance as needed, or there may be multiple capacity tiers.

[0033] Despite Figure 1In one implementation, the storage performance parameter imbalance rate calculation module 112 and the data layout selection module 114 are shown to be implemented inside the storage array 106-1 and outside the storage controller 110. However, in other implementations, one or both of the storage performance parameter imbalance rate calculation module 112 and the data layout selection module 114 may be implemented at least partially inside the storage controller 110, or at least partially outside the storage array 106-1, such as on one of the host devices 102, on one or more other storage arrays 106-2 to 106-M, or on one or more servers outside the host device 102 and storage array 106 (e.g., including implementation on a cloud computing platform or other types of information technology (IT) infrastructure). Furthermore, although... Figure 1 Not shown, but other storage arrays in storage arrays 106-2 to 106-M can implement corresponding instances of storage performance parameter imbalance rate calculation module 112 and data layout selection module 114.

[0034] At least some of the functionality of the storage performance parameter imbalance rate calculation module 112 and the data layout selection module 114 can be implemented, at least partially, in the form of software stored in memory and executed by the processor.

[0035] Figure 1 In the embodiments described, the host device 102 and storage array 106 are assumed to be implemented using at least one processing platform, wherein each processing platform includes one or more processing units, each processing unit having a processor coupled to memory. Such processing units may illustratively include specific arrangements of computing, storage, and networking resources. For example, in some embodiments, the processing units are implemented at least in part using virtual resources such as virtual machines (VMs) or Linux containers (LXCs) or a combination of both, such as in an arrangement where Docker containers or other types of LXCs are configured to run on VMs.

[0036] Although the host device 102 and the storage array 106 can be implemented on correspondingly different processing platforms, many other arrangements are possible. For example, in some embodiments, one or more of the host devices 102 and at least a portion of one or more of the storage arrays 106 are implemented on the same processing platform. One or more of the storage arrays 106 can therefore be implemented, at least in part, within at least one processing platform that implements at least a subset of the host devices 102.

[0037] Network 104 can be implemented using a variety of different types of networks to interconnect storage system components. For example, network 104 may include a SAN as part of a global computer network such as the Internet, but other types of networks may be part of a SAN, including wide area networks (WANs), local area networks (LANs), satellite networks, telephone or wired networks, cellular networks, wireless networks (such as WiFi or WiMAX networks), or various parts or combinations of these and other types of networks. Therefore, in some embodiments, network 104 includes a combination of a variety of different types of networks, each network including processing means configured to communicate using Internet Protocol (IP) or other relevant communication protocols.

[0038] As a more specific example, some implementations may utilize one or more high-speed local area networks, in which associated processing devices communicate with each other using peripheral interconnect high-speed (PCIe) cards and networking protocols such as InfiniBand, Gigabit Ethernet, or Fibre Channel. As those skilled in the art will understand, numerous alternative networking arrangements are possible in a given implementation.

[0039] While in some embodiments certain commands used by host device 102 to communicate with storage array 106 illustratively include SCSI commands, other types of commands and command formats may be used in other embodiments. For example, some embodiments may utilize command features and functionality associated with NVM Express (NVMe) as described in NVMe specification revision 1.3 of May 2017, which is incorporated herein by reference. Other storage protocols of this type that may be utilized in the illustrative embodiments disclosed herein include: architecture-based NVMe, also known as NVMoF; and Transmission Control Protocol (TCP)-based NVMe, also known as NVMe / TCP.

[0040] Assume that the memory array 106-1 in this embodiment includes persistent memory implemented using flash memory or other types of non-volatile memory. More specific examples include NAND-based flash memory or other types of non-volatile memory, such as resistive RAM, phase-change memory, spin torque transfer magnetoresistive RAM (STT-MRAM), and 3D XPoint-based... TM Intel Optane memory TMDevice. Further assuming that the persistent memory is separate from the storage device 108 of storage array 106-1, however, in other embodiments, the persistent memory may be implemented as one or more designated portions of one or more of the storage devices 108. For example, in some embodiments, such as those involving an all-flash memory array, storage device 108 may include a flash-based storage device, or may be implemented wholly or partially using other types of non-volatile memory.

[0041] As mentioned above, communication between host device 102 and storage array 106 can utilize PCIe connections or other types of connections implemented through one or more networks. For example, illustrative embodiments may use interfaces such as Internet SCSI (iSCSI), Serial Attached SCSI (SAS), and Serial ATA (SATA). In other embodiments, numerous other interfaces and associated communication protocols may be used.

[0042] In some implementations, storage array 106 can be implemented as part of a cloud-based system.

[0043] Therefore, it should be apparent that the term “memory array” as used herein is intended to be interpreted broadly and may encompass several different instances of commercially available memory arrays.

[0044] Other types of storage products that can be used to implement a given storage system in the illustrative embodiments include software-defined storage, cloud storage, object-based storage, and scale-out storage. In the illustrative embodiments, combinations of these and other storage types can also be used to implement a given storage system.

[0045] In some implementations, the storage system includes a first storage array and a second storage array arranged in an active-active configuration. For example, such an arrangement can be used to ensure that data stored in one storage array is replicated to the other using a synchronous replication process. This data replication across multiple storage arrays can be used to facilitate fault recovery in system 100. Thus, one of the storage arrays can act as a production storage array relative to another storage array that serves as a backup or recovery storage array.

[0046] However, it should be understood that the embodiments disclosed herein are not limited to active-active configuration or any other particular storage system arrangement. Therefore, the illustrative embodiments in this document can be configured using a variety of other arrangements, including, for example, active-passive arrangements, active-active asymmetric logical unit access (ALUA) arrangements, and other types of ALUA arrangements.

[0047] These and other storage systems may be part of what is more generally referred to herein as a processing platform, which includes one or more processing units, each including a processor coupled to memory. A given such processing unit may correspond to one or more virtual machines or other types of virtualization infrastructure, such as Docker containers or other types of LXC. As indicated above, communication between such elements of system 100 may take place on one or more networks.

[0048] As used herein, the term "processing platform" is intended to be interpreted broadly to include, for example, but not limited to, multiple sets of processing devices and one or more associated storage systems configured to communicate over one or more networks. For example, a distributed implementation of host device 102 is possible, wherein some host devices of host device 102 reside in a data center located in a first geographic location, while other host devices of host device 102 reside in one or more other data centers located in one or more other geographic locations that may be far from the first geographic location. Storage array 106 may be implemented at least partially in the first geographic location, the second geographic location, and one or more other geographic locations. Thus, in some implementations of system 100, different host devices of host device 102 and storage array 106 may reside in different data centers.

[0049] Numerous other distributed implementations of host device 102 and storage array 106 are possible. Therefore, host device 102 and storage array 106 can also be implemented in a distributed manner across multiple data centers.

[0050] The following will combine Figure 10 and Figure 11 Additional examples of the processing platform used to implement part of system 100 in the illustrative implementation are described in more detail.

[0051] It should be understood that Figure 1 The specific set of elements shown for selecting data layout among storage devices associated with nodes in a distributed file system cluster is presented by way of illustrative example only, and in other embodiments, additional or alternative elements may be used. Thus, another embodiment may include additional or alternative systems, devices, and other network entities, as well as different arrangements of modules and other components.

[0052] It should be understood that these and other features of the illustrative implementation are presented by way of example only and should not be construed as restrictive in any way.

[0053] Now refer to Figure 2The flowchart describes in more detail an exemplary process for selecting data layout among storage devices associated with nodes of a distributed file system cluster. It should be understood that this particular process is merely an example, and additional or alternative processes for selecting data layout among storage devices associated with nodes of a distributed file system cluster may be used in other implementations.

[0054] In this implementation, the process includes steps 200 to 208. These steps are assumed to be performed by the storage performance parameter imbalance rate calculation module 112 and the data layout selection module 114. The process begins at step 200: receiving a request to store one or more data portions at a given node of two or more nodes in the distributed file system cluster. The distributed file system cluster may include a Hadoop Distributed File System (Hadoop Distributed File System), and the given node includes a data node of the Hadoop Distributed File System. Requests may be received at the given node from additional nodes of the two or more nodes in the distributed file system cluster, the additional nodes including the name node of the Hadoop Distributed File System. The one or more data portions may include one or more data blocks of one or more files stored in the distributed file system of the distributed file system cluster.

[0055] In step 202, two or more performance parameters of each of the two or more storage devices associated with a given node are monitored. The two or more performance parameters include a first performance parameter characterizing data access frequency and at least a second performance parameter characterizing available space. Monitoring the first performance parameter may include monitoring a measure of the access frequency of data stored on each of the two or more storage devices associated with the given node. A given value for the first performance parameter of one of the two or more storage devices associated with the given node may include the ratio of a given access frequency measure of the given storage device to the sum of the access frequency measures of the two or more storage devices associated with the given node. Monitoring the second performance parameter may include monitoring a measure of the percentage of used space of each of the two or more storage devices associated with the given node, where the percentage of used space of one of the two or more storage devices includes the ratio of the used space of the given storage device to the total space of the given storage device. A given value for the second performance parameter of the given storage device may include the ratio of the percentage of used space of the given storage device to the sum of the percentage of used space measures of the two or more storage devices associated with the given node.

[0056] Figure 2The process continues to step 204: determining the overall performance metric for each of the two or more storage devices associated with a given node, at least in part, based on two or more monitored performance parameters. Determining the given value of the overall performance metric for one of the two or more storage devices associated with a given node may include calculating a weighted sum of the values ​​of the two or more performance parameters for the given storage device.

[0057] In step 206, at least one storage device from two or more storage devices associated with a given node is selected to store one or more data portions, based at least in part on the determined overall performance metrics. In step 208, the one or more data portions are stored on the selected at least one storage device from the two or more storage devices associated with the given node. In some embodiments, the request to store one or more data portions includes a request to store two or more blocks, each having a specified block size. In such embodiments, step 206 may include selecting a first storage device from two or more storage devices to store at least a first block of the two or more blocks and selecting a second storage device from two or more storage devices to store at least a second block of the two or more blocks. Step 206 may also or alternatively include: selecting at least a first storage device from two or more storage devices associated with a given node to store at least a first block of two or more blocks; updating the overall performance metrics of the two or more storage devices in response to determining changes in the values ​​of two or more parameters due to storing the at least first block of two or more blocks; and using the updated overall performance metrics to select at least a second storage device from two or more storage devices associated with a given node to store at least a second block of two or more blocks.

[0058] In some implementations, selecting at least one storage device associated with a given node to store one or more portions of data in step 206 is further based at least in part on reducing at least one imbalance rate associated with the given node. The at least one imbalance rate may include a first imbalance rate and a second imbalance rate, the first imbalance rate representing the difference in values ​​of a first parameter of two or more storage devices associated with the given node, the second imbalance rate representing the difference in values ​​of a second parameter of two or more storage devices associated with the given node, and step 206 may further be based at least in part on reducing the first imbalance rate relative to the second imbalance rate according to weights assigned to the first and second parameters.

[0059] This illustrative implementation provides a technique for selecting an optimal data layout strategy in a storage system. In some implementations, it is assumed that the storage system comprises a storage array (also called a data node) having multiple storage devices (also called disks) capable of storing new data. The optimal data layout strategy considers both the IO temperature and capacity / used space percentage of the multiple disks on the data node. The IO temperature, capacity, and used space of the different disks on the data node are monitored in real time to generate an overall performance metric for each disk. This overall performance metric characterizes the disk load and space availability. The disk with the lightest load and the largest available space is then selected for incoming data (also called incoming block volumes). In this way, IO load and capacity can be intelligently balanced among the different disks on the data node, improving efficiency compared to existing methods.

[0060] The following describes various implementation schemes in the context of the best block volume selection strategy for Hadoop Distributed File System (HDFS). Figure 3 An exemplary architecture for HDFS is shown, comprising one or more clients 302, a name node 304-1, an optional second-level name node 304-2, and multiple data nodes separated across multiple device racks 305-1 and 305-2 (collectively, rack 305). Rack 305-1 includes data nodes 306-1-1, 306-1-2, ... 306-1-D (collectively, data node 306-1), and rack 305-2 includes a set of data nodes 306-2. Data nodes 306-1 and 306-2 (collectively, data nodes 306) store data blocks. For example, data node 306-1-1 stores data block 361-1-1, data node 306-1-2 stores data block 361-1-2, data node 306-1-D stores data block 361-1-D, and data node 306-2 stores data block 361-2. Data blocks 361-1-1, 361-1-2, ... 361-1-D are collectively referred to as data block 361-1, and data blocks 361-1 and 361-2 are collectively referred to as data block 361.

[0061] HDFS utilizes a master / slave architecture, where NameNode 304-1 acts as the "master node," managing the file system namespace and controlling access to files by client 302. An optional second-level NameNode 304-2 is configured to retrieve checkpoints of the file system metadata maintained by NameNode 304-1. Such checkpoints can include an edit log indicating a series of changes made to the file system after NameNode 304-1 has started. Second-level NameNode 304-2 can periodically apply such edit logs to snapshots of the file system image to create a new file system image that is copied back to NameNode 304-1. NameNode 304-1 can then use the new file system image on its next restart, thereby reducing restart time because the number of file system edits to be merged is reduced (e.g., only those edits that occurred after the new file system image was recently copied to NameNode 304-1 need to be merged).

[0062] Data nodes 306 each manage a set of disks for storing data. HDFS exposes a file system namespace, allowing data to be stored in files. Internally, each file is divided into one or more blocks, which are stored by data nodes 306 across a set of disks managed by the file. Name node 304-1 manages file system namespace operations (e.g., opening, closing, and renaming files and directories) and also determines the mapping of data blocks 361 to different data nodes in data node 306. Data nodes 306 serve read and write requests from client 302. Data nodes perform block creation, deletion, and replication as instructed by name node 304-1 (e.g., within different racks 305 and across different data nodes in data node 306). Figure 3 This demonstrates how to perform metadata operations between client 302 and name node 304-1, block operations between name node 304-1 and data node 306, replication operations between different data nodes 306 and rack 305, and read / write operations between client 302 and data node 306.

[0063] Each data node 306 propagates its associated data blocks 361 to a local file system directory, which can be specified using "dfs.datanode.dara.dir" in a configuration file (e.g., an "hdfs-site.xml" file). In a typical installation, each directory (called a "volume" in HDFS terminology) resides on a different device (e.g., a separate HDD or SSD). When a new block is written to HDFS, the data node 306 uses a selection strategy to choose the disk to use for each block. Currently supported selection strategies include round-robin and available space. The round-robin selection strategy distributes new blocks evenly across available disks, while the available space strategy prioritizes writing data to the disk with the most free space (e.g., by percentage). Figure 4 A data node 406 with three disks 460-1, 460-2 and 460-3 (collectively referred to as disk 460) is shown, wherein the data node 406 implements a round-robin selection strategy to evenly distribute new blocks 461-1, 461-2 and 461-3 across disk 460. Figure 5 A data node 506 with three disks 560-1, 560-2 and 560-3 (collectively referred to as disk 560) is shown, wherein the data node 506 implements an available space selection strategy to preferentially write new data blocks 561-1, 561-2 and 561-3 to the disk 560 with the most free space (e.g., by percentage).

[0064] By default, data nodes in the HDFS architecture use a round-robin selection strategy to write new blocks. This strategy aims to distribute the new block write load evenly across all disks on the data nodes. However, in long-running data node clusters, various events (such as large-scale file deletions in HDFS, adding new disks to data nodes via hot-swap features, etc.) can still cause severe disk imbalance on one or more data nodes. This imbalance occurs in… Figure 4 The example illustrates three disks (460) with the same total capacity but storing significantly different amounts of data. Even if the data nodes were to switch to an available space selection strategy, disk imbalance could still lead to reduced operational efficiency (e.g., reduced disk I / O efficiency). For example, with an available space selection strategy, each new write might go to a newly added empty disk while other disks are idle. This creates a bottleneck on the newly added disk. This type of bottleneck and the resulting imbalance... Figure 5 The example is shown.

[0065] Using the techniques described in this paper, an improved data layout selection strategy is enabled for HDFS data nodes, where the improved volume selection strategy considers both disk I / O temperature and capacity / used space percentage. The system monitors the I / O temperature and used capacity of the disks in the data node in real time to generate overall disk metrics for each disk or storage volume in the data node. The data node then selects the disk or storage volume with the lightest load and available space for new data blocks. In this way, the data node can balance the I / O load and capacity among available disks and achieve better performance than round-robin and available space data layout selection strategies.

[0066] Let I / O temperature be represented by T, capacity by C, and overall disk performance by P. volume This indicates the I / O temperature of the block volume, and T disk This represents the disk I / O temperature, where:

[0067]

[0068] M represents the current volume count on the data node's disk. C represents the percentage of disk space used. disk It can be calculated as follows:

[0069]

[0070] Overall or total disk performance P is a weighted metric that combines multiple disk metrics. In some implementations, these metrics include I / O temperature (T) and percentage of used space (C). Higher I / O temperature values ​​and higher percentage of used space indicate a busier disk (and therefore lower performance for incoming I / O traffic). The weighted metric P can be calculated as follows:

[0071] P = ω T *T+ω C *C

[0072] Where ω T and ω C These are the weights of the I / O temperature and available space standards, respectively, and ω... T +ω C =1. By adjusting the weight ω T and ω C This can lead to better overall balance results. Although only two criteria are used to calculate P in the example above, it should be understood that in other implementations, various other criteria can be used to calculate P as a supplement to or alternative to the IO temperature and the space percentage standard.

[0073] Based on the weighted standard P, such as Figure 6As shown in process flow 600, the optimal block volume selection strategy algorithm is implemented. Process flow 600 begins at step 601, when a new block volume is being written to the HDFS of a specific data node, and the data node must choose which disk to lay out the new block volume on. In step 603, the real-time IO temperature ratio for each disk is calculated according to the following equation:

[0074]

[0075] Where N is the number of disks in the data node.

[0076] In step 605, the real-time used space ratio of each disk is calculated according to the following equation:

[0077]

[0078] In step 607, the overall disk computing performance metric P in the data node is calculated at least in part based on the real-time IO temperature ratio and used space ratio calculated for each disk. disk Calculate P for each disk i in the data node according to the following equation. disk :

[0079]

[0080] In step 609, at least in part, the performance metric P is based on calculation. disk To select the target disk for the incoming block volume. This can include selecting the disk with the smallest P... disk The disk with the minimum value is selected, or if multiple disks with the same minimum value exist, they are randomly selected from among them. Process flow 600 then ends at 611. In some implementations, process flow 600 may be repeated for each incoming block volume to obtain updated real-time IO temperature and used space ratio. In other implementations, process flow 600 is repeated when the previously calculated real-time IO temperature and used space ratio become outdated (e.g., according to a certain threshold time).

[0081] In order to effectively measure by Figure 6 The performance of the optimal block volume selection strategy algorithm shown in process flow 600 is used to determine the IO temperature imbalance rate and used space percentage imbalance rate of the data nodes. The average IO temperature of the disks in the data nodes is calculated based on the following:

[0082]

[0083] Where N is the disk count of the data node. The standard deviation of the disk I / O temperature in the data node is calculated based on the following:

[0084]

[0085] Standard deviation σ is a measure of how I / O temperature values ​​are distributed. A low standard deviation indicates that the disk's I / O temperature tends to be close to the set mean (also known as the expected value), while a high standard deviation indicates that the values ​​are distributed over a wider range. The I / O temperature imbalance rate (TIB) of a data node is calculated based on the following:

[0086]

[0087] The higher the TIB value, the more unbalanced the disk I / O temperature of the data node. The lower the TIB value, the more balanced the disk I / O temperature of the data node.

[0088] Calculate the average percentage of available disk space in a data node based on the following:

[0089]

[0090] Where N is the disk count of the data node. The standard deviation of the percentage of available disk space in the data node is calculated based on the following:

[0091]

[0092] Standard deviation μ is a measure of how the percentage of available space is distributed. A low standard deviation indicates that the percentage of available space on a disk tends to be close to the mean (also known as the expected value) of the set, while a high standard deviation indicates that the value is distributed over a wider range. The used space percentage imbalance rate (CIB) of a data node is calculated based on the following:

[0093]

[0094] The higher the CIB value, the more unbalanced the percentage of available disk space on the data nodes. The lower the CIB value, the more balanced the percentage of available disk space on the data nodes.

[0095] Now regarding Figures 7A to 7D Describe a specific example where it is assumed that data node 706 comprises three disks 760-1, 760-2, and 760-3 (collectively referred to as disk 760). In this example, it is assumed that the block volume size is 128 megabytes (MB), and that each newly introduced block volume write results in a real-time IO temperature of 20 units. Figure 7A Table 700 is shown, indicating the total size, used size, and real-time I / O temperature of disk 760. The total size and used size are listed in gigabytes (GB). Given the above, assume a large file write request, comprising 10 concurrent block volumes (denoted as blocks 761-1 to 761-10, collectively referred to as block 761) that need to be written to data node 706, and ω... T =C =0.5. Figure 7B The application of the cyclic selection strategy is shown, where blocks 761-1, 761-4, 761-7 and 761-10 are stored on disk 760-1, blocks 761-2, 761-5 and 761-8 are stored on disk 760-2, and blocks 761-3, 761-6 and 761-9 are stored on disk 760-3. Figure 7C The application of the available space selection strategy is shown, where blocks 761-1, 761-2, 761-4, 761-5, 761-7, 761-8, and 761-10 are stored on disk 760-2, and blocks 761-3, 761-6, and 761-9 are stored on disk 760-3. Figure 7D This demonstrates the application of the optimal data layout selection strategy (e.g., Figure 6 The process flow 600), wherein blocks 761-3, 761-5 and 761-9 are stored on disk 760-1, blocks 761-7 and 761-10 are stored on disk 760-2, and blocks 761-1, 761-2, 761-4, 761-6 and 761-8 are stored on disk 760-3.

[0096] Figure 8 A graph 800 shows the imbalance rates (TIB and CIB) for different data layout selection strategies. As shown, the cyclic selection strategy (such as...) Figure 7B (as shown) and available space selection strategies (such as) Figure 7C (As shown) a data node 706 does not achieve a sufficiently balanced I / O temperature and capacity. Optimal data layout selection strategy (e.g., Figure 7D (As shown in the figure) effectively improves the balance of block volume distribution for both IO temperature and capacity standards. Figure 9 Graph 900 shows the TIB and CIB values ​​of data node 706 after selecting a disk for each of the blocks in block 761. As shown, both TIB and CIB decrease with each block volume option. T and ω C Specific values ​​can be adjusted as needed (e.g., to emphasize the relative reduction of TIB and CIB).

[0097] It should be understood that the specific advantages described above and elsewhere herein are associated with specific illustrative embodiments and are not required to exist in other embodiments. Moreover, the features and functionality of the specific type of information processing system, as shown in the accompanying drawings and as described above, are merely exemplary, and numerous other arrangements may be used in other embodiments.

[0098] Now refer to Figure 10 and Figure 11A more detailed illustrative embodiment of a processing platform is described, which enables functionality for selecting data layouts among storage devices associated with nodes in a distributed file system cluster. Although described in the context of system 100, these platforms may also be used to implement at least a portion of other information processing systems in other embodiments.

[0099] Figure 10 An exemplary processing platform including cloud infrastructure 1000 is shown. Cloud infrastructure 1000 includes components that can be used to implement… Figure 1 The information processing system 100 comprises at least a portion of the physical and virtual processing resources. The cloud infrastructure 1000 includes multiple virtual machines (VMs) and / or container sets 1002-1, 1002-2, ... 1002-L implemented using virtualization infrastructure 1004. Virtualization infrastructure 1004 runs on physical infrastructure 1005 and illustratively includes one or more hypervisors and / or operating system-level virtualization infrastructures. Operating system-level virtualization infrastructure illustratively includes the kernel control group of a Linux operating system or other types of operating systems.

[0100] The cloud infrastructure 1000 further includes a set of applications 1010-1, 1010-2, ... 1010-L running on corresponding sets of VMs / containers in VM / container sets 1002-1, 1002-2, ... 1002-L, under the control of the virtualization infrastructure 1004. The VM / container set 1002 may include a corresponding VM, a corresponding set of one or more containers, or a corresponding set of one or more containers running in a VM.

[0101] exist Figure 10 In some implementations of the scheme, the VM / container set 1002 includes corresponding VMs implemented using a virtualization infrastructure 1004 including at least one hypervisor. A hypervisor platform can be used to implement the hypervisor within the virtualization infrastructure 1004, wherein the hypervisor platform has an associated virtual infrastructure management system. The underlying physical machines may include one or more distributed processing platforms, which include one or more storage systems.

[0102] exist Figure 10 In other implementations of the scheme, the VM / container set 1002 includes corresponding containers implemented using virtualization infrastructure 1004 that provides operating system-level virtualization functionality, such as support for Docker containers running on bare metal hosts or running on VMs. The containers are implemented illustratively using the corresponding kernel control group of the operating system.

[0103] As is evident from the above, one or more of the processing modules or other components of system 100 may each run on a computer, server, storage device, or other processing platform element. Such a given element may be considered as an example of what is more generally referred to herein as a "processing device". Figure 10 The cloud infrastructure 1000 shown can represent at least a portion of a processing platform. Another example of such a processing platform is... Figure 11 The processing platform 1100 shown in the figure.

[0104] In this embodiment, the processing platform 1100 includes a portion of the system 100 and includes a plurality of processing devices denoted as 1102-1, 1102-2, 1102-3, ..., 1102-K, which communicate with each other via a network 1104.

[0105] Network 1104 may include any type of network, such as global computer networks (such as the Internet), WANs, LANs, satellite networks, telephone or wired networks, cellular networks, wireless networks (such as WiFi or WiMAX networks), or portions or combinations of these and other types of networks.

[0106] The processing device 1102-1 in the processing platform 1100 includes a processor 1110 coupled to a memory 1112.

[0107] Processor 1110 may include a microprocessor, microcontroller, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), video processing unit (VPU) or other types of processing circuitry, as well as portions or combinations of such circuitry elements.

[0108] Memory 1112 may take any combination of forms including random access memory (RAM), read-only memory (ROM), flash memory, or other types of memory. Memory 1112 and other memories disclosed herein should be considered as illustrative examples of what is more generally referred to as a “processor-readable storage medium” storing executable program code of one or more software programs.

[0109] Articles of manufacture including such processor-readable storage media are considered illustrative embodiments. A given such article of manufacture may include, for example, a storage array, a storage disk, or an integrated circuit containing RAM, ROM, flash memory, or other electronic memory, or any of a variety of other types of computer program products. As used herein, the term "article of manufacture" should be understood to exclude transient propagated signals. Numerous other types of computer program products including processor-readable storage media may be used.

[0110] The processing device 1102-1 also includes a network interface circuit 1114 for interfacing the processing device with the network 1104 and other system components, and may include a conventional transceiver.

[0111] The other processing devices 1102 of the processing platform 1100 are assumed to be configured in a manner similar to that shown for the processing device 1102-1 in the figure.

[0112] Furthermore, the specific processing platform 1100 shown in the figure is presented only by way of example, and the system 100 may include additional or alternative processing platforms, as well as a number of different processing platforms in any combination, wherein each such platform includes one or more computers, servers, storage devices or other processing devices.

[0113] For example, other processing platforms used to implement the illustrative implementation scheme may include converged infrastructure.

[0114] Therefore, it should be understood that in other embodiments, different arrangements of additional or alternative elements may be used. At least a subset of these elements may be implemented together on a common processing platform, or each such element may be implemented on a separate processing platform.

[0115] As previously indicated, components of the information processing system disclosed herein can be implemented, at least in part, in the form of one or more software programs stored in memory and executed by a processor of a processing device. For example, at least part of the functionality disclosed herein for selecting data layout among storage devices associated with nodes of a distributed file system cluster is illustratively implemented in the form of software running on one or more processing devices.

[0116] It should be emphasized again that the above embodiments are presented for illustrative purposes only. Many variations and other alternative embodiments can be used. For example, the disclosed technology is applicable to a variety of other types of information processing systems, storage systems, etc. Moreover, the specific configurations of the system and apparatus elements illustrated in the accompanying drawings, as well as the associated processing operations, may change in other embodiments. Furthermore, the various assumptions made above in describing the illustrative embodiments should be considered exemplary and not as requirements or limitations of this disclosure. Numerous other alternative embodiments within the scope of the appended claims will be apparent to those skilled in the art.

Claims

1. An apparatus comprising: At least one processing device, the at least one processing device including a processor coupled to a memory; The at least one processing device is configured to perform the following steps: Receive a request to store two or more portions of data at a given node in a distributed file system cluster; The given node of the distributed file system cluster monitors two or more performance parameters of each of two or more storage devices associated with the given node, the two or more performance parameters including a first performance parameter characterizing the frequency of data access and at least a second performance parameter characterizing the available space; The overall performance metrics of each of the two or more storage devices associated with the given node are determined by the given node of the distributed file system cluster based at least in part on two or more monitored performance parameters; The imbalance rate associated with the given node is determined by the given node in the distributed file system cluster; The given node of the distributed file system cluster selects, at least a first storage device from the two or more storage devices associated with the given node, to store a first data portion of the two or more data portions, and at least a second storage device from the two or more storage devices associated with the given node, to store a second data portion of the two or more data portions, based at least in part on the determined overall performance metrics and the determined imbalance rate. as well as The given node of the distributed file system cluster stores the first and second data portions of the two or more data portions on selected first and second storage devices among the two or more storage devices associated with the given node.

2. The device of claim 1, wherein the at least one processing means comprises a storage controller of the given node.

3. The device of claim 1, wherein the distributed file system cluster includes a Hadoop Distributed File System, wherein the given node includes a data node of the Hadoop Distributed File System, and wherein the request is received at the given node from an additional node of one of the two or more nodes of the distributed file system cluster, the additional node including a name node of the Hadoop Distributed File System.

4. The device of claim 1, wherein the two or more data portions comprise two or more data blocks of two or more files stored in the distributed file system of the distributed file system cluster.

5. The device of claim 1, wherein monitoring the first performance parameter includes monitoring a measure of the access frequency of data stored on each of the two or more storage devices associated with the given node.

6. The device of claim 5, wherein a given value of the first performance parameter of one of the two or more storage devices associated with the given node includes the ratio of a given access frequency metric of the given storage device to the sum of the access frequency metrics of the two or more storage devices associated with the given node.

7. The device of claim 1, wherein monitoring the second performance parameter comprises monitoring a percentage of used space of each of the two or more storage devices associated with the given node, the percentage of used space of a given storage device comprising the ratio of the used space of the given storage device to the total space of the given storage device.

8. The device of claim 7, wherein the given value of the second performance parameter of the given storage device includes the ratio of the used space percentage measure of the given storage device to the sum of the used space percentage measures of the two or more storage devices associated with the given node.

9. The device of claim 1, wherein determining a given value for the overall performance metric of a given storage device among the two or more storage devices associated with the given node comprises calculating a weighted sum of the values ​​of the two or more performance parameters of the given storage device.

10. The device of claim 1, wherein the request to store the two or more data portions includes a request to store two or more blocks, each having a specified block size.

11. The device of claim 1, wherein selecting the first and second storage devices among the two or more storage devices associated with the given node for storing the two or more data portions comprises: In response to selecting a first storage device to store the first data portion of the two or more data portions, the overall performance metrics and the determined imbalance rate of the two or more storage devices are dynamically updated before selecting a second storage device to store the second data portion of the two or more data portions.

12. The device of claim 1, wherein the selection of the first and second storage devices among the two or more storage devices associated with the given node for storing the two or more data portions is further based at least in part on reducing the determined imbalance rate associated with the given node.

13. The device of claim 12, wherein determining the imbalance rate associated with the given node comprises determining a first imbalance rate and determining a second imbalance rate, the first imbalance rate characterizing the difference in values ​​of a first performance parameter of the two or more storage devices associated with the given node, the second imbalance rate characterizing the difference in values ​​of a second performance parameter of the two or more storage devices associated with the given node, and wherein selecting the first and second storage devices among the two or more storage devices associated with the given node to store the two or more data portions further adjusts the first imbalance rate relative to the second imbalance rate based at least in part on weights assigned to the first performance parameter and the second performance parameter.

14. A computer program product comprising a non-transitory processor-readable storage medium in which program code of one or more software programs is stored, wherein the program code, when executed by at least one processing device, causes the at least one processing device to perform the following steps: Receive a request to store two or more portions of data at a given node in a distributed file system cluster; The given node of the distributed file system cluster monitors two or more performance parameters of each of two or more storage devices associated with the given node, the two or more performance parameters including a first performance parameter characterizing the frequency of data access and at least a second performance parameter characterizing the available space; The overall performance metrics of each of the two or more storage devices associated with the given node are determined by the given node of the distributed file system cluster based at least in part on two or more monitored performance parameters; The imbalance rate associated with the given node is determined by the given node in the distributed file system cluster; The given node of the distributed file system cluster selects, at least a first storage device from the two or more storage devices associated with the given node, to store a first data portion of the two or more data portions, and at least a second storage device from the two or more storage devices associated with the given node, to store a second data portion of the two or more data portions, based at least in part on the determined overall performance metrics and the determined imbalance rate. as well as The given node of the distributed file system cluster stores the first and second data portions of the two or more data portions on selected first and second storage devices among the two or more storage devices associated with the given node.

15. The computer program product of claim 14, wherein the request to store the two or more data portions includes a request to store two or more blocks, each having a specified block size.

16. The computer program product of claim 14, wherein determining the imbalance rate associated with the given node comprises determining a first imbalance rate and determining a second imbalance rate, the first imbalance rate characterizing the difference in values ​​of a first performance parameter of the two or more storage devices associated with the given node, the second imbalance rate characterizing the difference in values ​​of a second performance parameter of the two or more storage devices associated with the given node, wherein selecting the first and second storage devices from the two or more storage devices associated with the given node to store the two or more data portions further adjusts the first imbalance rate at least in part based on the second imbalance rate.

17. A method comprising: Receive a request to store two or more portions of data at a given node in a distributed file system cluster; The given node of the distributed file system cluster monitors two or more performance parameters of each of two or more storage devices associated with the given node, the two or more performance parameters including a first performance parameter characterizing the frequency of data access and at least a second performance parameter characterizing the available space; The overall performance metrics of each of the two or more storage devices associated with the given node are determined by the given node of the distributed file system cluster based at least in part on two or more monitored performance parameters; The imbalance rate associated with the given node is determined by the given node in the distributed file system cluster; The given node of the distributed file system cluster selects, at least a first storage device from the two or more storage devices associated with the given node, to store a first data portion of the two or more data portions, and at least a second storage device from the two or more storage devices associated with the given node, to store a second data portion of the two or more data portions, based at least in part on the determined overall performance metrics and the determined imbalance rate. as well as The first data portion and the second data portion of the two or more data portions are stored by the given node of the distributed file system cluster on a selected first and second storage device among the two or more storage devices associated with the given node; The method is performed by at least one processing device, which includes a processor coupled to a memory.

18. The method of claim 17, wherein the request to store the two or more data portions includes a request to store two or more blocks, each having a specified block size.

19. The method of claim 17, wherein determining the imbalance rate associated with the given node comprises determining a first imbalance rate and determining a second imbalance rate, the first imbalance rate characterizing the difference in values ​​of a first performance parameter of the two or more storage devices associated with the given node, the second imbalance rate characterizing the difference in values ​​of a second performance parameter of the two or more storage devices associated with the given node, wherein selecting the first and second storage devices from the two or more storage devices associated with the given node to store the two or more data portions further adjusts the first imbalance rate at least in part based on the second imbalance rate.

Citation Information

Patent Citations

  • Cluster file system comprising data mover modules having associated quota manager for managing back-end user quotas

    US20180260398A1

  • Hot-pluggable file system interface

    US20200019621A1