Improving data performance by transferring data between storage layers using workload characteristics
By analyzing the characteristics of data workloads and generating corresponding data placement suggestions, the problem that traditional ILM solutions are difficult to manage striped data is solved, and efficient data transmission and processing are achieved in the cluster file system.
Patent Information
- Application Number
- CN201980088214.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-08
- Filing Date
- 2019-12-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-12-16
AI Technical Summary
Traditional information lifecycle management (ILM) solutions are difficult to effectively manage information distribution in a cluster's file system, especially when data is striped across different parts of storage.
By analyzing the workload characteristics of the data, one or more suggestions corresponding to the placement of data in the storage device are generated. These recommendations are used to identify and transmit portions of actual data in storage devices including shared nodes and no shared nodes, thereby performing data transmission between the first and second layers.
It realizes the support of striped and non-striped storage configurations in a single namespace, significantly improving the efficiency and performance of information lifecycle management, and accelerating data transmission and processing by utilizing dedicated hardware.
Smart Images

Figure CN113272781B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] The present invention relates to data storage systems, and more particularly, to using data workload characteristics to improve information lifecycle management.
[0002] Information typically undergoes multiple operations during the time it is maintained in a storage device. For example, after being stored (e.g., written) in storage, portions of the data undergo multiple read operations and / or modification operations. Different portions of the data are also typically moved (e.g., rewritten) to different locations in the storage over time. As storage capacity and data throughput increase over time, the desirability of a storage system that can perform these access operations in an efficient manner also increases.
[0003] In response, many storage systems implement information lifecycle management (ILM). ILM is an integrated approach for managing the flow of data and / or metadata included in an information system from a point of creation (e.g., initial storage) to a point of deletion. However, traditional ILM schemes have been unable to effectively manage the distribution of information in a clustered file system. This is particularly true for a clustered file system in which data is striped across different portions of the storage. SUMMARY OF THE INVENTION
[0004] According to one embodiment, a computer-implemented method includes: receiving one or more suggestions corresponding to placement of data in a storage device, wherein the one or more suggestions are based on data workload characteristics. The one or more suggestions are used to identify portions of actual data stored in an actual storage device corresponding to the one or more suggestions. The actual storage device includes: a first tier having two or more shared nodes; and a second tier having at least one non-shared node. For each portion of the identified portions of the actual data stored in the first tier, the one or more suggestions are also used to determine whether to transfer a given identified portion of the actual data to the second tier. Further, in response to determining to transfer at least one of the identified portions of the actual data to the second tier, sending one or more instructions to transfer the at least one of the identified portions of the actual data from the first tier to the second tier.
[0005] According to another embodiment, a computer program product includes a computer-readable storage medium having program instructions embodied therein. The program instructions are readable and / or executable by a processor to cause the processor to: perform the above method.
[0006] A system according to another embodiment includes: a processor; and logic integrated with, executable by, or integrated with and executable by the processor. Moreover, the logic is configured to: perform the above method.
[0007] According to another embodiment, a computer-implemented method includes: analyzing the workload characteristics of data stored in a file system of a cluster. The file system of the cluster is implemented in a storage device including a first tier having two or more shared nodes and a second tier having at least one shared-nothing node. Further, the first tier and the second tier are included in the same namespace. The analyzed workload characteristics are used to generate one or more recommendations corresponding to the placement of data in the storage device. Further, the one or more recommendations are used to transfer at least some of the data in the storage device between the first and second tiers.
[0008] According to another embodiment, a system includes: a storage device including a first tier having two or more shared nodes and a second tier having at least one shared-nothing node. Further, the first tier and the second tier are included in the same namespace. The system further includes a processor, and logic integrated with, executable by, or integrated with and executable by the processor. The logic is configured to: analyze, by the processor, the workload characteristics of data stored in the storage device. Further use, by the processor, the analyzed workload characteristics to generate one or more recommendations corresponding to the placement of data in the storage device. Further, use, by the processor, the one or more recommendations to transfer at least some of the data in the storage device between the first and second tiers.
[0009] Other aspects and embodiments of the present invention will become apparent from the following detailed description, which illustrates the principles of the present invention by way of example when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 Illustrates a network architecture according to one embodiment.
[0011] Figure 2 Illustrates a representative hardware environment that may be associated with Figure 1 the servers and / or clients of.
[0012] Figure 3 Illustrates a hierarchical data storage system according to one embodiment.
[0013] Figure 4 Is a partial schematic diagram of a file system of a cluster according to one embodiment.
[0014] Figure 5A Is a flowchart of a method according to one embodiment.
[0015] Figure 5B Is a flowchart of a method according to one embodiment.
[0016] Figure 5CFlowchart of a method according to an embodiment. Detailed implementation
[0017] The following description is made for the purpose of illustrating the general principles of the present invention and is not intended to limit the inventive concept claimed herein. Further, the specific features described herein can be used in combination with each of the other described features in different possible combinations and permutations.
[0018] Unless otherwise explicitly defined herein, all terms will be given their broadest possible interpretation, including meanings implied from the specification and understood by those skilled in the art and / or as defined in dictionaries, treatises, etc.
[0019] It must also be noted that, as used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural referents unless otherwise indicated. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0020] The following description discloses several preferred embodiments of systems, methods and computer program products for implementing a framework capable of supporting both striped and non-striped storage configurations in a single (same) namespace. In addition, some embodiments herein are capable of significantly enhancing the ILM scheme to appropriately place or migrate data between different layers based on data workload characteristics and the recommendations and / or data models derived therefrom, some of the different layers being implemented using dedicated hardware, as will be described in further detail below.
[0021] In a general embodiment, a computer-implemented method includes: receiving one or more recommendations corresponding to the placement of data in a storage device, wherein the one or more recommendations are based on data workload characteristics. The one or more recommendations are used to identify the portions of the actual data stored in the actual storage device that correspond to the one or more recommendations. The actual storage device includes: a first layer having two or more shared nodes; and a second layer having at least one shared-nothing node. For each portion of the identified portions of the actual data stored in the first tier, the one or more recommendations are also used to determine whether to transfer a given identified portion of the actual data to the second tier. Further, in response to determining that at least one of the identified portions of the actual data is to be transferred to the second layer, sending one or more instructions to transfer at least one of the identified portions of the actual data from the first layer to the second layer.
[0022] In another general embodiment, a computer program product includes a computer-readable storage medium having program instructions embodied therewith. The program instructions are readable and / or executable by a processor to cause the processor to: perform the method described above.
[0023] In another general embodiment, a system includes: a processor; and logic integrated with, executable by, or integrated with and executable by the processor. Moreover, the logic is configured to: perform the method described above.
[0024] In yet another general embodiment, a computer-implemented method includes: analyzing workload characteristics of data stored in a file system of a cluster. The file system of the cluster is implemented in a memory including a first tier having two or more shared nodes and a second tier having at least one shared-nothing node. In addition, the first tier and the second tier are included in the same namespace. The analyzed workload characteristics are used to generate one or more recommendations corresponding to placement of data in a storage device. Further, the one or more recommendations are used to transfer at least some of the data in the storage device between the first and second tiers.
[0025] In another general embodiment, a system includes: a storage device including a first tier having two or more shared nodes and a second tier having at least one shared-nothing node. In addition, the first tier and the second tier are included in the same namespace. The system further includes a processor, and logic integrated with, executable by, or integrated with and executable by the processor. The logic is configured to: analyze, by the processor, workload characteristics of data stored in the storage device. Further use, by the processor, the analyzed workload characteristics to generate one or more recommendations corresponding to placement of data in the storage device. In addition, use, by the processor, the one or more recommendations to transfer at least some of the data in the storage device between the first and second tiers.
[0026] Figure 1 Illustrates architecture 100 according to one embodiment. As Figure 1 shown, a plurality of remote networks 102 including a first remote network 104 and a second remote network 106 are provided. A gateway 101 may be coupled between the remote networks 102 and a neighboring network 108. In the context of this architecture 100, the networks 104, 106 may each take any form, including but not limited to a local area network (LAN), a wide area network (WAN) (such as the Internet), a public switched telephone network (PSTN), an internal telephone network, and the like.
[0027] In use, gateway 101 acts as an entry point from remote network 102 to adjacent network 108. As such, gateway 101 can act as a router and a switch, where the router is capable of directing a given data packet to gateway 101 and the switch provides an actual path for the given packet into and out of gateway 101.
[0028] Further included is at least one data server 114 that is coupled to adjacent network 108 and accessible from remote network 102 via gateway 101. It should be noted that the data server(s) 114 can include any type of computing device / groupware. Coupled to each data server 114 are a plurality of user devices 116. User devices 116 can also be directly connected via one of networks 104, 106, 108. Such user devices 116 can include desktop computers, laptop computers, handheld computers, printers, or any other type of logic. It should be noted that in one embodiment, user device 111 can also be directly coupled to any network.
[0029] Peripheral device 120 or a series of peripheral devices 120 (e.g., fax machines, printers, networked and / or local storage units or systems, etc.) can be coupled to one or more of networks 104, 106, 108. It should be noted that databases and / or additional components can be used with or integrated into any type of network element coupled to networks 104, 106, 108. In the context of this specification, a network element can refer to any component of a network.
[0030] According to some methods, the methods and systems described herein can be implemented and / or implemented thereon using virtual systems and / or systems that emulate one or more other systems, such as a UNIX system that emulates an IBM z / OS environment, a UNIX system that virtually hosts a MICROSOFT WINDOWS environment, a MICROSOFT WINDOWS system that emulates an IBM z / OS environment, etc. In some embodiments, such virtualization and / or emulation can be enhanced by using VMWARE software.
[0031] In more methods, one or more of networks 104, 106, 108 can represent a cluster of systems commonly referred to as the "cloud". In cloud computing, shared resources such as processing power, peripherals, software, data, servers, etc. are provided to any system in the cloud in an on-demand relationship, thus allowing access and distribution of services across many computing systems. Cloud computing typically involves an Internet connection between systems operating in the cloud, but other technologies for connecting systems can also be used.
[0032] Figure 2 Illustrated is in accordance with one embodiment withFigure 1 A representative hardware environment associated with user device 116 and / or server 114. This figure shows a typical hardware configuration of a workstation, which has a central processing unit 210 such as a microprocessor and multiple other units interconnected via a system bus 212.
[0033] Figure 2 The illustrated workstation includes random access memory (RAM) 214, read-only memory (ROM) 216, an input / output (I / O) adapter 218 for connecting peripheral devices such as a disk storage unit 220 to the bus 212, a user interface adapter 222 for connecting a keyboard 224, a mouse 226, speakers 228, a microphone 232, and / or other user interface devices such as a touch screen and a digital camera (not shown) to the bus 212, a communication adapter 234 for connecting the workstation to a communication network 235 (e.g., a data processing network), and a display adapter 236 for connecting the bus 212 to a display device 238.
[0034] The workstation may have an operating system residing thereon, such as Operating System (OS), MACOS, UNIX OS, etc. It will be appreciated that the preferred embodiments may also be implemented on platforms and operating systems other than those mentioned. The preferred embodiments may be written using Extensible Markup Language (XML), C, and / or C++ languages or other programming languages together with object-oriented programming methods. Object-oriented programming (OOP) can be used, which has become increasingly used for developing complex applications.
[0035] Now referring to Figure 3 , a storage system 300 according to one embodiment is shown. It should be noted that according to different embodiments, Figure 3Some of the components shown may be implemented as hardware and / or software. Storage system 300 may include a storage system manager 312 for communicating with a plurality of media and / or drives on at least one higher storage layer 302 and at least one lower storage layer 306. The higher storage layer 302 preferably may include one or more random access and / or direct access media 304, such as hard disks in a hard disk drive (HDD), non-volatile memory (NVM), solid state storage in a solid state drive (SSD), flash storage, SSD arrays, flash storage arrays, etc., and / or other media noted herein or known in the art. The lower storage layer 306 may preferably include one or more lower performance storage media 308, including sequential access media, such as tapes in a tape drive and / or optical media, slower access HDDs, slower access SSDs, etc., and / or other media noted herein or known in the art. One or more additional storage layers 316 may include any combination of storage media desired by the designer of system 300. Moreover, either the higher storage layer 302 and / or the lower storage layer 306 may contain a certain combination of storage devices and / or storage media.
[0036] The storage system manager 312 may communicate with the drives and / or storage media 304, 308 on the higher storage layer 302 and the lower storage layer 306 via a network 310 (e.g., a storage area network (SAN), as Figure 3 shown, or some other suitable network type). The storage system manager 312 may also communicate with one or more host systems (not shown) via a host interface 314, which may or may not be part of the storage system manager 312. The storage system manager 312 and / or any other components of the storage system 300 may be implemented in hardware and / or software and may utilize a processor (not shown) to execute commands of types known in the art, such as a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc. Of course, any arrangement of the storage system may be used, as will be apparent to those skilled in the art upon reading this specification.
[0037] In more embodiments, the storage system 300 may include any number of data storage tiers and may include the same or different storage media within each storage tier. For example, each data storage tier may contain the same type of storage media, such as HDDs, SSDs, sequential access media (tape in a tape drive, optical disc in an optical disc drive, etc.), direct access media (CD-ROM, DVD-ROM, etc.), or any combination of media storage types. In one such configuration, the higher storage tier 302 may include the majority of the SSD storage media for storing data in a higher performance storage environment, and the remaining storage tiers including the lower storage tier 306 and additional storage tier 316 may include any combination of SSDs, HDDs, tape drives, etc. for storing data in a lower performance storage environment. In this way, data that is accessed more frequently, data with a higher priority, data that requires faster access, etc. may be stored in the higher storage tier 302, while data that does not have one of these attributes may be stored in the additional storage tier 316, including the lower storage tier 306. Of course, those skilled in the art, after reading this specification, may design many other combinations of storage media types according to the embodiments presented herein to implement into different storage schemes.
[0038] According to some embodiments, a storage system (such as 300) may include logic configured to receive a request to open a data set, logic configured to determine whether the requested data set is stored in multiple associated portions in the lower storage tier 306 of the hierarchical data storage system 300, logic configured to move each associated portion of the requested data set to the higher storage tier 302 of the hierarchical data storage system 300, and logic configured to assemble the requested data set on the higher storage tier 302 of the hierarchical data storage system 300 from the associated portions.
[0039] Of course, according to different embodiments, the logic may be implemented as a method or a computer program product on any device and / or system.
[0040] As previously mentioned, many storage systems implement ILM to manage the flow of data and / or metadata in an information system from the point of creation (e.g., initial storage) to the point of deletion. However, traditional ILM schemes have been unable to effectively manage the distribution of information in a clustered file system. This is especially true for a clustered file system in which data is striped across different portions of storage. For example, while striping data across different disks of storage increases the achievable throughput, it also increases the latency in cases involving recalling each fragment from its respective disk, as will be described in further detail below.
[0041] See Figure 4, shows a schematic diagram of a cluster file system 400 according to an embodiment. As an option, the cluster file system 400 of the present cluster may be implemented in combination with features from any other embodiment listed herein, such as those described with reference to other figures (such as Figure 3 ). However, such a cluster file system 400 and other file systems presented herein may be used in different applications and / or permutations, which may or may not be specifically described in the illustrative embodiments listed herein. Further, the cluster file system 400 presented herein may be used in any desired environment. Thus, Figure 4 (and other figures) may be considered to include any possible permutations.
[0042] As shown, the cluster file system 400 includes a first layer 402 and a second layer 404. In a preferred approach, the first and second layers 402, 404 are included in the same namespace of the cluster file system 400. Thus, the first and second layers 402, 404 may be included in the same abstract container or environment that is created to hold a logical grouping of unique identifiers or symbols, for example, as will be recognized by those skilled in the art after reading this specification. Additionally, the first index node structure corresponding to the first layer 402 is preferably maintained separately from the second index node structure corresponding to the second layer 404. However, in some approaches, the first and second layers 402, 404 may each be assigned a unique portion of a combined index node structure.
[0043] Now looking specifically at the first layer 402, the shared nodes 406, 408 are coupled to each of a plurality of different storage components 410 that can be of any desired type. For example, in different approaches, any one or more of the storage components 410 may include HDDs, SSDs, tape libraries, etc. and / or combinations thereof. Thus, in some approaches, one or more of the shared nodes 406, 408 may actually be shared disk nodes. Each of the shared nodes 406, 408 also includes a controller 412 and a portion of the storage 414, which may be used as a cache, for example. It should also be noted that each of the controllers 412 may include or actually be any desired type of processing component, such as a processor, server, CPU, etc., depending on the desired approach.
[0044] The second layer 404 also includes a plurality of different storage components 416, each storage component being coupled to a corresponding shared-nothing node 418. Each of the shared-nothing nodes 418 has a shared-nothing architecture including a distributed computing architecture in which each of the nodes 418 is independent and self-sufficient relative to one another. In some illustrative methods, there is no single point of contention across the second layer 404. Thus, the shared-nothing nodes 418 do not allocate storage and / or computing resources among themselves, as will be recognized by those skilled in the art upon reading this specification.
[0045] Following that the first layer 402 is able to stripe data across the storage components 410 using the shared nodes 406, 408, while the second layer 404 is not able to stripe data across the storage components 416 using the shared-nothing nodes 418. Similarly, each of the shared nodes 406, 408 is coupled to each of the different storage components 410, which allows any one of the shared nodes 406, 408 to write data to and / or read data from any one of the storage components 410. As will be appreciated by those skilled in the art, data striping is a technique of segmenting logical sequential data (e.g., such as a file) such that consecutive segments are stored on different physical storage devices. Striping is useful when a processing device requests data faster than a single storage device can provide it. This is because spreading the segments across multiple devices that can be accessed simultaneously increases the total achievable data throughput. Data striping is also a useful process in order to balance the I / O load across an array of storage components. Additionally, some data striping processes involve interleaving sequential segments of data across storage devices in a cyclic manner starting from the beginning of the data sequence.
[0046] In contrast, each shared-nothing node 418 has a shared-nothing architecture and thus cannot implement data striping. Instead, each shared-nothing node 418 implements a data non-striping mode, such as, for example, the General Parallel File System for the shared-nothing cluster (GPFS-SNC) mode. According to an exemplary method (which is in no way intended to limit the present invention), the GPFS-SNC mode involves a scalable file system operating on a given cluster. It should also be noted that the number of components included in each of the first layer 402 and the second layer 404 is in no way intended to be limiting. Instead, any desired number of nodes, storage components, etc. may be implemented depending on the desired method.
[0047] While striping patterns are desirable in some cases given the parallelism they provide, non-striping patterns also offer benefits. For example, a non-striped architecture can achieve locality awareness that allows computational jobs to be scheduled on the nodes where the data resides. It also implements metachunks that allow large and small chunk sizes to coexist in the same file system, thereby satisfying the requests of different types of applications. Write affinity that allows applications to specify the layout of files across different nodes to maximize write and read bandwidth is also implemented in some approaches. Further, pipelined replication can be used to increase the use of network bandwidth for data replication, while distributed recovery can be used to reduce the impact of failures on ongoing computations. Thus, it is desirable to be able to effectively implement file systems with both striping and non-striping patterns.
[0048] Each of the shared-nothing nodes 418 in the shared-nothing node 418 includes a controller 420, a portion of the memory 422 (e.g., which can be used as a cache in some cases), and dedicated hardware 424. As general data usage and storage capacity continue to increase, any latency associated with performing data access operations is magnified for the overall system. This is especially true for file systems of prior clusters in which data is striped across different portions of storage. For example, while striping data across different storage disks can increase achievable throughput, it also increases latency in cases that involve recalling each segment from its respective storage location.
[0049] To offset this latency, some of the embodiments included herein implement dedicated hardware 424. The dedicated hardware 424 preferably can increase the speed at which the shared-nothing nodes 418 in the second tier 404 can perform data operations. In other words, the dedicated hardware 424 effectively increases the speed of data transfer performed between each of the shared-nothing nodes 418 and their respective storage components 416 coupled thereto. This allows data to be accessed more quickly from the storage component 416, thereby significantly reducing the latency associated with performing read operations, write operations, rewrite operations, etc.
[0050] An illustrative list of components that can be used to form the dedicated hardware 424 includes, but is not limited to, a graphics processing unit (GPU), an SSD cache, an ASIC, a fast non-volatile memory (NVMe), etc. and / or combinations thereof. By way of example, which is in no way intended to limit the invention, for instance, as will be described in further detail below, the dedicated hardware 424 including a GPU can be used to assist in developing a machine learning model. Additionally, each of the shared-nothing nodes 418 can include the same, similar, or different dedicated hardware 424 components, depending on the desired approach. For example, in some approaches, each of the shared-nothing nodes 418 includes dedicated hardware 424 for an SSD cache, while in other approaches, one of the shared-nothing nodes 418 includes dedicated hardware 424 for an SSD cache and another of the shared-nothing nodes 418 includes dedicated hardware 424 for a GPU.
[0051] In some approaches, the communication paths 426 extending between each of the shared-nothing nodes 418 and the corresponding storage components 416 coupled thereto can also accelerate data transfer speeds. By way of example, which is in no way intended to limit the invention, a high-speed peripheral component interconnect express (PCIe) bus serves as the communication path 426 coupling the shared-nothing nodes 418 to the corresponding storage components 416. Additionally, the dedicated hardware 424 can work in conjunction with the PCIe bus to further accelerate data transfer.
[0052] Still referring to Figure 4 , both the first layer 402 and the second layer 404 are connected to the network 428. The first layer 402 and / or the second layer 404 can be coupled to the network 428 using a wireless connection (e.g., WiFi, Bluetooth, cellular network, etc.); a wired connection (e.g., cable, fiber optic link, wire, etc.) or any other type of connection that will be apparent to those skilled in the art after reading this specification. Additionally, the network can be of any type, e.g., depending on the desired approach. For example, in some approaches, the network 428 is a WAN, such as the Internet. However, an illustrative list of other network types that the network 428 can implement includes, but is not limited to, a LAN, a PSTN, a SAN, an internal telephone network, etc. Thus, for example, although located at different geographical locations, the first layer 402 and the second layer 404 can communicate with each other regardless of the amount of separation that exists between them.
[0053] The central controller 430 and the user 432 (e.g., an administrator) are also coupled to the network 428. In some methods, the central controller 430 is used to manage the communication between the first layer 402 and the second layer 404. The central controller 430 can also manage the communication between the user 432 and the file system 400 of the cluster. According to some methods, the central controller 430 receives data, operation requests, commands, formatting instructions, etc. from the user 432 and directs the appropriate portions thereof to the first layer 402 and / or the second layer 404.
[0054] Again, as storage capacity and data throughput increase over time, the desirability of a storage system that can perform data access operations in an efficient manner also increases. Different embodiments among the embodiments included herein can achieve this desired improvement by implementing a storage architecture that allows machine learning and / or deep learning algorithms to be applied in an effective manner. Additionally, the performance characteristics specific to different layers in the file system can be intelligently utilized to further improve performance in real time. For example, now referring to Figure 5A FIG. 500, a flowchart of a computer-implemented method 500 according to one embodiment is shown. In different embodiments, the method 500 can be performed according to the present invention in any environment depicted in Figures 1-4 and so on. Of course, as those skilled in the art will understand after reading this specification, more or fewer operations than those specifically described in Figure 5A can be included in the method 500.
[0055] Each step of the method 500 can be performed by any suitable component of the operating environment. For example, each of the nodes 501, 502, 503 shown in the flowchart of the method 500 can correspond to one or more processors located at different positions in a multi-layer data storage system. Additionally, each of the one or more processors is preferably configured to communicate with each other.
[0056] In different embodiments, the method 500 can be performed partially or entirely by a controller, a processor, etc. or some other device having one or more processors therein. A processor (e.g., a processing circuit, chip, and / or module implemented in hardware and / or software and preferably having at least one hardware component) can be used in any device to perform one or more steps of the method 500. Illustrative processors include, but are not limited to, a central processing unit (CPU), an ASIC, a field programmable gate array (FPGA), etc., combinations thereof, or any other suitable computing device known in the art.
[0057] As described above, Figure 5AIncludes different nodes 501, 502, 503, each of which represents one or more processors, controllers, computers, etc. located at different positions in a multi-layer data storage system. For example, node 501 may include one or more processors (e.g., see 412, 420 above Figure 4 coupled to the layer of the file system of the cluster). Node 502 may include one or more processors that serve as the central controller of the file system of the cluster (e.g., see 430 above Figure 4 ). In addition, node 503 may include one or more processors located at the user location, and the user location communicates with one or more processors at each of nodes 501 and 503 (e.g., via a network connection). Thus, depending on the method, commands, data, requests, etc. may be sent between each of nodes 501, 502, 503. In addition, it should be noted that, for example, as will be understood by those skilled in the art after reading this specification, the different processes included in method 500 are in no way intended to be limiting. For example, in some methods, the data sent from node 502 to node 503 may start with a request sent from node 503 to node 502.
[0058] As shown, operation 504 of method 500 is performed by one or more processors at node 501. Again, it should be noted that one or more processors at node 501 are electrically coupled to a given layer of the file system of the cluster. Thus, in some methods, one or more processors at node 501 include one or more controllers 412 in shared nodes 406, 408. In other methods, one or more processors at node 501 include one or more of controllers 420 in non-shared node 418. In other methods, one or more processors at node 501 may represent one or more controllers 412 in shared nodes 406, 408 and one or more controllers 420 in non-shared node 418. In other words, the process performed by one or more processors at node 501 may be performed by Figure 4 any node among nodes 406, 408, 418 in any layer of layers 402, 404 of the file system of the cluster in
[0059] Still referring to method 500, operation 504 includes collecting data workload characteristics corresponding to the data stored in the file system of the cluster. As described above, the file system of the cluster is implemented in a storage including a first layer with two or more shared nodes and a second layer with at least one non-shared node. In a preferred method, the first layer and the second layer of the file system of the cluster are also included in the same namespace.
[0060] The characteristics of the data workloads collected vary according to the desired method. For example, an illustrative list of data workload characteristics that may be collected in operation 504 includes, but is not limited to, read and / or write patterns, the file type corresponding to the data, the specific portions of the file corresponding to the data (e.g., headers, footers, metadata sections, etc.), and the like. According to some methods, for example, as will be appreciated by those skilled in the art after reading this specification, a supervisor-assisted learning model may be used to identify the data workload characteristics collected in operation 504.
[0061] Operation 506 further includes sending the collected data workload characteristics to node 502. Depending on the method, the data workload characteristics may be sent to node 502 when they are collected, in batches of a predetermined size, periodically, etc.
[0062] In addition, operation 508 includes analyzing the data workload characteristics. The process of analyzing the data workload characteristics varies according to the quantity and / or type of the received workload characteristics. For example, in different approaches, operation 508 may be performed by analyzing the corresponding data type, data size, different read / write patterns, specific access patterns, the proposed framework estimate, the industry type of the file system from which the cluster is used, and / or the workload, etc.
[0063] Then, the analyzed workload characteristics are used to generate one or more suggestions and / or data models corresponding to the placement of specific portions of the data in the storage device. For example, in some methods, the suggestions (also referred to herein as "hints") and / or data models correspond to specific workloads. The generated suggestions are even used in some methods to, for example, identify certain machine learning and / or deep learning algorithms relevant to the situation based on the identified workload. These machine learning and / or deep learning algorithms may be developed over time to describe the data actually included in the storage device and / or the storage locations (e.g., nodes) where the data is stored. Thus, the machine learning and / or deep learning algorithms can be updated (e.g., improved) over time using or at least based on the data workload characteristics.
[0064] In operation 512, these generated suggestions are then sent to node 503 for approval. The user at node 503 (e.g., an administrator) has the ability to accept some, none, or all of the generated suggestions. The user at node 503 is also able to propose one or more additional suggestions and / or models according to the method. Thus, decision 514 includes determining whether the generated suggestions are accepted. In response to determining that the generated suggestions are not accepted, for example, as described above, method 500 proceeds to operation 516, whereby the user proposes one or more suggestions and / or data models. The suggestions and / or data models can be based on machine learning and / or deep learning algorithms running on the file system of the cluster from which the original suggestions were generated.
[0065] However, returning to decision 514, method 500 jumps to operation 518 in response to determining that the generated suggestions are accepted. However, it should be noted that operation 516 can be performed for some methods in which the generated suggestions are accepted. For example, the user can provide one or more suggestions and / or data models to supplement the generated suggestions.
[0066] As shown, operation 518 includes sending a reply to node 502 that indicates whether any of the generated suggestions have been accepted. The reply can also include one or more suggestions and / or data models as described above. In response to receiving the reply from node 503, operation 520 is performed. There, operation 520 includes using one or more approved suggestions and / or data models to transfer at least some of the data in the storage device between the first layer and the second layer. In other words, operation 520 includes applying the suggestions and / or data models to manage the data included in the file system of the cluster. In some methods, for example, as will be described in further detail below (e.g., see Figure 5C ), the suggestions and / or data models can also be applied to new data when new data is received.
[0067] Now referring to Figure 5B , an example sub - process of applying suggestions and / or data models to manage data included in a pre - populated, clustered file system according to one embodiment is shown, one or more of which can be used to perform Figure 5A 's operation 520. However, it should be noted that Figure 5B 's sub - process is shown according to one embodiment, which is in no way intended to limit the present invention. For example, although Figure 5A indicates that Figure 5B 's included example sub - process is executed by one or more processors at node 502, any one or more of the sub - processes can actually be executed by any processor among other processors in the file system of the cluster. Thus, Figure 5B 's included sub - processing, any one or more of which can be performed by the above Figure 4One or more of the controllers 412, 420 in
[0068] Continuing to refer to Figure 5B , sub-operation 540 includes receiving one or more suggestions and / or data models corresponding to the placement of data in the storage devices of the file system of the cluster. As described above, the file system of the cluster includes a memory having a first layer and a second layer, where the first layer has two or more shared nodes and the second layer has at least one shared-nothing node. The one or more suggestions and / or data models are also based on data workload characteristics. Although preferably the data workload characteristics are derived from the storage environment in which the suggestions and / or data models are applied, in some methods, other information may be used to derive one or more of the suggestions and / or data models. For example, one or more data models compiled using machine learning and / or deep learning algorithms executed on a similar cluster file system may be used.
[0069] The one or more suggestions and / or data models received in sub-operation 540 are also used to identify certain portions of the actual data stored in the actual storage devices, and certain portions of the actual data are predicted to benefit from being transferred to a specific layer in the layers of the storage devices. See sub-operation 542. For example, a specific portion of the data may be identified as having information included therein and / or corresponding thereto that would improve the accuracy of existing machine learning and / or deep learning algorithms. Thus, it may be desirable to transfer the identified data portion to one of the shared-nothing nodes in the second layer so that the dedicated hardware (e.g., GPU) included therein can be used to assist in using or at least updating the algorithm based on the data workload characteristics associated with that data portion. According to another example, the suggestions and / or data models may be used to identify a portion of the data that is expected to have an upcoming heavy data transfer workload. Given the upcoming heavy data transfer workload, it is predicted that this data portion will benefit from being stored in the second layer. The prediction is based on both the expected heavy data transfer workload and the configuration of the shared-nothing nodes in the second layer, which preferably includes dedicated hardware capable of achieving increased data transfer rates. According to yet another example, a portion of the data may be identified as being able to provide information valuable to machine learning and / or deep learning algorithms. Thus, for example, as will be appreciated by those skilled in the art after reading this specification, this portion of the data may be determined to have the potential to improve a customized ILM scheme for managing the placement of data in the file system of the cluster.
[0070] For each portion in the portion of the actual data identified in sub - operation 542, an actual determination is made as to whether a given portion of data should be transferred to a different layer in the storage device. See determination 544. According to some methods, performing determination 544 can actually involve determining the specific configuration of different layers in storage and comparing them with predictions made using usage recommendations and / or data models. For example, it can be determined whether any shared - nothing nodes in the second layer actually include dedicated hardware. Additionally, for those nodes determined to have dedicated hardware, a further determination can be made as to what performance level (e.g., such as increased data transfer rate capabilities) the dedicated hardware can achieve for a given node.
[0071] Similar determinations can be made regarding the shared nodes in the first layer. Although the shared nodes may not include added dedicated hardware, the way each of the shared nodes is interconnected across different storage components allows certain operations to be performed in parallel, thus achieving a higher processing rate than a single shared - nothing node might be able to achieve. It follows that, although considering the upcoming data - transfer heavy workload, it can be predicted that some portions of the data would benefit from being stored in the second layer with shared - nothing nodes, given the upcoming data - processing intensive workload, it can be predicted that other portions of the data would benefit from being stored in the first layer, which data - processing intensive workload can thus be performed in parallel by more than one processing component.
[0072] Figure 5B The flowchart of is shown to continue to sub - operation 546 in response to determining that a given portion of data should not be transferred to a different layer in the storage device. There, sub - operation 546 includes using a default ILM scheme to manage the placement of the given portion of data in the file system of the cluster. As described above, the ILM scheme automates the management processes involved in data storage, typically organizes data according to specified policies, and automates the migration of data from one layer to another based on these criteria. For example, newer data and / or more frequently accessed data are preferably stored on higher - performance media, while less - critical data is stored on lower - performance media. In some methods, the user can also specify specific storage policies in certain ILM schemes.
[0073] Returning to decision 544, in response to determining that a given portion of data should be transferred to a different tier in the storage device, the flow diagram proceeds to sub-operation 548. There, sub-operation 548 includes preparing source and destination tier information associated with performing the transfer. According to one example, transferring a portion of data stored in a first tier to a second tier involves preparing information identifying where in storage the data portion is currently stored (e.g., the logical and / or physical addresses where the data portion is striped across), the total size of the data portion, the metadata associated with the data portion, where the data portion will be stored in the second tier (e.g., the logical and / or physical address), and so on.
[0074] Additionally, sub-operation 550 includes sending one or more instructions to transfer (e.g., migrate) a given portion of data from the source tier to the destination tier. The one or more instructions can include the logical address of the expected storage location in the destination tier, or any other information that will be apparent to those skilled in the art after reading this specification and that is associated with actually performing the data transfer.
[0075] It should also be noted that the data portion is also preferably transferred back to the tier in which it was previously stored. This is particularly applicable to data portions transferred from the first tier to the second tier. Again, the dedicated hardware included in the second tier's shared-nothing nodes provides improved performance and is thus preferably reserved for the relevant data. According to an example, preferably, a shared-nothing node in the second tier that is converted to update a portion of data running a machine learning and / or deep learning algorithm is subsequently converted back to the first tier, thereby allowing the dedicated hardware to be used for additional processing.
[0076] Again, these sub-processes are performed for each of the portions of data identified in sub-operation 542. Thus, Figure 5B any one or more of the sub-processes included in can be repeated any number of times, e.g., depending on how many data portions are identified.
[0077] As described above, the recommendations and / or data models formed using some of the embodiments included herein are also preferably used to manage the process of storing newly received data in the storage device. For example, the recommendations and / or data models can be applied to data received from a user, a running application, another file system, etc., in order to determine where each portion of the incoming data should be stored in the file system of the cluster. Thus, Figure 5C FIG. 570 shows a flow diagram of a method 570 according to one embodiment. In different embodiments, method 570 can be performed according to the present invention in Figures 1-4 any environment depicted, etc. Of course, as those skilled in the art will understand after reading this specification, more or fewer Figure 5C than the operations specifically described in can be included in method 570.
[0078] Each step of method 570 can be performed by any suitable component of the operating environment. For example, in different embodiments, method 570 can be performed partially or fully by a controller, a processor, etc., or some other device having one or more processors therein. A processor (e.g., (multiple) processing circuits, (multiple) chips) and / or (multiple) modules implemented in hardware and / or software and preferably having at least one hardware component can be used in any device to perform one or more steps of method 570. Illustrative processors include, but are not limited to, a central processing unit (CPU), an ASIC, a field programmable gate array (FPGA), etc., combinations thereof, or any other suitable computing device known in the art.
[0079] As Figure 5C shown, operation 572 of method 570 includes receiving new data. Depending on the method, the data can be received continuously as a stream, in one or more packets, etc. Further, operation 574 includes using a suggestion and / or a data model to identify the corresponding portion of the newly received data. According to a particular method, the suggestion and / or the data model is preferably used to identify the portion of the newly received data that is predicted to benefit from being stored in a particular layer in the storage device.
[0080] As described above, the storage of the file system of the cluster in this embodiment has a first layer having two or more shared nodes therein and a second layer having at least one shared-nothing node therein. One or more suggestions and / or data models are also based on data workload characteristics. Thus, while in view of an upcoming data transfer heavy workload, it can be predicted that certain portions of the data benefit from being stored in the second layer having shared-nothing nodes, in view of an upcoming data processing intensive workload, it can be predicted that other portions of the data benefit from being stored in the first layer, and the data processing intensive workload can thus be performed in parallel by more than one processing component. Thus, for example, depending on the method, any one or more of the methods described above regarding performing sub-operation 542 can be implemented to perform operation 572.
[0081] Further, for each portion of the newly identified data in operation 574, an actual determination is made as to whether the given portion of data should actually be stored in a particular layer in the storage device. See decision 576. According to some methods, performing decision 576 can actually involve determining the specific configuration of the different layers in the storage device and comparing them with each of the identified portions of the new data. For example, it can be determined whether any of the shared-nothing nodes in the second layer actually include dedicated hardware. Further, for those nodes determined to have dedicated hardware, a further determination can be made as to what performance the dedicated hardware can achieve for a given node (e.g., such as an increased data transfer rate capability).
[0082] For example, according to any of the methods described herein, similar determinations can also be made regarding the shared nodes in the first tier. For example, although the shared nodes may not include added dedicated hardware, the way each of the shared nodes is interconnected across different storage components allows certain operations to be performed in parallel, thereby achieving a higher processing rate than a single shared-nothing node might be able to achieve. It follows that, although considering the upcoming data transfer heavy workload, certain portions of the data may benefit from being stored in the second tier with shared-nothing nodes, in view of the upcoming data processing intensive workload, other portions of the predictable data benefit from being stored in the first tier, and the data processing intensive workload can thereby be performed in parallel by more than one processing component.
[0083] In response to determining that a given portion of new data will not actually benefit from being stored in a particular tier of the storage device, method 570 proceeds to operation 578. There, operation 578 includes using a default ILM scheme to manage the placement of the given portion of new data in the file system of the cluster. By way of example, which is in no way intended to limit the invention, the default ILM scheme can provide that new data is preferably stored on the first tier by default, with the option of storing certain portions of the new data on the second tier. This allows the dedicated hardware included in the second tier to be reserved for relevant portions of the data, while using the first tier to manage the remainder of the data.
[0084] However, returning to decision 576, in response to determining that a given portion of new data will benefit from being stored in a particular tier of the storage device, the flowchart advances to sub-operation 580. There, sub-operation 580 includes sending one or more instructions to store the given identified portion of the newly received data in a particular one of the tiers. Again, the one or more instructions can include the logical address of the intended storage location of the data portion, or any other information that will be apparent to those skilled in the art after reading this specification.
[0085] According to one example, a particular portion of the data can be identified as having one or more particular characteristics that can be used to update a running machine learning and / or deep learning algorithm. The one or more instructions can thereby cause the identified data portion to be stored in the second tier, for example, such that the dedicated hardware included therein can be used to process the data based on a given method. According to another example, it is determined that a portion of the new data predicted to have an upcoming data transfer heavy workload will benefit from being stored in the second tier. Specifically, the dedicated hardware in the second tier can provide an increased level of performance that complements the expected upcoming data transfer heavy workload. The portion of the new data is thereby preferably stored in the second tier.
[0086] However, it should also be noted that the data portion stored in the second layer may ultimately be transferred to the first layer. Again, the dedicated hardware included in the shared-nothing nodes of the second layer provides improved performance and is thus preferably reserved for relevant data. For example, preferably, a portion of the data for updating a running machine learning and / or deep learning algorithm at the shared-nothing nodes of the second layer is subsequently converted back to the first layer, thereby allowing the dedicated hardware to be used for additional processing.
[0087] Again, these sub-processes are performed for each portion of the portion of new data identified in operation 574. Thus, Figure 5C any one or more of the processes included therein may be repeated any number of times, e.g., depending on how many data portions are identified and / or the amount of new data received.
[0088] It follows that different embodiments in the embodiments included herein are capable of implementing a framework that can support both striped and non-striped storage configurations in a single (same) namespace. Additionally, some embodiments herein are capable of significantly improving the accuracy with which machine learning and / or deep learning algorithms represent the internal workings of a given file system. In turn, these improved machine learning and / or deep learning algorithms are capable of enhancing the ILM scheme to appropriately place or migrate data between different layers in the file system based on data workload characteristics and the recommendations and / or data models derived therefrom, where some of the layers utilize dedicated hardware for implementation. Additionally, inode structures for interacting with the layer implementing the striped scheme and the layer implementing the non-striped scheme are maintained separately. Thus, replication and reliability are configured and maintained separately.
[0089] Some embodiments included herein are also capable of generating and presenting recommendations and / or data models based on an analysis of data workload characteristics. The analysis preferably employs data type, data volume, read and / or write access patterns, etc. These recommendations and / or data models can thus appropriately migrate data between different layers, some of which are enabled by dedicated hardware. In addition to migration, the recommendations and / or data models can also help develop placement rules for newly created files, newly received data, etc.
[0090] This is particularly desirable compared to the disadvantages experienced by conventional products. For example, conventional file systems face challenges in terms of latency that result from striping data across multiple disks, thereby causing the computing infrastructure to recall all fragments stored across different disks before being able to use the data for actual processing (e.g., via machine learning and / or deep learning algorithms).
[0091] The present invention may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0092] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0093] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a corresponding computing / processing device or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0094] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a LAN or WAN, or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing aspects of the present invention.
[0095] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0096] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, which, when executed by the processor of the computer or other programmable data processing apparatus, creates a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0097] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices that cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of the possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of an instruction, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, depending on the functionality involved, two consecutive blocks shown may actually be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order. It will also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a system based on dedicated hardware that performs the specified functions or actions or executes a combination of dedicated hardware and computer instructions.
[0099] In addition, a system according to various embodiments may include a processor and logic integrated with and / or executable by the processor, the logic being configured to perform one or more of the process steps described herein. The processor may have any configuration as described herein, such as a discrete processor or a processing circuit including many components such as processing hardware, memory, I / O interfaces, etc. By integrated, it means that the processor has logic embedded therein as hardware logic, such as an application-specific integrated circuit (ASIC), FPGA, etc. By processor-executable, it means that the logic is hardware logic; software logic, such as firmware, a part of an operating system, a part of an application; a combination of hardware and software logic that the processor can access and is configured to cause the processor to perform a certain function when executed by the processor, etc. As is known in the art, software logic can be stored on any type of local and / or remote memory. Any processor known in the art may be used, such as a software processor module and / or a hardware processor, such as an ASIC, FPGA, central processing unit (CPU), integrated circuit (IC), graphics processing unit (GPU), etc.
[0100] It will be clear that the different features of the foregoing systems and / or methods can be combined in any manner, thereby creating multiple combinations from the description presented above.
[0101] It will be further recognized that embodiments of the present invention may be provided in the form of customer-deployed services to provide on-demand services.
[0102] Although the different embodiments have been described above, it should be understood that these embodiments are presented by way of example only and not limitation. Thus, the breadth and scope of the preferred embodiments should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the appended claims and their equivalents.
Claims
1. A computer-implemented method, comprising: Receive one or more suggestions corresponding to the placement of data in a storage device, where the one or more suggestions are based on data workload characteristics; Use the one or more suggestions to identify portions of actual data stored in an actual storage device corresponding to the one or more suggestions, where the actual storage device includes: a first layer having two or more shared nodes, and a second layer having at least one non-shared node; For each identified portion of the identified portion of the actual data stored in the first layer, use the one or more suggestions to determine whether to transfer the identified portion of the actual data to the second layer; and In response to determining to transfer at least one identified portion of the identified portion of the actual data to the second layer, send one or more instructions to transfer the at least one identified portion of the identified portion of the actual data from the first layer to the second layer.
2. The computer-implemented method according to claim 1, wherein each of the at least one shared-nothing node in the second layer comprises dedicated hardware.
3. The computer-implemented method according to claim 2, wherein the dedicated hardware is selected from the group consisting of: a graphics processing unit, a solid state drive cache, an application specific integrated circuit, and a fast non-volatile memory.
4. The computer-implemented method according to claim 1, wherein the first layer is configured to stripe data across the two or more shared nodes, provided that the second layer is not configured to stripe data across two or more of the at least one shared-nothing node.
5. The computer-implemented method according to claim 1, wherein the first layer and the second layer are included in the same namespace.
6. The computer-implemented method according to claim 1, comprising: Use the one or more suggestions to identify corresponding portions of newly received data; For each identified portion of the newly received data, use the one or more suggestions to determine whether to store the identified portion of the newly received data in the second layer; In response to determining to store the identified portion of the newly received data in the second layer, send one or more instructions to store the identified portion of the newly received data in the second layer; And In response to determining not to store the identified portion of the newly received data in the second layer, send one or more instructions to store the identified portion of the newly received data in the first layer.
7. The computer-implemented method according to claim 1, wherein the data workload characteristics are generated using information selected from the group consisting of: read and / or write patterns, corresponding file types, and corresponding portions of files.
8. A computer system, comprising means adapted to perform all steps of the method according to any one of claims 1 to 7.
9. A computer program product, comprising instructions for performing all steps of the method according to any one of claims 1 to 7 when the computer program product is executed on a computer system.
Citation Information
Patent Citations
Dynamic load balancing method of distributed storage system
CN103701916A
Intelligence for controlling virtual storage appliance storage allocation
CN103907097A