Systems, methods, and apparatuses for storing query plans

CN114816237BActive Publication Date: 2026-09-29SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210059251.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-07
Filing Date
2022-01-19
Publication Date
2026-09-29
Estimated Expiration
2042-01-19

Smart Images

  • Figure CN114816237B_ABST
    Figure CN114816237B_ABST
Patent Text Reader

Abstract

A method can include receiving a request for storage resources to access a data set for a processing session, assigning one or more storage nodes for the processing session based on the data set, and mapping the one or more storage nodes to one or more compute nodes for the processing session through one or more network paths. The method can also include returning a resource map of the one or more storage nodes and the one or more compute nodes. The method can also include estimating available storage bandwidth for the processing session. The method can also include estimating available client bandwidth. The method can also include allocating bandwidth to a connection between at least one of the one or more storage nodes and at least one of the one or more compute nodes through one of the network paths.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to storage queries, and more specifically to systems, methods, and apparatus for storage query planning. Background Technology

[0002] A data processing session can read datasets that can be stored across multiple storage nodes. Data from different storage nodes can be accessed over the network and processed by different compute nodes.

[0003] The information disclosed in the background section is only for enhancing the understanding of the background of the present invention, and therefore may contain information that does not constitute prior art. Summary of the Invention

[0004] A method may include: receiving a request for storage resources to access a dataset used for processing a session; allocating one or more storage nodes for processing the session based on the dataset; and mapping the one or more storage nodes to one or more compute nodes for processing the session via one or more network paths. The method may also include returning a resource graph of the one or more storage nodes and one or more compute nodes. The resource graph may include the allocated resource graph. The resource graph may include an availability resource graph. The method may also include estimating the available storage bandwidth for processing the session. The method may also include estimating the available client bandwidth. The method may also include allocating bandwidth to a connection via one of the network paths between at least one of the one or more storage nodes and at least one of the one or more compute nodes. The available storage bandwidth for processing the session may be estimated based on baseline data of the one or more storage nodes. The available storage bandwidth for processing the session may be estimated based on historical data of the one or more storage nodes. The method may also include determining the performance of accessing the dataset used for processing the session. Determining the performance of accessing the dataset used for processing the session may include determining the Quality of Service (QoS) for processing the session. Determining the QoS for processing the session may include calculating a QoS probability based on either baseline data or historical data of the one or more storage nodes. The method may also include monitoring the actual performance of the one or more storage nodes used for processing the session. The processing session may include an artificial intelligence training session.

[0005] A system may include: one or more storage nodes configured to store a dataset for processing a session; one or more network paths configured to couple the one or more storage nodes to one or more compute nodes for processing the session; and a storage query manager configured to: receive a request for storage resources to access the dataset for processing the session, allocate at least one of the one or more storage nodes for processing the session based on the request, and map at least one allocated storage node to at least one of the one or more compute nodes for processing the session via at least one of the one or more network paths. The storage query manager may also be configured to allocate bandwidth to connections between at least one of the one or more storage nodes and at least one of the one or more compute nodes via at least one of the one or more network paths. The storage query manager may also be configured to estimate the available storage bandwidth for processing the session, estimate the available client bandwidth for processing the session, and return a resource graph based on the available storage bandwidth and available client bandwidth for processing the session. The storage query manager may also be configured to predict the Quality of Service (QoS) for processing the session.

[0006] A method may include: receiving a request for storage resources for processing a session, wherein the request includes information about a dataset and one or more compute nodes; allocating a storage node based on the dataset; allocating one of the compute nodes; allocating bandwidth for a network connection between the storage node and the allocated compute node; and returning a resource allocation graph for processing the session based on the storage node, the allocated compute node, and the network connection. The storage node may be a first storage node, the allocated compute node may be an allocated first compute node, the bandwidth may be a first bandwidth, and the network connection may be a first network connection. The method may further include allocating a second storage node based on the dataset; allocating a second compute node from among the compute nodes; and allocating a second bandwidth for a second network connection between the second storage node and the allocated second compute node. The resource allocation graph is also based on the second storage node, the allocated second compute node, and the second network connection. Attached Figure Description

[0007] The accompanying drawings are not necessarily drawn to scale, and for illustrative purposes, elements of similar structure or function are generally indicated by similar reference numerals or portions thereof throughout the drawings. The drawings are intended only to facilitate the description of the various embodiments described herein. The drawings do not depict every aspect of the teachings disclosed herein and do not limit the scope of the claims. To prevent obscurity, not all components, connections, etc., may be shown, and not all components may have reference numerals. However, the pattern of component configuration may readily be apparent from the drawings. The drawings, together with the specification, illustrate exemplary embodiments of this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0008] Figure 1 An embodiment of a storage query plan architecture according to an example embodiment of the present disclosure is shown.

[0009] Figure 2 An example embodiment of a storage query plan architecture according to an example embodiment of this disclosure is shown.

[0010] Figure 3 An example embodiment of a computing server according to an example embodiment of the present disclosure is shown.

[0011] Figure 4 An example embodiment of a storage server according to an example embodiment of the present disclosure is shown.

[0012] Figure 5 An example embodiment of a network and QoS management architecture according to an example embodiment of this disclosure is shown.

[0013] Figure 6 An embodiment of a method for initializing a storage query manager according to an example embodiment of the present disclosure is shown.

[0014] Figure 7 An embodiment of a method for operating a storage query manager according to an example embodiment of the present disclosure is shown.

[0015] Figure 8 An example embodiment of an availability resource graph according to an example embodiment of this disclosure is shown.

[0016] Figure 9 An example embodiment of a resource allocation diagram according to an example embodiment of this disclosure is shown.

[0017] Figure 10 An embodiment of a system including storage, networking, and computing resources according to an example embodiment of the present disclosure is shown.

[0018] Figure 11 An embodiment of a method for estimating performance according to an example embodiment of the present disclosure is shown.

[0019] Figure 12 A method according to an example embodiment of this disclosure is shown.

[0020] Figure 13 Another method according to an example embodiment of this disclosure is shown. Detailed Implementation

[0021] Overview

[0022] During processing sessions, such as artificial intelligence (AI) training and / or inference sessions, multiple computing resources can access datasets on the storage system through one or more network paths. If the storage system does not provide deterministic and / or predictable performance, some computing resources may lack data, while others may consume more storage bandwidth than required. This can lead to unpredictable and / or prolonged session completion times and / or underutilized storage resources, especially when running multiple concurrent processing sessions. It can also lead users to over-provision the storage system in an attempt to accelerate processing sessions.

[0023] According to example embodiments of this disclosure, a storage query plan (SQP) can be created for a processing session to allocate resources (such as storage bandwidth) before the session is initiated. Depending on the implementation details, this can enable more efficient use of storage resources and / or predictable and / or consistent runtime for the processing session.

[0024] In some embodiments, a user application may issue a request for resources to enable computing resources at one or more client nodes to access a dataset during a processing session. The request may include information such as the number of computing resources, the bandwidth of each computing resource, and information about the dataset. Based on the information in the request, the storage query manager can create a storage query plan for the processing session by allocating and / or scheduling computing, network, and / or storage resources. The storage query manager may allocate resources, for example, by estimating the total available storage bandwidth for the processing session and / or estimating the total available client bandwidth. In some embodiments, the storage query manager may allocate sufficient storage and / or network bandwidth to satisfy the total available client bandwidth, leaving resources available for other concurrent processing sessions.

[0025] In some embodiments, the storage query manager may predict the quality of service (QoS) used to process sessions based on, for example, baseline and / or historical performance data of the allocated resources, to determine whether the storage query plan is likely to provide sufficient performance.

[0026] In some embodiments, user applications can access the services of the storage query manager through an application programming interface (API). The API may include a set of commands that enable user applications to request resources for processing a session, check the status of a request, schedule resources, release resources, etc.

[0027] In some embodiments, the storage query manager may manage and / or monitor client computing resources, network resources, and / or storage resources during the execution of a processing session, for example, to determine whether pre-allocated resources and / or performance are being provided.

[0028] The principles disclosed herein have independent practicality and can be embodied independently, and not every embodiment can utilize every principle. However, these principles can also be embodied in various combinations, some of which can synergistically amplify the benefits of the individual principles.

[0029] Processing sessions

[0030] While the principles disclosed herein are not limited to any particular application, these techniques may be particularly beneficial in some embodiments when applied to AI training and / or inference sessions. For example, some AI training sessions may use one or more compute units (CUs) located on one or more client nodes, such as training servers, to run training algorithms. During a training session, each compute unit may have access to some or all of the training dataset, which may be distributed across one or more storage resources, such as nodes in a storage cluster.

[0031] According to example embodiments of this disclosure, in the absence of a storage query plan to reserve and / or manage performance across storage resources, individual computing resources may be unable to access relevant portions of the dataset during time slots when they might need those portions. This can be particularly problematic in data-parallel training sessions, where multiple sets of computing units running a parallel training session can share storage resources to access the same dataset. Furthermore, for one or more batch sizes in data-parallel training, gradients from one or more computing resources can be frequently aggregated. This can prevent relevant data from being consistently served to the individual computing resources at a deterministic rate.

[0032] However, some AI training sessions may tend to access data in a read-only manner and / or with predictable access patterns. For example, AI training applications can know before each training session which computing resources can access which parts of the training dataset during a specific time slot. This can facilitate the creation of storage query plans that coordinate resources for multiple concurrent AI training sessions so that they can share the same storage resources in a way that improves storage utilization.

[0033] Storage Query Plan Architecture

[0034] Figure 1 An embodiment of a storage query plan architecture according to an example embodiment of the present disclosure is shown. Figure 1 The illustrated embodiment may include one or more storage nodes 102 and a network 104, the storage nodes 102 being configured to store a dataset for processing sessions, and the network 104 including one or more network paths configured to connect one or more storage nodes 102 to one or more compute nodes 106 so that one or more compute nodes 106 can access at least a portion of the dataset. Figure 1 The illustrated embodiment may further include a storage query manager 108 configured to receive a request 110 for storage resources from application 112, enabling one or more compute nodes 106 to access at least a portion of a dataset used for processing a session. The storage query manager 108 may also be configured to allocate one or more storage nodes for processing a session based on the dataset. The storage query manager 108 may also be configured to map one or more allocated storage nodes 102 to one or more compute nodes 106 used for processing a session via one or more network paths in network 104.

[0035] One or more storage nodes 102 can be implemented using any type and / or configuration of storage resources. For example, in some embodiments, one or more of the storage nodes 102 can be implemented using one or more storage devices, such as hard disk drives (HDDs) that may include magnetic storage media, solid-state drives (SDDs) that may include solid-state storage media (such as NAND flash memory), optical drives, drives based on any type of persistent memory (such as cross-grid non-volatile memory, memory with varying bulk resistance, etc.), and / or any combination thereof. In some embodiments, one or more storage nodes 102 can be implemented using multiple storage devices arranged in, for example, one or more servers, such as server chassis, server racks, server rack groups, data rooms, data centers, edge data centers, mobile edge data centers, etc., and / or any combination thereof. In some embodiments, one or more storage nodes 102 can be implemented using one or more storage server clusters.

[0036] Network 104 can be implemented using any type and / or configured network resources. For example, in some embodiments, network 104 may include any type of network architecture (such as Ethernet, Fibre Channel, InfiniBand, etc.) using any type of network protocol (such as Transmission Control Protocol / Internet Protocol (TCP / IP), Remote Direct Memory Access over Converged Ethernet (RDMA) (RoCE), etc.), any type of storage interface and / or protocol (such as Serial ATA (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), Non-Volatile Memory High Speed ​​(NVMe), Architecture-based NVMe (NVMe-oF), etc.). In some embodiments, network 104 can be implemented using multiple networks and / or segments interconnected with one or more switches, routers, bridges, hubs, etc. Therefore, any part of a network or segment can be configured as one or more local area networks (LANs), wide area networks (WANs), storage area networks (SANs), etc., implemented using any type and / or configured network resources. In some embodiments, some or all of network 104 can be implemented using one or more virtual components (such as virtual LANs (VLANs), virtual WANs (VWANs), etc.).

[0037] One or more computing nodes 106 may be implemented using any type and / or configuration of computing resources. For example, in some embodiments, one or more of computing nodes 106 may be implemented using one or more computing units such as a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), etc. In some embodiments, one or more of computing nodes 106 and / or their computing units may be implemented using combinational logic, sequential logic, one or more timers, counters, registers, state machines, volatile memory (such as dynamic random access memory (DRAM) and / or static random access memory (SRAM)), non-volatile memory (such as flash memory), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), complex instruction set computer (CISC) processors (such as x86 processors), and / or reduced instruction set computer (RISC) processors (such as ARM processors), or the like that which executes instructions stored in any type of memory. In some embodiments, one or more of computing nodes and / or their computing units 106 may be implemented using any combination of the resources described herein. In some embodiments, one or more compute nodes 106 and / or their compute units may be implemented using multiple compute resources arranged in, for example, one or more servers, which may be configured in one or more server chassis, server racks, server rack groups, data rooms, data centers, edge data centers, mobile edge data centers, and / or any combination thereof. In some embodiments, one or more compute nodes 106 may be implemented using one or more compute server clusters.

[0038] The storage query manager 108 can be implemented in hardware, software, or any combination thereof. For example, in some embodiments, the storage query manager 108 can be implemented using combinational logic, sequential logic, one or more timers, counters, registers, state machines, volatile memory (such as DRAM and / or SRAM), non-volatile memory (such as flash memory), CPLD, FPGA, ASIC, CISC processor and / or RISC processor, the like of execution instructions, and GPU, NPU, TPU, etc.

[0039] Depending on the implementation details, Figure 1 The embodiments shown may provide any number of the following features and / or benefits.

[0040] In some embodiments, a cluster of storage nodes can provide deterministic input and / or output operations per second (IOPS) and / or bandwidth, enabling the creation of a storage query plan prior to each processing session to allocate storage bandwidth before the session is initiated. This allows for efficient use of storage resources and / or deterministic runtime for existing and / or new sessions.

[0041] Some embodiments can create storage query plans from a deterministic storage cluster that can provide consistent bandwidth performance for each supplied storage node. In some embodiments, this may involve a storage cluster that can be supplied for read-only performance and client nodes that can provide a specified number of connections, queue depth, and / or I / O (input / output) request size.

[0042] Some embodiments can achieve deterministic data parallelism and / or other processing session completion times. Some embodiments can efficiently utilize a storage cluster by intelligently distributing the load to some or all of the storage nodes in the cluster. This allows multiple concurrent processing sessions to share the same storage resources. Some embodiments can support the scheduling of storage queries, for example, by notifying users and / or applications when resources become available.

[0043] Some embodiments can effectively utilize some or all of the available performance of a storage cluster, which can reduce operating costs. In some embodiments, storage query planning can reduce or eliminate uncertainty and / or unpredictability of one or more processing sessions. This can be particularly effective when multiple concurrent processing sessions may use the same storage cluster. Some embodiments can simplify storage usage, for example, by providing a simplified user interface and / or storage performance management.

[0044] Some embodiments may manage storage cluster performance, for example, by estimating the overall performance capabilities of one or more computing units, storage resources, and / or network components, and by creating a database to manage the allocation and deallocation of bandwidth for one or more storage query plans. Some embodiments may provide API services that simplify the use of the storage cluster by allowing users and / or applications to provide sets of datasets and bandwidth requirements and receive mappings of storage resources to be used for processing sessions.

[0045] In some embodiments, providing a storage query plan to the user and / or application before the processing session enables the user and / or application to decide whether to execute the processing session. For example, a storage query plan that can reserve resources allows the user and / or application to determine whether storage and network resources can provide sufficient performance to successfully execute the processing session and can prevent interference with other currently running sessions. This is particularly beneficial for data-parallel AI training sessions, where operating multiple computing units (e.g., GPUs) may involve consistently accessing data at deterministic rates. Some embodiments can provide resource coordination and / or management that can improve storage and / or network utilization, for example, running multiple AI training sessions on a shared storage cluster.

[0046] Some embodiments may monitor and / or manage overall processing session resources, including client computing units (such as GPUs), network resources, and / or storage resources, during execution to ensure that pre-allocated resources provide performance that can be estimated before a processing session is initiated.

[0047] Figure 2 An example embodiment of a storage query plan architecture according to an example embodiment of this disclosure is shown. Figure 2 The illustrated embodiments can demonstrate Figure 1 Some possible example implementation details of the embodiments shown. Figure 2 The illustrated embodiment may include one or more storage nodes 202, one or more compute nodes 206, and a storage query manager 208. Each of the one or more storage nodes 202 may be implemented, for example, using any type of key-value (KV) storage scheme (such as object storage that can store all or part of one or more datasets as one or more objects). Each of the one or more compute nodes 206 (which may also be referred to as client nodes) may be implemented, for example, as a compute server. Figure 2 In the example shown, the computing server may include one or more computing units 214. If Figure 2 The illustrated embodiment is used for an AI training session, where one or more compute nodes 206 can be referred to as training nodes or training servers. (A more detailed example of compute node 206 can be found in...) Figure 3 As shown in [the image]. A more detailed example of storage node 202 can be found in [the image]. Figure 4 (As shown in the image.)

[0048] Network connections 216, which may represent ports, handles, etc., can be conceptually illustrated to show the connection between computing unit 214 and storage node 202 via network path 218, which may be established, for example, by storage query manager 208 as part of a storage query plan for processing sessions. In some embodiments, the actual connection between computing unit 214 and storage node 202 may be established via network interface controller (NIC) 220 on a corresponding computing server 206. In some embodiments, and depending on implementation details, a relatively large number of network connections 216 and network paths 218 may be configured to provide many-to-many data parallel operations that can be coordinated to achieve efficient use of storage resources and / or deterministic runtime behavior of multiple concurrent processing sessions.

[0049] Storage Query Manager 208 can collect resource information from one or more storage nodes 202 and / or transfer resource allocations to them via storage-side API 222. Storage Query Manager 208 can collect resource information from one or more compute servers 206 and / or transfer resource allocations to them via client-side API 224. APIs 222 and 224 can be implemented using any suitable type of API, such as a representational state transfer (REST) ​​API that facilitates interaction between Storage Query Manager 208 and client-side and storage-side components via network infrastructure. In some embodiments, the REST API can also enable users and / or applications to easily utilize QoS services available from one or more storage nodes 202. In some embodiments, one or more features of the API can be accessed, for example, through a library that can run on one or more compute servers 206 and handle the I / O (input and / or output) of one or more storage servers.

[0050] The storage query manager 208 can use one or more resource databases 226 to maintain information about the quantity, type, capacity, baseline performance data, historical performance data, etc. of resources existing in the system (including storage resources, computing resources, network resources, etc.). The storage query manager 208 can use one or more query configuration databases 228 to maintain information about resource requests, storage query plans, connections, configurations, etc.

[0051] Storage status monitor 230 can monitor and / or log information about the status and / or operation of any storage resources, network resources, computing resources, etc., in the system. For example, storage status monitor 230 can monitor one or more QoS metrics 232 of storage resources to enable storage query manager 208 to evaluate the performance of various components of the system, thereby determining whether the storage query plan is being executed as expected. As another example, storage status monitor 230 can collect historical performance data of various resources that can be used to more accurately estimate the performance of future storage query plans. Storage status monitor 230 can be implemented using any performance QoS monitoring resource, such as Graphite, Collectd, etc.

[0052] Storage state monitor 230, one or more resource databases 226, and one or more query configuration databases 228 can be located in any suitable location within the system. For example, in some embodiments, storage state monitor 230 and databases 226 and 228 can be located on a dedicated server. In some other embodiments, storage state monitor 230 and databases 226 and 228 can be integrated with and / or distributed therein with other components such as storage node 202 and compute server 206.

[0053] Figure 2 The illustrated embodiment may also include a representation of the software architecture 234 of the storage query manager 208. Architecture 234 may include a storage query API 236, which can be used by one or more user applications (e.g., processing applications such as AI training and / or inference applications) to access services provided by the storage status monitor 230. A query planning module 238 may generate a storage query plan in response to a resource request received via the storage query API 236. A query status module 240 may use information, for example, from one or more query configuration databases 228 to monitor the status and / or performance of one or more storage query plans. A storage / computing resource manager module 242 may use information, for example, from one or more resource databases 226 to manage storage and / or computing resources. A network resource manager module 244 may use information, for example, from one or more resource databases 226 to manage network resources.

[0054] Figure 3 An example embodiment of a computing server according to an example embodiment of the present disclosure is shown. Figure 3The illustrated compute server 306 can operate as one or more storage nodes and may include a NIC 320, memory 346, and any number of compute units 314. A data loader 348 (e.g., a data loader component) can be configured to load one or more portions of a dataset for a processing session received via the NIC 320 from one or more storage nodes into memory 346, and then distribute the dataset to the compute units 314 that can use the data. The data loader 348 can operate under the control of an application 350 (e.g., a user application that can implement processing sessions such as AI training and / or inference sessions). Some embodiments may also include a client wrapper 352 that enables the compute server 306 to operate as part of a decomposed storage system, as described below.

[0055] The computing unit 314, NIC 320, and / or data loader 348 may be implemented in hardware, software, or any combination thereof, including combinational logic, sequential logic, one or more timers, counters, registers, state machines, volatile memory (such as DRAM and / or SRAM), non-volatile memory (such as flash memory), CPLD, FPGA, ASIC, CISC processor and / or RISC processor, the like of execution instructions, and GPU, NPU, TPU, etc.

[0056] Figure 4 An example embodiment of a storage server according to an example embodiment of the present disclosure is shown. Figure 4 The storage server 402 shown can be used to implement, for example... Figure 1 One or more of the storage nodes 102 shown Figure 2 One or more of the storage nodes 202 shown. Storage server 402 may include a first dual-port NIC 420, a second dual-port NIC 421, a first storage server 452, a second storage server 454, a storage interface 456, a storage interface subsystem 458, and a KV storage pool 460.

[0057] Storage servers 452 and 454 can be implemented using, for example, any type of key-value (KV) storage scheme. In some embodiments, storage servers 452 and 454 can implement KV storage using an object storage scheme, in which all or part of one or more datasets can be stored as one or more objects. Examples of suitable object storage servers may include MinIO, OpenIO, etc. In some embodiments, a name server or other scheme can be used for item mapping. In some embodiments, storage servers 452 and 454 can be implemented using a key-value management scheme that can be scaled to store very large datasets (e.g., kilobytes (PB) and larger), and / or a name server capable of handling large numbers of item mappings (e.g., implementations that handle write-once, read-many).

[0058] In some embodiments, storage servers 452 and 454 may provide an object storage interface between NICs 420 and 421 and storage interface 456. Storage interface 456 and storage interface subsystem 458 may be implemented, for example, as an NVMe target and NVMe subsystem, respectively, to implement high-speed connectivity, queuing, input and / or output (I / O) requests, etc., for KV storage pool 460. In some embodiments, KV storage pool 460 may be implemented with one or more KV storage devices (e.g., KV SSDs) 462. Alternatively or additionally, KV storage pool 460 may implement conversion functions to convert KV (e.g., object) storage into file- or block-oriented storage devices 462.

[0059] In some embodiments, object storage servers 452 and 454 may implement one or more peer-to-peer (P2P) connections 464 and 466, for example, to enable storage server 402 to communicate directly with other devices (e.g., multiple instances of storage server 402) without consuming storage bandwidth, which would otherwise involve sending P2P traffic via NICs 420 and 421.

[0060] Network and QoS Management Architecture

[0061] Figure 5 An example embodiment of a network and QoS management architecture according to an example embodiment of this disclosure is shown. For example, Figure 5 The architecture shown can be used to implement any storage query planning techniques, QoS monitoring techniques, etc. disclosed herein.

[0062] Figure 5The illustrated embodiment may include N storage servers 502 of a first cluster, configured to be accessed as a cluster via a first cluster switch 568. One or more compute servers 506 may be configured to access the storage servers 502 of the first cluster via a network connection 567 and the first cluster switch 568. In some embodiments, one or more compute servers 506 may, for example, use... Figure 3 The computing server 306 shown is used for implementation, and the storage server 502 can be implemented, for example, using... Figure 4 The storage server 402 shown is used for implementation. Network connections 570 and 572 may be shown only for port 1 of NIC1 and port 1 of NIC2, respectively, to prevent the figures from becoming unclear, but the connections may also be implemented for port 2 of NIC1 and port 2 of NIC2. In some embodiments, one or more additional cluster switches 568 may be included for one or more additional clusters of storage server 502.

[0063] In some embodiments, each cluster switch 568 may further include cluster controller functionality to operate the N storage servers 502 of the first cluster as a storage cluster capable of implementing decomposed storage. Therefore, a client wrapper 552 incorporating the decomposed storage functionality in the respective cluster switch 568 can present the N storage servers 502 of the first cluster as a single storage node configured to improve performance, capacity, reliability, etc., to one or more compute servers 506. In some embodiments, P2P connections 564 and / or 566 enable storage servers 502 to implement data erasure coding across storage servers. Depending on the implementation details, this can be used to enable a more even distribution of the dataset used to process sessions across storage resources, which can improve overall storage resource (e.g., bandwidth) utilization, predictability and / or consistency of data access, etc.

[0064] In some embodiments, Figure 5 The network infrastructure shown can be segmented using one or more virtual network connections (e.g., VLANs), for example, using one or more compute units 514 (which in this example can be implemented as GPUs) and one or more storage servers 502 with one or more QoS metrics. Depending on the implementation details, this can prevent the front end of the network (e.g., the client side) from interfering with the back end of the network (e.g., the storage side).

[0065] In some embodiments, control and data traffic can be sent on separate planes over a network. For example, control commands (e.g., I / O requests) can be sent to storage server 502 on the control plane using, for example, TCP, while data can be sent back to one or more compute servers 506 on the data plane using, for example, RoCE. This allows data to be sent directly from storage server 502 to memory 546 in compute server 506 using RDMA, without any intermediate hops through any memory in storage server 502. Depending on the implementation details, this can improve latency, bandwidth, etc.

[0066] Figure 5 The illustrated embodiment may include a storage state monitor 530, which may monitor and / or collect QoS metrics from one or more compute servers 506, cluster switches 568, and / or storage servers 502 via connections 572, 574, and 576, respectively. Examples of client-side QoS metrics that may be collected for a user (e.g., User[1..N]) may include access / secret keys, bandwidth, IOPS, latency, priority, etc. Examples of client-side QoS metrics that may be collected for an application (e.g., Application[1..N]) may include application entity ID / name, bandwidth, IOPS, latency, priority, etc. Examples of client-side QoS metrics that may be collected for an accessible storage resource (which may be referred to as a bucket) (e.g., Bucket[1..N]) may include bucket name, bandwidth, IOPS, latency, etc.

[0067] Examples of storage-side QoS metrics that can be collected (e.g., cap[1..N]) may include cluster ID, server ingress endpoint, physical node name, unique user ID (UUID), bucket hosting and / or bucket name, bandwidth, IOPS, small IO cost (e.g., latency), large IO cost (e.g., latency), degraded small IO cost (e.g., latency), degraded large IO cost (e.g., latency), heuristic latency of small IO, heuristic latency of large IO, whether it is running at degraded performance, whether it is overloaded, whether the storage node or cluster is down, etc.

[0068] In some embodiments, the internal bandwidth of the storage server 502 (e.g., an NVMe-oF component) can be greater than that of the network, allowing QoS to be established primarily by suppressing data transfer rates from the storage server side of the system.

[0069] Storage Query Manager Initialization

[0070] Figure 6 An embodiment of a method for initializing a storage query manager according to an example embodiment of the present disclosure is shown. Figure 6 The illustrated embodiments can be used in Figure 5 The architecture shown is described in the context of the diagram, but Figure 6 The illustrated embodiments can also be implemented, for example, using any of the systems, methods, and / or apparatuses disclosed herein. In some embodiments, the goal of initializing the storage query manager may be to estimate the available storage bandwidth for processing sessions. Depending on the implementation details, this can enable the creation of a storage query plan that improves or maximizes the storage utilization of the session.

[0071] Figure 6 The method shown can begin at operation 602. At operation 604, the storage query manager can retrieve the performance bandwidth capabilities of one or more storage nodes. In some embodiments, the storage nodes may be implemented, for example, as a cluster of one or more storage servers. At operation 606, the storage query manager can retrieve one or more network topologies and / or bandwidth capabilities of network resources available for storage query planning. At operation 608, the storage query manager can retrieve the bandwidth capabilities of one or more compute servers. At operation 610, the storage query manager can load baseline data to serve as a performance baseline for one or more of the storage nodes, network resources, and / or compute servers. At operation 612, the storage query manager can retrieve information about one or more compute resources (such as compute units that may exist in one or more compute servers). In some embodiments, any or all of the data retrieved and / or loaded in operations 604, 606, 608, 610, and / or 612 may be, for example, from sources such as... Figure 2 The resource database 226 shown is obtained from the database.

[0072] In operation 614, the storage query manager may analyze and / or calculate the bandwidth available for one or more computing units of the computing server based on, for example, the topology and / or resources of the storage server, network, and / or compute server. In some embodiments, benchmark data may be used as a benchmark reference for the maximum performance capabilities of any or all of the resources used in operation 614. Benchmark data may be generated, for example, by running one or more tests on actual resources such as storage, network, and / or compute resources.

[0073] In operation 616, the storage query manager can then be prepared to process storage query requests from users and / or applications.

[0074] In some embodiments, Figure 6 One or more of the operations in the illustrated embodiments may be based on one or more assumptions about one or more storage servers, network capabilities and / or topology, computing servers, etc., for example, as referenced Figure 10 The subject of discussion.

[0075] Storage Query Manager Operations

[0076] Figure 7 An embodiment of a method for operating a storage query manager according to an example embodiment of the present disclosure is shown. Figure 7 The illustrated embodiments can be used in Figure 5 The architecture shown is described in the context of the diagram, but Figure 7 The embodiments shown can also be implemented, for example, using any of the systems, methods and / or apparatuses disclosed herein.

[0077] exist Figure 7 In the illustrated embodiments, the Storage Query Manager service can be accessed by users and / or applications via an API, as described above. Table 1 lists some exemplary commands that can be used to implement the API. The PALLOC command can be used to determine which resources are available to process a session. The PSELECT command can be used to select a resource, which may have been determined in response to the PALLOC command. The PSCHED command can be used to schedule a session using the selected resource. In some embodiments, the PSCHED command can initiate a session immediately if the resource is currently available, or once the resource becomes available. Additionally or alternatively, the Storage Query Manager can notify users and / or applications when a resource becomes available to determine whether the users and / or applications still want to continue the session. The PSTAT command can be used by users and / or applications to check (e.g., periodically) whether the session is running correctly. PREL can be used, for example, to release resources after a session completes. The PABORT command can be used to abort a session.

[0078] Table 1

[0079]

[0080] refer to Figure 7The method can begin at operation 702. At operation 704, the storage query manager can receive requests for resources from users and / or applications. In some embodiments, a request may include one or more of the request records shown in Table 2. The request ID of the record may provide an ID for a new request. The dataset of the record [] may include a list of objects, which may contain datasets for sessions that have submitted requests for them. In some embodiments, the list of objects may include some or all of the datasets to be loaded for the processing session. Additionally or alternatively, the dataset of the record [] may include one or more selection criteria that the storage query manager can use to select data for the dataset to be used in the session. For example, data may be selected from a pre-existing database or other potential data sources to be included in the dataset. The record for the number of compute units may indicate the total number of compute units that the user and / or application can use during the session. The compute units may be located, for example, in one or more compute servers. The record for NIC / CU bandwidth [index][BW] may include a list of specific bandwidths (BW) requested for each compute unit associated with each NIC, which may be identified by an index or the bandwidth may be automatically selected. In some embodiments, for data-parallel operations, the same amount of bandwidth may be requested for each computing unit, while for other types of operations, the bandwidth requested for each computing unit may be determined individually.

[0081] Table 2

[0082]

[0083] Refer again Figure 7 In operation 706, the storage query manager can analyze the request and create a storage query plan. Operation 706 may include sub-operations 706A through 706F. In sub-operation 706A, the storage query manager can locate and map dataset objects across one or more storage nodes (e.g., across one or more storage servers in one or more storage clusters). In sub-operation 706B, the storage query manager can allocate one or more storage nodes based on the location of the data in the dataset used for the session. In sub-operation 706C, the storage query manager can allocate compute units and / or corresponding NIC ports specified in the request.

[0084] In sub-operation 706D, the storage query manager can map one or more allocated storage nodes to one or more compute units via the NIC port and / or network path between the storage nodes and compute units. The storage query manager can then verify that the NIC port and network path can provide sufficient bandwidth to the compute units and the allocated storage nodes to support the requested bandwidth. In embodiments where the dataset is erase-coded across more than one storage node, verification can be applied to all nodes from which erase-coded data can be sent. In sub-operation 706E, the storage query manager can calculate the QoS probability of a storage query plan for a session based on, for example, one or more performance benchmarks and / or historical data. In sub-operation 706F, the storage query manager can return a storage query plan that maps one or more of the compute units identified in the request to one or more storage nodes using the allocated storage and / or network bandwidth that can satisfy one or more bandwidths requested in the request.

[0085] Table 3 shows the data structure with object examples. Figure 7 The illustrated embodiment can use this data structure to track the available bandwidth of one or more storage nodes, compute nodes, network paths, etc., as well as performance benchmark reference data and / or historical performance data.

[0086] Table 3

[0087]

[0088] In operation 708, if the storage query manager is unable to allocate resources to satisfy the request (e.g., the storage query plan fails), the storage query manager may generate an availability resource graph outlining the resources that may be available. In operation 710, if the storage query manager is able to allocate resources to satisfy the request (e.g., the storage query plan succeeds), it may return the allocated resource graph, which may specify the resources that the user and / or application can use during the processing session. Alternatively, if the storage query manager is unable to allocate resources to satisfy the request (e.g., the storage query plan fails), it may return an availability resource graph to inform the user and / or application of the resources that may be available. The user and / or application can use the availability resource graph, for example, to publish a revision request that the storage query manager may be able to satisfy. In operation 712, the storage query manager may wait for new requests and then return to operation 704.

[0089] Figure 8An example embodiment of an availability resource diagram according to an example embodiment of this disclosure is shown. Left side 800 may indicate the maximum bandwidth and allocated bandwidth of a combination of client nodes and ports. Right side 801 may indicate the maximum bandwidth and allocated bandwidth of a combination of storage nodes and ports. The maximum bandwidth can be determined, for example, by consulting device specifications, obtaining baseline data based on benchmark operations, or using historical performance data of the application device. The difference between the maximum bandwidth and the allocated bandwidth can represent the available bandwidth for each node / port combination. For example, a combination of client node CN1 and port P1 may have a maximum bandwidth of 100Gb / s, of which 10Gb / s can be allocated. Therefore, CN1 and P1 may have an available bandwidth of 90Gb / s.

[0090] Figure 9 An example embodiment of a resource allocation diagram according to an example embodiment of this disclosure is shown. For each row of the diagram, the three columns on the left may indicate compute nodes, network ports, and storage nodes that have been mapped to compute nodes. The rightmost column may indicate the amount of bandwidth that has been allocated to the combination of compute nodes and storage nodes. For example, 2Gb / s of bandwidth may have been allocated to the combination of client node CN1, port P1, and storage node SN5.

[0091] Figure 9 The embodiment of the resource allocation diagram 900 shown can be used, for example, to return resources to a user and / or application in response to a request for resources. As another example, it can be used for comparison with telemetry data 903, which may include one or more QoS metrics. The storage query manager can use the telemetry data 903 during execution to monitor and manage overall processing session resources, including compute (client) nodes, network resources, and / or storage nodes, to ensure that pre-allocated resources and performance are being provided. If actual performance does not meet the QoS estimated during storage query planning, the storage query manager can adjust system operation, for example, by reallocating one or more resources, including one or more storage nodes, network paths, etc.

[0092] Performance estimation

[0093] Figure 10 An embodiment of a system including storage, networking, and computing resources according to an example embodiment of the present disclosure is shown. Figure 10 The embodiments shown can be used to illustrate background information related to how bandwidth can be determined according to exemplary embodiments of this disclosure. For illustrative purposes, Figure 10 Some specific implementation details are shown in the document, but the principles of this disclosure are not limited to these details.

[0094] Figure 10The system shown may include a pool of one or more compute nodes 1006, one or more storage nodes 1002, and storage media 1062. Storage nodes 1002 may be grouped into one or more clusters, racks, etc. Figure 10 In the example shown, storage node 1002 can be divided into clusters 1 to N.

[0095] Network 1004 can be implemented using any type and / or configured network resource, wherein one or more network paths can be allocated between the allocated storage nodes and the allocated compute nodes. The pool of storage medium 1062 can be implemented using any storage medium such as flash memory, which can be interfaced with one or more of the storage nodes 1002 via a storage interconnect such as NVMe-oF and / or network 1080.

[0096] The maximum bandwidth (BW) of each compute (client) node can be indicated by BW=Ci, and the allocated bandwidth of each compute node can be indicated by BW=CAi, where i=0…n, and n can represent the number of compute nodes.

[0097] The maximum bandwidth of each storage node can be indicated by BW=Si, the maximum bandwidth of each corresponding NIC can be indicated by BW=Ni, and the allocated bandwidth of the combination of storage nodes and corresponding NICs can be indicated by BW=Ai, where i=0…n, and n can represent the number of storage nodes.

[0098] In some embodiments, it is possible to Figure 10 The components shown make one or more assumptions to simplify and / or facilitate, for example Figure 11 The performance estimates shown are as follows. For example, each compute node 1006 may have one or more corresponding NICs 1020, and in this example, one port may be allocated per compute unit in compute node 1006. Each storage node 1002 may have one or more corresponding NICs 1021, and each NIC may have one or more ports. The pooling and interconnect / network architecture of storage medium 1062 can provide consistent bandwidth support, which can be operated at any network speed that can be allocated to storage node 1002. Network 1004 can provide full point-to-point bandwidth support for any bandwidth that can be allocated between storage node 1002 and compute node 1006. Multiple concurrent processing sessions may be executed on the same server and / or server set. However, these assumptions are not required, and performance estimates can be implemented without these assumptions or any assumptions at all.

[0099] Figure 11 An embodiment of a method for estimating performance according to an example embodiment of the present disclosure is shown. Figure 11The embodiments shown can be based on, for example... Figure 10 The system shown.

[0100] refer to Figure 11 Applications A1, A2, and A3, which can run on each of compute nodes 1106, can access the corresponding portions of the dataset in storage node 1102 (shown in similar shades).

[0101] Aggregated client bandwidth (ACB) can be determined by summing Ci over all client nodes, as shown below:

[0102] Equation 1

[0103] Ci can indicate the maximum client node bandwidth (MCB), and n can be the number of clients.

[0104] Total client allocated bandwidth (TCAB) can be determined by summing the CAi of all client nodes, as shown below:

[0105] Equation 2

[0106] Here, CAi can indicate the allocated bandwidth for each client, and n can be the number of clients.

[0107] To determine the total storage bandwidth (TSB), the lower of the node's storage bandwidth (S) and the node's network bandwidth (N) can be used. This is because, for example, if the storage node's NIC has a lower maximum bandwidth, the entire maximum bandwidth of the storage node may be unavailable. Therefore, the total storage bandwidth (TSB) can be determined as follows:

[0108] Equation 3

[0109] Where N can indicate the maximum NIC bandwidth, S can indicate the maximum bandwidth of the corresponding storage node, and n can be the number of storage nodes.

[0110] The total allocated bandwidth (TAB) can be determined by summing the Ai values ​​of all storage nodes, as shown below:

[0111] Equation 4

[0112] Here, Ai can indicate the allocated bandwidth for each node, and n can be the number of storage nodes.

[0113] Then, the total available storage bandwidth (TASB) can be determined by TASB = TSB - TAB. Then, the total available client bandwidth (TACB) can be determined by TACB = ACB - TCAB.

[0114] Figure 12 A method according to an example embodiment of this disclosure is illustrated. The method may begin at operation 1202. At operation 1204, the method may receive a request for storage resources to access a dataset used for processing a session. At operation 1206, the method may allocate one or more storage nodes for processing the session based on the dataset. At operation 1208, the method may map one or more of the storage nodes to one or more compute nodes used for processing the session via one or more network paths. The method may end at operation 1210.

[0115] Figure 13 Another method according to an example embodiment of this disclosure is illustrated. The method may begin at operation 1302. At operation 1304, the method may receive a request for storage resources for processing a session, wherein the request includes information about a dataset and one or more compute nodes. At operation 1306, the method may allocate a storage node based on the dataset. At operation 1308, the method may allocate one of the compute nodes. At operation 1310, the method may allocate bandwidth for a network connection between the storage node and the allocated compute node. At operation 1312, the method may return a resource allocation graph for processing the session based on the storage node, the allocated compute node, and the network connection. The method may end at operation 1314.

[0116] Figure 12 and Figure 13 The embodiments shown herein, as well as other embodiments described herein, are all example operations and / or components. In some embodiments, some operations and / or components may be omitted, and / or other operations and / or components may be included. Furthermore, in some embodiments, the temporal and / or spatial order of operations and / or components may vary. Although some components and / or operations may be shown as separate components, in some embodiments, some components and / or operations shown separately may be integrated into a single component and / or operation, and / or some components and / or operations shown as single components and / or operations may be implemented with multiple components and / or operations.

[0117] The embodiments disclosed above have been described in the context of various implementation details, but the principles of this disclosure are not limited to these or any other specific details. For example, some functions have been described as being implemented by certain components, but in other embodiments, functions may be distributed among different systems and components in different locations and with various user interfaces. Some embodiments have been described as having specific processes, operations, etc., but these terms also include embodiments in which specific processes, operations, etc. may be implemented by multiple processes, operations, etc., or in which multiple processes, operations, etc. may be integrated into a single process, step, etc. References to components or elements may refer to only a portion of a component or element. For example, a reference to an integrated circuit may refer to all or only a portion of an integrated circuit, and a reference to a block may refer to the entire block or one or more sub-blocks. The use of terms such as “first” and “second” in this disclosure and claims may be merely for distinguishing what they modify and may not indicate any spatial or temporal order unless otherwise apparent from the context. In some embodiments, a reference to something may refer to at least a portion of something; for example, “based on” may mean “at least partially based on,” “access” may mean “at least partially access,” etc. A reference to a first element does not imply the existence of a second element. For convenience, various organizational aids may be provided, such as chapter titles, but the themes arranged according to these aids and the principles of this disclosure are not limited to these organizational aids.

[0118] Based on the inventive principles disclosed in this patent, the various details and embodiments described above can be combined to produce additional embodiments. Since the inventive principles disclosed in this patent can be modified in arrangement and detail without departing from the concept of the invention, such changes and modifications are considered to fall within the scope of the following claims.

Claims

1. A method for storing query plans, comprising: Receive a request for storage resources to access a training dataset of erase codes distributed across storage nodes for processing sessions, wherein the request includes bandwidth; Create a storage query plan, including: As part of the storage query plan, objects are located in the erased-encoded training dataset in the first set of storage nodes; A resource graph that maps the objects to storage nodes and one or more compute nodes; Based on the erased encoding training dataset, the location of the object, and the resource graph, a second set of storage nodes is allocated for processing the session; Partly based on the bandwidth, a second set of storage nodes is mapped to one or more compute nodes for the processing session via one or more network paths as part of a storage query plan, wherein the processing session includes a training dataset based on erasure encoding and a parallel training session allocating the second set of storage nodes; and Return the storage query plan.

2. The method of claim 1, further comprising returning a resource graph of the storage node and the one or more computing nodes.

3. The method according to claim 2, wherein, The resource map includes the allocated resource map.

4. The method according to claim 2, wherein, The resource map includes an availability resource map.

5. The method of claim 1, further comprising estimating the available storage bandwidth for the processing session.

6. The method of claim 5 further includes estimating available client bandwidth.

7. The method of claim 1, further comprising allocating bandwidth to a connection between at least one of the storage nodes and at least one of the one or more computing nodes via one of one or more network paths.

8. The method according to claim 5, wherein, The available storage bandwidth for the processing session is estimated based on the baseline data of the storage node.

9. The method according to claim 5, wherein, The available storage bandwidth for the processing session is estimated based on historical data from the storage node.

10. The method of claim 1, further comprising determining the performance of accessing the training dataset of the erasure encoding used for the processing session.

11. The method according to claim 10, wherein, Determining the performance of accessing the training dataset of the erasure encoding used for the processing session includes determining the Quality of Service (QoS) for the processing session.

12. The method according to claim 11, wherein, Determining the QoS for the processing session includes calculating the QoS probability based on either baseline data or historical data from the storage node.

13. The method of claim 1, further comprising monitoring the actual performance of the storage node used for the processing session.

14. The method according to claim 1, wherein, The storage query plan coordinates parallel training sessions of the training dataset with shared erasure codes.

15. A system for storing query plans, comprising: The first set of storage nodes is configured to store the erase encoding training dataset distributed across the storage nodes for processing sessions; One or more network paths are configured to couple the storage node to one or more compute nodes for the processing session; as well as The storage query manager is configured to create storage query plans using the following steps: Receive a request for storage resources to access a training dataset for erasure encoding used in the processing session, wherein the request includes bandwidth; Locate the object in the erased-encoded training dataset within the second set of storage nodes; A resource graph that maps the objects to storage nodes and one or more compute nodes; Based on the request, the location of the object, and the resource graph, allocate at least one of the storage nodes for the processing session; Partially based on the bandwidth, at least one storage node is mapped to at least one of the one or more compute nodes for the processing session via at least one of the one or more network paths, wherein the processing session includes a training dataset based on erasure encoding and a parallel training session allocating at least one storage node; and Return the storage query plan.

16. The system according to claim 15, wherein, The storage query manager is also configured to allocate bandwidth to the connection between at least one of the storage nodes and at least one of the one or more compute nodes via at least one of the one or more network paths.

17. The system according to claim 15, wherein, The storage query manager is also configured to: Estimate the available storage bandwidth for the processing session; Estimate the available client bandwidth for the processing session; and Return a resource graph based on the available storage bandwidth and available client bandwidth used for the processing session.

18. The system according to claim 15, wherein, The storage query manager is also configured to predict the Quality of Service (QoS) for the processing session.

19. A method for storing query plans, comprising: Receive a request for a first set of storage resources for processing a session, wherein the request includes bandwidth, a training dataset of erase codes distributed across the storage resources, and information about one or more computing nodes; Create a storage query plan, including: Locate the object in the erased-encoded training dataset within the second set of storage resources; Map the object to a resource graph of storage resources and one or more compute nodes; Based on the training dataset of the erased encoding, the location of the object, and the resource graph, storage nodes from the first set of storage resources are allocated as part of the storage query plan; Allocating one of the one or more compute nodes as part of a storage query plan; allocating bandwidth for the network connection between the storage node and the allocated compute node as part of the storage query plan; and Based on the storage node, the bandwidth, the allocated computing node, and the network connection, a resource allocation graph and storage query plan for the processing session are returned, wherein the processing session includes a parallel training session based on an erase-coded training dataset and allocated storage nodes.

20. The method according to claim 19, wherein, The storage node includes a first storage node, the allocated computing node includes an allocated first computing node, the bandwidth includes a first bandwidth, and the network connection includes a first network connection. The method further includes: A second storage node is allocated based on the training dataset containing the erasure code; A second computing node is allocated to the one or more computing nodes; and Allocate a second bandwidth for the second network connection between the second storage node and the second computing node; and The resource allocation graph is also based on the second storage node, the second computing node, and the second network connection.

Citation Information

Patent Citations

  • NON-VOLATILE MEMORY EXPRESS OVER FABRIC (NVMeOF) USING VOLUME MANAGEMENT DEVICE

    US20180337991A1