Methods, apparatuses, devices, media, and products for determining a node
Patent Information
- Application Number
- CN202510322652.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0029]可以理解地,上述提供的第二方面的装置、第三方面的电子设备、第四方面的计算机存储介质或第五方面的计算机程序产品用于执行第一方面所提供的方法。因此,关于第一方面的解释或者说明同样适用于第二方面、第三方面、第四方面和第五方面。此外,第二方面、第三方面、第四方面和第五方面所能达到的有益效果可参考对应方法中的有益效果,此处不再赘述。
Smart Images

Figure CN122824752A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application primarily relate to the field of node management. More specifically, the embodiments of this application relate to methods, apparatus, devices, media, and products for determining nodes. Background Technology
[0002] A distributed storage system is a computer system that distributes data across multiple nodes. This system connects multiple storage nodes via a network and employs technologies such as data redundancy, distributed file systems, object storage, and block storage. It utilizes methods like replication and erasure coding to ensure data reliability and employs consistent hashing and distributed hash tables to achieve balanced data distribution. It boasts advantages such as high scalability, high availability, high performance, and cost-effectiveness. It is now widely used in cloud computing, big data, artificial intelligence, and the Internet of Things, meeting the needs of large-scale data storage and processing and providing stable, efficient, and reliable data storage support for various applications.
[0003] In a distributed storage system, storage nodes typically include a Central Processing Unit (CPU), Random Access Memory (RAM), and a hard disk. Distributed storage systems allow for node expansion, which simultaneously scales performance and capacity, resulting in strong scalability. Furthermore, each node participates in data management and storage, distributing system load relatively evenly, and cross-node erasure coding (EC) can be used for data protection. Nodes in this storage system are interconnected via a network. However, many problems remain to be solved in the application of distributed storage systems. Summary of the Invention
[0004] The embodiments of this application provide a scheme for determining nodes.
[0005] According to a first aspect of this application, a method for determining a node is provided. The method includes receiving an access request for target data from a compute node in one of a plurality of clusters; determining topology information for a plurality of compute nodes and a plurality of storage nodes in the plurality of clusters and transmission performance requirements for the access request; and, based on the topology information and transmission performance requirements, determining a target storage node among the plurality of storage nodes that can be used for the access request to process the target data.
[0006] By utilizing the topology information of storage nodes across multiple clusters, this method can quickly determine the optimal target storage node from the cluster nodes based on the access request requirements. This improves the accuracy of node selection, speeds up data access, and enhances the user experience.
[0007] In some embodiments, determining the topology information for multiple compute nodes and multiple storage nodes in multiple clusters and the transmission performance requirements for access requests includes: collecting topology information for the multiple compute nodes and multiple storage nodes, the topology information including at least one of the following: the relationships between nodes, interconnect bandwidth, and interconnect latency; and determining the transmission performance requirements for access requests based on the type of access request and the amount of target data. This method allows for the rapid acquisition of topology information and accurate determination of the performance requirements for access requests.
[0008] In some embodiments, determining a target storage node among multiple storage nodes that can be used for access requests to process target data includes: determining multiple performance information items for the multiple storage nodes in the topology information; and determining the target storage node from the multiple storage nodes based on the multiple performance information items and transmission performance requirements. In this manner, a suitable target storage node is determined through performance comparison, enabling the determined storage node to achieve rapid processing of access requests.
[0009] In some embodiments, determining a target storage node from multiple storage nodes based on multiple performance information items and transmission performance requirements includes: determining a set of available storage nodes from the multiple storage nodes by matching the multiple performance information items and transmission performance requirements; and selecting a target storage node from the set of available storage nodes. This method determines available storage nodes through performance comparison, which in turn allows for the further determination of a preferred target storage node, resulting in more balanced service processing and faster request processing.
[0010] In some embodiments, the access request is a write request, and the method further includes: determining whether the target storage node is located in the cluster; if the target storage node is determined to be located in the cluster, causing the compute node to write the target data to the target storage node via the intra-cluster network. This approach allows the target data to be written to the storage node using the intra-cluster network when the target storage node is within the cluster, thus improving storage efficiency.
[0011] In some embodiments, the cluster is a first cluster, and the method further includes: if it is determined that the target storage node is not located in the first cluster, determining a first auxiliary storage node in the first cluster; causing a computing node to write the target data to the first auxiliary storage node in the first cluster; and using a network for multiple storage nodes, writing the target data from the first auxiliary storage node to a target storage node in a second cluster of multiple clusters. This method utilizes the network between storage nodes, enabling the target storage node to quickly write data to storage nodes even when it is in another cluster, thus improving storage efficiency.
[0012] In some embodiments, the access request is a read request, and the method includes: determining whether a target storage node is located in a cluster; if the target storage node is not located in a cluster, determining a second auxiliary storage node in the cluster; and causing the second auxiliary storage node to pull target data from the target storage node in a third cluster of multiple clusters using a network for multiple storage nodes; and providing the target data from the second auxiliary storage node to a compute node. This approach allows data to be pulled from storage nodes in different clusters using a network between storage nodes when executing a read request, improving data retrieval efficiency.
[0013] In some embodiments, the method further includes: if it is determined that the target storage node is located in a cluster, providing the target data from the target storage node to the compute node via the intra-cluster network. This method allows data to be retrieved from storage nodes within the same cluster using the intra-cluster network when executing a read request, improving data retrieval efficiency.
[0014] In some embodiments, multiple clusters are used to process tasks related to machine learning models. This approach can accelerate the processing of machine learning model tasks.
[0015] In some embodiments, the target data may be checkpoint file data or training data for a machine learning model. This approach can accelerate the inference and training of machine learning models.
[0016] According to a second aspect of this application, an apparatus for determining a node is provided. The apparatus includes: an access request receiving unit configured to receive an access request for target data from a compute node in one of a plurality of clusters; a topology and performance determination unit configured to determine topology information for a plurality of compute nodes and a plurality of storage nodes in the plurality of clusters and transmission performance requirements for the access request; and a target storage node determination unit configured to determine, based on the topology information and transmission performance requirements, a target storage node among the plurality of storage nodes that can be used for the access request to process the target data.
[0017] In some embodiments, the topology and performance determination unit includes: a topology information collection unit configured to collect topology information for multiple computing nodes and multiple storage nodes, the topology information including at least one of the following: the association between nodes, interconnect bandwidth, and interconnect latency; and a transmission performance requirement determination unit configured to determine the transmission performance requirements for the access request based on the type of access request and the amount of target data.
[0018] In some embodiments, the target storage node determination unit includes: an information item determination unit configured to determine multiple performance information items for multiple storage nodes in topology information; and a first node determination unit configured to determine a target storage node from multiple storage nodes based on the multiple performance information items and transmission performance requirements.
[0019] In some embodiments, the first node determination unit includes: a set of node determination units configured to determine a set of available storage nodes from a plurality of storage nodes by matching a plurality of performance information items and transmission performance requirements; and a storage node selection unit configured to select a target storage node from the set of available storage nodes.
[0020] In some embodiments, the access request is a write request, and the apparatus further includes: a first determination unit configured to determine whether the target storage node is located in the cluster; and a first data writing unit configured to, if it is determined that the target storage node is located in the cluster, cause the compute node to write the target data to the target storage node via the network within the cluster.
[0021] In some embodiments, the cluster is a first cluster, and the apparatus further includes: a first auxiliary storage node determining unit configured to determine a first auxiliary storage node in the first cluster if it is determined that the target storage node is not located in the first cluster; a first data writing unit configured to cause a computing node to write target data to the first auxiliary storage node in the first cluster; and a third data writing unit configured to use a network for multiple storage nodes to write target data from the first auxiliary storage node to a target storage node in a second cluster of multiple clusters.
[0022] In some embodiments, the access request is a read request, and the apparatus includes: a second determination unit configured to determine whether a target storage node is located in a cluster; a second auxiliary storage node determination unit configured to determine a second auxiliary storage node in the cluster if the target storage node is not located in the cluster; a target data retrieval unit configured to cause the second auxiliary storage node to retrieve target data from the target storage node in a third cluster of the plurality of clusters using a network for multiple storage nodes; and a first target data provision unit configured to provide the target data from the second auxiliary storage node to a compute node.
[0023] In some embodiments, the apparatus further includes a second target data providing unit configured to provide target data from the target storage node to the compute node via an intra-cluster network if it is determined that the target storage node is located in the cluster.
[0024] In some embodiments, multiple clusters are used to process machine learning models.
[0025] In some embodiments, the target data may be checkpoint file data or training data for a machine learning model.
[0026] According to a third aspect of this application, an electronic device is also provided, comprising: at least one computing unit; and at least one memory coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, the instructions, when executed by the at least one computing unit, causing the device to perform the method according to the first aspect of this application.
[0027] According to a fourth aspect of this application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the method described according to a first aspect of this application.
[0028] According to a fifth aspect of this application, a computer program product is also provided, including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method described according to a first aspect of this application.
[0029] Understandably, the apparatus of the second aspect, the electronic device of the third aspect, the computer storage medium of the fourth aspect, or the computer program product of the fifth aspect provided above are used to perform the method provided in the first aspect. Therefore, the explanations or descriptions regarding the first aspect also apply to the second, third, fourth, and fifth aspects. Furthermore, the beneficial effects achieved by the second, third, fourth, and fifth aspects can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0030] The above and other features, advantages and aspects of the embodiments of this application will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description.
[0031] In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0032] Figure 1 A schematic diagram illustrates an example environment in which several embodiments of this application can be implemented;
[0033] Figure 2 A schematic flowchart of a method for determining nodes according to some embodiments of this application is shown;
[0034] Figure 3 A schematic diagram of an example architecture for a cluster network according to some embodiments of this application is shown;
[0035] Figure 4 A schematic diagram illustrating an example of a task processing procedure according to some embodiments of this application is shown;
[0036] Figure 5 A schematic diagram of the software architecture of a distributed storage system according to some embodiments of this application is shown;
[0037] Figure 6 A schematic diagram of the node arrangement of a distributed storage system according to some embodiments of this application is shown;
[0038] Figure 7 A schematic diagram illustrating an example of a write workflow according to some embodiments of this application is shown;
[0039] Figure 8 A schematic diagram illustrating an example of a read workflow according to some embodiments of this application is shown;
[0040] Figure 9 A block diagram of an apparatus for determining nodes according to some embodiments of this application is shown; and
[0041] Figure 10 A block diagram of a computing device capable of implementing several embodiments of the present application is shown. Detailed Implementation
[0042] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0043] In the description of embodiments of this application, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0044] Traditional distributed storage systems can deploy multiple storage nodes. When expanding nodes within the cluster, both performance and capacity can be scaled. They offer strong scalability within the cluster, capable of scaling to thousands of nodes. Data is distributed across the nodes, resulting in near-linear performance growth and providing extremely high read / write bandwidth. Furthermore, each node can participate in data management and storage, distributing system load relatively evenly. This distributed storage system typically uses a minimum of three storage nodes, and fewer nodes result in lower capacity utilization, making it unsuitable for small-capacity scenarios. Capacities generally start at the hundreds of terabytes (TB) level, reaching the exabyte (EB) level. Cross-node EC (EC) can be used for data protection within this distributed storage system; if one node fails, other nodes continue operating, and the number of node failures depends on the number of EC redundancies. Additionally, compute and storage nodes are interconnected via a network.
[0045] However, existing distributed storage management systems do not consider the impact of the network layer on storage performance. In fully interconnected, non-converging networks, network bandwidth is sufficient and does not constitute a bottleneck for storage performance; therefore, they are typically deployed in a single cluster. However, in other networking scenarios, such as when the distributed storage system is stored in multiple clusters, and the clusters are only interconnected by storage nodes, the network may become a bottleneck.
[0046] To address at least some of the aforementioned problems and other potential issues, in embodiments of this application, a management device can receive access requests for target data from compute nodes in one of multiple clusters. A cluster includes a set of interconnected compute nodes and a set of storage nodes. Next, the management device needs to determine the topology information for the multiple compute nodes and storage nodes in the multiple clusters, as well as the transmission performance requirements for the access request. Using the obtained topology information and transmission performance requirements, the management device can determine the target storage node among the multiple storage nodes that can be used for the access request to process the target data. In this way, by utilizing the topology information for storage nodes in multiple clusters, the optimal target storage node to be used can be quickly determined from the nodes in the clusters according to the requirements of the access request, improving the accuracy of node selection, accelerating the access speed of data requests, and enhancing the user experience.
[0047] Figure 1 A schematic diagram of an example environment 100 in which various embodiments of this application can be implemented is shown. For example... Figure 1As shown, example environment 100 includes management device 104, which can be used to determine the appropriate storage node for data access requests in computing nodes. Management device 104 includes, but is not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants, media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices.
[0048] In example environment 100, there are also multiple clusters 102-1, ..., 102-N, where N is a positive integer. For ease of description, clusters 102-1, ..., 102-N can also be collectively referred to as cluster 102. Multiple clusters can include multiple compute nodes and multiple storage nodes. In one example, each of the multiple clusters can include a set of compute nodes and a set of storage nodes interconnected via an internal network; for example, cluster 102-1 includes a set of compute nodes 106 and a set of storage nodes 108. The number of compute nodes and storage nodes in each cluster can be set to any suitable number. In another example, some clusters may include both compute nodes and storage nodes, while others may only include storage nodes. These storage nodes within the multiple clusters constitute a distributed storage system, and data is exchanged between storage nodes in the clusters via an inter-cluster network built between the storage nodes. Due to the limitations of the inter-cluster network, there are network limitations, such as interconnect bandwidth and interconnect latency, when exchanging data between storage nodes in the clusters using the inter-cluster network. Additionally, the compute nodes in cluster 102 can be used to train machine learning models. In this case, the compute nodes may send data access requests to the distributed storage system, such as write requests for storing checkpoint files, or read requests for reading checkpoint files or training data.
[0049] The compute nodes in cluster 102 can be any suitable computing device. In one example, a compute node is a dedicated computing device that can be used to perform inference for artificial intelligence models, such as a neural-network processing unit (NPU) device or a GPU device. In another example, a compute node can also be a personal computer, a server, a minicomputer, a mainframe computer, or a distributed computing environment that includes any of the above systems or devices. The storage nodes in cluster 102 can be any storage device suitable for storing data. For example, the storage device may include a processor, RAM, and a hard disk.
[0050] When a compute node 110 in a set of compute nodes 106 is processing a task, such as training a machine learning model, it can send an access request to the management device 104, such as a write request or a read request for target data. The access request may include information such as the size of the data to be accessed and / or the required storage space.
[0051] Upon receiving access request 112, management device 104 can obtain topology information 116 for the storage nodes and compute nodes in cluster 102. This topology information 116 includes performance parameters for each node, such as interconnect bandwidth and interconnect latency information for each storage node. Interconnect bandwidth refers to the data transfer rate from the cluster to that node, typically measured in bits per second (bps), such as megabits per second (Mbps) or gigabits per second (Gbps). It represents the amount of data that can be transferred between nodes per unit time. Interconnect latency, also known as delay or latency time, refers to the time required for data to reach that node, typically measured in milliseconds (ms) or microseconds (μs). It reflects the speed and timeliness of data transmission between nodes.
[0052] In addition, the management device 104 will determine the transmission performance requirements for the access request based on the traffic planning of the distributed storage system. For example, it will determine the transmission performance requirements based on the size of the data in the access request and the bandwidth allocation strategy in the distributed storage system. The transmission performance requirements may refer to the requirements for transmission parameters or network parameters during transmission, such as the corresponding bandwidth and latency requirements. Then, the management device 104 will search for storage nodes that meet the performance transmission requirements 114 in the topology information 116 as target storage nodes 118 for storing or providing the data for the access request.
[0053] By utilizing the topology information of storage nodes across multiple clusters, this method can quickly determine the optimal target storage node from the cluster nodes based on the access request requirements. This improves the accuracy of node selection, speeds up data access, and enhances the user experience.
[0054] The above combination Figure 1 A schematic diagram illustrating an example environment 100 in which embodiments of this application can be implemented is described. The following is in conjunction with... Figure 2 This is a schematic flowchart illustrating example 200 of a method for controlling a node according to some embodiments of this application. Example 200 can be provided by... Figure 1 The management device 104 and any suitable computing device in the system are used for execution.
[0055] At box 202, management device 104 receives access requests for target data from compute nodes in one of multiple clusters. This cluster includes a set of compute nodes and a set of storage nodes interconnected via a network within the cluster. Management device 104 can receive various requests related to data storage or retrieval from compute nodes in the multiple clusters.
[0056] For example, these multiple clusters are used to process machine learning model tasks. Therefore, when a compute node trains a machine learning model, it can obtain and save checkpoint file data during the training process so that it can recover to the state corresponding to that checkpoint file in case of a training failure. To save the checkpoint file data, it needs to send a write request for the checkpoint file to the management device 104. If a failure occurs during training, the compute node will send a read request to the management device 104 to read the checkpoint file data. Alternatively, if the training data of the machine learning model is also stored in storage nodes within the cluster, other nodes can send requests to the management device 104 to read the training data. Alternatively, the compute node can send any suitable data access request to the management device 104.
[0057] At box 204, management device 104 determines topology information for multiple compute nodes and multiple storage nodes in multiple clusters, and transmission performance requirements for access requests. The multiple compute nodes include a group of compute nodes in one cluster, and the multiple storage nodes include a group of storage nodes in one cluster. To understand the connectivity between nodes in the cluster and the performance of each node, management device 104 can collect topology information for the multiple compute nodes and multiple storage nodes. In one example, this topology information includes interconnect bandwidth and interconnect latency for each storage node. Additionally, the topology information may further include the relationships between nodes. Furthermore, management device 104 can further determine the transmission performance requirements for an access request based on that access request. For example, management device 104 needs to obtain the type of access request and the amount of data of the target data to be accessed.
[0058] For example, management device 104 checks whether the access request is a read request or a write request, and the amount of data that the access request needs to process. Then, management device 104 determines the transmission performance requirements for the access request based on the type of access request and the amount of target data to be processed. Additionally, these transmission performance requirements may include bandwidth and latency requirements for the access request.
[0059] Alternatively, when determining transmission performance requirements, in addition to utilizing the type of access request and the amount of target data to be processed, it is also necessary to further obtain information such as the bandwidth allocation strategy and load balancing strategy for nodes in the management device to determine the transmission performance requirements for the access request.
[0060] At box 206, management device 104, based on topology information and transmission performance requirements, determines the target storage node from among multiple storage nodes that can be used for the access request to process the target data. After obtaining the topology information for the storage node and the transmission performance requirements of the access request, the target storage node for the access request can be determined from among the multiple storage nodes of the distributed storage system.
[0061] In some embodiments, when determining a target storage node available for an access request among a plurality of storage nodes, the management device 104 may determine a plurality of performance information items for the plurality of storage nodes in the topology information, wherein each performance information item corresponds to a storage node and may include interconnect bandwidth and interconnect latency information for that storage node. The management device 104 then uses the plurality of performance information items and transmission performance requirements to determine the target storage node from the plurality of storage nodes. For example, the management device matches the plurality of performance information items and transmission performance requirements to determine a set of available storage nodes from the plurality of storage nodes. The bandwidth and latency of this set of available storage nodes both meet the requirements of the access request. The management device then selects the target storage node from the set of available storage nodes. For example, the storage node with the lightest load is selected from the set of available storage nodes as the target storage node.
[0062] In some embodiments, the access request processed by the management device 104 may be a write request. After determining the target storage node, the management device 104 further determines whether the target storage node is located in the same cluster as the compute node. If the target storage node is located in the same cluster as the compute node, the management device 104 may send a message to the compute node to write the target data to the target storage node via the network within the cluster. For ease of description, the cluster in which the compute node issuing the access request is located is the first cluster. If the management device 104 determines that the target storage node is not located in the first cluster but in a second cluster among multiple clusters, the management device 104 needs to further determine a first auxiliary storage node from the first cluster. The first auxiliary storage node may be any suitable storage node in the first cluster, such as the lightest-loaded storage node. Then, the management device 104 notifies the compute node to write the target data to the first auxiliary storage node in the first cluster. Next, the first storage node uses the inter-cluster network for multiple storage nodes to write the target data from the first auxiliary storage node to the target storage node in the second cluster among multiple clusters.
[0063] In some embodiments, the access request processed by the management device 104 may be a read request. After determining the target storage node, the management device 104 further determines whether the target storage node is located in the first cluster where the compute node resides. If the target storage node is determined to be located in the first cluster where the compute node resides, the management device 104 may send a message to the compute node to cause the compute node to read the target data from the target storage node through the network within the cluster. If the target storage node is determined not to be located in the first cluster, the management device needs to determine a second auxiliary storage node in the first cluster. The second auxiliary storage node may be any suitable storage node in the first cluster, such as the lightest-loaded storage node. Next, the management device 104 may cause the second auxiliary storage node to use the inter-cluster network for multiple storage nodes to pull the target data from the target storage node in a third cluster, and then provide the pulled target data from the second auxiliary storage node to the compute node. Additionally, the target data is also stored in the second auxiliary storage node so that when the management device 104 receives a read request for the target data from a compute node in the first cluster, it can directly obtain the target data from the second auxiliary storage node.
[0064] By utilizing the topology information of storage nodes across multiple clusters, this method can quickly determine the optimal target storage node from the cluster nodes based on the access request requirements. This improves the accuracy of node selection, speeds up data access, and enhances the user experience.
[0065] The above combination Figure 2 A schematic flowchart illustrating a method for determining nodes according to some embodiments of this application is described below. Figure 3 This diagram illustrates an example architecture for cluster networking according to some embodiments of this application. In example architecture 300, there are two clusters, 302 and 304. Example architecture 300, including two clusters, is merely an example and may include any number of clusters. Additionally, the cluster may be an artificial intelligence (AI) cluster (e.g., POD).
[0066] In example architecture 300, the clusters are networked in a Dragonfly+ topology. Additionally, clusters can also be networked in other topologies. Cluster 302 includes a computing module 306, which performs computational tasks, such as computing machine learning models. Computing module 306 includes two computing nodes, each with 64 CPUs and 16 400G interfaces. Additionally, computing module 306 can include any suitable number of computing nodes, and each node can have any number of dedicated processing resources such as CPUs or GPUs, as well as any number of interfaces, with the interface bandwidth set to any suitable value.
[0067] In addition, the example architecture 300 also includes a storage module 314, which includes storage nodes 316 and 318. Each node has 32 interfaces with a bandwidth of 400G. Alternatively, the storage module 314 may also include any number of storage nodes, and each storage node may include any number of interfaces, and the bandwidth of the interfaces can also be set to any size. In this cluster 302, compute nodes and storage nodes communicate via an intra-cluster interconnect network 308. For example, they communicate via two switches in the intra-cluster interconnect network 308, such as switches 310 and 312. Additionally, each compute node and each storage node is connected to each switch in the intra-cluster interconnect network 308. Cluster 304 is interconnected with cluster 302, and cluster 304 also has a set of compute nodes and a set of storage nodes. The number of compute nodes and storage nodes included in cluster 304 can be set as needed. For example, cluster 304 includes a storage module 322, which includes two storage nodes 324 and 326. When a compute node in cluster 302 needs data from cluster 304, the data from the storage nodes in cluster 304 can first be loaded into a storage node in cluster 302, such as storage node 318, via the inter-cluster interconnection network 320, and then provided to the compute node through storage node 318. If a compute node needs data from a storage node within the same cluster, the data can be transferred directly through the intra-cluster interconnection network.
[0068] Within the cluster, data is transferred between compute nodes and storage nodes via switches. Additionally, cluster 302 communicates with cluster 304 via inter-cluster interconnection network 320. Inter-cluster interconnection network 320 includes one or more switches. The primary function of inter-cluster interconnection network 320 is to interconnect storage nodes between different clusters to facilitate data transfer.
[0069] The above combination Figure 3 A schematic diagram of an example architecture for cluster networking according to some embodiments of this application is described below, in conjunction with... Figure 4 A schematic diagram illustrating examples of task processing procedures according to some embodiments of this application.
[0070] In Example 400, the computation task scheduling system 402 is used to allocate computation tasks, such as assigning machine learning model training tasks to different computation nodes. It can also determine the checkpoint (CKPT) files for each task and the data to be loaded, and send storage requirements to the storage management system 404. Then, the storage management system 404 can select the nearest node for accessing the checkpoint files via the storage management network 406 based on the network topology and traffic planning for the storage nodes in the distributed storage system. For example, the traffic planning includes bandwidth and latency information for storage requirements, and the network topology information may also include interconnect bandwidth and interconnect latency for the storage nodes. Therefore, the storage management system can select a storage node based on the network topology information according to this traffic rule. For training data, if it crosses affinity domains, the data is actively pulled to the nearest storage system.
[0071] The above combination Figure 4 The following is a schematic diagram illustrating examples of task processing procedures according to some embodiments of this application; combined with... Figure 5 A schematic diagram illustrating the software architecture of a distributed storage system according to some embodiments of this application;
[0072] In this software architecture example 500, the distributed storage system 502 includes a protocol layer 504, a service layer 508, and an index layer 516. The protocol layer 504 includes a storage protocol 506, which can be a standard protocol for storing and retrieving data over a network. For example, this storage protocol may be for data objects. Additionally, this layer may also include the Network File System (NFS) protocol, the Server Message Block (SMB) protocol, and the Hadoop Distributed File System (HDFS) protocol. The NFS protocol is used for sharing files between systems over a network, supporting direct access to remote files across hosts (based on remote procedure calls). The Server Message Block protocol can be used for sharing resources such as files and printers between systems. The Hadoop Distributed File System (HDFS) protocol is designed for large-scale data storage, supporting the storage of files in blocks across multiple servers, providing high fault tolerance and high throughput.
[0073] Service layer 508 includes multiple modules, such as a Quality of Service (QoS) module 510, a storage topology module 512, and a topology-aware QoS system 514. The QoS module 510 controls service quality, for example, by queuing or prioritizing data access requests. The storage topology module 512 collects topology information between storage nodes in the distributed storage system 502. This topology information includes, for example, the bandwidth and latency corresponding to each storage node. Additionally, the topology also includes the compute nodes associated with the distributed storage system. The topology-aware QoS system 514 identifies suitable storage nodes from the topology information based on the bandwidth and latency of data access requests.
[0074] The index layer 516 includes buckets 518 corresponding to the storage protocol 506, which are used to store and organize data objects. For example, buckets 518 can perform data storage and management, as well as access control management. Additionally, the index layer also includes other modules, such as a distributed file system.
[0075] Below the index layer 516 are the storage nodes in the distributed storage management system 502 used to store data. This example shows five storage nodes. It is merely an example and not a specific limitation of this disclosure; a distributed storage system may include any number of storage nodes.
[0076] After receiving a data access request, the distributed storage system can transmit the request to the Quality of Service (QoS) module 510 via storage protocol 506. The request is then transmitted via QoS module 510 to the storage bucket 518 to access the corresponding data through the storage node. The storage node is an object storage device (OSD). Figure 5 This illustrates an example of accessing data at the software architecture level. Storage nodes in a distributed storage system consist of nodes located in different clusters. The storage topology module 512 can perceive the topology information for the storage nodes, and the topology-aware Quality of Service (QoS) system 514 determines the storage node that satisfies the data access request, allowing data to be stored or retrieved from that node. The determined storage node may be located in the same cluster as the computing device sending the access request, or it may be located in a different cluster.
[0077] The above combination Figure 5 A schematic diagram of the software architecture of a distributed storage system according to some embodiments of this application is described below; in conjunction with... Figure 6 A schematic diagram illustrating the node arrangement of a distributed storage system according to some embodiments of this application.
[0078] In Example 600, there are clusters 602 and 606. For example, clusters 602 and 606 can be AIPOD. Cluster 602 includes N compute cabinets and M storage cabinets, where N and M are both positive integers. Compute nodes in the compute cabinets within this cluster can access storage nodes in the storage cabinets through the cluster's internal switches. The interconnection method can be any suitable network topology, such as a fat tree network, which guarantees high-bandwidth interconnection between compute nodes and storage nodes. Since the network interfaces of all devices within a cluster can be fully connected through the cluster's internal switches, the data access speed within the cluster is not affected by the cluster's internal communication network.
[0079] Example 600 also includes cluster 606, which comprises N compute cabinets and M storage cabinets. Compute nodes in the compute cabinets within this cluster can access storage nodes in the storage cabinets via intra-cluster switches. The interconnection method can also be a fat tree network, which guarantees high-bandwidth interconnection between compute and storage nodes.
[0080] When scaling up the distributed storage system 610 to multiple clusters, you can configure storage nodes with... Figure 5 The distributed storage system 610 shown in the software architecture serves as the management node. Storage nodes in each cluster are interconnected individually, for example, via a switch 604. The interconnection bandwidth is narrower than the intra-cluster interconnection bandwidth, and the interconnection is achieved through a switch. The distributed storage system 610 can be deployed centrally on a single storage node or distributed across all storage nodes.
[0081] The storage nodes in the distributed storage system 610 are those added to clusters 602 and 606. To facilitate the management of storage nodes distributed across different clusters, a storage topology module 612 is included in the distributed storage system 610. This module collects topology information about storage nodes and compute nodes in different clusters, including interconnection relationships, interconnection bandwidth, and interconnection latency. Then, a topology-aware Quality of Service (QoS) system 614 selects a suitable storage node for access requests from compute nodes. For example, for data requests, a suitable storage node is found based on bandwidth, latency, and other requirements.
[0082] Figure 6 A schematic diagram of the node arrangement of a distributed storage system according to some embodiments of this application is described below; in conjunction with... Figure 7A schematic diagram illustrating an example of a write workflow according to some embodiments of this application is provided. In example 700, in the first step, a compute node in compute cabinet #1 of cluster 702 initiates a write request. This request is then transmitted via a switch in cluster 702 to a storage management system located in storage cabinet #M. The storage management system then selects a storage node based on topology, bandwidth, latency, and other information, and returns the information to the compute node. Then, in the second step, while the storage node is located in storage cabinet #1 of cluster 702, the compute node in compute cabinet #1 writes data to the storage node in storage cabinet #1.
[0083] In some embodiments, if the data is cold data, such as backup data, it may be stored on storage nodes located in other clusters. In this case, the storage management system in storage cabinet #M selects a storage node from cluster 702 as a secondary storage node, and then writes the data to be written to storage nodes in other clusters to that secondary storage node. Then, the network between storage nodes across clusters is further utilized to write the data from that secondary storage node to storage nodes in other clusters.
[0084] The above combination Figure 7 A schematic diagram illustrating an example of a write workflow according to some embodiments of this application is described below; in conjunction with Figure 8 A schematic diagram illustrating an example of a read workflow according to some embodiments of this application.
[0085] Example 800 illustrates a remote read process. In the first step, a compute node in compute cabinet #1 of cluster 802 initiates a read request, which is transmitted via the switch within cluster 802 to the storage management system in storage cabinet #M. The storage management system detects that the file resides on a remote node, such as a storage node in storage cabinet #M of cluster 806. In the second step, the storage management system selects a local storage node to retrieve the data from the remote storage system and caches the data in the local storage system. For example, it might first select a storage node in storage cabinet #M of cluster 802 to retrieve the data from the storage node in storage cabinet #M of cluster 806 via switch 804, which connects the inter-cluster storage nodes, and then return the data. This selected local storage node can be chosen based on load, such as the lightest-loaded storage node. In the third step, when the compute node accesses the data again, it can access it locally from the nearest node.
[0086] Figure 9 A block diagram of an apparatus 900 for determining nodes according to an embodiment of this application is further shown. The apparatus 900 is applied to a management device and may include multiple modules for performing tasks such as... Figure 2 The corresponding steps in example 200 of the method discussed herein. For example... Figure 9As shown, the apparatus 900 includes: an access request receiving unit 902, configured to receive an access request for target data from a compute node in one of a plurality of clusters; a topology and performance determination unit 904, configured to determine topology information for a plurality of compute nodes and a plurality of storage nodes in the plurality of clusters and transmission performance requirements for the access request; and a target storage node determination unit 906, configured to determine a target storage node among the plurality of storage nodes that can be used for the access request to process the target data based on the topology information and transmission performance requirements.
[0087] In some embodiments, the topology and performance determination unit 904 includes: a topology information collection unit configured to collect topology information for multiple computing nodes and multiple storage nodes, the topology information including at least one of the following: the association between nodes, interconnect bandwidth and interconnect latency; and a transmission performance requirement determination unit configured to determine the transmission performance requirements for the access request based on the type of access request and the amount of target data.
[0088] In some embodiments, the target storage node determination unit 906 includes: an information item determination unit configured to determine multiple performance information items for multiple storage nodes in topology information; and a first node determination unit configured to determine a target storage node from multiple storage nodes based on the multiple performance information items and transmission performance requirements.
[0089] In some embodiments, the first node determination unit includes: a set of node determination units configured to determine a set of available storage nodes from a plurality of storage nodes by matching a plurality of performance information items and transmission performance requirements; and a storage node selection unit configured to select a target storage node from the set of available storage nodes.
[0090] In some embodiments, the access request is a write request, and the apparatus further includes: a first determination unit configured to determine whether the target storage node is located in the cluster; and a first data writing unit configured to, if it is determined that the target storage node is located in the cluster, cause the compute node to write the target data to the target storage node via the network within the cluster.
[0091] In some embodiments, the cluster is a first cluster, and the apparatus further includes: a first auxiliary storage node determining unit configured to determine a first auxiliary storage node in the first cluster if it is determined that the target storage node is not located in the first cluster; a first data writing unit configured to cause a computing node to write target data to the first auxiliary storage node in the first cluster; and a third data writing unit configured to use a network for multiple storage nodes to write target data from the first auxiliary storage node to a target storage node in a second cluster of multiple clusters.
[0092] In some embodiments, the access request is a read request, and the apparatus includes: a second determination unit configured to determine whether a target storage node is located in a cluster; a second auxiliary storage node determination unit configured to determine a second auxiliary storage node in the cluster if the target storage node is not located in the cluster; a target data retrieval unit configured to cause the second auxiliary storage node to retrieve target data from the target storage node in a third cluster of multiple clusters using a network for multiple storage nodes; and a first target data provision unit configured to provide the target data from the second auxiliary storage node to a compute node.
[0093] In some embodiments, the apparatus further includes a second target data providing unit configured to provide target data from the target storage node to the compute node via an intra-cluster network if it is determined that the target storage node is located in the cluster.
[0094] In some embodiments, multiple clusters are used to process machine learning models.
[0095] In some embodiments, the target data may be checkpoint file data or training data for a machine learning model.
[0096] Figure 10 A schematic block diagram of an example device 1000 that can be used to implement embodiments of the present application is shown. For example, embodiments of the present application... Figure 1 The computing nodes, storage nodes, and management devices 104 in the example device 1000 can be implemented as such. As shown, device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 1002 or loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0097] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0098] The various processes and handling described above, such as method example 200, can be executed by processing unit 1001. For example, in some embodiments, method example 200 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by CPU 1001, one or more actions of method example 200 described above can be performed.
[0099] This application may be a method, apparatus, system, chip, and / or computer program product. A chip may include a processing unit and a communication interface, the processing unit being capable of processing program instructions received from the communication interface. A computer program product may include a computer-readable storage medium on which computer-readable program instructions for performing various aspects of this application are stored.
[0100] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0101] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0102] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.
[0103] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0104] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0105] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0107] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for determining nodes, characterized in that, The method includes: Receive access requests for target data from a compute node in one of multiple clusters; Determine the topology information for multiple compute nodes and multiple storage nodes in the multiple clusters, and the transmission performance requirements for the access request. Based on the topology information and the transmission performance requirements, a target storage node is determined from among the plurality of storage nodes that can be used for the access request to process the target data.
2. The method according to claim 1, characterized in that, Determining the topology information for multiple compute nodes and multiple storage nodes in the multiple clusters and the transmission performance requirements for the access request includes: Collect topology information for the plurality of computing nodes and the plurality of storage nodes, the topology information including at least one of the following: the association relationships between nodes, interconnect bandwidth, and interconnect latency; and Based on the type of the access request and the amount of the target data, the transmission performance requirements for the access request are determined.
3. The method according to claim 1, characterized in that, Determining the target storage node among the plurality of storage nodes that can be used for the access request to process the target data includes: Determine multiple performance information items for the multiple storage nodes in the topology information; and Based on the multiple performance information items and the transmission performance requirements, the target storage node is determined from the multiple storage nodes.
4. The method according to claim 3, characterized in that, Determining the target storage node from the plurality of storage nodes based on the plurality of performance information items and the transmission performance requirements includes: By matching the plurality of performance information items with the transmission performance requirements, a set of available storage nodes is determined from the plurality of storage nodes; and Select the target storage node from the set of available storage nodes.
5. The method according to claim 1, characterized in that, The access request is a write request, and the method further includes: Determine whether the target storage node is located in the cluster; and If it is determined that the target storage node is located in the cluster, the compute node writes the target data to the target storage node via the network within the cluster.
6. The method according to claim 5, characterized in that, The cluster is the first cluster, and the method further includes: If it is determined that the target storage node is not located in the first cluster, then determine the first auxiliary storage node in the first cluster; The computing node writes the target data to the first auxiliary storage node in the first cluster; and Using a network of the plurality of storage nodes, the target data is written from the first auxiliary storage node to the target storage node in the second cluster of the plurality of clusters.
7. The method according to claim 1, characterized in that, The access request is a read request, and the method includes: Determine whether the target storage node is located in the cluster; If it is determined that the target storage node is not located in the cluster, then a second auxiliary storage node in the cluster is determined. The second auxiliary storage node uses the network for the plurality of storage nodes to pull the target data from the target storage node in the third cluster of the plurality of clusters; and The target data is provided to the computing node by the second auxiliary storage node.
8. The method according to claim 7, characterized in that, Also includes: If it is determined that the target storage node is located in the cluster, the target data is provided from the target storage node to the computing node via the cluster network.
9. The method according to claim 1, characterized in that, The multiple clusters are used to process machine learning models.
10. The method according to claim 1, characterized in that, The target data can be checkpoint file data or training data for a machine learning model.
11. An apparatus for determining nodes, characterized in that, The device includes: The access request receiving unit is configured to receive access requests for target data from compute nodes in one of multiple clusters. A topology and performance determination unit is configured to determine topology information for multiple compute nodes and multiple storage nodes in the plurality of clusters and transmission performance requirements for the access request, and The target storage node determination unit is configured to determine, based on the topology information and the transmission performance requirements, a target storage node among the plurality of storage nodes that can be used for the access request to process the target data.
12. The apparatus according to claim 11, characterized in that, The topology and performance determination unit includes: A topology information collection unit is configured to collect topology information for the plurality of computing nodes and the plurality of storage nodes, the topology information including at least one of the following: the association relationships between nodes, interconnect bandwidth, and interconnect latency; and The transmission performance requirement determination unit is configured to determine the transmission performance requirement for the access request based on the type of the access request and the amount of target data.
13. The apparatus according to claim 11, characterized in that, The target storage node determination unit includes: An information item determination unit is configured to determine multiple performance information items for the plurality of storage nodes in the topology information; and The first node determination unit is configured to determine the target storage node from the plurality of storage nodes based on the plurality of performance information items and the transmission performance requirements.
14. The apparatus according to claim 13, characterized in that, The first node determination unit includes: A set of node determination units is configured to determine a set of available storage nodes from the plurality of storage nodes by matching the plurality of performance information items with the transmission performance requirements; and The storage node selection unit is configured to select the target storage node from the set of available storage nodes.
15. The apparatus according to claim 11, characterized in that, The access request is a write request, and the device further includes: The first determination unit is configured to determine whether the target storage node is located in the cluster; and The first data writing unit is configured to, if it is determined that the target storage node is located in the cluster, cause the computing node to write the target data to the target storage node via the intra-cluster network.
16. The apparatus according to claim 15, characterized in that, The cluster is a first cluster, and the apparatus further includes: The first auxiliary storage node determination unit is configured to determine the first auxiliary storage node in the first cluster if it is determined that the target storage node is not located in the first cluster. The data is sent to the data writing unit and configured to cause the computing node to write the target data to the first auxiliary storage node in the first cluster; and The third data writing unit is configured to use a network for the plurality of storage nodes to write the target data from the first auxiliary storage node to the target storage node in the second cluster of the plurality of clusters.
17. The apparatus according to claim 11, characterized in that, The access request is a read request, and the device includes: The second determination unit is configured to determine whether the target storage node is located in the cluster; The second auxiliary storage node determination unit is configured to determine a second auxiliary storage node in the cluster if it is determined that the target storage node is not located in the cluster. The target data retrieval unit is configured to cause the second auxiliary storage node to retrieve the target data from the target storage node in a third cluster of the plurality of clusters using a network for the plurality of storage nodes; and The first target data providing unit is configured to provide the target data from the second auxiliary storage node to the computing node.
18. The apparatus according to claim 17, characterized in that, Also includes: The second target data providing unit is configured to provide the target data from the target storage node to the computing node via an intra-cluster network if it is determined that the target storage node is located in the cluster.
19. The apparatus according to claim 11, characterized in that, The multiple clusters are used to process machine learning models.
20. The apparatus according to claim 11, characterized in that, The target data can be checkpoint file data or training data for a machine learning model.
21. An electronic device, comprising: At least one computing unit; as well as At least one memory coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, the instructions, when executed by the at least one computing unit, causing the device to perform the method according to any one of claims 1-10.
22. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1-10.
23. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1-10.