Multi-tenant set communication fabric

By building a virtual topology in a multi-tenant collective communication structure and utilizing indirect communication links between computing nodes, the problem of communication performance bottlenecks in distributed computing systems is solved, and the optimal utilization of system resources and the improvement of communication performance is achieved.

CN120144522APending Publication Date: 2025-06-13HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410763088.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-06-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to fully utilize the indirect communication links between computing nodes in distributed computing systems, resulting in a bottleneck in communication performance, reduced system utilization, and increased total cost of ownership.

Method used

By leveraging any available communication links, including direct and indirect connections, building a virtual topology, optimize communication paths between computing nodes, in order to achieve optimal system resource utilization and communication performance.

Benefits of technology

By utilizing indirect communication links, the communication performance between computing nodes is improved, the utilization of system resources is enhanced, the total cost of ownership is reduced, and the system throughput and end-to-end latency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144522A_ABST
    Figure CN120144522A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a multi-tenant set communication structure. Systems and methods for a multi-tenant collective communication fabric are provided for optimal utilization of communication links between compute nodes of an interconnected system. Examples include assigning a plurality of compute nodes to a first workload and obtaining a topology of the interconnected system representing indirect paths between the assigned compute nodes. The indirect path includes unallocated compute nodes of the interconnected system. Examples include creating a plurality of resource slices of unallocated compute nodes, a first slice dedicated to processing and forwarding data traffic along an indirect path, and a second slice configured to be allocated to a second workload. Examples also include executing a first workload by the allocated computing node and the unallocated computing node, wherein data traffic from the allocated plurality of computing nodes is transmitted via the indirect communication path.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] End users such as enterprises can use multi-tenant systems to run applications. These systems can be composed of multiple computing nodes (e.g., CPUs and GPUs) arranged according to the physical device topology within the system. Various communication link protocols can be used to implement the physical device topology, such as but not limited to Peripheral Component Interconnect Express (PCIe) and NVLink. The inter-system physical topology can connect multiple systems via high-speed network communication links.

[0002] Multi-tenant systems can assist in executing applications. Executing an application can include executing multiple computing workloads. Each workload consumes system resources such as memory and time. Multi-tenant systems can enhance the performance of applications by allocating multiple computing nodes to the workloads. By assigning computing nodes to different applications and workloads, different applications and different workloads can be distributed on a single multi-tenant system. In this way, the system resources of the multi-tenant system can be shared among multiple applications. Brief Description of the Drawings

[0003] The present disclosure according to one or more embodiments is described in detail with reference to the accompanying drawings. The drawings are provided for illustrative purposes only and depict typical or example embodiments.

[0004] Figure 1 is a schematic block diagram of the physical device topology of an interconnected computing system implemented according to the present disclosure.

[0005] Figure 2 is a schematic block diagram of an example architecture of a multi-tenant collective communication structure according to the present disclosure.

[0006] Figure 3 is a schematic block diagram of an example MCCF controller according to the present disclosure.

[0007] Figure 4 depicts an example virtual topology according to an example implementation of the present disclosure.

[0008] Figure 5 is a schematic block diagram of an example architecture of the computing node and MCCF data plane implementation on each computing node of an interconnected system according to the present disclosure.

[0009] Figure 6 is an example computing component that can be used to implement various features of multi-tenant collective communication in distributed computing according to the implementations disclosed herein.

[0010] Figure 7 is an example computer system that can be used to implement various features of multi-tenant collective communication in the distributed computing of the present disclosure.

[0011] The accompanying drawings are not exhaustive and do not limit the present disclosure to the exact forms disclosed. Detailed Description

[0012] Recent advancements in artificial intelligence (AI) and machine learning (ML) have not only revolutionized the technology industry but also transformed daily life. Increasingly complex ML models have been developed in various subject areas, enabling more powerful applications. As an example, GPT (Generative Pretrained Transformer) models have been developed in the field of natural language processing (NLP), driving services such as ChatGPT. The complexity of these GPT models has grown from 117 million parameters to 175 billion parameters.

[0013] Training such large-scale ML models may exceed the capabilities of a single computing node (e.g., central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), etc.), or even the capabilities of a single interconnected system of multiple computing nodes. To accelerate the ML training process, distributed training has been introduced, which can divide the training workload into sub-workloads (referred to herein as tasks) through various forms of parallelism, such as but not limited to, data parallelism, pipeline parallelism, tensor parallelism, etc. Each task can be assigned to a different computing node of the interconnected system. Training results from different computing nodes can be collected and synchronized. The exchange of training results between multiple computing nodes can be achieved through collective communication, such as Message Passing Interface (MPI) and NVIDIA-based collective communication. For example, the NVIDIA Collective Communication Library (NCCL) provides various inter-GPU communication primitives, including All-Reduce, Broadcast, Reduce, All-Gather, Reduce-Scatter, to name a few. Other example libraries include but are not limited to the Radeon Open Compute (ROCm) Communication Collective Library (RCCL) and the Microsoft Collective Communication Library (MSCCL), both of which are designed to achieve similar goals as NCCL. For example, MSCCL can be built on top of NCCL to provide a flexible and programmable interface to implement collective algorithms.

[0014] The above-mentioned interconnection system can provide benefits to end-users when executing applications. For example, in a pay-as-you-go model, the end-user pays for each workload executed on the interconnection system, which can reduce the total cost of ownership (TCO) for a given end-user across a range of applications and access patterns associated with operating their own dedicated proprietary system. Additionally, the interconnection system can provide various services that end-users can deploy, which may help shorten the time to market relative to proprietary dedicated systems. Further, the interconnection system can support the dynamic scaling of services that enable end-users to handle different system access patterns. The interconnection system can also provide end-users with different accelerators, such as GPUs (e.g., NVIDIA and AMD) and FPGAs (e.g., AMD Xilinx), which may otherwise be difficult to deploy and manage.

[0015] As computing power, ML models, and data grow, communication between computing nodes can become a performance bottleneck in the distributed training of ML models. Various methods have been proposed to establish high-speed interconnections between computing nodes of an interconnection system. However, many interconnections can be costly and may be limited to a small number of computing nodes in a specially designed physical device topology. For example, the NVIDIA DGX-1 system can adopt a specially designed hybrid cube mesh topology with multiple NVLinks that connect up to 8 GPUs on the same host.

[0016] To optimize communication performance, a collective communication library can utilize the physical device topology to create a logical topology (sometimes referred to herein as a virtual topology) among the computing nodes assigned to a workload. Collective communication libraries (e.g., NCCL, RCCL, MSCCL, etc.) can be used to generate various logical topologies (e.g., ring topology, tree topology, etc.) from the physical device topology to achieve enhanced communication performance based on communication primitive types (e.g., all-reduce, broadcast, reduce, all-gather, reduce-scatter, etc.) and available interconnection links of the physical device topology.

[0017] However, the formation of logical topologies by these traditional collective communication libraries may be restricted by the physical device topology. These methods may be limited to constructing logical topologies only using those computing nodes assigned to a specific workload and their direct connections therebetween. For example, it may not be possible to establish a logical topology by traditional methods where direct links between computing nodes are unavailable. In another example, communication links providing direct connections can have lower bandwidth and lower data transfer rates than indirect paths. However, since traditional methods may be limited to direct connections, the resulting logical topology may be limited to slower direct paths.

[0018] Due to the above limitations, indirect connections may not be fully utilized in a logical topology that could otherwise improve communication performance. In a multi-tenant system with heterogeneous infrastructure, the impact of the above limitations may be amplified. For example, heterogeneous infrastructure can include computing nodes with different computing capabilities and links with different speeds (e.g., data transfer rates), which are shared and allocated by different workloads. Some computing nodes and links may be underutilized or idle, while other computing nodes and links may be oversubscribed. As a result, the overall system utilization may be reduced, which may increase the TCO as system resources may be idle and may limit system performance such as its throughput and end-to-end latency.

[0019] Accordingly, the present disclosure provides a multi-tenant collective communication fabric (MCCF) that can enhance collective communication performance by optimal utilization of available communication links between computing nodes of an interconnected system. The techniques of the present disclosure can overcome the above technical deficiencies by leveraging any available communication link, whether directly connected or indirectly connected, to provide improved communication performance. The disclosed techniques can be well-suited for multi-tenant systems, in which, as described above, infrastructure with different computing nodes and data transfer rates can be shared by different workloads. The disclosed techniques can utilize different data transfer rates to improve communication performance between selected computing nodes.

[0020] For example, an interconnected system of computing nodes can be used for collective communication, in which multiple computing nodes of the interconnected system can be assigned to execute workloads. The workload can be divided into tasks, and each task in the tasks can be assigned to a computing node. The tasks assigned to a computing node can be regarded as the "tenants" of that computing node. Computing nodes can refer to CPUs, GPUs, TPUs, intelligent network interface controllers (e.g., NICs that can perform custom computing and communication tasks), servers, and any other computing devices that can be configured to execute assigned computing tasks. These computing nodes can be interconnected via a network of communication links that form a communication fabric. The communication links can include links that can have different speeds (e.g., different data transfer rates). Thus, for example, a first computing node can be connected to other computing nodes via a communication link with a first data transfer rate, while a second computing node can be connected to other computing nodes via a communication link with a second data transfer rate. The second data transfer rate can be faster than the first data rate. In some cases, there may be no communication link that directly connects one computing node to another computing node.

[0021] According to examples of the present disclosure, multiple computing nodes of an interconnect system can be assigned to a workload, and a logical topology of the interconnect system can be obtained. The interconnect system can include a physical device topology that can be logically represented by the logical topology. In various examples, the logical topology can include indirect communication paths between multiple computing nodes. According to the present disclosure, an indirect communication path can include one or more indirect communication links that form a connection between the assigned multiple computing nodes, and include at least one intermediate computing node communicatively connected between the assigned multiple computing nodes. As used herein, an indirect communication link can refer to a communication link that connects an assigned computing node to another unassigned computing node (e.g., an intermediate computing node) before connecting to another assigned computing node. For example, two assigned computing nodes can be directly connected to each other via a low-speed link and indirectly connected via a high-speed link through one or more intermediate unassigned computing nodes. In this case, in terms of data transfer speed and system resources, the optimal communication path can utilize the higher-speed indirect communication link as opposed to the lower-speed direct communication link.

[0022] According to examples of the present disclosure, the use of an indirect communication link can be provided by partitioning hardware resources (e.g., memory and computing resources) on the at least one intermediate computing node. For example, the hardware resources of a computing node can be logically partitioned to form hardware resource slices. Functions can be assigned to the slices according to desired operations. In an illustrative example, a slice can include a first slice dedicated to processing and forwarding data traffic along the indirect communication path and a second slice that can be assigned to a tenant (e.g., to perform a workload task).

[0023] As described above, examples of the present disclosure can execute a workload on multiple assigned computing nodes via an indirect communication path. Thus, data can be transmitted between the assigned computing nodes using an indirect communication link, which can include a communication link that can have a higher data transfer rate compared to a direct communication link. For example, a first assigned computing node can perform an assignment task of a workload and generate result data. The resulting data can be transmitted to a second assigned computing node via the indirect communication path. By utilizing the indirect communication path, an unassigned computing node can receive data traffic including the result data and forward the data traffic to the second assigned computing node for performing a task assigned to the second computing node. At the same time, the second slice of the unassigned computing node can be used for another task (tenant) assigned to the same or a different workload.

[0024] Thus, the implementation of the present disclosure can provide an optimal utilization of system resources within an interconnected system while also providing multi-tenant utilization of computing nodes. Different from existing methods, the techniques of the present disclosure can construct a virtual topology for a specific workload, which spans allocated and unallocated computing nodes and the communication links therebetween. The disclosed examples can provide a selection of any computing node in the physical device topology for the allocation of system resources. This can be achieved by slicing the hardware resources of the computing nodes and assigning a subset of the slices to perform the computing and forwarding functions of workload-related data. Other slices can also be utilized to simultaneously assign tasks of another workload to the computing nodes.

[0025] It should be noted that terms such as "optimize", "optimal", etc. used herein can be used to denote making or achieving performance that is as effective or perfect as possible. However, as will be recognized by one of ordinary skill in the art reading this document, it is not always possible to achieve perfection. Therefore, these terms can also cover making or achieving performance that is as good or effective or practical as possible in a given environment, or making or achieving performance that is better than that achievable with other settings or parameters.

[0026] Figure 1 is a schematic block diagram of a physical device topology of an interconnected computing system 100 according to an implementation of the present disclosure. The interconnected computing system 100 is an example of a multi-tenant system that can be used to execute one or more workloads. The (multiple) workloads can be attributed to a single end user (e.g., a customer, an organization, an enterprise, etc.) or multiple end users.

[0027] The interconnected computing system 100 includes a plurality of computing nodes that can be connected according to Figure 1 the physical device topology shown. The plurality of computing nodes can include various types of computing nodes, such as but not limited to CPUs, GPUs, intelligent NICs, TPUs, etc. In Figure 1 the example of, the interconnected computing system 100 includes a first plurality of computing nodes 102a - 102h (collectively referred to herein as the first computing nodes 102) provided as a first type of computing node, a second plurality of computing nodes 108a - 108b (collectively referred to herein as the second computing nodes 108) provided as a second type of computing node, and a third plurality of computing nodes 104a - 104d (collectively referred to herein as the third computing nodes 104) provided as a third type of computing node. In Figure 1 the illustrative example of, the first computing nodes 102 are provided as GPUs, the second computing nodes 108 are provided as CPUs, and the third computing nodes 104 are provided as intelligent NICs.

[0028] The interconnected computing system 100 also includes a plurality of communication links. In Figure 1In the example, the interconnected computing system 100 includes a first plurality of communication links 114a - 114h (collectively referred to herein as the first communication link 114), a second plurality of communication links 116a - 116i (collectively referred to herein as the second communication link 116), and a third plurality of communication links 112a - 112l (collectively referred to herein as the third communication link 112). The communication link 110 can be an example of a dedicated bus connecting the second computing node 108, such as a UltraPath Interconnect (UPI) or QuickPath Interconnect (QPI).

[0029] The plurality of communication links can include different speeds (e.g., bandwidth, data transfer rate, etc.). In Figure 1 the example, the communication link 114 can be configured for a first data transfer rate, the communication link 116 can be configured for a second data transfer rate, and the communication link 112 can be configured for a third data transfer rate. In an example implementation, the second data transfer rate can be faster than the first data transfer rate, and the first data transfer rate can be faster than the third data transfer rate.

[0030] In Figure 1 the example, the physical device topology of the interconnected computing system 100 includes two physical sub - topologies: a first sub - topology configured to support data transfer between the computing node 104 and the computing node 102, and a second sub - topology configured to support data transfer between the computing nodes 102. The first sub - topology can be implemented via the communication links 112 and switches 106a - 106d (collectively referred to herein as the switches 106). The second sub - topology can be implemented via the communication links 114 and 116 that form the interconnect between the computing nodes 102.

[0031] In an example, one or more protocols can be used to implement the communication links, such as but not limited to PCIe and NVIDIA NVLink. In an illustrative example, the third communication link 112 can be implemented as a PCIe link, and the switch 106 can be implemented as a PCIe switch. Additionally, the communication links 114 and 116 can be implemented as NVLink. According to various implementations, NVLink can have a higher data transfer rate relative to the data transfer rate implemented for PCIe links. Further, in this illustrative example, the NVLink of the communication link 114 can have a data transfer rate of 25 GB / s, while each communication link 116 can include two NVLinks providing a data transfer rate of 50 GB / s.

[0032] In Figure 1In the example, the first sub-topology may include a third plurality of computing nodes 108. The third plurality of computing nodes 108 may provide, for example, an inter-system connection between computing nodes 108a and 108b. In the case of the interconnected computing system 100, the third plurality of computing nodes 108 may be provided as NICs connected to the third plurality of computing nodes 104 via communication link 112 and connected to each other via communication link 110. For example, as Figure 1 shown, computing node 108a may be connected to computing node 104a via switch 106a and communication link 112a, and connected to computing node 104b via switch 106b and communication link 112l. Computing node 108b may be connected to computing node 104d via switch 106d and communication link 112f, and connected to computing node 104c via switch 106c and communication link 112g.

[0033] Although Figure 1 a specific physical device topology is depicted, the present invention is not limited to this specific topology. Other physical device topologies may be utilized within the scope of the present disclosure, as long as the physical device topology includes a plurality of computing nodes connected via a plurality of communication links to form an interconnected computing system. For example, the number of GPUs implemented may be reduced or increased. In addition, a subset of GPUs may be equipped with full connectivity to each other (e.g., all-to-all connectivity). In an illustrative example, an interconnected system consisting of sixteen GPUs may be provided, where every eight GPUs may be deployed to one of two GPU boards. These two GPU boards (and corresponding GPUs) may be interconnected via NVSwitch, with eight NVLinks used between every two NVSwitches. This configuration can support up to eight 100 Gbp / s NICs at most. These GPUs may be connected to additional GPUs via the NIC network.

[0034] The implemented physical device topology may be specifically configured to achieve certain throughput, latency, and connectivity goals of the topology desired by the operator. For example, an NVSwitch-based topology may achieve full connectivity between GPU computing nodes without the need for slower PCIe links, such as in the Figure 1 example shown.

[0035] In an example, the interconnected computing system 100 can allocate system resources to workloads based on a resource allocation policy. The system resources can include computing resources provided by computing nodes (e.g., hardware resources of the computing nodes) and networking resources such as communication bandwidth provided by communication links. The resource allocation policy can specify provisioning the system resources to achieve desired goals of a given system operator, such as but not limited to throughput and end-to-end latency. The resource allocation policy can be used to assign (e.g., allocate) one or more computing nodes and multiple communication links to a specific workload. The resource allocation policy can be managed and executed by a controller ( Figure 1 not shown in the figure), which can be a centralized controller or a distributed controller. Further details regarding the operation of an example controller are described below in conjunction with Figures 2 to 4 the figure.

[0036] Sharing of system resources can enable reuse of system resources across multiple different workloads. For example, different workloads can be assigned to respective computing nodes of the interconnected computing system 100 to share system resources. For example, sharing of computing resources can be achieved by running CPU virtualization technologies (e.g., Kernel-based Virtual Machine (KVM), Xen, HyperV, Amazon Web Services (AWS) Firecracker, etc.) and GPU sharing technologies (e.g., Multi-Process Service (MPS), Multi-Instance GPU (MIG), etc.). Sharing of network resources can be provided by running network virtualization protocols, such as but not limited to Virtual Local Area Network (VLAN) and Virtual eXtensible Local Area Network (VXLAN). The network virtualization protocol can operate to provide a network slice for each workload, where each network slice includes different networking elements (and bandwidth sharing of each networking element) and spans the allocated computing nodes.

[0037] While the foregoing techniques can provide provisioning and allocation of system resources of the interconnected computing system 100, they may not meet the requirements in achieving optimal utilization of the system resources. For example, when the computing nodes are allocated to perform their respective tasks, the workloads assigned to the computing nodes can exhibit complex communication patterns across the interconnected computing system 100. As an illustrative example with reference to Figure 1 the figure, consider a set of N computing nodes 102, where N is an integer greater than 1, that are allocated for the execution of a workload (e.g., an ML training workload). The workload can include one or more tasks that can be executed in parallel and / or sequentially, where each task can be assigned to a different one of the N computing nodes 102. Within the N allocated computing nodes 102, there can be a subset of M computing nodes, where M is an integer equal to or less than N, that may need to exchange data to perform their respective tasks.

[0038] This communication exchange among a subset of the M computing nodes 102 can be referred to as collective communication herein. Collective communication can be used to represent various communication patterns that may occur within an interconnected computing system (such as the interconnected computing system 100) for different workloads. The implementations disclosed herein can apply collective communication according to various patterns, such as but not limited to all-to-all collectives, where multiple assigned computing nodes can transmit their data to each of the multiple assigned computing nodes; reduction collectives, where multiple assigned computing nodes can apply an aggregation function (e.g., summation) while transmitting the result value to each of the multiple assigned computing nodes; broadcast collectives, where an assigned computing node can distribute its data to multiple other assigned computing nodes; reduce collectives, where data from multiple assigned computing nodes can be combined via an aggregation function and provided to an assigned computing node; gather collectives, where data from multiple assigned computing nodes can be provided to an assigned computing node; all-gather, where data from multiple assigned computing nodes can be collected at each of the assigned computing nodes; scatter collectives, where data on an assigned computing node can be split and distributed to multiple assigned computing nodes.

[0039] As described above, when performing the above communication collectives, traditional systems may be limited to those computing nodes and communication links assigned to a particular workload. Thus, the processing of the data traffic of the workload and data transmission can be restricted to the assigned computing nodes and their direct connections therebetween. This restriction can be imposed even if other computing nodes of the interconnected system are idle or underutilized. This restriction may result in reduced system utilization, increased TCO due to idle system resources, and limitations on the expected application performance (such as but not limited to, throughput).

[0040] To illustrate the above potential performance issues, reference will be made to Figure 1 a traditional implementation. In an illustrative example, the interconnected computing system 100 includes eight first computing nodes 102 interconnected via communication links 114 and 116, which may have different speeds. For example, communication link 114 may have a first data transmission rate (e.g., 25 GB / s) and communication link 116 may have a second data transmission rate (e.g., 50 GB / s). The data transmission rates mentioned herein are for illustrative purposes only, and other data transmission rates are possible depending on the actual physical device topology used.

[0041] In this example, the controller can allocate compute nodes 102b and 102d to a workload based on a resource allocation policy. The controller can execute a set of communications between compute nodes 102b and 102d to create a logical topology. Under traditional methods, the set of communications can utilize only communication link 114a that directly connects allocated compute node 102b to allocated compute node 102d, as shown by direct communication path 120. However, there can be multiple indirect communication paths between compute node 102b and compute node 102d. At least some of these can increase the data transfer rate relative to the direct communication path. For example, indirect communication path 118 connects compute node 102b to compute node 102d via unallocated compute nodes 102c and 102a. Indirect communication path 118 includes communication links 116a, 116b, and 116c, and each of these communication links can have a higher data transfer rate than communication link 114a. Compute nodes 102c and 102a may be idle or underutilized, and traditional methods may thus miss an opportunity to improve performance in terms of data exchange speed.

[0042] Accordingly, the implementations disclosed herein provide improved performance in collective communications by leveraging any available communication links between allocated compute nodes. For example, the implementations disclosed herein can be configured to identify an optimal communication path, whether composed of direct and / or indirect communication links, to provide improved performance in accelerating workload execution.

[0043] For example, referring to the above example, the implementations disclosed herein can provide functionality that allows the use of indirect communication path 118 to construct a virtual topology for executing a workload in a manner that provides optimal utilization of system resources. For example, idle or underutilized compute nodes can be utilized to execute certain processes and forward data along high-speed communication links to allow for optimal bandwidth utilization between compute nodes. The present disclosure can implement this functionality by slicing compute nodes and running compute and forwarding services on the slices to forward and process data traffic via communication links other than those only allocated to the workload. Additional details of slicing compute nodes are provided in connection with Figure 2 and Figure 5 provide additional details of slicing compute nodes.

[0044] Figure 2 is a schematic block diagram of an example architecture of a multi-tenant collective communication fabric (MCCF) 200 in accordance with the present disclosure. MCCF 200 includes a plurality of components, and each of these components can be implemented as Figure 7 a computer system 700. MCCF 200 includes a control plane 210 and a data plane 220.

[0045] The control plane 210 may include an MCCF manager 212, which is configured to execute a resource allocation policy in the resource allocation policy 214 for each workload. The MCCF manager 212 may store the resource allocation policy 214 in a memory or other data storage device. The MCCF manager 212 may be configured to allocate one or more compute nodes (referred to as allocated compute nodes) and communication links to a specific workload according to the resource allocation policy 214. The MCCF manager 212 may also be configured to construct a virtual topology from the physical device topology of an interconnect system (e.g., the interconnect computing system 100) based on the one or more compute nodes allocated for a specific workload. The virtual topology may include multiple allocated compute nodes and direct communication links between them, as well as one or more unallocated compute nodes of the interconnect system and indirect communication links between the one or more unallocated compute nodes and the multiple allocated compute nodes. According to some examples, the MCCF manager 212 may be an example of the controller mentioned above in connection with Figure 1 the example of the controller mentioned above.

[0046] The data plane 220 includes the physical device topology of an interconnect system (such as Figure 1 the interconnect computing system 100). The data plane 220 may include multiple compute nodes 224a - 224n (collectively referred to herein as compute nodes 224) and communication links 223a - 223n (collectively referred to herein as communication links 223) of the interconnect system. The compute nodes 224 may be interconnected via a network of communication links 223 that form a communication fabric 222. The compute nodes 224 may be implemented as Figure 1 any compute node, such as the first compute nodes 102, 104, and / or 108. The communication links 223 may be implemented as Figure 1 any communication link, such as the third communication links 112, 114, and / or 116. In various examples, the data plane 220 may be wired to the MCCF manager 212 via a communication link. In some examples, the data plane 220 may be wirelessly connected to the MCCF manager 212.

[0047] According to various examples, the MCCF manager creates multiple slices at each compute node 224. For example, the MCCF manager 212 may logically partition the hardware resources (e.g., compute and memory resources) of each compute node 224 to form multiple logical slices (referred to herein as slices) of the hardware resources. As an illustrative example, the MCCF manager 212 forms a first slice 226a (sometimes referred to herein as an MCCF slice) and a second slice 228a (sometimes referred to herein as a tenant slice) by partitioning the hardware resources of the compute node 224a.

[0048] The MCCF manager 212 may assign certain functions to slices 226a and 228a. For example, the first slice 226a may be dedicated to performing non-tenant services of the MCCF (referred to herein as MCCF services). In Figure 2 the example, the MCCF services include, but are not limited to, the computing service 221a that may be configured to process data traffic for a workload, and the forwarding service 225a that may be configured to deliver the data traffic to the next computing node. In various examples, when the computing node 224a is included in the virtual topology for a particular workload and is not one of the assigned computing nodes (e.g., the computing node 224a is an unassigned computing node), the MCCF services may be performed by the computing node 224a.

[0049] The second slice 228a may be assigned by the MCCF manager 212 to the logic for performing tasks for one or more workloads assigned to the computing node 224a. In this case, the tasks assigned to the second slice 228a may be for a workload different from the particular workload for which the computing node 224a is an unassigned computing node. Thus, the computing node 224a may be able to perform operations related to multiple workloads (e.g., multiple tenants). The resources of the second slice 228 may be further divided into multiple sub-slices 227a, 229a, 230a, and 232a. These sub-slices may be assigned to different tasks for one or more workloads. In this manner, the second slice 228a of the computing node 224a may be assigned to one or more workloads, while the first slice 226a may be used for the MCCF services.

[0050] Although the foregoing examples have been described with reference to the computing node 224a, other computing nodes 224b - 224n may be similar to the computing node 224a. For example, the computing node 224b may be divided into a first slice 226b and a second slice 228b, which may be further divided into sub-slices. Similarly, the computing nodes 224c through 224n may be divided into first slices 226c through 226n and second slices 228c through 228n.

[0051] Figure 3 is a schematic block diagram of an example MCCF controller 300 according to the present disclosure. The MCCF manager 300 may be Figure 2Example of the MCCF manager 212. As described above, the MCCF manager 300 operates in the control plane of the MCCF (e.g., the control plane 210 of the MCCF 200). According to various examples, the MCCF manager 300 can receive a resource allocation request 302 from an end user, allocate one or more computing nodes to a workload based on a resource allocation policy through an allocation resource module 306, and create a virtual topology to run the workload based on the allocated computing nodes through a virtual topology module 314. According to various examples, the virtual topology can include at least one unallocated computing node and indirect communication links.

[0052] According to some examples, the MCCF manager 300 can maintain a system state 304 of an interconnected system (such as the interconnected computing system 100). The system state 304 can be stored in a database or other storage device. For example, the system state 304 can include a physical device topology 310 and a load value status 308 of each computing node and communication link of the interconnected system. As combined above Figure 1 Examples of the physical device topology are provided. The physical device topology 310 can include an identifier of each computing node, information representing the computing resources of each computing node, an identifier of each communication link, an identifier of which computing nodes are connected to each communication link, and a data transmission rate of each communication link. The load value status 308 can represent, for example, the system resource utilization of any workload (if any) allocated to the interconnected system. In some examples, the load value status 308 can include measurements of resource utilization at each computing node and communication link, such as but not limited to computing resource utilization, available computing resources, available network resources, network resource utilization, etc. The measurements can be provided as a percentage of the total amount of each respective resource. The load value status 308 can be provided as a data structure or table of load values of the computing nodes and each associated computing node.

[0053] The MCCF manager 300 can refresh and update the system state 304 through a monitoring module 312. The monitoring module 312 can be configured to monitor the data plane 320 (e.g., Figure 2 an example of the data plane 220) and obtain the current system state of the physical device topology from the interconnected system. For example, the monitoring module 312 can communicate with each computing node (e.g., computing nodes 102, 104, and / or 106) through a vendor-specific application programming interface (API) to obtain the corresponding computing and communication utilization.

[0054] Figure 3 An example workflow of the MCCF manager 300 is also depicted. The example workflow will be described below with reference to Figure 1 This. More specifically, the workflow of the MCCF manager 300 will be described for using Figure 1Create a virtual topology using the indirect communication path 118 shown as an illustrative example.

[0055] The MCCF manager 300 can receive a resource allocation request 302 from an end user. The resource allocation request 302 can contain information requesting resources for executing a workload, such as metadata. In an example, the resource allocation request 302 can include information indicating the number of compute nodes requested for the workload. In some examples, the resource allocation request 302 can include information indicating the communication standard between the compute nodes requested for the workload. The communication standard can be, for example, a minimum network bandwidth requirement.

[0056] The MCCF manager 300 can execute an allocate resources module 306 to allocate compute nodes according to the resource allocation request 302. For example, the allocate resources module 306 can refer to a resource allocation policy (e.g., Figure 2 an example of the resource allocation policy 214) and obtain the current system state from the system state 304. The allocate resources module 306 can allocate specific system resources, such as compute nodes 102b and compute nodes 102d, for the workload mentioned in the resource allocation request 302 based on the resource allocation policy and the current system state. In the above example, the resource allocation request 302 may have requested two compute nodes for a specific workload, and the allocate resources module 306 can select certain compute nodes according to the resource allocation policy and the current system state. For example, the allocate resources module 306 can maintain a list of the least utilized compute nodes and network resources. The allocate resources module 306 can then execute an algorithm to iteratively exhaustively through the list to determine the selection of compute nodes (and the available direct and indirect communication paths between them) that meet or do not meet the request. The allocate resources module 306 can then reserve compute and memory slices on each selected compute node using a resource sharing method (e.g., multi-instance GPU (MIG) known in the art, etc.).

[0057] The MCCF manager 300 may execute a virtual topology module 314 to facilitate generation of a virtual topology for executing a workload using allocated computing nodes. The virtual topology module 314 may create a virtual topology based on various objectives, examples of which are described below. For example, using the allocated computing nodes, the virtual topology module 314 may identify direct communication paths between the allocated computing nodes, the data transfer rates therebetween, and the network resource utilization of the communication paths. The virtual topology module 314 may also identify indirect communication paths between the allocated computing nodes, the data transfer rates of the communication links forming the indirect communication paths, and their network resource utilization. Then, the virtual topology module 314 may calculate an optimal communication path in terms of data transfer rate from the direct and indirect communication paths based on the objectives described below. In some examples, the virtual topology module 314 may perform graph-based analysis, such as but not limited to minimum cost maximum flow analysis, to calculate the optimal communication path(s) of the virtual topology. This calculation may also consider the computing resource utilization at unallocated computing nodes to identify underutilized or idle computing nodes that may be optimal in terms of resource utilization for performing the MCCF service.

[0058] The virtual topology module 314 may consider multiple objectives when constructing a virtual topology. Example objectives may be to balance the resource load (e.g., utilization) between the computing nodes and communication links of the interconnection system. For example, unallocated computing nodes and associated communication links may be selected to balance the system-wide load. Another example objective may be to enforce quality of service (QoS) requirements, such as but not limited to minimizing end-to-end latency, maximizing throughput, or maximizing the resolution time within a time budget. As another example objective, the virtual topology module 314 may have fine-grained control over the type of communication links to be utilized. For example, the virtual topology module 314 may decide to utilize high-speed links instead of slower links for a particular task. The above objectives are non-exhaustive examples, and the virtual topology module 314 may consider other objectives when constructing a virtual topology.

[0059] Once the virtual topology module 314 creates a virtual topology, the MCCF manager 300 may configure the physical device topology to execute the workload according to the virtual topology. For example, the MCCF manager 300 may send an allocation control 316 that includes instructions to assign resources of the physical device topology to the workload according to the virtual topology. In an illustrative example, a workload task may be assigned to a tenant slice of an allocated computing node, while an MCCF slice of an unallocated computing node may be allocated for performing the MCCF service. Additional details are provided below in conjunction with Figure 5 Provide additional details.

[0060] Figure 4Depicts an example virtual topology 400 implemented in accordance with an example of the present disclosure. The virtual topology 400 is a schematic representation of a virtual topology computed by an MCCF manager (e.g., the virtual topology module 314).

[0061] In Figure 4 the example of, the virtual topology 400 includes allocated compute nodes 402b and 402d, and unallocated compute nodes 402c and 402a. Referring to Figure 1 , the allocated compute nodes 402b and 402d can be virtual representations of compute nodes 102b and 102d, respectively. The unallocated compute nodes 402c and 402a can be virtual representations of compute nodes 102c and 102a, respectively. The unallocated compute nodes 402c and 402a can be dedicated to compute and forwarding services (e.g., MCCF services).

[0062] The virtual topology 400 further includes communication links 416a - 416c that together form an indirect communication path. The communication links 416a - 416c can be used to perform forwarding services (e.g., MCCF services). Referring to Figure 1 , the communication links 416a - 416c can be virtual representations of communication links 116a - 116c. Figure 4 Also depicted is a direct communication path 414 determined from the allocation resource module 306. The direct communication path 414 can be Figure 1 an example of communication link 114a of. As described above, communication link 116 can have a higher data transfer rate than communication link 114. Thus, a virtual topology can be created based on the indirect communication path to provide an improved data transfer rate compared to the direct communication path.

[0063] As described above, each compute node can be partitioned into slices. Figure 4 Also schematically depicted is which slice each compute node uses. For example, compute nodes 402b and 402d are depicted with a first shading pattern that indicates that compute nodes 402b and 402d are allocated to perform tasks of a workload (e.g., tenant slices perform corresponding tasks). While compute nodes 402c and 402a include a second shading pattern indicating that the MCCF slices are dedicated to MCCF services.

[0064] Each computing node 402c and 402a also includes information indicating which MCCF services are performed by the corresponding computing node. For example, computing node 402c can perform a computing service 402c-1 and a forwarding service 402c-2, and computing node 402c can perform a forwarding service. In some examples, an unassigned computing node 402c can perform services by running a computing kernel (e.g., 3:c) and a forwarding kernel (e.g., 3:f) to transfer data to an unassigned computing node 402a, while an unassigned computing node 402a can run a forwarding kernel (e.g., 1:f) to transfer data to an assigned computing node 402d. Figure 4 The computing service and the forwarding service can be examples of the computing service 221a and the workload and forwarding service 225a, respectively.

[0065] Figure 5 is a schematic block diagram of an exemplary architecture of a data plane 500 of an MCCF according to the present disclosure. The data plane 500 can be Figure 2 an exemplary implementation of the data plane 220. More specifically, the data plane 500 includes a plurality of computing nodes, exemplarily shown as a computing node 510 and a computing node 520 connected via a communication link 530. Although Figure 5 two computing nodes are depicted, any number of computing nodes can be included in the data plane 500 according to the physical device topology of the discussed interconnection system.

[0066] The computing node 510 can include a memory resource 512 and a computing resource 509 connected via an internal connection 519 (e.g., a bus line). The memory resource 512 can be logically divided into a first slice 514a of the memory resource and a second slice 514b of the memory resource. Similarly, the computing resource 509 can be divided into a first slice 516a and a second slice 516b of the computing resource. As described above, the MCCF manager can assign slices of hardware resources to run MCCF services, while the remaining hardware resources can be used for tasks of the workload. For example, the MCCF manager can assign the first slices 514a and 516a to MCCF services (e.g., MCCF slices), while the second slices 514b and 516b can be used for tenants (e.g., workloads) assigned. The MCCF manager can configure and maintain the proportion of hardware resources occupied by each slice. For example, a slice can occupy any proportion of resources from 0% to 100% of the hardware resources.

[0067] The compute node 520 may include a memory resource 522 and a compute resource 508 connected via an internal connection 519 (e.g., a bus line). The memory resource 522 may be logically partitioned into a first slice 524a of the memory resource and a second slice 524b of the memory resource. Similarly, the compute resource 509 may be partitioned into a first slice 526a and a second slice 526b of the compute resource. As described above, the MCCF manager may assign the first slices 524a and 526a to the MCCF service (e.g., the MCCF slice), while the second slices 524b and 526b may be used for assignment to workloads.

[0068] The implementations disclosed herein may be applicable to most, if not all, computing technologies. As such and for simplicity, the term "kernel" as used herein may refer to any function that the MCCF service may be configured to perform. For example, GPU virtualization technology may be used to implement the MCCF service and slices. In this case, the MCCF service may execute a Compute Unified Device Architecture (CUDA) kernel. For CPU nodes, CPU virtualization technology (e.g., KVM) may be used to implement the MCCF service and slices. In this case, the MCCF executes CPU processes and threads.

[0069] As Figure 5 shown in the example of, a second slice 516a may be provided to execute the MCCF service. The second slice 516a may be configured to execute a compute service 518a and a forwarding service 518b. Each MCCF service may include one or more pairs of input kernel queues and kernel execution engines. For example, the compute service 518a may include input kernel queues 515a - 515n (collectively referred to herein as internal kernel queues 515) and corresponding kernel execution engines 513a - 513n (collectively referred to herein as kernel execution engines 513). Similarly, the forwarding service 518b may include input kernel queues 511a - 511n (collectively referred to herein as internal kernel queues 511) and kernel execution engines 517a - 517n (collectively referred to herein as kernel execution engines 517).

[0070] The input kernel queues 515 and 511 can be configured to help absorb the variations between the data arrival rate and the service execution rate of the kernels (e.g., kernel a, kernel b, etc.). The mapping between the kernels of the workload and the given internal kernel queues 515 and 511 can be accomplished and maintained by the MCCF manager (e.g., MCCF manager 212 and / or MCCF manager 300). Once a kernel dequeues from the internal kernel queue 515 or 511, the corresponding execution engine 513 or 517 can run the dequeued kernel respectively. For example, when kernel a dequeues from the internal kernel queue 515a, the execution engine 513a can run kernel a. As another example, when kernel b dequeues from the internal kernel queue 511n, the execution engine 517n can run kernel b. Each of the kernel execution engines 513 and 517 can be a general-purpose processing unit (e.g., a processor or other computing device) or a special-purpose processing unit. The allocation and sizing of the internal kernel queues 515 and 511 and the kernel execution engines 513 and 517 can be static or dynamic.

[0071] The compute node 520 can be similar to the compute node 510 in that a second slice 526a can be provided to execute the MCCF services. The second slice 526a can be configured to execute the compute services 528a and the forwarding services 528b. The compute services 528a can include the input kernel queues 525a - 525n (collectively referred to herein as the internal kernel queues 525) and the corresponding kernel execution engines 523a - 523n (collectively referred to herein as the kernel execution engines 523). The forwarding services 528b can include the input kernel queues 521a - 521n (collectively referred to herein as the internal kernel queues 521) and the kernel execution engines 527a - 527n (collectively referred to herein as the kernel execution engines 527).

[0072] The input kernel queues 525 and 521 can be configured to help absorb the variations between the data arrival rate and the service execution rate of the kernels (e.g., kernel c, etc.). The mapping between the kernels of the workload and the given internal kernel queues 525 and 521 can be accomplished and maintained by the MCCF manager. Once a kernel dequeues from the internal kernel queue 525 or 521, the corresponding execution engine 523 or 527 can run the dequeued kernel respectively. For example, when kernel c dequeues from the internal kernel queue 521a, the execution engine 527a can run kernel c. Each of the kernel execution engines 523 and 527 can be a general-purpose processing unit (e.g., a processor or other computing device) or a special-purpose processing unit. The allocation and sizing of the internal kernel queues 525 and 521 and the kernel execution engines 523 and 527 can be static or dynamic.

[0073] As an illustrative example of the above concepts, when implemented in the context of the CUDA API, the input kernel queues 515 and 511 can be represented as CUDA streams. In this example, the MCCF service can be executed using a GPU thread block grid. In the CUDA API, the CUDA streams and the thread blocks of the grid can be dynamically configured at runtime, and kernels can be launched on the specified CUDA streams using the configured thread block grid.

[0074] As described above, the MCCF service can include a compute service and a forwarding service. The kernels of the compute service can run operations local to the assigned compute node on one or more data buffers. For example, kernel a can run the operations of the compute service on the data buffer "ptr_a" of the first slice 514a of the memory resource 512. The implementation details can be technology-specific; however, at a high level, each compute kernel can be associated with a kernel type (e.g., sum, average, maximum, etc.) and multiple input data buffers and multiple output data buffers. When the data of kernel a arrives at the compute node 510, the data can be stored in the data buffer "ptr_a", and kernel a can be placed in the internal kernel queue 515 (e.g., internal kernel queue 515a in this example). When kernel a dequeues from the internal kernel queue 515a, the kernel execution engine 513a can be configured to check the kernel type and execute the operation specified by the kernel type. The result of the operation can be saved in the data buffer "ptr_a".

[0075] In the case of the forwarding service, the kernels of the forwarding service can initiate data copying or data transfer between local data buffers and remote data buffers. For example, kernel b can be executed to copy or transfer the data saved in the data buffer "ptr_a" to the data buffer "ptr_c" of the first slice 514a of the memory resource 512. When the data of kernel b arrives at the data buffer "ptr_a", kernel b can be placed in the internal kernel queue 511 (e.g., internal kernel queue 511n in this example). When kernel b dequeues from the internal kernel queue 511n, the kernel execution engine 517n can be configured to check the kernel type and copy or transfer the data to the data buffer "ptr_b" via the communication link 530 according to the kernel type. Then the data can be saved in the data buffer "ptr_c".

[0076] As an illustrative example, consider compute nodes 510 and 520, which can be unassigned compute nodes (e.g., Figure 4compute nodes 402c and 402a). The MCCF manager can assign compute nodes 510 and 520 to the virtual topology of a workload to increase throughput and data transfer rate. The virtual topology can include compute core a and forwarding core b at compute node 510, and forwarding core c at compute node 520. Compute core a can be configured to operate on data contained in data buffer "ptr_a", for example, by adding a fixed value to each element in the buffer. Once compute core a is run by execution engine 513a, forwarding core b can be executed to copy the resulting data to data buffer "ptr_c" on compute node 520. Forwarding core b can perform data transfer 540 to copy the data to data buffer "ptr_c". Similarly, to transfer data to another compute node in the virtual topology ( Figure 5 not shown in), forwarding core c can be executed to copy data buffer "ptr_c" to the next compute node via data transfer 542.

[0077] The MCCF service according to the present disclosure can support a wide range of use cases because the MCCF service enables the MCCF manager to build any virtual topology for any communication set. Three illustrative examples are provided below as a non-exhaustive list of these use cases.

[0078] In a first example use case, assume that a workload is assigned two compute nodes that will run a full reduction set. Since these two nodes may have limited or relatively slow direct bandwidth, a virtual topology spanning multiple indirect communication links can be built to increase bandwidth (e.g., Figure 1 indirect communication path 118). In this case, the full reduction set can be executed such that each of the assigned compute nodes in the assigned compute nodes performs a reduction function (e.g., aggregation) and the resulting values can be transmitted to each of the assigned compute nodes via indirect communication path 118, which can provide an improved data transfer rate on the direct communication path.

[0079] As a second example use case, communication time can be reduced by leveraging an additional N unassigned compute nodes between the sending compute node and the receiving compute node (e.g., the assigned compute node). For example, the sending compute node can perform a split (or, partition) set that divides the data traffic among the N unassigned compute nodes. The receiving compute node can then perform a gather set to gather the partitioned data from the N unassigned compute nodes. Since the data is divided into smaller parts and these smaller parts are transmitted to the receiving compute node via the N unassigned compute nodes simultaneously, the total time between sending and receiving the data can be reduced.

[0080] In a third example use case, the MCCF service can be used as a parameter server for a portion of a data traffic flow. In this case, unallocated compute nodes can receive a small portion of the data traffic and perform local reduction via the MCCF compute service. The unallocated compute nodes can then broadcast the reduced data traffic back to the allocated compute nodes.

[0081] Figure 6 Illustrated are example computing components that can be used to implement multi-tenant collective communication in distributed computing. Now referring to Figure 6 , the computing component 600 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. In an Figure 6 example implementation of

[0082] The hardware processor 602 can be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in the machine-readable storage medium 604. The hardware processor 602 can fetch, decode, and execute instructions, such as instructions 606 - 612, to control the processes or operations of multi-tenant collective communication in distributed computing. As an alternative or supplement to retrieving and executing instructions, the hardware processor 602 can include one or more electronic circuits that include electronic components for performing the functions of one or more instructions, such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuits.

[0083] A machine-readable storage medium, such as the machine-readable storage medium 604, can be any electrical, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, the machine-readable storage medium 604 can be, for example, a random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disc, etc. In some embodiments, the machine-readable storage medium 604 can be a non-transitory storage medium, where the term "non-transitory" does not cover transitory propagated signals. As described in detail below, the machine-readable storage medium 604 can be encoded with executable instructions, such as instructions 606 - 612.

[0084] The hardware processor 602 can execute instruction 606 to assign multiple compute nodes of an interconnect system to a first workload. In some examples, instruction 606 can include assigning compute and memory slices on the assigned nodes to the first workload.

[0085] The hardware processor 602 may execute instructions 608 to obtain the topology of the interconnection system. The topology may represent indirect communication paths between multiple allocated computing nodes. The indirect communication paths may include unallocated computing nodes of the interconnection system.

[0086] The hardware processor 602 may execute instructions 610 to create multiple slices of the hardware resources of the unallocated computing nodes. The first slice of the multiple slices may be dedicated to processing and forwarding data traffic along the indirect communication path. The second slice of the multiple slices may be configured to be allocated to a second workload.

[0087] The hardware processor 602 may execute instructions 612 to execute a first workload through the multiple allocated computing nodes and the unallocated computing nodes. Data traffic from the multiple allocated computing nodes may be transmitted via the indirect communication path.

[0088] Figure 7 A block diagram of an example computer system 700 in which various embodiments described herein may be implemented is depicted. The example computer system 700 may be an example implementation of any of the components disclosed herein, such as Figure 1 the first computing node 102, the third computing node 104, the switch 106, and / or the second computing node 108; Figure 2 the MCCF manager 212 and / or the computing node 224; Figure 3 the MCCF manager 300; and Figure 5 the computing node 510 and / or the computing node 520.

[0089] The computer system 700 includes a bus 702 or other communication mechanism for conveying information, and one or more hardware processors 704 coupled to the bus 702 for processing information. The (multiple) hardware processors 704 may be, for example, one or more general-purpose microprocessors.

[0090] The computer system 700 also includes a main memory 706, such as a random access memory (RAM), a cache, and / or other dynamic storage devices, coupled to the bus 702 for storing information and instructions to be executed by the processor 704. The main memory 706 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor 704. When such instructions are stored in a storage medium accessible to the processor 704, the computer system 700 is presented as a special-purpose machine customized to perform the operations specified in the instructions.

[0091] The computer system 700 also includes a read-only memory (ROM) 708 or other static storage device coupled to the bus 702 for storing static information and instructions for the processor 704. A storage device 710, such as a magnetic disk, an optical disk, or a USB thumb drive (flash drive), is provided and coupled to the bus 702 for storing information and instructions.

[0092] The computer system 700 may be coupled to a display 712, such as a liquid crystal display (LCD) (or, a touch screen), via the bus 702 for displaying information to a computer user. An input device 714, including alphanumeric keys and other keys, is coupled to the bus 702 for communicating information and command selections to the processor 704. Another type of user input device is a cursor control 716, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 704 and for controlling cursor movement on the display 712. In some embodiments, the same direction information and command selections as cursor control may be achieved by receiving touches on a touch screen without a cursor.

[0093] The computing system 700 may include a user interface module for implementing a GUI, which may be stored in the mass storage device as executable software code executed by the (one or more) computing devices. By way of example, this module and other modules may include components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables.

[0094] In general, terms such as "component", "engine", "system", "database", "data store", etc. as used herein may refer to logic embodied in hardware or firmware, or to a collection of software instructions that may have entry and exit points and are written in a programming language such as Java, C, or C++. Software components may be compiled and linked into an executable program, installed in a dynamic link library, or may be written in an interpreted programming language (e.g., BASIC, Perl, or Python). It should be understood that software components may call from other components or from themselves, and / or may be called in response to detected events or interrupts. Software components configured to execute on a computing device may be provided on a computer-readable medium, such as a compact disc, digital video disc, flash drive, magnetic disk, or any other tangible medium, or may be provided as a digital download (and may initially be stored in a compressed or installable format that requires installation, decompression, or decryption prior to execution). Such software code may be stored, in whole or in part, on the memory device of the executing computing device for execution by the computing device. Software instructions may be embedded in firmware, such as an EPROM. It should also be understood that hardware components may be composed of connected logic units (e.g., gates and flip-flops), and / or may be composed of programmable units (e.g., programmable gate arrays or processors).

[0095] Computer system 700 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, when combined with the computer system, cause or program the computer system 700 to be a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 700 in response to one or more sequences of one or more instructions contained in main memory 706 being executed by (one or more) processors 704. Such instructions may be read into main memory 706 from another storage medium, such as storage device 710. Execution of the sequence of instructions contained in main memory 706 causes (one or more) processors 704 to perform the processing steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of or in combination with software instructions.

[0096] As used herein, the term "non-transitory medium" and like terms refer to any medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-transitory media can include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 710. Volatile media includes dynamic memory, such as main memory 706. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage media, CD-ROM, any other optical data storage media, any physical media with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, and network versions thereof.

[0097] Non-transitory media is different from transmission media, but can be used in combination with transmission media. Transmission media participates in the transfer of information between non-transitory media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that form bus 702. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0098] Computer system 700 also includes a communication interface 718 coupled to bus 702. Network interface 718 provides two-way data communication coupling to one or more network links connected to one or more local networks. For example, communication interface 718 can be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, network interface 718 can be a Local Area Network (LAN) card to provide a data communication connection to a compatible LAN (or, a WAN component communicating with a WAN). A wireless link can also be implemented. In any such implementation, network interface 718 transmits and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0099] Network links typically provide data communication to other data devices through one or more networks. For example, a network link can provide a connection to a host computer through a local network or to a data device operated by an Internet Service Provider (ISP). The ISP, in turn, provides data communication services through the global packet data communication network now commonly referred to as the "Internet". Both local networks and the Internet use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks and signals on network links and through communication interface 718 are example forms of transmission media that carry digital data to and from computer system 700.

[0100] The computer system 700 can send messages and receive data, including program code, via (a) network(s), network link, and communication interface 718. In an Internet example, a server can send code requested for an application via the Internet, an ISP, a local network, and communication interface 718.

[0101] The received code can be executed by the processor 704 when it is received, and / or stored in the storage device 710 or other non-volatile storage for later execution.

[0102] Each of the processes, methods, and algorithms described in the foregoing sections can be embodied in code components executed by one or more computer systems or computer processors including computer hardware, and be fully or partially automated by the code components. The one or more computer systems or computer processors can also operate to support the performance of related operations in a “cloud computing” environment or as “software as a service” (SaaS). These processes and algorithms can be implemented in part or in whole in dedicated circuitry. The various features and processes described above can be used independently of one another or can be combined in various ways. Different combinations and sub-combinations are intended to fall within the scope of this disclosure, and certain method or process blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any particular order, and the blocks or states associated therewith can be executed in other suitable orders, or can be executed in parallel or in some other manner. Blocks or states can be added to or deleted from the disclosed example embodiments. The performance of certain operations or processes can be distributed among computer systems or computer processors, not only residing within a single machine but deployed across multiple machines.

[0103] As used herein, circuitry can be implemented using any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms can be implemented to constitute circuitry. In implementation, the various circuits described herein can be implemented as discrete circuits, or the described functions and features can be partially or fully shared among one or more circuits. Although various features or functional elements may be described or claimed separately as discrete circuits, these features and functions can be shared among one or more common circuits, and such description should not require or imply the need for separate circuits to implement such features or functions. In cases where the circuitry is implemented in whole or in part using software, such software can be implemented to operate with a computing or processing system (e.g., computer system 700) capable of performing the functions described thereof.

[0104] As used herein, the term "or" may be construed in an inclusive or exclusive sense. In addition, the description of a singular resource, operation, or structure should not be construed as excluding a plurality. Conditional language, such as "able to", "can", "may", or "could", unless specifically stated otherwise or otherwise understood in the context in which it is used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements, and / or steps.

[0105] Unless otherwise expressly stated, the terms and phrases used in this document and their variants should be construed as open-ended rather than limiting. Adjectives such as "conventional", "traditional", "normal", "standard", "known", and terms of similar import should not be construed as limiting the item described to a given time period or to items available as of a given time, but should be understood to encompass conventional, traditional, normal, or standard techniques that may be available or known at any time, present or future. In some instances, the presence of broad words and phrases such as "one or more", "at least", "but not limited to", or other similar phrases should not be taken to imply or require the use of a narrower case where such broad phrases might be absent.

Claims

1. A method for multi-tenant set communication in distributed computing, the method comprising: assigning a plurality of computing nodes of the interconnected system to a first workload; obtaining a topology of the interconnect system, the topology representing indirect communication paths between the allocated plurality of computing nodes, wherein the indirect communication paths include unallocated computing nodes of the interconnect system; creating a plurality of slices of hardware resources of the unallocated compute node, a first slice of the plurality of slices being dedicated to processing and forwarding data traffic along the indirect communication path, and a second slice of the plurality of slices being configured for allocation to a second workload; as well as The first workload is executed by the allocated plurality of computing nodes and the unallocated computing nodes, wherein data traffic from the allocated plurality of computing nodes is transmitted via the indirect communication path. 2 . The method of claim 1 , wherein the indirect communication path comprises a plurality of communication links between the unassigned computing node and the assigned plurality of computing nodes. 3 . The method of claim 2 , wherein the plurality of communication links comprise a higher data transmission rate than communication links of direct communication paths connecting the plurality of allocated computing nodes. The method of claim 2 , wherein a direct communication path between the allocated plurality of computing nodes is not available.

5. The method according to claim 1, further comprising: receiving a resource allocation request, the resource allocation request including information requesting a number of computing nodes of the interconnected system for executing the first workload, Wherein allocating the plurality of computing nodes is based on the number of computing nodes requested. The method of claim 5 , wherein the resource allocation includes an allocated communication standard between the plurality of computing nodes.

7. The method according to claim 1, further comprising: identifying the allocated direct communication paths between the plurality of computing nodes; identifying the indirect communication path; determining that the indirect communication path includes a second data transmission rate that is optimal relative to a first data transmission rate of the direct communication path; as well as The topology is generated based on the determination.

8. The method of claim 1, wherein the interconnect system comprises a physical device topology including the plurality of computing nodes and a plurality of communication links, wherein the acquired topology is a virtual topology for executing the first workload.

9. The method of claim 1 , wherein executing the first workload comprises: transmitting a first data traffic from a first allocated computing node of the plurality of allocated computing nodes to the unallocated computing node via the indirect communication path; as well as The second data service is forwarded to a second allocated computing node among the allocated multiple computing nodes using hardware resources of the unallocated computing node corresponding to the first slice, the second data service being based on the first data service.

10. The method of claim 9, wherein executing the first workload comprises: The first data service is processed using hardware resources of the unallocated computing nodes corresponding to the first slice to generate the second data service.

11. The method according to claim 1, further comprising: Compute and memory slices on the allocated plurality of compute nodes are assigned to the first workload.

12. A system comprising: a memory configured to store instructions; as well as at least one processor communicatively coupled to the memory and configured to execute the instructions to: assigning a plurality of computing nodes of the interconnected system to a first workload; obtaining a topology of the interconnect system, the topology representing indirect communication paths between the allocated plurality of computing nodes, wherein the indirect communication paths include unallocated computing nodes of the interconnect system; creating a plurality of slices of hardware resources of the unallocated compute node, a first slice of the plurality of slices being dedicated to processing and forwarding data traffic along the indirect communication path, and a second slice of the plurality of slices being configured for allocation to a second workload; as well as The first workload is executed by the allocated plurality of computing nodes and the unallocated computing nodes, wherein data traffic from the allocated plurality of computing nodes is transmitted via the indirect communication path.

13. The system of claim 12, wherein the indirect communication path comprises a plurality of communication links between the unassigned computing nodes and the assigned plurality of computing nodes.

14. The system of claim 12, wherein the at least one processor is further configured to execute the instructions to: receiving a resource allocation request, the resource allocation request including information requesting a number of computing nodes of the interconnected system for executing the first workload, Wherein allocating the plurality of computing nodes is based on the number of computing nodes requested.

15. The system of claim 12, wherein the at least one processor is further configured to execute the instructions to: identifying the allocated direct communication paths between the plurality of computing nodes; identifying the indirect communication path; determining that the indirect communication path includes a second data transmission rate that is optimal relative to a first data transmission rate of the direct communication path; as well as The topology is generated based on the determination.

16. The system of claim 12, wherein the interconnect system comprises a physical device topology including the plurality of computing nodes and a plurality of communication links, wherein the acquired topology is a virtual topology for executing the first workload.

17. The system of claim 12, wherein the at least one processor is further configured to execute the instructions to: transmitting a first data traffic from a first allocated computing node of the plurality of allocated computing nodes to the unallocated computing node via the indirect communication path; and The second data service is forwarded to a second allocated computing node among the allocated multiple computing nodes using hardware resources of the unallocated computing node corresponding to the first slice, the second data service being based on the first data service.

18. An interconnected computing system comprising: Multiple communication links; a first plurality of computing nodes configured to be assigned to a first workload; as well as a second computing node connected to a first subset of communication links, wherein the first subset of communication links forms a first communication path including the first plurality of computing nodes and the second computing node, The second computing node includes hardware resources divided into a plurality of slices, and a first slice of the plurality of slices is configured to be dedicated to transmitting data traffic of the first workload to the first plurality of computing nodes via the first subset of communication links.

19. The interconnect system of claim 18, wherein the plurality of communication links includes a second plurality of communication links forming a second communication path, the second communication path directly connecting the first plurality of computing nodes.

20. The interconnect system of claim 18, wherein the plurality of slices includes a second slice configured to be assigned to a second workload.

Citation Information

Patent Citations

  • Task execution method and device, equipment and storage medium

    CN114024858A

  • Server GPU computing power distribution system and method and server

    CN115904699A

  • Distributed system communication scheduling method and distributed machine learning system

    CN116204327A