Computing resource cluster network architecture oriented to multi-dimensional slices
By using a 3D torus network and a nested cube topology computing resource cluster architecture, the problems of small scale and low communication efficiency of direct connection of computing resources in existing technologies are solved, realizing computing resource sharing and efficient communication for multi-domain applications.
Patent Information
- Application Number
- CN202511441887.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-02-13
AI Technical Summary
Existing advanced computing cluster network architectures have a small physical direct connection scale, low communication efficiency between computing resources, high construction costs, and are difficult to adapt to multi-domain applications and computing resource sharing.
A computing resource cluster network architecture oriented towards multi-dimensional slicing is adopted, using a 3D torus network and nested cube topology, combined with redundant connections and management control modules, to achieve efficient interconnection and management of computing resources.
It improves the communication efficiency and reliability between computing resources, supports the partitioning and sharing of computing resources for multi-domain applications, and reduces construction costs.
Smart Images

Figure CN121530872A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a kind of mimicry computing, advanced computing, heterogeneous computing technology, in particular to a kind of multi-dimensional slice-oriented computing resource cluster network architecture. BACKGROUND
[0002] With the artificial intelligence driven industry development empowerment to come explosive period, countries for seizing the new technology route highland, carried out "computing power arms race", million card cluster, super million card cluster, hundred thousand card cluster etc. Computing power infrastructure is under construction. Computing cluster computing power resources are no longer limited to general CPU computing power, GPU, ASIC, FPGA etc. High-performance computing power scale development trend far exceeds CPU, but the scale of high-performance computing power under traditional computing architecture is limited to communication and management mode, construction cost is higher, large-scale integration is more difficult. At present, large-scale high-performance computing board card cluster interconnection communication becomes the research focus in the direction of computing equipment development, computing center integration, scale computing power supply etc.
[0003] The existing advanced computing cluster communication network architecture is mainly based on CLOS-like topology, which is interconnected by high-speed connection and high-speed dedicated switch to form a large-scale cluster. With the help of high-efficiency data read-write technology such as RDMA, direct communication between high-performance computing boards is realized, and the purpose of large-scale efficient computing is achieved by breaking through the physical size limit of the chip in logic. Nvidia has researched the communication of computing boards in the cluster and proposed NVLink and NVLink Switch products. Through multi-layer NVLink Switch and NVLink high-speed interconnection between NVLink Switch and GPU, a TB / s-level high-speed interconnection communication network of 576 GPUs is realized. Tesla's Cortex supercomputing cluster uses Nvidia chips and interconnection technology to build and deploy a large-scale cluster network with more than 100,000 GPUs. Byte Toutiao uses a three-layer switch to connect in a CLOS-like topology to form a GPU cluster with more than 10,000 cards, and based on Kubenetes, it identifies and recovers faults to improve the fault tolerance of the cluster. Alibaba's HPN7 computing cluster introduces a two-layer, dual-platform network architecture, using high-bandwidth NICs and high-performance Ethernet switches to build a single-pod GPU interconnection with more than 10,000 cards. Each pod is interconnected by high-performance Ethernet switches to 15 segments, each segment is interconnected by 16 ToR switches through high-bandwidth NICs to 136 Host hosts, and each Host host has 8 GPU computing boards, forming a GPU cluster of 1024 main computing nodes + 64 backup computing nodes in each segment. Huawei proposes UB-Mesh network architecture and nD-FullMesh network topology, which directly connects 8 NPUs in the server backplane in 1D-FullMesh, and interconnects the servers in the same rack in 2D-FullMesh. Each row of racks and each room is interconnected by switches recursively in FullMesh. The communication equipment used is self-developed by Huawei, achieving the purpose of building a 10,000-card cluster. However, these cluster solutions require a large number of high-performance switches and other interconnection devices, with high construction cost and reduced reliability as the cluster size increases. The efficiency of large-scale deployment of computing application tasks is reduced, and it is difficult to reuse and share resources among application tasks. In addition, existing cluster solutions are only for accelerating artificial intelligence, neural network, and deep learning computing algorithms, and have certain limitations in multi-scene application in the field.
[0004] In the research of new high-speed interconnection network technology, due to the closed source of Nvidia NVLink and NVLink Switch technology product system, AMD, Google, Microsoft and other companies have established UVLink Alliance to achieve single-channel 200Gbps transmission rate and 1024 computing accelerator / high-performance computing board interconnection. Linux Foundation, Meta, Microsoft, AMD and other companies have established Super Ethernet Alliance to achieve million-node computing cluster interconnection network based on RDMA technology on the basis of Ethernet. The above alliances are still in the initial stage, and the standards, technologies, and open-source reusable tools and software libraries need to be further improved. Domestic new high-speed interconnection network technology is also following up. Tsinghua University team[4] proposed an intelligent optical computing chip cluster architecture and system, which builds optical paths between multiple optical computing chips for high-speed optical communication through spatial phase modulation and time-domain intensity adjustment. However, this technology is only suitable for optical technology systems, and the network communication problem solved is concentrated in the physical layer interconnection. Huawei proposed a message transmission, network equipment and computing equipment cluster in the field of cloud computing[5], but this technology mainly solves the message transmission load sharing demand in cluster networking, focusing on data link layer protocol and switch development. Most of the above technologies are still in the initial stage of technical demonstration and research, and more research is needed in the aspects of technical system improvement and implementation verification.
[0005] In summary, high-speed interconnection communication has become a bottleneck restricting the construction of advanced computing infrastructure, and the research of computing resource interconnection, sharing and reliability guarantee in large-scale clusters has become the key to breaking the bottleneck. For advanced computing technology theories such as heterogeneous computing and analog computing, the construction of large-scale cluster network and cluster resource sharing management are particularly important. Especially for the research of advanced computing cluster interconnection network for field efficient computing, it can lay a practical foundation for the supply of computing power resources and industrial application.
[0006] [1] Ziheng Jiang, Haibin Lin, etc. 2024. MegaScale: Scaling LargeLanguage Model Training to More Than 10,000 GPUs. arXiv:2402.15627. https: / / doi.org / 10.48550 / arXiv.2402.15627 [2] Kun Qian, Yongqing Xi, etc. 2024. Alibaba HPN: A Data CenterNetwork for Large Language Model Training. In Proceedings of the ACM SIGCOMM2024 Conference (ACM SIGCOMM '24). Association for Computing Machinery, NewYork, NY, USA, 691–706. https: / / doi.org / 10.1145 / 3651890.3672265 [3]Heng Liao, Bingyang Liu, etc. 2025. UB-Mesh: a HierarchicallyLocalized nD-FullMesh Datacenter Network Architecture. arXiv:2503.20377.https: / / doi.org / 10.48550 / arXiv.2503.20377 [4] Tsinghua University. Intelligent optical computing chip cluster architecture and system: 202510379745.1[P]. 2025-06-13. [5] Shenzhen Huawei Cloud Computing Technology Co., Ltd. Message Transmission Methods, Systems, Network Equipment and Computing Equipment Clusters: 202510659658.1 [P]. 2025-07-25. Existing large-scale advanced computing high-efficiency computing cluster interconnection networks mainly consist of high-speed optical fibers and high-performance switching equipment. The network topology they use is usually a Clos-like architecture, which uses multi-level switching equipment to build a switching network to interconnect computing devices. Several technical problems exist in the large-scale construction of computing resource clusters such as domain-general high-efficiency computing / heterogeneous computing and mimicry computing based on heterogeneous resource hardware and software collaborative variable structure computing. The network architecture of existing advanced computing clusters is mainly designed for artificial intelligence algorithms. It adopts a multi-layered architecture similar to neural networks to interconnect computing resources at multiple levels. It is difficult to adapt to efficient computing for other application tasks and is difficult to apply to fields such as variable structure computing and heterogeneous computing.
[0007] Currently, advanced computing cluster network architectures and management methods are mainly used for scenarios such as training large artificial intelligence models, making it difficult to partition computing resources for multi-scenario applications and share computing resources among application tasks.
[0008] The existing network architecture of advanced computing clusters relies on a large number of high-performance switching devices, resulting in high construction costs. The constructed computing network is an abstract network with a small physical direct connection scale, and multi-level routing and forwarding reduces the communication efficiency between computing resources. Summary of the Invention
[0009] To address the issues of small physical direct connection scale and low communication efficiency between computing resources in existing advanced computing cluster network architectures, a computing resource cluster network architecture oriented towards multi-dimensional slicing is proposed.
[0010] The technical solution of this invention is as follows: A computing resource cluster network architecture for multi-dimensional slicing includes a high-speed interconnection network for high-efficiency computing nodes / accelerated computing node clusters, computing node drivers for node data communication forwarding, and cluster network control devices and management control modules. The topology of the high-speed interconnection network of high-efficiency computing nodes / accelerated computing node clusters is a generalized hypercube—3D torus network composed of cluster network units based on cubic topology; Compute node drivers are used for data transmission between directly connected physical compute nodes; The cluster network control device performs topology discovery and monitoring of the high-speed interconnection network, and controls and manages the computing nodes in the high-speed interconnection network according to the application requirements of the domain scenario and the resource scheduling results, thereby improving the application scope and resource utilization of the high-speed interconnection network. At the same time, the cluster network control device performs congestion discovery and fault monitoring according to the status of the high-speed interconnection network, and uses redundant connections for processing to form a corresponding data forwarding scheme for configuration and enabling. The management and control module periodically or passively performs status detection on computing nodes or communication paths based on the threshold settings for collecting communication data from computing nodes, diagnoses potential faulty nodes or connections, and uploads the relevant information to the cluster network control device.
[0011] Furthermore, the specific structure of the 3D torus network is described as follows: A cube is formed by eight adjacent high-efficiency computing nodes / accelerated computing nodes, serving as the basic cluster network unit. Three of the four interconnect ports of each computing node are used for building the cube network topology, and the remaining port is used as the external connection port of the basic cluster network unit. Multiple basic cluster network units are then interconnected to form a 3D torus network. Each basic cluster network unit selects six of the eight high-efficiency computing nodes / accelerated computing nodes, and uses the remaining port on each of these six high-efficiency computing nodes / accelerated computing nodes to perform torus network interconnection in six directions: front-back, up-down, and left-right. The remaining two ports on the two high-efficiency computing nodes / accelerated computing nodes are used for redundant interconnection.
[0012] Furthermore, in each sub-cube of the 3D torus, the two furthest endpoints are connected using the remaining two ports on each basic cluster network unit for redundant connections.
[0013] Furthermore, in the high-speed interconnect network of the cluster, each computing node can be identified by a 6-tuple: {x,y,z,a,b,c}, where x / y / z are the coordinates of the basic cluster network unit in the 3D torus network, and a / b / c are the coordinates of the computing node within the basic cluster network unit. This high-speed interconnect network of the cluster can use the four ports on each computing node to extend the simple cubic structure of the basic cluster network unit into a generalized hypercube structure. The two-layer nested cubic structure has shorter addressing and communication paths when the cluster size is large, and combined with redundant interconnects, it can provide efficient and reliable links.
[0014] Furthermore, the compute node driver is specifically as follows: A compute node includes a compute device and its driver software. The computing power required by the driver software is supplied to compute devices including FPGAs and ASICs, dedicated compute chips such as onboard Arm and MCUs, and the CPU computing power of the server. The compute node driver is used for high-speed data forwarding and data transmission for data unpacking / packing between directly connected physical compute nodes. It supports RDMA functionality and has data filtering configuration capabilities. Through the compute node driver, each compute node in the basic cluster network unit realizes direct and efficient data communication between nodes.
[0015] Furthermore, the cluster network control device is specifically configured as follows: The cluster network control device connects to each bearer server / backplane via a low-speed Ethernet switch to form a management network. For multi-scenario applications, the cluster network control device first obtains information about each computing node from the management control module, perceiving the computing node topology. Then, based on the number of applications, it divides several parallel planes from the 3D torus network cube structure for coarse-grained computing resource allocation. Next, based on the application scale, it starts fine-grained resource allocation from the same parallel coordinates on the plane, prioritizing the allocation of application resources to fill a sub-cube of a basic cluster network unit to improve resource utilization and communication distance. Based on the allocated computing resource coordinates, it performs route planning for each application and distributes the routing methods and constraints to each computing node in the form of a routing table through the management control module, enabling the computing nodes to perform data forwarding and filtering operations.
[0016] Furthermore, the specific operations after the management and control module uploads relevant information to the cluster network control device are as follows: The cluster network control device updates and recalculates the corresponding computing node status information and topology, finds the nearest computing node with the same function for node sharing, and updates and distributes the route plan; the computing node attempts to retransmit data based on the new route information to restore communication; if it is impossible to find a computing node with the same function, especially if the computing node is a reconfigurable mimic computing node, the user is notified to reconfigure the computing node and migrate the application locally, and then the above operations are performed again to update the computing node topology and distribute the route replanning.
[0017] Furthermore, the data transmission is as follows: When the computing nodes in the cluster communicate, the following steps are followed for data transmission: Step (1) First, according to the routing and forwarding algorithm in the computing node driver, the data is forwarded to the computing node with the corresponding external connection, i.e., the internal forwarding node, in the sub-cube of the basic cluster network unit; Step (2) Then, the data is forwarded to the computing node of the next hop basic network unit, i.e., the external forwarding node, through the cluster interconnection for data reception; Step (3) The external forwarding node determines whether the coordinates of the basic network unit it is in is the coordinates of the destination node. If the coordinates of the basic network unit are the coordinates of the destination node, then internal forwarding is performed in the sub-cube of the unit to forward the data to the target computing node; if the coordinates of the basic network unit are not the coordinates of the destination node, then steps (1) and (2) are iterated.
[0018] Furthermore, the computing nodes are integrated by the host server / backplane, and the host server / backplane is equipped with cluster network control equipment and management control modules.
[0019] Preferably, the specific implementation is as follows: it is divided into cluster network control equipment, bearer server / backplane, management control module, computing nodes, and cluster high-speed interconnection network; The cluster network control device includes a management server and a multi-port Ethernet switch. The management server is equipped with cluster management tools / platforms to perform cluster resource allocation and computing resource monitoring; or it can implement cluster resource monitoring, diagnosis, allocation and management functions based on secondary development or self-developed cluster management tools / platforms. The computing node topology generation and routing planning use the form of static routing tables. In addition to generating physical topology and routing plans, it also includes generating corresponding local resource topology and routing plans for each application based on these. The host server / backplane is a multi-PCIe expansion bit computing server, and the connection between the server and the cluster network control device is an Ethernet cable connection; each computing node is powered, monitored for status, updated for drivers, and distributed for reconfigurable images through the host server / backplane. The management and control module is based on Cyborg and is a secondary development to realize operations such as monitoring the status of each computing node in the server, managing and distributing route planning, updating and configuring computing node drivers, identifying nodes and connections and diagnosing faults; and uploading the acquired information to the cluster network control device. Computing nodes include physical computing devices and computing node drivers. For GPU-accelerated clusters, an adapted GPU is selected, the corresponding GPU driver is installed and configured, and a computing node driver layer is added between the management and control module and the GPU driver to implement data forwarding functionality. For mimicry computing clusters, reconfigurable mimicry computing nodes can be implemented by FPGA integration, and computing node drivers can be implemented based on the corresponding FPGA driver and open-source / self-developed communication management IP. The high-speed interconnect network of the cluster is composed of four 25G full-duplex optical modules on each computing node interconnected by optical fiber. The interconnection method uses a 6D network formed by nested cubes. When the cluster is expanded, optical fiber is used for interconnection according to the 6D network structure.
[0020] The beneficial effects of this invention are as follows: The 6D network formed by the nested cubes proposed in this invention can perform slice-based network partitioning, which facilitates the logical generation of multi-level, multi-interconnection structure of computing resources interconnection, and is suitable for the allocation and management of computing resources such as neural network accelerated computing, mimicry application accelerated computing, and heterogeneous computing.
[0021] By using a nested cube 6D network architecture, computing resources can be allocated and managed in a fine-grained manner within the server and in a coarse-grained manner within the basic cluster network unit. Computing resources can be allocated to different domain applications based on slices of different dimensions, realizing the partitioning of computing resources for multi-scenario applications in a domain and the sharing of computing resources among application tasks.
[0022] By directly connecting computing resources through a nested 6D cube network architecture, the overall communication distance of the cluster is shortened. Combined with DMA / RDMA technology, the communication efficiency between computing nodes can be effectively improved. Furthermore, the reliability of the cluster communication link is enhanced through multiple redundant torus loops spanning multiple dimensions. Attached Figure Description
[0023] Figure 1 This is a diagram of the cluster network architecture of the present invention; Figure 2 This is a network topology diagram of the cluster in this invention. Detailed Implementation
[0024] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0025] Addressing the needs of multi-scenario application deployment, efficient computing resource sharing, and large-scale direct connection of computing resources for advanced computing resources, this invention supports efficient computing in multi-scenario applications, partitioning and sharing of physical / logical computing resources, and efficient interconnection and communication of computing resources. It includes a high-speed interconnection network for high-efficiency computing nodes / accelerated computing node clusters, computing node drivers for node data communication forwarding, and cluster network control devices and management control modules (such as…). Figure 1 (As shown).
[0026] The high-speed interconnect network topology of high-efficiency computing nodes / accelerated computing node clusters is a 6D network formed by nested cubes, that is, a generalized hypercube—3D torus network—composed of cluster network units based on cube topology (such as...). Figure 2 (As shown).
[0027] The network consists of eight adjacent high-efficiency computing nodes / accelerated computing nodes forming a cube, which serves as the basic cluster network unit. Three of the four interconnect ports on each computing node are used for building the cube network topology, and the remaining port serves as the external connection port for the basic cluster network unit. Therefore, each basic cluster network unit has eight interconnected computing nodes and eight external interconnect ports.
[0028] Multiple basic cluster network units are interconnected to form a 3D torus network. Each basic cluster network unit selects 6 out of 8 high-efficiency computing nodes / accelerated computing nodes. The remaining 1 port on each of these 6 high-efficiency computing nodes / accelerated computing nodes (6 nodes × 1 port, a total of 6 ports) is used to interconnect the torus network in six directions: front-back, up-down, and left-right. The remaining 2 ports on the 2 high-efficiency computing nodes / accelerated computing nodes are used for redundant interconnection to improve cluster reliability and congestion handling efficiency during resource sharing.
[0029] To make it easier to understand, here is a brief explanation: One basic cluster network unit = 8 compute nodes; One compute node has 4 ports: 3 internal ports + 1 external port. One basic cluster network unit = 8 compute nodes = 8*1, with a total of 8 external ports; Six of the eight external ports can be selected for external interconnection in six directions: front, back, left, right, up, and down. The remaining two ports also need to be connected to the remaining two external ports of another basic cluster network unit to form two redundant connections (requiring four ports). Redundant connections are formed to create links, used for fault-tolerant routing and as temporary communication connections in situations like fault rerouting. Redundant connections aren't simply reconnected when not in use; they are established beforehand and switched at the routing level when needed. Each connection requires two ports. The remaining two ports on each basic cluster network unit are used for one connection to the lower left and one to the upper right. Each aspect forms a loop, creating a torus structure.
[0030] In each sub-cube of the 3D torus, the two furthest endpoints are connected using the remaining two ports on each basic cluster network unit for redundant connections. For example, the (0,0,0) endpoint and the (3,3,3) endpoint of each sub-cube are connected. The redundant interconnections in the entire 3D torus network also form a torus loop.
[0031] The basic cluster network unit, formed by interconnecting computing nodes in the above manner, enables large-scale interconnection of computing nodes without relying on a large number of multi-port, high-speed, and high-performance switches for networking, providing a foundation for direct data transmission between computing nodes. In the high-speed interconnection network, each computing node can be identified by a 6-tuple: {x,y,z,a,b,c}, where x / y / z are the coordinates of the basic cluster network unit in the 3D torus network, and a / b / c are the coordinates of the computing node within the basic cluster network unit. This high-speed interconnection network can utilize the four ports on each computing node, extending the simple cubic structure of the basic cluster network unit into a generalized hypercube structure. The two-layer nested cube structure provides shorter addressing and communication paths when the cluster size is large, and combined with redundant interconnection, it can provide efficient and reliable links.
[0032] Computing nodes include computing devices such as GPUs and FPGAs, along with their driver software. The computing power required for this driver software can be supplied by computing devices such as FPGAs and ASICs, or by dedicated computing chips such as onboard Arm and MCUs, or even by leveraging the CPU power of the hosting server. The computing node driver is primarily used for high-speed data forwarding, data unpacking / repackaging, and other data transmission between directly connected physical computing nodes. It supports RDMA (Remote Direct Memory Access) and has configuration capabilities such as data filtering. Through the computing node driver, each computing node in the basic cluster network unit achieves direct and efficient data communication between nodes.
[0033] Compute nodes are typically integrated within a host server / backplane. Currently, mainstream servers can generally support eight high-efficiency or accelerated compute nodes. The host server / backplane usually has independent computing units such as CPUs. To support high-speed interconnection of compute node clusters, cluster network control devices and management control modules need to be installed and deployed on the host server / backplane. These modules monitor the status of all compute nodes on the server, update and configure compute node drivers, and issue data forwarding rules to manage the communication and status of the compute nodes.
[0034] Cluster network control devices are used for status monitoring and management of the entire computing cluster. They typically connect to each host server / backplane via Ethernet switches using low-speed connections, forming a management network. These devices can discover and monitor the high-speed interconnect network topology and, based on application requirements and resource scheduling, logically partition, network, and control the computing nodes within the high-speed interconnect network, improving its application scope and resource utilization. Simultaneously, they perform congestion detection and fault monitoring based on the high-speed interconnect network status, utilizing redundant connections to process data and develop corresponding data forwarding schemes for configuration and enabling. Because the communication data volume of cluster network control devices is relatively small and does not require high-speed network connections, their construction costs are low.
[0035] For multi-scenario applications, the cluster network control device first obtains information about each computing node from the management control module to perceive the computing node topology. Then, based on the number of applications, it divides several parallel planes from the 3D torus network cube structure for coarse-grained computing resource allocation. Next, based on the application scale, it starts fine-grained resource allocation from the same parallel coordinates on the plane, prioritizing the allocation of application resources to fill a sub-cube of a basic cluster network unit to improve resource utilization and communication distance. Based on the allocated computing resource coordinates, it performs route planning for each application and distributes the routing methods and constraints to each computing node through the management control module in the form of routing tables, etc., for the computing nodes to perform data forwarding, filtering, and other operations.
[0036] When computing nodes communicate in a cluster, data transmission is performed according to the following steps: (1) First, according to the routing and forwarding algorithm in the computing node (sending node) driver, the data is forwarded to the computing node (internal forwarding node) with the corresponding external connection in the sub-cube of the basic cluster network unit. (2) Then, the data is forwarded to the computing node (external forwarding node) of the next hop basic network unit through the cluster interconnection for data reception. (3) The external forwarding node determines whether the coordinates of the basic network unit it is in is the coordinates of the destination node. If the coordinates of the basic network unit are the coordinates of the destination node, internal forwarding is performed in the sub-cube of the unit to forward the data to the target computing node; if the coordinates of the basic network unit are not the coordinates of the destination node, steps (1) and (2) are iterated.
[0037] The management and control module can periodically or passively perform status checks on compute nodes or communication paths based on compute node communication data collection thresholds, diagnose potential faulty nodes or connections, and upload relevant information to the cluster network control device. The cluster network control device updates and recalculates the corresponding compute node status information and topology, finding the nearest comparable compute nodes with the shortest distance for node sharing, and updating and distributing the updated routing plan. The compute nodes attempt data retransmission based on the new routing information to restore communication. If a comparable compute node cannot be found, especially if the compute node is a reconfigurable mimicry compute node, the user is notified to reconfigure the compute node and migrate the application locally, before the aforementioned operations of updating the compute node topology and re-planning the routing plan are performed again.
[0038] The embodiments of this invention are correspondingly divided into cluster network control equipment, bearer server / backplane, management control module, computing nodes, and cluster high-speed interconnection network. Under current mainstream commercial equipment and technology conditions, wherein: The cluster network control device can include a management server and a multi-port Ethernet switch. The management server runs cluster management tools / platforms, such as OpenStack, for cluster resource allocation and computing resource monitoring. Alternatively, it can be based on OpenStack or Kubernetes with secondary development or a self-developed cluster management tool / platform to implement functions such as cluster resource monitoring, diagnosis, and allocation management. Compute node topology generation and routing planning can use static routing tables. In addition to generating physical topology and routing plans, it can also generate corresponding local resource topology and routing plans for each application.
[0039] The host server / backplane can be a mainstream multi-PCIe expansion bit computing server, and the connection between the server and the cluster network control device can be a network cable connection. Each computing node performs operations such as power supply, status monitoring, driver updates, and reconfigurable image distribution through the host server / backplane.
[0040] The management and control module can be further developed based on Cyborg to perform operations such as monitoring the status of each compute node in the server, managing and distributing routing plans, updating and configuring compute node drivers, identifying nodes and connections, and diagnosing faults. It can also upload the acquired information to the cluster network control device.
[0041] Compute nodes consist of physical computing devices and compute node drivers. For GPU-accelerated clusters, commercial GPUs such as NVIDIA can be used, and the corresponding GPU drivers can be installed and configured. A compute node driver layer is added between the management and control module and the GPU driver to implement functions such as data forwarding. For mimicry computing clusters, reconfigurable mimicry computing nodes can be implemented using FPGA integration. The compute node drivers can be implemented based on corresponding FPGA drivers and open-source / self-developed communication management IPs, such as basic RDMA / DMA drivers, custom data formats, and forwarding algorithms.
[0042] The high-speed interconnect network of the cluster mainly uses optical fibers to interconnect four 25G full-duplex optical modules on each computing node (including GPU-accelerated clusters and mimicry computing clusters). The interconnection method uses a 6D network formed by nested cubes proposed in this invention. When the cluster is expanded, it is only necessary to use optical fibers for interconnection according to the 6D network structure of this invention.
[0043] Key technical points: 1. The nested cube 6D network architecture for direct connection to computing resources shortens the average communication distance of large-scale clusters. Combined with node / connection fault diagnosis and location and link recovery based on redundant Torus loops, it ensures the efficiency and reliability of cluster connections.
[0044] 2. Based on multi-dimensional slicing, the allocation and sharing of computing resources supports multi-domain application structure adaptation and sharing of common computing resources among domain applications. The multi-dimensional slicing method simplifies resource allocation and sharing strategies, and realizes the partitioning of computing resources for multi-scenario applications and the sharing of computing resources among application tasks.
[0045] The above-described embodiments are merely one implementation of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention should be determined by the appended claims.
Claims
1. A multi-dimensional slice oriented computing resource cluster network architecture, characterized in that, The high-efficiency computing node / accelerated computing node cluster high-speed interconnection network, a computing node driver for node data communication forwarding, and a cluster network control device and a management control module are included. The topology architecture of the high-efficiency computing node / accelerated computing node cluster high-speed interconnection network is a generalized hypercube, i.e., a 3D torus network, composed of cubic topology-based cluster network units. The computing node driver is used for data transmission of each directly connected physical computing node. The cluster network control device performs cluster high-speed interconnection network topology discovery and monitoring, and controls and manages the computing nodes in the cluster high-speed interconnection network according to the field scenario application requirements and resource scheduling results. Meanwhile, the cluster network control device performs congestion discovery and fault monitoring according to the cluster high-speed interconnection network state, and processes using redundant connections to form a corresponding data forwarding scheme for configuration and enablement. The management control module actively or passively detects the state of the computing nodes or communication paths according to the computing node communication data collection threshold setting, diagnoses potential fault nodes or connections, and uploads relevant information to the cluster network control device.
2. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, The specific structure of the 3D torus network is as follows: eight adjacent high-efficiency computing nodes / accelerated computing nodes form a cube as a basic cluster network unit, and three of the four interconnection ports of each computing node are used for cubic network topology construction, and the remaining one port is used as an external connection port of the basic cluster network unit. Multiple basic cluster network units are connected to each other to form a 3D torus network, and six of the eight high-efficiency computing nodes / accelerated computing nodes in each basic cluster network unit are selected to use the remaining one port in each of the six high-efficiency computing nodes / accelerated computing nodes for torus network interconnection in the front-back-up-down-left six directions, and the remaining two ports in the remaining two high-efficiency computing nodes / accelerated computing nodes are used for redundant interconnection.
3. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 2, wherein, In each sub-cube in the 3D torus, the remaining two ports for redundant connection in each basic cluster network unit are used to connect two farthest endpoints.
4. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, In the cluster high-speed interconnection network, each computing node can be identified by a 6-tuple: {x, y, z, a, b, c}, where x / y / z is the coordinate position of the basic cluster network unit where the computing node is located in the 3D torus network, and a / b / c is the coordinate position of the computing node in the basic cluster network unit. The cluster high-speed interconnection network can use the four ports on each computing node to expand the simple cube structure of the basic cluster network unit to a generalized hypercube structure; the two-layer nested cube structure has a shorter addressing communication path when the cluster size is large, and in combination with redundant interconnection, it can provide efficient and reliable links.
5. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, The computing node driver is as follows: the computing node includes a computing device and its driver software, the driver software required for the computing device including FPGA, ASIC, dedicated computing chip supply of on-board Arm, MCU, and CPU computing power of a server. The computing node drive is used for data high-speed forwarding, data unpacking / packing of each direct-connection physical computing node, supports RDMA function, and has function configuration ability of data filtering; each computing node in the basic cluster network unit realizes direct and efficient data communication between nodes through the computing node drive.
6. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, The cluster network control device is connected with each bearing server / backplane through an Ethernet switch to form a management network; for field multi-scene application, the cluster network control device firstly acquires computing node information according to the management control module, senses computing node topology, then cuts a plurality of parallel planes from a 3D torus network cube structure according to the number of applications to perform coarse-grained computing resource division, then performs fine-grained resource allocation from the same parallel coordinates on the plane, and preferentially allocates application resources to fill a sub-cube of the basic cluster network unit to improve resource utilization and resource communication distance; the allocated computing resource coordinates are used for routing planning of each application, and the routing mode and constraint are issued to each computing node in the form of a routing table through the management control module for data forwarding and filtering operation of the computing node.
7. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, The specific operation of the management control module after uploading the related information to the cluster network control device is as follows: the cluster network control device updates and recalculates the planning of the corresponding computing node state information and topology, finds the nearest same-function computing node for node sharing, and updates and issues the routing planning; The computing node attempts data retransmission according to the new routing information to restore communication; If the same-function computing node cannot be found, especially when the computing node is a reconfigurable quasiparticle computing node, the user is notified to perform computing node reconstruction and application local migration, and then the above operation is performed to update the computing node topology and issue the routing re-planning.
8. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, The data transmission is as follows: when the computing nodes in the cluster communicate, the following steps are followed for data transmission: step (1) firstly, according to the routing forwarding algorithm in the computing node drive, the data is forwarded to the computing node with corresponding external connection in the sub-cube of the basic cluster network unit, i.e. the internal forwarding node; step (2) then, the data is forwarded to the computing node of the next hop basic network unit, i.e. the external forwarding node, through the cluster interconnection for data reception; step (3) the external forwarding node determines whether the coordinate of the basic network unit is the destination node coordinate, if the coordinate of the basic network unit is the destination node coordinate, the internal forwarding is performed in the sub-cube of the unit to forward the data to the target computing node; if the coordinate of the basic network unit is not the destination node coordinate, the steps (1) and (2) are iterated.
9. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, The computing node is integrated by the bearing server / backplane, and the cluster network control device and the management control module are installed and deployed on the bearing server / backplane.
10. The multi-dimensional slice oriented computing resource cluster network architecture according to claim 1, wherein, The specific implementation is as follows: the cluster network control device, the bearing server / backplane, the management control module, the computing node, and the cluster high-speed interconnection network part are divided. The cluster network control device comprises a management server and a multi-port Ethernet switch, wherein the management server installs a cluster management tool / platform to perform cluster resource allocation and computing resource monitoring; or a secondary development or self-developed cluster management tool / platform is used to realize the functions of cluster resource monitoring, diagnosis, allocation management; the computing node topology generation and routing planning use a static routing table form, in addition to generating a physical topology and routing planning, it also includes generating a corresponding local resource topology and routing planning for each application on this basis; The carrying server / backplane selects a multi-PCIe expansion bit computing server, and an Ethernet network cable is used to connect the server and the cluster network control device; Each computing node performs power supply, state monitoring, driver update, and reconfigurable image distribution operation through the carrying server / backplane; The management control module is based on Cyborg for secondary development to realize the operations of computing node state monitoring, routing planning management and distribution, computing node driver update and configuration, node and connection identification and fault diagnosis in the server; The obtained information is uploaded to the cluster network control device; The computing node comprises a physical computing device and a computing node driver; for a GPU accelerated cluster, an adaptive GPU is selected, a corresponding GPU driver is installed and configured, and a computing node driver layer is added between the management control module and the GPU driver to realize the function of data forwarding; for a quasiparticle computing cluster, a reconfigurable quasiparticle computing node can be realized by FPGA integration, and the computing node driver can be realized based on a corresponding FPGA driver and an open source / self-developed communication management IP; The cluster high-speed interconnection network uses optical fibers to interconnect four 25G full duplex optical modules on each computing node to form a 6D network formed by a nested cube; when the cluster is expanded, optical fibers are used for interconnection according to the 6D network structure.
Citation Information
Patent Citations
Intelligent optical computing chip cluster architecture and system
CN120146128B
Message transmission method and system, network equipment and computing equipment cluster
CN120200962A