Wide-area computing power cascade network architecture based on hierarchical Dragonfly + Fat-Tree

By using a hierarchical cascaded network architecture, utilizing tree topology and direct chip connection technology, computing chips are connected to form an ultra-large-scale network, which solves the limitations of traditional architectures in large-scale computing and data transmission, and achieves efficient communication and scalability.

CN121940336APending Publication Date: 2026-04-28HAINAN SHILIAN ZHIXIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HAINAN SHILIAN ZHIXIN TECHNOLOGY CO LTD
Filing Date
2025-12-22
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional data center network architectures cannot meet the demands of handling large-scale concurrent computing and massive data transmission. How to extend local computing power networks to ultra-large-scale network architectures remains an urgent problem to be solved.

Method used

A hierarchical cascaded network architecture is adopted, which connects computing chips to form a computing cluster through a tree topology and connects to the first-level autonomous network architecture in the region, forming a topology that is vertically layered and horizontally decoupled. High bandwidth and low latency communication are achieved by using chip direct connection technology.

Benefits of technology

It ensures communication speed in ultra-large-scale network architecture, reduces backbone network pressure, achieves microsecond-level collaboration, and improves communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940336A_ABST
    Figure CN121940336A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a wide-area computing power cascade network architecture based on hierarchical Dragonfly + Fat-Tree, and the network architecture comprises at least one computing power cluster which comprises a plurality of computing power chips in communication connection; the at least one first network architecture comprises a plurality of first nodes, and the first nodes in the same first network architecture are in communication connection to form a tree topology structure; the first-level autonomous architecture comprises a plurality of second nodes, the second nodes in the same first-level autonomous architecture are in communication connection to form a direct connection topological structure, and different first-level autonomous architectures belong to different areas; wherein at least one computing power cluster is in communication connection with a first-stage autonomous architecture through a first network architecture, and the first-stage autonomous architecture is connected with the computing power cluster in an area to which the first-stage autonomous architecture belongs. Therefore, according to the embodiment of the invention, a super-large-scale network architecture can be constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a hierarchical cascaded network architecture. Background Technology

[0002] With the explosive growth in demand for high-performance computing, such as artificial intelligence and big data analytics, the need for large-scale concurrent computing and massive data transmission is increasing. Traditional data center network architectures are no longer sufficient to handle these demands. Against this backdrop, the industry has begun exploring new network architectures to meet this growing need.

[0003] Among these, network architectures based on chip-to-chip (CPC) technology have attracted widespread attention. This architecture can form local computing networks through computing chips, possessing potential for high bandwidth and low latency. However, how to scale up such local networks to ultra-large-scale network architectures remains a problem that urgently needs to be solved. Summary of the Invention

[0004] The embodiments of this application provide a wide-area computing power cascaded network architecture based on a hierarchical Dragonfly+Fat-Tree to construct an ultra-large-scale network architecture.

[0005] In a first aspect, embodiments of this application provide a hierarchical cascaded network architecture, including: At least one computing power cluster, the computing power cluster comprising multiple computing power chips connected in communication; At least one first network architecture, the first network architecture including multiple first nodes, and the first nodes in the same first network architecture communicate to form a tree topology; At least one first-level autonomous architecture, wherein the first-level autonomous architecture includes multiple second nodes, and the second nodes in the same first-level autonomous architecture communicate to form a direct connection topology, and different first-level autonomous architectures belong to different regions; In this configuration, at least one of the computing power clusters communicates with a first-level autonomous architecture through a first network architecture, and the first-level autonomous architecture connects to computing power clusters within its respective region.

[0006] In some embodiments of this application, a hierarchical cascaded network architecture is provided, comprising: at least one computing power cluster, at least one first network architecture, and at least one first-level autonomous architecture; wherein, the computing power cluster includes multiple computing power chips with communication connections; the first network architecture includes multiple first nodes, and the first nodes in the same first network architecture are connected to form a tree topology; the first-level autonomous architecture includes multiple second nodes, and the second nodes in the same first-level autonomous architecture are connected to form a direct-connection topology, and different first-level autonomous architectures belong to different regions; at least one computing power cluster is connected to a first-level autonomous architecture through a first network architecture, and the first-level autonomous architecture connects to computing power clusters within its own region.

[0007] As can be seen, in some embodiments of this application, a computing power cluster formed by connecting multiple first nodes is connected to the first-level autonomous network architecture in its region through a tree topology, thereby forming a hierarchical and cascaded ultra-large-scale network architecture.

[0008] In some embodiments of this application, the three-level topology of tree-structured aggregation network, directly connected autonomous systems (AAS) in regions, and distributed computing power clusters are deeply coupled to form a "vertically layered and horizontally decoupled" topology structure. Specifically, vertically, the tree-structured network efficiently compresses uplink traffic, which can reduce the pressure on the backbone network; horizontally, direct connections within the AAS enable microsecond-level collaboration; and globally, the AAS connects to computing power clusters within its own region, allowing computing power clusters to access the AAS from the nearest location, thereby improving communication speed.

[0009] Therefore, the hierarchical cascaded network architecture provided in this application embodiment can guarantee communication speed while achieving ultra-large scale. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the hierarchical cascaded network architecture according to an embodiment of this application; Figure 2 This is a schematic diagram of the first-level autonomous architecture or the second-level autonomous architecture in the embodiments of this application; Figure 3 This is one of the connection diagrams of a computing chip within a network unit in an embodiment of this application; Figure 4 This is the second schematic diagram showing the connection of a computing chip within a network unit in an embodiment of this application; Figure 5 This is the third schematic diagram showing the connection of a computing chip within a network unit in an embodiment of this application; Figure 6 This is a schematic diagram illustrating a specific implementation of the hierarchical cascaded network architecture in the embodiments of this application. Detailed Implementation

[0011] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0012] In existing technologies, network architectures based on chip-to-chip (CPC) technology have attracted widespread attention. This architecture can form local computing networks through computing chips, possessing potential high bandwidth and low latency characteristics. However, how to scale up such local networks to ultra-large-scale network architectures remains a problem that urgently needs to be solved.

[0013] In some embodiments of this application, a tree-like topology formed by multiple first nodes connects computing power clusters connected by computing chips, connecting them to the first-level autonomous network architecture in their respective regions, thus forming a hierarchical, cascaded, ultra-large-scale network architecture. Specifically, in some embodiments of this application, the tree-like aggregation network, directly connected autonomous regions, and distributed computing power clusters are deeply coupled to form a "vertically layered, horizontally decoupled" topology: vertically, the tree-like network efficiently compresses uplink traffic, reducing backbone network pressure; horizontally, direct connections within autonomous regions enable microsecond-level collaboration; globally, connecting computing power clusters within an autonomous region allows clusters to access the nearest autonomous region, thereby improving communication speed. Therefore, the hierarchical, cascaded network architecture provided by this application can achieve ultra-large-scale operation while maintaining communication speed.

[0014] The hierarchical cascaded network architecture of the embodiments of this application will be described in detail below.

[0015] Firstly, such as Figure 1 As shown, embodiments of this application provide a hierarchical cascaded network architecture; this hierarchical cascaded network architecture may include: At least one computing power cluster, the computing power cluster comprising multiple computing power chips connected in communication; At least one first network architecture, the first network architecture including multiple first nodes, and the first nodes in the same first network architecture communicate to form a tree topology; At least one first-level autonomous architecture, wherein the first-level autonomous architecture includes multiple second nodes, and the second nodes in the same first-level autonomous architecture communicate to form a direct connection topology, and different first-level autonomous architectures belong to different regions; In this configuration, at least one of the computing power clusters communicates with a first-level autonomous architecture through a first network architecture, and the first-level autonomous architecture connects to computing power clusters within its respective region.

[0016] First, the above-mentioned computing power cluster is explained as follows: For any computing power cluster, each computing power chip may include at least one switching core. The computing power chips in the same computing power cluster are connected to each other through the switching core, which is a core that encapsulates the communication protocol.

[0017] It should be noted that in traditional solutions, cluster networking is achieved through a series of network devices such as network interface cards (NICs) and optical modules, resulting in high costs for network architectures built on computing chips. However, in some embodiments of this application, the communication protocol can be encapsulated into a switching core and embedded within the computing chip. This allows computing chips to communicate and connect via the switching core, enabling addressing between them. Consequently, computing chips within the computing cluster can communicate without additional network equipment, thus reducing costs.

[0018] Optionally, the communication protocol can be the V2V protocol or the Internet Protocol (IP) protocol. That is, the V2V protocol can be encapsulated into a switching core and installed in the computing chip, enabling communication between computing chips based on the V2V protocol; alternatively, the IP protocol can also be encapsulated into a switching core and installed in the computing chip, enabling communication between computing chips based on the IP protocol. The V2V protocol is primarily used to solve the problems of instantaneous addressing, communication establishment, data transmission, quality assurance, and security assurance among hundreds of millions of devices over a wide area.

[0019] The computing chip may include at least one of the following: a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), and a field-programmable gate array (FPGA).

[0020] It should be noted that the number of computing chips included in different computing power clusters may be the same or different, and the specific topology formed by the communication connection of computing power chips in different computing power clusters may be the same or different.

[0021] Secondly, the first network architecture described above is explained as follows: In each first network architecture, the communication connections of the first node form a tree topology. This tree topology is a network architecture formed through hierarchical connections, its core feature being the expansion of multiple sub-topologies in a tree-like structure. This structure uses a unidirectional transmission medium to construct a non-closed loop, supports bidirectional data transmission in broadcast communication, and possesses hierarchical expansion capabilities and fault isolation characteristics.

[0022] It should be noted that the number of first nodes included in different first network architectures may be the same or different, and the specific form of the tree topology formed by the communication connections of the first nodes in different first network architectures may be the same or different.

[0023] Furthermore, the above-mentioned first-level autonomous architecture is explained as follows: In each Level 1 autonomous architecture, the communication connections of the second nodes form a direct-connect topology. In a direct-connect topology, all second nodes communicate point-to-point. For example, in a direct-connect topology, one second node communicates point-to-point with each other (i.e., all nodes other than itself). The direct-connect topology is a layout where network nodes are directly connected via point-to-point physical links, characterized by its simplicity and low latency.

[0024] It should be noted that the number of second nodes included in different first-level autonomous architectures may be the same or different, and the specific form of the direct connection topology formed by the communication connections of the second nodes in different first-level autonomous architectures may be the same or different.

[0025] Finally, the connection relationships between the aforementioned computing power cluster, the first network architecture, and the first-level autonomous architecture are explained as follows: In some embodiments of this application, a region can be divided into multiple regions. Each region can have a first-level autonomous architecture and at least one computing power cluster. Each first-level autonomous architecture can have a corresponding first network architecture (i.e., it can also be understood that a first network architecture is set up in each region). In this way, computing power clusters in the same region can access the first-level autonomous architecture in that region through the first network architecture in that region.

[0026] As described above, in some embodiments of this application, a hierarchical cascaded network architecture is provided, comprising: at least one computing power cluster, at least one first network architecture, and at least one first-level autonomous architecture; wherein, the computing power cluster includes multiple computing power chips with communication connections; the first network architecture includes multiple first nodes, and the first nodes in the same first network architecture are connected to form a tree topology; the first-level autonomous architecture includes multiple second nodes, and the second nodes in the same first-level autonomous architecture are connected to form a direct-connection topology, and different first-level autonomous architectures belong to different regions; at least one computing power cluster is connected to a first-level autonomous architecture through a first network architecture, and the first-level autonomous architecture connects to computing power clusters within its respective region.

[0027] As can be seen, in some embodiments of this application, a computing power cluster formed by connecting multiple first nodes is connected to the first-level autonomous network architecture in its region through a tree topology, thereby forming a hierarchical and cascaded ultra-large-scale network architecture.

[0028] In some embodiments of this application, the three-level topology of tree-structured aggregation network, directly connected autonomous systems (AAS) in regions, and distributed computing power clusters are deeply coupled to form a "vertically layered and horizontally decoupled" topology structure. Specifically, vertically, the tree-structured network efficiently compresses uplink traffic, which can reduce the pressure on the backbone network; horizontally, direct connections within the AAS enable microsecond-level collaboration; and globally, the AAS connects to computing power clusters within its own region, allowing computing power clusters to access the AAS from the nearest location, thereby improving communication speed.

[0029] Therefore, the hierarchical cascaded network architecture provided in this application embodiment can guarantee communication speed while achieving ultra-large scale.

[0030] Optionally, such as Figure 1 As shown, the hierarchical cascaded network architecture also includes: At least one second-level autonomous architecture, wherein the second-level autonomous architecture includes multiple third nodes, the third nodes in the same second-level autonomous architecture communicate to form a direct connection topology, and different second-level autonomous architectures belong to different regions; The second-level autonomous architecture connects to the first-level autonomous architecture within at least one region of its own region.

[0031] As can be seen, in some embodiments of this application, the communication connections of the third nodes in each second-level autonomous architecture form a direct-connection topology. It should be noted that the number of third nodes included in different second-level autonomous architectures may be the same or different, and the specific form of the direct-connection topology formed by the communication connections of the third nodes in different second-level autonomous architectures may be the same or different.

[0032] Therefore, if different regions belong to the same region, the first-level autonomous architecture in the different regions can be connected to the second-level autonomous architecture in the region. This further expands the scale of the hierarchical cascaded network architecture of this application embodiment, thereby enabling scheduling between nodes in a larger area.

[0033] Optionally, the second nodes are connected via a GlobalLink communication link.

[0034] Optionally, the third nodes are connected via GlobalLink communication.

[0035] As can be seen, the nodes in a direct-connect topology can communicate and connect with each other through a global bus.

[0036] The following uses V2V video network as an example to introduce the specific structure of the first-level autonomous architecture or the second-level autonomous architecture: like Figure 2 As shown, each node in a direct-connect topology can include intelligent computing resources and storage resources. Within a node, devices serving as intelligent computing resources and devices serving as storage resources can communicate with at least one switch 201, while different nodes can communicate with each other through the switches 201 included in the node. Specifically, devices serving as intelligent computing resources can communicate with switches via AccessLink, devices serving as storage resources can communicate with switches via AccessLink, switches within a node can communicate with each other via LocalLink, and nodes can communicate with each other via GlobalLink. Furthermore, in a direct-connect topology, any node can communicate with all other nodes except itself.

[0037] It should be noted that in some embodiments, when the first-level autonomous architecture or the second-level autonomous architecture adopts a direct connection topology based on V2V video networking, it can achieve network configuration of any size without considering the number of links; moreover, it adopts a non-IP routing mechanism, based on device identifier point-to-point direct connection, and the latency can approach the speed of light; in addition, it can also take into account the overall network situation, realize resource reservation based on thermal sensing mechanism, and achieve packet loss-free communication.

[0038] Optionally, the target object includes at least one network unit, the network unit includes multiple computing chips, the computing chips in the same network unit are connected to form multiple loops, and one computing chip is in two loops; In the case where the target object includes multiple network units, the target object also includes at least one switch, and the network units are communicatively connected through the switch; The target object includes at least one of the following: the computing power cluster, at least one first node, at least one second node, and at least one third node.

[0039] In some embodiments, the switch can be a VRB switch. Thus, in this embodiment, a network can be formed using a VRB switch to obtain the aforementioned network unit based on a computing chip. VRB is a Remote Direct Memory Access (RDMA) technology implemented using the V2V communication protocol; RDMA is a technology for directly accessing the memory of a remote computer in a computer network. RDMA allows a computer to directly access the memory of another computer, bypassing most of the intervention of the operating system kernel. The sending end can directly send data from its own memory to the receiving end's memory without going through complex intermediate processing steps, greatly reducing data transmission latency.

[0040] As described above, each of the aforementioned at least part of the computing power cluster, at least part of the first node, at least part of the second node, and at least part of the third node may include at least one of the aforementioned network units. It should be noted that when the second or third node includes at least one of the aforementioned network units, at least some of the computing power chips within the second or third node can serve as intelligent computing resources.

[0041] It is understood that at least some of the first node, at least some of the second node, and at least some of the third node may also be network devices.

[0042] Optionally, where each computing chip may include at least one switching core, the computing chips within the same network unit can communicate with each other via the switching cores. For example, the number of switching cores within a single computing chip is 2. n n is an integer greater than or equal to 2; a computing chip has 2 n-2 One switching core and another computing chip 2 n-2 If two switching cores are connected, then a computing chip is on two loops; for example, when n=3, if two switching cores of one computing chip are connected to two switching cores of another computing chip, then a computing chip is in two loops.

[0043] Optionally, the computing chips within the same network unit are arranged in a matrix, with computing chips in the same row of the matrix connected end to end to form a loop, and computing chips in the same column of the matrix connected end to end to form a loop.

[0044] Therefore, the computing chips within a network unit can be arranged in a matrix. In this case, computing chips in the same row form a loop, and computing chips in the same column form another loop. For example, when n=2, multiple computing chips within a network unit can be arranged as follows: Figure 3 In the matrix form shown, computing chips in the same row are connected end-to-end to form a loop, and computing chips in the same column are connected end-to-end to form another loop. Thus, a single computing chip is located in two loops. Figure 3 One circle in the diagram represents a computing chip.

[0045] It is understood that the connection method between multiple computing chips in a network unit is not limited to the above description. For example, in some embodiments, the communication connections between computing chips within a network unit form multiple loops, and one computing chip is in a 2 n-1 Within a loop.

[0046] Optionally, in a computing chip at 2 n-1 In a loop, with n=3 (i.e., when one computing chip is in 4 loops), the computing chips in a network unit are distributed on each edge (i.e., 12 edges) of a cuboid and on the four first diagonals of the cuboid. The computing chips on the same line of each edge of the cuboid and the four first diagonals are connected end to end to form a loop. In the cuboid, the vertices of the four first diagonals are all different.

[0047] For example, when n=3, the computing chips within a network unit can be located on the 12 edges of a cuboid and the four diagonals with distinct vertices. Furthermore, the computing chips on any one edge or diagonal can be connected end-to-end to form a loop, such as... Figure 4 or Figure 5 As shown. It should be noted that, Figure 4 and Figure 5 Only one loop is shown in the diagram; examples of the other loops are similar.

[0048] It is understandable that the four object lines with distinct vertices in a cuboid are not limited to... Figure 4 or Figure 5 Other situations may exist as shown, which will not be listed here.

[0049] It should also be noted that, Figure 4 and Figure 5 Only the connection method for 8 computing chips is shown. If more computing chips exist, they can be connected together. Figure 4 or Figure 5 The units shown are then stacked and connected to obtain a network unit that includes more computing chips.

[0050] Optionally, the computing chips within the same network unit are connected via optical fiber communication; that is, a dedicated network interface is provided on the computing chip, which is directly connected to the network interfaces of other computing chips via optical fiber to realize data transmission between chips. Therefore, in this embodiment, chip-to-chip technology can be used to establish connections between multiple computing chips directly via optical fiber, forming a fully interconnected local computing network. For example, when the computing chips are GPUs, the network can support up to 4096 GPUs.

[0051] Furthermore, optionally, each computing chip may also include at least one of the following: a processor core, a memory controller, a direct memory access interface, and a GPU core module including a multi-channel direct memory access (DMA) interface and a hash. The processor core may adopt an ARM Cortex-A57 architecture, the memory controller may support multiple memory modes, and the DMA interface may support 8-channel memory access. It should be noted that ARM Cortex-A57 represents a high-performance application processor.

[0052] Optionally, the first node includes a root node, leaf nodes, and intermediate nodes, and the leaf nodes are communicatively connected to the root node through the intermediate nodes; that is, each first network architecture can be formed by the root node, leaf nodes, and intermediate nodes communicating with each other to form a tree topology.

[0053] As can be seen, in some embodiments of this application, when the scale of the hierarchical cascaded network architecture exceeds that of a single computing power cluster, a tree topology can be used to connect multiple computing power clusters to the corresponding first-level autonomous architecture to build a super-large-scale cluster. The roles of each node in the tree topology can be defined: the root node is a Spine node, the leaf nodes are Leaf nodes, and the intermediate nodes are Intermediate nodes. Spine nodes connect to Leaf nodes / Intermediate nodes, and Intermediate nodes connect to Spine nodes / Leaf nodes. Furthermore, the number of layers in the tree topology is configurable, and all nodes have routing capabilities.

[0054] For example, in the tree topology, the root node uses 4 Spine nodes, the leaf nodes use 64 Leaf nodes, and the intermediate nodes use 16 Intermediate nodes. Spine nodes connect Leaf nodes and Intermediate nodes, and Intermediate nodes connect Spine nodes and Leaf nodes; the tree topology has 3 levels.

[0055] Optionally, both Leaf and Intermediate nodes are configured with host routing tables and network segment routing tables, while Spine nodes are configured with only network segment routing tables. The host routing table stores host routing entries, and the network segment routing table stores network segment routing entries. When a data packet arrives at a Spine node from a Leaf node, the Spine node makes a routing decision based on the source and destination addresses of the packet and forwards it to the destination address. During the forwarding process between Leaf and Spine nodes, a routing identifier field is added to the packet header to indicate the nodes the packet passed through during transmission. For example, if the data packet's transmission path is: Leaf node 1 → Intermediate node 1 → Root node 1 → Root node 2 → Intermediate node 2 → Leaf node 2, then each node on this path will add a field to its packet header to identify itself when forwarding the data packet. Thus, the transmission path of the data packet can be determined from the packet header.

[0056] Optionally, the network interfaces of the leaf nodes and the intermediate nodes are pluggable, thereby enabling the first network architecture to support the addition or removal of corresponding nodes as needed.

[0057] Optionally, when the first nodes (i.e., nodes in the tree topology) are connected via optical fiber, the network interfaces between the first nodes are hot-swappable. Hot-swappable design refers to the ability of a hardware or software system to allow a user to safely insert or remove components without powering off or interrupting service. In some embodiments of this application, if the network interfaces between the first nodes support hot-swapping when they are connected via optical fiber, the first network architecture can support more flexible configurations.

[0058] Optionally, the root node, leaf nodes, and intermediate nodes are each provided with a cache region. For example, leaf nodes are equipped with a level-one cache, intermediate nodes with a level-two cache, and the root node with a level-three cache. The cache at each level can employ a hybrid strategy combining static and dynamic methods. Therefore, in some embodiments of this application, the first network architecture can employ a hierarchical multi-level cache.

[0059] Optionally, in the first network architecture, a distributed task scheduling mechanism can be adopted to dynamically adjust the task allocation strategy based on the computing power and real-time load of the nodes.

[0060] Optionally, each computing chip and node in the hierarchical cascaded network architecture can also support multiple network protocol stacks (such as InfiniBand protocol, Ethernet protocol, Remote Direct Memory Access (RDMA) protocol), and can be selected according to the scenario; thus, the hierarchical cascading of the embodiments of this application can be applied to more scenarios.

[0061] Optionally, each computing chip and node in the hierarchical cascaded network architecture can also be flexibly configured with various acceleration engines, such as encryption engines, data compression engines, and task scheduling engines. Among them, the encryption engine is used to encrypt data packets during forwarding to improve data transmission security; the data compression engine is used to compress data during forwarding to reduce data size, thereby improving transmission rate; and the task scheduling engine is used to schedule corresponding nodes to execute corresponding tasks.

[0062] Optionally, the hierarchical cascaded network architecture can support a variety of scheduling algorithms (such as round-robin scheduling, priority scheduling, and performance-based scheduling), and can select appropriate strategies as needed. For example, users can set up different computing power clusters, different first-level network architectures, different first-level autonomous architectures, or different second-level autonomous architectures, and use different scheduling algorithms.

[0063] In summary, the specific implementation of the hierarchical cascaded network architecture of this application can be described as follows: First, it's important to note that traditional data center network architectures are no longer sufficient to handle large-scale concurrent computing and massive data transmission. Against this backdrop, the industry has begun exploring new network architectures to meet these growing demands. Among these, network architectures based on direct chip-to-chip (CPC) connectivity have attracted widespread attention. This architecture allows for the formation of local area networks (LANs) directly between computing chips via fiber optic cables, offering potentially high bandwidth and low latency. However, how to extend such LANs to ultra-large-scale clusters and achieve high-throughput, fully connected wide-area computing networks remains a critical challenge.

[0064] To address the aforementioned problems, this embodiment provides a hierarchical cascaded network architecture, such as... Figure 6 As shown, the hierarchical cascaded network architecture includes: At least one computing power cluster 601 is provided, which includes at least one network unit. If the computing power cluster 601 includes multiple network units, it also includes at least one switch. The network units within the computing power cluster 601 are communicatively connected via the switch. It should be noted that... Figure 6 In this context, AIDC represents a computing power cluster; At least one first network architecture 602, the first network architecture 602 includes multiple first nodes, and the first nodes in the same first network architecture 602 are connected to form a tree topology; at least some of the first nodes include at least one network unit, wherein, when the first node includes multiple network units, the first node also includes at least one switch, and the network units within the first node are connected to each other through the switch; at least some of the first nodes are network devices. At least one Level 1 Autonomous Architecture 603, which includes multiple Second Nodes, and the Second Nodes in the same Level 1 Autonomous Architecture 603 are connected to form a direct-connect topology. Different Level 1 Autonomous Architectures 603 belong to different regions. At least some of the Second Nodes include at least one Network Unit. In the case where a Second Node includes multiple Network Units, the Second Node also includes at least one Switch. The Network Units within the Second Node are connected to each other through the Switch. At least some of the Second Nodes are network devices. At least one Level 2 Autonomous Architecture 604, wherein the Level 2 Autonomous Architecture 604 includes multiple Third Nodes, and the Third Nodes in the same Level 2 Autonomous Architecture 604 are connected to form a direct-connect topology. Different Level 2 Autonomous Architectures 604 belong to different regions. At least some of the Third Nodes include at least one Network Unit, wherein, in the case where the Third Node includes multiple Network Units, the Third Node also includes at least one Switch, and the Network Units within the Third Node are connected to each other through the Switch. At least some of the Third Nodes are network devices. In this configuration, at least one computing power cluster 601 is connected to a first-level autonomous architecture 603 via a first network architecture 602, and the first-level autonomous architecture 603 is connected to the computing power cluster 601 within its region; the second-level autonomous architecture 604 is connected to the first-level autonomous architecture 603 within at least one region within its region.

[0065] In addition, for the above-mentioned network unit, a network unit may include multiple computing chips. The computing chips in the same network unit can be arranged in a matrix, and the computing chips in the same row of the matrix are connected end to end to form a loop. The computing chips in the same column of the matrix are connected end to end to form a loop. The computing chips in the same network unit can be connected to each other through optical fiber communication.

[0066] Furthermore, regarding the aforementioned first network architecture 602: The first node in the first network architecture 602 may include a root node, leaf nodes, and intermediate nodes. The leaf nodes communicate with the root node through the intermediate nodes. The network interfaces of the leaf nodes and the intermediate nodes are pluggable. When the first nodes are connected by optical fiber, the network interfaces between the first nodes are hot-swappable. Buffer areas can be set in the root node, leaf nodes, and intermediate nodes respectively.

[0067] In addition, host routing tables and network segment routing tables are configured in both leaf nodes and intermediate nodes, while the root node only has a network segment routing table. The host routing table stores host routing entries, and the network segment routing table stores network segment routing entries. When a data packet arrives at the root node from a leaf node, the root node makes a routing decision based on the source and destination addresses of the data packet and forwards the data packet to the destination address. During the forwarding process of the data packet in the leaf and root nodes, a routing identifier field can be added to the packet header to indicate the nodes that the data packet passes through during transmission.

[0068] Furthermore, in the first network architecture 602, a distributed task scheduling mechanism can be adopted to dynamically adjust the task allocation strategy based on the computing power and real-time load of the nodes.

[0069] For the aforementioned Level 1 Autonomous Architecture 603 or Level 2 Autonomous Architecture 604, each node may include intelligent computing resources and storage resources. Within a node, devices serving as intelligent computing resources and devices serving as storage resources can communicate with at least one switch, while different nodes can communicate with each other through the switches included in the node. Specifically, devices serving as intelligent computing resources and switches can communicate via AccessLink, devices serving as storage resources and switches can communicate via AccessLink, switches within a node can communicate via LocalLink, and nodes can communicate with each other via GlobalLink.

[0070] It should also be noted that the various computing chips and nodes in the hierarchical cascaded network architecture of this embodiment can also support multiple network protocol stacks (such as InfiniBand protocol, Ethernet protocol, RDMA protocol), and can be selected according to the scenario.

[0071] In this embodiment, the various computing chips and nodes in the hierarchical cascaded network architecture can also be flexibly configured with a variety of acceleration engines, such as encryption engines, data compression engines, and task scheduling engines.

[0072] The hierarchical cascaded network architecture in this embodiment can support a variety of scheduling algorithms (such as round-robin scheduling, priority scheduling, and performance-based scheduling), and appropriate strategies can be selected as needed.

[0073] As can be seen from the above, this embodiment has the following advantages: 1. By utilizing chip-to-chip connection technology, a local computing network (i.e., computing cluster) is formed directly between computing chips via optical fiber, which breaks through the limitations of traditional data center network architecture in handling large-scale concurrent computing and massive data transmission, and effectively meets the increasingly complex business needs.

[0074] 2. Based on VRB technology, multiple computing power clusters are networked in a tree topology to form an ultra-large-scale chip cluster, realizing a high-throughput, fully connected wide-area computing power network, overcoming the limitations of existing network-on-chip (NoC) architecture in handling large-scale concurrent computing and high bandwidth requirements.

[0075] 3. By accessing the V2V autonomous cloud through a tree topology, the network achieves flexible expansion and efficient utilization, solving the problem that the existing network architecture is difficult to adapt to different application scenarios and changing needs.

[0076] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0077] Furthermore, although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0078] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0079] The above provides a detailed description of a hierarchical cascaded network architecture provided by this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the structure and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A hierarchical cascaded network architecture, characterized in that, include: At least one computing power cluster, the computing power cluster comprising multiple computing power chips connected in communication; At least one first network architecture, the first network architecture including multiple first nodes, and the first nodes in the same first network architecture communicate to form a tree topology; At least one first-level autonomous architecture, wherein the first-level autonomous architecture includes multiple second nodes, and the second nodes in the same first-level autonomous architecture communicate to form a direct connection topology, and different first-level autonomous architectures belong to different regions; In this configuration, at least one of the computing power clusters communicates with a first-level autonomous architecture through a first network architecture, and the first-level autonomous architecture connects to computing power clusters within its respective region.

2. The hierarchical cascaded network architecture according to claim 1, characterized in that, The hierarchical cascaded network architecture also includes: At least one second-level autonomous architecture, wherein the second-level autonomous architecture includes multiple third nodes, the third nodes in the same second-level autonomous architecture communicate to form a direct connection topology, and different second-level autonomous architectures belong to different regions; The second-level autonomous architecture connects to the first-level autonomous architecture within at least one region of its own region.

3. The hierarchical cascaded network architecture according to claim 1 or 2, characterized in that, The target object includes at least one network unit, the network unit includes multiple computing chips, the computing chips within the same network unit are connected to form multiple loops, and one computing chip is in two loops; In the case where the target object includes multiple network units, the target object also includes at least one switch, and the network units are communicatively connected through the switch; The target object includes at least one of the following: the computing power cluster, at least one first node, at least one second node, and at least one third node.

4. The hierarchical cascaded network architecture according to claim 3, characterized in that, The computing chips within the same network unit are arranged in a matrix, with computing chips in the same row of the matrix connected end to end to form a loop, and computing chips in the same column of the matrix connected end to end to form a loop.

5. The hierarchical cascaded network architecture according to claim 3, characterized in that, The computing chips within the same network unit are connected via optical fiber communication.

6. The hierarchical cascaded network architecture according to claim 1, characterized in that, The first node includes a root node, leaf nodes, and intermediate nodes, and the leaf nodes are communicatively connected to the root node through the intermediate nodes.

7. The hierarchical cascaded network architecture according to claim 6, characterized in that, Both the network interfaces of the leaf nodes and the network interfaces of the intermediate nodes are pluggable.

8. The hierarchical cascaded network architecture according to claim 1, characterized in that, When the first nodes are connected via optical fiber, the network interfaces between the first nodes are hot-swappable.

9. The hierarchical cascaded network architecture according to claim 6, characterized in that, The root node, the leaf node, and the intermediate node are each provided with a cache area.

10. The hierarchical cascaded network architecture according to claim 1, characterized in that, The second node is connected to each other via a global interconnect link, GlobalLink.