Network-on-chip architecture for super-large-scale neural mimicry computing
By adopting hybrid routing modules, multi-stage pipeline routers and scalable Tile hierarchical structures in the on-chip network architecture, combining dual-mode data packets and hybrid flow control mechanisms, the problems of insufficient communication efficiency, scalability bottlenecks and hardware performance limitations in the existing technology are solved, and the effects of efficient communication and large-scale expansion are achieved.
Patent Information
- Application Number
- CN202510577007.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-13
AI Technical Summary
When building a neuromimicry computing system with a scale of hundreds of billions of neurons, the existing on-chip network architecture faces problems such as insufficient communication efficiency, scalability bottlenecks and hardware performance limitations.
A network architecture for ultra-large-scale neuromimicry computing is proposed, using a hybrid routing module, a multi-stage pipeline router and an extensible Tile hierarchical structure, combining dual-mode data packets and hybrid flow control mechanisms to support differentiated processing of pulsed data multicast and non-pulse data unicast.
It significantly improves the communication efficiency and scalability of brain-like computing systems, supports the expansion of neuron scale from ten thousand to tens of billions of levels, and controls the hardware complexity, meeting the needs of real-time computing.
Smart Images

Figure CN120144524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neuromorphic computing, and in particular to an on-chip network architecture for ultra-large-scale neuromorphic computing. Background Art
[0002] Neuromorphic computing realizes low-power and high-parallel computing by simulating the spiking neural network (SNN) of the biological brain, and is a key direction to break through the "memory wall" and high-power consumption bottleneck of the traditional von Neumann architecture. However, when constructing a brain-inspired computing system with a scale of hundreds of billions of neurons, the existing on-chip network (NoC) faces the following core challenges: Insufficient communication efficiency: The spiking data of SNN has sparsity, asynchrony, and multicast requirements. Traditional single routing strategies (such as XY dimension-order routing or source-driven routing) are difficult to simultaneously meet the efficient distribution of spiking data and the deterministic transmission of non-spiking data, resulting in low link utilization and high latency.
[0003] Scalability bottleneck: As the neuron scale expands, the storage capacity of the routing table and the hardware complexity increase exponentially. Existing hardware designs are difficult to support large-scale routing paths in limited SRAM, and the address space and path planning efficiency are limited when cascading across chips.
[0004] Hardware performance limitation: Traditional NoC routers adopt synchronous or asynchronous architectures, with insufficient working frequency and throughput, and cannot meet the strict requirements of SNN real-time computing for the processing speed of spiking data (such as peak throughput rate, latency).
[0005] Therefore, a new technical solution is urgently needed to solve the above technical problems. Summary of the Invention
[0006] The purpose of the present invention is to overcome the problems of the above existing technologies, and provides an on-chip network architecture for ultra-large-scale neuromorphic computing to solve the technical problems of data transmission and routing optimization in ultra-large-scale neuromorphic computing.
[0007] The above purpose is achieved by the following technical solutions: An on-chip network architecture for ultra-large-scale neuromorphic computing, comprising: A hybrid routing module for implementing a three-layer hybrid strategy of source-driven routing, XY dimension-order routing, and explicit deterministic routing, and supporting differential processing of spiking data multicast and non-spiking data unicast; A multi-stage pipelined router, adopting pipeline inter-stage cache optimization and bubble-free flow control design, and supporting a peak working frequency above 1 GHz; An extensible Tile hierarchical structure for realizing the scale expansion of hundreds of billions of neurons through 2D-Mesh topology cascading; Dual-mode data packets and hybrid flow control mechanism, which distinguish between pulsed and non-pulsed data types, and optimize data parsing and forwarding efficiency.
[0008] Furthermore, the source-driven routing adopts a clustering-based address event representation method, splitting the source neuron ID into a routing path group number and a source neuron sub-number, and improving the storage utilization rate through routing table clustering compression.
[0009] Furthermore, the XY-order routing is based on a 2D-Mesh topology, and realizes deterministic path planning through coordinate difference calculation, supporting the X / Y priority strategy and large-bitwidth coordinate addresses (≥30 bits).
[0010] Furthermore, the explicit deterministic routing embeds complete path information in the data packet, supports dynamic traffic allocation and path detouring, and is applicable to irregular topologies and fault-tolerant scenarios.
[0011] Furthermore, the multi-stage pipelined router includes parsing, routing calculation, switching, and output pipelines. The single-node peak pulse throughput rate is ≥800 Mspikes / s, and the processing delay is ≤15 ns.
[0012] Furthermore, the scalable Tile structure includes a router, a processing unit, and local memory. Dies are formed by cascading 3×3 Tiles, and then chips are formed by cascading 2×2 Dies, supporting cross-Die cascaded communication.
[0013] Furthermore, the dual-mode data packet uses a 128-bit width, including a routing packet header and a data packet header. The routing packet header carries a routing mode identifier, address information, and control fields, and the data packet header includes read / write addresses and data length information.
[0014] Furthermore, the hybrid flow control mechanism includes: the on-chip router adopts the AXI-Stream protocol, supporting handshake and backpressure flow control; the local module realizes high-speed data interaction between the CPU / BPU and the NoC through DMA technology, including bpu_NI and cpu_NI interfaces.
[0015] Furthermore, the router microarchitecture adopts a multi-stage pipeline design, including parsing, routing calculation, switching, and output pipelines. FIFO caches and bubble-free flow control mechanisms are integrated between pipeline stages.
[0016] The on-chip network architecture for ultra-large-scale neuromorphic computing provided by the present invention can significantly improve the communication efficiency and scalability of the brain-like computing system, provide key technical support for ultra-large-scale neuromorphic chips, and also include the following beneficial effects: Efficient Communication: The hybrid routing strategy is optimized for the characteristics of pulsed and non-pulsed data, improving link utilization and transmission efficiency to meet the asynchronous, sparse, and multicast communication requirements of SNNs.
[0017] High Scalability: Through clustering compression, large-bitwidth addresses, and a hierarchical Tile architecture, it supports scaling from tens of thousands to hundreds of billions of neurons, with controllable hardware complexity.
[0018] High-Performance Hardware: The multi-stage pipeline design significantly improves the operating frequency and throughput rate, and the low-latency feature meets the requirements of real-time computing, providing an efficient communication infrastructure for neuromorphic chips.
[0019] Versatility and Flexibility: It supports 2D-Mesh topologies and multiple routing modes, is compatible with existing neuromorphic chip architectures, and can be extended to 3D integration or optical interconnection (ONoC) to adapt to future technological evolution. Description of the Drawings
[0020] Figure 1 This is a schematic diagram of the hybrid routing architecture in the on-chip network architecture for ultra-large-scale neuromorphic computing according to the present invention, showing the data flow paths of source-driven routing (red), XY dimension-order routing (green), and explicit deterministic routing (blue), as well as the differential processing of pulsed data multicast and non-pulsed data unicast; Figure 2 This is a diagram of the core composition architecture of a Tile in the on-chip network architecture for ultra-large-scale neuromorphic computing according to the present invention, including a router, processing units (BPU / CPU), local memory, and interfaces, reflecting the hierarchical expansion structure; Figure 3 This is a micro-architecture diagram of the router hardware in the on-chip network architecture for ultra-large-scale neuromorphic computing according to the present invention, showing the multi-stage pipeline design, input / output ports, switching structure, and routing calculation unit (RAU); Figure 4 This is an architecture diagram of the routing calculation unit (RAU) in the on-chip network architecture for ultra-large-scale neuromorphic computing according to the present invention, showing the hybrid routing hardware architecture. Detailed Embodiments
[0021] The present invention will be further described in detail below with reference to the drawings and embodiments. The described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0022] This solution provides an on-chip network architecture for ultra-large-scale neuromorphic computing, proposing a hybrid strategy of source-driven routing, XY-order routing, and explicit deterministic routing. Combining multi-stage pipelined hardware design with a scalable Tile architecture, it solves the problems of efficient transmission and large-scale expansion of SNN pulse data. The hybrid routing is optimized for data characteristics, the pipelining technology improves throughput and latency performance, and the hierarchical structure supports a scale of hundreds of billions of neurons; including: A hybrid routing module for implementing a three-layer hybrid strategy of source-driven routing, XY-order routing, and explicit deterministic routing, supporting differential processing of pulse data multicast and non-pulse data unicast; A multi-stage pipelined router, using pipeline inter-stage cache optimization and bubble-free flow control design, supporting a peak operating frequency above 1 GHz; A scalable Tile hierarchical structure, realizing the scale expansion of hundreds of billions of neurons through 2D-Mesh topology cascading; A dual-modal data packet and hybrid flow control mechanism, distinguishing between pulse and non-pulse data types, and optimizing data parsing and forwarding efficiency.
[0023] As Figure 1 shown, as the hybrid routing architecture design in this embodiment, a three-layer hybrid routing strategy including source-driven routing, XY-order routing, and explicit deterministic routing is proposed, supporting differential processing of pulse data and non-pulse data, where: Source-driven routing: For the multicast requirements of pulse data, using the clustering-based address event representation method, splitting the source neuron ID into a routing path group number and a source neuron sub-number, compressing the routing table entries through clustering, storing more routing paths in limited SRAM, and improving the fan-out scale. The routing packet header contains fields such as Type, Spike flag, RID, PID, and NID, and the multicast path can be quickly determined by looking up the table, which is suitable for the pulse distribution scenario of the fully connected layer.
[0024] XY-order routing: For the deterministic transmission of non-pulse data (such as control signals, weight updates), based on the 2D-Mesh topology, by calculating the coordinate difference (Delta X / Delta Y) between the source node and the destination node, using the X-first or Y-first strategy for deterministic routing, avoiding deadlocks and simplifying the hardware design, which is suitable for inter-layer neuron communication. The routing packet header contains fields such as Type, X / Y first flag, DstX, DstY, SrcX, and SrcY, supporting large-bitwidth addresses for cross-Die cascading.
[0025] Explicit Routing Determination: For irregular topologies or fault-tolerant scenarios, complete path information (DstPath) is embedded in the data packet to support dynamic traffic allocation and path detouring. The routing packet header contains fields such as Type, Last flag, and DstPath. By parsing the path information hop by hop, free path transmission is achieved, enhancing network flexibility and fault tolerance.
[0026] In this embodiment, the multi-stage pipeline hardware architecture includes: As Figure 3 shown, the router microarchitecture: Adopts a multi-stage pipeline design (such as parsing, routing calculation, switching, output pipeline), supports a peak operating frequency above 1 GHz, the actual application frequency reaches 800 MHz, the peak pulse throughput rate of a single node reaches 800 Mspikes / s, and the processing delay is as low as 15 ns. FIFO caches and bubble-free flow control mechanisms are integrated between pipeline stages to optimize the critical path delay.
[0027] As Figure 2 shown, the Tile hierarchical expansion structure: Using Tile as the basic unit, each Tile integrates a router, processing units (BPU / CPU), and local memory. A 3×3 Tile Die-level array is constructed through a 2D-Mesh topology, and then a chip-level architecture is formed through 2×2 Die cascading, supporting the expansion of a scale of hundreds of billions of neurons. A loose coupling design is adopted within the Tile, and high-speed data interaction is achieved through the AXI-Stream interface.
[0028] As the optimization of data flow and communication protocol in this embodiment: Dual-mode data packets: Distinguish between pulse data (multicast) and non-pulse data (unicast), design a 128-bit wide packet, including a routing packet header and a data packet header. The routing packet header carries a routing mode identifier, address information, and control fields, and the data packet header contains read / write addresses, data lengths, etc., supporting efficient parsing and forwarding.
[0029] Hybrid flow control mechanism: The AXI-Stream protocol is adopted between on-chip routers, supporting handshake and backpressure flow control; local modules (Local) achieve high-speed data interaction between the CPU / BPU and the NoC through DMA technology, including bpu_NI and cpu_NI interfaces, supporting pulse data unpacking / packeting and memory reading / writing of non-pulse data.
[0030] This solution also involves technologies for enhancing scalability, including: Routing table compression: Through source route clustering technology, similar routing paths are merged into routing path groups, reducing the number of routing entries in SRAM, improving storage utilization, and supporting large-scale neuron mapping.
[0031] Large-bitwidth address design: In the XY dimension-order routing, 30-bit coordinate differences (DstX / DstY) are adopted, supporting long-distance communication across Dies for cascading; explicitly determine the routing to extend the path length through multi-cycle routing header extensions, theoretically supporting infinite path extension.
[0032] As the implementation of the hybrid routing strategy in this embodiment; Source-driven routing: When pulse data enters the router, parse the RID / PID / NID in the routing header. By looking up the routing entry matching the RID in the local routing table, if valid, copy the data packet to multiple output ports in parallel according to the routing information to achieve multicast forwarding. Clustering compression enables a single routing entry to cover the same paths of multiple source neurons, reducing the lookup time and storage overhead.
[0033] XY dimension-order routing: After non-pulse data enters the router, calculate DstX / DstY. According to the X / Y priority strategy (configurable register or header identifier), first forward it in the X direction to the target row, and then forward it in the Y direction to the target node to ensure deterministic path planning and avoid congestion and deadlocks.
[0034] Explicitly determined routing: The data packet carries the complete DstPath path information. Every time it passes through a router, strip the highest-bit path direction (e.g., 0 = west, 1 = north, 2 = east, 3 = south), shift the remaining path information to the left, and forward it hop by hop until it reaches the destination node, supporting dynamic bypass of faulty nodes or optimizing load balancing.
[0035] As Figure 4 shown, as the implementation of the hardware architecture in this embodiment; Router pipeline: Divided into three-level pipelines of Deframe (parse the header and distribute it to the corresponding routing module), Arb (arbitration and data extension), and Demux (output according to the routing direction). Each level of pipeline registers is isolated, supporting a clock frequency above 1GHz, and reducing the critical path delay through timing optimization.
[0036] Local module: Process local data packets, convert the AXI-Stream protocol to the AXI protocol through the axis2axi module to achieve on-chip memory read and write; the bpu_NI module processes pulse data and forwards it to the BPU after unpacking; the cpu_NI module realizes efficient communication between the CPU and the NoC through a double-buffering mechanism (CPU_RCV / CPU_TRANS), supporting priority queues and flow control.
[0037] The above is only to illustrate the embodiments of the present invention and is not intended to limit the present invention. For those skilled in the art, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A network-on-chip architecture for ultra-large-scale neuromorphic computing, characterized in that: include: Hybrid routing module, used to implement the three-layer hybrid strategy of source-driven routing, XY-order routing and explicit routing, supporting differentiated processing of pulse data multicast and non-pulse data unicast; Multi-stage pipeline router, using pipeline inter-stage cache optimization and bubble-free flow control design, supports peak operating frequency above 1GHz; The scalable Tile hierarchical structure can achieve the expansion of hundreds of billions of neurons through 2D-Mesh topological cascading; Dual-mode data messages and hybrid flow control mechanism distinguish between pulse and non-pulse data types, and optimize data parsing and forwarding efficiency.
2. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The source driven routing adopts a cluster-based address event representation method, splits the source neuron ID into a routing path group number and a source neuron sub-number, and improves storage utilization by clustering and compressing the routing table.
3. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The XY dimensional order routing is based on 2D-Mesh topology, realizes deterministic path planning through coordinate difference calculation, and supports X / Y priority strategy and large bit width coordinate address (≥30 bits).
4. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The explicit routing embeds complete path information in the data packet, supports dynamic traffic allocation and path detour, and is suitable for irregular topology and fault-tolerant scenarios.
5. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The multi-stage pipeline router includes analysis, routing calculation, switching, and output pipelines, with a single-node peak pulse throughput rate of ≥800Mspikes / s and a processing delay of ≤15ns.
6. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The scalable Tile structure includes a router, a processing unit, and a local memory. A Die is formed by cascading 3×3 Tiles, and then a chip is formed by cascading 2×2 Dies, supporting cross-Die cascade communication.
7. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The bimodal data message adopts a 128-bit width and includes a routing header and a data header. The routing header carries a routing mode identifier, address information and a control field, and the data header includes a read / write address and data length information.
8. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The hybrid flow control mechanism includes: The on-chip router uses the AXI-Stream protocol and supports handshake and back-pressure flow control; The local module uses DMA technology to realize high-speed data interaction between CPU / BPU and NoC, including bpu_NI and cpu_NI interfaces.
9. The network-on-chip architecture for ultra-large-scale neuromorphic computing according to claim 1, characterized in that: The router micro-architecture adopts a multi-stage pipeline design, including parsing, routing calculation, switching, and output pipelines, and integrates FIFO cache and bubble-free flow control mechanism between pipeline stages.