Network-on-chip device and routing method

Through the hexagonal topology structure, XY-XY dimension order routing algorithm and congestion-aware routing algorithm, the scalability and flexibility problems of the on-chip network communication system are solved, and efficient and reliable data transmission is achieved, which is suitable for neuromorphic computing.

CN120017567BActive Publication Date: 2025-09-19PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510476261.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-09-19
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing on-chip network communication systems in neuromorphic computing have problems such as low communication bandwidth, poor scalability, routing deadlock, and inability to flexibly adapt to different congestion scenarios.

Method used

The on-chip network device adopts a hexagonal topology, combines a deterministic XY-XY dimension-order routing algorithm and a fully adaptive congestion-aware routing algorithm, supports asynchronous communication across clock domains, achieves local synchronization through a group thread control mechanism, and designs flexible unicast and multicast routing mechanisms.

Benefits of technology

It improves communication bandwidth and throughput, supports two-dimensional planar and three-dimensional expansion, avoids routing deadlock, dynamically adapts to network congestion, optimizes data transmission paths, reduces latency and improves network efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017567B_ABST
    Figure CN120017567B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of communications and provides an on-chip network device and routing method. The device includes: multiple core nodes in a hexagonal topology structure, each core node includes six communication channel modules and a routing module, and the six communication channel modules are respectively connected to six adjacent core nodes; the routing module includes an input distributor, an output arbiter, a synchronous first-in-first-out queue unit and a local output cache, the input distributor is used to distribute data packets to specific output channels based on the target address of the data packet and the routing algorithm, and the output arbiter is used to select a data packet with a higher priority for transmission when multiple data packets compete for the same output channel; the six communication channel modules are arranged in the XY+, X+, Y-, XY-, X- and Y+ directions in a clockwise direction. The present invention solves the problems of low on-chip network communication bandwidth and difficulty in adapting to different congestion scenarios in the prior art, and realizes a high communication bandwidth and deadlock-free routing mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a network-on-chip device and a routing method. Background Art

[0002] Neuromorphic computing, a cutting-edge technology in the development of artificial intelligence systems, strives to mimic the structure and function of biological neurons to achieve more flexible and adaptive computing models. This computing paradigm offers significant advantages over traditional instruction-sequence-based artificial intelligence algorithms, particularly in applications requiring massively parallel tasks and extremely low power consumption. However, neuromorphic computing also faces significant challenges, particularly in hardware implementation. Due to the complexity of biological neural connections, the unstructured nature of spike data streams, and the unbalanced distribution of model workloads, it is extremely difficult to integrate millions of neurons and billions of synaptic connections on a single neuromorphic chip.

[0003] To address communication challenges in neuromorphic computing, network-on-chip (NoC) architectures have been widely adopted in neuromorphic systems. With advantages such as resource reuse, scalability, distributed parallelism, and event-driven nature, NoC architectures are ideal for communication systems in neuromorphic computing. Among existing NoC technologies, 2D mesh and tree topologies are the two most popular. 2D mesh NoC circuits interconnect communication nodes by arranging them in a two-dimensional grid. However, this structure typically lacks support for three-dimensional inter-chip scalability, and each communication node is connected only to four surrounding nodes, resulting in limited communication bandwidth and severely restricting system throughput and energy efficiency. On the other hand, tree NoC circuits utilize a tree-like structure for node interconnection. While this improves communication efficiency to some extent, it also lacks support for three-dimensional inter-chip scalability and has relatively low intra-chip communication bandwidth. Other NoC circuits based on other topologies, such as single deterministic routing or adaptive routing, typically support only single-mode unicast and multicast routing algorithms, making them difficult to adapt to diverse congestion scenarios. Summary of the Invention

[0004] The present invention provides an on-chip network device and routing method, which solves the problems of low on-chip network communication bandwidth, poor scalability, routing deadlock and inability to flexibly adapt to different congestion scenarios in the prior art, and realizes a high communication bandwidth, deadlock-free routing mechanism and flexible routing algorithm switching.

[0005] The present invention provides a network-on-chip device, comprising:

[0006] Multiple core nodes in a hexagonal topology, each core node comprising six communication channel modules and a routing module, the six communication channel modules being respectively connected to six adjacent core nodes; each core node being located in an independent clock domain, and communication across clock domains being achieved through asynchronous handshaking;

[0007] The routing module includes an input distributor, an output arbiter, a synchronous first-in-first-out queue unit, and a local output buffer. The input distributor is used to distribute the data packet to a specific output channel according to the target address of the data packet and the routing algorithm. The output arbiter is used to select a data packet with a higher priority for transmission when multiple data packets compete for the same output channel.

[0008] The six communication channel modules are arranged in the clockwise direction as XY+, X+, Y-, XY-, X- and Y+ directions; each communication channel module is connected to the routing module, and includes an asynchronous first-in-first-out queue unit, a channel output buffer and an asynchronous handshake unit.

[0009] According to a network-on-chip device provided by the present invention, the communication channel modules in the XY+ and XY- directions include two asynchronous first-in-first-out queue units.

[0010] The present invention further provides a routing method based on any one of the above-mentioned network-on-chip devices, comprising:

[0011] Get the destination address information of the data packet;

[0012] At least one routing algorithm is used to determine the transmission path of the data packet according to the target address information; the routing algorithm includes a deterministic XY-XY dimensional order routing algorithm and a fully adaptive congestion-aware routing algorithm.

[0013] According to the present invention, a routing method based on the network-on-chip device described in any one of the above items is provided, which uses a deterministic XY-XY dimensional order routing algorithm to determine the transmission path of a data packet. Specifically, the method includes: first performing routing in the XY direction according to the target address information to determine the transmission path of the data packet; when the data bits representing the XY direction are cleared, performing routing in the X direction to determine the transmission path of the data packet; when the data bits representing the X direction are cleared, performing routing in the Y direction to determine the transmission path of the data packet; when the data bits representing the XY, X, and Y directions are all cleared, using the local direction as the transmission direction of the data packet to determine the transmission path of the data packet.

[0014] According to the present invention, a routing method based on a network-on-chip device as described in any one of the above items is provided, which uses a fully adaptive congestion-aware routing algorithm to determine the transmission path of a data packet. Specifically, the method includes: determining the direction in which the data packet needs to be transmitted based on the target address information; among the directions in which the data packet needs to be transmitted, selecting the direction with the smallest congestion signal value as the transmission direction of the data packet based on the values ​​of congestion signals from different directions; the smaller the value of the congestion signal, the lower the congestion level in the corresponding direction.

[0015] According to the present invention, a routing method for a network-on-chip device based on any one of the above items, when generating a congestion signal representing the degree of congestion, includes: multiplying the congestion signal value input by the upstream core node by a scaling factor to obtain a first congestion value; multiplying the number of data packets in the first-in-first-out queues of other input directions that exceed a preset probability of constituting an arbitration conflict by the scaling factor to obtain a second congestion value; when a data packet is read out of the first-in-first-out queue of the current input direction, subtracting a movement compensation value from the congestion signal value output by the current input direction to obtain a third congestion value; and adding the number of data packets in the current input direction FIFO, the first congestion value, the second congestion value, and the third congestion value to obtain a congestion signal representing the degree of congestion in the current direction.

[0016] According to a routing method based on an on-chip network device described in any one of the above items provided by the present invention, during the routing process, the method further includes: achieving local synchronization between core nodes in a small group through a group thread control mechanism; the group thread control mechanism monitors and indicates the working status of the computing core nodes and the on-chip network routing by configuring multiple control signals; the control signals include a sync_all signal, an initial_all signal, a done signal and a busy signal; the sync_all signal is used to ensure that the data flow of the core nodes in the group is executed in the order of calculation; the initial_all signal is used to initialize the operation of the core nodes in the group; the done signal is used to reflect the working status of the core nodes; and the busy signal is used to reflect the working status of the on-chip network routing.

[0017] According to a routing method based on a network-on-chip device described in any one of the above items, the present invention provides a method for routing, during the routing process, further comprising: determining the type of multicast based on the duplicate address information in the data packet; the multicast types include one-dimensional multicast, two-dimensional multicast, three-dimensional multicast, and three-dimensional multicast in a two-dimensional plane; transmitting the data packet along the corresponding direction according to the multicast type, and completing the multicast when the data packet is transmitted to all target core nodes.

[0018] According to a routing method for a network-on-chip device based on any one of the above items provided by the present invention, before transmitting a data packet along a corresponding direction according to the type of multicast, the method further includes: performing unicast routing through the current core node to transmit the data packet to the core node in the target direction until the relative address is equal to zero; the relative address is the address offset of the data packet between the current core node and the target core node; when the relative address is cleared, the core node receiving the data packet replicates the data packet; and based on the replicated data packet, continuing to transmit the data packet according to the type of multicast.

[0019] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the routing method based on the on-chip network device as described above is implemented.

[0020] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the routing method based on the on-chip network device as described above is implemented.

[0021] The present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned routing methods based on the network-on-chip device.

[0022] The present invention provides a network-on-chip (NoC) device and routing method, which offer the following benefits: Each core node can communicate with six adjacent nodes through a hexagonal topology, significantly improving communication bandwidth and throughput compared to traditional 2D mesh structures. The hexagonal topology supports both two-dimensional and three-dimensional expansion, enhancing the system's scalability and flexibility, making it suitable for large-scale neuromorphic computing. By introducing a deterministic XY-XY dimensional order routing algorithm, routing deadlocks are avoided, ensuring efficient and reliable data packet transmission. Deterministic routing and a fully adaptive congestion-aware routing algorithm are supported, enabling flexible switching of routing modes based on network congestion, optimizing data transmission paths, reducing latency, and improving network efficiency. Each core node resides in an independent clock domain, and cross-clock domain communication is achieved through asynchronous handshaking, improving resource utilization. Local synchronization ensures the precise computational order of data streams, making it suitable for multi-task parallel processing in neuromorphic computing. The fully adaptive congestion-aware routing algorithm dynamically detects network congestion and selects the optimal path for data packet transmission, reducing average routing latency and improving overall network performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 This is a diagram of the on-chip network structure of the hexagonal topology provided by the present invention.

[0025] Figure 2 It is a structural diagram of the routing module provided by the present invention.

[0026] Figure 3 It is a structural diagram of the input distributor and output arbitrator provided by the present invention.

[0027] Figure 4 It is a flow chart of a routing method based on an on-chip network device provided by the present invention.

[0028] Figure 5 This is an example diagram of a routing deadlock provided by the present invention.

[0029] Figure 6 It is a schematic diagram of the algorithm of the deterministic XY-XY dimensional order routing provided by the present invention.

[0030] Figure 7 This is a schematic diagram of the fully adaptive congestion-aware routing algorithm provided by the present invention.

[0031] Figure 8 It is a schematic diagram of the congestion level characterization signal generation algorithm and output direction selection logic optimization provided by the present invention.

[0032] Figure 9 It is a schematic diagram of the group control signal and multi-thread operation provided by the present invention.

[0033] Figure 10 This is a schematic diagram of the multi-dimensional multicast mechanism provided by the present invention.

[0034] Figure 11 It is a numerical diagram of various indicators of the routing experiment provided by the present invention.

[0035] Figure 12 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0036] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0037] Neuromorphic computing is a key technology in the development of artificial intelligence systems. It aims to develop a novel biomimetic computing architecture to implement specific functions of the human brain, including massively parallel processing and extremely low-power computing paradigms. Based on the structure and function of biological neurons, neuromorphic computing systems can achieve flexibility and adaptability unmatched by traditional artificial intelligence algorithms. However, given the complexity of biological neural connections, unstructured spike data flows, and unbalanced model workload distribution, these promising advantages may be offset by the mismatched communication infrastructure required for large-scale hardware implementations.

[0038] A huge challenge it faces stems from integrating millions of neurons and billions of synaptic connections on a single neuromorphic chip. Unlike continuous one-dimensional instruction sequences, data interactions between neurons are usually performed by discrete pulse sequences over huge spatiotemporal dimensions, which poses great difficulties for the processing of unstructured data streams in communication systems. On-chip network circuits are often used for communication in neuromorphic systems, mainly because of the many advantages of on-chip network architecture, including resource reuse, good scalability, distributed parallelism, and event-driven nature. Common NoC topologies include 2DMesh structure, Torus structure, Ring structure, Star structure, Octagon structure, Spidergon structure, Tree structure, and Butterfly structure. Among them, the most popular ones suitable for neuromorphic processors are 2D Mesh and Tree structures.

[0039] There are three main types of existing on-chip network circuits:

[0040] One is based on 2D Mesh-type on-chip network circuits. However, 2D Mesh-type on-chip network circuits usually do not support three-dimensional expansion between chips, and because each communication node is only connected to four surrounding communication nodes, it has a low communication bandwidth, which seriously restricts the system's throughput and energy efficiency.

[0041] The second is based on the Tree-type on-chip network circuit. However, the Tree-type on-chip network circuit also does not support three-dimensional expansion between chips and has a low intra-chip communication bandwidth.

[0042] The third is the on-chip network circuit based on other topological structures of single deterministic routing or adaptive routing. However, this type of on-chip network circuit only supports single-mode routing unicast and multicast algorithms, and is difficult to adapt to different congestion scenarios.

[0043] 2D mesh-based and tree-based NoC circuits suffer from poor scalability, poor versatility, and low intra-chip communication bandwidth. NoC circuits based on single deterministic routing or other adaptive routing topologies have poor awareness of different congestion scenarios and struggle to maximize communication efficiency. Therefore, this paper proposes a novel hexagonal topology that optimizes throughput, latency, and cost, in which each node can exchange data with its six neighboring nodes. By adhering to the globally asynchronous, locally synchronous (GALS) design principle, processing nodes in the NoC can be controlled as group threads in independent clock domains, running at independent speeds to improve resource utilization. The proposed NoC architecture supports flexible unicast and multicast routing mechanisms, and supports flexible switching between deterministic XY-XY routing and a fully adaptive congestion-aware routing algorithm to adapt to different congestion scenarios.

[0044] The present invention aims to explore a network on chip (NoC) with a hexagonal topology suitable for multi-core neuromorphic computing from the perspective of digital chip design. Based on a sophisticated communication link design, the NoC architecture proposed in the present invention can be expanded in two dimensions along the plane as well as in three dimensions, and adopts the global asynchronous and local synchronous (GALS) design concept to further improve resource utilization. In order to meet the different requirements of data reuse on the chip, the present invention also proposes a flexible unicast and multicast routing mechanism that can freely switch between deterministic routing and adaptive routing modes. The network on chip circuit proposed in the present invention was evaluated in 28nm CMOS and tested in 0.0226mm 2 With a small area overhead, it achieves a maximum throughput of 179.2Gbps and an optimal energy efficiency of 4.872pJ / packet.

[0045] The following combination Figures 1-12 The embodiments of the present invention are described in detail.

[0046] Figure 1 FIG. 1 shows a diagram of a network-on-chip (NoC) structure with a hexagonal topology structure according to an embodiment of the present invention, including:

[0047] Multiple core nodes in a hexagonal topology, each core node includes six communication channel modules and a routing module. The six communication channel modules are connected to six adjacent core nodes respectively. Each core node is located in an independent clock domain, and communication between clock domains is achieved through asynchronous handshake.

[0048] The routing module includes an input distributor, an output arbiter, a synchronous first-in-first-out queue unit, and a local output buffer. The input distributor is used to distribute data packets to specific output channels based on the destination address of the data packet and the routing algorithm. The output arbiter is used to select the data packet with higher priority for transmission when multiple data packets compete for the same output channel.

[0049] The six communication channel modules are XY+, X+, Y-, XY-, X- and Y+ in clockwise direction; each communication channel module is connected to the routing module, including an asynchronous first-in-first-out queue unit, a channel output buffer and an asynchronous handshake unit.

[0050] According to a network-on-chip device provided by the present invention, the communication channel modules in the XY+ and XY- directions include two asynchronous first-in-first-out queue units.

[0051] Specifically, the NoC on a chip with a hexagonal topology structure proposed in the present invention is as follows: Figure 1 Each core node is surrounded by six channels, interconnecting with cores in six directions (defined clockwise as XY+, X+, Y-, XY-, X-, and Y+). Each core node can operate independently in a different clock domain, and cross-domain communication is achieved using asynchronous handshaking. Figure 1 (b) is a two-dimensional plane expansion diagram between chips. When two-dimensional plane expansion is performed between chips, six directions are expanded along the two-dimensional plane, that is, the XY+ direction is connected to the XY- direction of the adjacent chip, the X+ direction is connected to the X- direction of the adjacent chip, and the Y+ direction is connected to the Y- direction of the adjacent chip. Figure 1 Middle (c) is a three-dimensional expansion diagram between chips. When three-dimensional expansion is performed between chips, the XY± directions are used for cross-plane wiring, while the X± and Y± directions are consistent with the two-dimensional plane expansion.

[0052] like Figure 2 As shown in Figure 1, the routing module of each core node has an input distributor and an output arbiter, as well as asynchronous first-in, first-out queues (FIFOs) for six directions, a synchronous FIFO for the local direction, and an output buffer. Input packets first complete clock domain cross-operations in the input FIFO based on their source direction. The input distributor then distributes them to specific output channels based on their tagged address information. Since multiple packets may compete for a single channel, the output arbiter selects the packet with the highest priority as the winner and sends it to the output buffer. Once the adjacent or local FIFO returns an available signal, the output buffer can finally transmit the packet to its desired direction.

[0053] Input distributor such as Figure 3As shown on the left. When the FIFO not empty signal is high, the input distributor will read the FIFO and parse the data according to its marked data type and the configured routing algorithm. Data packets can be divided into two types: single data frames or continuous data frames, the latter of which only carries necessary pulse data to improve information density and follows the same direction as the previous data frame. Next, by analyzing its destination address, the data packet can be executed as unicast or other multicast routing. The output channel of each data packet can be determined by a deterministic XY-XY dimensional order routing algorithm or a congestion-aware routing algorithm for fully adaptive routing. The output arbitrator is as follows Figure 3 As shown on the right, it uses a round-robin scheduling priority queue mechanism, but data in the local direction is given the highest priority to avoid interrupting the local core's computation. In addition, routing requests from the XY±, X±, and Y± directions undergo a mandatory prioritization process. Notably, the data input distributors and output arbiters for all directions can operate in parallel, fully utilizing the communication bandwidth in all directions.

[0054] Figure 4 FIG. 1 is a flow chart of a routing method based on an on-chip network device provided by the present invention, such as Figure 4 As shown, the method includes the following steps:

[0055] S410: Obtain target address information of the data packet.

[0056] S420: Determine a transmission path for the data packet using at least one routing algorithm according to the target address information.

[0057] Routing algorithms include deterministic XY-XY dimensional order routing algorithm and fully adaptive congestion-aware routing algorithm.

[0058] According to the present invention, a routing method for a network-on-chip device based on any of the above items is provided, which uses a deterministic XY-XY dimensional order routing algorithm to determine the transmission path of a data packet. Specifically, the method includes: first performing routing in the XY direction based on the target address information to determine the transmission path of the data packet; when the data bits representing the XY direction are cleared, performing routing in the X direction to determine the transmission path of the data packet; when the data bits representing the X direction are cleared, performing routing in the Y direction to determine the transmission path of the data packet; when the data bits representing the XY, X, and Y directions are all cleared, using the local direction as the transmission direction of the data packet to determine the transmission path of the data packet.

[0059] Specifically, in a NoC with a hexagonal topology, each routing module has limited cache space to store received but unprocessed data packets. When the cache space of node A is full, the data packet sending request sent by node B to node A will not be answered. Node B's data packet will wait until the cache space of node A is available before being sent, thus continuing to occupy the cache space of node B. If a "request-occupancy" cache resource dependency loop is formed between multiple nodes, the so-called deadlock phenomenon will occur. A typical routing deadlock is as follows: Figure 5 As shown in the figure, all six nodes take data from their cache and send it to their neighboring nodes in a clockwise direction. However, the caches of all six nodes are full, so they all enter a waiting state. In a deadlock state, a portion of the data packets in the network will be blocked forever.

[0060] The present invention adopts a deterministic XY-XY dimensional order routing algorithm and a fully adaptive congestion-aware routing algorithm to avoid routing deadlock. The two routing algorithms can be flexibly switched according to the congestion scenario. Each core actually has no concept of coordinates, only the three-dimensional relative coordinate information of XY-XY on the data packet. The coordinate information of the data packet carries a sign bit, indicating the positive and negative directions. For deterministic XY-XY dimensional order routing, routing in the XY direction is performed first, and after the data bits representing the XY direction are cleared, routing in the X direction is performed, and after the data bits representing the X direction are cleared, routing in the Y direction is performed. Finally, after the data bits representing the three directions are cleared, the data packet will enter the local core. In XY-XY dimensional order routing, due to the routing order of XY first, then X and Y, some turns are prohibited, so the routing data packet will not form a loop, and deadlock will not occur. The algorithm of XY-XY dimensional order routing adopted by the present invention is as follows Figure 6 shown.

[0061] According to the present invention, a routing method for a network-on-chip device based on any of the above items is provided, which uses a fully adaptive congestion-aware routing algorithm to determine the transmission path of a data packet, specifically including: determining the direction in which the data packet needs to be transmitted based on the target address information; among the directions in which the data packet needs to be transmitted, selecting the direction with the smallest congestion signal value as the transmission direction of the data packet based on the values ​​of congestion signals from different directions; the smaller the value of the congestion signal, the lower the congestion level in the corresponding direction.

[0062] Specifically, the present invention proposes a fully adaptive routing algorithm that can achieve lower average routing delay under typical multi-core neuromorphic chip data flow compared to traditional adaptive routing algorithms, so as to address the on-chip network congestion problem caused by uneven data flow distribution in neuromorphic chips.

[0063] For the fully adaptive congestion-aware routing algorithm, it uses a pair of counters to indicate the busyness of adjacent routing nodes. When there are multiple choices for the output direction, that is, the (△XYn, △Xn, △Yn) or (#XYn, #Xn, #Yn) address is equal to non-zero, the router can transmit the data packet to the adjacent node with the lowest busy level. In addition, since fully adaptive routing may fall into deadlock without any turning restrictions, the router is also equipped with two virtual channels for the XY± directions to ensure deadlock-free, where data packets in the X+ and Y- directions must be sent to the first virtual channel, while data packets in the X- and Y+ directions must be sent to the second virtual channel. The fully adaptive congestion-aware routing algorithm adopted by the present invention is as follows Figure 7 shown.

[0064] According to the present invention, a routing method for a network-on-chip device based on any one of the above items, when generating a congestion signal representing the degree of congestion, includes: multiplying the congestion signal value input by the upstream core node by a scaling factor to obtain a first congestion value; multiplying the number of data packets in the first-in-first-out queues of other input directions that exceed a preset probability of constituting an arbitration conflict by the scaling factor to obtain a second congestion value; when a data packet is read out of the first-in-first-out queue of the current input direction, subtracting a movement compensation value from the congestion signal value output by the current input direction to obtain a third congestion value; and adding the number of data packets in the current input direction FIFO, the first congestion value, the second congestion value, and the third congestion value to obtain a congestion signal representing the degree of congestion in the current direction.

[0065] Specifically, in traditional adaptive routing algorithms, the congestion degree characterization signal congestion_credit (crd) transmitted between adjacent routing nodes only includes the number of data packets in a certain input direction FIFO, which results in the limited ability of the on-chip network to perceive and avoid congestion. The present invention has made several improvements to this. First, in larger-scale on-chip networks, the average routing distance of data packets is longer. Therefore, when measuring whether a certain direction is congested, it is necessary not only to consider whether the adjacent nodes in that direction are congested, but also whether the nodes farther away in that direction are congested. Therefore, when each node transmits crd_out downstream, it is necessary to multiply the crd_in of the upstream input by the scaling factor α and add it to the crd generated by the number of data packets in the current input direction FIFO. The present invention also pays attention to the delay caused by the arbitration waiting time. Therefore, when generating the crd signal value, the number of data packets in the FIFO of the input direction that may constitute an arbitration conflict will be multiplied by the scaling factor β and added to characterize the degree of congestion that may be caused by the arbitration waiting. In addition, when a data packet is read out of the current input direction FIFO, the crd value output in that direction will be subtracted from the moving_reward value to reflect the difference in congestion level caused by whether the data packet can be read out. Figure 8(a) illustrates the algorithm using crd_out_XY- as an example. In this embodiment, setting α = 0.5, β = 0.5, and moving_reward = 2 can achieve good results while maintaining a low hardware cost.

[0066] The present invention also optimizes the output port selection logic in the input distributor. In traditional adaptive routing, the input distributor directly compares the CRDs sent by several neighboring nodes and selects the minimum value among the legal output directions. However, this will cause several input directions to compete for the same output direction, causing arbitration waiting. Therefore, in the present invention, after receiving the CRD, the input distributor Figure 8 As shown in (b), the straight_reward value will be subtracted from the crd received in the output direction opposite to the input direction, and then each crd value will be sent to the comparator to select the output direction. The selection algorithm is as follows Figure 7 This slightly biases packets towards going straight through the node to avoid arbitration waits. A straight_reward value of 3 works best.

[0067] According to a routing method for an on-chip network device based on any of the above items provided by the present invention, during the routing process, the method further includes: achieving local synchronization between core nodes in a small group through a group thread control mechanism; the group thread control mechanism monitors and indicates the working status of the computing core nodes and the on-chip network routing by configuring multiple control signals; the control signals include a sync_all signal, an initial_all signal, a done signal and a busy signal; the sync_all signal is used to ensure that the data flow of the core nodes in the group is executed in the order of calculation; the initial_all signal is used to initialize the operation of the core nodes in the group; the done signal is used to reflect the working status of the core nodes; and the busy signal is used to reflect the working status of the on-chip network routing.

[0068] Specifically, the present invention follows the design concept of global asynchronous local synchronous, where each node is located in a separate clock domain, thus having inter-node asynchrony. Considering the common situation in neuromorphic computing where a group of core nodes request to interact with data dependencies, the NoC architecture also includes a concise and effective local synchronization mechanism called group thread control. Figure 9As shown in (a), the group thread control mechanism allocates multiple control signals to monitor and instruct the computing cores within the NoC. These signals can be flexibly configured to be connected to adjacent core nodes. Within the scope of the connection, they can be relayed to and affect a certain scale of core nodes, namely the so-called group threads. In each group thread, local synchronization is performed by the sync_all signal, which ensures the precise calculation order of the data flow; the initial operation is performed by the initial_all signal; the working status of the processing core and NoC routing is reflected by the done and busy signals. Multi-threaded operation is as follows Figure 9 (b) of the , which may span multiple chips. In this case, different threads are controlled by different global signal pairs with independent time steps, thus achieving local synchronous intra-group and global asynchronous inter-group communication. Through the group thread control mechanism, this invention provides a synchronous multitasking solution for neuromorphic chips, helping to balance the overall workload on large-scale systems.

[0069] According to a routing method for a network-on-chip device based on any of the above items provided by the present invention, during the routing process, the method further includes: judging the type of multicast based on the duplicate address information in the data packet; the types of multicast include one-dimensional multicast, two-dimensional multicast, three-dimensional multicast, and three-dimensional multicast in a two-dimensional plane; transmitting the data packet along the corresponding direction according to the type of multicast, and completing the multicast when the data packet is transmitted to all target core nodes.

[0070] According to a routing method for a network-on-chip device based on any of the above items provided by the present invention, before transmitting a data packet along a corresponding direction according to the type of multicast, the method further includes: performing unicast routing through the current core node to transmit the data packet to the core node in the target direction until the relative address is equal to zero; the relative address is the address offset of the data packet between the current core node and the target core node; when the relative address is cleared, the core node that receives the data packet copies the data packet; based on the copied data packet, continuing to transmit the data packet according to the type of multicast.

[0071] Specifically, the present invention utilizes the replication address to indicate how many packets should be replicated in each direction. Figure 10 As shown, when there is one, two or three replication addresses equal to non-zero, we can call them one-dimensional multicast, two-dimensional multicast or three-dimensional multicast respectively. In each dimension, the routing node will first perform unicast routing until the relative address is equal to zero. Then, each time the packet is replicated into two copies, one sent to the local core and the other sent to the desired multicast direction. In this way, by performing multiple replications in the intermediate nodes, flexible multicast of any number of core nodes can be achieved, in the form of A, A×B and A×B×C, corresponding to the proposed one-dimensional, two-dimensional and three-dimensional multicast respectively. Note that in Figure 10In (c), by splitting the 3D multicast into three 2D multicasts, 3D multicast can even be performed on a 2D plane.

[0072] The present invention has been successfully designed, synthesized, laid out and taped out in 28nm complementary metal oxide semiconductor (CMOS) technology. Figure 11 shown.

[0073] Through the above scheme, the present invention has the following beneficial effects:

[0074] 1) Compared with 2D Mesh-type NoC circuits, the NoC circuit designed in this invention has higher communication bandwidth, throughput, and energy efficiency, and supports both two-dimensional planar expansion and three-dimensional expansion.

[0075] 2) Compared with the Tree-type NoC circuit, the NoC circuit designed in the present invention has higher communication bandwidth, throughput and energy efficiency, and supports both two-dimensional planar expansion and three-dimensional stereo expansion.

[0076] 3) Compared with other NoC circuits based on single deterministic routing or adaptive routing topologies, the routing algorithm designed in this invention can achieve lower average packet routing delay in adaptive routing mode, and supports free switching between deterministic routing and adaptive routing modes, which can flexibly adapt to different congestion scenarios.

[0077] Figure 12 An example of a physical structure diagram of an electronic device is shown below. Figure 12 As shown, the electronic device may include: a processor 1210, a communications interface 1220, a memory 1230, and a communication bus 1240, wherein the processor 1210, the communications interface 1220, and the memory 1230 communicate with each other via the communication bus 1240. The processor 1210 may invoke logic instructions in the memory 1230 to execute a routing method based on an on-chip network device, the method comprising: obtaining destination address information of a data packet; and determining a transmission path for the data packet using at least one routing algorithm based on the destination address information; the routing algorithms include a deterministic XY-XY dimensional order routing algorithm and a fully adaptive congestion-aware routing algorithm.

[0078] Furthermore, the logic instructions in the aforementioned memory 1230 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0079] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the routing method based on the on-chip network device provided by the above methods, the method including obtaining the destination address information of the data packet; according to the destination address information, using at least one routing algorithm to determine the transmission path of the data packet; the routing algorithm includes a deterministic XY-XY dimensional order routing algorithm and a fully adaptive congestion-aware routing algorithm.

[0080] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the routing method based on the on-chip network device provided by the above-mentioned methods, the method comprising: obtaining the destination address information of the data packet; according to the destination address information, using at least one routing algorithm to determine the transmission path of the data packet; the routing algorithm comprises a deterministic XY-XY dimensional order routing algorithm and a fully adaptive congestion-aware routing algorithm.

[0081] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0082] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A network-on-chip device, characterized in that: include: Multiple core nodes in a hexagonal topology, each core node comprising six communication channel modules and a routing module, the six communication channel modules being respectively connected to six adjacent core nodes; each core node being located in an independent clock domain, and communication across clock domains being achieved through asynchronous handshaking; The routing module includes an input distributor, an output arbiter, a synchronous first-in-first-out queue unit, and a local output buffer. The input distributor is used to distribute the data packet to a specific output channel according to the target address of the data packet and the routing algorithm. The output arbiter is used to select a data packet with a higher priority for transmission when multiple data packets compete for the same output channel. The six communication channel modules are arranged in the clockwise direction as XY+, X+, Y-, XY-, X- and Y+ directions; each communication channel module is connected to the routing module, and includes an asynchronous first-in-first-out queue unit, a channel output buffer and an asynchronous handshake unit.

2. The network-on-chip device according to claim 1, wherein: The communication channel modules in the XY+ and XY- directions include two asynchronous first-in first-out queue units.

3. A routing method based on the on-chip network device according to any one of claims 1-2, characterized in that: include: Get the destination address information of the data packet; Determining a transmission path for the data packet using at least one routing algorithm based on the target address information; The routing algorithms include a deterministic XY-XY dimensional order routing algorithm and a fully adaptive congestion-aware routing algorithm; During the routing process, local synchronization is achieved between core nodes within a group through a group thread control mechanism; the group thread control mechanism monitors and indicates the working status of the computing core nodes and the on-chip network routing by configuring multiple control signals; During the routing process, the type of multicast is determined based on the duplicate address information in the data packet; the multicast types include one-dimensional multicast, two-dimensional multicast, three-dimensional multicast, and three-dimensional multicast on a two-dimensional plane; the data packet is transmitted along the corresponding direction according to the multicast type, and the multicast is completed when the data packet is transmitted to all target core nodes.

4. The routing method according to claim 3, wherein: A deterministic XY-XY dimensional order routing algorithm is used to determine the transmission path of the data packet, including: Perform XY routing based on the target address information to determine the transmission path of the data packet; When the data bits representing the XY directions are cleared, routing in the X direction is performed to determine the transmission path of the data packet; When the data bits representing the X direction are cleared, routing in the Y direction is performed to determine the transmission path of the data packet; When the data bits representing the XY, X, and Y directions are all cleared, the local direction is used as the transmission direction of the data packet to determine the transmission path of the data packet.

5. The routing method according to claim 3, wherein: A fully adaptive congestion-aware routing algorithm is used to determine the transmission path of data packets, including: Determining the direction in which the data packet needs to be transmitted according to the target address information; Among the directions that need to be transmitted, according to the values ​​of the congestion signals from different directions, the direction with the smallest congestion signal value is selected as the transmission direction of the data packet; The smaller the value of the congestion signal is, the lower the congestion level in the corresponding direction is.

6. The routing method according to claim 5, wherein: When generating a congestion signal representing the degree of congestion, it includes: The first congestion value is obtained by multiplying the congestion signal value input by the upstream core node by the scaling factor. The second congestion value is obtained by multiplying the number of packets in the first-in-first-out queue of other input directions that exceed the preset probability of causing arbitration conflicts by the scaling factor; When a data packet is read out of the FIFO queue in the current input direction, the congestion signal value output in the current input direction is subtracted from the motion compensation value to obtain the third congestion value; The number of data packets in the current input direction FIFO, the first congestion value, the second congestion value, and the third congestion value are added together to obtain a congestion signal representing the congestion degree of the current direction.

7. The routing method according to claim 3, wherein: The control signals include a sync_all signal, an initial_all signal, a done signal and a busy signal; The sync_all signal is used to ensure that the data flow of the core nodes in the group is executed in the order of calculation; The initial_all signal is used to initialize the operation of the core nodes in the group; The done signal is used to reflect the working status of the core node; The busy signal is used to reflect the working status of the on-chip network routing.

8. The routing method according to claim 3, wherein: Before transmitting the data packet in a corresponding direction according to the multicast type, the method further includes: Perform unicast routing through the current core node to transmit the data packet to the core node in the target direction until the relative address is equal to zero; the relative address is the address offset of the data packet between the current core node and the target core node; When the relative address is cleared, the core node that receives the data packet replicates the data packet; Based on the copied data packet, data packet transmission is continued according to the multicast type.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the routing method according to any one of claims 3 to 8 is implemented.

Citation Information

Patent Citations

  • Two-dimensional network-on-chip structure, routing method and device thereof, equipment and storage medium

    CN116545960A

  • Hexagonal on-chip network topology structure

    CN118869493A