Ultra-low latency, multi-mode on-chip routing method
By adopting multi-mode on-chip routing methods in brain-like computing with multicast and unicast routing modes, the challenges of different data types transmission requirements in brain-like computing are solved, and low-latency and efficient data transmission are achieved.
Patent Information
- Application Number
- PCT/CN2023/140408
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-12
AI Technical Summary
The prior art is difficult to effectively solve the transmission needs of different data types in brain-like computing, resulting in the problems of global transmission load and data communication delay of the chip network.
The ultra-low-latency multi-mode on-chip routing method is used to process pulsed data and non-pulsed data through multicast routing mode and unicast routing mode respectively. The multicast routing mode uses the cached pulse weight index method and the load-aware multicast routing method, and the unicast routing mode adopts a general unicast traffic control method.
It realizes efficient transmission of different data types, reduces the global transmission load and data communication delay of the chip network, and ensures balanced and overlap-free transmission of the data stream.
Smart Images

Figure CN2023140408_12062025_PF_FP_ABST
Abstract
Description
An ultra-low latency multi-mode on-chip routing method Technical Field
[0001] The present invention relates to the field of brain-inspired computing technology, and in particular to an ultra-low latency multi-mode on-chip routing method. Background Art
[0002] Brain-inspired computing, also known as neuromorphic computing, is a general term for computing theories, architectures, chip designs, and application models and algorithms that draw on the information processing patterns and structures of biological neural systems. As a new computing paradigm, it aims to achieve higher-level intelligent computing tasks with greater energy efficiency by mimicking the brain's structure and information processing mechanisms.
[0003] Large-scale brain-inspired computing requires multi-chip networks to deliver high computing power. Inter-chip communication challenges include data transmission latency, traffic congestion, and data synchronization. Designing efficient multi-chip data transmission solutions is crucial to the performance of brain-inspired computing platforms.
[0004] Currently, efficient routing solutions include reducing the global transmission load or data communication delay of chip networks through enhanced multicast routing, localized multicast communication methods, and priority scheduling mechanisms. Among them, multicast routing methods are widely used to transmit pulse data in brain-inspired computing. From the perspective of pulse data transmission, multicast methods have been proven to be more suitable than unicast methods. However, in addition to pulse data, brain-inspired computing also involves other data types such as neuron states, weights, and control words. These data types have regular transmission patterns and are suitable for unicast transmission.
[0005] Therefore, a new technical solution is urgently needed to meet a wider range of data transmission requirements.
[0006] Summary of the Invention
[0007] The purpose of the present invention is to overcome the above-mentioned problems of the prior art and provide an ultra-low latency multi-mode on-chip routing method to solve the technical problems of global transmission load or data communication delay of the chip network caused by using the same method for different pulse data in the prior art.
[0008] The above objectives are achieved through the following technical solutions:
[0009] An ultra-low latency multi-mode on-chip routing method, including multicast routing mode and unicast routing mode:
[0010] The pulse data is transmitted through the multicast routing mode, which uses a cache-like pulse weight indexing method to enable the pulse signal to find target neurons dispersed in different chips, establishes virtual synaptic links between the target neurons, and generates a routing table by connecting the provided target neurons through a load-aware multicast routing method;
[0011] The non-pulse data is transmitted via the unicast routing mode, which uses a general unicast flow control method to establish non-overlapping unicast routing paths for non-pulse signals.
[0012] Furthermore, the pulse data is a data packet carrying information of neurons that transmit pulses; and the transmission of the pulse data is accomplished using a source-driven mechanism.
[0013] Furthermore, the non-pulse data is a data packet of the initial state of the neuron, the synaptic weight and the simulation result, and the transmission of the non-pulse data is completed by adopting a target-driven mechanism.
[0014] Furthermore, in the class cache pulse weight indexing method, the pulse data includes a constant bit, a second bit and a single neuron identification bit, the constant bit is 0, the second bit is effective in the case of non-pulse data unicast, and the single neuron identification bit represents a uniquely identified single neuron.
[0015] Furthermore, the single neuron identification bit includes a routing path group and a source number, and the routing path group represents an inter-chip routing path predefined by a routing table in the platform.
[0016] Furthermore, the routing path group includes a flag bit and an index bit, the routing path group occupies 20 bits, the flag bit occupies 10 bits, and the index bit occupies 10 bits.
[0017] Furthermore, in the general unicast flow control method, the non-pulse data includes a constant bit, a routing priority and a data packet, the constant bit is 1, and the data packet selects a different coordinate comparison order according to the routing priority to change the routing path.
[0018] Furthermore, the data packet includes a chip identifier and a data payload.
[0019] Furthermore, the chip identifier occupies 14 bits, and the data payload occupies 16 bits.
[0020] Furthermore, the generation of the routing table is achieved by running a software program on a server machine. Beneficial effects
[0021] The present invention provides an ultra-low latency multi-mode on-chip routing method, which divides data into pulse data and non-pulse data according to data type. Pulse data uses a cache-like pulse weight indexing method to quickly search the multicast routing table, and generates a multicast routing table through a load-aware multicast routing method to balance the global inter-chip communication load and reduce peak traffic. Non-pulse data is transmitted through a universal unicast flow control method, which can ensure that when multiple external communication links are exported simultaneously, the data streams of different external communication links will not overlap in the same direction of the same chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] FIG1 is an architecture diagram of an ultra-low latency multi-mode on-chip routing method according to the present invention;
[0023] FIG2 is a schematic diagram of generating a routing table in a load-aware multicast routing method in an ultra-low latency multi-mode on-chip routing method according to the present invention;
[0024] FIG3 is a schematic diagram of pulse bits in a cache-like pulse weight indexing method in an ultra-low latency multi-mode on-chip routing method according to the present invention;
[0025] FIG4 is a schematic diagram of a non-pulse bit in a general unicast flow control method in an ultra-low latency multi-mode on-chip routing method according to the present invention;
[0026] FIG5 is a schematic diagram of a data flow path in a general unicast flow control method in an ultra-low latency multi-mode on-chip routing method described in the present invention. DETAILED DESCRIPTION
[0027] The present invention will be further described in detail below with reference to the accompanying drawings and examples. The described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.
[0028] This solution provides an ultra-low latency multi-mode on-chip routing method, including multicast routing mode and unicast routing mode:
[0029] According to the type of data to be transmitted between chips, the data is divided into pulse data and non-pulse data, where:
[0030] The pulse data is a data packet carrying information about neurons that transmit pulses; in this embodiment, the transmission of the pulse data is accomplished using a source-driven mechanism;
[0031] The non-pulse data is a data packet of the initial state of the neuron, the synaptic weight and the simulation result; the transmission of the non-pulse data in this embodiment is completed by a target-driven mechanism;
[0032] The pulse data is transmitted through the multicast routing mode, which uses a cache-like spike-weight indexing method (CSWI) to enable the pulse signal to find target neurons distributed in different chips, establish virtual synaptic links between the target neurons (i.e., virtual synaptic links are established between neurons through a cache-like three-layer indexing data structure, which helps index multiple weight data triggered by a single received pulse), and generate a routing table by connecting the provided target neurons through a load-aware multicast routing method (LAMR).
[0033] The non-pulse data is transmitted through the unicast routing mode, and the unicast routing mode adopts a general unicast flow control method (GUFC-General Unicast Flow Control) to establish a non-overlapping unicast routing path for the non-pulse signal.
[0034] In brain-inspired computing, millions of neurons may interact simultaneously. Neurons are mapped to multiple chips and communicate with each other through chip networks. A spike generated by a single neuron in a chip may reach multiple target chips. Neurons are dispersed across different chips.
[0035] As a specific introduction to pulse data transmission, the pulse data refers to a data packet carrying information about neurons that transmit pulses.
[0036] Address Event Representation (AER) is a classic approach to emulating the behavior of biological neurons on multi-chip platforms. In this approach, the physical connections between neurons are replaced by virtual links in the routing information of each chip.
[0037] Currently, address event representation mainly uses two types of routing mechanisms: source-driven and target-driven to establish virtual links between chips.
[0038] When the pulse signal transmission adopts the source-driven mechanism, all chips must store all possible received routing paths in local memory, and the source chip only transmits its source address;
[0039] When using the target-driven mechanism, the chip does not have to store all routing paths.
[0040] The transmission of the pulse data in this embodiment is accomplished by using a source drive mechanism.
[0041] As a detailed introduction to multicast routing of pulse data, it can be seen from the above that the pulse data transmission in this scheme is based on source-driven packet representation and multicast mode. Therefore, the routing scheme must store a routing table for all possible routing paths.
[0042] The proposed CSWI-Cache-like Spike-Weight Indexing scheme adopts a cache-like structure to achieve fast search of multicast routing tables.
[0043] The load-aware multicast routing method (LAMR) is used to generate a multicast routing table to balance the global inter-chip communication load and reduce peak traffic.
[0044] Specifically, the load-aware multicast routing method (LAMR) is an algorithm for generating a routing table based on provided neuron connections.
[0045] Assuming that neurons have the same firing rate, the historically generated routing paths can be used to estimate the load of the network, which allows the routing algorithm to search for a load-balanced routing path while ensuring that it is also the shortest.
[0046] As a detailed introduction to the load-aware multicast routing method (LAMR), its main ideas include:
[0047] (1) A chip that obtains routing packets in the multiplexing path;
[0048] (2) Iteratively select the destination chip so that the path grows from the source;
[0049] (3) Select the path through breadth-first search to ensure load balancing.
[0050] As shown in Figure 2, an example is used to illustrate how the load-aware multicast routing method (LAMR) generates routing paths. The routing table is formed through an iterative process of breadth-first search based on historical load information.
[0051] src is the source chip that transmits the pulse signal.
[0052] dst is the chipset that should receive the peak,
[0053] obt represents the chipset on the historical routing path.
[0054] In the figure, chip 4 is the source chip that transmits the pulse, and the other gray chips (chips 1, 7, 12, 27, 28, 32, 36, 40, 44, 48, 55, 56, 57, 58, 61 and 62) are all target chips, which should receive the pulse signal sent by chip 4.
[0055] This algorithm will iteratively perform a breadth-first search (BFS) method until all paths to the target chip are established.
[0056] The load-aware generation mode is primarily implemented using a breath-first search (BFS) method based on prior load information. This problem involves finding a path between nodes, which can be visualized as a directed graph or tree structure. In this embodiment, the search is hop-by-hop, which can be constructed into a tree structure. Therefore, the classic breath-first search (BFS) method was chosen.
[0057] Furthermore, the routing algorithm reuses routing paths by updating the chipsets along historical routing paths. The routing table is generated by a software program running on a server machine and should not change during a series of brain-inspired computing simulations. When neuronal connectivity changes, the routing table must be regenerated and downloaded to the platform's routing table storage device before a new series of simulations can begin.
[0058] The cache-like spike-weight indexing method (CSWI) described in this embodiment is a method for enabling a spike signal to find target neurons dispersed in different chips.
[0059] This embodiment is based on this approach, so that the physical connections between biological neurons can be replaced by virtual links maintained by the routing table.
[0060] As shown in FIG3 , in the cache-like pulse weight indexing method, the pulse data includes a constant bit, a second bit, and a single neuron identification bit;
[0061] Wherein, the constant bit is 0 to distinguish pulse data from non-pulse data;
[0062] The second bit (Routing Priority (uncast)) is effective only in the case of non-pulse data unicast;
[0063] The single neuron ID represents a unique identifier of a single neuron.
[0064] The individual neuron identification bits include a routing path group and a source number. The routing path group represents the inter-chip routing path predefined by the routing table in the platform. The routing path group includes a tag bit and an index bit. The routing path group occupies 20 bits, the tag bit occupies 10 bits, and the index bit occupies 10 bits.
[0065] Specifically, from a high-level perspective, the remaining bits can be regarded as a single neuron ID in the platform. This ID is unique for each neuron and can represent up to 230 neurons.
[0066] From the bottom layer, pulse data can be divided into two parts: routing path group and source number.
[0067] The routing path group occupies 20 bits and represents the inter-chip routing path pre-defined by the routing table in the platform.
[0068] Therefore, the underlying representation of a spike can be interpreted as the following statement: neurons whose number is equal to the number of sources will transmit spikes through the routing paths indexed by the tag bit and the index bit.
[0069] As a specific introduction to pulse data transmission, the platform also has non-pulse data to support brain-like computing simulation. The data types include the initial state of neurons, synaptic weights, and simulation results.
[0070] Specifically, the transmission of non-pulse data mainly occurs when most chip networks are idle; and non-pulse data has the inherent properties of simple routing patterns and low latency requirements.
[0071] The transmission of the non-pulse data in this embodiment is accomplished by using a target-driven mechanism.
[0072] As a detailed introduction to unicast routing for non-pulse data, the general unicast flow control method (GUFC) adopted in this solution is designed to support non-pulse data routing, establish non-overlapping unicast routing paths, and share communication resources with multicast routing.
[0073] Among them, the routing direction of the unicast message is obtained through the coordinate bit in the message, without accessing the routing table; the routing priority bit in the message allows the message to select different routing paths.
[0074] As shown in FIG4 , in the general unicast flow control method (GUFC), the representation bits of the non-pulse data include a constant bit, a routing priority, and a data packet;
[0075] Among them, the most significant bit constant 1 indicates that it is a non-pulse data packet;
[0076] The second bit is the routing priority, which is used by the data packet to select a different coordinate comparison order to change the routing path.
[0077] The data packet includes a chip identifier and a data payload. The chip identifier occupies 14 bits and carries the destination coordinate information. This scheme can represent up to 214.
[0078] This isn't a huge number when it comes to large-scale brain-inspired computing simulations. However, when there are multiple external interfaces, the chip can be partitioned and the encoding can be restarted from the origin (0,0) to avoid bit representation overflow. Therefore, the solution can still be expanded based on some external connection requirements.
[0079] Assuming that the brain-inspired computing chip network is square, the row number (Row) and the column number (Column) have the same bit width, and the remaining bits of the data packet are used as the data payload (Data Payload), which occupies 16 bits.
[0080] Fine-grained management of routing paths can prevent data loss caused by communication congestion. The General Unicast Flow Control (GUFC) solution includes a cross-shaped unicast routing pattern. This pattern ensures that when multiple external communication links are simultaneously exported, data flows from different external communication links do not overlap in the same direction on the same chip. The key trick is that all data input by each external communication link is routed to only the chip in a specific row.
[0081] As a specific embodiment of this solution, FIG5 shows a cross-shaped path of download and upload data flows in a 4×8 network with four external links. The cross points are chip0, chip5, chip10, and chip15.
[0082] There are no overlapping paths between different external links, and the number of chips in different external links is also equal, which means that the traffic load of each link is also the same.
[0083] In the figure, for a 4×8 two-dimensional grid chip array, if four external communication links are connected, the destination chip for the data sent by the first external link will only be all the data in the first row;
[0084] Similarly, the second link corresponds to the second row, the third link corresponds to the third row, and the fourth link corresponds to the fourth row; forming a cross-shaped communication link with multiple chips as intersections.
[0085] Based on this data flow path, the download and upload data flows between different links will not overlap.
[0086] This routing approach is also scalable when the number of rows in the network is different from the number of external interfaces:
[0087] (1) When the number of rows is less than the number of external links, only some of the external links are used; each row will be assigned a separate external link, following the same pattern as the same case.
[0088] (2) When the number of rows is greater than the number of external links, the general unicast flow control GUFC scheme can still be applied to the platform by cascading the extra rows.
[0089] This is a common situation because chips are often scaled up to form large arrays for large-scale brain-like computing simulations. For example, an 8-row network can be viewed as a cascade of two 4-row networks.
[0090] The above description is only for explaining the embodiments of the present invention and is not intended to limit the present invention. For those skilled in the art, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An ultra-low latency multi-mode on-chip routing method, characterized in that, it includes a multicast routing mode and a unicast routing mode: The pulse data is transmitted through the multicast routing mode. The multicast routing mode enables the pulse signal to find the target neurons scattered in different chips through a cache-like pulse weight indexing method, establishes virtual synaptic links between the target neurons, and generates a routing table by connecting the provided target neurons through a load-aware multicast routing method; The non-pulse data is transmitted through the unicast routing mode. The unicast routing mode uses a general unicast traffic control method to establish non-overlapping unicast routing paths for non-pulse signals.
2. An ultra-low latency multi-mode on-chip routing method according to claim 1, characterized in that, the pulse data is a data packet carrying neuron information with emitted pulses; the transmission of the pulse data is completed by a source-driven mechanism.
3. An ultra-low latency multi-mode on-chip routing method according to claim 1, characterized in that, the non-pulse data is a data packet of neuron initial state, synaptic weight, and simulation results, and the transmission of the non-pulse data is completed by a target-driven mechanism.
4. An ultra-low latency multi-mode on-chip routing method according to claim 1, characterized in that, in the cache-like pulse weight indexing method, the pulse data includes a constant bit, a second bit, and a single neuron identification bit. The constant bit is 0, the second bit functions in the case of unicast of non-pulse data, and the single neuron identification bit represents a unique single neuron.
5. An ultra-low latency multi-mode on-chip routing method according to claim 4, characterized in that, the single neuron identification bit includes a routing path group and a source number, and the routing path group represents an inter-chip routing path predefined by the routing table in this platform.
6. An ultra-low latency multi-mode on-chip routing method according to claim 5, characterized in that, the routing path group includes a flag bit and an index bit. The routing path group occupies 20 bits, the flag bit occupies 10 bits, and the index bit occupies 10 bits.
7. An ultra-low latency multi-mode on-chip routing method according to claim 1, characterized in that, in the general unicast traffic control method, the non-pulse data includes a constant bit, a routing priority, and a data packet. The constant bit is 1, and the data packet selects different coordinate comparison orders through the routing priority to change the routing path.
8. An ultra-low latency multi-mode on-chip routing method according to claim 7, characterized in that, the data packet includes a chip identifier and a data payload.
9. An ultra-low latency multi-mode on-chip routing method according to claim 8, characterized in that, the chip identifier occupies 14 bits, and the data payload occupies 16 bits.
10. An ultra-low latency multi-mode on-chip routing method according to claim 1, characterized in that, the generation of the routing table is realized by running a software program on a server machine.
Citation Information
Patent Citations
Method for efficiently transmitting pulse data packet in brain-like computer
CN111082949A
Network-on-chip routing communication method for brain-like processor and network-on-chip
CN112468401A
Shortest path route generation method based on load weighting
CN112511445A
Conflict-free, stall-free, broadcast network on chip
US20220121951A1
Neurosynaptic core and method of operating the same
WO2023107004A2