Method for running and deploying spiking neural networks on hardware in brain-like computers
By using a combination of spatial fill curves and force-guided graph algorithms on brain-like hardware, the problems of scalability and low computing efficiency in large-scale pulse neural network mapping problems are solved, and efficient mapping allocation and low power consumption calculation are achieved.
Patent Information
- Application Number
- CN202210593127.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-27
AI Technical Summary
The prior art is difficult to efficiently deploy pulsed neural networks on brain-like hardware, especially when large-scale network mapping problems, and lacks scalability and portability mapping algorithms, resulting in low computing efficiency and high power consumption.
By combining the spatial fill curve and the force-guided graph algorithm, a directed graph is constructed and topologically sorted to form an initial mapping scheme. Then, through the force-guided graph algorithm iteratively optimized multiple times to obtain the final mapping allocation from neurons to the computing core.
It realizes the acquisition of high-quality mapping solutions in a short time, improves the computing efficiency of brain-like hardware, reduces computing power consumption, and has scalability in hyper-large-scale network mapping problems.
Smart Images

Figure CN115081587B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of neuromorphic computing technology, and specifically relates to a method for running and deploying a spiking neural network on hardware in a brain-like computer. Background Art
[0002] Spiking neural networks are known as the third generation of artificial neural networks. With their core feature of "event-driven" and biological interpretability, they are becoming a hot topic in current artificial intelligence research. Many machine learning applications based on spiking neural networks have been proposed by researchers.
[0003] In order to efficiently research, deploy and run pulse neural network applications, many neuromorphic computing chips or brain-like computing hardware platforms have been designed and produced. These new brain-like hardware architectures vary, and the chip manufacturing process and underlying implementation principles are also different. However, they all follow the same design concept to deploy and run pulse neural network applications, namely:
[0004] 1) Use a large number of specialized brain-like computing cores to simulate neuronal dynamics in parallel.
[0005] 2) Use a high-density on-chip network organization to achieve pulse communication between computing cores.
[0006] Based on the above design concepts, brain-like hardware can realize efficient simulation and operation of pulse neural network applications.
[0007] In order to efficiently utilize these new brain-like hardware resources, it is necessary to map the neurons in the spiking neural network to the computing core of the underlying brain-like hardware. The quality of the spiking neural network mapping scheme will greatly affect the operating efficiency and power consumption of the brain-like hardware.
[0008] Different from the task mapping problem of traditional multi-core systems, the pulse neural network mapping of the pulse neural network side computing process based on pulse transmission to complete reasoning is generally modeled as a multi-objective nonlinear programming problem under constraints. It is necessary to consider multiple indicators such as the total power consumption of pulse transmission between all nodes in the chip (that is, the communication cost), the longest pulse transmission distance, the degree of congestion of communication routes, and the time required to solve the mapping scheme.
[0009] For the graph topology of pulse neural networks, the mapping algorithm is generally divided into two stages: 1) first divide the original pulse neural network and reconstruct the topology; 2) assign and bind each network block to the physical computing core. Since solving the optimal mapping solution under a given optimization goal has been proven to be an NP-Hard problem, how to find an approximate solution in a shorter time under the hardware constraints of the chip is a technical difficulty. Another problem that needs to be solved is the scalability of the mapping algorithm. With the increase in the scale of pulse neural networks that need to be deployed and the increase in the number of underlying chip computing cores and the enhancement of core performance, the search space of the mapping problem will grow exponentially. Therefore, the mapping optimization algorithm needs to have extremely strong scalability to cope with possible ultra-large-scale pulse neural network model mapping problems.
[0010] There are many types of mapping algorithms and frameworks in the industry, but most of them are designed for specific modeling and optimization algorithms for specific brain-like hardware, so they lack portability. At the same time, the existing mapping algorithms lack scalability. When faced with large-scale network mapping problems, they either cannot be solved within an acceptable time or the quality of the solutions obtained is low. Summary of the invention
[0011] In view of the above, the purpose of the present invention is to provide a method for running and deploying a pulse neural network in a brain-like computer on hardware, and to find the optimal mapping scheme from the neurons of the pulse neural network to the computing core of the underlying brain-like hardware through space filling curves and force-guided graphs.
[0012] To achieve the above-mentioned purpose of the invention, an embodiment provides a method for running and deploying a spiking neural network on hardware in a brain-like computer, comprising the following steps:
[0013] Step 1: According to the limitations of the underlying brain-like computing hardware, the neurons of the spiking neural network are divided into clusters, and the clusters are used as nodes. The edges between the nodes are determined according to the directed pulse connection relationship between the clusters to construct a directed graph;
[0014] Step 2, use a level-first topological sorting algorithm to sort the nodes of the directed graph to obtain a node sequence;
[0015] Step 3: label the computing cores of the underlying brain-like computing hardware in the order of the Hilbert space filling curve, and assign each node in the node sequence to the computing core in a one-to-one manner according to the label to form a preliminary mapping scheme;
[0016] Step 4, based on the preliminary mapping scheme, the force-guided graph algorithm is used to iteratively optimize the mapping scheme multiple times to obtain the final mapping scheme;
[0017] Step 5: Implement the mapping allocation of neurons of the pulse neural network to computing cores according to the final mapping scheme.
[0018] In step 1 of one embodiment, neurons of the spiking neural network are clustered according to the limitations of the underlying brain-like computing hardware, including:
[0019] According to the number of loaded neurons of each computing core of the underlying brain-like computing hardware, the neurons of the pulse neural network are divided into clusters to ensure that the number of neurons contained in each cluster does not exceed the number of loaded neurons of the computing core, so as to ensure that each cluster can be assigned to any computing core.
[0020] In step 1 of one embodiment, clusters are used as nodes, and edges between nodes are determined according to impulse connection relationships between clusters to construct a directed graph, including:
[0021] Taking clusters as nodes, for any two first clusters A and second clusters B, if there is at least one directed pulse connection relationship (a, b) in the spiking neural network, and a is a presynaptic neuron and belongs to the first cluster A; b is a postsynaptic neuron and belongs to the second cluster B, then a hyperedge from the first cluster A to the second cluster B is constructed between the first cluster A and the second cluster B as the edge between the nodes; the strength of the edge is set to the total amount of pulses emitted by the neurons in the first cluster A to the neurons in the second cluster B, so as to construct a directed graph.
[0022] In one embodiment, step 2 includes:
[0023] Step 2-1: Count the in-degree and out-degree of each node in the directed graph, and select nodes with in-degree 0 to add to the node sequence, where the number of outward-pointing edges of the node is used as the out-degree of the node, and the number of inward-pointing edges of the node is used as the in-degree of the node;
[0024] Step 2-2, starting with the first node in the node queue, performs the following operations for each node in the node queue: traverse each adjacent node connected to the node, and reduce the in-degree of the adjacent node by 1. If the in-degree of the adjacent node is 0 at this time, add the adjacent node to the end of the node sequence; after completing the above operations for all nodes, the node sequence after the nodes are sorted is obtained.
[0025] In one embodiment, step 3 includes:
[0026] Step 3-1, based on the distribution state of the computing cores of the underlying brain-like computing hardware, a Hilbert space filling curve that conforms to the distribution state is calculated and generated, and a sequential number of each computing core under the Hilbert space filling curve is obtained to form a number sequence;
[0027] Step 3-2, matching the label sequence with the node sequence, that is, according to the order of the labels in the label sequence, allocating the cluster corresponding to each node in the node sequence to the computing core corresponding to each label to form a preliminary mapping scheme.
[0028] In one embodiment, step 4 includes:
[0029] Step 4-1, converting the connection strength between nodes in the preliminary mapping scheme into tensile strength according to the energy field model;
[0030] Step 4-2, for a node pair consisting of an adjacent first node and a second node, calculate the sum of the tensile strengths between the nodes as the tension of the node pair. At this time, the tension of the node pair is equal to the reduction in the total energy of the system after the corresponding nodes are exchanged. Filter all node pairs with positive tension and sort the node pairs in descending order according to the tension to form an exchange sequence;
[0031] Step 4-3, select the first λ% node pairs from the exchange sequence for exchange. For the node pair to be exchanged, first check whether the tension of the node pair is positive. If the tension is positive, exchange the positions of the computing cores corresponding to the two nodes in the node pair, update the tension strength of all other nodes connected to the exchanged node, and record the nodes whose tension strength has changed;
[0032] Step 4-4, after the exchange of the first λ% node pairs is completed, the nodes recorded in step 4-3 and all nodes of the exchange sequence are counted as statistical nodes, and after the tension of the statistical node pairs related to the statistical nodes is calculated, all statistical node pairs with positive tension are screened and sorted in descending order according to the tension to generate a new exchange sequence;
[0033] Step 4-5, using the new exchange sequence as the exchange sequence in step 4-3, repeating steps 4-3 and 4-4 until the new exchange sequence is empty, thereby optimizing the preliminary allocation and obtaining a final mapping solution.
[0034] In one embodiment, step 4-1 includes:
[0035] Step 4-1-1, for each current node, take the computing core position corresponding to each connection node connected to the current node as the origin to build an energy field model, and calculate the energy coefficient of the current node in each energy field model according to the relative coordinates of the computing core corresponding to the current node relative to the origin of each energy field;
[0036] Step 4-1-2, taking the product of the connection strength between the current node and each connected node and the energy coefficient as the energy of the current node relative to each connected node;
[0037] Step 4-1-3, count the sum of the energy of the current node relative to all connected nodes as the energy of the node;
[0038] Step 4-1-4, update the calculation core position corresponding to the current node to the adjacent calculation core positions in the four directions of up, down, left, and right, and calculate the four new relative coordinates relative to the origin of each energy field according to the four new calculation core positions corresponding to the current node;
[0039] Step 4-1-5, for each direction, calculate the new energy coefficient of the current node in each energy field model according to the new relative coordinates, take the product of the connection strength between the current node and each connected node and the new energy coefficient as the new energy of the current node relative to each connected node, and count the sum of the new energies of the current node relative to all connected nodes as the energy of the corresponding direction;
[0040] Step 4-1-6, calculate the difference between the energy of the current node and the energy in the four directions to obtain the tensile strength of the current node in the four directions.
[0041] In step 4-1-1 of an embodiment, the energy field models include three types, and the calculation formulas of the energy coefficient Z corresponding to the three energy field models are respectively:
[0042] The energy coefficient corresponding to the first energy field model is Z=|X|+|Y|;
[0043] The energy coefficient Z corresponding to the second energy field model is Z = (|X| + |Y|) 2
[0044] The energy coefficient corresponding to the third energy field model is Z=|X| 2 +|Y| 2
[0045] Among them, (X, Y) is the relative coordinate of the computing core corresponding to each adjacent node relative to the origin of the energy field;
[0046] When the goal of optimizing the mapping scheme is to reduce the total pulse distance, the first energy field model is selected to calculate the energy coefficient;
[0047] When the goal of optimizing the mapping scheme is to reduce the total pulse distance in addition to reducing the longest pulse distance, the second energy field model or the third energy field model is selected to calculate the energy coefficient.
[0048] In step 4-2 of one embodiment, for a node pair consisting of an adjacent first node and a second node, there will be a tensile strength directed from the first node to the second node, and there will also be a tensile strength directed from the second node to the first node. The tension of the node pair is obtained by summing these two tensile strengths.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] The Hilbert space filling curve used makes full use of the locality of the pulse neural network connection, shortens the physical distance between neurons with pulse connection relationships as much as possible at different distance scales, and provides a high-quality initial allocation plan in a relatively short computing time.
[0051] The force-directed graph algorithm optimizes and adjusts the allocation plan based on the initial allocation plan. Since the tension calculation of node pairs is also local, the complexity upper limit of the force-directed graph algorithm is effectively limited, so it has scalability, so that it can efficiently calculate high-quality allocation plans when facing ultra-large-scale pulse neural network mapping problems, and can simultaneously optimize multiple optimization goals such as minimizing the total pulse distance, reducing the longest pulse distance, and reducing routing congestion. Ultimately, it achieves the goal of reducing the allocation solution time, improving the computing efficiency of brain-like hardware, and reducing computing power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0053] Figure 1 It is a flow chart of a method for running and deploying a spiking neural network on hardware in a brain-like computer provided by an embodiment;
[0054] Figure 2 is a flow chart of an original spiking neural network, a preliminary mapping scheme, and a final mapping scheme provided by an embodiment;
[0055] Figure 3 These are three energy field models provided in the embodiment;
[0056] Figure 4 It is a flowchart of an embodiment of the invention that uses a force-directed graph algorithm to iteratively optimize a mapping solution. DETAILED DESCRIPTION
[0057] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0058] When faced with large-scale spiking neural network mapping problems, there are technical problems such as lack of scalability, inability to effectively calculate mapping solutions, and the lack of portability of mapping algorithms for specialized brain-like hardware. To address this technical problem, an embodiment provides a method for running and deploying a spiking neural network on hardware in a brain-like computer.
[0059] Figure 1 1 is a flow chart of a method for running and deploying a pulse neural network on hardware in a brain-like computer provided by an embodiment. Figure 1 As shown, the method for running and deploying a pulse neural network on hardware in a brain-like computer provided by the embodiment includes the following steps:
[0060] Step 1: According to the limitations of the underlying brain-like computing hardware, the neurons of the pulse neural network are clustered, and the clusters are used as nodes. The edges between the nodes are determined according to the directed pulse connection relationship between the clusters to construct a directed graph.
[0061] In an embodiment, when clustering neurons of a spiking neural network, the neurons of the spiking neural network are clustered according to the number of loaded neurons of each computing core of the underlying brain-like computing hardware, ensuring that the number of neurons contained in each cluster does not exceed the number of loaded neurons of the computing core, so as to ensure that each cluster can be assigned to any computing core.
[0062] In the embodiment, when constructing a directed graph, clusters are used as nodes. For any two first clusters A and second clusters B, if there is at least one directed pulse connection relationship (a, b) in the pulse neural network, and a is a presynaptic neuron and belongs to the first cluster A; b is a postsynaptic neuron and belongs to the second cluster B, then a hyperedge from the first cluster A to the second cluster B is constructed between the first cluster A and the second cluster B as the edge between the nodes. The strength of the edge is set to the total amount of pulses emitted by the neurons in the first cluster A to the neurons in the second cluster B. Specifically, suppose there are N directed pulse connection relationships (a 1 ,b 1 ),(a 2 ,b 2 ),…(a N ,b N ), satisfying a i is the presynaptic neuron and a i Belongs to the first cluster A; b i is the postsynaptic neuron and b i Belong to the second cluster B, and these pre-pulse neurons will each emit a number of pulses, the number of pulses emitted is recorded as w 1 ,w 2 ,…w N , then the total amount of pulses emitted by the neurons in the first cluster A to the neurons in the second cluster B is
[0063] Step 2: Use a level-first topological sorting algorithm to sort the nodes of the directed graph to obtain a node sequence.
[0064] In the embodiment, in step 2, the directed graph is sorted using a level-first topological sorting algorithm to obtain a node sequence, including:
[0065] Step 2-1, for each node in the directed graph, count the in-degree and out-degree of the node, and select nodes with in-degree 0 to add to the node sequence, where the number of outward-pointing edges of the node is used as the out-degree of the node, and the number of inward-pointing edges of the node is used as the in-degree of the node. Assume that there are nodes 1, 2, and 3, where there is an edge L1 from node 1 to node 2 between node 1 and node 2, and there is an edge L2 from node 2 to node 3 between node 2 and node 3. For node 2, edge L1 is an inward-pointing edge, and edge L2 is an outward-pointing edge.
[0066] Step 2-2, starting from the first node in the node queue, perform the following operations for each node in the node queue: traverse each adjacent node connected to the node, and reduce the in-degree of the adjacent node by 1. If the in-degree of the adjacent node is 0 at this time, add the adjacent node to the end of the node sequence; after completing the above operations for all nodes, the node sequence after the nodes are sorted is obtained.
[0067] Step 3: Label the computing cores of the underlying brain-like computing hardware according to the order of the Hilbert space filling curve, and assign each node in the node sequence to the computing core in a one-to-one manner according to the label to form a preliminary mapping scheme.
[0068] In the embodiment, step 3 comprises:
[0069] Step 3-1, based on the distribution state of the computing cores of the underlying brain-like computing hardware, a Hilbert space filling curve that conforms to the distribution state is calculated, and the sequential number of each computing core under the Hilbert space filling curve is obtained to form a number sequence, such as Figure 2 As shown in b, the direction of the arrow is the direction of the label sequence.
[0070] Step 3-2, matching the label sequence with the node sequence, that is, according to the order of the labels in the label sequence, allocating the cluster corresponding to each node in the node sequence to the computing core corresponding to each label to form a preliminary mapping scheme.
[0071] Step 4: Based on the preliminary mapping scheme, the force-guided graph algorithm is used to iteratively optimize the mapping scheme multiple times to obtain the final mapping scheme.
[0072] In the embodiment, step 4 comprises:
[0073] Step 4-1, convert the connection strength between nodes in the preliminary mapping scheme into tensile strength according to the energy field model.
[0074] Specifically, Figure 4 As shown, step 4-1 includes:
[0075] Step 4-1-1, Step 4-1-1, for each current node, take the computing core position corresponding to each connection node connected to the current node as the origin to build an energy field model, and calculate the energy coefficient of the current node in each energy field model according to the relative coordinates of the computing core corresponding to the current node relative to the origin of each energy field;
[0076] In the embodiment, the energy size of the node on different computing cores can be obtained according to different energy field models. In simple terms, the farther away from other nodes with a connection relationship and the greater the connection strength, the higher the energy generated. At this time, the problem of reducing the pulse distance in the optimization mapping scheme can be converted into the problem of reducing the total energy of the system.
[0077] In the embodiment, Figure 3 As shown, three energy field models are provided, among which, Figure 3 As shown in (a), in the first energy field model, the Manhattan distance of the energy and pulse connection relationship is linearly related. This modeling is the simplest and has the fastest calculation speed. Specifically, the energy coefficient Z corresponding to the first energy field model is |X|+|Y|. Figure 3 As shown in (b), in the second energy field model, the square of the Manhattan distance is taken as the calculation base of the energy field, and the energy coefficient corresponding to the second energy field model is Z = (|X| + |Y|) 2 .like Figure 3 As shown in (c), in the third energy field model, the square of the Euclidean distance is taken as the calculation basis of the energy field. The energy coefficient Z corresponding to the third energy field model is |X| 2 +|Y| 2 In the second and third energy field models, the energy will increase significantly when the distance is too far, so as to punish long-distance pulses. Different energy field modeling schemes can be selected according to different optimization target requirements, including: when the goal of the optimization mapping scheme is to reduce the total pulse distance, the first energy field model is selected to calculate the energy coefficient; when the goal of the optimization mapping scheme is to reduce the total pulse distance in addition to reducing the longest pulse distance, the second energy field model or the third energy field model is selected to calculate the energy coefficient.
[0078] Step 4-1-2, taking the product of the connection strength between the current node and each connected node and the energy coefficient as the energy of the current node relative to each connected node;
[0079] Step 4-1-3, count the sum of the energy of the current node relative to all connected nodes as the energy of the node. Assume that, in a directed graph, node 1 has three connected nodes with which it is connected. It should be noted that the connection relationship here considers both inward pointing connections and outward pointing connections. Then the sum of the energy of node 1 relative to the three connected nodes is taken as the energy of node 1.
[0080] Step 4-1-4, update the calculation core position corresponding to the current node to the adjacent calculation core positions in the four directions of up, down, left, and right, and calculate the four new relative coordinates relative to the origin of each energy field according to the four new calculation core positions corresponding to the current node;
[0081] Step 4-1-5, for each direction, calculate the new energy coefficient of the current node in each energy field model according to the new relative coordinates, take the product of the connection strength between the current node and each connected node and the new energy coefficient as the new energy of the current node relative to each connected node, and count the sum of the new energies of the current node relative to all connected nodes as the energy of the corresponding direction;
[0082] Step 4-1-6, calculate the difference between the energy of the current node and the energy in the four directions to obtain the tensile strength of the current node in the four directions.
[0083] Step 4-2, for a node pair consisting of an adjacent first node and a second node, calculate the sum of the tensile strengths between the nodes as the tension of the node pair. At this time, the tension of the node pair is equal to the reduction in the total energy of the system after the corresponding nodes are exchanged. Filter all node pairs with positive tension and sort the node pairs in descending order according to the tension to form an exchange sequence.
[0084] It should be noted that the first node and the second node constituting the node pair are adjacent. The adjacent here is understood to be adjacent in the physical position of the computing core corresponding to the node. There may be a connection relationship between the first node and the second node, that is, there is a directed edge, or there may not be a connection relationship, that is, there is no directed edge.
[0085] After the calculation in step 4-1, the tensile strength of each node in four directions can be obtained. Since the tensile strength is directional, for a node pair consisting of an adjacent first node 1 and a second node 2, there will be a tensile strength from the first node 1 to the second node 2, and there will also be a tensile strength from the second node 2 to the first node 1. The tension of the node pair is obtained by summing these two tensile strengths.
[0086] Step 4-3, select the first λ% node pairs from the exchange sequence for exchange. For the node pair to be exchanged, first check whether the tension of the node pair is positive. If the tension is positive, it means that the total energy of the system can be reduced when the two nodes are exchanged. Otherwise, the total energy of the system will be increased. When the tension is positive, the computing cores corresponding to the two nodes in the node pair are exchanged, and the tension strength of all other nodes connected to the exchanged nodes is updated, and the nodes whose tension strength has changed are recorded.
[0087] In the embodiment, λ can be customized as a hyperparameter. When the first λ% node pairs of the exchange queue are exchanged, the smaller λ is, the more accurate the calculation is, but the lower the calculation efficiency is. The larger λ is, the less accurate the calculation is, and the higher the calculation efficiency is. The appropriate λ value can be set according to the needs. Therefore, according to experimental exploration, the empirical value of λ with better effect is 50.
[0088] In the embodiment, the node pairs are exchanged in real time. After the current node pair is exchanged, it may affect the stress conditions of other node pairs. If the tension of the current node pair is not positive, it means that the results of the exchange of other node pairs have affected the stress conditions of the current exchanged node pair. Continuing the exchange cannot reduce the total energy of the system. Therefore, for the current node pair, before the exchange, it is necessary to judge whether the tension of the node pair is still positive. When it is not positive, the node pair is discarded. When the tension of the node pair is positive, the computing cores corresponding to the two nodes in the node pair are exchanged, and the tension strength of all other nodes connected to the exchanged node is updated, and the nodes whose tension strength changes are recorded.
[0089] It should be noted that the tensile strength of other nodes connected to the exchange node is updated in the manner of step 4-1, and the details are not repeated here.
[0090] Step 4-4, after the exchange of the first λ% node pairs is completed, the nodes recorded in step 4-3 and all nodes of the exchange sequence are counted as statistical nodes, and after the tension of the statistical node pairs related to the statistical nodes is calculated, all statistical node pairs with positive tension are screened and sorted in descending order according to the tension to generate a new exchange sequence.
[0091] Step 4-5, using the new exchange sequence as the exchange sequence in step 4-3, repeating steps 4-3 and 4-4 until the new exchange sequence is empty, thereby optimizing the preliminary allocation and obtaining a final mapping solution.
[0092] like Figure 2 The original pulse neural network to be allocated shown in a is initially allocated through the Hilbert space filling curve in step 3, and the following is obtained: Figure 2 The preliminary mapping scheme shown in b is then locally adjusted using the force-guided graph algorithm in step 4 to obtain Figure 2 The final mapping scheme is shown in c.
[0093] Step 5: Implement the mapping allocation of neurons of the pulse neural network to computing cores according to the final mapping scheme.
[0094] In the embodiment, the neurons in the cluster corresponding to the node are allocated to the computing core corresponding to the node to complete the mapping allocation.
[0095] The above and embodiments provide a method for running and deploying a pulse neural network on hardware in a brain-like computer. Based on the obtained high-quality initial mapping scheme, further adjustments and optimizations are made, and the method has scalability, so that it can efficiently calculate a high-quality mapping scheme when facing the problem of ultra-large-scale pulse neural network mapping, and can simultaneously optimize multiple optimization goals such as minimizing the total pulse distance, reducing the longest pulse distance, and reducing routing congestion. Ultimately, the time for solving the allocation scheme is reduced, the computing efficiency of brain-like hardware is improved, and computing power consumption is reduced.
[0096] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for running and deploying a spiking neural network on hardware in a brain-like computer. It is characterized in that The following steps are involved: Step 1: According to the limitations of the underlying brain-like computing hardware, the neurons of the spiking neural network are divided into clusters, and the clusters are used as nodes. The edges between the nodes are determined according to the directed pulse connection relationship between the clusters to construct a directed graph; Step 2, use a level-first topological sorting algorithm to sort the nodes of the directed graph to obtain a node sequence; Step 3: label the computing cores of the underlying brain-like computing hardware in the order of the Hilbert space filling curve, and assign each node in the node sequence to the computing core in a one-to-one manner according to the label to form a preliminary mapping scheme; Step 4, based on the preliminary mapping scheme, the force-guided graph algorithm is used to iteratively optimize the mapping scheme multiple times to obtain the final mapping scheme; Step 5: Implement the mapping allocation of neurons of the pulse neural network to computing cores according to the final mapping scheme.
2. The method for running and deploying a spiking neural network on hardware in a brain-like computer according to claim 1, It is characterized in that In step 1, the neurons of the spiking neural network are clustered according to the limitations of the underlying brain-like computing hardware, including: According to the number of loaded neurons of each computing core of the underlying brain-like computing hardware, the neurons of the pulse neural network are divided into clusters to ensure that the number of neurons contained in each cluster does not exceed the number of loaded neurons of the computing core, so as to ensure that each cluster can be assigned to any computing core.
3. The method for running and deploying a pulse neural network on hardware in a brain-like computer according to claim 1, It is characterized in that In step 1, clusters are used as nodes, and the edges between nodes are determined according to the impulse connection relationship between clusters to construct a directed graph, including: Taking clusters as nodes, for any two first clusters A and second clusters B, if there is at least one directed pulse connection relationship (a, b) in the spiking neural network, and a is a presynaptic neuron and belongs to the first cluster A; b is a postsynaptic neuron and belongs to the second cluster B, then a hyperedge from the first cluster A to the second cluster B is constructed between the first cluster A and the second cluster B as the edge between the nodes; the strength of the edge is set to the total amount of pulses emitted by the neurons in the first cluster A to the neurons in the second cluster B, so as to construct a directed graph.
4. The method for running and deploying a pulse neural network on hardware in a brain-like computer according to claim 1, It is characterized in that Step 2 includes: Step 2-1: Count the in-degree and out-degree of each node in the directed graph, and select nodes with in-degree 0 to add to the node sequence, where the number of outward-pointing edges of the node is used as the out-degree of the node, and the number of inward-pointing edges of the node is used as the in-degree of the node; Step 2-2, starting with the first node in the node queue, performs the following operations for each node in the node queue: traverse each adjacent node connected to the node, and reduce the in-degree of the adjacent node by 1. If the in-degree of the adjacent node is 0 at this time, add the adjacent node to the end of the node sequence; after completing the above operations for all nodes, the node sequence after the nodes are sorted is obtained.
5. The method for running and deploying a spiking neural network on hardware in a brain-like computer according to claim 1, It is characterized in that Step 3 includes: Step 3-1, based on the distribution state of the computing cores of the underlying brain-like computing hardware, a Hilbert space filling curve that conforms to the distribution state is calculated and generated, and a sequential number of each computing core under the Hilbert space filling curve is obtained to form a number sequence; Step 3-2, matching the label sequence with the node sequence, that is, according to the order of the labels in the label sequence, allocating the cluster corresponding to each node in the node sequence to the computing core corresponding to each label to form a preliminary mapping scheme.
6. The method for running and deploying a spiking neural network on hardware in a brain-like computer according to claim 1, It is characterized in that Step 4 includes: Step 4-1, converting the connection strength between nodes in the preliminary mapping scheme into tensile strength according to the energy field model; Step 4-2, for a node pair consisting of an adjacent first node and a second node, calculate the sum of the tensile strengths between the nodes as the tension of the node pair, select all node pairs with positive tension and sort the node pairs in descending order according to the tension to form an exchange sequence; Step 4-3, select the first λ% node pairs from the exchange sequence for exchange. For the node pair to be exchanged, first check whether the tension of the node pair is positive. If the tension is positive, exchange the positions of the computing cores corresponding to the two nodes in the node pair, update the tension strength of all other nodes connected to the exchanged node, and record the nodes whose tension strength has changed; Step 4-4, after the exchange of the first λ% node pairs is completed, the nodes recorded in step 4-3 and all nodes of the exchange sequence are counted as statistical nodes, and after the tension of the statistical node pairs related to the statistical nodes is calculated, all statistical node pairs with positive tension are screened and sorted in descending order according to the tension to generate a new exchange sequence; Step 4-5, using the new exchange sequence as the exchange sequence in step 4-3, repeating steps 4-3 and 4-4 until the new exchange sequence is empty, thereby optimizing the preliminary allocation and obtaining a final mapping solution.
7. The method for running and deploying a spiking neural network in a brain-like computer on hardware according to claim 6, step 4-1 include: Step 4-1-1, for each current node, take the computing core position corresponding to each connection node connected to the current node as the origin to build an energy field model, and calculate the energy coefficient of the current node in each energy field model according to the relative coordinates of the computing core corresponding to the current node relative to the origin of each energy field; Step 4-1-2, taking the product of the connection strength between the current node and each connected node and the energy coefficient as the energy of the current node relative to each connected node; Step 4-1-3, count the sum of the energy of the current node relative to all connected nodes as the energy of the node; Step 4-1-4, update the calculation core position corresponding to the current node to the adjacent calculation core positions in the four directions of up, down, left, and right, and calculate the four new relative coordinates relative to the origin of each energy field according to the four new calculation core positions corresponding to the current node; Step 4-1-5, for each direction, calculate the new energy coefficient of the current node in each energy field model according to the new relative coordinates, take the product of the connection strength between the current node and each connected node and the new energy coefficient as the new energy of the current node relative to each connected node, and count the sum of the new energies of the current node relative to all connected nodes as the energy of the corresponding direction; Step 4-1-6, calculate the difference between the energy of the current node and the energy in the four directions to obtain the tensile strength of the current node in the four directions.
8. According to the method for running and deploying a pulse neural network in a brain-like computer on hardware in claim 7, in step 4-1-1, the energy field model includes three types, and the calculation formulas of the energy coefficient Z corresponding to the three energy field models are respectively: The energy coefficient corresponding to the first energy field model is Z=|X|+|Y|; The energy coefficient Z corresponding to the second energy field model is Z = (|X| + |Y|) 2 The energy coefficient corresponding to the third energy field model is Z=|X| 2 +|Y| 2 in, (X,Y) are the relative coordinates of the computing core corresponding to each adjacent node relative to the origin of the energy field; When the goal of optimizing the mapping scheme is to reduce the total pulse distance, the first energy field model is selected to calculate the energy coefficient; When the goal of the optimization mapping scheme is to reduce the total pulse distance in addition to reducing the longest pulse distance, the second energy field model or the third energy field model is selected to calculate the energy coefficient.
9. According to the method for running and deploying a pulse neural network in a brain-like computer on hardware in claim 6, in step 4-2, for a node pair consisting of an adjacent first node and a second node, there will be a tensile strength from the first node to the second node, and there will also be a tensile strength from the second node to the first node. The two tensile strengths are summed to obtain the tension of the node pair.
Citation Information
Patent Citations
Gas turbine inlet guide vane system fault diagnosis method based on feature information fusion
CN113850181A
Neuron information visualization method for operating system of brain-like computer
WO2022099557A1