A Data Mapping Method from Pulse Convolutional Neural Network to Brain-Inspired Computing Chip
By mapping pulse convolution neural networks to the data mapping method of brain-like computing chips, the problems of high algorithm complexity and low solution efficiency in the existing methods are solved through computational graph segmentation and inter-core communication optimization, and the effect of reducing communication delay and improving computing efficiency is achieved.
Patent Information
- Application Number
- CN202510378277.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing methods of mapping pulse convolutional neural networks (SCNNs) to brain-like computing chips are difficult to make full use of the structural characteristics of CNNs, resulting in high algorithm complexity and low solution efficiency.
A data mapping method from pulse convolution neural network to brain-like computing chip is proposed. Through the steps of calculating graph segmentation, inter-core communication modeling, calculating cost function and optimizing cost function, the automatic parallel deployment of SCNN is realized. This method will synthesize the computational graphs in parallel, optimize inter-core communication using a three-dimensional packing model, and select the optimal segmentation scheme with the lowest communication latency.
Through this method, communication delay can be reduced, computing efficiency can be improved, SCNN automatic parallel deployment based on brain-like computing chips can be realized, and communication efficiency can be improved.
Smart Images

Figure CN119886235B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network chips, and in particular relates to a data mapping method from a pulse convolutional neural network to a brain-like computing chip. Background Art
[0002] Spiking Neural Networks (SNN) are called the "third generation of neural networks". They simulate the pulse emission mechanism of biological neurons, which enables them to more accurately simulate the information processing method of the human brain. Compared with traditional artificial neural networks, spiking neural networks show higher low power consumption and flexibility in processing timing information and dynamic changes. By deploying on brain-like computing chips, they can achieve low power consumption and real-time performance. Convolutional Neural Networks (CNN) capture local features through convolution kernels, and through multi-layer convolution, they can gradually extract multi-scale features from low-level to high-level. Therefore, many researchers have combined spiking neural networks and convolutional neural networks to construct spiking convolutional neural networks (SCNN), which have achieved good results in tasks such as target detection.
[0003] Most of the brain-like computing chips that have been developed are built on a chip network (NoC), which efficiently connects the many neuron cores in the platform. Mapping SCNN to brain-like computing chips is a key step in realizing SCNN deployment. The general mapping solution is to divide SCNN into multiple sub-blocks and place them in the neuron core. It is necessary to ensure that the number of neurons in each sub-block cannot exceed the capacity of the neuron core, and the communication delay between cores is as low as possible. At present, some mapping methods have emerged to deploy SNN on brain-like platforms, but these methods are difficult to fully utilize the structural characteristics of CNN, which makes their algorithm complexity high and the solution efficiency low.
[0004] According to the structural characteristics of CNN, one method of parallelization is data parallelism. Data parallelism refers to the division of operator input data, thereby splitting operators, distributing data from multiple slices to each core of the same device, and achieving parallelism through multi-core computing. Through data parallelism, the input and output neuron clusters of the convolutional layer and pooling layer are divided into multiple slices and placed in multiple neuron cores, thereby distributing the communication load of a single neuron core to multiple neuron cores, improving the utilization of communication resources in NoC, and thus reducing communication delays. As a dedicated integrated circuit, brain-like chips lack a convenient and efficient parallel programming interface, making it difficult to efficiently implement SCNN parallelization. Summary of the invention
[0005] To solve the above problems, the present invention proposes a data mapping method from a spiking convolutional neural network to a brain-inspired computing chip, which realizes the automatic parallel deployment of the spiking convolutional neural network based on the brain-inspired computing chip, can reduce communication latency, and improve computing efficiency.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A data mapping method from a spiking convolutional neural network to a brain-inspired computing chip includes the following steps:
[0008] S1. Computational graph partitioning: Parallelly partition the data of the computational graph to obtain each subgraph;
[0009] S2. Inter-core communication modeling: Model the inter-core communication as a three-dimensional bin-packing model, and the three-dimensional bin-packing model includes mapping neuron cores and statistical communication loads;
[0010] The specific process of step S2 is as follows:
[0011] S21. Mapping neuron cores: Place the neuron clusters in the partitioned subgraphs on adjacent neuron cores in topological order, so as to satisfy the neuron core capacity constraint for the partitioned subgraphs and maximize the communication efficiency between adjacent neuron clusters;
[0012] S22. Statistical communication load: According to the mapping of neuron groups to neuron cores and the static routing algorithm, statistically calculate the communication load of each node on the on-chip network; among them, for the on-chip network with a Mesh network structure of size , the communication load is located in the two-dimensional space of the on-chip network. According to the static routing algorithm, the calculation formula for the communication load of each node is: = , , where in the formula, is the communication load from core to core j ; is the sending pulse core; j is the receiving pulse core; is the number of pulses sent from core to core j ; is the communication load of node ; is the abscissa of the two-dimensional grid NoC, is the ordinate of the two-dimensional grid NoC, is the node communication load coordinate, constitutes a three-dimensional space; is the summation function; is the abscissa interval of the two-dimensional grid NoC; is the vertical coordinate interval of the two-dimensional grid NoC;
[0013] S3. Calculate the cost function:
[0014] Using three-dimensional space, model the communication optimization problem as a three-dimensional bin-packing problem. The optimization objective is to minimize the maximum height. The calculation formula of the cost function is: , where is the maximum communication load of all nodes; is the function to find the maximum value;
[0015] S4. Optimize the cost function: Based on the cost function, use an enumeration algorithm or a heuristic algorithm to select the optimal segmentation scheme, and then execute step S1 to obtain the final parallel data flow graph.
[0016] Preferably, the specific process of step S1 is as follows:
[0017] S11. Calculate the shape of the neuron cluster: Model the neural network as a directed acyclic graph. The points in the directed acyclic graph are used to execute operators, and the edges represent the data dependencies between operators. According to the directed acyclic graph model, starting from the input and the first layer, the shapes of each neuron cluster are inferred layer by layer. Then, based on the input and parameters of each layer, calculate the output size of the sub-operator to obtain the shape of the neuron cluster;
[0018] S12. Calculate the input size: Calculating the input size is the inverse operation of calculating the output size. According to the output shape of the sub-operator, calculate the input size of the sub-operator;
[0019] S13. Perform splitting: According to the output size and input size of the sub-operator, perform operator splitting, cut the input and allocate it to each sub-operator, and analyze whether the output size of the previous sub-operator matches the input size of the current sub-operator; if the output size of the previous sub-operator does not match the input size of the current sub-operator, the current sub-operator needs to use part of the output of the previous sub-operator in the adjacent subgraph as input and introduce a connection between the two subgraphs; if the output size of the previous sub-operator matches the input size of the current sub-operator, there is no need to obtain input from the adjacent subgraph; among them, when initializing the splitting, cut the directed acyclic graph into n subgraphs, and n is the minimum value that satisfies the neuron core capacity constraint.
[0020] Preferably, the specific process of step S4 is as follows: Determine the total number of operator segmentation schemes according to the calculated shape of the neuron cluster; Use the total number of operator segmentation schemes as the solution space size of single-objective optimization, and give a threshold of the solution space size as a hyperparameter; if the total number of operator segmentation schemes is less than or equal to the threshold and the solution time is acceptable, use an enumeration algorithm to traverse the solution space to find the optimal solution; if the total number of operator segmentation schemes is greater than the threshold and the solution time is unacceptable, use a heuristic algorithm to find a feasible solution.
[0021] Preferably, the value range of the threshold is 1000 to 10000.
[0022] Preferably, the specific calculation process of the enumeration algorithm is as follows: traverse all partitioning schemes, and in each iteration, calculate the communication load under the current scheme according to the communication load calculation formula of the on-chip network in step S2, so as to find the minimum communication load and the corresponding partitioning scheme.
[0023] Preferably, the heuristic algorithm is implemented by using an adaptive hill-climbing algorithm, and the specific calculation process is as follows:
[0024] S41. Select the best neighbor solution: Select the solution with the highest evaluation value from the neighbor solutions as the best neighbor solution;
[0025] S42. Update the current solution: Set the best neighbor solution as the new current solution;
[0026] S43. Adaptive adjustment: Adjust the search strategy according to the performance of the current solution and the neighbor solutions. The adjustment of the search strategy includes changing the step size and changing the neighborhood structure;
[0027] S44. Check the termination condition: Judge whether the termination condition is satisfied. The termination condition is to reach the maximum number of iterations or the quality of the solution reaches the preset threshold;
[0028] S45. Output the optimal solution: If the termination condition is satisfied, output the current optimal solution and end the process; if the termination condition is not satisfied, return to step S1 to continue the search.
[0029] After adopting the above technical solutions, the present invention has the following beneficial effects: The present invention provides a data mapping method from a pulse convolutional neural network to a brain-inspired computing chip. The ultimate goal of data parallel mapping is to explore the parallel computing scheme of pulse convolutional neural network operators during the compilation process. According to the proposed communication cost function, select the optimal partitioning scheme with the lowest communication delay, so as to map the pulse convolutional neural network to multi-core computing resources. All the partitioned sub-operators are executed in parallel in the brain-inspired computing chip according to the parallel data flow graph, which is used to increase the utilization rate of communication resources and improve the inter-core communication efficiency, realize the automatic parallel deployment of the pulse convolutional neural network based on the brain-inspired computing chip, reduce the communication delay, and improve the computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flowchart of the present invention;
[0031] Figure 2 is a process framework diagram of the present invention;
[0032] Figure 3 is a diagram of the inter-core communication modeling process of the present invention;
[0033] Figure 4 It is a process diagram for selecting the optimal segmentation scheme of the present invention;
[0034] Figure 5 It is a process diagram of the adaptive hill climbing algorithm of the present invention;
[0035] Figure 6 It is an example diagram of data parallelism of the pulsed convolutional neural network of the present invention. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] As Figures 1 to 6 shown, a data mapping method from a pulsed convolutional neural network to a brain-like computing chip includes the following steps:
[0038] S1. Computational graph segmentation: Parallelly segment the data of the computational graph to obtain each subgraph;
[0039] The specific process of step S1 is as follows:
[0040] S11. Calculate the shape of neuron clusters: Model the neural network as a directed acyclic graph. The points in the directed acyclic graph are used to execute operators, and the edges represent the data dependencies between operators; According to the directed acyclic graph model, starting from the input and the first layer, the shapes of each neuron cluster are inferred layer by layer, and then according to the input and parameters of each layer, the output size of the sub-operator is calculated to obtain the shape of the neuron cluster;
[0041] S12. Calculate the input size: Calculating the input size is the inverse operation of calculating the output size. According to the output shape of the sub-operator, the input size of the sub-operator is calculated;
[0042] S13. Perform segmentation: According to the output size and input size of the sub-operator, perform operator segmentation, cut the input and allocate it to each sub-operator, and analyze whether the output size of the previous sub-operator matches the input size of the current sub-operator; If the output size of the previous sub-operator does not match the input size of the current sub-operator, the current sub-operator needs part of the output of the previous sub-operator in the adjacent subgraph as input, and a connection is introduced between the two subgraphs; If the output size of the previous sub-operator matches the input size of the current sub-operator, there is no need to obtain input from the adjacent subgraph; Among them, when initializing the segmentation, the directed acyclic graph is segmented into n subgraphs, and n is the minimum value that satisfies the neuron core capacity constraint;
[0043] S2. Inter-core communication modeling: Model the inter-core communication as a three-dimensional bin packing model, and the three-dimensional bin packing model includes mapping neuron cores and statistical communication loads;
[0044] The specific process of step S2 is as follows:
[0045] S21. Map neuron cores: Place the neuron clusters in the segmented subgraphs onto adjacent neuron cores in topological order, so that the segmented subgraphs meet the neuron core capacity constraint and maximize the communication efficiency between adjacent neuron clusters;
[0046] S22. Statistically calculate the communication load: According to the mapping of neuron groups to neuron cores and the static routing algorithm, statistically calculate the communication load of each node on the on-chip network; among them, for the on-chip network with a Mesh network structure of size , the communication load is located in the two-dimensional space of the on-chip network. According to the static routing algorithm, the calculation formula for the communication load of each node is: = , , where in the formula, is the communication load from core to core j ; is the core that sends pulses; j is the core that receives pulses; is the number of pulses sent from core to core j ; is the communication load of node ; is the abscissa of the two-dimensional grid NoC, is the ordinate of the two-dimensional grid NoC, is the coordinate of the node communication load, constitutes a three-dimensional space; is the summation function; is the abscissa interval of the two-dimensional grid NoC; is the ordinate interval of the two-dimensional grid NoC;
[0047] S3. Calculate the cost function:
[0048] Using the three-dimensional space, model the communication optimization problem as a three-dimensional bin packing problem. The optimization goal is to minimize the maximum height. The calculation formula for the cost function is: , where in the formula, is the maximum communication load of all nodes; is the function to find the maximum value;
[0049] S4. Optimize the cost function: Based on the cost function, use an enumeration algorithm or a heuristic algorithm to select the optimal segmentation scheme, and then execute step S1 to obtain the final parallel data flow graph;
[0050] The specific process of step S4 is as follows: Determine the total number of operator segmentation schemes according to the calculated shape of the neuron cluster; Use the total number of operator segmentation schemes as the solution space size of single-objective optimization, and give a threshold of the solution space size as a hyperparameter; If the total number of operator segmentation schemes is less than or equal to the threshold and the solution time is acceptable, use the enumeration algorithm to traverse the solution space to find the optimal solution; If the total number of operator segmentation schemes is greater than the threshold and the solution time is unacceptable, use the heuristic algorithm to find a feasible solution;
[0051] The value range of the threshold is 1000~10000;
[0052] The specific calculation process of the enumeration algorithm is as follows: Traverse all segmentation schemes, and in each iteration, calculate the communication load under the current scheme according to the communication load calculation formula of the on-chip network in step S2 to find the minimum communication load and the corresponding segmentation scheme;
[0053] The heuristic algorithm is implemented using the adaptive hill climbing algorithm, and the specific calculation process is as follows:
[0054] S41. Select the best neighborhood solution: Select the solution with the highest evaluation value from the neighborhood solutions as the best neighborhood solution;
[0055] S42. Update the current solution: Set the best neighborhood solution as the new current solution;
[0056] S43. Adaptive adjustment: Adjust the search strategy according to the performance of the current solution and the neighborhood solutions. The adjusted search strategy includes changing the step size and changing the neighborhood structure;
[0057] S44. Check the termination condition: Judge whether the termination condition is satisfied. The termination condition is to reach the maximum number of iterations or the quality of the solution reaches the preset threshold;
[0058] S45. Output the optimal solution: If the termination condition is satisfied, output the current optimal solution and end the process; If the termination condition is not satisfied, return to step S1 to continue the search.
[0059] As mentioned above, only the preferred specific embodiments of the present invention are described, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A data mapping method from a pulse convolutional neural network to a brain-like computing chip, characterized in that: The following steps are involved: S1. Computation graph segmentation: The data of the computation graph is segmented in parallel to obtain various subgraphs; S2. Inter-core communication modeling: The inter-core communication is modeled as a three-dimensional box packing model, which includes mapping neuron cores and statistical communication loads; The specific process of step S2 is: S21, mapping neuron cores: placing the neuron clusters in the segmented subgraphs onto adjacent neuron cores in topological order, so as to make the segmented subgraphs satisfy the neuron core capacity constraints and maximize the communication efficiency between adjacent neuron clusters; S22. Statistical communication load: According to the mapping of neuron groups to neuron cores and the static routing algorithm, the communication load of each node on the on-chip network is counted; In a Mesh network structured network on chip, the communication load is located in the two-dimensional space of the network on chip. According to the static routing algorithm, the calculation formula for the communication load of each node is: = , , where Core To the core j The communication load; is the sending pulse core; j is the receiving pulse core; Core Send to core j The number of pulses; For Node The communication load; is the horizontal coordinate of the two-dimensional grid NoC, is the ordinate of the two-dimensional grid NoC, is the node communication load coordinate, Constitute a three-dimensional space; is the sum function; is the horizontal coordinate interval of the two-dimensional grid NoC; is the ordinate interval of the two-dimensional grid NoC; S3. Calculate the cost function: Using three-dimensional space, the communication optimization problem is modeled as a three-dimensional packing problem. The optimization goal is to minimize the maximum height. The calculation formula of the cost function is: , where is the maximum communication load of all nodes; To find the maximum value function; S4, optimize the cost function: select the optimal segmentation scheme based on the cost function using an enumeration algorithm or a heuristic algorithm, and then execute step S1 to obtain the final parallel data flow graph.
2. The method for mapping data from a pulse convolutional neural network to a brain-like computing chip according to claim 1, characterized in that: The specific process of step S1 is: S11. Calculate the shape of neuron clusters: Model the neural network as a directed acyclic graph. The points in the directed acyclic graph are used to execute operators, and the edges represent the data dependencies between operators. According to the directed acyclic graph model, starting from the input and the first layer, the shapes of each neuron cluster are inferred layer by layer. Then, according to the input and parameters of each layer, the output size of the sub-operator is calculated to obtain the shape of the neuron cluster. S12, calculating input size: calculating input size is the inverse operation of calculating output size, and the input size of the sub-operator is calculated according to the output shape of the sub-operator; S13, perform segmentation: perform operator segmentation according to the output size and input size of the sub-operator, cut and distribute the input to each sub-operator, and analyze whether the output size of the previous sub-operator matches the input size of the current sub-operator; if the output size of the previous sub-operator does not match the input size of the current sub-operator, the current sub-operator requires part of the output of the previous sub-operator in the adjacent sub-graph as input, and introduces a connection between the two sub-graphs; if the output size of the previous sub-operator matches the input size of the current sub-operator, there is no need to obtain input from the adjacent sub-graph; wherein, when initializing the segmentation, the directed acyclic graph is segmented into n sub-graphs, and n is the minimum value that satisfies the neuron core capacity constraint.
3. The data mapping method from a pulse convolutional neural network to a brain-like computing chip according to claim 1, characterized in that: The specific process of step S4 is as follows: according to the calculated shape of the neuron cluster, determine the total number of operator partitioning schemes; use the total number of operator partitioning schemes as the size of the solution space for single-objective optimization, and give a threshold of the solution space size as a hyperparameter; if the total number of operator partitioning schemes is less than or equal to the threshold, the solution time is acceptable, and an enumeration algorithm is used to traverse the solution space to find the optimal solution; if the total number of operator partitioning schemes is greater than the threshold, the solution time is unacceptable, and a heuristic algorithm is used to find a feasible solution.
4. The method for mapping data from a pulse convolutional neural network to a brain-like computing chip according to claim 3, characterized in that: The threshold value ranges from 1000 to 10000.
5. The method for mapping data from a pulse convolutional neural network to a brain-like computing chip as claimed in claim 3, characterized in that: The specific calculation process of the enumeration algorithm is: traverse all partitioning schemes, and in each iteration, calculate the communication load under the current scheme according to the communication load calculation formula of the on-chip network in step S2 to find the minimum communication load and the corresponding partitioning scheme.
6. The method for mapping data from a pulse convolutional neural network to a brain-like computing chip as claimed in claim 3, characterized in that: The heuristic algorithm is implemented using an adaptive hill climbing algorithm, and the specific calculation process is: S41, selecting the best neighborhood solution: selecting the solution with the highest evaluation value from the neighborhood solutions as the best neighborhood solution; S42, updating the current solution: setting the best neighborhood solution as a new current solution; S43, adaptive adjustment: adjusting the search strategy according to the performance of the current solution and the neighborhood solution, the adjustment of the search strategy includes changing the step size and changing the neighborhood structure; S44, checking termination conditions: determining whether termination conditions are met, the termination conditions being that a maximum number of iterations is reached or the quality of the solution reaches a preset threshold; S45, output the optimal solution: if the termination condition is met, output the current optimal solution and end the process; If the termination condition is not met, return to step S1 to continue searching.
Citation Information
Patent Citations
Network-on-chip resource mapping method based on dynamic characteristics of neural network
CN105469143A
On-chip core compiling and mapping method and device of neural network based on reinforcement learning
CN114492782A