Neural network deployment method and device and electronic equipment

By generating target calculation diagrams and performing simulation verification, and generating configuration information of brain-like chip computing cores, the problem of inefficient deployment of neural networks in the existing technology is solved, and efficient and accurate deployment results are achieved.

CN120087428APending Publication Date: 2025-06-03PEKING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510002957.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and efficiently deploy neural network models into brain-like chips, resulting in inefficient deployment and error-prone.

Method used

By generating a target calculation diagram, simulation calculation is performed based on the calculation diagram and its correctness is verified. Then, configuration information of the computing core in the brain-like chip is generated, and an executable file is finally generated, including binary instructions that can be recognized by the brain-like chip.

Benefits of technology

It realizes more accurate and efficient deployment of neural network models into brain-like chips, optimizes computing resource allocation, and improves computing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087428A_ABST
    Figure CN120087428A_ABST
Patent Text Reader

Abstract

The invention provides a neural network deployment method and device and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: generating a target calculation graph based on a network structure of a to-be-deployed neural network and a target mapping relation, carrying out simulation calculation based on the target calculation graph, and carrying out simulation calculation on the basis of the target calculation graph under the condition of determining that the target calculation graph passes verification. And generating configuration information of a calculation core in the brain-like chip on the basis of a topological connection relationship between neuron nodes and synaptic connection nodes in the target calculation graph, and further generating an executable file comprising a binary instruction which can be identified by the brain-like chip on the basis of the configuration information. According to the neural network deployment method and device and the electronic equipment provided by the invention, the neural network can be deployed into the brain-like chip more accurately and more efficiently, the computing resource allocation of the brain-like chip can be optimized, the computing resource utilization rate of the brain-like chip can be improved, and the application prospect is wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a neural network deployment method, device and electronic equipment. Background Art

[0002] Neuro-inspired chips are chips that use circuits to simulate the neural network architecture of the human brain. Neuro-inspired chips combine microelectronics technology and new neuromorphic devices and are designed to imitate the computing principles of the human brain's nervous system to achieve ultra-low power consumption and parallel information processing capabilities similar to the human brain.

[0003] Compared with traditional chips, brain-like chips have greater advantages in power consumption and learning ability, and can achieve deep integration of storage and computing, greatly improving computing performance, increasing integration, and reducing energy consumption. Therefore, deploying neural network models on brain-like chips can bring many advantages such as improved energy efficiency, enhanced real-time processing capabilities, adaptive learning, hardware acceleration, and promoting the development of artificial intelligence.

[0004] However, most of the commonly used neural networks are designed and trained based on mainstream deep learning frameworks (such as TensorFlow, PyTorch, and SpikingJelly). The main target hardware platforms of the above frameworks are usually central processing units (CPUs) and graphics processing units (GPUs), rather than brain-like chips. The neural networks designed and trained based on the above mainstream deep learning frameworks cannot be directly deployed and run on brain-like chips, and additional conversion and adaptation work is required. When converting and adapting neural networks, the efficiency of converting and adapting neural networks is low and prone to errors because the structure of neural networks is usually complex and changeable. Therefore, it is difficult to accurately and efficiently deploy neural networks into brain-like chips in related technologies. How to deploy neural networks into brain-like chips more accurately and efficiently is a technical problem that needs to be solved in this field. Summary of the invention

[0005] The present invention provides a neural network deployment method, device and electronic device to solve the problem in the prior art that it is difficult to accurately and efficiently deploy a neural network model into a brain-like chip, thereby achieving more accurate and efficient deployment of the neural network model into a brain-like chip.

[0006] The present invention provides a neural network deployment method, comprising the following steps.

[0007] Generate a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, where the target mapping relationship is used to describe the mapping relationship between operators in the neural network and hardware units in the brain-inspired chip, the type of the hardware unit is a neuron unit or a synaptic connection unit, any neuron node in the target computation graph represents one of the neuron units, any synaptic connection node in the target computation graph represents one of the synaptic connection units, and a unidirectional arrow used to connect two hardware nodes in the target computation graph represents the input-output relationship between the two hardware nodes; Perform simulation calculations based on the target computation graph and obtain the simulation calculation results. When it is determined based on the simulation calculation results that the target computation graph passes the verification, generate configuration information for the computing cores in the brain-inspired chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target computation graph; Generate an executable file based on the configuration information of the computing cores, where the executable file includes binary instructions recognizable by the brain-inspired chip.

[0008] The present invention also provides a neural network deployment device, including the following modules: A computation graph generation module, configured to generate a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, where the target mapping relationship is used to describe the mapping relationship between operators in the neural network and hardware units in the brain-inspired chip, the type of the hardware unit is a neuron unit or a synaptic connection unit, any neuron node in the target computation graph represents one of the neuron units, any synaptic connection node in the target computation graph represents one of the synaptic connection units, and a unidirectional arrow used to connect two hardware nodes in the target computation graph represents the input-output relationship between the two hardware nodes; A configuration generation module, configured to perform simulation calculations based on the target computation graph and obtain the simulation calculation results. When it is determined based on the simulation calculation results that the target computation graph passes the verification, generate configuration information for the computing cores in the brain-inspired chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target computation graph; A file generation module, configured to generate an executable file based on the configuration information of the computing cores, where the executable file includes binary instructions recognizable by the brain-inspired chip.

[0009] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the neural network deployment method described in any one of the above is implemented.

[0010] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the neural network deployment method described in any one of the above is implemented.

[0011] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the neural network deployment method described in any one of the above is implemented.

[0012] The neural network deployment method, device and electronic device provided by the present invention, after generating a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, perform simulation calculation based on the target computation graph and obtain the simulation calculation result. When it is determined that the target computation graph passes the verification based on the simulation calculation result, generate the configuration information of the computing cores in the brain-inspired chip based on the topological connection relationship between the neuron nodes and synapse connection nodes in the target computation graph. Furthermore, based on the configuration information of the computing cores, generate an executable file including binary instructions recognizable by the brain-inspired chip, which can deploy the neural network to the brain-inspired chip more accurately and efficiently, can optimize the computing resource allocation of the brain-inspired chip, improve the utilization rate of the computing resources of the brain-inspired chip, and has broad application prospects. Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 is one of the schematic flowcharts of the neural network deployment method provided by the present invention.

[0015] Figure 2 is one of the connection diagrams of the target computation graph in the neural network deployment method provided by the present invention.

[0016] Figure 3 is the second connection diagram of the target computation graph in the neural network deployment method provided by the present invention.

[0017] Figure 4 is one of the combination ways of the hardware unit combinations corresponding to the complex operators in the neural network deployment method provided by the present invention.

[0018] Figure 5 is the second combination way of the hardware unit combinations corresponding to the complex operators in the neural network deployment method provided by the present invention.

[0019] Figure 6It is the third combination method of the hardware unit combinations corresponding to the complex operators in the neural network deployment method provided by the present invention.

[0020] Figure 7 It is a schematic flow chart of generating a target computation graph in the neural network deployment method provided by the present invention.

[0021] Figure 8 It is a schematic diagram of the evolution of the target computation graph in the neural network deployment method provided by the present invention.

[0022] Figure 9 It is a schematic flow chart of performing simulation calculation based on the target computation graph in the neural network deployment method provided by the present invention.

[0023] Figure 10 It is the second schematic flow chart of the neural network deployment method provided by the present invention.

[0024] Figure 11 It is a schematic structural diagram of the target routing group in the neural network deployment method provided by the present invention.

[0025] Figure 12 It is a corresponding relationship diagram between the target computation graph and the structure of the computing cores in the neuromorphic chip in the neural network deployment method provided by the present invention.

[0026] Figure 13 It is a schematic flow chart of segmenting and optimizing the target synaptic connection nodes in the neural network deployment method provided by the present invention.

[0027] Figure 14 It is a schematic flow chart of performing row-column transformation on the target synaptic connection nodes in the neural network deployment method provided by the present invention.

[0028] Figure 15 It is a schematic structural diagram of the neural network deployment device provided by the present invention.

[0029] Figure 16 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0030] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] In the description of the invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", and "linked" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0032] In the description of this application, the terms "first", "second", etc. are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually of the same kind, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, in the description of this application, "and / or" means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0033] It should be noted that as the carrier of brain-like computing, the design principle of the brain-like chip is deeply inspired by the structure of the biological brain, aiming to achieve efficient large-scale parallel computing by simulating neurons and their synaptic connections. Such a chip usually integrates multiple hardware modules, including neuron computing circuits, synaptic computing circuits, storage modules for storing weight data and neuron attribute information, etc., as well as control modules, scheduling modules, and routing modules responsible for coordinating the work of each component. These modules together form a computing core, and multiple computing cores are connected through a Network on Chip (NoC) to form a complete brain-like chip architecture. The computing cores are arranged in an array form to achieve data communication with each other through the on-chip network; the output generated by the computing cores can be directly transmitted to one or more other cores, or transmitted to the outside of the chip. In addition, multiple brain-like chips can also be cascaded in an array form and communicate efficiently between chips through a dedicated protocol, thereby constructing a larger-scale computing system.

[0034] When deploying neural networks to brain-like chips based on traditional neural network deployment methods in related technologies, there are still the following shortcomings: First, most of the commonly used neural networks are designed and trained based on mainstream deep learning frameworks (such as TensorFlow, PyTorch, and SpikingJelly). The main target hardware platforms of the above frameworks are usually central processing units (CPUs) and graphics processing units (GPUs), not brain-like chips. The neural networks designed and trained based on the above mainstream deep learning frameworks cannot be directly deployed and run on brain-like chips, and additional conversion and adaptation work is required.

[0035] Secondly, when compiling brain-like chips, it is necessary to deeply understand the hardware architecture of brain-like chips and the computing requirements of the application model in order to form the correct mapping and calculation. However, since the structure of neural networks is usually complex and changeable, there is a huge design space for mapping neural networks to brain-like chips. As a result, the traditional neural network deployment method based on related technologies needs to be configured in detail for each neuron and synapse when applied to the deployment of brain-like chips. This is not only inefficient but also prone to configuration errors, which in turn affects the reliability and performance of the deployed neural network.

[0036] Finally, the traditional neural network deployment methods in related technologies lack optimization and adaptation for the hardware characteristics of brain-like chips. Brain-like chips have unique hardware architectures, such as large-scale parallel computing units, on-chip network communications, and low-power design. If these hardware characteristics are not fully utilized when the neural network model is deployed on the brain-like chip, the hardware resources of the brain-like chip will not be fully utilized, and its unique architectural advantages will not be fully utilized.

[0037] Therefore, it is difficult to automatically, accurately and efficiently convert neural networks into executable instructions for brain-like chips based on traditional neural network deployment methods in related technologies. Especially when dealing with large-scale and complex neural networks, the optimization and scheduling problems in the process of deploying neural networks in brain-like chips are particularly prominent. Therefore, how to deploy neural network models in brain-like chips more accurately and efficiently is a technical problem that needs to be solved in this field.

[0038] Combine the following Figures 1-14 The neural network deployment method provided by the present invention is described.

[0039] Figure 1 This is one of the flow charts of the neural network deployment method provided by the present invention. Figure 1As shown, the method includes the following: Step 101, based on the network structure of the neural network to be deployed and the target mapping relationship, generate a target computation graph corresponding to the neural network to be deployed. The target mapping relationship is used to describe the mapping relationship between the operators in the neural network and the hardware units in the brain-inspired chip. The type of the hardware unit is a neuron unit or a synaptic connection unit. Any neuron node in the target computation graph represents a neuron unit, and any synaptic connection node in the target computation graph represents a synaptic connection unit. The unidirectional arrow used to connect two hardware nodes in the target computation graph represents the input-output relationship between the two hardware nodes.

[0040] It should be noted that the execution subject of the embodiments of the present invention is a neural network deployment device. The above neural network deployment device can be configured in an electronic device such as a computer or a server.

[0041] Specifically, the neural network deployment method provided by the present invention is used to convert the neural network to be deployed into binary instructions executable by the brain-inspired chip, so as to realize the deployment of the neural network to be deployed in the brain-inspired chip.

[0042] It should be noted that the neural network to be deployed in the embodiments of the present invention can be designed and trained based on mainstream deep learning frameworks (such as TensorFlow, PyTorch, and SpikingJelly, etc.).

[0043] It can be understood that the neural network to be deployed in the embodiments of the present invention can be determined based on actual needs. There is no specific limitation on the neural network to be deployed in the embodiments of the present invention.

[0044] It can be understood that in a neural network, an operator usually refers to a function or operator that performs specific operations on data. The above specific operations can include but are not limited to mathematical operations, transformations, and activation functions, etc., which are used to implement various functions in the neural network. The types of operators in the neural network can include but are not limited to convolution operators (ConvolutionOperator), pooling operators (Pooling Operator), fully connected operators (Fully Connected Operator), activation operators (Activation Function), Dropout operators, Softmax operators, loss function operators, arithmetic operators (including addition, subtraction, multiplication, division, etc., used to perform basic mathematical operations), and linear algebra operators (such as matrix multiplication, transpose, inversion, etc., used to process vectors and matrices), etc. By combining these operators, an efficient and accurate neural network can be constructed.

[0045] It should be noted that the types of operators in the neural network in the embodiments of the present invention may further include input node operators and output node operators. The input node operator represents the data input port of the neural network, and the output node operator represents the data output port of the neural network. The input node operator is responsible for receiving external data and passing the data to other operators, and the output node operator is used to send the received data to external devices.

[0046] It should be noted that the types of operators in the neural network in the embodiments of the present invention may further include complex operators. The complex operator is an encapsulation of operators with specific logical functions in the neural network.

[0047] It should be noted that the neuron unit and the synaptic connection unit are important hardware units in the brain-inspired chip. In the brain-inspired chip, the combination of the neuron unit and the synaptic connection unit constitutes the basic architecture of the brain-inspired chip.

[0048] Among them, the neuron unit is usually implemented as a circuit with multiple inputs. The input signal acts on the membrane level accumulation process of the neuron, thereby affecting its output state. The output of the neuron unit is usually an analog signal or a digital signal, depending on the design and implementation method of the chip.

[0049] The types of neuron units in the embodiments of the present invention may include, but are not limited to, basic (Neuron) neuron units, IF (integrate-and-fire) neuron units, LIF (leaky integrate-and-fire) neuron units, continuous spiking neuron units, periodic spiking neuron units, etc. The encapsulation of the neuron unit includes not only input-output characteristics but also the update logic of the internal state. Each neuron unit will ultimately call the update method of the parent class MetaNeuron.

[0050] The synaptic connection unit is used to connect neuron units and transmit signals in the brain-inspired chip. The synaptic connection unit simulates the synaptic connection between biological neurons, meaning that signals can only be transmitted from one neuron unit to another neuron unit via this synaptic connection unit and cannot be transmitted in the reverse direction. The synaptic connection unit is usually implemented as a circuit with programmable weights. These weights can be adjusted through training to change the transmission strength of the synaptic connection unit for input signals.

[0051] The types of synaptic connection units in the embodiments of the present invention may include basic synaptic connection units, fully-connected synaptic connection units, convolutional synaptic connection units, 2D matrix multiplication synaptic connection units, etc. Different types of synaptic connection units are applicable to different network connection methods and neural network models. For example, fully-connected synaptic connection units correspond to full connections in neural networks, and convolutional synaptic connection units correspond to convolutional connections in neural networks.

[0052] In the embodiments of the present invention, based on the hardware architecture of the brain-inspired chip, as well as prior knowledge and / or actual situations, the mapping relationship between the operators in the neural network and the hardware units in the brain-inspired chip is obtained as the target mapping relationship.

[0053] The network structure of the neural network to be deployed in the embodiments of the present invention may include the operators in the neural network to be deployed and the connection relationships between the operators. Based on the network structure of the neural network to be deployed and the target mapping relationship, the target computational graph corresponding to the neural network to be deployed can be generated.

[0054] Figure 2 is one of the connection schematic diagrams of the target computational graph in the neural network deployment method provided by the present invention. Figure 3 is the second connection schematic diagram of the target computational graph in the neural network deployment method provided by the present invention. As Figure 2 and Figure 3 shown, the hardware nodes in the target computational graph in the embodiments of the present invention may include neuron nodes and synaptic connection nodes. Any neuron node in the target computational graph represents a neuron unit in the brain-inspired chip, and any synaptic connection node in the target computational graph represents a synaptic connection unit in the brain-inspired chip.

[0055] The input side of any synaptic connection node can only be connected to one neuron node, and the output side of the above synaptic connection node can also only be connected to one neuron node, indicating that the forward direction of any synaptic connection unit can only be connected to one neuron unit, and the backward direction of the above synaptic connection unit can also only be connected to one neuron unit.

[0056] The input side and the output side of any neuron node can be respectively connected to one or more synaptic connection nodes, indicating that the neuron unit represented by the above neuron node can have one or more inputs and one or more outputs.

[0057] It should be noted that the input node in the target computational graph represents the input node operator in the neural network to be deployed, and the output node in the target computational graph represents the output node operator in the neural network to be deployed.

[0058] It should be noted that in the target mapping relationship, complex operators in the neural network can be equivalently replaced by a combination of hardware units formed by connecting one or more synaptic connection units and one or more neuron units according to the conversion method. When the hardware unit combination corresponding to the complex operator is regarded as a whole, the above-mentioned hardware unit combination can receive the outputs of one or more neuron units and output them to one or more synaptic connection units. It can be understood that the input-side behavior of the hardware unit combination corresponding to the complex operator is similar to that of a synaptic connection unit because it receives the outputs from neuron units; the output-side behavior of the hardware unit combination corresponding to the complex operator is similar to that of a neuron unit because it transmits the output to synaptic connection units. The combination method of the hardware unit combination corresponding to the complex operator can be any combination method that satisfies the connection rules of neuron units and synaptic connection units.

[0059] It should be noted that Figure 2 and Figure 3 the numbers in represent the identification information of hardware nodes in the target computation graph, Figure 2 and Figure 3 the identification information of the hardware nodes shown in is only illustrative.

[0060] Figure 4 is one of the combination methods of the hardware unit combination corresponding to the complex operator in the neural network deployment method provided by the present invention. As shown in Figure 4 , the hardware unit combination corresponding to the complex operator can be a combination of one synaptic connection unit and one neuron unit.

[0061] Figure 5 is the second combination method of the hardware unit combination corresponding to the complex operator in the neural network deployment method provided by the present invention. As shown in Figure 5 , the hardware unit combination corresponding to the complex operator can be a combination of two synaptic connection units and one neuron unit.

[0062] Figure 6 is the third combination method of the hardware unit combination corresponding to the complex operator in the neural network deployment method provided by the present invention. As shown in Figure 6 , the hardware unit combination corresponding to the complex operator can be a combination of multiple synaptic connection units and multiple neuron units.

[0063] For example, bitwise operation operators, as a type of complex operator, include Bitwise AND operator, Bitwise OR operator, and Bitwise XOR operator. Arithmetic operators, as a type of complex operator, include SpikingAdd operator, SpikingSub operator, SpikingAvgPool1d operator, SpikingMaxPool1d operator, etc. These complex operators as a whole implement specific functional logics. The internal build method predefines transformation methods for equivalently replacing themselves with a combination of specific hardware units while maintaining the same external connection relationships as the original operators.

[0064] It should be noted that Figures 4-6 the numbers in Figures 4-6 represent the identification information of the hardware units, and the identification information of the hardware units shown in

[0065] Figure 7 is a schematic flow diagram of generating a target computation graph in the neural network deployment method provided by the present invention. As shown in Figure 7 , as an optional embodiment, based on the network structure of the neural network to be deployed and the target mapping relationship, generating a target computation graph corresponding to the neural network to be deployed includes: generating an original computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed. Any operator node in the original computation graph represents an operator in the neural network to be deployed, and the unidirectional arrow used to connect two operator nodes in the original computation graph represents the input-output relationship between the two operator nodes.

[0066] It should be noted that a computation graph is a data structure used to represent the composition structure and composition relationship of a neural network. A computation graph is a directed graph. The nodes in the computation graph represent the operators in the neural network, and the edges represent the connection relationships between the operators. Except for the input nodes and output nodes, the in-degree and out-degree of each node are both greater than 0. The edges represent the execution order and data flow direction of the operators through unidirectional arrows.

[0067] By parsing the network structure of the neural network to be deployed, an original computation graph corresponding to the neural network to be deployed can be generated.

[0068] For a type of operator nodes in the original computation graph, based on the first mapping sub-relation in the target mapping relation, replace the type of operator nodes with hardware unit nodes. For a second type of operator nodes in the original computation graph, based on the second mapping sub-relation in the target mapping relation, replace the second type of operator nodes with complex operator nodes, to obtain an intermediate computation graph corresponding to the neural network to be deployed. The first mapping sub-relation is used to describe the mapping relation between a type of operators in the neural network and the hardware units in the brain-inspired chip. The second mapping sub-relation is used to describe the mapping relation between the second type of operators in the neural network and the complex operators. The first type of operators and the second type of operators in the neural network are predefined. Any type of operator node in the original computation graph represents a type of operator. Any second type of operator node in the original computation graph represents a second type of operator. Any complex operator node in the intermediate computation graph represents a complex operator.

[0069] It should be noted that the first mapping sub-relation in the embodiments of the present invention can be determined based on prior knowledge and / or actual situations. The first mapping sub-relation is shown in Table 1.

[0070] Table 1 Schematic Table of the First Mapping Sub-relation

[0071] It should be noted that some operations are defined through objects in Pytorch or SpikingJelly, and this operator can be called and perform operations; while others are defined using mathematical operators (+, -, ×, / ). The operators in the embodiments of the present invention also support calculations in some operator forms, that is, the brain-inspired chip can perform these forms of calculations.

[0072] For a type of operator nodes in the original computation graph, based on the first mapping sub-relation in the target mapping relation, replace the type of operator nodes one-to-one with hardware units. During the replacement process, all attributes and configuration information of the type of operator nodes will be accurately transferred to the hardware unit nodes, ensuring that the converted hardware unit nodes are completely consistent with the original type of operator nodes in function.

[0073] The second mapping sub-relation in the embodiments of the present invention can be determined based on prior knowledge and / or actual situations. The second mapping sub-relation is shown in Table 2.

[0074] Table 2 Schematic Table of the Second Mapping Sub-relation

[0075] For the second - type operator nodes in the original computational graph, there are no hardware unit nodes that can be directly and equivalently replaced for the above - mentioned second - type operators. Therefore, during the construction stage of the neural network to be deployed, the second - type operators are regarded as an indivisible whole. Before the neural network to be deployed is deployed to the brain - like chip, the second - type operators are approximately replaced with complex operators to ensure that the neural network to be deployed can run smoothly on the brain - like chip after being deployed.

[0076] The conversion between the second - type operator nodes and the complex operator nodes strictly follows the principle of equivalent replacement, that is, ensuring that the output results under the same input conditions are consistent with the original neural network to be deployed. The specific conversion rules are shown in Table 2. During the conversion process, the structural information and weight data of the second - type operators will be appropriately adjusted and mapped to be applicable to the complex operators and the corresponding hardware unit groups, while maintaining the consistency of the data - flow structure and the accuracy of the operation results.

[0077] It should be noted that when replacing the first - type operator nodes in the original computational graph with hardware unit nodes and replacing the second - type operator nodes with complex operator nodes, the external connection relationships of the first - type operator nodes and the second - type operator nodes are retained.

[0078] Based on the third mapping sub - relationship in the target mapping relationship, the complex operator nodes in the intermediate computational graph are replaced with hardware units to obtain the target computational graph. The third mapping sub - relationship is used to describe the mapping relationship between the complex operators in the neural network and the hardware units in the brain - like chip.

[0079] It should be noted that the third mapping sub - relationship in the embodiments of the present invention can be determined based on prior knowledge and / or actual situations. The third mapping sub - relationship is shown in Table 3.

[0080] Table 3 Schematic Table of the Third Mapping Sub - relationship

[0081] Figure 8 is the schematic diagram of the evolution of the target computational graph in the neural network deployment method provided by the present invention. As Figure 8 shown, the original computational graph corresponding to the neural network to be deployed evolves into the intermediate computational graph, and then the evolution process into the target computational graph is as Figure 8 shown. Figure 8 The dashed arrows in Figure 8 represent the replacement correspondence relationships between the first - type operators and the hardware units, as well as between the second - type operators and the complex operators and the hardware units. The dashed box in

[0082] It can be understood that Figure 8 In Figure 8 , the IFNode node 1, IFNode node 2, and IFNode node 3 in the original computational graph all represent an IFNode operator, the addition node 1 represents an addition operator, and the 2D convolution node represents a convolution node. As shown in Table 1, the IFNode node 1, IFNode node 2, IFNode node 3, and the 2D convolution node are all operator nodes of one type. As shown in Table 2, the addition node 1 is an operator node of the second type.

[0083] It should be noted that when replacing the complex operator nodes in the intermediate computational graph with hardware unit nodes, the external connection relationships of the complex operator nodes are retained.

[0084] Through the abstract design of the hardware unit in the embodiments of the present invention, modular encapsulation of operators in the neural network is realized, and each operator in the neural network is an independent module with clear functions and interfaces, which is convenient for management and reuse. On this basis, the embodiments of the present invention formulate connection rules between hardware units, realizing the conversion between hardware units and operators in the neural network, not only improving the flexibility of neural network deployment, but also simplifying the construction and optimization of the neural network.

[0085] Step 102: Perform simulation calculations based on the target computational graph and obtain the simulation calculation results. When it is determined that the target computational graph passes the verification based on the simulation calculation results, generate the configuration information of the computing cores in the brain-like chip based on the topological connection relationships of the neuron nodes and synaptic connection nodes in the target computational graph.

[0086] It should be noted that since when generating the target computational graph, the second-type operators in the neural network to be deployed can only be approximately converted into the hardware unit group corresponding to the complex operator, which is mathematically different from the second-type operators. In order to ensure that the target computational graph can match the hardware structure and functions of the brain-like chip and avoid the situation of deployment failure due to the hardware function limitations of the brain-like chip (for example, the brain-like chip cannot process topological structures such as loops), in the embodiments of the present invention, before performing the next deployment work, by performing simulation calculations based on the target computational graph, the numerical correctness of the target computational graph running on the brain-like chip can be verified, and then it can be verified whether the neural network to be deployed can work properly after being deployed to the brain-like chip.

[0087] Therefore, in the embodiments of the present invention, after obtaining the target computational graph, perform simulation calculations based on the target computational graph, and when it is determined that the numerical values of the target computational graph running on the brain-like chip are correct based on the simulation calculation results, determine that the target computational graph passes the verification.

[0088] As an optional embodiment, performing simulation calculations based on the target computational graph and obtaining simulation calculation results includes: adding probes to the hardware unit nodes to be verified and the properties to be verified in the target computational graph to obtain the computational graph to be simulated.

[0089] Input the computational graph to be simulated into the simulator for the simulator to perform simulation calculations based on the computational graph to be simulated and trigger the recording of the global number of simulation times.

[0090] In the case where the global number of simulation times has not reached the predefined number threshold, call the update method of the simulator to sequentially update the states of the input nodes, synaptic connection nodes, and neuron nodes in the computational graph to be simulated, and save the attribute information of the hardware unit nodes to be verified in the target computational graph as the simulation calculation result of this simulation calculation. If the global number of simulation times reaches the predefined number threshold, it is determined that the simulation ends. The input nodes in the target computational graph are used to input data.

[0091] Figure 9 It is a schematic flowchart of performing simulation calculations based on the target computational graph in the neural network deployment method provided by the present invention. As Figure 9 shown, in step 901, add probes (Probe) to the hardware unit nodes to be verified and the properties to be verified in the target computational graph to obtain the computational graph to be simulated.

[0092] In step 902, input the neural network to be simulated into the simulator (Simulator) as the simulation target of the simulator.

[0093] In step 903, start performing simulation calculations. The simulator internally records the global number of simulation times, with an initial value of 0.

[0094] In step 904, determine whether the current simulation time has reached the predefined number threshold. If so, end the simulation. If not, execute step 905.

[0095] In step 905, the simulator calls its update method to update the states of all input nodes in the computational graph to be simulated.

[0096] In step 906, the simulator calls its update method to update the states of all synaptic connection nodes in the computational graph to be simulated.

[0097] In step 907, the simulator calls its update method to update the states of all neuron nodes in the computational graph to be simulated.

[0098] In step 908, the simulator obtains the simulation end result of this simulation calculation according to the operator objects and properties to be monitored specified by all the probes in the computational graph to be simulated, and copies and saves the simulation end result of this simulation calculation inside the simulator.

[0099] Step 909: increment the global simulation count recorded inside the emulator by 1, and then return to execute Step 904.

[0100] It should be noted that between Step 902 and Step 903, external probes can also be added to the hardware unit nodes to be verified and the attributes to be verified in the computation graph to be simulated.

[0101] After the emulator finishes the simulation computation, based on the simulation computation results of each simulation computation, it can be verified whether the numerical values are correct when the target computation graph runs on the brain-like chip. And when it is determined that the numerical values are correct when the target computation graph runs on the brain-like chip, it is determined that the target computation graph passes the verification.

[0102] The simulation computation method in the embodiments of the present invention follows the same computational principle as the brain-like chip. For example, there is an input buffer between neurons and synapses, and the emulator will also follow the same behavior as the brain-like chip for the read and write behaviors of data in the input buffer; for another example, each computing core in the brain-like chip can set the start time and end time of working, and this hardware feature is reflected in two parameters of the neuron, tick_wait_start (abbreviated as tws) and tick_wait_end (abbreviated as twe), which define the start time and end time of the neuron's work. In the simulation computation, the emulator will compare the global simulation count with these two parameters to determine whether the neuron works in each simulation computation.

[0103] Based on the simulation computation method in the embodiments of the present invention, before deploying the target computation graph to the brain-like chip, it is possible to verify whether the operation results of the target computation graph are correct through software, and based on the verification results, determine whether subsequent deployment processes need to be carried out and start hardware-related debugging work such as starting the brain-like chip hardware, saving manpower, improving work efficiency, and shortening the development cycle.

[0104] It should be noted that the routing structure of the brain-like chip supports the multicast function, that is, the output of the neuron unit in a certain computing core can be input to the neuron units in multiple computing cores through the on-chip network. The neuron unit as the destination needs to meet certain routing rules, such as being continuously arranged in the brain-like chip, etc. Therefore, in the embodiments of the present invention, the target computation graph is parsed, and the neuron units with the same output destination are put into the same routing group and allocated to the computing cores on the brain-like chip as a whole to meet the multicast rules. For a specific network structure, nesting may occur between multiple routing groups, and at this time, each routing group needs to meet the routing rules.

[0105] When it is determined that the target computational graph passes the verification based on the simulation calculation results, the configuration information of the computing cores in the brain-inspired chip can be generated by means of numerical calculation, mathematical statistics, etc. based on the topological connection relationship between the neuron nodes and synapse connection nodes in the target computational graph.

[0106] Figure 10 It is the second flow schematic diagram of the neural network deployment method provided by the present invention. As Figure 10 shown, as an optional embodiment, generating the configuration information of the computing cores in the brain-inspired chip based on the topological connection relationship between the neuron nodes and synapse connection nodes in the target computational graph includes: splitting and optimizing the weight matrix of the target synapse connection nodes in the target computational graph to obtain the optimized weight matrix of the target synapse connection nodes.

[0107] Assign routing coordinates for each of the target routing groups to indicate the positions of each target routing group in the computing cores of the brain-inspired chip, and then generate the configuration information of the computing cores based on the neuron units included in each target routing group, the nested relationship between the target routing groups, the routing coordinates of each target routing group, and the optimized weight matrix of the target synapse connection nodes.

[0108] Generate an executable file based on the configuration information of the computing cores, and the executable file includes binary instructions recognizable by the brain-inspired chip.

[0109] As an optional embodiment, grouping the neuron unit nodes in the target computational graph to obtain multiple target routing groups includes: for each neuron node in the target computational graph, determining each successor neuron node of each neuron node as an original routing group, and determining each neuron node as the input neuron node of the original routing group. The successor neuron node of a neuron node is a neuron node located on the output side of the neuron node and connected through a synapse connection node. Any synapse connection node in the target computational graph represents a synapse connection unit in the brain-inspired chip.

[0110] Traverse each original routing group. When there are the same neuron nodes in any two original routing groups, merge and deduplicate the two original routing groups to obtain an intermediate routing group, and determine the input neuron nodes that are not in the intermediate routing group among the input neuron nodes of the two original routing groups as the input neuron nodes of the intermediate routing group.

[0111] Traverse each intermediate routing group. When any intermediate routing group includes any neuron node and the successor neuron node of any neuron node, determine the successor neuron node of any neuron node as a sub-routing group in any intermediate routing group, and determine any neuron node as the input neuron node of the sub-routing group in any intermediate routing group.

[0112] Determine each original routing group, each intermediate routing group, and each sub-routing group in each intermediate routing group as each target routing group.

[0113] Figure 11 It is a schematic diagram of the structure of the target routing group in the neural network deployment method provided by the present invention. As Figure 11 shown, the successor neuron node of neuron node A is neuron node B, the successor neuron nodes of neuron node B are neuron node C and neuron node D respectively, and the successor neuron node of neuron node C is neuron node D. Therefore, neuron node B can be determined as the original routing group 3, and the input neuron node of the original routing group 3 can be determined as neuron node A, neuron nodes C and D can be determined as the original routing group 1, and the input neuron node of the original routing group 1 can be determined as neuron node B, neuron node D can be determined as the original routing group 2, and the input node of the original routing group 2 can be determined as neuron node C.

[0114] Since both the original routing group 1 and the original routing group 2 include neuron node D, a situation of routing group nesting occurs. Therefore, the original routing group 1 and the original routing group 2 can be merged and de-duplicated to obtain an intermediate routing group including neuron nodes C and D, and neuron node B can be determined as the input neuron node of the above intermediate routing group.

[0115] Since the above intermediate routing group includes neuron node C and the successor neuron node D of neuron node C, therefore, neuron node D can be determined as a sub-routing group in the above intermediate routing group, and neuron node C can be determined as the input neuron node of the sub-routing group in the above intermediate routing group.

[0116] Furthermore, the original routing group 3 including neuron node B can be determined as the target routing group 3, the intermediate routing group including neuron nodes C and D can be determined as the target routing group 1, and the one including neuron node D (the sub-routing group in the above intermediate routing group 1) can be determined as the target routing group 2.

[0117] It should be noted that Figure 11 the unidirectional arrow in represents the synaptic connection node between two neural nodes. Since the synaptic connection node is not involved, it is omitted in the figure.

[0118] In the embodiment of the present invention, by reasonably arranging and aligning the sub-packets in the original routing group, it is ensured that each target routing group meets the routing rules.

[0119] As an optional embodiment, after obtaining a plurality of target routing groups, the method further includes: obtaining the optimal weight bit-width of the weight matrix of each target routing group.

[0120] Obtaining the optimal weight bit-width of the weight matrix of each target routing group includes: obtaining the maximum value and the minimum value of the weight matrix of each target routing group.

[0121] When the maximum value of the weight matrix of each target routing group is in the first numerical interval and the minimum value of the weight matrix of each target routing group is in the second numerical interval, it is determined that the optimal weight bit-width of the weight matrix of each target routing group is the first preset value.

[0122] When the maximum value of the weight matrix of each target routing group is in the third numerical interval and the minimum value of the weight matrix of each target routing group is in the fourth numerical interval, it is determined that the optimal weight bit-width of the weight matrix of each target routing group is the second preset value.

[0123] When the maximum value of the weight matrix of each target routing group is in the fifth numerical interval and the minimum value of the weight matrix of each target routing group is in the sixth numerical interval, it is determined that the optimal weight bit-width of the weight matrix of each target routing group is the third preset value.

[0124] When the maximum value of the weight matrix of each target routing group is in the seventh numerical interval and the minimum value of the weight matrix of each target routing group is in the eighth numerical interval, it is determined that the optimal weight bit-width of the weight matrix of each target routing group is the fourth preset value.

[0125] Figure 12 It is a correspondence diagram between the target computation graph and the structure of the computing cores in the neuromorphic chip provided by the present invention. As Figure 12 shown, the neuron nodes in the target computation graph represent the neuron units in the neuromorphic chip, and the neuron unit is the neuron computing circuit of the computing core in the neuromorphic chip. The synaptic connection nodes in the target computation graph represent the synaptic connection units in the neuromorphic chip, and the attribute weight matrix of the synaptic connection unit is stored in the crossbar array of the computing core in the neuromorphic chip.

[0126] In the embodiment of the present invention, by utilizing the hardware characteristic that the neuromorphic chip can support multiple weight bit-widths, the optimal weight bit-width of each target routing group is obtained by obtaining the maximum value and the minimum value of the weight matrix of each target routing group, so that under the condition of the same size of crossbar array hardware resources, a larger weight matrix can be stored, saving the hardware resources of the computing core in the neuromorphic chip.

[0127] It should be noted that the first numerical range may include a numerical range not greater than 1; the second numerical range includes a numerical range greater than 0; the value of the first preset value may be 1 bit.

[0128] The third numerical range may include a numerical range not greater than 1; the fourth numerical range may include a numerical range less than 0 and not less than -2; the value of the second preset value may be 2 bits.

[0129] The fifth numerical range may include a numerical range greater than 1 and not greater than 7; the sixth numerical range may include a numerical range less than -2 and not less than -8; the value of the third preset value may be 4 bits.

[0130] The seventh numerical range may include a numerical range greater than 7 and not greater than 127, and the eighth numerical range may include a numerical range less than -8 and not less than -128; the value of the fourth preset value may be 8 bits.

[0131] In the embodiments of the present invention, is used to identify the target routing group,

[0132] For the th target routing group, the optimal weight bit width of the weight matrix of the th target routing group is preset to be 8 bits.

[0133] The maximum value of the weight matrix of the th target routing group and the minimum value are obtained through numerical calculation.

[0134] In the case of and , it is determined that the optimal weight bit width of the weight matrix of the th target routing group is 1 bit.

[0135] In the case of and , it is determined that the optimal weight bit width of the weight matrix of the th target routing group is 2 bits.

[0136] In the case of and , it is determined that the weight bit width of the weight matrix of the th target routing group is 4 bits.

[0137] In the case of and , it is determined that the optimal weight bit width of the weight matrix of the th target routing group is 8 bits.

[0138] After obtaining the optimal weight bit-width of the weight matrix of the th target routing group, through numerical calculation, based on the optimal weight bit-width of the weight matrix of the th target routing group, the weight matrix of the th target routing group can be segmented and optimized to obtain the optimized weight matrix of the th target routing group.

[0139] Specifically, in the embodiments of the present invention, synaptic connection nodes that meet preset conditions in the target computational graph can be determined as target synaptic connection nodes. Furthermore, through numerical calculation, the weight matrix of the target synaptic connection nodes in the target computational graph can be segmented and optimized to obtain the optimized weight matrix of the target synaptic connection nodes.

[0140] As an optional embodiment, segmenting and optimizing the weight matrix of the target synaptic connection nodes in the target computational graph to obtain the optimized weight matrix of the target synaptic connection nodes includes: for any neuron node in the target computational graph, when there are several forward neuron nodes for any neuron node, each synaptic connection node between any neuron node and each forward neuron node of any neuron node is determined as each original synaptic connection node corresponding to any neuron node, and the forward neuron node of any neuron node is a neuron node connected to the input side of any neuron node through a synaptic connection node.

[0141] When the weight matrix of any original synaptic connection node corresponding to any neuron node is a block diagonal matrix, any original synaptic connection node corresponding to any neuron node is determined as the first target synaptic connection node among the target synaptic connection nodes corresponding to any neuron node, and other original synaptic connection nodes except any original synaptic connection node corresponding to any neuron node are determined as the second target synaptic connection nodes among the target synaptic connection nodes corresponding to any neuron node. The elements in the non-zero sub-matrix located on the main diagonal of the block diagonal matrix are not all 0 values, and all other elements except the non-zero sub-matrix located on the main diagonal of the block diagonal matrix are 0 values.

[0142] Each sub-matrix on the main diagonal of the weight matrix of the first target synaptic connection node corresponding to any neuron node is respectively determined as the optimized weight matrix of the first target synaptic connection node corresponding to any neuron node.

[0143] Based on the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any neuron node, the weight matrix of the second target synaptic connection node corresponding to any neuron node is divided into multiple sub-matrices in the column dimension as the optimized weight matrix of the second target synaptic connection node corresponding to any neuron node, and the number of the optimized weight matrices of the second target synaptic connection node corresponding to any neuron node is the same as the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any neuron node.

[0144] Based on the column dimension, determine the corresponding relationship between the columns of the optimized weight matrix of the first target synaptic connection node corresponding to any neuron node and the optimized weight matrix of the second target synaptic connection node.

[0145] Based on the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any neuron node, divide the output of the forward neuron node connected to the input side of any neuron node through the first target synaptic connection node into multiple sub-outputs, and determine the one-to-one correspondence between each sub-output and the optimized weight matrix of the target synaptic connection node corresponding to any neuron node.

[0146] Based on the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any neuron node, create multiple sub-synaptic connection nodes between any neuron node and each neuron node of any neuron node in the target computational graph, and determine the one-to-one correspondence between each sub-synaptic connection node and each sub-output. The number of sub-synaptic connection nodes between any neuron node and each neuron node of any neuron node is the same as the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any neuron node. Each sub-synaptic connection node between any neuron node and each neuron node of any neuron node represents the computational task between each sub-output and the optimized weight matrix of the target synaptic connection node corresponding to any neuron node.

[0147] It should be noted that the synaptic connection unit in the brain-inspired chip is used to connect neuron units. The output of a neuron unit is a vector. If the output of neuron unit 1 has a length of , and the output of neuron unit 2 has a length of , and synaptic connection unit 1 is used to connect neuron 1 and neuron 2, then synaptic connection unit 1 has a weight matrix with a size of . For any neuron unit with several forward neuron units, its output is , where is the output of the forward neuron unit, is the weight matrix of the synaptic connection unit connecting the forward neuron unit and this neuron unit, represents the activation function of this neuron unit, is used to identify the forward neuron unit of this neuron unit.

[0148] Figure 13 is a schematic flow chart for segmenting and optimizing the target synaptic connection node in the neural network deployment method provided by the present invention. As shown in Figure 13 Figure (1) therein, if the input sides of neuron nodes A and B are respectively connected to neuron node C through synaptic connection nodes in the target computation graph, then both neuron nodes A and B are the forward neuron nodes of neuron node C.

[0149] Correspondingly, the synaptic connection node 1 between neuron node A and neuron node C and the synaptic connection node 2 between neuron node B and neuron node C can be respectively determined as the original synaptic connection nodes corresponding to neuron node C.

[0150] It can be understood that after the operation results obtained by the output of neuron node A through the operation corresponding to synaptic connection node 1 and the output of neuron node B through the operation corresponding to synaptic connection node 2 are respectively input into neuron node C, neuron node C can calculate the sum of the above operation results and substitute it into the activation function of neuron node C for calculation. Figure 13 The operation corresponding to synaptic connection node 1 in can be a 2D matrix multiplication operation, and the operation corresponding to synaptic connection node 2 can be a fully connected operation.

[0151] Respectively obtain the weight matrices of synaptic connection node 1 and synaptic connection node 2. As shown in Figure 13 Figure (2) therein, the weight matrix of synaptic connection node 1 can be represented as multiple sub-matrices, and only the sub-matrix on the main diagonal is a non-zero matrix, and the remaining sub-matrices are all zero matrices. Therefore, it can be determined that the weight matrix of synaptic connection node 1 is a block diagonal matrix, indicating that there is a waste of computing resources in the weight matrix of synaptic connection node 1. Synaptic connection node 1 can be determined as the first target synaptic connection node corresponding to neuron node C.

[0152] It should be noted that the weight matrix of synaptic connection node 2 is not a block diagonal matrix. Therefore, synaptic connection node 2 can be determined as the second target synaptic connection node corresponding to neuron node C.

[0153] As shown in Figure 13As shown in Figure (2) therein, when the weight matrix of synaptic connection node 1 can be divided into three block diagonal matrices, each non-zero sub-matrix on the main diagonal of the weight matrix of synaptic connection node 1: non-zero sub-matrix S1, non-zero sub-matrix S2, and non-zero sub-matrix S3, can be respectively determined as the optimized weight matrix of synaptic connection node 1. The weight matrix of synaptic connection node 2 can be divided into 3 sub-matrices in the column dimension, which are sub-matrix S4-1, sub-matrix S4-2, and sub-matrix S4-3 respectively.

[0154] As Figure 13 As shown in Figure (2) therein, S1 of the weight matrix of synaptic connection node 1 and sub-matrix S4-1 in the weight matrix of synaptic connection node 2 are in the same column dimension, and the corresponding relationship between non-zero sub-matrix S1 of the weight matrix of synaptic connection node 1 and sub-matrix S4-1 in the weight matrix of synaptic connection node 2 can be determined; similarly, non-zero sub-matrix S2 of the weight matrix of synaptic connection node 1 and sub-matrix S4-2 in the weight matrix of synaptic connection node 2 are in the same column dimension, and the corresponding relationship between S2 of the weight matrix of synaptic connection node 1 and sub-matrix S4-2 in the weight matrix of synaptic connection node 2 can be determined. Non-zero sub-matrix S3 of the weight matrix of synaptic connection node 1 and sub-matrix S4-3 in the weight matrix of synaptic connection node 2 are in the same column dimension, and the corresponding relationship between S3 of the weight matrix of synaptic connection node 1 and sub-matrix S4-3 in the weight matrix of synaptic connection node 2 can be determined.

[0155] Since neuron node A is connected to neuron node C through the first target synaptic connection node (synaptic connection node 1), when the weight matrix of synaptic connection node 1 is composed of three non-zero sub-matrices located on the main diagonal, and all other elements in the weight matrix of synaptic connection node 1 except the above three non-zero sub-matrices are 0 values, the output of neuron node A can be divided into three sub-outputs, which are sub-output A-1, sub-output A-2, and sub-output A-3 respectively. And the corresponding relationship between sub-output A-1 and non-zero sub-matrix S1 of the weight matrix of synaptic connection node 1 and sub-matrix S4-1 in the weight matrix of synaptic connection node 2 can be determined, the corresponding relationship between sub-output A-2 and non-zero sub-matrix S2 of the weight matrix of synaptic connection node 1 and sub-matrix S4-2 in the weight matrix of synaptic connection node 2 can be determined, and the corresponding relationship between sub-output A-3 and non-zero sub-matrix S3 of the weight matrix of synaptic connection node 1 and sub-matrix S4-3 in the weight matrix of synaptic connection node 2 can be determined.

[0156] As Figure 13As shown in Figure (3), after obtaining the optimized weight matrices of synaptic connection node 1 and synaptic connection node 2, and determining the one-to-one correspondence between the sub-output of neuron node A and the optimized weight matrices of synaptic connection node 1 and synaptic connection node 2, when the weight matrix of synaptic connection node 1 can be divided into a block diagonal matrix, three sub-synaptic connection nodes can be created between neuron node C and neuron node A and neuron node B in the target computational graph, namely sub-synaptic connection node 1, sub-synaptic connection node 2, and sub-synaptic connection node 3, and determine the correspondence between sub-synaptic connection node 1 and the sub-output A-1 of neuron node A, determine the correspondence between sub-synaptic connection node 2 and the sub-output A-2 of neuron node A, and determine the correspondence between sub-synaptic connection node 3 and the sub-output A-3 of neuron node A.

[0157] After creating three sub-synaptic connection nodes between neuron node C and neuron node A and neuron node B in the target computational graph, the sum of the calculation result obtained by performing a 2D matrix multiplication on the sub-output A-1 of neuron node A and the non-zero sub-matrix S1 of the weight matrix of synaptic connection node 1, and the calculation result obtained by performing a fully connected operation on the output B of neuron node B and the sub-matrix S4-1 in the weight matrix of synaptic connection node 2 can be calculated using sub-synaptic connection node 1 as the sub-input C-1 of neuron node C.

[0158] It should be noted that the output of neuron node A is the matrix , to calculate the matrix multiplication , then tile the matrix by row dimension. Then copy several copies and arrange them on the diagonal of the weight matrix of the synaptic connection node. The non-zero sub-matrices S1, S2, and S3 of the weight matrix of synaptic connection node 1 are several copies generated by replication. . The sub-output A-1 of neuron node A is only the first row of the matrix , the sub-output A-2 of neuron node A is the second row of the matrix , and the sub-output A-3 of neuron node A is the third row of the matrix .

[0159] Similarly, the sub-input C-2 of neuron node C can be calculated by adding the result obtained from the 2D matrix multiplication of the sub-output A-2 of neuron node A and the non-zero sub-matrix S2 of the weight matrix of synapse connection node 1 with the result obtained from the full connection operation of the output B of neuron node B and the sub-matrix S4-2 in the weight matrix of synapse connection node 2. The sub-input C-3 of neuron node C can be calculated by adding the result obtained from the 2D matrix multiplication of the sub-output A-3 of neuron node A and the non-zero sub-matrix S3 of the weight matrix of synapse connection node 1 with the result obtained from the full connection operation of the output B of neuron node B and the sub-matrix S4-3 in the weight matrix of synapse connection node 2.

[0160] In an embodiment of the present invention, for a specific synapse connection node, such as a synapse connection node representing 2D matrix multiplication elements, its weight matrix is a block diagonal matrix with multiple non-zero sub-matrices arranged diagonally and the rest being zero, so it has the property of being splittable.

[0161] For a general synapse connection node, if its weight matrix is a sparse matrix, the weight matrix of the above synapse connection node can be converted into a block diagonal matrix through the following embodiments. At this time, the synapse connection node can be split into multiple sub-synapse connection nodes, and the redundant zeros in the weight matrix can be removed. By the above method, while ensuring that the calculation structure remains unchanged, as much as possible of the non-0 part of the weight matrix can be deployed to the crossbar array of the computing core. The scale of the weight matrix deployed to the crossbar array of the computing core by the synapse connection node is significantly reduced. Under the condition of the same size of crossbar array hardware resources, a larger weight matrix can be stored, saving the hardware resources of the computing core and improving the computing resource utilization rate and computing efficiency of the brain-inspired chip.

[0162] As an optional embodiment, for any neuron node in the target computation graph, when there are several forward neuron nodes for any neuron node, each synapse connection node between any neuron node and each forward neuron node of any neuron node is determined as each original synapse connection node corresponding to any neuron node. Before the forward neuron node of any neuron node is a neuron node connected to the input side of any neuron node through a synapse connection node, the method further includes: performing row and column transformation on the weight matrix of the synapse connection node in the target computation graph.

[0163] Perform row and column transformations on the weight matrix of the synaptic connection nodes in the target computational graph, including: for any synaptic connection node in the target computational graph, when the weight matrix of any synaptic connection node is a sparse matrix, group the columns in the weight matrix of any synaptic connection node based on the union-find algorithm to obtain the column grouping result of the weight matrix of any synaptic connection node.

[0164] Based on the column grouping result of the weight matrix of any synaptic connection node, perform column transformation on the weight matrix of any synaptic connection node, arrange the columns belonging to the same group in the weight matrix of any synaptic connection node together, and obtain each column transformation matrix corresponding to any synaptic connection node.

[0165] Perform row transformation on each column transformation matrix corresponding to any synaptic connection node column, arrange the rows with all elements being 0 values in each column transformation matrix corresponding to any synaptic connection node column together, and arrange the rows with not all elements being 0 values in each column transformation matrix corresponding to any synaptic connection node column together, to obtain each row transformation matrix corresponding to any synaptic connection node.

[0166] Determine the matrix with not all elements being 0 values in each row transformation matrix corresponding to any synaptic connection node as the non-zero submatrix in each row transformation matrix corresponding to any synaptic connection node. When the number of non-zero submatrices in each row transformation matrix corresponding to any synaptic connection node is 1, splice the row transformation matrices corresponding to any synaptic connection node, and through row transformation, arrange the non-zero submatrices in each row transformation matrix corresponding to any synaptic connection node along the main diagonal to obtain the weight matrix after row and column transformation of any synaptic connection node.

[0167] Figure 14 It is a schematic flow chart of row and column transformation of synaptic connection nodes in the neural network deployment method provided by the present invention. As Figure 14 shown, for any synaptic connection node in the target computational graph, when the size of the weight matrix of the above synaptic connection node is 12×12, if the elements in the weight matrix of the above synaptic connection node include 0 values and non-0 values, it can be determined that the above two arbitrary columns belong to the same group when the 0 values in any two columns in the weight matrix of the above synaptic connection node are in the same row, and then the column grouping result of the weight matrix of the above synaptic connection node can be obtained.

[0168] By grouping the columns in the weight matrix of the above synaptic connection node, it can be identified which rows in the weight matrix of the above synaptic connection node can be combined to form a non-zero submatrix for independent processing.

[0169] Based on the column grouping result of the weight matrix of the above synaptic connection nodes, perform column transformation on the weight matrix of the above synaptic connection nodes, arrange the columns belonging to the same group in the weight matrix of the above synaptic connection nodes together, and obtain each column transformation matrix corresponding to the above synaptic connection nodes.

[0170] Perform row transformation on each column transformation matrix corresponding to the above synaptic connection nodes, arrange the rows in which all elements in each column transformation matrix corresponding to the above synaptic connection nodes are 0 values together, and arrange the rows in which all elements in each column transformation matrix corresponding to the above synaptic connection nodes are not all 0 values together, and obtain each row transformation matrix corresponding to the above synaptic connection nodes.

[0171] Determine the matrix in which all elements in each row transformation matrix corresponding to the above synaptic connection nodes are not all 0 values as the non-zero sub-matrix in each row transformation matrix corresponding to the above synaptic connection nodes. When the number of non-zero sub-matrices in each row transformation matrix corresponding to the above synaptic connection nodes is 1, splice the row transformation matrices corresponding to the above synaptic connection nodes, and through row transformation, arrange the non-zero sub-matrices in each row transformation matrix corresponding to the above synaptic connection nodes along the main diagonal of the matrix to obtain the weight matrix after row and column transformation of the above synaptic connection nodes.

[0172] As Figure 14 shown, the weight matrix after row and column transformation corresponding to the above synaptic connection nodes is a block diagonal matrix, and the non-zero sub-matrices in each row transformation matrix corresponding to the above synaptic connection nodes are equivalent to Figure 14 the non-zero sub-matrix S1, non-zero sub-matrix S2, and non-zero sub-matrix S3 in

[0173] In the embodiments of the present invention, by appropriately adjusting the weight matrix of the synaptic connection nodes in the target computational graph, the weight matrix of the synaptic connection nodes can be effectively deployed on the crossbar array of the computing core, thereby achieving the purpose of improving computing efficiency and reducing resource waste.

[0174] It should be noted that in the embodiments of the present invention, routing grouping is performed on the neuron nodes in the target computational graph to determine the arrangement of the neuron nodes in each target routing group in the computing core, and then an address in the neuron RAM of the computing core can be assigned to each neuron.

[0175] In the embodiments of the present invention, grouping the synaptic connection nodes in the target computational graph is to determine the arrangement of the inputs of each target routing group in the inputs of the computing core and assign the order of the synaptic connection nodes in the inputs of the computing core. The synaptic connection unit of the computing core corresponds to the output destination of the forward neuron nodes in the target computational graph. At the same time, the same address is assigned to the inputs shared by all neuron nodes or sub-routing groups in each target routing group, and the above operations are recursively performed on the sub-routing groups.

[0176] Due to the characteristics of the brain-inspired chip, the routing coordinates need to be allocated according to certain rules.

[0177] The routing structure in the brain-inspired chip is a five-level quad-tree. Each computing core can be indexed by a two-dimensional routing coordinate (x, y), where both x and y are 5-bit numbers. When setting the destination of a neuron, multicast can be achieved by setting several bits at the end of the routing coordinate of the target computing core to be undetermined (denoted by the symbol "*"). For example, setting the routing coordinate of the target computing core to (010**, 100**) means that the data will be sent to (01000, 10000) to (01000, 10111), (01001, 10000) to (01001, 10011), (01010, 10000) to (01001, 10011), (01011, 10000) to (01001, 10011), a total of 16 computing cores.

[0178] In view of this function, the brain-inspired chip requires that the number of computing cores occupied by each routing group needs to be a power of 2. At the same time, the routing coordinate serial number of the first computing core it uses needs to end with several 0s (if using computing cores, the number of 0s at the end needs to be greater than

[0179] At the same time, the brain-inspired chip also requires that in each level of the five-level quad-tree, the data flow between routing nodes cannot form a loop. For example, routing group 1 → routing group 2 → routing group 3 → routing group 1. This requires allocating routing coordinates to each routing group in the order of topological sorting.

[0180] In the embodiments of the present invention, based on the nested relationship between target routing groups, routing coordinates for indicating the positions of computing cores of each target routing group in the brain-inspired chip can be allocated through numerical calculation, mathematical statistics, conditional judgment, and other means.

[0181] As an optional embodiment, allocating routing coordinates for indicating the positions of computing cores of each target routing group in the brain-inspired chip includes: determining the number of computing cores required for each target routing group.

[0182] Based on the topological relationship between neuron nodes and synapse connection nodes in the target computation graph and the number of computing cores required for each target routing group, allocate routing coordinates for each target routing group.

[0183] Specifically, when allocating routing coordinates to a target routing group in an embodiment of the present invention, considering that the target routing group has a nested structure, in order to ensure that the routing coordinates of each target routing group meet the above requirements, the routing coordinate allocation of the target routing group should be carried out according to the following process: First, determine the number of computing cores required for each target routing group.

[0184] Due to the multicast function of the neuromorphic chip, the number of computing cores required for each target routing group is not a simple sum of the computing cores required for all sub-routing groups and neuron nodes in the target routing group. After performing a topological sort on the neuron nodes in the target routing group, the positions of the neuron nodes in the target routing group are determined in sequence according to the alignment requirements, and finally the number of computing cores required for the target routing group is determined. Since the target routing group is a nested structure, this step is a bottom-up recursive process.

[0185] It should be noted that, in order to meet the rules of the routing structure of the neuromorphic chip, in an embodiment of the present invention, based on the topological sort of the target routing group, the number of computing cores required for the target routing group is determined. For any routing group, the number of computing cores required for each sub-routing group is rounded up to the nearest power of 2 in sequence. For any sub-routing group, the subscript of its first computing core in the above target routing group needs to be an integer multiple of the rounded-up number of computing cores required for the sub-routing group. After the positions of all sub-routing groups of the target routing group are determined, the number of computing cores required for the above target routing group can be further determined.

[0186] Since the number of computing cores of a single neuromorphic chip has an upper limit. Therefore, when the remaining computing cores of the neuromorphic chip are not enough to accommodate all target routing groups, the allocation can continue from the first computing core of the next neuromorphic chip.

[0187] Considering the existence of sub-routing groups, the number of computing cores required for a target routing group is not just a simple sum of the computing cores required for the neuron nodes it contains. During the topological sort and alignment process of the sub-routing groups, additional computing core requirements will be generated. After adding this part of the additional requirements to the computing cores required for the neuron nodes, the total number of computing cores required for the target routing group is obtained.

[0188] According to the topological relationship between neuron nodes and synaptic connection nodes in the target computation graph, set the starting coordinates of the computing cores for each target routing group, and then the routing coordinates of each target routing group can be obtained based on the number of computing cores required for each target routing group.

[0189] It should be noted that since the target routing group is a nested structure, this step is a top-down recursive process.

[0190] After assigning routing coordinates to each target routing group, configuration information may be generated for each used computing core, including but not limited to weight data, computing core attributes, and neuron unit information.

[0191] The weight data can be determined based on the input neurons and output neurons connected within the computational core. If there is a synapse that directly connects the input and the output, the weight matrix data of the synapse is used; otherwise, the weight of this part is set to zero.

[0192] All computing cores in the same routing group share the same property settings, such as random number seeds, etc. These properties have been defined in the previous steps, and the computing core only needs to apply the corresponding properties of the routing group to which it belongs.

[0193] Neuron unit information includes the properties of the neuron unit and the details of its output destination. The properties of the neuron unit are determined by the neural network to be deployed. The output destination information involves the routing coordinates of the computational kernel and its order in the input. By locating the routing group where the output destination of the neuron unit is located, the routing coordinates of the group are obtained as part of the target information; at the same time, the synaptic distribution within the relevant routing group is found to obtain the order of the neurons in the input of the routing group to form the target information of the complete neuron.

[0194] Step 103: Generate an executable file based on the configuration information of the computing core, where the executable file includes binary instructions recognizable by the brain-like chip.

[0195] Although the configuration information of the computing core comprehensively describes the neural network to be deployed, including the weight data, neuron unit properties and data flow information of the neural network to be deployed, this configuration information is only a data structure that is easy for humans to understand and check, and cannot be directly executed on brain-like chips.

[0196] Therefore, in the embodiment of the present invention, after generating the configuration information of the computing core in the brain-like chip, the configuration information of the computing core can be converted into corresponding machine code according to the instruction set of the brain-like chip and following specific rules. The generated machine code is stored in the form of a file. These data can be downloaded to the brain-like chip and executed, thereby realizing the deployment and operation of the neural network to be deployed.

[0197] In the embodiment of the present invention, after generating a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, performing simulation calculation based on the target computation graph and obtaining the simulation calculation result, and when it is determined that the target computation graph passes the verification based on the simulation calculation result, generating configuration information of the computing cores in the brain-inspired chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target computation graph, and further generating an executable file including binary instructions recognizable by the brain-inspired chip based on the configuration information of the computing cores, the neural network can be deployed to the brain-inspired chip more accurately and efficiently, the computing resource allocation of the brain-inspired chip can be optimized, and it has broad application prospects.

[0198] In the neural network deployment method provided by the present invention, modular encapsulation of operators in the neural network, connection rules between operators, and a conversion method from the original computation graph to a computation graph constructed by operators in an operator set supported by the brain-inspired chip hardware are realized. By calling these modular operators, users can flexibly combine according to specific task requirements to construct the required neural network, greatly increasing the degree of freedom in design. Moreover, the neural network deployment method provided by the present invention clarifies the connection rules between operators, can effectively prevent connection errors that may occur in the design process, and improves the development efficiency of the neural network. Through this conversion method, a neural network constructed using other frameworks can maintain its original network structure and parameter information on the premise of meeting the hardware limitations of the brain-inspired chip, and realize migration and deployment to the brain-inspired chip, enhancing the applicability and flexibility of the brain-inspired chip.

[0199] In the neural network deployment method provided by the present invention, by recording the states of neurons and synapses in the neural network at each simulation time step, cycle-accurate simulation of the neural network based on the hardware computing principle of the brain-inspired chip can be realized, which can provide detailed data support for subsequent analysis and optimization, not only improving the efficiency of neural network debugging and performance optimization, but also effectively shortening the entire development cycle.

[0200] In the neural network deployment method provided by the present invention, the unique hardware architecture of the brain-inspired chip and the advantages of routing communication are fully utilized, and the computing core resources occupied by the deployment of multiple specific operators are reduced. These improvements enhance the deployment efficiency of the neural network, and a larger-scale neural network can be deployed in the same number of computing cores. Moreover, according to the neural network deployment method provided by the present invention, a larger-scale neural network can be deployed to multiple brain-inspired chips, significantly enhancing the scalability of the system.

[0201] The neural network deployment method provided by the present invention is particularly applicable to neural networks with complex network topologies. This not only significantly enhances the flexibility and adaptability of routing but also enables networks with complex topological structures to fully utilize the multicast function of neuromorphic chips. In the field of modern deep learning, the residual structure has become a core component in neural network design and is widely seen in various network architectures such as Residual Network (ResNet). The present invention expands the application scope of neuromorphic chips. By supporting the deployment of branch network structures, the residual structure can also be deployed according to the method provided by the present invention.

[0202] Figure 15 It is a schematic structural diagram of the neural network deployment device provided by the present invention. The following combines Figure 15 to describe the neural network deployment device provided by the present invention. The neural network deployment device described below can be correspondingly referred to the neural network deployment method provided by the present invention described above. As Figure 15 shown, the device includes: a computation graph generation module 1501, a configuration generation module 1502, and a file generation module 1503.

[0203] The computation graph generation module 1501 is configured to generate a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and a target mapping relationship. The target mapping relationship is used to describe the mapping relationship between operators in the neural network and hardware units in the neuromorphic chip. The type of the hardware unit is a neuron unit or a synaptic connection unit. Any neuron node in the target computation graph represents one of the neuron units, any synaptic connection node in the target computation graph represents one of the synaptic connection units, and a unidirectional arrow used to connect two hardware nodes in the target computation graph represents the input-output relationship between the two hardware nodes; The configuration generation module 1502 is configured to perform simulation calculations based on the target computation graph and obtain simulation calculation results. When it is determined that the target computation graph passes the verification based on the simulation calculation results, generate configuration information of the computing cores in the neuromorphic chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target computation graph.

[0204] The file generation module 1503 is configured to generate an executable file based on the configuration information of the computing cores. The executable file includes binary instructions recognizable by the neuromorphic chip.

[0205] Specifically, the computation graph generation module 1501, the configuration generation module 1502, and the file generation module 1503 are electrically connected.

[0206] In the neural network deployment device according to the embodiments of the present invention, after generating a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, performing simulation computation based on the target computation graph and obtaining the simulation computation result, and when it is determined that the target computation graph passes the verification based on the simulation computation result, generating configuration information of computation cores in the brain-inspired chip based on the topological connection relationship between neuron nodes and synaptic connection nodes in the target computation graph, and further generating an executable file including binary instructions recognizable by the brain-inspired chip based on the configuration information of the computation cores, it is possible to deploy the neural network to the brain-inspired chip more accurately and efficiently, optimize the allocation of computation resources of the brain-inspired chip, improve the utilization rate of the computation resources of the brain-inspired chip, and have broad application prospects.

[0207] Figure 16 An entity structure diagram of an electronic device is exemplified, as Figure 16 shown. The electronic device may include: a processor 1610, a communication interface 1620, a memory 1630, and a communication bus 1640. Among them, the processor 1610, the communication interface 1620, and the memory 1630 complete communication with each other through the communication bus 1640. The processor 1610 may call logic instructions in the memory 1630 to execute a neural network deployment method, and the method includes: generating a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, where the target mapping relationship is used to describe the mapping relationship between operators in the neural network and hardware units in the brain-inspired chip, the type of the hardware unit is a neuron unit or a synaptic connection unit, any neuron node in the target computation graph represents a neuron unit, any synaptic connection node in the target computation graph represents a synaptic connection unit, and a one-way arrow used to connect two hardware nodes in the target computation graph represents the input-output relationship between the two hardware nodes; performing simulation computation based on the target computation graph and obtaining the simulation computation result, and when it is determined that the target computation graph passes the verification based on the simulation computation result, generating configuration information of computation cores in the brain-inspired chip based on the topological connection relationship between neuron nodes and synaptic connection nodes in the target computation graph; generating an executable file based on the configuration information of the computation cores, and the executable file includes binary instructions recognizable by the brain-inspired chip.

[0208] In addition, when the logical instructions in the above-mentioned memory 1630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0209] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the neural network deployment method provided by the above-mentioned various methods. The method includes: generating a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, where the target mapping relationship is used to describe the mapping relationship between the operators in the neural network and the hardware units in the brain-inspired chip, the type of the hardware unit is a neuron unit or a synaptic connection unit, any neuron node in the target computation graph represents a neuron unit, any synaptic connection node in the target computation graph represents a synaptic connection unit, and the unidirectional arrow used to connect two hardware nodes in the target computation graph represents the input-output relationship between the two hardware nodes; performing simulation calculations based on the target computation graph and obtaining the simulation calculation results. When it is determined that the target computation graph passes the verification based on the simulation calculation results, generating configuration information of the computing cores in the brain-inspired chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target computation graph; generating an executable file based on the configuration information of the computing cores, where the executable file includes binary instructions recognizable by the brain-inspired chip.

[0210] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the neural network deployment method provided by the above-mentioned various methods. The method includes: generating a target computation graph corresponding to the neural network to be deployed based on the network structure of the neural network to be deployed and the target mapping relationship, where the target mapping relationship is used to describe the mapping relationship between the operators in the neural network and the hardware units in the brain-inspired chip. The type of the hardware unit is a neuron unit or a synaptic connection unit. Any neuron node in the target computation graph represents a neuron unit, and any synaptic connection node in the target computation graph represents a synaptic connection unit. The unidirectional arrow used to connect two hardware nodes in the target computation graph represents the input-output relationship between the two hardware nodes; performing simulation calculations based on the target computation graph and obtaining the simulation calculation results. When it is determined based on the simulation calculation results that the target computation graph passes the verification, generating the configuration information of the computation cores in the brain-inspired chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target computation graph; generating an executable file based on the configuration information of the computation cores, and the executable file includes binary instructions recognizable by the brain-inspired chip.

[0211] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0212] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0213] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A neural network deployment method, characterized in that: include: Based on the network structure and target mapping relationship of the neural network to be deployed, a target computation graph corresponding to the neural network to be deployed is generated, wherein the target mapping relationship is used to describe the mapping relationship between operators in the neural network and hardware units in the brain-like chip, wherein the type of the hardware unit is a neuron unit or a synaptic connection unit, and any neuron node in the target computation graph represents a neuron unit, and any synaptic connection node in the target computation graph represents a synaptic connection unit, and a one-way arrow in the target computation graph used to connect two hardware nodes represents an input-output relationship between the two hardware nodes; Perform simulation calculation based on the target calculation graph and obtain simulation calculation results, and when it is determined that the target calculation graph passes the verification based on the simulation calculation results, generate configuration information of the computing core in the brain-like chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target calculation graph; Based on the configuration information of the computing core, an executable file is generated, wherein the executable file includes binary instructions recognizable by the brain-like chip.

2. The neural network deployment method according to claim 1, characterized in that: The generating a target calculation graph corresponding to the neural network to be deployed based on the network structure and the target mapping relationship of the neural network to be deployed includes: Based on the network structure of the neural network to be deployed, generating an original calculation graph corresponding to the neural network to be deployed, wherein any operator node in the original calculation graph represents an operator in the neural network to be deployed, and a one-way arrow in the original calculation graph used to connect two operator nodes represents an input-output relationship between the two operator nodes; For a type I operator node in the original calculation graph, based on the first mapping sub-relationship in the target mapping relationship, the type I operator node is replaced with the hardware unit node; for a type II operator node in the original calculation graph, based on the second mapping sub-relationship in the target mapping relationship, the type II operator node is replaced with a complex operator node, to obtain an intermediate calculation graph corresponding to the neural network to be deployed, wherein the first mapping sub-relationship is used to describe the mapping relationship between a type I operator in the neural network and a hardware unit in the brain-like chip, and the second mapping sub-relationship is used to describe the mapping relationship between a type II operator in the neural network and a complex operator, the type I operator and type II operator in the neural network are predefined, any type I operator node in the original calculation graph represents a type I operator, any type II operator node in the original calculation graph represents a type II operator, and any complex operator node in the intermediate calculation graph represents a complex operator; Based on the third mapping sub-relationship in the target mapping relationship, the complex operator nodes in the intermediate calculation graph are replaced with the hardware unit nodes to obtain the target calculation graph. The third mapping sub-relationship is used to describe the mapping relationship between the complex operators in the neural network and the hardware units in the brain-like chip.

3. The neural network deployment method according to claim 1, characterized in that: The generating of configuration information of the computing core in the brain-like chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target computing graph includes: Grouping each neuron unit node in the target computation graph to obtain a plurality of target routing groups; Segmenting and optimizing the weight matrix of the target synaptic connection node in the target calculation graph to obtain the optimized weight matrix of the target synaptic connection node; Routing coordinates for indicating the location of the computing core of each target routing group in the brain-like chip are assigned to each of the target routing groups, and then the configuration information of the computing core is generated based on the neuronal units included in each of the target routing groups, the nested relationship between the target routing groups, the routing coordinates of each of the target routing groups, and the optimized weight matrix of the target synaptic connection nodes.

4. The neural network deployment method according to claim 3, characterized in that: The step of grouping the neuron unit nodes in the target computation graph to obtain a plurality of target routing groups includes: For each neuron node in the target computation graph, each successor neuron node of each neuron node is determined as an original routing group, and each neuron node is determined as an input neuron node of the original routing group, the successor neuron node of the neuron node is a neuron node located at the output side of the neuron node and connected to the neuron node through a synaptic connection node, and any synaptic connection node in the target computation graph represents a synaptic connection unit in the brain-like chip; Traversing each of the original routing groups, if there are identical neuron nodes in any two original routing groups, merging the two original routing groups and removing duplicates to obtain an intermediate routing group, and determining the input neuron nodes in the input neuron nodes of the two original routing groups that are not in the intermediate routing group as the input neuron nodes of the intermediate routing group; Traversing each of the intermediate routing groups, in the case that any intermediate routing group includes any neuron node and a successor neuron node of the any neuron node, determining the successor neuron node of the any neuron node as a sub-routing group in the any intermediate routing group, and determining the any neuron node as an input neuron node of the sub-routing group in the any intermediate routing group; Each of the original routing groups, each of the intermediate routing groups, and each of the sub-routing groups in each of the intermediate routing groups are determined as each of the target routing groups.

5. The neural network deployment method according to claim 3, characterized in that: The segmenting and optimizing the weight matrix of the target synaptic connection node in the target computation graph to obtain the optimized weight matrix of the target synaptic connection node includes: For any neuron node in the target computation graph, when there are several forward neuron nodes in the any neuron node, each synaptic connection node between the any neuron node and each forward neuron node of the any neuron node is determined as each original synaptic connection node corresponding to the any neuron node, and the forward neuron node of the any neuron node is a neuron node connected to the input side of the any neuron node through a synaptic connection node; In the case where the weight matrix of any original synaptic connection node corresponding to any neuron node is a block diagonal matrix, the any original synaptic connection node corresponding to any neuron node is determined as a first target synaptic connection node among target synaptic connection nodes corresponding to any neuron node, and the other original synaptic connection nodes corresponding to any neuron node except the any original synaptic connection node are determined as second target synaptic connection nodes among target synaptic connection nodes corresponding to any neuron node, the elements in the submatrix located on the main diagonal of the block diagonal matrix are not all 0, and the elements in the block diagonal matrix except the non-zero submatrix located on the main diagonal are all 0; Determine each submatrix on the main diagonal of the weight matrix of the first target synaptic connection node corresponding to any neuron node as the optimized weight matrix of the first target synaptic connection node corresponding to any neuron node; Based on the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any one of the neuron nodes, the weight matrix of the second target synaptic connection node corresponding to any one of the neuron nodes is divided into a plurality of sub-matrices in the column dimension as the optimized weight matrix of the second target synaptic connection node corresponding to any one of the neuron nodes, and the number of the optimized weight matrices of the second target synaptic connection node corresponding to any one of the neuron nodes is the same as the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any one of the neuron nodes; Determine, based on the column dimension, a correspondence between a weight matrix after optimization of a first target synaptic connection node and a column of a weight matrix after optimization of a second target synaptic connection node corresponding to any one of the neuron nodes; Based on the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to the any neuron node, the output of the forward neuron node connected to the input side of the any neuron node through the first target synaptic connection node is divided into a plurality of sub-outputs, and a one-to-one correspondence between each of the sub-outputs and the optimized weight matrix of the target synaptic connection node corresponding to the any neuron node is determined; Based on the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any one of the neuron nodes, multiple sub-synaptic connection nodes are created between any one of the neuron nodes and each of the neuron nodes of any one of the neuron nodes in the target calculation graph, and a one-to-one correspondence between each of the sub-synaptic connection nodes and each of the sub-outputs is determined, the number of sub-synaptic connection nodes between any one of the neuron nodes and each of the neuron nodes of any one of the neuron nodes is the same as the number of non-zero sub-matrices in the weight matrix of the first target synaptic connection node corresponding to any one of the neuron nodes, and each sub-synaptic connection node between any one of the neuron nodes and each of the neuron nodes of any one of the neuron nodes represents a calculation task between each of the sub-outputs and the optimized weight matrix of the target synaptic connection node corresponding to any one of the neuron nodes.

6. The neural network deployment method according to claim 5, characterized in that: For any neuron node in the target computation graph, when there are several forward neuron nodes in the any neuron node, each synaptic connection node between the any neuron node and each forward neuron node of the any neuron node is determined as each original synaptic connection node corresponding to the any neuron node, and the forward neuron node of the any neuron node is before the neuron node connected to the input side of the any neuron node through the synaptic connection node, the method further includes: Performing a row-column transformation on the weight matrix of the synaptic connection nodes in the target computation graph; The performing row-column transformation on the weight matrix of the synaptic connection nodes in the target computation graph includes: For any synaptic connection node in the target computation graph, when the weight matrix of any synaptic connection node is a sparse matrix, grouping the columns in the weight matrix of any synaptic connection node based on a union-find algorithm to obtain a column grouping result of the weight matrix of any synaptic connection node; Based on the column grouping result of the weight matrix of any synaptic connection node, the weight matrix of any synaptic connection node is transformed into a column, the columns belonging to the same group in the weight matrix of any synaptic connection node are arranged together, and each column transformation matrix corresponding to the any synaptic connection node is obtained; Performing row transformation on each column transformation matrix corresponding to any synaptic connection node column, arranging rows in which all elements in each column transformation matrix corresponding to any synaptic connection node column are 0 values ​​together, arranging rows in which all elements in each column transformation matrix corresponding to any synaptic connection node column are not all 0 values ​​together, and obtaining each row transformation matrix corresponding to any synaptic connection node; A matrix in which all elements in each row transformation matrix corresponding to any synaptic connection node are not all zero values ​​is determined as a non-zero submatrix in each row transformation matrix corresponding to any synaptic connection node. When the number of non-zero submatrices in each row transformation matrix corresponding to any synaptic connection node is 1, the row transformation matrices corresponding to any synaptic connection node are concatenated, and through row transformation, the non-zero submatrices in each row transformation matrix corresponding to any synaptic connection node are arranged along the main diagonal to obtain a weight matrix after row-column transformation of any synaptic connection node.

7. The neural network deployment method according to claim 3, characterized in that: The step of allocating routing coordinates for indicating the location of a computing core of each target routing group in the brain-like chip includes: Determining the number of computing cores required for each of the target routing groups; The routing coordinates are allocated to each of the target routing groups based on the topological relationship between the neuron nodes and the synaptic connection nodes in the target computing graph and the number of computing cores required for each of the target routing groups.

8. The neural network deployment method according to any one of claims 1 to 7, characterized in that: The performing simulation calculation based on the target calculation graph and obtaining the simulation calculation result includes: Adding probes to the hardware unit nodes and properties that need to be verified in the target computation graph to obtain the computation graph to be simulated; Inputting the calculation graph to be simulated into the simulator, so that the simulator performs simulation calculation based on the calculation graph to be simulated, and triggering global simulation number recording; When the number of global simulations does not reach a predefined number threshold, the simulator's update method is called to update the states of the input nodes, synaptic connection nodes, and neuron nodes in the computational graph to be simulated in turn, and save the attribute information of the hardware unit nodes that need to be verified in the target computational graph as the simulation calculation result of this simulation calculation. If the number of global simulations reaches a predefined number threshold, the simulation is determined to be finished, and the input nodes in the target computational graph are used to input data.

9. A neural network deployment device, characterized in that: include: A computation graph generation module, used to generate a target computation graph corresponding to the neural network to be deployed based on the network structure and target mapping relationship of the neural network to be deployed, wherein the target mapping relationship is used to describe the mapping relationship between operators in the neural network and hardware units in the brain-like chip, wherein the type of the hardware unit is a neuron unit or a synaptic connection unit, and any neuron node in the target computation graph represents a neuron unit, and any synaptic connection node in the target computation graph represents a synaptic connection unit, and a one-way arrow in the target computation graph used to connect two hardware nodes represents an input-output relationship between the two hardware nodes; A configuration generation module, used to perform simulation calculations based on the target calculation graph and obtain simulation calculation results, and generate configuration information of the computing core in the brain-like chip based on the topological connection relationship between the neuron nodes and the synaptic connection nodes in the target calculation graph when it is determined that the target calculation graph passes the verification based on the simulation calculation results; The file generation module is used to generate an executable file based on the configuration information of the computing core, and the executable file includes binary instructions recognizable by the brain-like chip.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the neural network deployment method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Multilevel intermediate representation compiling method of heterogeneous neural network and electronic equipment

    CN122489079A