Heterogeneous cluster-oriented spiking neural network distributed deployment and simulation method
By building a reversible pulse neural network and a hierarchical cluster communication architecture in the pulse neural network simulator, the problems of performance bottlenecks and communication redundancy in large-scale neural network simulation are solved, and efficient distributed deployment and simulation methods are realized.
Patent Information
- Application Number
- CN202411941295.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
AI Technical Summary
The existing pulsed neural network (SNN) emulators have performance bottlenecks in large-scale neural network simulations, especially in cluster systems, where the total number of communication between servers and high message redundancy is difficult to achieve efficient simulation of super-large-scale neural networks.
A distributed deployment and simulation method for pulsed neural networks for heterogeneous clusters is proposed. By determining the bidirectional connection relationship between target presynaptic neurons and target postsynaptic neurons based on preset synaptic numbers and neuron numbers, a reversible pulsed neural network is constructed; heterogeneous computing units are divided, a hierarchical cluster communication architecture is constructed, and computing efficiency is evaluated to perform network segmentation and task division operations.
By reducing the coupling between cluster servers, the communication efficiency and scalability of the system are improved, and the problem of incompatibility of forward and reverse connection representation forms in the existing SNN network simulation process is solved, and the simulation performance of large-scale neural networks is significantly improved.
Smart Images

Figure CN120012841A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large-scale spiking neural networks, and in particular to a spiking neural network distributed deployment and simulation method for heterogeneous clusters. Background Art
[0002] Compared with traditional computing-intensive load applications, the simulation of spiking neural networks (SNNs) has a unique computing mode. First, during the SNN calculation process, all neurons and synapses in the network can be calculated simultaneously. Therefore, it has parallelism and can be used for distributed simulation. Secondly, the simulation of SNNs is based on time steps and continuously updates itself. Its calculation process is rhythmic and needs to be synchronized.
[0003] However, with the continuous deepening of brain neurology research, traditional SNN simulators are unable to complete large-scale neural network simulations under limited computing power platforms. Current technologies, such as multi-threading and GPU, have certain bottlenecks and deficiencies, which greatly limit the simulation and verification of large-scale network models by SNN simulators. Therefore, in view of the above bottleneck problems, the present invention discloses a cluster-oriented large-scale SNN distributed deployment and simulation method.
[0004] The simulator is the main software tool in the field of SNN simulation. It uses a designed software platform and computing library to achieve efficient SNN simulation on a general-purpose computing CPU / GPU platform. Typical simulators include NEURON, NEST, Brain2, Nengo, Carlsim, etc.
[0005] NEST (Neural Simulation Tool) is an advanced open source software toolkit. When using it, users can freely choose or define different neuron dynamics models and synaptic connection rules according to their own needs. In terms of application interface, NEST provides a detailed Python interface. Although NEST uses C++ as the development language, it provides the Python interface of PyNEST. Through PyNEST, users can use Python language to configure, control simulation, and analyze simulation results to improve the convenience and flexibility of use.
[0006] Existing simulators clearly have performance bottlenecks. As the size of the SNN network increases, the computational load of the simulator also increases. However, there is still room for improvement in the design of the current simulator architecture, especially in terms of cluster and heterogeneous cluster support. Currently, for the current popular GPU / CPU heterogeneous architecture, most simulators cannot simultaneously utilize the computing power provided by the CPU and GPU.
[0007] In addition, cluster acceleration systems are another solution to the current problem of insufficient computing power. Such systems utilize the computing power of multiple integrated systems in a server cluster environment, and can achieve higher performance during actual operation. In the latest SNN simulators, this type of architectural design is widely used. For example, NEST supports parallel computing and distributed computing in terms of architectural design, and can run efficiently on hardware platforms such as multi-core processors and computer clusters to increase simulation speed. SINABS is developed based on PyTorch. Users can use all PyTorch computing models and support multi-GPU operation. PyNN is a python library that provides SNN simulation. Its architecture adopts a multi-core parallel design, covering everything from single CPUs to large-scale parallel computing systems. The solution can select the most appropriate resources for simulation experiments as needed.
[0008] Although good performance benefits can be obtained through the cluster system and the technical bottleneck of improving performance can be converted into the hardware cost of the cluster server; but in essence, the infinite superposition of cluster scale will put higher requirements on cluster communication and data exchange; further, in order to pursue higher performance indicators, as the cluster scale increases, the scale of data exchange will also increase accordingly. In a limited network environment, network communication capacity will become a new bottleneck restricting the continuous expansion of cluster scale.
[0009] NEST can be deployed on FPGA systems, and a special version of Nengo also supports deploying trained network models on hardware platforms such as FPGA, ASIC, etc. to achieve efficient real-time processing. SINABS can accurately implement SNN simulation based on aiCTX processors, and provides functions such as conversion of pulse sequences to analog values. BrainScaleS is a large-scale simulation system for neurological simulation that implements mixed signals.
[0010] Although customized hardware can speed up the simulation process and has a high energy efficiency ratio, it will bring new costs, especially ASIC chips, which have long design and development cycles and large initial investments, which may restrict the development of such simulation methods. At the same time, customized hardware is not very flexible. When the demand changes, the difficulty of hardware modification and high technical cost will also become a new adaptability bottleneck.
[0011] As research continues to deepen, neuron models and synapse models have also changed, from simple IF neurons to complex HH models, which has put forward higher requirements on the adaptability of simulation systems. To meet such requirements, various simulation platforms have proposed different solutions. Existing solutions often directly use neuron clusters as the basic unit for SNN network connection relationships and data storage representation.
[0012] Traditional SNN models have high parameter integration and strong structural coupling, which leads to complex data storage, large communication volume, and redundant transmission. For example, neuron clusters are often of different sizes, and directly using clusters as task division units will lead to unbalanced load. Especially under limited network conditions, other constraints such as communication bandwidth and communication delay will become new performance bottlenecks, thereby reducing the efficiency of simulation.
[0013] Therefore, the prior art mainly has the following problems:
[0014] 1. Simulation performance is restricted
[0015] In the acceleration strategy based on multi-threading, the simulation performance is constrained by the performance of the processor, making it impossible to simulate large-scale neural networks.
[0016] 2. Network communication quality becomes a new bottleneck;
[0017] In the strategy based on cluster servers, although a larger-scale neural network simulation can be achieved, the increase in the total amount of communication between cluster servers and the deepening of data redundancy hinders the realization of large-scale neural network simulation.
[0018] 3. In the traditional method, the coupling degree between models and algorithms is high at the SNN data representation layer, which leads to strong correlation between various components and modules in the simulator, hindering its simulation of ultra-large-scale neural networks.
[0019] In summary, the forward connection of the existing SNN network simulation process and the reverse connection representation of the SNN learning algorithm are incompatible, and in the cluster SNN simulation system, the total amount of communication between servers is large and the message redundancy is high, which needs to be solved urgently. Summary of the invention
[0020] The present application provides a distributed deployment and simulation method of a pulse neural network for heterogeneous clusters to solve the problems of incompatibility between the forward connection representation of the existing SNN network simulation process and the reverse connection representation of the SNN learning algorithm, and the large total amount of communication between servers and high message redundancy in the cluster SNN simulation system.
[0021] The first aspect of the present application provides a method for distributed deployment and simulation of a pulse neural network for heterogeneous clusters, comprising the following steps: determining a bidirectional connection relationship between a target presynaptic neuron and a target postsynaptic neuron based on a preset synapse number and a neuron number, so as to construct a reversible pulse neural network according to the bidirectional connection relationship; dividing a preset plurality of heterogeneous computing units to obtain a plurality of sub-communication architectures, and constructing a hierarchical cluster communication architecture through the plurality of sub-communication architectures and a preset communication strategy; evaluating the computational efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture, so as to perform network segmentation and task division operations of the reversible pulse neural network based on the computational efficiency and the hierarchical cluster communication architecture.
[0022] Optionally, in one embodiment of the present application, the determining of the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron based on the preset synapse number and neuron number includes: obtaining and judging the connection requirement between the target presynaptic neuron and the target postsynaptic neuron; when the connection requirement is a forward connection requirement, determining the number of synapses and synapse numbers of all post-synaptic synapses corresponding to each neuron, and constructing at least one first target array corresponding to the forward connection requirement according to the number of synapses and synapse numbers of all post-synapses; when the connection requirement is a reverse connection requirement, determining the number of synapses and synapse numbers of all post-synapses corresponding to each neuron; The invention relates to a method for obtaining a first target array corresponding to the reverse connection requirement according to the number of synapses and synapse numbers of all the corresponding pre-connection synapses of the neurons, and constructing at least one first target array corresponding to the reverse connection requirement according to the number of synapses and synapse numbers of all the pre-connection synapses of the neurons; obtaining the address and size information of each first target array in the at least one first target array corresponding to the number of neurons in the reversible pulse neural network and the forward connection requirement or the reverse connection requirement, and determining a second target array according to the number of neurons and the address and size information; determining the bidirectional connection relationship based on the second target array and the at least one first target array corresponding to the forward connection requirement or the reverse connection requirement.
[0023] Optionally, in one embodiment of the present application, the preset multiple heterogeneous computing units are divided to obtain multiple sub-communication architectures, and a hierarchical cluster communication architecture is constructed through the multiple sub-communication architectures and the preset communication strategy, including: dividing the multiple heterogeneous computing units according to a preset physical classification strategy or logical classification strategy to generate the multiple sub-communication architectures, and determining the key units in each of the multiple sub-communication architectures; when communicating within each sub-communication architecture, the target communication data is passed to the target heterogeneous computing unit in the current sub-communication architecture through the heterogeneous computing unit in the current sub-communication architecture; when communicating between different sub-communication architectures, the target communication data is passed to the key unit in the current sub-communication architecture in the current sub-communication architecture, and the target communication data is passed to the key unit in the target sub-communication architecture using the key unit in the current sub-communication architecture, and the target communication data is sent to the target heterogeneous computing unit in the target sub-communication architecture through the key unit in the target sub-communication architecture.
[0024] Optionally, in one embodiment of the present application, the computing efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture is evaluated to perform network segmentation and task division operations of the reversible pulse neural network based on the computing efficiency and the hierarchical cluster communication architecture, including: determining the scale of the network segmentation of the reversible pulse neural network according to the computing efficiency, and calculating the sum of the computing efficiencies of all heterogeneous computing units in each sub-communication architecture, so as to perform a first task division operation using the sum of the computing efficiency and the scale to obtain multiple sub-networks corresponding to the reversible pulse neural network; calculating the sum of the sub-network computing efficiencies of all heterogeneous computing units within each sub-communication architecture in each of the multiple sub-networks, so as to perform a second task division on each sub-network according to the sum of the sub-network computing efficiencies.
[0025] The second aspect of the present application provides a distributed deployment and simulation device of a pulse neural network for heterogeneous clusters, including: a construction module, which is used to determine the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron based on the preset synapse number and neuron number, so as to construct a reversible pulse neural network according to the bidirectional connection relationship; a division module, which is used to divide the preset multiple heterogeneous computing units to obtain multiple sub-communication architectures, and construct a hierarchical cluster communication architecture through the multiple sub-communication architectures and preset communication strategies; an evaluation module, which is used to evaluate the computing efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture, so as to perform network segmentation and task division operations of the reversible pulse neural network based on the computing efficiency and the hierarchical cluster communication architecture.
[0026] Optionally, in one embodiment of the present application, the construction module includes: a judgment unit, which is used to obtain and judge the connection requirement between the target pre-synaptic neuron and the target post-synaptic neuron; a forward connection unit, which is used to determine the number of synapses and synapse numbers of all post-neuron connecting synapses corresponding to each neuron when the connection requirement is a forward connection requirement, and construct at least one first target array corresponding to the forward connection requirement based on the number of synapses and synapse numbers of all post-neuron connecting synapses; a reverse connection unit, which is used to determine the number of synapses and synapse numbers of all pre-neuron connecting synapses corresponding to each neuron when the connection requirement is a reverse connection requirement. The invention relates to a method for determining a bidirectional connection relationship based on the number of neurons in the reversible spiking neural network and the address and size information of each first target array in the at least one first target array corresponding to the forward connection requirement or the reverse connection requirement, and determining a second target array based on the number of neurons and the address and size information. The invention also relates to a method for determining a bidirectional connection relationship based on the second target array and the at least one first target array corresponding to the forward connection requirement or the reverse connection requirement.
[0027] Optionally, in one embodiment of the present application, the division module includes: a generation unit, which is used to divide the multiple heterogeneous computing units according to a preset physical classification strategy or logical classification strategy to generate the multiple sub-communication architectures and determine the key units in each of the multiple sub-communication architectures; a first communication unit, which is used to pass the target communication data to the target heterogeneous computing unit in the current sub-communication architecture through the heterogeneous computing unit in the current sub-communication architecture when communicating within each sub-communication architecture; a second communication unit, which is used to pass the target communication data to the key unit in the current sub-communication architecture in the current sub-communication architecture when communicating between different sub-communication architectures, and use the key unit in the current sub-communication architecture to pass the target communication data to the key unit in the target sub-communication architecture, and send the target communication data to the target heterogeneous computing unit in the target sub-communication architecture through the key unit in the target sub-communication architecture.
[0028] Optionally, in one embodiment of the present application, the evaluation module includes: a first computing unit, used to determine the scale of the network segmentation of the reversible pulse neural network according to the computing efficiency, and calculate the sum of the computing efficiencies of all heterogeneous computing units in each sub-communication architecture, so as to use the sum of the computing efficiency and the scale to perform a first task division operation to obtain multiple sub-networks corresponding to the reversible pulse neural network; a second computing unit, used to calculate the sum of the subnetwork computing efficiencies of all heterogeneous computing units within each sub-communication architecture in each of the multiple subnetworks, so as to perform a second task division on each subnetwork according to the sum of the subnetwork computing efficiencies.
[0029] The third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the distributed deployment and simulation method of pulse neural networks for heterogeneous clusters as described in the above embodiments.
[0030] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned distributed deployment and simulation method of pulse neural networks for heterogeneous clusters.
[0031] The fifth aspect of the present application provides a computer program product, including a computer program, which is executed to implement the above-mentioned distributed deployment and simulation method of pulse neural networks for heterogeneous clusters.
[0032] Therefore, the embodiments of the present application have the following beneficial effects:
[0033] The embodiments of the present application can determine the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron based on the preset synapse number and neuron number, so as to construct a reversible pulse neural network according to the bidirectional connection relationship; divide the preset multiple heterogeneous computing units to obtain multiple sub-communication architectures, and construct a hierarchical cluster communication architecture through multiple sub-communication architectures and preset communication strategies; evaluate the computing efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture, so as to perform network segmentation and task division operations of the reversible pulse neural network based on the computing efficiency and the hierarchical cluster communication architecture. The present application classifies heterogeneous computing units according to physical and logical characteristics, thereby reducing the coupling degree between cluster servers and improving the communication efficiency and scalability of the system. As a result, the problems of incompatibility between the forward connection of the existing SNN network simulation process and the reverse connection representation of the SNN learning algorithm, and the large amount of communication between servers and high message redundancy in the cluster SNN simulation system are solved.
[0034] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0036] Figure 1 A flowchart of a distributed deployment and simulation method of a spiking neural network for heterogeneous clusters provided according to an embodiment of the present application;
[0037] Figure 2 A schematic diagram of a representation form supporting forward and reverse connections provided according to an embodiment of the present application;
[0038] Figure 3 A schematic diagram of a hierarchical communication architecture provided according to an embodiment of the present application;
[0039] Figure 4 This is an example diagram of a spiking neural network distributed deployment and simulation device for heterogeneous clusters according to an embodiment of the present application;
[0040] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0041] Among them, 10-distributed deployment and simulation device of pulse neural network for heterogeneous clusters; 100-building module, 200-partitioning module, 300-evaluation module; 501-memory, 502-processor, 503-communication interface. DETAILED DESCRIPTION
[0042] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0043] The following describes the distributed deployment and simulation method of a spiking neural network for heterogeneous clusters of the embodiment of the present application with reference to the accompanying drawings. In view of the problems mentioned in the above background technology, the present application provides a distributed deployment and simulation method of a spiking neural network for heterogeneous clusters, in which the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron is determined based on the preset synapse number and neuron number, so as to construct a reversible spiking neural network according to the bidirectional connection relationship; the preset multiple heterogeneous computing units are divided to obtain multiple sub-communication architectures, and a hierarchical cluster communication architecture is constructed through multiple sub-communication architectures and preset communication strategies; the computational efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture is evaluated, so as to perform network segmentation and task division operations of the reversible spiking neural network based on the computational efficiency and the hierarchical cluster communication architecture. The present application classifies heterogeneous computing units according to physical characteristics and logical characteristics, thereby reducing the coupling degree between cluster servers and improving the communication efficiency and scalability of the system. As a result, the problems that the forward connection of the existing SNN network simulation process and the reverse connection representation of the SNN learning algorithm are incompatible, and the total amount of communication between servers and the high message redundancy in the cluster SNN simulation system are solved.
[0044] Specifically, Figure 1 A flowchart of a distributed deployment and simulation method of a spiking neural network for heterogeneous clusters provided in an embodiment of the present application.
[0045] like Figure 1 As shown, the distributed deployment and simulation method of the pulse neural network for heterogeneous clusters includes the following steps:
[0046] In step S101, based on preset synapse numbers and neuron numbers, a bidirectional connection relationship between a target presynaptic neuron and a target postsynaptic neuron is determined to construct a reversible spiking neural network according to the bidirectional connection relationship.
[0047] Those skilled in the art should understand that the multi-layer pulse neural network simulation process can usually transfer data (pulse signals, etc.) from presynaptic neurons to postsynaptic neurons. At this time, only the connection relationship from the presynaptic neuron to the postsynaptic neuron is needed, and this process is called forward connection; but in some learning algorithms (such as STDP), the postsynaptic neuron needs to actively obtain the information of the presynaptic neuron (release time, etc.). At this time, it is necessary to obtain the connection relationship from the postsynaptic neuron to the presynaptic neuron, and this process is called reverse connection.
[0048] It should be noted that the embodiments of the present application can determine the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron to construct a reversible spiking neural network, thereby being compatible with the spiking neural network representation forms that require both forward connection and reverse connection representation requirements, and can support application requirements such as synapse generation that require dynamic changes in connection relationships.
[0049] Therefore, the embodiments of the present application support multiple representation requirements and provide dynamic modification capabilities by separating indexes and records based on a flexible and reversible pulse neural network representation method.
[0050] Optionally, in one embodiment of the present application, based on preset synapse numbers and neuron numbers, determining a bidirectional connection relationship between a target presynaptic neuron and a target postsynaptic neuron includes: obtaining and determining a connection requirement between the target presynaptic neuron and the target postsynaptic neuron; when the connection requirement is a forward connection requirement, determining the number of synapses and synapse numbers of all post-neuronal connecting synapses corresponding to each neuron, and constructing at least one first target array corresponding to the forward connection requirement based on the number of synapses and synapse numbers of all post-neuronal connecting synapses; when the connection requirement is a reverse connection requirement, determining the number of synapses and synapse numbers of all pre-neuronal connecting synapses corresponding to each neuron, and constructing at least one first target array corresponding to the reverse connection requirement based on the number of synapses and synapse numbers of all pre-neuronal connecting synapses; obtaining the number of neurons in the reversible spiking neural network and the address and size information of each first target array in at least one first target array corresponding to the forward connection requirement or the reverse connection requirement, and determining a second target array based on the number of neurons and the address and size information; determining a bidirectional connection relationship based on a second target array and at least one first target array corresponding to the forward connection requirement or the reverse connection requirement.
[0051] Specifically, the embodiments of the present application can represent the neural network connection relationship as two types of arrays, there is only one first type array (i.e., the second target array), the number of its elements is the same as the number of neurons, and each element in the array records the address and size of the second type array (i.e., the first target array) corresponding to the neuron.
[0052] In addition, in the embodiments of the present application, there may be one or more second-type arrays, and the number of arrays, the meaning of array elements, and the size of each array depend on the specific connection relationship. Figure 2 As shown, when representing a forward connection, each element is the number of the subsequent synaptic connection of the neuron, and its number is the number of all synapses emitted by the neuron; when representing a reverse connection, each element is the number of the input synaptic connection of the neuron, and its number is the number of all synapses input to the neuron.
[0053] It is understandable that the representation method of the embodiment of the present application can support the representation of forward connection and reverse connection, and the specific elements can select synapse number, neuron number and other code information, so as to respectively record the neuron-synapse connection relationship, neuron-neuron connection relationship, synapse-neuron connection relationship, etc. In addition, different arrays of the second type of array can be merged according to demand (the arrays are stored continuously) to achieve continuous memory access.
[0054] During the actual execution process, if the connection relationship needs to be dynamically updated, the embodiment of the present application only needs to change the actual content corresponding to the second type of array. If new storage space cannot be allocated within the second type of array, it is only necessary to create a larger array, copy the original content to the new array, and update the corresponding records within the first type of array.
[0055] Therefore, the embodiments of the present application optimize the SNN data expression method, so that it can support multiple different representation forms (forward connection and reverse connection) as needed. At the same time, it can also realize the ability to dynamically modify the network connection structure through a flexible array structure.
[0056] In step S102, the preset multiple heterogeneous computing units are divided to obtain multiple sub-communication architectures, and a hierarchical cluster communication architecture is constructed through the multiple sub-communication architectures and the preset communication strategy.
[0057] Furthermore, the embodiments of the present application can classify computing units according to physical characteristics and logical characteristics, that is, divide a variety of heterogeneous computing units to reduce the coupling between cluster servers and improve the communication efficiency and scalability of the system.
[0058] Optionally, in one embodiment of the present application, a plurality of preset heterogeneous computing units are divided to obtain a plurality of sub-communication architectures, and a hierarchical cluster communication architecture is constructed through the plurality of sub-communication architectures and preset communication strategies, including: dividing the plurality of heterogeneous computing units according to a preset physical classification strategy or a logical classification strategy to generate a plurality of sub-communication architectures, and determining the key units in each of the plurality of sub-communication architectures; when communicating within each sub-communication architecture, passing the target communication data to the target heterogeneous computing unit in the current sub-communication architecture through the heterogeneous computing unit in the current sub-communication architecture; when communicating between different sub-communication architectures, passing the target communication data to the key unit in the current sub-communication architecture in the current sub-communication architecture, and using the key unit in the current sub-communication architecture to pass the target communication data to the key unit in the target sub-communication architecture, and sending the target communication data to the target heterogeneous computing unit in the target sub-communication architecture through the key unit in the target sub-communication architecture.
[0059] It should be noted that the hierarchical cluster communication architecture (such as GPU / CPU heterogeneous cluster architecture) of the embodiment of the present application is as follows Figure 3 As shown, the architecture divides various heterogeneous computing units according to physical classification (such as connection relationship or communication delay) or logical classification (such as computing unit type, computing efficiency).
[0060] Among them, a category (corresponding to Figure 3 The dotted box in the figure can select a computing unit as the key unit. All computing units in the category can communicate directly with each other; however, the communication between computing units and computing units outside the category needs to be carried out through the key unit. When sending communication data, the communication data must first be sent to the key unit of this category. The key unit of this category can communicate directly with the key units of other categories. After receiving the data, the key units of other categories further send it to the target computing unit.
[0061] Therefore, the embodiments of the present application can reduce communication redundancy and improve message transmission efficiency by constructing a hierarchical cluster communication architecture.
[0062] In step S103, the computational efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture is evaluated to perform network segmentation and task division operations of the reversible spiking neural network based on the computational efficiency and the hierarchical cluster communication architecture.
[0063] Furthermore, the embodiments of the present application can perform pulse neural network segmentation and task division operations based on a hierarchical communication architecture.
[0064] Optionally, in one embodiment of the present application, the computational efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture is evaluated to perform network segmentation and task division operations of the reversible pulse neural network based on the computational efficiency and the hierarchical cluster communication architecture, including: determining the scale of the network segmentation of the reversible pulse neural network according to the computational efficiency, and calculating the sum of the computational efficiencies of all heterogeneous computing units in each sub-communication architecture, so as to perform a first task division operation using the sum of the computational efficiency and the scale to obtain multiple sub-networks corresponding to the reversible pulse neural network; calculating the sum of the sub-network computational efficiencies of all heterogeneous computing units within each sub-communication architecture in each of the multiple sub-networks, so as to perform a second task division on each sub-network according to the sum of the sub-network computational efficiencies.
[0065] In the specific implementation process, the embodiments of the present application can first evaluate the computational efficiency of different computing units; secondly, determine the size of the network segmentation scale based on the computational efficiency; thirdly, regard multiple computing units in the same category as the same heterogeneous unit, and use the sum of their computational efficiencies as the total computational efficiency for the first task division; after obtaining the corresponding sub-network, the embodiments of the present application can perform a second task division based on the computational efficiency of the internal computing units, thereby realizing the distributed deployment and simulation of large-scale pulse neural networks.
[0066] Therefore, the embodiments of the present application fully utilize the advantages of the hierarchical cluster communication architecture through hierarchical network segmentation and task division strategies.
[0067] According to the distributed deployment and simulation method of pulse neural networks for heterogeneous clusters proposed in the embodiment of the present application, the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron is determined based on the preset synapse number and neuron number, so as to construct a reversible pulse neural network according to the bidirectional connection relationship; the preset multiple heterogeneous computing units are divided to obtain multiple sub-communication architectures, and a hierarchical cluster communication architecture is constructed through multiple sub-communication architectures and preset communication strategies; the computing efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture is evaluated, so as to perform network segmentation and task division operations of the reversible pulse neural network based on the computing efficiency and the hierarchical cluster communication architecture. The present application reduces the coupling degree between cluster servers and improves the communication efficiency and scalability of the system by classifying heterogeneous computing units according to physical and logical characteristics.
[0068] Secondly, the distributed deployment and simulation device of the pulse neural network for heterogeneous clusters proposed in the embodiment of the present application is described with reference to the accompanying drawings.
[0069] Figure 4 It is a block diagram of a distributed deployment and simulation device of a spiking neural network for heterogeneous clusters according to an embodiment of the present application.
[0070] like Figure 4 As shown, the pulse neural network distributed deployment and simulation device 10 for heterogeneous clusters includes: a construction module 100, a partitioning module 200 and an evaluation module 300.
[0071] Among them, the construction module 100 is used to determine the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron based on the preset synapse number and neuron number, so as to construct a reversible pulse neural network according to the bidirectional connection relationship.
[0072] The division module 200 is used to divide the preset multiple heterogeneous computing units to obtain multiple sub-communication architectures, and to construct a hierarchical cluster communication architecture through the multiple sub-communication architectures and preset communication strategies.
[0073] The evaluation module 300 is used to evaluate the computational efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture, so as to perform network segmentation and task division operations of the reversible pulse neural network based on the computational efficiency and the hierarchical cluster communication architecture.
[0074] Optionally, in one embodiment of the present application, the construction module 100 includes: a judgment unit, a forward connection unit, a reverse connection unit, a first determination unit and a second determination unit.
[0075] The determination unit is used to obtain and determine the connection requirement between the target presynaptic neuron and the target postsynaptic neuron.
[0076] A forward connection unit is used to determine the number of synapses and synapse numbers of all post-connection synapses of neurons corresponding to each neuron when the connection requirement is a forward connection requirement, and to construct at least one first target array corresponding to the forward connection requirement based on the number of synapses and synapse numbers of all post-connection synapses of neurons.
[0077] A reverse connection unit is used to determine the number of synapses and synapse numbers of all neuron pre-connection synapses corresponding to each neuron when the connection requirement is a reverse connection requirement, and to construct at least one first target array corresponding to the reverse connection requirement based on the number of synapses and synapse numbers of all neuron pre-connection synapses.
[0078] The first determination unit is used to obtain the address and size information of each first target array in at least one first target array corresponding to the number of neurons in the reversible pulse neural network and the forward connection requirement or the reverse connection requirement, and determine a second target array according to the number of neurons and the address and size information.
[0079] The second determining unit is used to determine a bidirectional connection relationship based on a second target array and at least one first target array corresponding to a forward connection requirement or a reverse connection requirement.
[0080] Optionally, in one embodiment of the present application, the division module 200 includes: a generation unit, a first communication unit and a second communication unit.
[0081] Among them, the generation unit is used to divide a variety of heterogeneous computing units according to a preset physical classification strategy or logical classification strategy to generate multiple sub-communication architectures and determine the key units in each of the multiple sub-communication architectures.
[0082] The first communication unit is used to transmit target communication data to a target heterogeneous computing unit in the current sub-communication architecture through the heterogeneous computing unit in the current sub-communication architecture when communicating within each sub-communication architecture.
[0083] The second communication unit is used to transmit the target communication data to the key unit in the current sub-communication architecture when communicating between different sub-communication architectures, and to use the key unit in the current sub-communication architecture to transmit the target communication data to the key unit in the target sub-communication architecture, and to send the target communication data to the target heterogeneous computing unit in the target sub-communication architecture through the key unit in the target sub-communication architecture.
[0084] Optionally, in one embodiment of the present application, the evaluation module 300 includes: a first calculation unit and a second calculation unit.
[0085] Among them, the first computing unit is used to determine the scale of network segmentation of the reversible pulse neural network according to the computing efficiency, and calculate the sum of the computing efficiency of all heterogeneous computing units in each sub-communication architecture, so as to use the sum of the computing efficiency and the scale to perform the first task division operation to obtain multiple sub-networks corresponding to the reversible pulse neural network.
[0086] The second computing unit is used to calculate the sum of the subnetwork computing efficiencies of all heterogeneous computing units within each subcommunication architecture in each subnetwork of the multiple subnetworks, so as to perform a second task division on each subnetwork according to the sum of the subnetwork computing efficiencies.
[0087] It should be noted that the aforementioned explanation of the embodiment of the distributed deployment and simulation method of pulse neural networks for heterogeneous clusters is also applicable to the distributed deployment and simulation device of pulse neural networks for heterogeneous clusters of this embodiment, and will not be repeated here.
[0088] According to the embodiment of the present application, a distributed deployment and simulation device of a pulse neural network for heterogeneous clusters is proposed, including a construction module for determining the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron based on the preset synapse number and neuron number, so as to construct a reversible pulse neural network according to the bidirectional connection relationship; a division module for dividing the preset multiple heterogeneous computing units to obtain multiple sub-communication architectures, and constructing a hierarchical cluster communication architecture through multiple sub-communication architectures and preset communication strategies; an evaluation module for evaluating the computing efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture, so as to perform network segmentation and task division operations of the reversible pulse neural network based on the computing efficiency and the hierarchical cluster communication architecture. The present application classifies heterogeneous computing units according to physical and logical characteristics, thereby reducing the coupling between cluster servers and improving the communication efficiency and scalability of the system.
[0089] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0090] A memory 501 , a processor 502 , and a computer program stored in the memory 501 and executable on the processor 502 .
[0091] When the processor 502 executes the program, the distributed deployment and simulation method of the spiking neural network for heterogeneous clusters provided in the above embodiment is implemented.
[0092] Furthermore, the electronic device further comprises:
[0093] The communication interface 505 is used for communication between the memory 501 and the processor 502 .
[0094] The memory 501 is used to store computer programs that can be executed on the processor 502 .
[0095] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0096] If the memory 501, the processor 502 and the communication interface 505 are implemented independently, the communication interface 505, the memory 501 and the processor 502 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0097] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 505 are integrated on a chip, the memory 501, the processor 502 and the communication interface 505 can communicate with each other through an internal interface.
[0098] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0099] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned distributed deployment and simulation method of pulse neural networks for heterogeneous clusters.
[0100] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned distributed deployment and simulation method of pulse neural networks for heterogeneous clusters.
[0101] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0102] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0103] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0104] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0105] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0106] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0107] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0108] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A distributed deployment and simulation method of spiking neural networks for heterogeneous clusters, characterized in that: The following steps are involved: Based on the preset synapse number and neuron number, determining the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron, so as to construct a reversible spiking neural network according to the bidirectional connection relationship; Dividing the preset multiple heterogeneous computing units to obtain multiple sub-communication architectures, and constructing a hierarchical cluster communication architecture through the multiple sub-communication architectures and the preset communication strategy; The computational efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture is evaluated to perform network segmentation and task division operations of the reversible spiking neural network based on the computational efficiency and the hierarchical cluster communication architecture.
2. The method according to claim 1, characterized in that The determining of the bidirectional connection relationship between the target presynaptic neuron and the target postsynaptic neuron based on the preset synapse number and neuron number includes: Acquiring and determining a connection requirement between the target presynaptic neuron and the target postsynaptic neuron; When the connection requirement is a forward connection requirement, determining the number of synapses and synapse numbers of all post-neuron connected synapses corresponding to each neuron, and constructing at least one first target array corresponding to the forward connection requirement according to the number of synapses and synapse numbers of all post-neuron connected synapses; When the connection requirement is a reverse connection requirement, determining the number of synapses and synapse numbers of all neuron pre-connection synapses corresponding to each neuron, and constructing at least one first target array corresponding to the reverse connection requirement according to the number of synapses and synapse numbers of all neuron pre-connection synapses; Acquire the number of neurons in the reversible spiking neural network and the address and size information of each first target array in at least one first target array corresponding to the forward connection requirement or the reverse connection requirement, and determine a second target array according to the number of neurons and the address and size information; The bidirectional connection relationship is determined based on the one second target array and at least one first target array corresponding to the forward connection requirement or the reverse connection requirement.
3. The method according to claim 2, characterized in that The method of dividing the preset multiple heterogeneous computing units to obtain multiple sub-communication architectures, and constructing a hierarchical cluster communication architecture through the multiple sub-communication architectures and the preset communication strategy includes: Dividing the plurality of heterogeneous computing units according to a preset physical classification strategy or a logical classification strategy to generate the plurality of sub-communication architectures, and determining a key unit in each of the plurality of sub-communication architectures; When communicating within each sub-communication architecture, the target communication data is transmitted to the target heterogeneous computing unit in the current sub-communication architecture through the heterogeneous computing unit in the current sub-communication architecture; When communicating between different sub-communication architectures, the target communication data is passed to the key unit in the current sub-communication architecture in the current sub-communication architecture, and the key unit in the current sub-communication architecture is used to pass the target communication data to the key unit in the target sub-communication architecture, and the target communication data is sent to the target heterogeneous computing unit in the target sub-communication architecture through the key unit in the target sub-communication architecture.
4. The method according to claim 3, characterized in that The evaluating the computational efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture to perform network segmentation and task division operations of the reversible spiking neural network based on the computational efficiency and the hierarchical cluster communication architecture includes: Determine the scale of the network segmentation of the reversible spiking neural network according to the computational efficiency, and calculate the sum of the computational efficiencies of all heterogeneous computing units in each sub-communication architecture, so as to perform a first task division operation using the sum of the computational efficiencies and the scale, so as to obtain a plurality of sub-networks corresponding to the reversible spiking neural network; The sum of the subnetwork computing efficiencies of all heterogeneous computing units within each subcommunication architecture in each subnetwork of the multiple subnetworks is calculated, so as to perform a second task division on each subnetwork according to the sum of the subnetwork computing efficiencies.
5. A distributed deployment and simulation device for spiking neural networks in heterogeneous clusters, characterized in that: include: A construction module, used to determine a bidirectional connection relationship between a target presynaptic neuron and a target postsynaptic neuron based on a preset synapse number and a neuron number, so as to construct a reversible spiking neural network according to the bidirectional connection relationship; A partitioning module, used to partition the preset multiple heterogeneous computing units to obtain multiple sub-communication architectures, and to construct a hierarchical cluster communication architecture through the multiple sub-communication architectures and preset communication strategies; An evaluation module is used to evaluate the computing efficiency of each heterogeneous computing unit in the hierarchical cluster communication architecture, so as to perform network segmentation and task division operations of the reversible spiking neural network based on the computing efficiency and the hierarchical cluster communication architecture.
6. The device according to claim 5, characterized in that The building blocks include: A determination unit, used to obtain and determine a connection requirement between the target pre-synaptic neuron and the target post-synaptic neuron; A forward connection unit, used for determining the number of synapses and synapse numbers of all post-neuron connection synapses corresponding to each neuron when the connection requirement is a forward connection requirement, and constructing at least one first target array corresponding to the forward connection requirement according to the number of synapses and synapse numbers of all post-neuron connection synapses; A reverse connection unit, used for determining the number of synapses and synapse numbers of all neuron pre-connection synapses corresponding to each neuron when the connection requirement is a reverse connection requirement, and constructing at least one first target array corresponding to the reverse connection requirement according to the number of synapses and synapse numbers of all neuron pre-connection synapses; a first determining unit, configured to obtain the number of neurons in the reversible spiking neural network and the address and size information of each first target array in at least one first target array corresponding to the forward connection requirement or the reverse connection requirement, and determine a second target array according to the number of neurons and the address and size information; The second determining unit is configured to determine the bidirectional connection relationship based on the one second target array and at least one first target array corresponding to the forward connection requirement or the reverse connection requirement.
7. The device according to claim 6, characterized in that The division module comprises: A generating unit, configured to divide the plurality of heterogeneous computing units according to a preset physical classification strategy or a logical classification strategy to generate the plurality of sub-communication architectures, and determine a key unit in each of the plurality of sub-communication architectures; A first communication unit, configured to transmit the target communication data to a target heterogeneous computing unit in the current sub-communication architecture through a heterogeneous computing unit in the current sub-communication architecture when communicating within each sub-communication architecture; The second communication unit is used to pass the target communication data to the key unit in the current sub-communication architecture when communicating between different sub-communication architectures, and use the key unit in the current sub-communication architecture to pass the target communication data to the key unit in the target sub-communication architecture, and send the target communication data to the target heterogeneous computing unit in the target sub-communication architecture through the key unit in the target sub-communication architecture.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the distributed deployment and simulation method of a pulse neural network for heterogeneous clusters as described in any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the distributed deployment and simulation method of a pulse neural network for heterogeneous clusters as described in any one of claims 1 to 4.
10. A computer program product, comprising a computer program, characterized in that The computer program is executed to implement the distributed deployment and simulation method of pulse neural networks for heterogeneous clusters as described in any one of claims 1 to 4.