Compilation scheduling method and device for neural network, electronic equipment and storage medium
Patent Information
- Application Number
- CN202511237401.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-09-01
AI Technical Summary
然而,这些复杂模型的计算量和能耗巨大,对传统计算硬件提出了严峻挑战
[0011]根据本公开的实施例,可以通过将待部署的神经网络中的部分或全部ANN算子基于精度和硬件效率的评估自动替换为SNN算子,不仅大幅降低神经网络推理的整体能耗,而且实现了异构部署过程的自动化和智能化,降低了开发和部署的门槛。
Smart Images

Figure CN120762679B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a compilation and scheduling method, apparatus, electronic device, and computer-readable storage medium for neural networks. Background Technology
[0002] With the development of artificial intelligence technology, artificial neural network (ANN) models have achieved great success in fields such as image recognition and natural language processing. However, these complex models have enormous computational and energy consumption, posing a severe challenge to traditional computing hardware. Spiking neural networks (SNNs), due to their event-driven and sparse computation characteristics, have shown great potential in low-power scenarios. Meanwhile, to balance the high accuracy of ANNs and the high energy efficiency of SNNs, heterogeneous computing architectures have emerged, whose processors can simultaneously support the deployment and computation of both ANN and SNN neurons.
[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0004] This disclosure provides compilation and scheduling methods, apparatus, electronic devices, and storage media for neural networks.
[0005] According to one aspect of this disclosure, a compilation and scheduling method for neural networks is provided, comprising: receiving and parsing a first neural network model to be deployed on heterogeneous acceleration hardware for artificial neural networks (ANN) and spiking neural networks (SNN); identifying one or more ANN operators in the first neural network model that can be converted into SNN operators; converting the first neural network model into one or more second neural network models, wherein the one or more second neural network models are obtained by replacing at least one ANN operator in the one or more ANN operators in the first neural network model with a corresponding SNN operator; performing multi-dimensional evaluation on the first neural network model and the one or more second neural network models to obtain evaluation results, wherein the multi-dimensional evaluation includes hardware efficiency evaluation and accuracy evaluation; selecting at least one second neural network model from the one or more second neural network models as a target neural network based on the evaluation results; and deploying the target neural network on the ANN-SNN heterogeneous acceleration hardware.
[0006] According to another aspect of this disclosure, a compilation scheduling apparatus for neural networks is provided, comprising: a first unit configured to receive and parse a first neural network model to be deployed on heterogeneous acceleration hardware for artificial neural networks (ANN) and spiking neural networks (SNN); a second unit configured to identify one or more ANN operators in the first neural network model that can be converted into SNN operators; a third unit configured to convert the first neural network model into one or more second neural network models, wherein the one or more second neural network models are obtained by replacing at least one ANN operator in the one or more ANN operators in the first neural network model with a corresponding SNN operator; a fourth unit configured to perform multi-dimensional evaluation on the first neural network model and the one or more second neural network models to obtain an evaluation result, wherein the multi-dimensional evaluation includes hardware efficiency evaluation and accuracy evaluation; a fifth unit configured to select at least one second neural network model as a target neural network based on the evaluation result; and a sixth unit configured to deploy the target neural network on the ANN-SNN heterogeneous acceleration hardware.
[0007] According to another aspect of this disclosure, an electronic circuit is provided, comprising: a circuit configured to perform the steps of the above-described method.
[0008] According to another aspect of this disclosure, an electronic device is provided. The electronic device includes: a processor; and a memory storing a program including instructions that, when executed by the processor, cause the processor to perform the methods described above.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing a program is provided. The program includes instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the methods described above.
[0010] According to another aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the above-described method.
[0011] According to embodiments of this disclosure, by automatically replacing some or all ANN operators in a neural network to be deployed with SNN operators based on accuracy and hardware efficiency evaluations, not only can the overall energy consumption of neural network inference be significantly reduced, but the heterogeneous deployment process can also be automated and intelligentized, lowering the threshold for development and deployment.
[0012] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A flowchart illustrating an exemplary process of a compilation scheduling method for ANN and SNN according to embodiments of the present disclosure is shown; Figure 2 An exemplary block diagram of a compilation scheduling apparatus for ANN and SNN according to embodiments of the present disclosure is shown; Figure 3 The diagram shows a detailed flowchart of an exemplary process for a compilation scheduling method for ANN and SNN according to embodiments of the present disclosure; Figure 4 An exemplary block diagram of a compiler according to embodiments of the present disclosure is shown; Figures 5A-5C Examples of operator substitution processes in compilation scheduling methods for ANNs and SNNs according to embodiments of this disclosure are shown; and Figure 6 This is a block diagram illustrating an example of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0015] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0016] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0017] In related technologies, directly deploying traditional ANN models to SNNs often leads to a loss of accuracy, while simple operator replacement and manual optimization are difficult to adapt to diverse network structures and complex performance metrics (e.g., multiple optimization objectives such as hardware efficiency and accuracy).
[0018] To address the aforementioned problems in related technologies, this disclosure provides a compilation and scheduling method for neural networks. This method can automatically replace some or all ANN operators in a neural network to be deployed with SNN operators based on an evaluation of accuracy and hardware efficiency. This not only significantly reduces the overall energy consumption of neural network inference and better balances accuracy and hardware efficiency (e.g., avoiding blindly sacrificing accuracy for low power consumption or excessively retaining accuracy while wasting energy), but also eliminates the limitations of manual tuning and experience-based judgment in traditional ANN-to-SNN conversion. It automates and intelligently implements the heterogeneous deployment process, lowering the development and deployment threshold. Embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0019] Figure 1 A flowchart illustrating an exemplary process of a compilation scheduling method 100 for ANN and SNN according to embodiments of the present disclosure is shown.
[0020] In step S102, the first neural network model to be deployed on the ANN-SNN heterogeneous acceleration hardware can be received and parsed.
[0021] In step S104, one or more ANN operators in the first neural network model that can be converted into SNN operators can be identified.
[0022] In step S106, the first neural network model can be transformed into one or more second neural network models, wherein the one or more second neural network models are obtained by replacing at least one ANN operator in one or more ANN operators in the first neural network model with the corresponding SNN operator.
[0023] In step S108, a multi-dimensional evaluation can be performed on the first neural network model and one or more second neural network models to obtain evaluation results. The multi-dimensional evaluation includes hardware efficiency evaluation and accuracy evaluation.
[0024] In step S110, at least one of the one or more second neural network models can be selected as the target neural network based on the evaluation results.
[0025] In step S112, the target neural network can be deployed on ANN-SNN heterogeneous acceleration hardware.
[0026] The compilation and scheduling method for ANN and SNN provided by the embodiments of this disclosure can automatically replace some or all ANN operators in the neural network to be deployed with SNN operators based on the evaluation of accuracy and hardware efficiency. This not only significantly reduces the overall energy consumption of neural network inference, but also better balances accuracy and hardware efficiency, avoiding blindly sacrificing accuracy to pursue low power consumption or excessively retaining accuracy and wasting energy. Furthermore, it eliminates the limitations of manual tuning and experience-based judgment in the traditional ANN to SNN conversion, realizing the automation and intelligence of the heterogeneous deployment process and lowering the threshold for development and deployment.
[0027] The steps of method 100 are described in detail below.
[0028] In step S102, the first neural network model to be deployed on the ANN-SNN heterogeneous acceleration hardware can be received and parsed.
[0029] In some examples, the first neural network model can consist of an ANN or a combination of an ANN and a SNN. In these examples, the processor in the ANN-SNN heterogeneous acceleration hardware (or ANN-SNN heterogeneous chip) can switch between ANN and SNN modes, enabling simultaneous deployment and computation of both ANN and SNN neurons.
[0030] In some examples, the first neural network model can be a pre-trained network model, such as PyTorch / ONNX, and the front-end format is not limited here.
[0031] In some examples, parsing the first neural network model can resolve the first neural network model into an intermediate representation, enabling the machine to identify the operators in the neural network. Furthermore, it can perform static analysis on each layer of the first neural network model to extract structural information of the first neural network model (e.g., the input / output dimensions, operation types, and parameter features of each layer).
[0032] In step S104, one or more ANN operators in the first neural network model that can be converted into SNN operators can be identified.
[0033] refer to Figure 5A The figure shows an ANN operator that can be converted into an SNN operator in the first neural network model shown in the figure. This ANN operator is obtained by merging Conv and ReLU.
[0034] In some embodiments, the ANN operator that can be converted into an SNN operator can be a single operator such as Conv, MLP, transformer, etc., or a fusion operator composed of multiple single operators, such as... Figure 5BAs shown, the ANN operator and its corresponding SNN operator are both fusion operators. The ANN operator is obtained by merging Conv and ReLU, and the corresponding SNN operator is obtained by merging Encoder, SNN Conv and Decoder.
[0035] In step S106, the first neural network model can be transformed into one or more second neural network models, wherein the one or more second neural network models are obtained by replacing at least one ANN operator in one or more ANN operators in the first neural network model with the corresponding SNN operator.
[0036] refer to Figure 5C ,Will Figure 5B In the first neural network model shown in the diagram, the ANN operators that can be converted into SNN operators are replaced with the corresponding SNN operators, thereby obtaining... Figure 5C The example in the second neural network model.
[0037] It should be understood that, Figures 5A-5C The operator replacement process in this example is only for illustration, and the type and number of operators to be replaced are not limited here.
[0038] In step S108, a multi-dimensional evaluation can be performed on the first neural network model and one or more second neural network models to obtain evaluation results. The multi-dimensional evaluation includes hardware efficiency evaluation and accuracy evaluation.
[0039] In some examples, hardware efficiency evaluation can be performed by a trained hardware efficiency evaluation model, and accuracy evaluation can be performed by a trained accuracy evaluation model. Specifically, the hardware efficiency evaluation model can model heterogeneous ANN / SNN hardware, and obtain the energy consumption, latency, and other indicators of ANN cores and SNN cores under different precision and different input data sparsity by running performance simulations at compile time. Meanwhile, the accuracy analysis model can implement a computation library based on different SNN encoding schemes to collect the accuracy loss of each layer of the neural network in ANN / SNN mode.
[0040] In some examples, the evaluation result can be obtained by weighting the results of the hardware efficiency evaluation and the accuracy evaluation, where the specific weights can be set manually according to the actual situation. Alternatively, the evaluation result can also be obtained by fusing the results of the hardware efficiency evaluation and the accuracy evaluation in other ways, such as exponential summation, reciprocal summation, etc., which are not limited here.
[0041] In some embodiments, hardware efficiency assessment may include one or more of energy consumption assessment, throughput assessment, bandwidth assessment, and latency assessment.
[0042] In some examples, energy consumption prediction models for each type of neural network layer in both ANN and SNN modes can be built based on hardware characteristics for energy consumption assessment. Specifically, the energy consumption of the SNN mode can take into account factors such as the sparsity of the input data, the pulse firing frequency, and the time step.
[0043] In some examples, latency assessment can be performed by building computational latency prediction models for each type of neural network layer in both ANN and SNN modes, based on hardware characteristics.
[0044] Therefore, by conducting hardware efficiency and accuracy assessments before deployment, we can avoid blindly sacrificing accuracy in pursuit of low power consumption and low latency, or excessively retaining accuracy and wasting energy.
[0045] In some embodiments, method 100 may further include obtaining input data for each layer of a first neural network, and the energy consumption assessment includes performing an energy consumption assessment for each layer of the neural network based on the sparsity of the obtained input data for each layer of the first neural network.
[0046] In some examples, the input data for each layer of the first neural network can be reference samples obtained from an external source (such as deep learning frameworks like PyTorch). This type of data has the same distribution as the original training dataset of the first neural network and can also be called a "calibration dataset".
[0047] Since the evaluation metrics of neural networks after ANN operators are converted into SNN operators are directly related to the distribution of the input data, for example, in energy consumption evaluation, the input distribution directly affects the number of pulses fired after SNN encoding, thus affecting the energy consumption evaluation result; in accuracy evaluation, the input distribution affects the activation range, which in turn affects the determination of scaling and discretization parameters, thereby affecting the final output accuracy. Therefore, introducing a calibration dataset with the same distribution as the training dataset can improve the accuracy of the evaluation results.
[0048] In some examples, the first neural network model can be parsed to obtain its structural information, such as input / output dimensions, operation types, and parameter features. Based on the obtained structural information of the first neural network model, the input data of each layer of the first neural network can be obtained. Based on the sparsity of the input data of each layer of the first neural network (e.g., the proportion of zero values in the activation values), the energy consumption of each layer of the first neural network model and one or more second neural network models can be evaluated. For example, if the proportion of zero values in the input data of a layer is larger, it means that the input data is more sparse. In this case, the layer can be converted to an SNN mode to save more energy.
[0049] Therefore, layers suitable for sparse computing are automatically deployed to the SNN mode, significantly reducing the overall energy consumption of neural network inference.
[0050] In some examples, besides using the structural information of the obtained first neural network model, the input data for each layer of the first neural network can also be obtained by analyzing representative input data of each layer. For example, the sparsity of the representative input data of each layer of the first neural network can be analyzed to generate data with consistent or approximate sparsity as the input data for each layer of the first neural network. The above methods of obtaining input data are merely examples and are not intended to be limiting.
[0051] In some embodiments, accuracy assessment may include truncation error assessment and discretization error assessment.
[0052] Clipping error refers to the activation value truncation error that may occur during the process of simulating ANN activation values to SNN pulse encoding. Discretization error refers to the error range between the output of the continuous activation function of ANN and the pulse output of SNN under a specified input distribution. Specifically, error models can be established for different encoding methods from continuous activation functions of ANN (such as ReLU) to discrete pulses of SNN, and statistical methods can be used to analyze the error range between the output of the continuous activation function and the pulse output under a specified input distribution. These errors can lead to a decrease in the accuracy of neural networks. For example, the more ANN operators are converted to SNN operators, the greater the decrease in the accuracy of the neural network tends to be.
[0053] It is important to note that the error in this paper is not the error between the neural network output and the true value provided by the dataset during the neural network training process, but rather the error caused by the conversion process of the neural network input and output from simulated activation values to impulses after the ANN operator is converted to the SNN operator during the neural network deployment process (at which point the neural network has been trained).
[0054] In step S110, at least one of the one or more second neural network models can be selected as the target neural network based on the evaluation results.
[0055] In some examples, a target neural network can be selected based on the evaluation results, exhibiting minimal accuracy loss and high hardware efficiency (low latency, low power consumption). Alternatively, the target neural network can be automatically selected based on preset conditions (e.g., achieving a certain level of accuracy and power consumption below a certain threshold), or it can be manually selected. In some examples, the target neural network can also be one or more neural networks that meet the preset conditions; this is not a limitation.
[0056] In step S112, the target neural network can be deployed on ANN-SNN heterogeneous acceleration hardware.
[0057] Specifically, based on the selected target neural network, the operators in the matching operator library can be mapped for each layer of the target neural network, and a heterogeneous execution program can be automatically generated for deployment on ANN-SNN heterogeneous acceleration hardware. This automates and automates the heterogeneous deployment process, lowering the development and deployment threshold.
[0058] In some embodiments, method 100 further includes inserting quantization layers, dequantization layers, encoding layers, and decoding layers into the target neural network before deploying the target neural network onto ANN-SNN heterogeneous acceleration hardware.
[0059] In some examples, the data for the first neural network is typically stored using floating-point types, and the deployment of the SNN operator requires discretizing this data. Inserting quantization / dequantization layers can simulate quantization errors when calculating the transformed SNN operator parameters, improving the consistency of the model's expressive power before and after transformation. It can also connect ANN subgraphs and SNN subgraphs with different precisions.
[0060] In some examples, the quantization layer can map floating-point data (e.g., FP32) to an integer space (e.g., INT4). Specifically, the quantization layer can determine the range of the input data, map the floating-point data to the representation range of an integer type by calculating a scaling factor and zeros, and then obtain discrete data through rounding.
[0061] In some examples, the dequantization layer can map integer data (e.g., INT8) back to floating-point space (e.g., FP32). Specifically, the dequantization layer can use the scaling factor and zeros determined during quantization to subtract the zeros from the integer data and then multiply it by the scaling factor, thereby restoring an approximation of a floating-point number, ensuring numerical continuity and computational precision requirements.
[0062] In some examples, the quantization techniques used in the quantization layer may be quantization-aware training (QAT) or post-training quantization (PTQ), etc., without limitation.
[0063] According to embodiments of this disclosure, a compilation scheduling apparatus for ANN and SNN is also provided. Figure 2An exemplary block diagram of a compilation scheduling apparatus for ANN and SNN according to embodiments of the present disclosure is shown. The compilation scheduling apparatus 200 for ANN and SNN may include: a first unit 210 configured to receive and parse a first neural network model to be deployed on ANN-SNN heterogeneous acceleration hardware; a second unit 220 configured to identify one or more ANN operators in the first neural network model that can be converted into SNN operators; a third unit 230 configured to convert the first neural network model into one or more second neural network models, wherein the one or more second neural network models are obtained by replacing at least one ANN operator in the one or more ANN operators in the first neural network model with a corresponding SNN operator; a fourth unit 240 configured to perform a multi-dimensional evaluation on the first neural network model and the one or more second neural network models to obtain an evaluation result, wherein the multi-dimensional evaluation includes hardware efficiency evaluation and accuracy evaluation; a fifth unit 250 configured to select at least one second neural network model as a target neural network based on the evaluation result; and a sixth unit 260 configured to deploy the target neural network onto the ANN-SNN heterogeneous acceleration hardware.
[0064] Here, the operation of each unit of the device is similar to the operation of steps S102 to S112 described above, and will not be repeated here.
[0065] According to another aspect of this disclosure, an electronic circuit is also provided, including a circuit configured to perform the steps of the above-described method.
[0066] According to another aspect of this disclosure, an electronic device is also provided, comprising: a processor; and a memory storing a program, the program including instructions that, when executed by the processor, cause the processor to perform the methods described above.
[0067] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing a program is also provided, the program including instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the method described above.
[0068] According to another aspect of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described method.
[0069] Figure 3 The diagram shows a detailed flowchart of an exemplary process 300 of a compilation scheduling method 100 for ANN and SNN according to embodiments of the present disclosure.
[0070] The specific flow of an exemplary process 300 for a compilation scheduling method 100 for ANN and SNN according to embodiments of the present disclosure may include the following steps: S301, Network Input: Input a pre-trained network model (PyTorch / ONNX, etc., the form of the compiled front-end is not limited here), and parse it into a unified intermediate representation; S302, SNN operator identification, which identifies ANN operators that can be converted into SNN operators; S303, ANN network diagram compilation; S304. Hardware efficiency evaluation and accuracy evaluation are performed using hardware efficiency evaluation model and accuracy evaluation model. The hardware efficiency evaluation model can model heterogeneous hardware of ANN / SNN and obtain the energy consumption, latency and other indicators of ANN core and SNN core under different accuracies and sparsities by running performance simulation at compile time. The accuracy analysis model can implement the computing library based on different encoding schemes of SNN and collect the accuracy loss of each layer in ANN / SNN mode. S305. Obtain the evaluation results of the hardware efficiency evaluation model and the accuracy evaluation model; S306 Deployment Decision: Based on the evaluation results, some network layers are marked as SNN mode and converted in the subsequent graph compilation process. The marking may include information such as execution mode (ANN / SNN) and SNN encoding method. S307, SNN Operator Transformation: Based on the deployment decision of the SNN multi-dimensional evaluation model output, some layers in the ANN network are transformed into forms adapted to SNN operators, and necessary quantization / dequantization layers and encoding / decoding layers are inserted; and S308, Operator Mapping: The compiler matches operators in the operator library with the final mode selection result (e.g., whether each layer uses ANN mode or SNN mode) for mapping, and generates the final heterogeneous executable program for deployment on hardware. The operator library can be implemented as an independent software and can be extended independently of the compilation framework to improve performance and architecture compatibility, avoid directly generating instructions, and thus improve the generalization of the scheduling model.
[0071] Figure 4 An exemplary block diagram of a compiler 400 according to an embodiment of the present disclosure is shown.
[0072] Compiler 400 is a hybrid architecture compiler that supports the compilation of both ANNs and SNNs. Compiler 400 includes a compilation front-end 400, an SNN graph compiler 420, an ANN graph compiler 430, an SNN operator transformation 440, a quantization / encoding / decoding layer insertion 450, an SNN / ANN operator library 460, and an operator mapping 470. The SNN / ANN operator library 460 can be an existing operator library used for operator mapping 470. The SNN / ANN operator library 460 can be implemented as independent software and can be extended independently of the compilation framework to improve performance and architectural compatibility, avoid directly generating instructions, and thus improve the generalization of the scheduling model. The SNN graph compiler 420 may include subgraph recognition 421 (or layer / operator recognition), a hardware efficiency evaluation module 422, a precision evaluation module 423, a multi-dimensional weighted evaluation 424, and a deployment decision 425.
[0073] It should be understood that, Figure 3 The specific process and Figure 4 The compilers shown are for illustrative purposes only, not for limiting purposes. For example, the multi-dimensional weighted evaluation 424 is only for illustrative purposes, and the dimensions can also be evaluated in a comprehensive manner other than weighting.
[0074] See Figure 6 Electronic device 600 will now be described as an example of a hardware device (electronic device) that can be applied to various aspects of this disclosure. Electronic device 600 can be any machine configured to perform processing and / or computation, and can be, but is not limited to, a workstation, server, desktop computer, laptop computer, tablet computer, personal digital assistant, robot, smartphone, in-vehicle computer, or any combination thereof. The compilation and scheduling method 100 described above for ANN and SNN can be implemented wholly or at least partially by electronic device 600 or similar devices or systems.
[0075] Electronic device 600 may include elements that connect to or communicate with bus 602 (possibly via one or more interfaces). For example, electronic device 600 may include bus 602, one or more processors 604, one or more input devices 606, and one or more output devices 608. The one or more processors 604 may be any type of processor and may include, but are not limited to, one or more general-purpose processors and / or one or more dedicated processors (e.g., special-purpose chips). Input devices 606 may be any type of device capable of inputting information to electronic device 600 and may include, but are not limited to, a mouse, keyboard, touchscreen, microphone, and / or remote control. Output devices 608 may be any type of device capable of presenting information and may include, but are not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Electronic device 600 may also include a non-transitory storage device 610. The non-transitory storage device can be any storage device that is non-transitory and capable of storing data, including but not limited to disk drives, optical storage devices, solid-state storage, floppy disks, flexible disks, hard disks, magnetic tapes or any other magnetic media, optical discs or any other optical media, ROM (read-only memory), RAM (random access memory), cache memory and / or any other memory chip or cartridge, and / or any other medium from which a computer can read data, instructions, and / or code. The non-transitory storage device 610 can be detached from an interface. The non-transitory storage device 610 may have data / programs (including instructions) / code for implementing the methods and steps described above. Electronic device 600 may also include a communication device 612. The communication device 612 can be any type of device or system enabling communication with external devices and / or with a network, and may include, but is not limited to, modems, network interface cards, infrared communication devices, wireless communication devices and / or chipsets, such as Bluetooth™ devices, 802.11 devices, Wi-Fi devices, Wi-Max devices, cellular communication devices, and / or the like.
[0076] Electronic device 600 may also include working memory 614, which may be any type of working memory that can store programs (including instructions) and / or data useful for the operation of processor 604, and may include, but is not limited to, random access memory and / or read-only memory devices.
[0077] Software elements (programs) may reside in working memory 614, including but not limited to operating system 616, one or more application programs 618, drivers, and / or other data and code. Instructions for performing the above methods and steps may be included in one or more application programs 618, and the above-described compilation scheduling method 100 for ANN and SNN can be implemented by processor 604 reading and executing the instructions of one or more application programs 618. More specifically, in the above-described compilation scheduling method 100 for ANN and SNN, steps S102-S112 can be implemented, for example, by processor 604 executing application programs 618 having instructions for steps S102-S112. Furthermore, other steps in the above-described compilation scheduling method 600 for ANN and SNN can be implemented, for example, by processor 604 executing application programs 618 having instructions for executing the corresponding steps. The executable code or source code of the instructions of the software elements (programs) may be stored in a non-transitory computer-readable storage medium (e.g., the storage device 610 described above) and may be stored in working memory 614 during execution (possibly for compilation and / or installation). The executable code or source code of the instructions for software elements (programs) can also be downloaded from a remote location.
[0078] It should also be understood that various modifications can be made depending on specific requirements. For example, custom hardware can also be used, and / or specific elements can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. For example, some or all of the disclosed methods and apparatus can be implemented by programming hardware (e.g., programmable logic circuits including field-programmable gate arrays (FPGAs) and / or programmable logic arrays (PLAs)) using logic and algorithms according to this disclosure in assembly language or hardware programming languages (such as Verilog, VHDL, C++).
[0079] It should also be understood that the aforementioned methods can be implemented using a server-client model. For example, the client can receive user input data and send it to the server. Alternatively, the client can receive user input data, perform a portion of the processing described in the aforementioned methods, and send the resulting data to the server. The server can receive data from the client, execute the aforementioned methods or a portion thereof, and return the execution result to the client. The client can receive the execution result from the server and, for example, present it to the user via an output device.
[0080] It should also be understood that the components of electronic device 600 can be distributed across a network. For example, some processing can be performed using one processor, while other processing can be performed simultaneously by another processor located far away from that processor. Other components of computing system 600 can also be distributed similarly. Thus, electronic device 600 can be interpreted as a distributed computing system that performs processing in multiple locations.
[0081] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A compilation scheduling method for neural networks, comprising: Receive and parse the first neural network model to be deployed on heterogeneous acceleration hardware of artificial neural network (ANN) and spiking neural network (SNN); Identify one or more ANN operators in the first neural network model that can be converted into SNN operators; The first neural network model is transformed into one or more second neural network models, wherein the one or more second neural network models are obtained by replacing at least one ANN operator in the one or more ANN operators in the first neural network model with the corresponding SNN operator; The first neural network model and the one or more second neural network models are evaluated in multiple dimensions to obtain evaluation results. The multi-dimensional evaluation includes hardware efficiency evaluation and accuracy evaluation. The hardware efficiency evaluation includes energy consumption evaluation, throughput evaluation, bandwidth evaluation and latency evaluation. The energy consumption evaluation includes constructing an energy consumption prediction model for each neural network layer type in ANN mode and SNN mode based on hardware characteristics for the energy consumption evaluation. The latency evaluation includes constructing a computational latency prediction model for each neural network layer type in ANN mode and SNN mode based on hardware characteristics for the latency evaluation. Based on the evaluation results, at least one of the one or more second neural network models is selected as the target neural network; and The target neural network is deployed on the ANN-SNN heterogeneous acceleration hardware.
2. The method as described in claim 1, wherein, The one or more ANN operators that can be converted to SNN form include single ANN operators that can be converted to SNN form and / or fused ANN operators that can be converted to SNN form, wherein the fused ANN operator is obtained by merging multiple single ANN operators.
3. The method of claim 1, further comprising obtaining input data for each layer of the first neural network, and the energy consumption assessment comprising: Energy consumption is evaluated for each layer of the first neural network model and the one or more second neural network models based on the sparsity of the input data of each layer of the first neural network.
4. The method of claim 1, wherein, The accuracy assessment includes truncation error assessment and discretization error assessment. The truncation error is the activation value truncation error generated during the process of converting simulated ANN activation values to SNN pulse encoding, and the discretization error is the error between the ANN continuous activation function output and the SNN pulse output.
5. The method of claim 1, further comprising inserting a quantization layer, a dequantization layer, an encoding layer, and a decoding layer into the target neural network before deploying the target neural network onto the ANN-SNN heterogeneous acceleration hardware.
6. A compiler scheduling apparatus for neural networks, comprising: The first unit is configured to receive and parse the first neural network model to be deployed on the heterogeneous acceleration hardware of the artificial neural network ANN-spiking neural network SNN. The second unit is configured to identify one or more ANN operators in the first neural network model that can be converted into SNN operators; The third unit is configured to transform the first neural network model into one or more second neural network models, wherein the one or more second neural network models are obtained by replacing at least one ANN operator in the one or more ANN operators in the first neural network model with a corresponding SNN operator. The fourth unit is configured to perform multi-dimensional evaluations on the first neural network model and the one or more second neural network models to obtain evaluation results. The multi-dimensional evaluation includes hardware efficiency evaluation and accuracy evaluation. The hardware efficiency evaluation includes energy consumption evaluation, throughput evaluation, bandwidth evaluation, and latency evaluation. The energy consumption evaluation includes constructing an energy consumption prediction model for each neural network layer type in ANN mode and SNN mode based on hardware characteristics for the energy consumption evaluation. The latency evaluation includes constructing a computational latency prediction model for each neural network layer type in ANN mode and SNN mode based on hardware characteristics for the latency evaluation. The fifth unit is configured to select at least one of the one or more second neural network models as the target neural network based on the evaluation results; and The sixth unit is configured to deploy the target neural network onto the ANN-SNN heterogeneous acceleration hardware.
7. An electronic circuit, comprising: A circuit configured to perform the steps of the method according to any one of claims 1 to 5.
8. An electronic device, comprising: processor; as well as A memory storing a program, the program comprising instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium storing a program, the program comprising instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 5.
10. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Impulse neural network construction method
CN119990197A
Low-power ai processing system and method combining artificial neural network and spiking neural network
US20240320472A1