Neuromorphic hardware acceleration method, device and module and readable medium
By obtaining the maximum and second-largest difference of the result vector in real time in the pulse neural network, the problems of power consumption and computational complexity in traditional methods are solved, more efficient hardware acceleration is achieved, recognition time is shortened and accuracy is maintained.
Patent Information
- Application Number
- CN202510468151.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional neuromorphic hardware acceleration methods have excessive power consumption and computational complexity in spiking neural networks, especially in the time dimension, where they fail to effectively reduce time steps, resulting in repeated operations and excessive computation.
By obtaining the current time step result vector of the spiking neural network in real time, extracting the maximum and second largest values, and calculating their difference and comparing it with the credibility threshold, if the difference is greater than the threshold, the calculation of subsequent time steps is terminated and the current result is output as the final inference result.
It simplifies the calculation process, saves invalid time and space resources, brings significant computing power and energy efficiency benefits, shortens the recognition process time, and maintains recognition accuracy.
Smart Images

Figure CN120597954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a neuromorphic hardware acceleration method, device, module, and readable medium. Background Art
[0002] The computational process of neural network models requires hardware support. Because neural network training and inference require extensive computation and storage, traditional central processing units (CPUs) often cannot meet the high-performance and low-latency requirements, while large GPUs cannot meet the power consumption and latency constraints of the edge. Therefore, to overcome this limitation, deploying neural network models on mobile devices or edge devices requires hardware accelerators specifically designed for neural network computation. Compared to general-purpose processors, neural network accelerators can accelerate core operations such as convolution, pooling, and nonlinear activation through model-hardware co-optimization, thereby providing faster computation speed and lower power consumption. In the design of pulse-based neural network accelerators, the main design and optimization targets are multiplication-accumulation units and some nonlinear functions. For processing time-dimensional information, the same components are generally reused across different time steps, which can reduce the duplication of hardware resources to a certain extent. However, in actual inference, some tasks do not require the same number of time steps as during training and can be completed in a very short time. This supports the idea of reducing the time step of network operation to save time and energy. Traditional pulse-based neural network algorithms and their corresponding accelerators lack the ability to process time-dimensional information or reduce time steps. This repeated computation has no impact on operational efficiency. Furthermore, the computational output from the first few time steps is wasted and not effectively utilized, creating unnecessary overhead.
[0003] Related technologies achieve hardware acceleration by applying an attention mechanism to the temporal dimension to identify the more important input time steps. However, during operation, two fully connected networks are required to extract information from each time step, which results in a very large amount of computation even for a small number of time steps. A sigmoid function is then used for scoring and evaluation. However, the nonlinear function used here has no good hardware implementation, and a lookup table (LUT) is typically used for approximate calculations. Both of these methods result in significant area and power consumption overhead. Summary of the Invention
[0004] The present invention provides a neuromorphic hardware acceleration method, device, module and readable medium to address the defects of traditional neuromorphic hardware acceleration methods, such as excessive power consumption and complex calculations.
[0005] The present invention provides a neuromorphic hardware acceleration method, comprising:
[0006] During the operation of the spiking neural network, the result vector of the current time step is obtained in real time;
[0007] Extracting the maximum value and the second maximum value from the result vector;
[0008] Calculate the difference between the maximum value and the second largest value, and compare it with the credibility threshold;
[0009] If the difference is greater than the credibility threshold, the calculation of the subsequent time step is terminated, and the result of the current time step is output as the final inference result.
[0010] According to the neuromorphic hardware acceleration method provided by the present invention, the confidence threshold is calculated based on the confidence requirement reaching the threshold and the number of dimensions of the result vector.
[0011] According to the neuromorphic hardware acceleration method provided by the present invention, the credibility threshold calculation method is:
[0012] Confidence threshold
[0013] Where x is the confidence threshold and n is the number of dimensions of the result vector.
[0014] According to the neuromorphic hardware acceleration method provided by the present invention, the credibility threshold is dynamically configured according to the requirements of the model.
[0015] According to the neuromorphic hardware acceleration method provided by the present invention, extracting the maximum value and the second largest value from the result vector includes:
[0016] If the result vectors arrive in sequence, the received result vectors are compared item by item, and the maximum value and the second largest value are dynamically updated.
[0017] According to the neuromorphic hardware acceleration method provided by the present invention, extracting the maximum value and the second largest value from the result vector includes:
[0018] If the result vectors arrive at the same time, the maximum value and the second largest value are extracted from the result vectors in parallel.
[0019] The present invention also provides a neuromorphic hardware acceleration device, comprising:
[0020] The acquisition module is used to obtain the result vector of the current time step in real time during the operation of the spiking neural network;
[0021] An extraction module, configured to extract a maximum value and a second maximum value from the result vector;
[0022] a comparison module, configured to calculate a difference between the maximum value and the second largest value, and compare the difference with a credibility threshold;
[0023] The output module is used to terminate the calculation of subsequent time steps if the difference is greater than the credibility threshold, and output the result of the current time step as the final inference result.
[0024] The present invention further provides a serial arbitration module, applicable to the above-mentioned neuromorphic hardware acceleration method, comprising:
[0025] Multiple registers, respectively used to store the maximum value and the second maximum value of the currently received results, as well as the pre-stored confidence threshold;
[0026] The first comparator is used to compare a classification result input each time with the current maximum value and the second maximum value to update the maximum value and the second maximum value;
[0027] A subtractor is used to calculate the difference between the maximum value and the second maximum value after obtaining the final maximum value and the second maximum value;
[0028] The second comparator is used to compare the subtractor result with the credibility threshold to determine whether the model operation can be terminated early.
[0029] The present invention further provides a parallel arbitration module applicable to the above-mentioned neuromorphic hardware acceleration method, comprising:
[0030] Multiple process comparators are used to group multiple input result vectors into pairs and distinguish the size values within the groups;
[0031] A plurality of compressors, the plurality of compressors being connected in cascade and configured to find a maximum value and a minimum value in a plurality of input groups;
[0032] The result comparator is used to subtract the output value of the last stage compressor and compare the result with the pre-stored credibility threshold to determine whether the model operation can be interrupted.
[0033] The present invention also provides a non-transitory computer-readable storage medium, which stores computer-readable instructions. The computer-readable instructions can be executed by a computer having one or more processors to enable the processor to perform the neuromorphic hardware acceleration method as described in any one of the above items.
[0034] The neuromorphic hardware acceleration method, device, module and readable medium provided by the present invention obtain the result vector of the current time step in real time during the operation of the spiking neural network; extract the maximum value and the second largest value from the result vector; calculate the difference between the maximum value and the second largest value and compare it with the credibility threshold; if the difference is greater than the credibility threshold, terminate the calculation of the subsequent time step and output the result of the current time step as the final inference result. The present invention simplifies the confidence judgment to only compare the difference between the maximum value and the second largest value in the result vector, avoiding exponential operations and summation, and directly terminates the calculation of the subsequent time step based on the judgment result, saving invalid time and space resources and bringing considerable computing power and energy efficiency benefits. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0036] Figure 1 is a flowchart of a neuromorphic hardware acceleration method provided by an embodiment of the present invention;
[0037] Figure 2 This is a schematic diagram of the principle of the traditional non-reduced time step technology;
[0038] Figure 3 is a functional structural diagram of a neuromorphic hardware acceleration device provided by an embodiment of the present invention;
[0039] Figure 4 1 is a functional structure diagram of a serial arbitration module provided by an embodiment of the present invention;
[0040] Figure 5 2 is a functional structure diagram of a parallel arbitration module provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0041] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0042] Figure 1 A flowchart of a neuromorphic hardware acceleration method provided by an embodiment of the present invention is shown in FIG. Figure 1As shown, the neuromorphic hardware acceleration method provided by the embodiment of the present invention includes:
[0043] Step 101: During the operation of the spiking neural network, the result vector of the current time step is obtained in real time;
[0044] Step 102: extract the maximum value and the second maximum value from the result vector;
[0045] Step 103: Calculate the difference between the maximum value and the second maximum value, and compare it with the credibility threshold;
[0046] Step 104: If the difference is greater than the credibility threshold, the calculation of the subsequent time step is terminated, and the result of the current time step is output as the final inference result.
[0047] Traditional neuromorphic hardware acceleration methods such as Figure 2 As shown, this is accomplished by simply reusing the same components at different time steps. Some proposed acceleration methods employ attention mechanisms across the temporal dimension to identify more important input time steps for hardware acceleration. However, in operation, two fully connected networks are required to extract information from each time step, which results in a very large computational load even for a small number of time steps. Scoring and evaluation are then performed using a sigmoid function. The nonlinear function used here has no good hardware implementation, and a lookup table (LUT) is typically used for approximate calculations. Both of these factors result in significant area and power consumption overhead.
[0048] The neuromorphic hardware acceleration method provided by an embodiment of the present invention obtains the result vector of the current time step in real time during the operation of the spiking neural network; extracts the maximum value and the second largest value from the result vector; calculates the difference between the maximum value and the second largest value and compares it with a credibility threshold; if the difference is greater than the credibility threshold, terminates the calculation of subsequent time steps and outputs the result of the current time step as the final inference result. The embodiment of the present invention simplifies the confidence judgment to only compare the difference between the maximum value and the second largest value in the result vector, avoiding exponential operations and summations, and directly terminates the calculation of subsequent time steps based on the judgment result, saving invalid time and space resources and bringing considerable computing power and energy efficiency benefits.
[0049] Based on any of the above embodiments, the credibility threshold is calculated based on the confidence requirement reaching the threshold and the number of dimensions of the result vector.
[0050] The early retirement mechanism of the spiking neural network is mainly to find the maximum value of the result vector after the calculation of each time step is completed by softmax operation. This value represents the reasoning result of the spiking neural network model at the current time step to a certain extent. Evaluating the credibility of this result can help judge the reasoning of the model. The calculation method of softmax is (a i For each value in the result vector):
[0051]
[0052] Let the maximum value in the result vector be a max , the second largest value is a submax , in order to perform early retreat of the spiking neural network, the confidence threshold required to be reached is x%, and the conditions that need to be met are:
[0053]
[0054] After moving the terms, it can be expressed as:
[0055]
[0056] (3) can be simplified as:
[0057]
[0058] (4) The right side satisfies:
[0059]
[0060] Combining (4) and (5), we can get:
[0061]
[0062] (6) Taking the logarithm of the left and right sides, we can get:
[0063]
[0064] (7) can be expressed as:
[0065] a max >β+a submax ,
[0066] Where x is the confidence threshold and n is the number of dimensions of the result vector.
[0067] Therefore, to determine whether the results of each time step of the spiking neural network meet the criteria for the early-exit mechanism, it is only necessary to observe the maximum and second-largest values in the result vector. When the absolute value of the difference between the two is greater than a certain threshold β, the complete calculation process can be stopped at the current time step, where β is a value related to the required confidence threshold and the number of dimensions of the result vector. Through the above derivation, the calculation process of early-exit can be simplified to finding the maximum and second-largest values and comparing the difference between the two with the threshold β. This calculation scheme is much simpler than the original calculation scheme that requires a small network or nonlinear function.
[0068] In some embodiments of the present invention, the credibility threshold is dynamically configured according to the requirements of the model.
[0069] The calculation of the above confidence threshold derivation process is more stringent. Therefore, in actual use, according to the needs of the model, the user can reconfigure the confidence threshold β, loosen or tighten it, to fine-tune the model's reasoning process and early exit progress.
[0070] Based on any of the foregoing embodiments, extracting the maximum value and the second largest value from the result vector includes:
[0071] If the result vectors arrive in sequence, the received result vectors are compared item by item, and the maximum value and the second largest value are dynamically updated.
[0072] If the result vectors arrive at the same time, the maximum value and the second largest value are extracted from the result vectors in parallel.
[0073] There are relatively few existing algorithms that can reduce the time step of pulse neural networks, and most of them require the use of complex calculation processes to extract features and assist in judgment, such as fully connected networks, reinforcement learning networks, etc., or the use of some nonlinear functions, such as sigmoid, softmax, etc. These algorithms generally bring a considerable amount of new calculations, are not suitable for hardware deployment, and have certain usage limitations and low versatility. The embodiment of the present invention derives and simplifies the calculation process of the original algorithm by observing the numerical results after the calculation of each time step is completed. It only needs to compare the maximum and second largest values in the results to determine whether the result of the current time step has reached a sufficiently high credibility and can be output as the final result. Based on this, the unfinished pulse calculation is terminated, saving reasoning time. Compared with the original technology, the algorithm designed in the embodiment of the present invention is simpler, does not require complex calculations, and is more hardware-friendly.
[0074] The neuromorphic hardware acceleration method provided by the embodiment of the present invention is applicable to both ASIC and FPGA. Tools have been used to implement ASIC design and have passed comprehensive testing. Before cropping, the recognition process of the ImageNet dataset took approximately 6 time steps, and the recognition process of the CIFAR10 dataset took approximately 4 time steps. The simplified time step cropping algorithm proposed in the embodiment of the present invention can shorten the recognition process of the ImageNet dataset to an average of 3 time steps, and the recognition process of the CIFAR10 dataset to an average of 2 time steps, while basically maintaining the optimal recognition accuracy.
[0075] The neuromorphic hardware acceleration method provided by the embodiment of the present invention is applicable to the calculation process of various pulse-based neural network models. Taking the classification task as an example: first, according to the task requirements, the credibility threshold of the calculation result is set. During the operation, the calculation and classification results of the current time step are obtained, and the maximum and second maximum values are found. If the difference between the maximum value and the second maximum value is greater than the preset credibility threshold, it can be considered that the result of the current time step is sufficiently credible, and the operation of the network model can be terminated, and the current result is used as the final output; otherwise, it is considered that the credibility of the result of the current time step is not high enough, and the calculation of the next time step needs to be performed. This algorithm that judges based on the two values in the current time step result of the pulse neural network can judge the calculation result in real time during the model operation at a relatively low computational cost, eliminating the complex multiplication and accumulation operations and nonlinear functions in other algorithms, and is applicable to all data sets and task scenarios.
[0076] The neuromorphic hardware acceleration device provided by the present invention is described below. The neuromorphic hardware acceleration device described below and the neuromorphic hardware acceleration method described above can be referenced to each other.
[0077] Figure 3 A schematic diagram of the structure of a neuromorphic hardware acceleration device provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the neuromorphic hardware acceleration device provided by the embodiment of the present invention includes:
[0078] The acquisition module 301 is used to obtain the result vector of the current time step in real time during the operation of the spiking neural network;
[0079] An extraction module 302 is configured to extract a maximum value and a second maximum value from the result vector;
[0080] Comparison module 303, configured to calculate the difference between the maximum value and the second largest value, and compare the difference with a credibility threshold;
[0081] The output module 304 is configured to terminate the calculation of subsequent time steps if the difference is greater than the credibility threshold, and output the result of the current time step as the final inference result.
[0082] In traditional solutions, lookup tables (LUTs) and multi-stage multiplication and accumulation units are used for nonlinear functions (such as sigmoid and softmax). In the embodiment of the present invention, only a comparator and a subtractor or a compressor are required to complete the judgment, which significantly reduces power consumption.
[0083] The neuromorphic hardware acceleration device provided by an embodiment of the present invention obtains the result vector of the current time step in real time during the operation of the spiking neural network; extracts the maximum value and the second largest value from the result vector; calculates the difference between the maximum value and the second largest value and compares it with a credibility threshold; if the difference is greater than the credibility threshold, terminates the calculation of subsequent time steps and outputs the result of the current time step as the final inference result. The embodiment of the present invention simplifies the confidence judgment to only compare the difference between the maximum value and the second largest value in the result vector, avoiding exponential operations and summations, and directly terminates the calculation of subsequent time steps based on the judgment result, saving invalid time and space resources and bringing considerable computing power and energy efficiency benefits.
[0084] An embodiment of the present invention further provides a serial arbitration module, including:
[0085] Multiple registers, respectively used to store the maximum value and the second maximum value of the currently received results, as well as the pre-stored confidence threshold;
[0086] The first comparator is used to compare a classification result input each time with the current maximum value and the second maximum value to update the maximum value and the second maximum value;
[0087] A subtractor is used to calculate the difference between the maximum value and the second maximum value after obtaining the final maximum value and the second maximum value;
[0088] The second comparator is used to compare the subtractor result with the credibility threshold to determine whether the model operation can be terminated early.
[0089] In the embodiment of the present invention, Figure 4As shown, the serial arbitration module includes three registers, one for storing the maximum and second-largest values of the currently received results, and a pre-stored confidence threshold β; a subtractor for calculating the difference between the final maximum and second-largest values after obtaining them; and two comparators: one for comparing the subtractor result with the confidence threshold to determine whether to terminate the model early, and the other for comparing each input classification result with the current maximum and second-largest values to update the maximum and second-largest values. The module's operation is controlled by input start and end signals. Before operation, the maximum and second-largest registers are pre-written with a value of 0. During operation, the input start signal is first turned on. When this state is on, whenever other circuits calculate a value in the result vector, this value is input to the module and compared with the stored maximum and second-largest values in comparator 1. If the value is greater than the maximum value, the maximum value is updated to that value; if it is between the maximum and second-largest values, the second-largest value is updated to that value; if it is less than the second-largest value, the second-largest value is discarded. After all values in the result vector have been calculated, the stored maximum and second-largest values are fed into a subtractor for subtraction. This difference is then compared in comparator 2 with a pre-stored confidence threshold β. If the difference is greater than the threshold, a model termination signal is output; otherwise, the signal remains at 0, and calculations continue for the next time step. After completing the judgment for the current time step, the serial arbitration module can be shut down using an end signal to save energy. A serial arbitration module can have a smaller footprint, more flexible judgment functionality, and support result vectors of varying dimensions. However, due to the serial input of data, the judgment time is longer. However, this characteristic may actually be more consistent with actual operation, where calculation results arrive sequentially.
[0090] An embodiment of the present invention further provides a parallel arbitration module, including:
[0091] Multiple process comparators are used to group multiple input result vectors into pairs and distinguish the size values within the groups;
[0092] A plurality of compressors, the plurality of compressors being connected in cascade and configured to find a maximum value and a minimum value in a plurality of input groups;
[0093] The result comparator is used to subtract the output value of the last stage compressor and compare the result with the pre-stored credibility threshold to determine whether the model operation can be interrupted.
[0094] like Figure 5As shown, the parallel adjudication module (10-classification) specifically includes: five comparators that divide the 10 input classification results into five groups, each with its own value; four 4-2 compressors that find the maximum and minimum values within the four input data points; and a result comparator that subtracts the output of the last 4-2 compressor (large minus small) and compares the result with a pre-stored confidence threshold β to determine whether to interrupt the model run. If the classification results calculated for the current time step are available simultaneously, the parallel adjudication module can also be used for this determination. Take the 10-classification inference task as an example. The input result vector contains 10 values, which are grouped into two groups, and each group is distinguished between the larger and smaller values using a comparator. The first four data sets are then divided equally into two groups. The maximum and minimum values of each group (assuming a and b are in one group, with a>b, and c and d are in another group, with c>d) are selected using a 4-2 compressor. The two maximum values are first compared; the larger one is the maximum (assuming a). The remaining larger value (c) is then compared with the smaller value (b) from the other group. If c>b, the next largest value is c; otherwise, b. This method only requires two comparisons to determine the maximum and next largest values among the four numbers. The newly obtained four values are then passed through the 4-2 compressor to obtain the new maximum and next largest values. Finally, these two values and the remaining group of numbers are passed through the 4-2 compressor again to obtain the maximum and next largest values among the ten numbers. These two values are then fed into a comparator for subtraction and compared with a pre-stored confidence threshold β. If the difference is greater than the threshold, a model termination signal is output; otherwise, the signal remains at 0, and the calculation continues at the next time step. Compared to the serial arbitration module, the parallel arbitration module operates on data in parallel, resulting in faster results. However, more internal comparators will bring some area and power consumption costs. At the same time, due to its fixed size, it has poor flexibility and cannot have the ability to judge the termination of any model.
[0095] Existing pulse-based neural network accelerators do not have a solution for reducing time steps to achieve acceleration effects, and all time steps must be fully calculated to obtain the final result. The embodiments of the present invention design two circuit units: a serial arbitration module and a parallel arbitration module. These modules can determine the credibility of the current calculation based on the inference results of the current time step obtained by real-time hardware operation, and promptly interrupt certain calculation and inference processes whose results can be determined in advance, thereby significantly reducing the number of time steps required to run the neural network model and bringing considerable computing power and energy efficiency benefits.
[0096] In addition, in terms of hardware implementation, the serial arbitration module provided by the embodiment of the present invention has a module area of about 10055 μm after integration. 2 The actual power consumption in operation is about 0.3mW, accounting for less than 1% of the entire system, which can save 50% of the inference time with almost no additional overhead.
[0097] On the other hand, the present invention also provides a non-transitory computer-readable storage medium. The present invention also provides a non-transitory computer-readable storage medium, which stores computer-readable instructions. The computer-readable instructions can be executed by a computer having one or more processors to enable the processor to execute the neuromorphic hardware acceleration method as described in the above embodiment, the method comprising: obtaining the result vector of the current time step in real time during the operation of the pulse neural network; extracting the maximum value and the second largest value from the result vector; calculating the difference between the maximum value and the second largest value, and comparing it with the credibility threshold; if the difference is greater than the credibility threshold, terminating the calculation of subsequent time steps, and outputting the result of the current time step as the final inference result.
[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A neuromorphic hardware acceleration method, characterized in that: include: During the operation of the spiking neural network, the result vector of the current time step is obtained in real time; Extracting the maximum value and the second maximum value from the result vector; Calculate the difference between the maximum value and the second largest value, and compare it with the credibility threshold; If the difference is greater than the credibility threshold, the calculation of the subsequent time step is terminated, and the result of the current time step is output as the final inference result.
2. The neuromorphic hardware acceleration method according to claim 1, wherein: The confidence threshold is calculated based on the confidence requirement reaching the threshold and the number of dimensions of the result vector.
3. The neuromorphic hardware acceleration method according to claim 1, wherein: The credibility threshold calculation method is: Confidence threshold Where x is the confidence threshold and n is the number of dimensions of the result vector.
4. The neuromorphic hardware acceleration method according to claim 1, wherein: The credibility threshold is dynamically configured according to the requirements of the model.
5. The neuromorphic hardware acceleration method according to claim 1, wherein: The extracting the maximum value and the second largest value from the result vector includes: If the result vectors arrive in sequence, the received result vectors are compared item by item, and the maximum value and the second largest value are dynamically updated.
6. The neuromorphic hardware acceleration method according to claim 1, wherein: The extracting the maximum value and the second largest value from the result vector includes: If the result vectors arrive at the same time, the maximum value and the second largest value are extracted from the result vectors in parallel.
7. A neuromorphic hardware acceleration device, characterized in that include: The acquisition module is used to obtain the result vector of the current time step in real time during the operation of the spiking neural network; An extraction module, configured to extract a maximum value and a second maximum value from the result vector; a comparison module, configured to calculate a difference between the maximum value and the second largest value, and compare the difference with a credibility threshold; The output module is used to terminate the calculation of subsequent time steps if the difference is greater than the credibility threshold, and output the result of the current time step as the final inference result.
8. A serial arbitration module, applicable to the neuromorphic hardware acceleration method according to claim 5, characterized in that: include: Multiple registers, respectively used to store the maximum value and the second maximum value of the currently received results, as well as the pre-stored confidence threshold; The first comparator is used to compare a classification result input each time with the current maximum value and the second maximum value to update the maximum value and the second maximum value; A subtractor is used to calculate the difference between the maximum value and the second maximum value after obtaining the final maximum value and the second maximum value; The second comparator is used to compare the subtractor result with the credibility threshold to determine whether the model operation can be terminated early.
9. A parallel arbitration module, applicable to the neuromorphic hardware acceleration method according to claim 6, characterized in that: include: Multiple process comparators are used to group multiple input result vectors into pairs and distinguish the size values within the groups; A plurality of compressors, the plurality of compressors being connected in cascade and configured to find a maximum value and a minimum value in a plurality of input groups; The result comparator is used to subtract the output value of the last stage compressor and compare the result with the pre-stored credibility threshold to determine whether the model operation can be interrupted.
10. A non-transitory readable storage medium, characterized in that: The non-transitory computer-readable medium stores computer-readable instructions, which can be executed by a computer having one or more processors to enable the processor to perform the neuromorphic hardware acceleration method according to any one of claims 1 to 6.