Pure pulse-driven Transform neural network automatic architecture search method
Through SN-ConvBN pulse residual connection and TransSpikingformer network architecture, the problem of non-pulse calculation in the deep SNN model is solved, and a low-power consumption and efficient neural network design is realized, which is suitable for a variety of device resource situations.
Patent Information
- Application Number
- CN202510411451.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
AI Technical Summary
The existing deep SNN model has non-pulsive calculations in the residual connection structure, resulting in high energy consumption and difficulty in deploying on neuromorphic hardware, while traditional NAS methods are time-consuming and difficult to find the optimal architecture.
Using the SN-ConvBN pulse residual connection structure, the TransSpikingformer network architecture is designed, combining pulse word segmentation, pulse Transformer block and classification head, and using the self-attention mechanism and pulse neuron characteristics, the optimal model is quickly screened through the training-free neural architecture search method.
It realizes the design of pulsed neural network with low power consumption and high performance, adapts to different resource constraints, and improves the deployment efficiency and accuracy of the model on neuromorphic hardware.
Smart Images

Figure CN120354890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to neural network technology in the field of artificial intelligence, particularly to the innovative integration of spiking neural networks (SNNs) and the Transformer architecture, and the application of the neural architecture search method without training in spiking neural networks. Background Art
[0002] With the continuous progress of artificial intelligence technology, neural networks have achieved remarkable achievements in many fields. However, traditional artificial neural networks (ANNs) based on floating-point operations have certain limitations in terms of energy efficiency. In contrast, spiking neural networks (SNNs), as a new generation of neural network models, exhibit many advantages. The core units of SNNs are spiking neurons, which release spike signals when the membrane potential reaches a specific threshold by simulating the behavior of real biological neurons. This spike-generation-based mechanism endows SNNs with event-driven characteristics. In the network, the computing and communication processes only occur at the spike-generation moment, effectively reducing the waste of computing resources, especially when dealing with sparse and dynamic data. In addition, the spike coding method of SNNs allows them to adopt low-power accumulate operations (AC) instead of traditional high-power multiply-accumulate operations (MAC). This characteristic has significantly improved the energy efficiency of SNNs. On neuromorphic hardware, this advantage in energy efficiency is even more prominent. Neuromorphic hardware is designed specifically to support the spike operations of SNNs, which can further reduce power consumption and improve computing speed, opening up a new path for the development of artificial intelligence technology.
[0003] The Transformer architecture was initially designed for natural language processing tasks, and its core is the self-attention mechanism, which can effectively capture the dependencies between different positions in a sequence. This architecture has also achieved great success in many computer vision tasks, including image classification, object detection, and semantic segmentation, etc. The success of the Transformer mainly owes to its ability to model global information through the self-attention mechanism, without being restricted by the local receptive fields in convolutional neural networks.
[0004] With the continuous in-depth research on SNN, its performance has been significantly improved. For example, ResNet has conducted extensive research on expanding the depth of SNN by introducing jump connections. SEW ResNet, as a representative convolutional SNN, successfully overcame the gradient vanishing / exploding problem in Spiking ResNet through a simple identity mapping mechanism, and became the first deep SNN to directly train more than 100 layers. On the other hand, Spikformer, as a Transformer-based SNN, became the first attempt to successfully introduce the Transformer architecture into SNN design by utilizing the self-attention ability and the biological characteristics of SNN, and showed strong performance.
[0005] However, existing deep SNNs, including Spikformer and SEW ResNet, face a common challenge: non-spiking computation. Specifically, the residual connection structure of these networks leads to non-spiking computation (such as integer-floating point multiplication), which not only limits their advantages in energy efficiency, but also makes them difficult to deploy on mainstream neuromorphic hardware that only supports spiking operations. The root cause of this non-spiking computation problem lies in the ADD residual connection, which introduces non-spiking data in the convolutional batch normalization (ConvBN) layer, thus destroying the pure pulse-driven nature of SNNs.
[0006] In the design process of SNN architecture, both manual design and traditional Neural Architecture Search (NAS) methods have obvious limitations. Manual design is not only time-consuming and laborious, but also difficult to guarantee the optimal network architecture. Traditional NAS methods usually require a large number of training phases or a super network training containing all architecture candidates, which leads to longer convergence time, especially when the training speed of SNN is generally slower than ANNs, this time cost becomes even more unacceptable.
[0007] In addition, existing NAS methods face other challenges when applied to SNNs. For example, the non-differentiability and high sparsity of SNNs make it difficult to directly apply traditional gradient-based NAS methods. At the same time, manual design or traditional NAS methods often find it difficult to fully utilize the unique characteristics of SNNs, such as event-driven computing and pulse coding, when searching for the optimal SNN architecture, thus limiting the further improvement of SNN performance.
[0008] In summary, it is of great significance to develop a pure SNN architecture that can avoid non-pulse calculations while maintaining high performance, and to design an efficient architecture search method for promoting the development and application of SNNs. Based on this background, the present invention proposes an implementation method of a spiking neural network based on automatic architecture search and Transformer, aiming to solve the above problems in the prior art. Summary of the Invention
[0009] The present invention proposes an automatic architecture search method for a pure pulse-driven Transformer neural network, aiming to solve the problems of non-pulse calculation energy consumption and long training time of traditional NAS methods in the prior art.
[0010] An automatic architecture search method for a pure pulse-driven Transformer neural network includes the following steps:
[0011] Step S1: Construct a pulse-driven residual connection structure, adjust the position of neuron operation, adopt the SN-ConvBN sequence, convert the input signal into a pulse signal through a spiking neuron layer first, and then perform convolution and batch normalization operations to ensure that the input and output of each layer of the network are pulse signals and eliminate non-pulse data;
[0012] Step S2: Design a TransSpikingformer network architecture, including a pulse tokenizer, multiple pulse Transformer blocks, and a classification head; wherein, the pulse Transformer block is composed of a pulse self-attention mechanism and a pulse multi-layer perceptron, and the self-attention mechanism uses a scaling factor to control the numerical range of the calculation result;
[0013] Step S3: Define the search space, including the embedding dimension, the number of multi-head attention heads, the MLP expansion ratio, and the network depth, and generate three types of candidate architectures, namely Tiny_Model, Small_Model, and Base_Model, according to the parameter range;
[0014] Step S4: Based on the neural architecture search method without training, using FLOPs as the performance index, combined with the evolutionary algorithm for iterative optimization, screen out efficient architectures that meet the resource constraints, print the top K network structure parameters, and save the top 1 network structure parameters;
[0015] Step S5: Select the top 1 network structure parameters, generate a TransSpikingformer network architecture according to this parameter, then verify the accuracy of the screened architecture on the target dataset, or optimize the model through retraining and deploy it to neuromorphic hardware.
[0016] The present invention proposes an automatic architecture search method for Transformer neural networks driven by pure pulses. Aiming at problems such as high energy consumption of traditional neural networks, time-consuming and laborious manual design, and long training time of traditional NAS methods, through the innovative SN-ConvBN pulse residual connection structure, non-pulse calculations in traditional residual connections are avoided, ensuring that all operations conform to the event-driven principle and significantly reducing energy consumption. Combining with the TransSpikingformer network structure, which includes a pulse tokenizer, pulse Transformer blocks, and a classification head, efficient feature extraction is achieved by using the self-attention mechanism and the characteristics of pulse neurons. Further, a method based on neural architecture search (NAS) without training is proposed. By defining the search space, generating candidate architectures, quickly screening the optimal model using FLOPs as an indicator, and iteratively optimizing with an evolutionary algorithm, it adapts to different resource constraints (Tiny / Small / Base models). This method shows higher accuracy and significant energy consumption advantages on different data sets, and can automatically find the optimal model structure according to resource constraints, suitable for various device resource situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the SN-ConvBN pulse residual connection.
[0018] Figure 2 It is a schematic diagram of the overall structure of the TransSpikingformer model.
[0019] Figure 3 It is a schematic flow of generating the TransSpikingformer structure based on the NAS method. DETAILED DESCRIPTION OF THE INVENTION
[0020] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0021] An automatic architecture search method for Transformer neural networks driven by pure pulses includes the following steps:
[0022] Step S1: Construct a pulse-driven residual connection structure, adjust the working positions of neurons, and adopt the SN-ConvBN sequence. First, convert the input signal into a pulse signal through a pulse neuron layer, and then perform convolution and batch normalization operations to ensure that the input and output of each layer of the network are pulse signals and eliminate non-pulse data.
[0023] The pulse-driven residual connection structure is SN-ConvBN, and the calculation satisfies the following formula:
[0024] X = SN(ConvBN(SN(X)))
[0025] where X represents the input, SN represents the neuron, and ConvBN is the combined operation of convolution and batch normalization; this calculation formula ensures that all inputs and outputs are pulse signals, and it becomes a pure addition operation on the convolutional layer, realizing pure pulse-driven calculation.
[0026] Step S2: Design the TransSpikingformer network architecture, including a pulse tokenizer, multiple pulse Transformer blocks, and a classification head; among them, the pulse Transformer block consists of a pulse self-attention mechanism and a pulse multi-layer perceptron, and the self-attention mechanism uses a scaling factor to control the numerical range of the calculation result.
[0027] The TransSpikingformer network structure specifically includes:
[0028] Pulse tokenizer: Downsample and embed the input two-dimensional image sequence, and map the input to pulse features;
[0029] Multiple pulse Transformer blocks: Each block includes a pulse attention and a pulse MLP block, which are used to encode and transform the pulse features;
[0030] Classification head: Consists of a fully connected layer, which is used to classify the encoded features.
[0031] The non-pulse calculations are mainly concentrated in the first layer convolutional feature extraction layer of the model, so the MACs operations in the model are mainly in the first layer. For different parameter models, the number of operations for calculating the theoretical MACs should be the same.
[0032] Step S3: Define the search space, including the embedding dimension, the number of multi-head attention heads, the MLP expansion ratio, and the network depth, and generate three types of candidate architectures, namely Tiny_Model, Small_Model, and Base_Model, according to the parameter range.
[0033] The parameter range of the search space is 4 - 5M, 11 - 15M, and 25 - 35M.
[0034] Step S4: Based on the neural architecture search method without training, using FLOPs as the performance metric, and combined with the evolutionary algorithm for iterative optimization, screen out efficient architectures that meet the resource constraints, print the top K network structure parameters, and save the top 1 network structure parameters.
[0035] The iterative optimization of the evolutionary algorithm includes the following steps:
[0036] Step S41: Initialize the population, generate random candidate architectures, and perform legality checks;
[0037] Step S42: Adjust the MLP expansion ratio, number of attention heads, and embedding dimension through mutation operations;
[0038] Step S43: Generate offspring architectures by fusing the parameters of two parent architectures through crossover operations;
[0039] Step S44: Iteratively screen the candidate architectures with the optimal FLOPs until the preset number of iterations is reached.
[0040] Step S5: Select the top 1 network structure parameters, generate the TransSpikingformer network architecture according to these parameters, then verify the accuracy of the screened architecture on the target dataset, or optimize the model through retraining and deploy it to neuromorphic hardware.
[0041] This method avoids non-spiking calculations through an innovative spiking residual learning mechanism, proposes the TransSpikingformer network structure, and combines a neural architecture search method without training to achieve a high-performance and low-power spiking neural network design. The specific content is as follows:
[0042] 1) Pulse-driven residual connections
[0043] The core idea of the residual connection is to directly add the input to the output of the network, enabling the network to directly learn the difference between the input and output. As shown in Figure 1 Figure -a, this design of skip connection not only avoids gradient vanishing but also suppresses gradient explosion, thus maintaining the effective transmission of gradients. The residual connection can be expressed by the following formulas (1) and (2). First, perform convolution and batch normalization operations, and then input the result into the SN layer for pulse-driven calculation, that is, the ConvBN-SN (ConvBN, Convolution + Batch Normalization; SN, Spiking Neuron) structure.
[0044] O l = SN l (ConvBN l (O l-1 )) + O l-1 = S l + O l-1 (1)
[0045] O l+1 = SN l+1 (ConvBN l+1 (Ol ))+O l =S l+1 +O l (2)
[0046] Among them, O l represents the output at layer l, S l Indicates the pulse emission in this layer.
[0047] However, the current residual design will inevitably bring non-pulse data, which will generate MAC operations in the next layer. l and O l-1 is a pulse signal, that is, {0, 1}, and their sum will produce O l , is a non-pulse signal, that is, {0, 1, 2}. This non-pulse signal data will have a huge impact on the spiking neural network. First, the introduction of non-pulse data causes SNN to perform unnecessary calculations, such as multiplication of non-pulse data in the convolution layer, which violates the event-driven principle. Second, non-pulse data (such as integers and floating-point numbers) need to be multiplied, and the energy consumption of multiplication is much higher than the accumulation operation (AC) performed by pulse data. As the depth of the network increases, the range of non-pulse data will continue to expand, resulting in further increase in energy consumption. This makes the energy consumption of SNN close to that of ANN with the same structure, losing its low power consumption advantage on neuromorphic hardware. Third, mainstream neuromorphic hardware only supports pulse operations and cannot directly run SNN models containing non-pulse data. Therefore, non-pulse data limits the deployment and application of SNN on neuromorphic hardware. It can be seen that this non-pulse signal will continue to participate in the operation in the next layer, destroying the pulse characteristics of the spiking neural network, and the pulse data is proportional to the residual block in the spiking neural network model.
[0048] Therefore, in order to avoid the generation of the above non-pulse data, the event-driven principle is guaranteed, such as Figure 1 -b, adjust the working position of the neuron, adopt the structure of SN-ConvBN, and use formula (3) and formula (4) to calculate. First, convert the input signal into a pulse signal, then perform convolution operation and batch normalization on the pulse signal, extract features and reduce the dimension. It can be seen that the structure of SN-ConvBN ensures that all inputs and outputs are pulse signals, and it becomes a pure addition operation on the convolution layer, thus realizing pure pulse-driven calculation.
[0049] O l =ConvBN l (SN l (O l-1 ))+O l-1 =S l +O l-1 (3)
[0050] O l+1 = ConvBN l+1 (SN l+1 (O l )) + O l = S l+1 + O l (4)
[0051] 2) TransSpikingformer Network Structure
[0052] As Figure 2 shown, the overall model includes three parts: a Spiking Tokenizer (ST), multiple Spiking Transformer Blocks, and a Classification Head. Given a two-dimensional image sequence Use the Spiking Tokenizer to downsample and perform petch embedding on the input, mapping the input to where the first layer in the Spiking Tokenizer serves as a spiking feature extractor. The X after being mapped by the Spiking Tokenizer will be sent to L Spiking Transformer Blocks. Similar to the classical ViT encoder structure, a Spiking Transformer Block in the present invention also includes a Spike Self-Attention (SSA) and a Spike Multi-layer Perceptron (SMLP) block. At the end of the model, the output of the Spiking Transformer Blocks is sent to the Classification Head. The Classification Head is mainly composed of a fully connected layer (FC). Global average pooling (GAP) is used before the fully connected layer to reduce the number of parameters of the FC and improve the classification ability of Spikingformer. All processes can be given by the following formula:
[0053]
[0054] Y = FC(GAP(X L )) or Y = FC(GAP(SN(X L ))) (8)
[0055] where T is the time window, C is the number of channels, H and W are the image sizes, is the input, N is the number of tiles after the input is segmented, D is the embedding dimension, representing the features of each tile; X l ' is the output of the spiking attention in the l-th Spiking Transformer Blocks, that is, the input of the l-th SMLP; X l is the processed output of the l-th SMLP, that is, the output of the l-th Spiking Transformer Blocks; Y is the output of the Classification Head, that is, the final classification result of the model.
[0056] 3) Search for the TransSpikingformer architecture based on the NAS method
[0057] As Figure 3 shown, using the number of operations (FLOPs) as the performance metric, quickly screen out high-performance TransSpikingformer architectures in the predefined search space without training.
[0058] Search space definition: Define the search space of the TransSpikingformer architecture, including key variables such as embedding size, number of heads in the multi-head attention mechanism, expansion ratio of the multi-layer perceptron, and network depth. And according to the above variables, the model is divided into three types: Tiny_Model, Small_Model, and Base_Model, as shown in Table 1, to correspond to the limited situation of model parameters under different device resources.
[0059] Table 1 Search space of the TransSpikingformer architecture
[0060] Tiny_model small_model base_model EmbedSize (192,384,64) (256,512,64) (384,768,64) MlpRation (3,5,1) (3,5,1) (3,6,1) HeadNum (4,8,4) (4,8,4) (4,8,4) Depth (1,8,1) (2,12,1) (4,15,1) ParamsRange 4-5M 11-15M 25-35M
[0061] Candidate architecture generation: Generate multiple candidate architectures in the predefined search space.
[0062] Performance evaluation based on FLOPs: Calculate the FLOPs of each candidate architecture and use it as the performance evaluation metric.
[0063] Architecture selection: According to the correlation between FLOPs and performance, screen out high-performance candidate architectures, and then iteratively optimize the architecture through evolutionary algorithms (such as tournament selection, crossover and mutation).
[0064] To verify the above invention content, the CIFAR10 and CIFAR100 datasets are selected for detailed step description.
[0065] Step 1: Modification of the Pulse Transformer based on the classical Transformer. For the classical self-attention mechanism (Class Self Attention, CSA), given an input feature sequence Query (Q) matrix, Key (K) matrix, and Value (V) matrix, these three quantities can all be calculated by three linear matrices and the input as follows,
[0066] Q F = XW Q , K F = XW K , V F = XW V (9)
[0067] where F represents the floating-point form, then Q F represents the query matrix in floating-point form, K F represents the key matrix in floating-point form, and V F represents the value matrix in floating-point form. Therefore, the output of the classical self-attention mechanism can be calculated by formula (10):
[0068]
[0069] where d = D / H, representing the feature dimension of each head, which is determined by the embedding dimension D and the number of heads H in the model's multi-head attention mechanism.
[0070] The softmax function normalizes the vector, which involves exponential operations and division operations. The operations of the SNN are based on the state update of neurons, including integration and threshold comparison, rather than floating-point multiplication and addition. Therefore, in order to make Q F , K F , V F become the pulse forms Q, K, V that can be used in the spiking neural network, this invention designs to use learnable matrices to convert the input feature sequence into Q, K, V matrices, and then these matrices pass through different spiking neuron layers, thereby converting the floating-point forms of Q F , K F , V F into the pulse forms of Q, K, V, and the calculation is shown in formula (11):
[0071] Q = SN Q (BN(XW Q ))), K = SN K (BN(XW K ))), V = SN V (BN(XWV )) (11)
[0072] Among them, SN Q 、SN K and SN V are all neurons used to transform the pulse sequence in the model, and they act on the linear layers W Q 、W K 、W V ; represent the query matrix, key matrix, and value matrix in the form of pulses respectively; BN() represents the batch normalization operation.
[0073] Therefore, the pulse self-attention mechanism SSA of the pulse transformer block can be expressed as
[0074] SSA'(Q, K, V) = SN(QK T V * s) (12)
[0075] SSA(Q, K, V) = SN(ConvBN(SSA'(Q, K, V))) (13)
[0076] Among them, SSA'(Q, K, V) is the result of the pulse self-attention calculation of the three matrices Q, K, and V, and s is the scaling factor to control the larger value of the matrix calculation result.
[0077] According to the content of the pulse-driven residual connection in the invention content, in order to obtain a pure pulse-driven transformer network, the calculations of the above Q, K, and V are further modified, and the calculation formulas are as follows
[0078] X' = SN(X) (14)
[0079] Q = SN Q (ConvBN Q (X′)) (15)
[0080] K = SN K (ConvBN K (X′)) (16)
[0081] V = SN V (ConvBN V (X′)) (17)
[0082] SSA(Q, K, V) = SN(ConvBN(SN(QK T V * s))) (18)
[0083] Step 2: Parse the parameters, initialize the data loader and the searcher, and start the search.
[0084] Step 3: Perform evolutionary search. First, initialize the population. The main task of this step is to generate initial candidate architectures and ensure their legality. Specifically, this step includes the following key operations:
[0085] Generate random candidate architectures: Randomly select parameters such as depth, mlp_ratio, num_heads, and embed_dim according to the definition of the search space in Table 1 to generate a specified number of random candidate architectures, and each candidate architecture conforms to the TransSpikingformer model framework.
[0086] Legality check: For each generated candidate architecture, perform a legality check. Check whether the candidate architecture has been visited (to avoid duplicate calculations). Calculate the complexity of the model (such as FLOPs), calculate the FLOPs parameter quantity of the candidate architecture according to formulas (19) and (20), and determine whether the candidate architecture is legal. Only legal candidate architectures will be added to the population. If it is legal, mark it as legal and record it as visited to prevent duplicate calculation of legality. If it is not legal, it will be skipped. This process continues until the number of candidate architectures in the population list reaches the specified population size. (The population size is a hyperparameter that needs to be given manually and is mainly determined by three points. First, the size of the search space. If the search space is large, a larger population size may be required to maintain diversity and avoid premature convergence to a local optimal solution. Second, the limitation of computing resources. The larger the population size, the higher the evaluation and operation cost per generation. Therefore, the setting of the population size needs to consider the available computing resources. Third, the characteristics of the search algorithm. Different evolutionary algorithms may have different requirements for the population size. For example, some algorithms may perform well with a smaller population size, while others may require a larger population size to maintain the diversity of the search. In this case, it is set to 50)
[0087] FLOPs SSA = L * n * d model *(2 * d model + n) (19)
[0088] FLOPs SMLP = L * n * (2 * d model * d mlp ) (20)
[0089] Where, L represents the number of layers, n represents the sequence length, d model represents the embedding dimension, d mlp represents the hidden layer dimension; FLOPs SSA represents the total number of calculations of the SSA module; FLOPs SMLP represents the total number of calculations of the SMLP module.
[0090] Step 4: After initializing the population, enter the iterative loop search (the number of iterations is a hyperparameter, set to 20 in this case). First, start updating the top k, and the specific operations are as follows:
[0091] Save the current candidate to memory: Save the candidate architectures in the current population to memory for subsequent analysis and recording.
[0092] Update the top k candidate list: According to the candidate architectures in the current population, sort the combined list according to metrics (such as the maximum FLOPs under Tiny_Model), and update and maintain a top k candidate list. The purpose of this step is to select the top k candidate architectures with the best performance in the current population.
[0093] Step 5: Generate a new generation of candidate frameworks, mainly including mutation, crossover operations, and random supplementation. The specific operations are as follows:
[0094] Mutation operation: Select the parent candidate architecture from the current top k candidate list and perform mutation operations on it. Conceptually mutate mlp_ratio, num_heads, and embed_dim to ensure that the generated candidate architecture is different from the parent, thereby increasing the diversity of the population.
[0095] Crossover operation: Randomly select two parent candidate architectures from the current top k candidate list and perform crossover operations on them. If the lengths of the two parent candidate architectures are different, reselect until the lengths are the same or the maximum number of attempts is reached. Randomly select one of each bit of the two parent candidate architectures as the corresponding bit of the offspring. Legality checks need to be performed to ensure that the generated crossover candidate architecture is legal.
[0096] Step 6: Add the new candidate architectures generated by mutation and crossover to the population list.
[0097] Step 7: To maintain diversity, it is necessary to perform the operations in Step 3, randomly select candidate architectures, analyze their legality, and add them to the population list.
[0098] Step 8: According to the performance metrics, select the top k candidate architectures with the best performance from the current population in the above population list and update the top k list.
[0099] Step 9: Repeat Steps 4 - 8 until the number of iterations reaches 20.
[0100] Step 10: After the iteration ends, print the final top k candidate boxes. Mark the top 1 as the optimal candidate box and save the optimal candidate architecture as a YAML file for subsequent model training and evaluation.
[0101] Step 11: Load the results finally saved in Step 10, generate a TransSpikingformer structure according to the parameters in the YAML file, and perform final dataset testing or retraining.
[0102] Technical effects:
[0103] 1) The energy consumption comparison of different self-attention mechanisms is shown in Table 2. Five self-attention mechanisms are compared. The first row in the data is the calculation of Relu, which only retains the positive values of Q, K, and V and sets the negative values to 0; the second row is LeakRelu, which partially retains the negative values on the basis of the positive values of Relu; the third row is the generation of the attention map using softmax according to the standard self-attention mechanism (CSA); the fourth row is the calculation result of the SSA method without corrected residual connection, and the fifth row is the SSA after the correction of the residual connection.
[0104] Table 2 Energy consumption comparison of pulse-driven residual connection transformers
[0105]
[0106]
[0107] From the analysis of accuracy, the corrected SSA pulse self-attention mechanism is more suitable for the pulse neural network model, achieving the highest accuracy on both datasets. On the CIFAR100 dataset, the pulse neural network model with the pulse SSA algorithm is 3.16% higher than softmax. From the analysis of MACs and ACs operation counts, based on the pulse-driven residual connection, most operations are calculated based on ACs, and only a small part of the operations are MACs. For example, on the CIFAR100 dataset, for the pulse neural network in the form of pulse SSA, the MACs operation is only 0.01G, and the ACs is 0.48G. However, for other pulse neural networks without modified residual connections, most of them are performing MACs operations. For example, the MACs operation of the pulse neural network in the form of SSA on the CIFAR100 dataset is 1.07G, which is 100 times that of pulse SSA. Due to the pulse-driven residual connection, most operations are calculated based on ACs, so the energy consumption advantage is more obvious. For example, other neural networks not integrated into the pulse SSA structure consumed 2.70mJ and 5.18mJ on the two datasets respectively, while the pulse neural networks composed of pulse SSA consumed 0.32mJ and 0.46mJ respectively, which is 88.15% and 91.92% less. Therefore, the experimental results show that the SSA method and the corrected pulse residual algorithm are more suitable for the pulse neural network.
[0108] 2) Performance comparison under different resource constraints
[0109] Table 3 Model Comparison under Different Resources
[0110] Methods Param(M) Timesteps CIFAR10 CIFAR100 DesignType Spikformer-4-26 4.15 4 93.94 75.96 Manual AutoSNN 5.44 8 92.54 69.16 Auto SNASNet-Bw - 5 93.64 73.04 Auto Tiny_Model 4.20 4 94.18 77.00 Auto TET 12.60 6 94.50 74.72 Manual DSR 11.20 20 95.40 78.20 Manual Spikformer-5-384 11.32 4 95.24 78.12 Manual Smal_Model 11.52 4 95.86 79.05 Auto Diet-SNN 39.90 5 93.44 69.67 Manual RMP 39.90 2048 93.63 70.93 Manual Spikformer-8-512 29.68 4 95.53 78.48 Manual base_Model 29.64 4 95.90 79.21 Auto
[0111] Table 3 shows the performance comparison of each model under different resource parameter requirements. According to the different floating-point operation counts (FLOPs), we divide the models into three categories: Tiny_Model, Small_Model, and Base_Model. Under the given resource constraints, all three models achieve the best performance. The size of the model is mainly affected by two factors: dimension (dim) and depth (depth). In resource-constrained platforms, manually designing the optimal model structure will consume a large amount of time and computing resources for repeated experiments. However, the algorithm proposed in the present invention can directly and proportionally automatically find the optimal model according to the size of the FLOPs parameter. Whether it is in terms of model capacity, computational complexity, or various requirements of different resource devices, this algorithm can effectively find the best model structure that meets the conditions.
Claims
1. A method for automatically searching the architecture of a Transformer neural network driven by pure pulses, characterized in that, It includes the following steps: Step S1: Construct a pulse-driven residual connection structure, adjust the working position of neurons, adopt the SN-ConvBN sequence, convert the input signal into a pulse signal through the pulse neuron layer first, and then perform convolution and batch normalization operations to ensure that the input and output of each layer of the network are pulse signals and eliminate non-pulse data; Step S2: Design the TransSpikingformer network architecture, including a pulse tokenizer, multiple pulse Transformer blocks, and a classification head; among them, the pulse Transformer block is composed of a pulse self-attention mechanism and a pulse multi-layer perceptron, and the self-attention mechanism uses a scaling factor to control the numerical range of the calculation result; Step S3: Define the search space, including the embedding dimension, the number of multi-head attention heads, the MLP expansion ratio, and the network depth, and generate three types of candidate architectures, namely Tiny_Model, Small_Model, and Base_Model, according to the parameter range; Step S4: Based on the neural architecture search method without training, using FLOPs as the performance metric, and iteratively optimize in combination with the evolutionary algorithm to screen out efficient architectures that meet the resource constraints, print the top K network structure parameters, and save the top 1 network structure parameters; Step S5: Select the top 1 network structure parameters, generate the TransSpikingformer network architecture according to this parameter, then verify the accuracy of the screened architecture on the target dataset, or optimize the model through retraining and deploy it to neuromorphic hardware.
2. The method for automatically searching the architecture of a Transformer neural network driven by pure pulses according to claim 1, wherein The pulse-driven residual connection structure is SN-ConvBN, and the calculation satisfies the following formula: X = SN(ConvBN(SN(X))) where X represents the input, SN represents the neuron, and ConvBN is the combined operation of convolution and batch normalization; this calculation formula ensures that all inputs and outputs are pulse signals, and it becomes a pure addition operation on the convolution layer, realizing pure pulse-driven calculation.
3. The method for automatically searching the architecture of a Transformer neural network driven by pure pulses according to claim 1, characterized in that, The TransSpikingformer network structure specifically includes: Pulse tokenizer: Downsample and embed the input two-dimensional image sequence, and map the input to pulse features; Multiple pulse Transformer blocks: Each block includes a pulse attention and a pulse MLP block, which are used to encode and transform the pulse features; Classification head: Composed of fully connected layers, which are used to classify the encoded features.
4. The method for automatically searching the architecture of a Transformer neural network driven by pure pulses according to claim 1, wherein, The non-pulse calculations are concentrated in the first-layer convolutional feature extraction layer of the model, so the MACs operations in the model are mainly in the first-layer convolutional feature extraction layer. For different parameter models, the number of operations for calculating the theoretical MACs is the same.
5. The method for automatically searching the architecture of a Transformer neural network driven by a pure pulse according to claim 1, characterized in that The parameter range of the search space is 4 - 5M, 11 - 15M, and 25 - 35M.
6. The method for automatically searching the architecture of a Transformer neural network driven by pure pulses according to claim 1, characterized in that: The iterative optimization of the evolutionary algorithm includes the following steps: Step S41: Initialize the population, generate random candidate architectures and perform legality checks; Step S42: Adjust the MLP expansion ratio, the number of attention heads, and the embedding dimension through mutation operations; Step S43: Fuse the parameters of two parent architectures through crossover operations to generate offspring architectures; Step S44: Iteratively screen the candidate architectures with the optimal FLOPs until the preset number of iterations is reached.