Visual Image Classification Method Based on Pulse Transformer
By building the pulse Transformer model on the Transformer model and fine-tuning it in the target domain, the problem of insufficient accuracy in the visual image classification of existing pulse neural networks is solved, and a low-energy consumption and high-accuracy visual image classification is achieved.
Patent Information
- Application Number
- CN202510415078.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Existing pulsed neural networks have problems with insufficient accuracy in visual image classification, which is difficult to meet actual needs.
By constructing a pulsed Transformer model based on the pre-trained Transformer model, it includes the construction and adjustment of the quantitative Transformer model, converting it into a pulsed Transformer model, and fine-tuning of the image classification task in the target domain.
It realizes low-energy consumption and high-quality target domain visual image classification, improves the accuracy of image classification, and meets the high accuracy requirements of visual classification.
Smart Images

Figure CN119942244B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the combination of brain-inspired computing and visual image classification, and specifically relates to a visual image classification method based on a pulsed Transformer. Background Art
[0002] In recent years, large language models based on the Transformer architecture have been rapidly developed in the field of visual image classification, and their performance has even surpassed that of traditional deep learning models. However, running such models requires consuming a large amount of computing and storage resources, so that they cannot be deployed on edge devices and thus cannot be widely applied in daily life.
[0003] On the other hand, brain-inspired computing technology centered on spiking neural networks (SNNs) has gradually attracted people's interest due to its unique computing architecture. Compared with traditional artificial neural networks (ANNs), SNNs adopt sparse spike information representation and event-driven computing paradigms, and can give full play to their high energy efficiency and low power consumption computing advantages in combination with neuromorphic devices, and can achieve high-quality visual image classification with low power consumption in visual image classification. However, existing spiking neural networks have technical problems in that the accuracy of visual image classification cannot meet actual requirements in visual image classification. Summary of the Invention
[0004] In view of the above, the purpose of the present invention is to provide a visual image classification method based on a pulsed Transformer, introducing a new construction method for a pulsed Transformer model, that is, on the basis of a pre-trained initial Transformer model, by replacing different computing modules in the initial Transformer model, constructing a quantized Transformer model as an intermediate body, and using a quantized activation function to process the model data of the quantized Transformer model, and finally reasonably adjusting the quantized Transformer model to convert it into a pulsed Transformer model, and using the pulsed Transformer model constructed in this way to retrain the image classification task in the target domain, so as to realize the use of the pulsed Transformer model for low-energy consumption and high-quality visual image classification in the target domain to meet the high-accuracy requirements of visual classification.
[0005] To achieve the above invention purpose, a visual image classification method based on a pulsed Transformer provided by an embodiment includes the following steps:
[0006] Pre-train a Transformer model for visual image classification based on an image sample set;
[0007] Construct a spiking Transformer model based on the pre-trained Transformer model, including: adjusting the multi-layer perceptron, layer normalization, and multi-head attention included in the standard Transformer encoder in the Transformer model through activation function adjustment, normalization adjustment, attention conversion, and quantization function replacement to obtain a quantized Transformer encoder including a quantized perceptron, batch normalization, and quantized attention; converting the quantized Transformer encoder through layer fusion, spiking neuron parameter setting, variable-to-matrix multiplication conversion, and quantization function-to-spiking neuron node replacement to obtain a spiking Transformer encoder including a spiking perceptron, spiking neuron nodes, and spiking attention, and further obtaining a spiking Transformer model composed of multiple spiking Transformer encoders and the classification head of the original Transformer model;
[0008] Deploy the spiking Transformer model onto a neuromorphic device, fine-tune the classification head in the spiking Transformer model deployed on the neuromorphic device using image samples from the target domain, and use the fine-tuned spiking Transformer model for image classification of the target domain.
[0009] Preferably, for the multi-layer perceptron, after adjusting the Gaussian error linear unit (GELU) in the multi-layer perceptron to the rectified linear unit (ReLU), use a quantization function to replace the rectified linear unit (ReLU) to obtain a quantized perceptron;
[0010] For layer normalization, replace layer normalization with batch normalization through normalization adjustment.
[0011] Preferably, for multi-head attention, first add batch normalization after the linear layer in multi-head attention, and then convert the activation function softmax () to the activation function , then the converted multi-head attention is:
[0012] ;
[0013] wherein, respectively represent the query vector, key vector, and value vector in multi-head attention, the superscript T represents transpose, the symbol represents matrix multiplication, represents the feature dimension of the key vector, N represents the sequence length;
[0014] Finally, a quantization function is used to replace the multi-head attention in the linear rectifier function ReLU to obtain quantization attention.
[0015] Preferably, the quantization function is expressed as:
[0016] ;
[0017] where represents the quantization function for the independent variable , s represents the quantization step size, and represent the data clipping range: given the quantization data bit width as b , for unsigned data and ; for signed data and , represents rounding to the nearest integer, represents restricting the value between the specified minimum and maximum values, and are two different computational representation forms;
[0018] When using the quantization function to replace the linear rectifier function ReLU, the independent variable in the quantization function is replaced with the independent variable in the linear rectifier function ReLU.
[0019] Preferably, the quantization perceptron includes a linear layer, and the adjacent batch normalization layer and the linear layer are fused to obtain a pulse perceptron, including:
[0020] When the linear layer is before the batch normalization layer, the linear layer realizes the mapping relationship of the input and through the parameters as: the output ; the batch normalization layer realizes the mapping relationship of the input , , and through the parameters as: the output ; by adjusting the parameters of the linear layer to , the batch normalization layer operation is discarded to obtain , realizing the fusion of the two layers;
[0021] When the linear layer is after the batch normalization layer, the batch normalization layer realizes the mapping relationship of the input , , and through the parameters The mapping relationship is: output ; The linear layer passes through parameters and to implement the mapping relationship of the input as: output ; By adjusting the parameters of the linear layer to , it is also possible to discard the batch normalization layer operation to obtain and achieve the fusion of two layers.
[0022] Preferably, the quantization attention includes a linear layer. According to the above two-layer fusion method for quantization attention and the adjacent batch normalization, the pulse neuron parameter setting and the matrix multiplication conversion between variables are performed simultaneously to obtain the pulse attention.
[0023] Preferably, when setting the pulse neuron parameters, a time window is set, where is the upper limit of the data clipping range in the quantization function, and the pulse neuron threshold is set, where represents the scaling factor of each layer, and represents the step size in the quantization function.
[0024] Preferably, when performing the matrix multiplication conversion between variables, the variables are first defined as follows:
[0025] ;
[0026] Among them, , , , respectively represent the quantization outputs of the quantization function for , , , . These quantization outputs are equivalent to the pulse firing rate, , , , respectively represent the different pulse outputs at time , and represents the time window;
[0027] Based on the above-defined variables, using the time splitting method, is converted into a multi-pulse matrix multiplication, expressed as:
[0028] ;
[0029] Among them, the intermediate variable , the intermediate variable , and the time ;
[0030] Similarly, is converted into a multi-term pulse matrix multiplication, expressed as:
[0031] ;
[0032] where the intermediate variable , and the intermediate variable ;
[0033] In this way, the input data at each time step in the Transformer encoder is converted into a form of summing up each fractional term.
[0034] Preferably, the quantization function in the quantized Transformer encoder is converted into a spiking neuron node by replacing the quantization function with a spiking neuron node, including:
[0035] When converting, the cumulative firing neuron model is adopted, and the iterative calculation formula is:
[0036] ;
[0037] where and respectively represent the membrane potential of the l -th layer and the i -th neuron node after cumulative charging and reset discharging at time t , and are the firing thresholds of the neurons in the l -th layer and the l- +1-th layer, and the thresholds of different layers are different, represents the connection weight from the l -th layer to the output of the i -th neuron node to the j -th neuron node, represents whether the i -th neuron fires a pulse at time t , taking the value of 1 when firing and 0 when not firing, is the pulse firing function, satisfying .
[0038] Preferably, the classification head in the spiking Transformer model deployed on the neuromorphic device is fine-tuned using the image samples in the target domain, including:
[0039] Fix the network parameters of the multi-pulse Transformer encoder in the fixed-pulse Transformer model, and input the image samples in the target domain into the pulse Transformer model. After the image samples in the target domain are processed by the multi-pulse Transformer encoder to extract visual features, the classification head is used to perform image classification based on the visual features, obtain the image classification result, and construct a loss function based on the difference between the image classification result and the true label corresponding to the image sample. The loss function is used to fine-tune only the parameters of the classification head.
[0040] Compared with the prior art, the beneficial effects of the present invention at least include:
[0041] Based on the Transformer model, the present invention constructs a pulse Transformer through model conversion. The constructed pulse Transformer model not only takes into account the visual classification performance of the Transformer model but also the low-power performance of the neuromorphic device-based brain-like neural network. Moreover, during construction, the pulse Transformer model is fine-tuned for the visual image classification task in the target domain, which can achieve low-cost and high-quality classification of visual images in the target domain, improve the accuracy of image classification, and meet the high-accuracy requirements for visual classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 is a flowchart of a visual image classification method based on a pulse Transformer provided by an embodiment;
[0044] Figure 2 is a schematic structural diagram of a Transformer model for visual image classification provided by an embodiment;
[0045] Figure 3 is a schematic conversion diagram from a standard Transformer encoder to a quantized Transformer encoder provided by an embodiment;
[0046] Figure 4 is a curve comparison diagram of the ReLU and GELU functions provided by an embodiment;
[0047] Figure 5 is a schematic conversion diagram from a standard attention to a quantized attention provided by an embodiment;
[0048] Figure 6 It is a schematic diagram of the conversion from a quantized Transformer encoder to a pulsed Transformer encoder provided by the embodiment;
[0049] Figure 7 It is a schematic diagram of the conversion from quantized attention to pulsed attention. Detailed implementation manners
[0050] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0051] As Figure 1 shown, a visual image classification method based on a pulsed Transformer provided by the embodiment includes the following steps:
[0052] S1, pre-train a Transformer model for visual image classification based on an image sample set.
[0053] In the embodiment, during the visual classification task, first pre-train a standard Transformer model (Vision Transformer, ViT) for visual image classification based on an image sample set. As Figure 2 shown is a schematic diagram of the standard ViT model, which includes various word embeddings, a Transformer encoder, and a multi-layer perceptron used as a classification head. Among them, the word embeddings involve image patch embeddings, position embeddings, and learnable [[category]] token embeddings; the Transformer encoder is composed of multiple Transformer blocks, and each Transformer block includes two layer normalizations, a multi-head self-attention, and a multi-layer perceptron. The last head in the multi-head self-attention is usually a single fully-connected linear layer.
[0054] In the embodiment, the visual classification task can be a classification task applicable to Cifar-10 classification. At this time, the standard Transformer encoder in the backbone part of the standard ViT model is responsible for extracting task-irrelevant low-level features, while the multi-layer perceptron in the top layer part extracts task-related high-level features and performs classification. Therefore, the embodiment uses a transfer learning strategy: first construct a ViT model and load the pre-trained parameters, and then adjust the top layer part of the model to adapt to the task before starting training. Specifically, construct and load a ViT model trained on a large-scale dataset ImageNet, and then change the number of neurons in the output layer to 10, so as to adapt to the ten-class classification task of Cifar-10.
[0055] S2. Build a pulsed Transformer model based on the pre-trained Transformer model.
[0056] In the embodiment, when building a pulsed Transformer model based on the pre-trained Transformer model, first convert the standard Transformer encoder in the pre-trained Transformer model into a quantized Transformer encoder, and then convert the quantized Transformer encoder into a pulsed Transformer encoder, thereby obtaining a pulsed Transformer model.
[0057] In the embodiment, convert the standard Transformer encoder in the pre-trained Transformer model into a quantized Transformer encoder, adjust the pre-trained Transformer model for visual image classification to adapt to the neuromorphic device and the operation mechanism of the pulsed network, and then perform quantization fine-tuning training on the adjusted model. Figure 3 The figure shows the conversion schematic diagram from the standard Transformer encoder to the quantized Transformer encoder. When converting the standard Transformer encoder, the multi-layer perceptron, layer normalization, and multi-head attention included in the standard Transformer encoder in the pre-trained Transformer model are obtained by activation function adjustment, normalization adjustment, attention conversion, and quantization function replacement, resulting in a quantized Transformer encoder including a quantized perceptron, batch normalization, and quantized attention.
[0058] (S2-a) Activation function adjustment
[0059] In the conversion method, the activation function in the original neural network model usually uses the Rectified Linear Unit (ReLU), while the multi-layer perceptron in the Transformer encoder uses the Gaussian Error Linear Unit (GELU). There are differences in their calculations:
[0060] ;
[0061] Among them, x represents the independent variable input in the activation function, is the cumulative distribution function of the standard normal distribution. Figure 4 The figure shows the curves of the ReLU function and the GELU function. In actual operation, the GELU function is usually approximated by the tanh or sigmoid function. From Figure 4It can be seen that the output values of GELU and ReLU functions are very close under the same input. To ensure the efficiency of the conversion method, the embodiment uses the ReLU activation function to replace the original GELU and performs subsequent steps.
[0062] (S2-b)Normalization adjustment
[0063] Transformer emerged in the field of natural language processing. Therefore, layer normalization (LN), which is commonly used for processing language tasks, is adopted for the normalization operation in the model. However, LN requires a large amount of complex calculations (such as exponential, division, square root operations, etc.) during the inference stage, making it difficult to be applied to deep SNNs and even more impossible to be implemented on low-power dedicated neuromorphic devices, which also poses challenges to the conversion of ViT. In contrast, batch normalization (BN) uses the parameter values statistically calculated during training during the inference stage, enabling it to be fused with linear layers to reduce the computational amount. On the other hand, BN has been successfully applied to deep learning models for visual tasks. Therefore, the embodiment selects to replace the LNs in the original model with BNs and adjusts the position of the normalization layer to facilitate the fusion of subsequent modules.
[0064] (S2-c)Attention conversion
[0065] The multi-head attention module is the most important part of the Transformer model. By calculating the degree of attention of the model to different positions in the sequence, useful information in the sequence can be obtained while ignoring invalid information, thus efficiently processing long sequence tasks. The specific calculation is as follows:
[0066] ;
[0067] Among them, is the word embedding after the input is linearly transformed, and Similarly, they are the query vector, key vector, and value vector respectively, is the feature dimension of the key vector, represents matrix multiplication, and the superscript T represents transpose. It can be seen from the formula that the difficulty of multi-head attention conversion lies in the conversion of the softmax function, matrix multiplication between variables, and division calculation. Currently, only the conversion of softmax is discussed, and the conversion strategies for the other two points are described in subsequent steps. Figure 5 It is a schematic diagram of the conversion principle from standard attention to quantized attention.
[0068] As Figure 5As shown in the figure, first, a BN layer is added after the linear layer in the standard multi-head attention to make the model training converge faster and more stably. The main function of the softmax function in multi-head attention is normalization, but it requires many complex operations (including division and exponential calculations, etc.), which brings great challenges to conversion and needs to be transformed into other simple calculations. This function makes the output data non-negative and the sum is 1, so the expectation of each output value is 1 / N, where N is the sequence length. Therefore, the activation function is used to replace the softmax function (which can be further replaced by a quantization function later), thus solving the problem of high computational complexity while keeping the output data non-negative and the mathematical expectation being 1 / N. In addition, some additional ReLU functions are added to convert the floating-point input into a spike input later. The converted multi-head attention is as follows:
[0069] .
[0070] Replacement with quantization function
[0071] The conversion idea from an artificial neural network (ANN) to a spiking neural network (SNN) is to make the spike firing rate of the converted spiking neurons approximate the activation output value of the artificial neurons before conversion. However, ANN usually uses floating-point data with very high precision, which makes SNN need a long time step to simulate its exact value (for example, the floating-point number 0.1 needs to fire a single spike within ten time steps to be equivalent, while 0.01 needs to fire a single spike within a hundred time steps), seriously damaging the energy efficiency and performance of SNN. The quantization model serves as a good connection between the two. Through effective quantization of ANN, SNN can achieve performance comparable to ANN by running within a very small number of time steps (such as only less than ten time steps for classification tasks), greatly reducing the running latency and power consumption of SNN and ensuring good application performance.
[0072] Therefore, the embodiment of the present invention adopts the learned step size quantization technology (LSQ) to quantize the original ANN. Specifically, all ReLUs in the standard Transformer encoder are replaced with quantization functions, including the ReLUs in the multi-layer perceptron and the ReLUs in the multi-head attention. The quantization function is calculated as follows:
[0073] ;
[0074] where represents the quantization function for the independent variable , s represents the quantization step size, and Indicates the data clipping range: Given the quantization data bit width as b For unsigned data and ; For signed data and , Indicates rounding to the nearest integer, Indicates restricting the value between the specified minimum and maximum values, and are two different computational representation forms. The LSQ technique also provides reasonable initialization parameters and gradient scaling factors to ensure the generation of an effective quantized Transformer encoder.
[0075] Based on the above activation function adjustment, normalization adjustment, attention transformation, and quantization function replacement, when converting the quantized Transformer encoder, for the multi-layer perceptron, after adjusting the Gaussian error linear unit (GELU) in the multi-layer perceptron to the rectified linear unit (ReLU), the quantization function is used to replace the rectified linear unit (ReLU) to obtain the quantized perceptron; for layer normalization, batch normalization is used for replacement; as Figure 5 shown, for the multi-head attention, first add a batch normalization layer after the linear layer in the multi-head attention, and then convert the activation function softmax () in the multi-head attention to the activation function , to obtain the converted multi-head attention , and then use the quantization function to replace the rectified linear unit (ReLU) in the multi-head attention to obtain the quantized attention. In this way, a quantized Transformer encoder including a quantized perceptron, batch normalization, and quantized attention can be obtained. Subsequently, the quantized Transformer encoder can be further converted into a spiking Transformer encoder, and then a spiking model can be obtained.
[0076] In an embodiment, when the quantized Transformer encoder for a specific task is obtained, the quantized Transformer encoder is further converted into a spiking Transformer encoder, and then a spiking Transformer model is obtained. Figure 6 Is a schematic diagram of the conversion from the quantized Transformer encoder to the spiking Transformer encoder; Figure 7Schematic diagram for quantifying the conversion of attention to spiking attention. When converting the quantized Transformer encoder, through layer fusion, spiking neuron parameter setting, matrix multiplication conversion between variables, and replacement of the quantization function with spiking neuron nodes, a spiking Transformer encoder including spiking perceptrons, spiking neuron nodes, and spiking attention is obtained, and then a spiking Transformer model composed of multiple spiking Transformer encoders and the classification head of the original Transformer model is obtained.
[0077] (S2-a’) Layer fusion
[0078] Since the batch normalization layer calculates standardized parameters during training and uses fixed parameters for calculation during inference, it can be fused with the linear layer to avoid calculations such as floating-point multiplication and division, reduce the computational complexity, and facilitate the implementation of spiking neural networks and neuromorphic devices. There are two parameter setting schemes according to the different orders:
[0079] When the linear layer is before the batch normalization layer, the linear layer uses parameters and to implement the mapping relationship for the input as: output ; the batch normalization layer uses parameters , , and to implement the mapping relationship for the input as: output ; by adjusting the parameters of the linear layer to , discarding the batch normalization layer operation to obtain , realizing the fusion of the two layers;
[0080] When the linear layer is after the batch normalization layer, the batch normalization layer uses parameters , , and to implement the mapping relationship for the input as: output ; the linear layer uses parameters and to implement the mapping relationship for the input as: output ; by adjusting the parameters of the linear layer to , it can also discard the batch normalization layer operation to obtain , realizing the fusion of the two layers. Note that the actual affine transformation parameters should be tensors and the dimensions need to be specifically considered.
[0081] (S2-b’) Spiking neuron parameter setting
[0082] By setting a time window and the threshold of the spiking neuron , the traditional convolutional neural network can directly perform parameter mapping after quantization. However, as can be seen from step S2, the Transformer encoder contains a constant scaling transformation, such as and in the calculation), the threshold needs to be further adjusted.
[0083] Specifically, the input relationship before and after scaling is . The spiking neuron satisfies the cumulative input of the membrane potential and fires a pulse after reaching the threshold. Therefore, the judgment condition can be regarded as , represents the weight parameter. Since the scaling constant does not change with time, and each layer of neurons has the same scaling constant and the same firing threshold, by adjusting the neuron threshold the equivalent judgment condition can be obtained, avoiding the division operation caused by scaling. When the scaling transformation of each layer corresponds to the scaling factor , set the threshold of the spiking neuron .
[0084] (S2-c') Conversion of matrix multiplication between variables
[0085] After model quantization, the quantized activation output can be mapped to the pulse firing rate. Since the attention contains the multiplication operation between variables, this leads to floating-point multiplication operations, which are difficult to adapt to the pulse-driven spiking neural network and neuromorphic devices. Therefore, the matrix multiplication between variables (especially in the attention calculation) is converted to pulse matrix multiplication. Pulse matrix multiplication: when two variables are multiplied, one of the variables is a binary pulse with values of 0 / 1, that is, the pulse matrix multiplication only contains the accumulation operation. Based on this, the embodiment uses the time splitting method to process the accumulation operation. First, define the variables as follows:
[0086] ;
[0087] Among them, , , , respectively represent the quantization outputs of the quantization function for , , , , and these quantization outputs are equivalent to the pulse firing rate. , , , respectively represent the different pulse outputs at time . represents a time window;
[0088] Based on the above-defined variables and using the time splitting method, is transformed into a multi-term pulse matrix multiplication, expressed as:
[0089] ;
[0090] where the intermediate variable , the intermediate variable , and the moment ;
[0091] Similarly, is transformed into a multi-term pulse matrix multiplication, expressed as:
[0092] ;
[0093] where the intermediate variable , the intermediate variable ;
[0094] In this way, the input data at each time step in the Transformer encoder is transformed into a form of summation of fractional terms. Each term in the numerator is a pulse matrix multiplication; the denominator is a constant scaling, which can be combined according to the threshold adjustment scheme.
[0095] Replacement of the quantization function with a pulsed neuron node in (S2-d’)
[0096] Based on the above transformation, finally, only the quantization function needs to be replaced with a pulsed neuron node on the basis of the quantization model, and fine-tuning can be performed according to the above steps to obtain the pulsed Transformer model. This step is used to describe the adopted pulsed neuron model. Specifically, the integrate-and-fire (IF) neuron model is adopted, and the iterative calculation formula is as follows:
[0097] ;
[0098] where and respectively represent the membrane potential of the l th neuron node in the i th layer after cumulative charging and reset discharge at the moment t , and are the firing thresholds of the neurons in the l th layer and the l- +1 layer, and the thresholds of different layers are different. represents the l th layer for the i th neuron node to the jThe connection weights output by a neuron node. Indicates the i neuron at time t whether a pulse is emitted, taking the value 1 when emitted and 0 when not emitted. is the pulse emission function, satisfying . To reduce the model conversion loss, a soft reset method of subtracting the threshold is adopted.
[0099] Based on the above layer fusion, pulse neuron parameter setting, and matrix multiplication conversion between variables, as well as the replacement of the quantization function with a pulse neuron node, when performing the conversion of the pulse Transformer encoder, for the quantization perceptron, its linear layer and batch normalization are fused according to the above two-layer fusion method to obtain a pulse perceptron; for the quantization attention, the linear layer and the adjacent batch normalization in the quantization attention are fused according to the above two-layer fusion method, and at the same time, the pulse neuron parameter setting and the matrix multiplication conversion between variables are performed to obtain a pulse attention; for the quantization function in the quantization Transformer encoder, it is converted into a pulse neuron node through the replacement of the quantization function with a pulse neuron node to obtain a pulse Transformer encoder including a pulse perceptron, a pulse neuron node, and a pulse attention. On this basis, the multi-pulse Transformer encoder and the classification head of the original standard ViT model are combined to form a pulse ViT model.
[0100] So far, the standard ViT model based on the Transformer architecture can be converted into a pulse ViT model. Theoretically speaking, this conversion method is also applicable to the conversion of common deep models (including ANN, CNN, Transformer, etc.) into pulse models.
[0101] By adjusting and reconstructing different components of the standard Transformer model for visual image classification, and quantifying and then fine-tuning the entire model, the standard Transformer model is efficiently converted into a pulse Transformer model. The overall conversion process is simple and clear, which is extremely easy for researchers to operate and implement, significantly reducing the development time and cost of developing a new Transformer model on neuromorphic devices.
[0102] S3. Deploy the pulse Transformer model on the neuromorphic device, fine-tune the classification head in the pulse Transformer model deployed on the neuromorphic device using the image samples in the target domain, and use the fine-tuned pulse Transformer model for image classification in the target domain.
[0103] In the embodiment, the pulsed Transformer model is also deployed on the neuromorphic device, and the classification head in the pulsed Transformer model deployed on the neuromorphic device is fine-tuned using the image samples in the target domain. The specific fine-tuning process is as follows: Fix the network parameters of the multi-pulsed Transformer encoder in the pulsed Transformer model, and input the image samples in the target domain into the pulsed Transformer model. After the image samples in the target domain are used by the multi-pulsed Transformer encoder to extract visual features, the classification head is used to perform image classification based on the visual features to obtain the image classification result, and a loss function is constructed based on the difference between the image classification result and the true label corresponding to the image sample, and only the classification head is fine-tuned using the loss function. When fine-tuning, the target domain can be any image classification domain of concern, including vehicle classification, animal classification, and face classification in life scenarios, etc. After fine-tuning, the fine-tuned pulsed Transformer model is used to perform image classification in the target domain.
[0104] The converted pulsed Transformer model only needs to be fine-tuned for the non-Transformer encoder part (i.e., the classification head) for the target domain. This method greatly reduces the time and resource costs required to directly construct and train a pulsed Transformer model for visual image classification tasks, and avoids complex calculations that are difficult to implement, facilitating deployment and application on neuromorphic devices, and enabling low-power and high-performance visual image classification tasks on neuromorphic devices.
[0105] In summary, the visual image classification method based on pulsed Transformer provided by the embodiment of the present invention realizes low-cost and high-quality classification of visual images in the target domain by constructing a new pulsed Transformer model, improves the accuracy of image classification, and meets the high-accuracy requirements for visual classification.
[0106] The above specific embodiments have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A visual image classification method based on pulse Transformer, characterized in that: The following steps are involved: Pre-training the Transformer model for visual image classification based on the image sample set; A pulse Transformer model is constructed based on the pre-trained Transformer model, including: adjusting the multi-layer perceptron, layer normalization, and multi-head attention included in the standard Transformer encoder in the Transformer model through activation function adjustment, normalization adjustment, attention conversion, and quantization function replacement to obtain a quantized Transformer encoder including quantized perceptron, batch normalization, and quantized attention; the quantized Transformer encoder is obtained through layer fusion, pulse neuron parameter setting, matrix multiplication conversion between variables, and replacement of quantization function to pulse neuron node to obtain a pulse Transformer encoder including pulse perceptron, pulse neuron node, and pulse attention, and then a pulse Transformer model consisting of a multi-pulse Transformer encoder and a classification head of the original Transformer model is obtained; Deploy the spiking Transformer model on a neuromorphic device, use the image samples of the target domain to fine-tune the classification head in the spiking Transformer model deployed on the neuromorphic device, and use the fine-tuned spiking Transformer model to classify images in the target domain; The quantization function is expressed as: ; in, Indicates that for the independent variable quantization function, s represents the quantization step size, and Indicates the data clipping range: Given the quantized data width is b , for unsigned data and ; For signed data and , Indicates rounding to the nearest integer. Indicates that the value is limited between the specified minimum and maximum values. and There are two different calculation forms.
2. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: For the multi-layer perceptron, after adjusting the Gaussian linear error function GELU in the multi-layer perceptron to the linear rectification function ReLU, the linear rectification function ReLU is replaced by the quantization function to obtain the quantized perceptron; For layer normalization, replace layer normalization with batch normalization through normalization adjustment.
3. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: For multi-head attention, we first add batch normalization after the linear layer in the multi-head attention, and then change the activation function in the multi-head attention softmax () is converted to an activation function , then the converted multi-head attention for: ; in, Represent the query vector, key vector, and value vector in multi-head attention, respectively. T Indicates transposition, symbol represents matrix multiplication, represents the feature dimension of the key vector, N Indicates the length of the sequence; Finally, the quantization function is used to replace the multi-head attention The linear rectification function ReLU in , we get the quantized attention.
4. The pulse Transformer-based visual image classification method according to claim 2 or 3, characterized in that: When the quantization function is used to replace the linear rectification function ReLU, the independent variable in the quantization function is replaced by the independent variable in the linear rectification function ReLU.
5. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: The quantized perceptron contains a linear layer. The adjacent batch normalization layer and the linear layer are fused to obtain a pulse perceptron, including: When a Linear layer precedes a Batch Normalization layer, the Linear layer is and Implementing input The mapping relationship is: Output ; The batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , discarding the batch normalization layer operation to obtain , to achieve the fusion of two layers; When a linear layer is placed after a batch normalization layer, the batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; The linear layer passes the parameter and Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , we can also discard the batch normalization layer operation to obtain , realizing the fusion of two layers.
6. The method for visual image classification based on pulse Transformer according to claim 5, characterized in that: The quantized attention includes a linear layer. The quantized attention and the adjacent batch normalization are fused according to the above two-layer method. The pulse neuron parameters are set and the matrix multiplication conversion between variables is performed at the same time to obtain the pulse attention.
7. The method for visual image classification based on pulse Transformer according to claim 6, characterized in that: When setting the pulse neuron parameters, set the time window ,in, Set the spike neuron threshold as the upper limit of the data clipping range in the quantization function ,in represents the scaling factor of each layer, Represents the step size in the quantized function.
8. The pulse Transformer-based visual image classification method according to claim 4, characterized in that: When performing matrix multiplication conversion between variables, first define the variables as follows: ; in, , , , Respectively represent the quantization function , , , The quantized outputs are equivalent to the pulse emission rate, , , , Respectively indicate time Different pulse outputs, represents a time window; Based on the above defined variables, the time split method is used to Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ,time ; Likewise Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ; This converts the input data of each time step in the Transformer encoder into the sum of various fractions.
9. The method for visual image classification based on pulse Transformer according to claim 4, characterized in that: The quantization function in the quantized Transformer encoder is converted into a spiking neuron node by replacing the quantization function with the spiking neuron node, including: The cumulative firing neuron model is used during conversion, and the iterative calculation formula is: ; in, and Respectively represent l Layer i The neuron node at time t The membrane potential after cumulative charge and reset discharge, and For the l Layer and l- The threshold of the neurons in the first layer is different in different layers. Indicates l Layer for the i The neuron node to the j The connection weights of the outputs of the neuron nodes are Indicates i Neurons at time t Whether to issue a pulse. If it is issued, the value is 1, and if it is not issued, the value is 0. is the pulse emission function, satisfying .
10. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: Fine-tune the classification head in the spiking Transformer model deployed on the neuromorphic device using image samples from the target domain, including: The network parameters of the multi-pulse Transformer encoder in the pulse Transformer model are fixed, and the image samples of the target domain are input into the pulse Transformer model. After the visual features of the image samples in the target domain are extracted by the multi-pulse Transformer encoder, the classification head is used to perform image classification based on the visual features to obtain the image classification results. A loss function is constructed based on the difference between the image classification results and the true labels corresponding to the image samples, and the loss function is used to fine-tune the parameters of the classification head only.
Citation Information
Patent Citations
Neural network quantification method for solving regression problem
CN114676826A
Target detection method and device, storage medium and electronic equipment
CN116403097A