Visual image classification method based on pulse Transform
By building the pulse Transformer model on the Transformer model and fine-tuning it in the target domain, the problem of insufficient accuracy in the visual image classification of existing pulse neural networks is solved, and a low-energy consumption and high-accuracy visual image classification is achieved.
Patent Information
- Application Number
- CN202510415078.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Existing pulsed neural networks have problems with insufficient accuracy in visual image classification, which is difficult to meet actual needs.
By constructing a pulsed Transformer model based on the pre-trained Transformer model, it includes the construction and adjustment of the quantitative Transformer model, converting it into a pulsed Transformer model, and fine-tuning of the image classification task in the target domain.
It realizes low-energy consumption and high-quality target domain visual image classification, improves the accuracy of image classification, and meets the high accuracy requirements of visual classification.
Smart Images

Figure CN119942244A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the combination of brain-like computing and visual image classification, and specifically relates to a visual image classification method based on pulse Transformer. Background Art
[0002] In recent years, large language models based on the Transformer architecture have developed rapidly in the field of visual image classification, and their performance has surpassed traditional deep learning models. However, running such models requires a lot of computing and storage resources, making it impossible to deploy them on edge devices and then promote them to be widely used in daily life.
[0003] On the other hand, brain-like computing technology with spiking neural network (SNN) as the core has gradually attracted people's interest due to its unique computing architecture. Compared with traditional artificial neural network (ANN), SNN uses sparse pulse information representation and event-driven computing paradigm, which can be combined with neuromorphic devices to give full play to its high energy efficiency and low power computing advantages. In visual image classification, it can achieve low power consumption and high-quality visual image classification. However, the existing spiking neural network has technical problems in visual image classification, and the accuracy of visual image classification cannot meet actual needs. Summary of the invention
[0004] In view of the above, the purpose of the present invention is to provide a visual image classification method based on a pulse Transformer, and introduce a new construction method of a pulse Transformer model, that is, on the basis of a pre-trained initial Transformer model, by replacing different computing modules in the initial Transformer model, a quantized Transformer model is constructed as an intermediate, and a quantized activation function is used to process the model data of the quantized Transformer model, and finally the quantized Transformer model is reasonably adjusted and converted into a pulse Transformer model, and the pulse Transformer model constructed in this way is used to retrain the image classification task in the target domain, thereby realizing low-energy consumption and high-quality target domain visual image classification using the pulse Transformer model to meet the high accuracy requirements of visual classification.
[0005] To achieve the above-mentioned purpose of the invention, an embodiment provides a visual image classification method based on pulse Transformer, comprising the following steps: Pre-training the Transformer model for visual image classification based on the image sample set; A pulse Transformer model is constructed based on the pre-trained Transformer model, including: adjusting the multi-layer perceptron, layer normalization, and multi-head attention included in the standard Transformer encoder in the Transformer model through activation function adjustment, normalization adjustment, attention conversion, and quantization function replacement to obtain a quantized Transformer encoder including quantized perceptron, batch normalization, and quantized attention; the quantized Transformer encoder is obtained through layer fusion, pulse neuron parameter setting, matrix multiplication conversion between variables, and replacement of quantization function to pulse neuron node to obtain a pulse Transformer encoder including pulse perceptron, pulse neuron node, and pulse attention, and then a pulse Transformer model consisting of a multi-pulse Transformer encoder and a classification head of the original Transformer model is obtained; The spiking Transformer model is deployed on a neuromorphic device, and the classification head in the spiking Transformer model deployed on the neuromorphic device is fine-tuned using image samples of the target domain. The fine-tuned spiking Transformer model is used to classify images in the target domain.
[0006] Preferably, for the multi-layer perceptron, after adjusting the Gaussian linear error function GELU in the multi-layer perceptron to a linear rectification function ReLU, a quantization function is used to replace the linear rectification function ReLU to obtain a quantized perceptron; For layer normalization, replace layer normalization with batch normalization through normalization adjustment.
[0007] Preferably, for multi-head attention, batch normalization is added after the linear layer in the multi-head attention, and then the activation function in the multi-head attention is softmax () is converted to an activation function , then the converted multi-head attention for: ; in, Represent the query vector, key vector, and value vector in multi-head attention, respectively. T Indicates transposition, symbol represents matrix multiplication, represents the feature dimension of the key vector, N Indicates the length of the sequence; Finally, the quantization function is used to replace the multi-head attention The linear rectification function ReLU in , we get the quantized attention.
[0008] Preferably, the quantization function is expressed as: ; in, Indicates that for the independent variable quantization function, s represents the quantization step size, and Indicates the data clipping range: Given the quantized data width is b , for unsigned data and ; For signed data and , Indicates rounding to the nearest integer. Indicates that the value is limited between the specified minimum and maximum values. and For two different computational representations; When the quantization function is used to replace the linear rectification function ReLU, the independent variable in the quantization function is replaced by the independent variable in the linear rectification function ReLU.
[0009] Preferably, the quantized perceptron includes a linear layer, and adjacent batch normalization layers and linear layers are fused to obtain a pulse perceptron, including: When a Linear layer precedes a Batch Normalization layer, the Linear layer is and Implementing input The mapping relationship is: Output ; The batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , discarding the batch normalization layer operation to obtain , to achieve the fusion of two layers; When a linear layer is placed after a batch normalization layer, the batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; The linear layer passes the parameter and Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , we can also discard the batch normalization layer operation to obtain , realizing the fusion of two layers.
[0010] Preferably, the quantized attention includes a linear layer, and the quantized attention and the adjacent batch normalization are fused in the above-mentioned two-layer manner, and the pulse neuron parameters are set and the matrix multiplication conversion between variables is performed at the same time to obtain the pulse attention.
[0011] Preferably, when setting the pulse neuron parameters, set the time window ,in, Set the spike neuron threshold as the upper limit of the data clipping range in the quantization function ,in represents the scaling factor of each layer, Represents the step size in the quantized function.
[0012] Preferably, when performing matrix multiplication conversion between variables, the variables are first defined as follows: ; in, , , , Respectively represent the quantization function , , , The quantized outputs are equivalent to the pulse emission rate, , , , Respectively indicate time Different pulse outputs, represents a time window; Based on the above defined variables, the time split method is used to Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ,time ; Likewise Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ; This converts the input data of each time step in the Transformer encoder into the sum of various fractions.
[0013] Preferably, the quantization function in the quantized Transformer encoder is converted into a spiking neuron node by replacing the quantization function with the spiking neuron node, comprising: The cumulative firing neuron model is used during conversion, and the iterative calculation formula is: ; in, and Respectively represent l Layer i The neuron node at time t The membrane potential after cumulative charge and reset discharge, and For the l Layer and l- The threshold of the neurons in the first layer is different in different layers. Indicates l Layer for the i The neuron node to the j The connection weights of the outputs of the neuron nodes are Indicates i Neurons at time t Whether to issue a pulse. If it is issued, the value is 1, and if it is not issued, the value is 0. is the pulse emission function, satisfying .
[0014] Preferably, fine-tuning a classification head in a spiking Transformer model deployed on a neuromorphic device using image samples of the target domain comprises: The network parameters of the multi-pulse Transformer encoder in the pulse Transformer model are fixed, and the image samples of the target domain are input into the pulse Transformer model. After the visual features of the image samples in the target domain are extracted by the multi-pulse Transformer encoder, the classification head is used to perform image classification based on the visual features to obtain the image classification results. A loss function is constructed based on the difference between the image classification results and the true labels corresponding to the image samples, and the loss function is used to fine-tune the parameters of the classification head only.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention is based on the Transformer model and constructs a pulse Transformer through model conversion. The pulse Transformer model constructed in this way takes into account both the visual classification performance of the Transformer model and the low power consumption performance of the brain-like neural network deployed in neuromorphic devices. During the construction, the pulse Transformer model is also fine-tuned for the target domain visual image classification task in the target domain. In this way, low-cost and high-quality classification of target domain visual images can be achieved, the accuracy of image classification can be improved, and the high accuracy requirements of visual classification can be met. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 is a flow chart of a visual image classification method based on pulse Transformer provided in an embodiment; Figure 2 is a schematic diagram of the structure of a Transformer model for visual image classification provided in an embodiment; Figure 3 is a schematic diagram of conversion from a standard Transformer encoder to a quantized Transformer encoder provided in an embodiment; Figure 4 It is a curve comparison diagram of ReLU and GELU functions provided in the embodiment; Figure 5 is a schematic diagram of conversion from standard attention to quantized attention provided by an embodiment; Figure 6 is a schematic diagram of conversion from a quantized Transformer encoder to a pulse Transformer encoder provided in an embodiment; Figure 7 It is a schematic diagram of the conversion from quantized attention to pulse attention provided by an embodiment. DETAILED DESCRIPTION
[0018] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0019] like Figure 1As shown, the embodiment provides a pulse Transformer-based visual image classification method, comprising the following steps: S1, pre-training the Transformer model for visual image classification based on the image sample set.
[0020] In the embodiment, in the visual classification task, the standard Transformer model (Vision Transformer, ViT) for visual image classification is first pre-trained based on the image sample set. Figure 2 The figure shows a schematic diagram of the standard ViT model, which includes multiple word embeddings, a Transformer encoder, and a multi-layer perceptron used as a classification head, wherein the word embedding involves image block embedding, position embedding, and learnable [category] word embedding; the Transformer encoder is composed of multiple Transformer blocks, each of which includes two layer normalizations, a multi-head self-attention, and a multi-layer perceptron. The last head in the multi-head self-attention is usually a single fully connected linear layer.
[0021] In an embodiment, the visual classification task can be a classification task suitable for Cifar-10 classification. At this time, the standard Transformer encoder in the backbone of the standard ViT model is responsible for extracting low-level features that are irrelevant to the task, while the multi-layer perceptron in the top layer extracts high-level features related to the task and performs classification. Therefore, the embodiment uses a transfer learning strategy: first build a ViT model and load pre-trained parameters, and then adjust the top layer of the model to the adaptation task and start training. Specifically, build and load a ViT model trained on the large-scale dataset ImageNet, and then change the number of neurons in the output layer to 10 to adapt to the ten classification tasks of Cifar-10.
[0022] S2, builds a spiking Transformer model based on the pre-trained Transformer model.
[0023] In an embodiment, when constructing a pulse Transformer model based on a pre-trained Transformer model, the standard Transformer encoder in the pre-trained Transformer model is first converted into a quantized Transformer encoder, and then the quantized Transformer encoder is converted into a pulse Transformer encoder, thereby obtaining the pulse Transformer model.
[0024] In an embodiment, a standard Transformer encoder in a pre-trained Transformer model is converted into a quantized Transformer encoder, the pre-trained Transformer model for visual image classification is adjusted to adapt to the neuromorphic device and the pulse network operation mechanism, and then the adjusted model is subjected to quantized fine-tuning training. Figure 3 The figure shows the conversion diagram of the standard Transformer encoder to the quantized Transformer encoder. When converting the standard Transformer encoder, the multi-layer perceptron, layer normalization, and multi-head attention included in the standard Transformer encoder in the pre-trained Transformer model are adjusted by activation function, normalization adjustment, attention conversion, and quantization function replacement to obtain a quantized Transformer encoder including quantized perceptron, batch normalization, and quantized attention.
[0025] (S2-a) Activation function adjustment In the conversion method, the activation function in the original neural network model usually uses the linear rectified linear unit (ReLU), while the multi-layer perceptron in the Transformer encoder uses the Gaussian error linear unit (GELU). The two are different in calculation: ; in, x represents the independent variable input into the activation function, is the cumulative distribution function of the standard normal distribution. Figure 4 The following is a graph of the ReLU function and the GELU function. In practice, the tanh or sigmoid function is usually used to approximate the GELU function. Figure 4 It can be seen that the output values of GELU and ReLU functions under the same input are very close. In order to ensure the efficiency of the conversion method, the embodiment uses the ReLU activation function to replace the original GELU and perform subsequent steps.
[0026] (S2-b) Normalization adjustment Transformer emerged in the field of natural language processing, so the normalization operation in the model uses layer normalization (LN), which is commonly used to process language tasks. However, LN requires a lot of complex calculations (such as exponentiation, division, square root operations, etc.) in the reasoning stage, which makes it difficult to apply to deep SNN, and it cannot be implemented on low-power dedicated neuromorphic devices, which also brings challenges to the conversion of ViT. In contrast, batch normalization (BatchNormalization, BN) uses the parameter values statisticed during training in the reasoning stage, so that it can be fused with the linear layer to reduce the amount of calculation. On the other hand, BN has been successfully applied to deep learning models for visual tasks. Therefore, the embodiment chooses to replace the LNs in the original model with BNs, and adjusts the position of the normalization layer to facilitate the fusion of subsequent modules.
[0027] (S2-c) Attention switching The multi-head attention module is the most important part of the Transformer model. By calculating the degree of attention the model pays to different positions in the sequence, it can obtain useful information in the sequence and ignore invalid information, thereby efficiently processing long sequence tasks. The specific calculation is: ; in, is the word embedding after the linear transformation of the input, and Similarly, they are query vector, key vector, and value vector, respectively. is the characteristic dimension of the key vector, Indicates matrix multiplication, the superscript T It can be seen from the formula that the difficulty of multi-head attention conversion lies in the softmax function, matrix multiplication between variables and Conversion of division calculation. Currently, only the conversion of softmax is discussed, and the conversion strategies of the other two points will be described in subsequent steps. Figure 5 Schematic diagram of the conversion principle from standard attention to quantized attention.
[0028] like Figure 5 As shown in the figure, we first add a BN layer after the linear layer in the standard multi-head attention to make the model training converge faster and more stable. The main function of the softmax function in the multi-head attention is normalization, but it needs to perform many complex operations (including division and exponential calculations, etc.), which brings great challenges to the conversion and needs to be converted into other simple calculations. This function makes the output data non-negative and the sum is 1, so the expectation of each output value is 1 / N, where N is the sequence length. Therefore, the activation function is used The softmax function is replaced (which can be further replaced by a quantized function later), thus solving the problem of high computational complexity while keeping the output data non-negative and the mathematical expectation value 1 / N. In addition, some additional ReLU functions are added to convert floating-point inputs into pulse inputs later. for: .
[0029] (S2-d) Quantization function replacement The idea of converting artificial neural networks (ANN) to spiking neural networks (SNN) is to make the pulse firing rate of the converted spiking neurons approximate to the activation output value of the artificial neurons before the conversion. However, ANN usually uses very high-precision floating-point data, which requires SNN to simulate its exact value in a long time step (for example, a floating-point number 0.1 requires a single pulse to be issued in ten time steps to be equivalent, while 0.01 requires a single pulse to be issued in a hundred time steps), which seriously damages the energy efficiency and performance of SNN. The quantization model serves as a good link between the two. By effectively quantizing the ANN, the SNN can achieve performance comparable to that of the ANN by running in very few time steps (for example, classification tasks only require less than ten time steps), greatly reducing the running delay and power consumption of the SNN, and ensuring good application performance.
[0030] Therefore, the embodiment of the present invention adopts the learned step size quantization technology (LSQ) to quantize the original ANN. Specifically, the quantization function is used to replace all ReLUs in the standard Transformer encoder, including the ReLU in the multi-layer perceptron and the ReLU in the multi-head attention. The quantization function is calculated as follows: ; in, Indicates that for the independent variable quantization function, s represents the quantization step size, and Indicates the data clipping range: Given the quantized data width is b , for unsigned data and ; For signed data and , Indicates rounding to the nearest integer. Indicates that the value is limited between the specified minimum and maximum values. and The LSQ technique also provides reasonable initialization parameters and gradient scaling factors to ensure the generation of effective quantized Transformer encoders.
[0031] Based on the above activation function adjustment, normalization adjustment, attention conversion and quantization function replacement, when performing quantized Transformer encoder conversion, for the multi-layer perceptron, the Gaussian linear error function GELU in the multi-layer perceptron is adjusted to the linear rectifier function ReLU, and then the linear rectifier function ReLU is replaced by the quantization function to obtain the quantized perceptron; for layer normalization, batch normalization is used for replacement; Figure 5 As shown in the figure, for multi-head attention, first add a batch normalization layer after the linear layer in the multi-head attention, and then change the activation function in the multi-head attention softmax () is converted to an activation function , get the converted multi-head attention , and then use the quantization function to replace the multi-head attention The linear rectifier function ReLU in is used to obtain quantized attention. In this way, a quantized Transformer encoder including quantized perceptron, batch normalization and quantized attention is obtained. Subsequently, the quantized Transformer encoder can be further converted into a pulse Transformer encoder to obtain a pulse model.
[0032] In an embodiment, after obtaining a quantized Transformer encoder for a specific task, the quantized Transformer encoder is converted into a pulse Transformer encoder, thereby obtaining a pulse Transformer model. Figure 6 Schematic diagram of the conversion from quantized Transformer encoder to pulse Transformer encoder; Figure 7 The figure is a schematic diagram of the conversion from quantized attention to pulse attention. When converting the quantized Transformer encoder, through layer fusion, pulse neuron parameter setting, matrix multiplication conversion between variables, and replacement of quantization functions to pulse neuron nodes, a pulse Transformer encoder containing pulse perceptrons, pulse neuron nodes, and pulse attention is obtained, and then a pulse Transformer model consisting of a multi-pulse Transformer encoder and the classification head of the original Transformer model is obtained.
[0033] (S2-a') layer fusion Since the batch normalization layer statistically standardizes parameters during training and uses fixed parameters for calculation during inference, it can be fused with the linear layer to avoid floating-point multiplication and division, reducing computational complexity and facilitating the implementation of spiking neural networks and neuromorphic devices. There are two parameter setting schemes based on the order of precedence: When a Linear layer precedes a Batch Normalization layer, the Linear layer is and Implementing input The mapping relationship is: Output ; The batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , discarding the batch normalization layer operation to obtain , to achieve the fusion of two layers; When a linear layer is placed after a batch normalization layer, the batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; The linear layer passes the parameter and Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , we can also discard the batch normalization layer operation to obtain , to achieve the fusion of the two layers. Note that the actual affine transformation parameters should be tensors, and the dimensions need to be specifically considered.
[0034] (S2-b') Spiking neuron parameter setting By setting the time window Threshold of spiking neurons , the traditional convolutional neural network can directly perform parameter mapping after quantization. However, as can be seen from step S2, the Transformer encoder contains constant scaling transformations, such as and Calculated ), the threshold needs to be further adjusted.
[0035] Specifically, the input relationship before and after scaling is The pulse neuron emits a pulse after the membrane potential cumulative input reaches the threshold, so the judgment condition can be regarded as , Table weight parameters, since the scaling constant does not change over time, and each layer of neurons has the same scaling constant and the same firing threshold, by adjusting the neuron threshold Equivalence judgment conditions can be obtained , avoiding the division operation caused by scaling, when the scaling transformation of each layer corresponds to the scaling factor When setting the spike neuron threshold .
[0036] (S2-c') matrix multiplication conversion between variables After the model is quantized, the quantized activation output can be mapped to the pulse emission rate. Since attention involves multiplication operations between variables, this results in floating-point multiplication operations, which are difficult to adapt to pulse-driven pulse neural networks and neuromorphic devices. Therefore, the matrix multiplication between variables (especially in attention calculations) is converted to pulse matrix multiplication. Pulse matrix multiplication: two variables are multiplied, one of which is a binary pulse with a value of 0 / 1, that is, the pulse matrix multiplication only includes accumulation operations. Based on this, the embodiment uses a time splitting method to process the accumulation operation, and the variables are preferably defined as follows: ; in, , , , Respectively represent the quantization function , , , The quantized outputs are equivalent to the pulse emission rate, , , , Respectively indicate time Different pulse outputs, represents a time window; Based on the above defined variables, the time split method is used to Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ,time ; Likewise Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ; This converts the input data of each time step in the Transformer encoder into a sum of fractions. Each term in the numerator is a pulse matrix multiplication; the denominator is a constant scaling, which can be combined according to the threshold adjustment scheme.
[0037] (S2-d') Replacement of quantization function to spiking neuron node Based on the above conversion, we only need to replace the quantization function with the pulse neuron node on the basis of the quantization model, and fine-tune it according to the above steps to obtain the pulse Transformer model. This step is used to describe the pulse neuron model used. Specifically, the Integrate-and-Fire (IF) neuron model is used, and the iterative calculation formula is as follows: ; in, and Respectively represent l Layer i The neuron node at time t The membrane potential after cumulative charge and reset discharge, and For the l Layer and l- The threshold of the neurons in the first layer is different in different layers. Indicates l Layer for the i The neuron node to the j The connection weights of the outputs of the neuron nodes. Indicates i Neurons at time t Whether to issue a pulse. If it is issued, the value is 1, and if it is not issued, the value is 0. is the pulse emission function, satisfying In order to reduce the model conversion loss, a soft reset method of subtracting the threshold is adopted.
[0038] Based on the above-mentioned layer fusion, pulse neuron parameter setting, matrix multiplication conversion between variables, and replacement of quantization function to pulse neuron node, when performing pulse Transformer encoder conversion, for the quantized perceptron, the linear layer and batch normalization included in it are fused according to the above-mentioned two-layer fusion method to obtain the pulse perceptron; for the quantized attention, the linear layer and the adjacent batch normalization in the quantized attention are fused according to the above-mentioned two-layer fusion method, and the pulse neuron parameters are set and the matrix multiplication conversion between variables is performed at the same time to obtain the pulse attention; for the quantized function in the quantized Transformer encoder, it is converted into a pulse neuron node by replacing the quantization function to the pulse neuron node, and a pulse Transformer encoder including a pulse perceptron, a pulse neuron node, and a pulse attention is obtained. On this basis, the multi-pulse Transformer encoder is combined with the classification head of the original standard ViT model to form a pulse ViT model.
[0039] So far, the standard ViT model based on the Transformer architecture has been converted to a pulse ViT model. In theory, this conversion method is also applicable to the conversion of common deep models (including ANN, CNN, Transformer, etc.) to pulse models.
[0040] By adjusting and reconstructing the different components of the standard Transformer model for visual image classification, and quantizing and fine-tuning the entire model, the standard Transformer model can be efficiently converted to a pulse Transformer model. The overall conversion process is simple and clear, and it is extremely easy for researchers to operate and implement, which significantly reduces the development time and cost of developing new Transformer models on neuromorphic devices.
[0041] S3, deploying the spiking Transformer model on the neuromorphic device, using the image samples of the target domain to fine-tune the classification head in the spiking Transformer model deployed on the neuromorphic device, and using the fine-tuned spiking Transformer model to classify the images of the target domain.
[0042] In the embodiment, the pulse Transformer model is also deployed on the neuromorphic device, and the classification head in the pulse Transformer model deployed on the neuromorphic device is fine-tuned using the image samples of the target domain. The specific fine-tuning process is as follows: the network parameters of the multi-pulse Transformer encoder in the pulse Transformer model are fixed, and the image samples of the target domain are input into the pulse Transformer model. After the image samples of the target domain are extracted with visual features by the multi-pulse Transformer encoder, the classification head is used to perform image classification based on the visual features to obtain the image classification results, and a loss function is constructed based on the difference between the image classification results and the true labels corresponding to the image samples. The loss function is used to fine-tune the parameters of the classification head only. During fine-tuning, the target domain can be any image classification domain of interest, including vehicle classification, animal classification, and face classification in life scenes. After fine-tuning, the fine-tuned pulse Transformer model is used to perform image classification in the target domain.
[0043] The converted pulse Transformer model only needs to fine-tune the non-Transformer encoder part (i.e., the classification head) for the target domain. This method greatly reduces the time and resource cost required to directly build and train the pulse Transformer model for visual image classification tasks, and avoids difficult-to-achieve complex calculations, making it easy to deploy and apply on neuromorphic devices. It can achieve low power consumption and computing resources on neuromorphic devices, and high-performance visual image classification tasks.
[0044] In summary, the pulse Transformer-based visual image classification method provided by the embodiment of the present invention realizes low-cost and high-quality classification of target domain visual images by constructing a new pulse Transformer model, thereby improving the accuracy of image classification and meeting the high accuracy requirements of visual classification.
[0045] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A visual image classification method based on pulse Transformer, characterized in that: The following steps are involved: Pre-training the Transformer model for visual image classification based on the image sample set; A pulse Transformer model is constructed based on the pre-trained Transformer model, including: adjusting the multi-layer perceptron, layer normalization, and multi-head attention included in the standard Transformer encoder in the Transformer model, and obtaining a quantized Transformer encoder including a quantized perceptron, batch normalization, and quantized attention through activation function adjustment, normalization adjustment, attention conversion, and quantization function replacement; The quantized Transformer encoder is transformed into a spiking Transformer encoder including spiking perceptrons, spiking neuron nodes, and spiking attention through layer fusion, spiking neuron parameter setting, matrix multiplication conversion between variables, and replacement of quantization functions into spiking neuron nodes. This leads to a spiking Transformer model consisting of a multi-spiking Transformer encoder and the classification head of the original Transformer model. The spiking Transformer model is deployed on a neuromorphic device, and the classification head in the spiking Transformer model deployed on the neuromorphic device is fine-tuned using image samples of the target domain. The fine-tuned spiking Transformer model is used to classify images in the target domain.
2. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: For the multi-layer perceptron, after adjusting the Gaussian linear error function GELU in the multi-layer perceptron to the linear rectification function ReLU, the linear rectification function ReLU is replaced by the quantization function to obtain the quantized perceptron; For layer normalization, replace layer normalization with batch normalization through normalization adjustment.
3. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: For multi-head attention, we first add batch normalization after the linear layer in the multi-head attention, and then change the activation function in the multi-head attention softmax () is converted to an activation function , then the converted multi-head attention for: ; in, Represent the query vector, key vector, and value vector in multi-head attention, respectively. T Indicates transposition, symbol represents matrix multiplication, represents the feature dimension of the key vector, N Indicates the length of the sequence; Finally, the quantization function is used to replace the multi-head attention The linear rectification function ReLU in , we get the quantized attention.
4. The pulse Transformer-based visual image classification method according to claim 2 or 3, characterized in that: The quantization function is expressed as: ; in, Indicates that for the independent variable quantization function, s represents the quantization step size, and Indicates the data clipping range: Given the quantized data width is b , for unsigned data and ; For signed data and , Indicates rounding to the nearest integer. Indicates that the value is limited between the specified minimum and maximum values. and For two different computational representations; When the quantization function is used to replace the linear rectification function ReLU, the independent variable in the quantization function is replaced by the independent variable in the linear rectification function ReLU.
5. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: The quantized perceptron contains a linear layer. The adjacent batch normalization layer and the linear layer are fused to obtain a pulse perceptron, including: When a Linear layer precedes a Batch Normalization layer, the Linear layer is and Implementing input The mapping relationship is: Output ; The batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , discarding the batch normalization layer operation to obtain , to achieve the fusion of two layers; When a linear layer is placed after a batch normalization layer, the batch normalization layer is parameterized , , as well as Implementing input The mapping relationship is: Output ; The linear layer passes the parameter and Implementing input The mapping relationship is: Output ; By adjusting the parameters of the linear layer to , we can also discard the batch normalization layer operation to obtain , realizing the fusion of two layers.
6. The method for visual image classification based on pulse Transformer according to claim 5, characterized in that: The quantized attention includes a linear layer. The quantized attention and the adjacent batch normalization are fused according to the above two-layer method. The pulse neuron parameters are set and the matrix multiplication conversion between variables is performed at the same time to obtain the pulse attention.
7. The method for visual image classification based on pulse Transformer according to claim 6, characterized in that: When setting the pulse neuron parameters, set the time window ,in, Set the spike neuron threshold as the upper limit of the data clipping range in the quantization function ,in represents the scaling factor of each layer, Represents the step size in the quantized function.
8. The pulse Transformer-based visual image classification method according to claim 4, characterized in that: When performing matrix multiplication conversion between variables, first define the variables as follows: ; in, , , , Respectively represent the quantization function , , , The quantized outputs are equivalent to the pulse emission rate, , , , Respectively indicate time Different pulse outputs, represents a time window; Based on the above defined variables, the time split method is used to Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ,time ; Likewise Converted to multinomial pulse matrix multiplication, expressed as: ; Among them, the intermediate variable , intermediate variable ; This converts the input data of each time step in the Transformer encoder into the sum of various fractions.
9. The method for visual image classification based on pulse Transformer according to claim 4, characterized in that: The quantization function in the quantized Transformer encoder is converted into a spiking neuron node by replacing the quantization function with the spiking neuron node, including: The cumulative firing neuron model is used during conversion, and the iterative calculation formula is: ; in, and Respectively represent l Layer i The neuron node at time t The membrane potential after cumulative charge and reset discharge, and For the l Layer and l- The threshold of the neurons in the first layer is different in different layers. Indicates l Layer for the i The neuron node to the j The connection weights of the outputs of the neuron nodes are Indicates i Neurons at time t Whether to issue a pulse. If it is issued, the value is 1, and if it is not issued, the value is 0. is the pulse emission function, satisfying .
10. The method for visual image classification based on pulse Transformer according to claim 1, characterized in that: Fine-tune the classification head in the spiking Transformer model deployed on the neuromorphic device using image samples from the target domain, including: The network parameters of the multi-pulse Transformer encoder in the pulse Transformer model are fixed, and the image samples of the target domain are input into the pulse Transformer model. After the visual features of the image samples in the target domain are extracted by the multi-pulse Transformer encoder, the classification head is used to perform image classification based on the visual features to obtain the image classification results. A loss function is constructed based on the difference between the image classification results and the true labels corresponding to the image samples, and the loss function is used to fine-tune the parameters of the classification head only.
Citation Information
Patent Citations
Neural network quantification method for solving regression problem
CN114676826A
Target detection method and device, storage medium and electronic equipment
CN116403097A
Method and system for performing target detection on pulse signal
CN117252239A
Image classification method and device based on time reversible pulse Transform
CN118506094A
Picture annotation neural network model construction method and system
CN119089964A
Cited By
Automatic driving target classification method based on pulse deep learning
CN121640162A
An automatic driving target classification method based on pulse deep learning
CN121640162B
Visual sensor data classification system, training method and classification method
CN122574538A