Target detection method based on pulse neural network and Transform
By combining pulse neural networks and Transformer modules, a new pulse neuron and feature extraction network is designed, which solves the problem of low performance of existing pulse neural networks and realizes efficient and low-energy-consuming object detection, and is suitable for edge small computing devices and brain-like chips.
Patent Information
- Application Number
- CN202510332760.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The existing target detection network based on pulsed neural networks has low performance, too many time steps and is difficult to train, and the network is poorly scalable, so it is impossible to fully combine the advantages of pulsed neural networks and artificial neural networks.
The object detection method based on pulsed neural network and Transformer module is adopted, and a new pulsed neuron and Transformer module are designed through feature extraction networks, and the fast Fourier convolution, spatial pyramid pooling layer and decoupled object detection head are combined to achieve high-performance object detection.
While reducing network operation power consumption and computing volume, it improves target detection performance and can be deployed on edge small computing devices and brain-like chip devices to achieve large-scale applications.
Smart Images

Figure CN120259630A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the application fields of object detection such as image object detection and underwater object detection, and discloses an object detection method based on the combination of a spiking neural network and a Transformer, and particularly relates to an object detection method for an end-to-end trained spiking neural network. Background Art
[0002] Object Detection is a research topic in computer vision, which has been fully studied in recent years, and its achievements have been widely applied at present. Since most traditional object detection algorithms are based on convolutional neural networks, the power consumption is relatively high during the deployment and application processes, and this point needs to be optimized specifically. Therefore, how to achieve a more energy-efficient and fast object detection method has been a research hotspot in recent years.
[0003] The Transformer module is a popular module in natural language processing. Due to the calculation of its self-attention, the network has a global receptive field, so it has been widely applied to computer vision tasks in recent years, such as image classification, object detection, semantic segmentation, image inpainting, etc. The Transformer module obtains image features through a full-image self-information mechanism. This module consists of an attention layer, a linear connection layer, and a normalization layer, and uses a skip-layer connection method for output to improve the stability of the module.
[0004] The spiking neural network (SNN) is currently regarded as the next-generation neural network. Compared with the artificial neural network (ANN) that uses convolutional operations or Transformer modules for calculation, the spiking neural network generates a pulse sequence mathematically to simulate the process of neurons firing nerve pulses in organisms, and realizes more efficient neural network calculations. The spiking neural network simplifies the amount of calculation by converting continuous data into discrete numerical pulses, thereby greatly reducing the energy consumption required to run the network. In the above process, the simplification of data values will reduce the network performance. Therefore, the spiking neural network introduces multiple time steps, and uses the information interaction of multiple time steps to enhance the information transmission inside the network, thereby reducing the performance gap between the spiking neural network and the traditional artificial neural network.
[0005] The existing object detection networks based on spiking neural networks mainly have the following problems: 1) Poor performance, and there is a large performance gap between the existing networks and the mature object detection algorithms based on artificial neural networks; 2) Too many time steps, in order to achieve high performance, the existing networks have too many time steps and are difficult to train; 3) Poor network scalability, currently it is very difficult to directly apply the spiking neural network module to the mature artificial neural network, and the advantages of the two cannot be fully combined. Summary of the Invention
[0006] The present invention discloses an object detection algorithm based on a spiking neural network and a Transformer module, including a feature extraction network based on a spiking neural network and a Transformer, the design of a new spiking neuron, and the design of a Transformer module based on spiking neurons, mainly solving the problem of the too low performance of the existing object detection network based on a spiking neural network.
[0007] The technical solution of the present invention:
[0008] An object detection method based on a spiking neural network and a Transformer is as follows:
[0009] (1) Feature extraction: Use the designed feature extraction network to perform various convolutional and spiking activation operations on the input image to obtain feature maps of the input image at different scales, and then use them for subsequent object detection;
[0010] The feature extraction network is based on the YOLOX object detection network and has a total of five layers. The first layer is the Focus convolution, which realizes the preliminary and rapid extraction of the feature map by performing block convolution on the input image. The second to fourth layers are the deep feature extraction network, which consists of a Depthwise Convolution (DW Conv) layer and a Cross Stage Partial (CSP) layer. The feature map obtained from the previous layer is sequentially input into the depth convolution layer and the cross-stage branch layer, and the feature map obtained from the previous layer is gradually further processed to generate a primary deep feature map with more details, laying a foundation for the fifth layer to further process the feature map. The depth convolution layer performs depth convolution on the feature map. The cross-stage branch layer has two branches inside, namely the large branch and the skip layer branch. The cross-stage branch layer divides the feature map extracted by the depth convolution layer into two parts. The first part of the feature map undergoes multi-layer convolution through the large branch inside to obtain more detailed feature extraction, and the other part of the feature map is merged with the output of the large branch through the skip layer branch to output the primary deep feature map. The fifth layer consists of five modules: the first depth convolution layer, the Spatial Pyramid Pooling layer, the CSP-FFC-SNN layer, the PSA-SNN layer, and the second depth convolution layer. The primary deep feature map output by the fourth layer is first processed by the first depth convolution layer, the Spatial Pyramid Pooling layer, and the CSP-FFC-SNN layer in sequence to obtain a high-level deep feature map. The high-level deep feature map is then used to extract the self-information map based on spiking neuron calculations through the PSA-SNN layer. Finally, the high-level deep feature map and the self-information map are merged through a depth convolution layer to obtain the final feature map. Among them, the CSP-FFC-SNN layer adds a Fast Fourier Convolution (FFC) and a Signed Spiking Neuron outside the cross-stage branch. The Spatial Pyramid Pooling (SPP) layer further extracts and optimizes the input feature map, and finally adds a SNN-based Partial Self-Attention (PSA-SNN) module. The Fast Fourier Convolution divides the input feature map into two layers by channel and then inputs them into two branches for processing respectively: one branch is a fully convolutional layer, which includes four convolutional layers to extract the local detail part in the feature map, and the other branch performs Fourier transform on the feature map to obtain the frequency domain feature, then performs convolution processing on the frequency domain feature in the frequency domain, and finally becomes the time domain feature through inverse Fourier transform. Finally, the feature maps of the two branches are merged as the final output of this layer.
[0011] The specific steps of FFC are as follows:
[0012] (1) Perform a fast Fourier transform on the input feature map and combine the real and imaginary parts:
[0013]
[0014] (2) Frequency-domain convolution, which obtains a global receptive field in the time domain under this operation:
[0015]
[0016] (3) Perform an inverse fast Fourier transform to restore the time-domain features:
[0017]
[0018]
[0019] (2) Object detection: For the third, fourth, and fifth layer feature maps obtained in step (1), first input them into the feature pyramid integration layer. The feature pyramid integration layer contains multiple convolutional layers, which fuse and interact with feature maps of multiple scales to fully fuse the feature maps at different scales and make the details included in the feature maps more abundant. Then input the fused feature maps into the decoupled detection head to obtain the object classification, object category, and object detection box results, and process them to obtain the final object detection results.
[0020] Furthermore, in the CSP-FFC-SNN layer, the neurons that generate pulses are redesigned. To add an additional pulse generation threshold, the neurons generate three different values of pulses (-1, 0, 1), and the expression for generating pulses is as follows:
[0021]
[0022] Among them, θ and θ ′ are the threshold values for generating positive pulses respectively, N is the number of times the neuron has generated positive pulses, is the membrane potential of each layer of neurons; in this method, a neuron can generate negative pulses only when it has already generated positive pulses;
[0023] The membrane potential update function is used to simulate the process of neurons accumulating membrane potential and emitting nerve pulses in the organism; when the membrane potential accumulates to the set threshold, it will emit pulses, and at the same time, the neuron membrane potential returns to the resting state to prepare for the next membrane potential accumulation before emitting pulses; due to the introduction of negative pulses, the membrane potential update function also needs to consider the situation of generating negative pulses; for the membrane potential update of the l-th layer of neurons at time t Its expression is:
[0024]
[0025] Among them, Ml-1 is a pulsed neuron in the (l - 1)-th layer. In the calculation, the levels of all neurons and their pulses need to be considered. is the weight of the l-th layer. is the bias. is the pulse of the (l - 1)-th layer; due to the introduction of negative pulses, the membrane potential of the neuron needs to consider the conditions for generating negative pulses, and its update expression is:
[0026]
[0027] Furthermore, the spatial pyramid pooling layer is to solve the problem that the convolutional neural network is sensitive to the size of the input image; through multi-scale pooling operations, the spatial pyramid pooling layer allows the network to process inputs of arbitrary sizes while generating feature maps of a fixed size; the spatial pyramid pooling layer includes multiple parallel max-pooling layers, and by varying the core size of the pooling layer, it is used to extract features at different levels; in this method, three parallel pooling layers are used, with pooling kernel sizes of 5×5, 9×9, and 13×13 respectively, to further extract the detailed features of the feature map from local to global; after passing through the spatial pyramid pooling layer, a convolutional layer is then used to concatenate the pooling outputs of different layers for subsequent further processing.
[0028] Furthermore, in the partial self-attention module based on the pulsed neural network, the pulsed neural network is combined with the Transformer module. The Transformer module consists of two modules: the Token mixer and the multi-layer perceptron, and pulsed neurons are added on the basis of these two modules; in the Token mixer, after the feature map output by the CSP-FFC-SNN layer is pulse serialized, the sequences for calculating self-information in the Transformer module are extracted through reparameterized convolution: the key value K, the query Q, and the true value V; the true value sequence is further pulse serialized to simplify the calculation of attention; the calculation formula for self-attention is as follows:
[0029]
[0030] where K S , Q S , V S are pulse sequences, scale is the control coefficient in the Transformer module, LIF(·) is the Leaky Integrate-and-Fire (LIF) pulse module, and RepConv(·) is the reparameterized convolution, which can improve the operation speed of the network and reduce the network calculation amount, and at the same time will not reduce the model performance.
[0031] The multi-layer perceptron uses a feedforward network. Different from the traditional feedforward network that uses one-dimensional convolution for full connection processing, it uses two-dimensional convolution for full connection processing, further combines features of multiple levels, so as to achieve better object detection performance.
[0032] Furthermore, for the object detection head, a decoupled object detection head is used, which adopts a two-branch design to separately detect the classification task and the regression task; after each object detection head extracts features further through two convolutional layers, the final results of each task are obtained through a separate detection head; for the classification task, the confidence of each classification is finally obtained, and the classification with the highest confidence is selected as the classification result of object detection; for the regression task, two detection heads are respectively used to output whether there is an object of a certain classification and the position detection box of the object; the three are combined to obtain the final object detection result.
[0033] The beneficial effect of the present invention is to propose a fast and effective object detection method. By introducing spiking neurons, the object detection network reduces the network operation power consumption, the number of parameters and the amount of computation while achieving high performance, and can effectively train the network in various object detection scenarios and deploy the network weights to edge small computing power devices and brain-like chip devices, realizing the large-scale application of the object detection algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the object detection network structure of the method of the present invention.
[0035] Figure 2 It is a schematic diagram of the PSA-SNN structure in the method of the present invention.
[0036] Figure 3 It is a schematic diagram of the object detection head structure of the method of the present invention.
[0037] Figure 4 It is a schematic diagram of the object detection result of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0038] The following further describes the specific embodiments of the present invention in combination with the drawings and technical solutions.
[0039] In the present invention, the training data is collected using the MS-COCO 2017 dataset, and the input image size is 640×640 pixels. The network in the present invention is implemented by programming in Python language with the PyTorch framework, and the spiking neurons are implemented using the SpikingJelly library. When training the network, the Adam optimizer is used, the initial learning rate of the network is set to 0.01, and the learning rate decay strategy uses the step decay method. The network is trained using an end-to-end method, that is, in step (1), the image to be detected is directly input to directly obtain the object detection result finally output in step (2). The setting of the network training loss function is different according to the different tasks of the decoupled detection head: the cross-entropy loss function is used for the classification task, and the intersection over union loss function is used for the regression task. During the training process, the data augmentation strategies used are Mosaic and MixUP. The Mosaic strategy adds a mosaic to the input image while keeping the ground truth unchanged, enhancing the generalization of the network through the blurred input image. This data augmentation method is adopted in the first half of the training process. The MixUP strategy can fuse multiple different images and ground truths in a certain proportion, thus enriching the number of objects in the input image. The combination of the MixUP strategy and the Mosaic strategy greatly improves the object detection performance of the network.
[0040] The present invention trains object detection networks of 5 different sizes, and the network sizes from small to large are: Nano, Tiny, S, M, L. These 5 network structures have different numbers of weights and computational amounts according to the different depths of repetition of each layer module of the network, which can meet the requirements of different devices for different object detection tasks, and greatly increase the deployment and application scope of the present invention.
[0041] Example 1
[0042] As Figure 4 shown, this figure shows the object detection results of the test set in the MS-COCO 2017 dataset. By comparing the object detection results with the feature maps in the detection head, it can be seen that the object detection network adopted by the present invention can accurately detect the specified objects (such as cats, cups, keyboards, etc.) in the input image. Table 1 shows the quantitative test results of different size models on the MS-COCO dataset, and the results show that the object detection method in the present invention reaches the top level in the industry.
[0043] Table 1 Test results of different size models on the MS-COCO dataset
[0044]
Claims
1. A target detection method based on spiking neural network and Transformer, characterized in that The steps are as follows: (1) Feature extraction: Use the designed feature extraction network to perform various convolutional and pulse activation operations on the input image to obtain feature maps of the input image at different scales, which are then used for subsequent object detection; The feature extraction network is based on the YOLOX object detection network and has a total of five layers. The first layer is the focus convolution, which realizes the preliminary and rapid extraction of the feature map by performing block convolution processing on the input image. The second to fourth layers are the deep feature extraction network, which are both composed of a depth convolution layer and a cross-stage branch layer. The feature map obtained from the previous layer is successively input into the depth convolution layer and the cross-stage branch layer to gradually perform further feature extraction on the feature map obtained from the previous layer, generating a primary deep feature map with more details, laying the foundation for the fifth layer to further process the feature map. The depth convolution layer performs depth convolution on the feature map; The cross-stage branch layer has two branches inside, namely the large branch and the skip layer branch; The cross-stage branch layer divides the feature map extracted by the depth convolution layer into two parts. The first part of the feature map undergoes multi-layer convolution through the internal large branch to obtain more detailed feature extraction, and the other part of the feature map is merged with the output of the large branch through the skip layer branch to output the primary deep feature map; The fifth layer consists of five modules: the first depth convolution layer, the spatial pyramid pooling layer, the CSP-FFC-SNN layer, the PSA-SNN layer, and the second depth convolution layer. The primary deep feature map output by the fourth layer is first processed by the first depth convolution layer, the spatial pyramid pooling layer, and the CSP-FFC-SNN layer in sequence to obtain the high-level deep feature map. The high-level deep feature map is then used to extract the self-information map based on pulse neuron calculation through the PSA-SNN layer. Finally, the high-level deep feature map and the self-information map are merged through a depth convolution layer to obtain the final feature map. Among them, the CSP-FFC-SNN layer adds a fast Fourier convolution and sign pulse neurons outside the cross-stage branch. The spatial pyramid pooling layer further extracts and optimizes the input feature map and adds a partial self-attention module based on the pulse neural network at the end. The fast Fourier convolution divides the input feature map into two layers by channel and then inputs them into two branches for processing respectively: one branch is a fully convolutional layer, which includes four convolutional layers to extract the local detail part in the feature map, and the other branch performs Fourier transform on the feature map to obtain the frequency domain feature, then performs convolution processing on the frequency domain feature in the frequency domain, and finally becomes the time domain feature through inverse Fourier transform. Finally, the feature maps of the two branches are merged as the final output of this layer; (2) Object detection: Input the feature maps of the third, fourth, and fifth layers obtained in step (1) into the feature pyramid integration layer first. The feature pyramid integration layer contains multiple convolutional layers to fuse and interact the feature maps of multiple scales, enabling the feature maps at different scales to be fully fused and making the details included in the feature map more abundant. Then, input the fused feature map into the decoupled detection head to obtain the object classification, object category, and object detection box results, and process them to obtain the final object detection result.
2. The object detection method based on pulsed neural network and Transformer according to claim 1, wherein In the CSP-FFC-SNN layer, neurons that generate pulses are redesigned. By adding an additional pulse generation threshold, the neurons generate three different values of pulses (-1, 0, 1). The expression for generating pulses is as follows: where θ and θ ′ are the threshold values for generating positive pulses respectively, N is the number of times the neuron has generated positive pulses, is the membrane potential of each layer of neurons; in this method, a neuron can generate a negative pulse if and only if it has already generated a positive pulse; The membrane potential update function is used to simulate the process of neurons accumulating membrane potential and firing nerve impulses in the body; when the membrane potential accumulates to the set threshold, an impulse will be fired, and at the same time, the neuron membrane potential returns to the resting state to prepare for the next membrane potential accumulation before firing an impulse; due to the introduction of negative impulses, the membrane potential update function also needs to consider the situation of generating negative impulses; for the membrane potential update of the neurons in the l-th layer at time t Its expression is: Among them, M l-1 is the spiking neuron of the (l - 1)-th layer. In the calculation, the levels of all neurons and their spikes need to be considered. is the weight of the l-th layer, is the bias, is the spike of the (l - 1)-th layer; due to the introduction of negative spikes, the membrane potential of the neuron needs to consider the conditions for generating negative spikes, and its update expression is:
3. The object detection method based on pulsed neural network and Transformer according to claim 1, wherein The spatial pyramid pooling layer is to solve the problem that the convolutional neural network is sensitive to the input image size. Through multi-scale pooling operations, the spatial pyramid pooling layer allows the network to process inputs of any size and simultaneously generates feature maps of a fixed size. The spatial pyramid pooling layer includes multiple parallel max pooling layers. By using different core sizes of the pooling layer, it is used to extract features at different levels. In this method, three parallel pooling layers are used, and their pooling kernel sizes are 5×5, 9×9, and 13×13 respectively, to further extract the detailed features of the feature map from local to global. After passing through the spatial pyramid pooling layer, a convolutional layer is used to splice the pooling outputs of different layers for subsequent further processing.
4. The object detection method based on pulsed neural network and Transformer according to claim 1, wherein In the partial self-attention module based on the pulsed neural network, the pulsed neural network is combined with the Transformer module. The Transformer module consists of two modules: the Token mixer and the multi-layer perceptron. Pulse neurons are added on the basis of these two modules. In the Token mixer, after the feature map output by the CSP-FFC-SNN layer is pulse serialized, the sequences for calculating self-information in the Transformer module are extracted through reparameterized convolution: the key value K, the query Q, and the true value V. The true value sequence is further pulse serialized to simplify the calculation of attention. The calculation formula for self-attention is as follows: Among them, K S , Q S , V S are pulsed sequences, scale is the control coefficient in the Transformer module, LIF(·) is the leaky integrate-and-fire pulse module, and RepConv(·) is the reparameterized convolution, which can improve the operation speed of the network and reduce the network computation amount, and at the same time will not reduce the model performance; The multi-layer perceptron uses a feed-forward network and performs fully connected processing using two-dimensional convolution to further combine features at multiple levels, thereby achieving better object detection performance.
5. The object detection method based on pulsed neural network and Transformer according to claim 1, wherein For the object detection head, a decoupled object detection head is used, and a two-branch design is adopted. The classification task and the regression task are detected separately. After further feature extraction by two convolutional layers for each object detection head, the final results of each task are obtained through a separate detection head. For the classification task, the confidence of each classification is finally obtained, and the classification with the highest confidence is selected as the classification result of the object detection. For the regression task, two detection heads are used to separately output whether there is an object of a certain classification and the position detection box of the object. The combination of the three obtains the final object detection result.
Citation Information
Cited By
Channel buoy detection method based on fusion of multi-mode pulse neural network and visual Transform
CN121616952A
Lightweight target trajectory prediction method based on twin pulse neural network
CN121659997A
Lightweight target trajectory prediction method based on twin-pulse neural network
CN121659997B