Range adaptive pulse neural network target detection method based on YOLO and Transform bridging

By introducing a range adaptive pulse neural network method based on YOLO and Transformer in SNN target detection, the problems of increased energy consumption and insufficient feature fusion in low-power scenarios are solved, and efficient target detection and scene perception capabilities are improved.

CN120014243APending Publication Date: 2025-05-16DALIAN UNIV OF TECH
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510170802.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing SNN object detection method is difficult to maintain low energy consumption characteristics in low-power scenarios, and the introduction of Transformer has led to an increase in energy consumption, lacking effective feature fusion and self-attention scaling factors, which affects model performance.

Method used

A range adaptive pulse neural network object detection method based on YOLO bridged with Transformer is proposed. Through the top attention mix feature fusion module and the range adaptive pulse attention module, efficient feature fusion and self-attention operation are achieved to avoid gradient vanishing.

Benefits of technology

While maintaining low power consumption, the accuracy and performance of SNN target detection are improved, the problem of I-LIF neurons lacking reasonable self-attention scaling factors is solved, and the scene perception ability of the model is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014243A_ABST
    Figure CN120014243A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence and target detection, and discloses a range adaptive spiking neural network target detection method based on YOLO and Transform bridging, and the method comprises the following steps: data preprocessing; extracting low-layer detail features; extracting high-level semantic information; carrying out feature channel dimension alignment; performing interaction in the scale; carrying out cross-scale feature fusion; training the network by using a joint loss function until convergence; and detecting a test set target. The method not only can effectively combine the global modeling capability of Transform and the local sensing characteristic of YOLO under the condition of low power consumption, but also can be matched with I-LIF neurons to adapt to any integer pulse to complete self-attention operation, and the target detection performance of the spiking neural network is remarkably improved under the condition of low power consumption. The range adaptive pulse neural network target detection method based on YOLO and Transform bridging can be widely applied to the field of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and target detection technology, and in particular to a range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging. Background Art

[0002] Object detection aims to automatically identify and locate objects of interest in images, accurately identify objects in images, and mark their positions and categories in images. Object detection plays a vital role in smart manufacturing, autonomous driving, and smart security. It can be applied in areas such as defect detection on production lines, automatic identification of vehicles and traffic signs, and intelligent alarms for illegal intruders.

[0003] In recent years, artificial neural networks (ANNs) have made significant progress in the field of target detection. However, in practical application scenarios, ANN methods still face many challenges due to limited computing resources and energy requirements of devices. Especially in low-power scenarios such as embedded devices and mobile devices, the high computing requirements of ANNs often lead to energy efficiency issues, which greatly limits their application in these scenarios. In contrast, spiking neural networks (SNNs) can complete tasks with higher computing efficiency and lower energy consumption through event-driven computing, so SNN target detection methods have received widespread attention from researchers. This type of method uses sparse pulse signals to transmit information, significantly reducing the amount of computation and energy consumption, and is regarded as an important direction for promoting the application of artificial intelligence technology on low-power devices.

[0004] At present, the research methods in the field of SNN target detection can be roughly divided into three categories: (1) unsupervised learning algorithms, which use the Hebbian learning rule of SNN to modify the connections between neurons according to the relative time of neuron pulses to learn the model. (2) ANN-to-SNN conversion methods, which map the parameters of the pre-trained ANN model to its corresponding SNN model based on the matching of ANN activation values ​​and SNN average firing rate. (3) Direct SNN training methods, which use proxy gradients to solve the non-differentiable problem of pulses and train SNN models from scratch.

[0005] Direct training of SNN uses proxy gradients to solve the non-differentiable problem of pulses, and is trained directly on data sets. It can achieve high performance in fewer time steps and effectively process static images and event data, while avoiding the limitations of the accuracy and structure of the ANN pre-training model, and showing higher flexibility in model construction. Compared with unsupervised learning algorithms, the direct training method can build deeper complex neural networks with a higher performance ceiling. Compared with the ANN-to-SNN conversion method, the direct training method avoids the limitations of the pre-trained ANN in model structure and performance, and can better process event data. Therefore, the direct training SNN method has stronger practicality in actual detection scenarios and has occupied a dominant position in past research. The present invention is mainly aimed at the field of target detection of direct training SNN, and proposes a range-adaptive SNN target detection method based on YOLO and Transformer bridging.

[0006] In recent years, the popular direct training SNN target detection methods mainly include pure YOLO architecture and hybrid Transformer architecture. Most SNNs used for target detection are based on pure YOLO architecture. This type of method builds a SNN-based YOLO model that can be directly trained by adding spike neurons and gradient replacement functions. For example, "Su Q, Chou Y, Hu Y, et al. Deep directly-trained spiking neural networks for object detection [C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023: 6555-6565" proposed a full pulse energy-efficient residual block EMS-ResNet, avoiding the redundant MAC operations caused by non-pulse convolution, thereby constructing the framework EMS-YOLO for direct training of SNN target detection, which has higher performance and lower energy consumption; "Luo X, Yao M, Chou Y, et al. Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection [C] / / European Conference on Computer Vision. Springer, Cham, 2025: 253-272" SpikeYOLO was proposed, which simplified the YOLOv8 design to make it more suitable for SNN target detection tasks. At the same time, I-LIF (integer leaky integral-firing) neurons were proposed, which reduced the quantization error of spike neurons and reduced power consumption through integer training and pulse reasoning. It is currently the most advanced spike neuron design and SNN target detection framework. However, due to the small receptive field characteristics of the pure convolution structure, these methods still have insufficient global modeling capabilities. Therefore, some methods proposed a hybrid Transformer architecture SNN target detection framework. These methods improve the global modeling capability by introducing Transformer, thereby improving model performance.For example, "Yao M, Hu J, Hu T, et al. Spike-driven transformer v2: Meta spiking neural network architecture inspiring the design of next-generation neuromorphic chips [J]. arXiv preprint arXiv: 2404.03663, 2024" introduces the Transformer architecture into SNN target detection to construct Meta-Transformer. The model consists of a series of Conv-based SNN Blocks and Transformer-based SNN Blocks, and finally directly connects the detection head. The performance of the model is improved by Transformer. Although the existing methods are effective, they ignore three important factors: (1) The main application advantage of SNN lies in its low energy consumption characteristics, which can be fully utilized in low-power scenarios such as embedded devices and mobile devices. However, the current Transformer hybrid architecture has introduced too many self-attention mechanisms, which makes the model energy consumption increase or even approach the power consumption of ANN in the same scenario, which cannot reflect the low power consumption advantage of SNN. (2) Although the addition of Transformer improves the global modeling capability of the model, the hybrid architecture removes the feature fusion stage, which will lead to bottlenecks in effectively extracting local features and affect model performance. (3) The introduction of the Transformer architecture in SNN requires corresponding pulse self-attention for intra-scale interaction. Although existing methods have proposed pulse self-attention compatible with SNN, they lack reasonable scaling factors for I-LIF neurons. This results in these attention mechanisms being only applicable to LIF neurons with 0 / 1 binary outputs. When faced with I-LIF neurons that output multiple integer pulses, gradient vanishing may occur, affecting the improvement of model accuracy.

[0007] In view of the above problems, the present invention proposes a range-adaptive pulse neural network target detection method based on the bridge of YOLO and Transformer, which aims to combine the global modeling capability of Transformer and the local perception characteristics of YOLO while maintaining the low power consumption characteristics of SNN, and cooperate with I-LIF neurons to complete the process of integer training and pulse reasoning, thereby completing the effective bridge of YOLO and Transformer while maintaining low power consumption, so that the model can perform accurate target detection on static images and event data. Summary of the invention

[0008] In order to solve the above technical problems, the purpose of the present invention is to provide a range-adaptive pulse neural network target detection method based on YOLO and Transformer bridge. The method uses a top-attention hybrid feature fusion (TAHFF) module, including the use of scale-based interactions based on self-attention operations to globally model high-level features, capture the relationship between semantic information and conceptual entities, and effectively reduce energy consumption; and adopts a cross-scale feature fusion method to make up for the shortcomings of Transformer in local feature extraction, and realizes the efficient fusion of YOLO's local modeling ability and Transformer's global modeling ability under low power conditions, thereby improving model performance. In addition, the present invention also proposes a range-adaptive pulse attention (RASA) module for scale-based interaction, which avoids the gradient vanishing phenomenon by adjusting the scaling factor, and cooperates with I-LIF neurons to adapt arbitrary integer outputs for integer training and pulse reasoning self-attention operations, thereby improving the performance of the model under the premise of low power consumption. Therefore, this method not only improves the accuracy of SNN target detection by bridging YOLO and Transformer while effectively maintaining the low power consumption advantage of SNN, but also solves the problem of I-LIF neurons lacking a reasonable self-attention scaling factor, improves the effectiveness of self-attention operations, and significantly improves the performance of SNN target detection.

[0009] The technical solution of the present invention is a range-adaptive pulse neural network target detection method based on YOLO and Transformer bridge, comprising the following steps:

[0010] Step 1: Perform data preprocessing on the static image dataset and the dynamic image dataset to obtain the pulse feature matrix Input T×C×H×W , where T represents the time step, C represents the number of channels, and H×W represents the spatial resolution;

[0011] Step 2: Process the pulse feature matrix Input obtained in step 1 T×C×H×W , extract low-level detail features through the downsampling layer and the pulse-based convolution block SNNBlock to obtain the low-level pulse feature matrix Feature low-stage2 ;

[0012] Step 3: Process high-order features by using the downsampling layer and the multi-scale feature fusion module MSFF to process the low-level pulse feature matrix Feature obtained in step 2 low-stage2 Extract high-level semantic information and obtain the pulse feature matrix Feature high-stage3and the matrix of spatial pyramid pooling output;

[0013] Step 4: Align the feature channel dimension of the low-level detail features in step 2 and the high-level semantic information in step 3, and the low-level pulse feature matrix Feature low-stage2 , pulse feature matrix Feature high-stage3 The matrices output by the spatial pyramid pooling are aligned and are respectively the S3 feature matrix, the S4 feature matrix, and the S5 feature matrix, which are used as feature matrices for subsequent intra-scale interaction and feature fusion;

[0014] Step 5: Apply the intra-scale interaction based on the self-attention operation to the S5 feature matrix output in step 4;

[0015] Step 6: Perform cross-scale feature fusion on S3, S4 obtained in step 4 and F5 obtained in step 5;

[0016] Step 7: Use the joint loss function to train the network until convergence; after convergence, the entire network performs target detection.

[0017] The process of obtaining the pulse feature matrix from the static image data set is as follows: repeating the static image along the time dimension and using it as the input value of each time step T; encoding the continuous input values ​​into pulse signals through pulse neurons to obtain the pulse feature matrix.

[0018] The dynamic image dataset is neuromorphic data, and each data point is an event, including pixel coordinates, timestamp and polarity. The process of obtaining the pulse feature matrix is ​​as follows: aggregate the event stream within a fixed time window into frames; given a spatiotemporal window ζ, the asynchronous event stream E = {e n ∈ζ:n=1,...,N} represents a sparse grid of points in three-dimensional space. A constant time window dt is used to divide E into time intervals, and the events are mapped into a two-dimensional matrix representation of the image. T fixed time steps are processed each time, and the total sequence Γ=T×dt, where dt and T are the constant time window and time step respectively. The continuous input values ​​of the total sequence Γ are encoded into pulse signals through pulse neurons to obtain a pulse feature matrix.

[0019] The pulse feature matrix Input T×C×H×W Preliminary feature extraction is performed in sequence through stage-1 and stage-2; the structures of stage-1 and stage-2 are the same, both consisting of a downsampling layer and an SNNBlock;

[0020] The stage-1 processing flow is as follows: Pulse feature matrix Input T×C×H×W First, pass the downsampling layer to get the matrix Input into SNNBlock for feature extraction, the structure of SNNBlock is as follows:

[0021]

[0022] Feature low-stage1 =Input′+ChannelConv(Input′) (2)

[0023] Among them, + indicates that the matrix addition operation completes the residual connection, Input′∈R T×C×H×W Represents the intermediate output matrix after the first residual, Feature low-stage1 ∈R T×C×H×W represents the output matrix of SNNBlock, T represents the time step, C represents the number of channels, and H×W represents the spatial resolution;

[0024] SepConv(·) is a reverse separable convolution module with a 7×7 convolution kernel for capturing global features, and a 3×3 convolution is added to further perform spatial feature fusion; the SepConv(·) is specifically expressed as:

[0025]

[0026] Among them, Conv pw1 (·) and Conv pw2 (·) is point-wise convolution, Conv dw1 (·) and Conv dw2 (·) is the depthwise convolution, and BN(·) represents the batch normalization operation;

[0027] ChannelConv(·) is used as a channel mixer to realize the fusion of information between channels, which is expressed as:

[0028] ChannelConv(Input′)=Conv 3×3 (SN(Conv 3×3 (SN(Input′)))) (4)

[0029] Among them, Conv 3×3 (·) represents a standard convolution operation with a kernel size of 3×3;

[0030] SN(·) represents the spike neuron layer, using I-LIF neurons, and its specific calculation method is as follows:

[0031] U[t]=H[t-1]+X[t] (5)

[0032] H[t]=β(U[t]-S[t]) (6)

[0033] S[t]=Clip(round(U[t]),0,D) (7)

[0034] Where t is the pulse time step, U[t] is the membrane potential that combines the temporal information H[t-1] of the previous time step t and the spatial information X[t] input at the current time t, S[t] is the integer pulse matrix, round(·) is the rounding function, Clip(x,min,max) means clipping x to the range [min,max], and D is a hyperparameter representing the maximum integer value that an I-LIF neuron can emit; the membrane potential U[t] decays by a factor of β and is reset by subtracting S[t] after the pulse S[t] is emitted, otherwise H[t] remains unchanged;

[0035] In the inference phase, the integer pulse values ​​emitted by the I-LIF neurons are converted into binary pulses by the following formula (8), thus ensuring that the inference phase is pulse-driven;

[0036]

[0037] Where X l [t] represents the input of the lth layer of neurons, and S l [t,d] represents a pulse sequence, which contains only 0 / 1, and W l represents the coefficient matrix extracted during the expansion process;

[0038] The entire I-LIF spiking neuron model is expressed as:

[0039] S=SN(U) (9)

[0040] Where SN(·) is the spiking neuron layer mentioned above, whose input is the membrane potential tensor U and output is the spiking tensor S;

[0041] The structure and processing flow of stage-2 and stage-1 are exactly the same. low-stage1 The low-level pulse feature matrix Feature is formed through stage-2 processing low-stage2 .

[0042] The low-level pulse feature matrix Feature low-stage2 The multi-dimensional features of the object are captured from different levels through stage-3 and stage-4. The structures of stage-3 and stage-4 are the same, both of which consist of a downsampling layer and a multi-scale feature fusion module. Finally, the pulse feature matrix Feature output by stage-4 is processed by spatial pyramid pooling SPPF. high-stage4 Perform multi-scale spatial pooling processing;

[0043] The stage-3 processing flow is as follows: Low-level pulse feature matrix Feature low-stage2 First, the matrix is ​​obtained by downsampling layer processing Then input into the multi-scale feature fusion module, whose structure is as follows:

[0044]

[0045] Feature high-stage3 =F3′+ChannelConv2(F3′) (11)

[0046] Among them, F3′∈R T×C×H×W Represents the intermediate output matrix after the first residual, Feature high-stage3 ∈R T ×C×H×W Represents the output matrix of the multi-scale feature fusion module;

[0047] DMSFF(·) is a dilated multi-scale feature fusion module, which is mainly composed of four parallel dilated grouped convolutions with different dilation rates. The four dilated grouped convolution outputs are channel-connected and channel-downsampled using 1×1 convolution blocks. The dilated multi-scale feature fusion module is specifically described as:

[0048]

[0049] in, Indicates that the convolution is a dilated convolution of size 3×3, represents the feature matrix obtained by dilated convolution, d represents the dilation rate, g represents the group, c represents the number of channels, SN(·) represents the spike neuron layer, Concat(·) represents the matrix concatenation operation, Conv 1×1 (·) indicates that the convolution is a standard convolution of size 1×1; the result of the dilated multi-scale feature fusion module is input into SepConv(·) to refine the features;

[0050] ChannelConv2(·) is used as a channel mixer to achieve information fusion between channels. It uses a re-parameterized convolution with a kernel size of 3×3 to minimize the parameter count, which is described as:

[0051] ChannelConv2(F′3)=BN(RepConv(SN(BN(RepConv(SN(F′3)))))) (14)

[0052] RepConv(U′)=Conv pw2 (Conv dw1 (Conv pw1 (U′))) (15)

[0053] Among them, RepConv(·) represents a parameterized convolution with a kernel size of 3×3, U′∈R T×C×H×W RepConv(·) represents the input matrix, which is reparameterized as a standard convolution during inference; Conv pw1 (·) and Conv pw2 (·) is point-wise convolution, Conv dw1 (·) is a depthwise convolution, SN(·) represents a spiking neuron layer, and BN(·) represents a batch normalization operation;

[0054] The structure and processing flow of stage-4 and stage-3 are exactly the same. high-stage3 The pulse feature matrix Feature is formed through stage-4 processing high-stage4 ;

[0055] The spatial pyramid pooling is described as:

[0056] y1=MaxPool(Feature high-stage4 )y2=MaxPool(y1)y3=MaxPool(y2) (16)

[0057] SPPF(Feature high-stage4 )=Conv 1×1 (SN(Concat(Feature high-stage4 ,y1,y2,y3)))(17)

[0058] Among them, MaxPool(·) represents the maximum pooling operation with a kernel size of 5×5, y1, y2, and y3 represent the output matrices after the pooling operation, and Conv 1×1 (·) represents a standard convolution operation with a convolution kernel of 1×1, SN(·) represents a spike neuron layer, and Concat(·) represents a matrix concatenation operation.

[0059] The step 4 is specifically as follows: inputting the output of stage-2 in step 2, the output matrix of stage-3 and SPPF in step 3 into the LCB block for channel alignment, and its structure is as follows:

[0060] LCB(U in )=BN(Conv(SN(U in ))) (18)

[0061] Among them U in ∈R T×C×H×Wrepresents the layer input matrix, BN(·) represents batch normalization operation, Conv(·) represents standard convolution operation, and SN(·) represents spiking neuron layer; Conv(·) adopts convolution with kernel 1×1, and LCB block is a spiking convolution with kernel 1×1, which is used to adjust the channel dimension of feature matrix;

[0062] The feature matrices of stage-2, stage-3 and SPPF after LCB alignment are recorded as S3, S4 and S5 respectively, which are used for intra-scale interaction and feature fusion operations in subsequent processes.

[0063] The intra-scale interaction is accomplished through a Transformer-based SNN block; the Transformer-based SNN block includes a range-adaptive spike attention RASA and a SpikeMLP; the Transformer-based SNN block for intra-scale interaction is expressed as:

[0064] Attention=S5+RASA(Q,K,V) (19)

[0065] F5=Attention+SpikeMLP(Attention) (20)

[0066] SpikeMLP(Attention)=SN(SN(Attention)W1)W2 (21)

[0067] Among them, Attention∈R T×C×H×W represents the intermediate output matrix after the first residual, F5∈R T×C×H×W Represents the output matrix of Transformer-based SNN Block, W1∈R C×rC and W2∈R C×rC are the learnable parameters of the spiking MLP with expansion ratio r = 4; RASA is expressed as follows:

[0068] Q s =SN(Conv 1×1 (Q)),K f =Conv 1×1 (K),V f =Conv 1×1 (V) (22)

[0069]

[0070] RASA(Q,K,V)=SN(AttnMap·V f *c2) (24)

[0071] Among them, Q, K, V represent the Query, Key, and Value matrices involved in the self-attention operation, which are essentially the input matrix S5∈R T×C×H×W , Conv 1×1 (·) represents a standard convolution operation with a convolution kernel of 1×1, SN(·) represents a spike neuron layer, · represents matrix multiplication, * represents the multiplication of a matrix and a coefficient, AttnMap is the attention map output in the process, c1 and c2 represent two scaling factors to prevent the gradient vanishing problem in the spike attention process, where:

[0072]

[0073] and f Attn Q s and the average firing rate of the spike attention map Attn, d represents the embedding dimension, H and W represent the spatial height and width of the input, p represents the stride of the convolution used, and D is a hyperparameter representing the maximum integer value of I-LIF spike.

[0074] The cross-scale feature fusion is as follows: first, F5 is upsampled and spatially aligned with the S4 feature matrix, then concatenated and input into MSFF to fuse features of different scales, and then channel-aligned with the S3 feature matrix through LCB:

[0075] Fusion1=LCB(MSFF(Concat(S4,UpSampling(F5)))) (27)

[0076] Then the spatial dimension is aligned with the S3 feature matrix through upsampling and input into MSFF:

[0077] Fusion2=MSFF(Concat(S3,UpSampling(Fusion1))) (28)

[0078] Among them, UpSampling(·) represents the upsampling operation; Concat(·) represents the matrix concatenation operation;

[0079] Then, LCB is used to align the dimensions, and then MSFF is used to fuse the features:

[0080] Fusion3=MSFF(Concat(Fusion1,LCB(Fusion2))) (29)

[0081] Fusion4=MSFF(Concat(UpSampling(F5),LCB(Fusion3))) (30)

[0082] Among them, UpSampling(·) represents the upsampling operation; Concat(·) represents the matrix concatenation operation; MSFF(·) is consistent with the description in step 3.

[0083] Using the detection head and loss function of the YOLOv8 network, Fusion2, Fusion3, and Fusion4 obtained in step 6 are input into the detection head to convert them into specific prediction results, including the coordinates, confidence, and category probability of the bounding box, and then the loss is calculated through the joint loss function; the joint loss function of YOLOv8 is defined as follows:

[0084] Loss = λ1·L BCE +λ2·L CIoU +λ3·L DFL Concat(·) (31)

[0085] Among them, L BCE is the binary cross entropy loss used to calculate the classification loss; L CIoU is the complete intersection-union loss, L DFL is the distribution focus loss, and the two together serve as the positioning loss; λ1, λ2, and λ3 are the weights of the loss function.

[0086] The beneficial effect of the present invention is that the present invention designs an efficient top-attention hybrid feature fusion (TAHFF) module based on the sparsity and stability of high-level pulses, including scale-interaction and cross-scale feature fusion based on self-attention operation. Considering that the shallow pulse signals of the network often lack stability and semantic information, there is a risk of redundancy and confusion when performing scale-interaction on them. Therefore, feature interaction based on attention is only performed on stable and semantically rich high-level features, which reduces energy consumption while giving full play to the global modeling advantage of attention. Subsequently, effective cross-scale feature fusion is performed with low-level features to make up for the bottleneck of Transformer in local feature extraction. In this way, the effective bridging of YOLO and Transformer is completed under the condition of maintaining the low power consumption advantage of SNN, and the SNN target detection task is efficiently completed. In addition, the present invention designs a range-adaptive spiking attention (RASA) for scale-interaction, which achieves the effect of avoiding gradient vanishing and adapting any integer pulse value to perform self-attention operation by adjusting the scaling factor, and enhances the scene perception ability of the model under the condition of maintaining low power consumption, so that the model can effectively detect targets of interest. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1It is a flow chart of the range adaptive SNN target detection method based on YOLO and Transformer bridging of the present invention;

[0088] Figure 2 It is an overall framework diagram of a specific embodiment of the present invention;

[0089] Figure 3 Schematic diagram of the structure of the multi-scale feature fusion module in the present invention;

[0090] Figure 4 It is a schematic diagram of the structure of the Transformer-based SNN block in the present invention;

[0091] Figure 5 The figure is a flow chart of the test steps of a specific embodiment of the present invention. DETAILED DESCRIPTION

[0092] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The present invention includes but is not limited to the following embodiments.

[0093] like Figure 1 As shown, the present invention provides a range-adaptive SNN target detection method based on YOLO and Transformer bridging, and its specific implementation process is as follows:

[0094] 1. Data preprocessing

[0095] The input of the model can be uniformly expressed as Input T×C×H×W , where T represents the time step, C represents the number of channels, and H×W represents the spatial resolution. For static image input, in order to fully utilize the spatiotemporal processing capabilities of SNN, the usual approach is to repeat the static image along the time dimension and use it as input for each time step T. This strategy is called direct input encoding, in which the first layer of spike neurons in the network encodes continuous input values ​​into spike signals; the spike feature matrix Input T ×C×H×W .

[0096] The preprocessing strategy for event-based input is usually to aggregate the event stream within a fixed time window to obtain the pulse feature matrix Input T×C×H×W The event-based input (also called neuromorphic data) is generated by a dynamic vision sensor (DVS). A pixel is detected only when the logarithmic change of the light intensity I(x,y,t) at a pixel exceeds a predefined threshold θ. th At timestamp t n Pixels on (x n ,y n ) will trigger an event e n =(x n ,yn ,t n ,p n ), that is, a peak value is generated. Polarity p n ∈{-1,1} indicates an increase or decrease in light intensity. Event-based input has many advantages, such as low resource requirements, high temporal resolution, and strong robustness. The spike-driven characteristics make SNN very suitable for processing event streams. The preprocessing strategy adopted by event-based input is usually to aggregate the event stream within a fixed time window. Specifically, given a spatiotemporal window ζ, the asynchronous event stream E={e n ∈ζ:n=1,...,N} represents a sparse grid of points in three-dimensional space. E is divided into time intervals using a constant time window dt, and the events are mapped into a two-dimensional representation similar to an image. Each time T fixed time steps are processed, the total sequence Γ=T×dt, where dt and T are the constant time window and time step, respectively. The continuous input values ​​of the total sequence Γ are encoded into pulse signals through pulse neurons to obtain a pulse feature matrix.

[0097] 2. Detailed feature extraction at the lower layer of the backbone network

[0098] First, the pulse feature matrix obtained in step 1 is Figure 2 The stage-1 and stage-2 shown in the figure extract low-level detail features, such as local details such as texture, edge, shape, etc.; the structure of stage-1 and stage-2 is the same, both consisting of a downsampling layer and a SNNBlock. The processing flow of stage-1 is as follows. Pulse matrix Input T×C×H×W First, pass the downsampling layer to get the matrix Reduce the resolution of the feature map while maintaining the expression of key information, so that the feature map is more suitable for subsequent target detection tasks; then Input to SNNBlock for feature extraction. The structure of SNNBlock is as follows:

[0099]

[0100] Feature low-stage1 =Input′+ChannelConv(Input′) (33)

[0101] in, Represents the layer input matrix, + represents the matrix addition operation to complete the residual connection, Input′∈R T×C×H×W Represents the intermediate output matrix after the first residual, Feature low-stage1 ∈R T×C×H×Wrepresents the output matrix of SNNBlock, T represents the time step, C represents the number of channels, and H×W represents the spatial resolution; SepConv(·) is a reverse separable convolution module with a 7×7 convolution kernel for capturing global features, and a 3×3 convolution is added to further perform spatial feature fusion. The SepConv(·) is specifically expressed as:

[0102]

[0103] Among them, Conv pw1 (·) and Conv pw2 (·) is point-wise convolution, Conv dw1 (·) and Conv dw2 (·) is a depthwise convolution, and BN(·) represents batch normalization.

[0104] ChannelConv(·) is a channel mixer that realizes the fusion of information between channels and can be expressed as:

[0105] ChannelConv(Input′)=Conv 3×3 (SN(Conv 3×3 (SN(Input′)))) (35)

[0106] Among them, Conv 3×3 (·) represents a standard convolution operation with a kernel size of 3×3.

[0107] SN(·) represents the spiking neuron layer, using I-LIF spiking neurons, which are calculated as follows:

[0108] U[t]=H[t-1]+X[t] (36)

[0109] H[t]=β(U[t]-S[t]) (37)

[0110] S[t]=Clip(round(U[t]),0,D) (38)

[0111] Where t represents the pulse time step, U[t] represents the membrane potential that integrates the time information H[t-1] of the previous time step t and the spatial information X[t] input at the current time t, S[t] represents the integer pulse matrix, round(·) represents the rounding function, Clip(x,min,max) represents clipping x to the range [min,max], and D is a hyperparameter representing the maximum integer value that can be emitted by the I-LIF neuron. The membrane potential U[t] decays by a factor of β and is reset by subtracting S[t] after emitting a pulse S[t], otherwise H[t] remains unchanged. In the inference stage, the integer pulse values ​​emitted by the I-LIF neuron are converted to binary pulses by the following formula (39), thereby ensuring that the inference stage is pulse-driven.

[0112]

[0113] Where X l [t] represents the input of the lth layer of neurons, and S l [t,d] represents a pulse sequence, which contains only 0 / 1, and W l Represents the coefficient matrix extracted during the expansion process.

[0114] For convenience of representation, the present invention represents the entire I-LIF spiking neuron model as follows:

[0115] S=SN(U) (40)

[0116] Where SN(·) is the spiking neuron layer mentioned above, whose input is the membrane potential tensor U and whose output is the spiking tensor S.

[0117] The structure and processing flow of stage-2 and stage-1 are exactly the same. low-stage1 The low-level pulse feature matrix Feature is formed through stage-2 processing low-stage2 .

[0118] 3. Extract semantic information at the high level of the backbone network

[0119] The low-level pulse feature matrix Feature output by stage-2 low-stage2 Will go through Figure 2 The stage-3 and stage-4 shown in the figure extract high-level semantic information to enhance the network's ability to recognize objects of different sizes and positions. Finally, the pulse feature matrix Feature output by stage-4 is processed through spatial pyramid pooling SPPF. high-stage4 Perform multi-scale spatial pooling processing.

[0120] The structures of stage-3 and stage-4 are the same, both consisting of a downsampling layer and a multi-scale feature fusion module. The processing flow of stage-3 is as follows: Low-level pulse feature matrix Feature low-stage2 After the downsampling layer, the matrix is ​​obtained Then input into the multi-scale feature fusion MSFF module, whose structure is as follows:

[0121]

[0122] Feature high-stage3 =F3′+ChannelConv2(F3′) (42)

[0123] in, Represents the layer input matrix, + represents the matrix addition operation to complete the residual connection, F3′∈R T×C×H×W Represents the intermediate output matrix after the first residual, Feature high-stage3 ∈R T×C×H×W Represents the output matrix of the MSFF module.

[0124] DMSFF(·) is a dilated multi-scale feature fusion module (Dilated Multi-Scale Feature Fusion, DMSFF), which first consists of four parallel dilated group convolutions with different dilation rates. The dilated convolutions with different dilation rates perform more detailed and diversified sampling and integration of features, extract more robust local features, and facilitate the subsequent attention mechanism to perform more targeted global modeling; then the four convolution outputs are channel-connected and channel-downsampled using 1×1 convolution blocks. The DMSFF module can be specifically described as:

[0125]

[0126] in, Indicates that the convolution is an expanded convolution of size 3×3; represents the feature matrix obtained by dilated convolution; d represents the dilation rate, and the dilation rates of the four groups of dilated convolution are 1, 2, 3, and 4 respectively; g represents the group, and here there are 4 groups; c represents the number of channels, SN(·) represents the spike neuron layer, Concat(·) represents the matrix concatenation operation, and Conv 1×1 (·) indicates that the convolution is a standard convolution of size 1×1. Then it is input into SepConv(·) to refine the features. SepConv(·) is consistent with that described in step 2.

[0127] ChannelConv2(·) is used as a channel mixer to achieve information fusion between channels. It uses a re-parameterized convolution with a kernel size of 3×3 to minimize the parameter count, which can be described as:

[0128] ChannelConv2(F′3)=BN(RepConv(SN(BN(RepConv(SN(F′3)))))) (45)

[0129] RepConv(U′)=Conv pw2 (Conv dw1 (Conv pw1 (U′))) (46)

[0130] Among them, RepConv(·) represents a reparameterized convolution with a kernel size of 3×3, which can be reparameterized into a standard convolution during inference, U′∈R T×C×H×W Represents the input matrix of RepConv(·). Conv pw1 (·) and Conv pw2 (·) is point-wise convolution, Conv dw1 (·) is a depthwise convolution, SN(·) denotes a spiking neuron layer, and BN(·) denotes a batch normalization operation.

[0131] The structure and processing flow of stage-4 and stage-3 are exactly the same. high-stage3 The pulse feature matrix Feature is formed through stage-4 processing high-stage4 ;

[0132] Finally, the feature matrix is ​​processed by multi-scale spatial pooling through SPPF to capture feature information of different scales. SSPPF extracts more contextual information at different scales through multiple maximum pooling operations, so that the network can better identify objects of different sizes. Specifically, it can be described as:

[0133] y1=MaxPool(Feature high-stage4 )y2=MaxPool(y1)y3=MaxPool(y2) (47)

[0134] SPPF(Feature high-stage4 )=Conv 1×1 (SN(Concat(Feature high-stage4 ,y1,y2,y3)))(48)

[0135] Among them, Feature high-stage4 ∈R T×C×H×Wrepresents the layer input matrix, MaxPool(·) represents the maximum pooling operation with a kernel size of 5×5, y1, y2, y3 represent the output matrix after the pooling operation, Conv 1×1 (·) represents a standard convolution operation with a convolution kernel of 1×1, SN(·) represents a spike neuron layer, and Concat(·) represents a matrix concatenation operation.

[0136] 4. Feature channel dimension alignment

[0137] The output matrices of stage-2, stage-3, and SPPF are input into the LCB block for channel alignment for subsequent intra-scale interaction and feature fusion. The structure of LCB is as follows;

[0138] LCB(U in )=BN(Conv(SN(U in ))) (49)

[0139] Among them U in ∈R T×C×H×W Represents the layer input matrix, BN(·) represents batch normalization operation, Conv(·) represents standard convolution operation, and SN(·) represents the spike neuron layer. Here, Conv(·) uses a convolution kernel of 1×1. At this time, the LCB block can be regarded as a spike convolution operation with a convolution kernel of 1×1, which is used to adjust the channel dimension of the feature matrix.

[0140] The feature matrices of stage-2, stage-3 and SPPF after LCB alignment are recorded as S3, S4 and S5 respectively, which are used for intra-scale interaction and feature fusion operations in subsequent processes.

[0141] 5. Complete in-scale interaction with low power consumption

[0142] The intra-scale interaction is accomplished by the Transformer-based SNN Block proposed in this invention, which includes a Range-Adaptive Spiking Attention (RASA) and a SpikeMLP, which can give full play to the global modeling capability of Transformer while maintaining the low power consumption advantage of SNN. The Transformer-based SNN Block applies self-attention operation to S5 for intra-scale interaction, effectively capturing the connection between conceptual entities and facilitating the positioning and recognition of objects by subsequent modules.

[0143] The Transformer-based SNN block consists of a Range-AdaptiveSpiking Attention (RASA) and a SpikeMLP, which can be expressed as:

[0144] Attention=S5+RASA(Q,K,V) (60)

[0145] F5=Attention+SpikeMLP(Attention) (51)

[0146] SpikeMLP(Attention)=SN(SN(Attention)W1)W2 (52)

[0147] Among them, S5∈R T×C×H×W Represents the layer input matrix, + represents the matrix addition operation to complete the residual connection, Attention∈R T×C×H×W represents the intermediate output matrix after the first residual, F5∈R T×C×H×W represents the output matrix of Transformer-based SNN Block, SN(·) represents the spike neuron layer, W1∈R C×rC and W2∈R C ×rC is the learnable parameter of the pulse MLP with expansion ratio r=4, RASA is the range-adaptive pulse self-attention proposed in the present invention, which can be expressed as follows:

[0148] Q s =SN(Conv 1×1 (Q)),K f =Conv 1×1 (K),V f =Conv 1×1 (V) (53)

[0149]

[0150] RASA(Q,K,V)=SN(AttnMap·V f *c2) (55)

[0151] Among them, Q, K, V represent the Query, Key, and Value matrices involved in the self-attention operation, which are essentially the input matrix S5∈R T×C×H×W , Conv 1×1(·) represents a standard convolution operation with a convolution kernel of 1×1, SN(·) represents a spike neuron layer, · represents matrix multiplication, * represents the multiplication of a matrix and a coefficient, AttnMap is the attention map output in the process, c1 and c2 represent two scaling factors to prevent the gradient vanishing problem in the spike attention process, where:

[0152]

[0153] here and f Attn Q s and the average trigger rate of the pulse attention map Attn, d represents the embedding dimension, H and W represent the spatial height and width of the input, p represents the stride of the convolution used, and D is a hyperparameter representing the maximum integer value of I-LIF. This hyperparameter ensures that the values ​​after matrix multiplication are properly scaled during training, thereby avoiding gradient vanishing and adapting to any integer pulses.

[0154] 6. Cross-scale feature fusion

[0155] Cross-scale feature fusion is performed on S3, S4 and F5 to make up for the bottleneck of Transformer in local feature extraction. First, F5 is upsampled and spatially aligned with S4, then concatenated and input into MSFF to fuse features of different scales, and then channel-aligned with S3 through LCB:

[0156] Fusion1=LCB(MSFF(Concat(S4,UpSampling(F5)))) (58)

[0157] Then the spatial dimension is aligned with S3 through upsampling and input into MSFF:

[0158] Fusion2=MSFF(Concat(S3,UpSampling(Fusion1))) (59)

[0159] Among them, UpSampling(·) represents the upsampling operation; Concat(·) represents the matrix concatenation operation; MSFF(·) is consistent with the description in step 3, using dilated convolution to expand the effective receptive field with minimal structural changes, while maintaining sparse and low power characteristics while enhancing feature representation capabilities; LCB is consistent with the description in step 4, where a 1×1 convolution kernel is used for channel dimension adjustment;

[0160] Then, we still use LCB to align the dimensions. LCB is the same as described in step 4. Here, we use a convolution with a kernel size of 2×2 to align the channel dimension and the spatial dimension at the same time, and then combine it with MSFF for feature fusion:

[0161] Fusion3=MSFF(Concat(Fusion1,LCB(Fusion2))) (60)

[0162] Fusion4=MSFF(Concat(UpSampling(F5),LCB(Fusion3))) (61)

[0163] Among them, UpSampling(·) represents the upsampling operation; LCB is consistent with the description in step 4, where a 2×2 convolution kernel is used to adjust the channel and spatial dimensions at the same time; Concat(·) represents the matrix concatenation operation; MSFF(·) is consistent with the description in step 3.

[0164] 7. Use the joint loss function to train the network until convergence

[0165] The present invention uses the detection head and loss function of the YOLOv8 network, and converts Fusion2, Fusion3, and Fusion4 into specific prediction results, including the coordinates, confidence, and category probability of the bounding box, and then calculates the loss through the joint loss function. The joint loss function of YOLOv8 is defined as follows:

[0166] Loss = λ1·L BCE +λ2·L CIoU +λ3·L DFL Concat(·) (62)

[0167] Among them, L BCE It is the binary cross-entropy loss (Binary Cross-Entropy Loss) used to calculate the classification loss; L CIoU is the Complete Intersection over Union Loss, L DFL is the distribution focal loss, and the two together serve as the positioning loss; λ1, λ2, λ3 are the weights of the loss function, which default to 0.5, 7.5, and 1.5 in YOLOv8.

[0168] 8. Test set object detection

[0169] like Figure 5As shown, the test image is input into step 1 for preprocessing to obtain an input matrix that can be used for the SNN network, and Fusion2, Fusion3, and Fusion4 are obtained through steps 2-6 and input into the YOLOv8 detection head to convert the feature map into a specific prediction result, including the coordinates, confidence, and category probability of the bounding box. In the output result after completing the target detection, the non-maximum suppression (NMS) algorithm is used for post-processing to remove duplicate detection frames and retain only the optimal target detection frame to complete the target detection task.

[0170] In summary, the present invention discloses a range-adaptive spiking neural network target detection method based on YOLO and Transformer bridging. The present invention focuses on analyzing the pulse characteristics in the network model, and finds that shallow pulse signals often lack stability and semantic information, and the attention operation is too redundant and has the risk of feature confusion, so only the attention-based feature interaction is performed on stable and semantically rich high-level features. Using a single self-attention module, the global modeling advantage of Transformer can be used to effectively capture global dependencies, and the increase in computational costs caused by too many self-attention operations can be avoided, reducing energy consumption. Subsequently, effective cross-scale feature fusion is performed with low-level features, and the local perception characteristics of convolution are used to compensate for the shortcomings of Transformer in local modeling capabilities, and the model expression ability is enhanced, thereby combining the advantages of Transformer and YOLO. In addition, the present invention designs a range-adaptive spiking attention (RASA) for intra-scale interaction. By adjusting the scaling factor, the gradient vanishing phenomenon can be avoided, and any integer pulse value can be adapted for self-attention operations. Effective intra-scale feature interaction is performed on high-level features while maintaining low power consumption, enhancing the scene perception ability of the model, so that the model can effectively detect targets of interest.

Claims

1. A range-adaptive pulse neural network target detection method based on YOLO and Transformer bridge, characterized in that: The following steps are involved: Step 1: Perform data preprocessing on the static image dataset and the dynamic image dataset to obtain the pulse feature matrix Input T×C×H×W , where T represents the time step, C represents the number of channels, and H×W represents the spatial resolution; Step 2: Process the pulse feature matrix Input obtained in step 1 T×C×H×W , extract low-level detail features through the downsampling layer and the pulse-based convolution block SNNBlock to obtain the low-level pulse feature matrix Feature low-stage2 ; Step 3: Process high-order features by using the downsampling layer and the multi-scale feature fusion module MSFF to process the low-level pulse feature matrix Feature obtained in step 2 low-stage2 Extract high-level semantic information and obtain the pulse feature matrix Feature high-stage3 and the matrix of spatial pyramid pooling output; Step 4: Align the feature channel dimension of the low-level detail features in step 2 and the high-level semantic information in step 3, and the low-level pulse feature matrix Feature low-stage2 , pulse feature matrix Feature high-stage3 The matrices output by the spatial pyramid pooling are aligned and are respectively the S3 feature matrix, the S4 feature matrix, and the S5 feature matrix, which are used as feature matrices for subsequent intra-scale interaction and feature fusion; Step 5: Apply the intra-scale interaction based on the self-attention operation to the S5 feature matrix output in step 4; Step 6: Perform cross-scale feature fusion on S3, S4 obtained in step 4 and F5 obtained in step 5; Step 7: Use the joint loss function to train the network until convergence; after convergence, the entire network performs target detection.

2. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 1 is characterized in that: The process of obtaining the pulse feature matrix from the static image data set is as follows: repeating the static image along the time dimension and using it as the input value of each time step T; encoding the continuous input values ​​into pulse signals through pulse neurons to obtain the pulse feature matrix.

3. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 1 is characterized in that: The dynamic image dataset is neuromorphic data, and each data point is an event, including pixel coordinates, timestamp and polarity. The process of obtaining the pulse feature matrix is ​​as follows: aggregate the event stream within a fixed time window into frames; given a spatiotemporal window ζ, the asynchronous event stream E = {e n ∈ζ:n=1,...,N} represents a sparse grid of points in three-dimensional space. A constant time window dt is used to divide E into time intervals, and the events are mapped into a two-dimensional matrix representation of the image. T fixed time steps are processed each time, and the total sequence Γ=T×dt, where dt and T are the constant time window and time step respectively. The continuous input values ​​of the total sequence Γ are encoded into pulse signals through pulse neurons to obtain a pulse feature matrix.

4. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 1 is characterized in that: The pulse feature matrix Input T×C×H×W Preliminary feature extraction is performed in sequence through stage-1 and stage-2; the structures of stage-1 and stage-2 are the same, both consisting of a downsampling layer and an SNNBlock; The stage-1 processing flow is as follows: Pulse feature matrix Input T×C×H×W First, pass the downsampling layer to get the matrix Input into SNNBlock for feature extraction, the structure of SNNBlock is as follows: Feature low-stage1 =Input′+ChannelConv(Input′) (2) Among them, + indicates that the matrix addition operation completes the residual connection, Input′∈R T×C×H×W Represents the intermediate output matrix after the first residual, Feature low-stage1 ∈R T×C×H×W represents the output matrix of SNNBlock, T represents the time step, C represents the number of channels, and H×W represents the spatial resolution; SepConv(·) is a reverse separable convolution module with a 7×7 convolution kernel for capturing global features, and a 3×3 convolution is added to further perform spatial feature fusion; the SepConv(·) is specifically expressed as: Among them, Conv pw1 (·) and Conv pw2 (·) is point-wise convolution, Conv dw1 (·) and Conv dw2 (·) is the depthwise convolution, and BN(·) represents the batch normalization operation; ChannelConv(·) is used as a channel mixer to realize the fusion of information between channels, which is expressed as: ChannelConv(Input′)=Conv 3×3 (SN(Conv 3×3 (SN(Input′)))) (4) Among them, Conv 3×3 (·) represents a standard convolution operation with a kernel size of 3×3; SN(·) represents the spike neuron layer, using I-LIF neurons, and its specific calculation method is as follows: U[t]=H[t-1]+X[t] (5) H[t]=β(U[t]-S[t]) (6) S[t]=Clip(round(U[t]),0,D) (7) Where t is the pulse time step, U[t] is the membrane potential that combines the temporal information H[t-1] of the previous time step t and the spatial information X[t] input at the current time t, S[t] is the integer pulse matrix, round(·) is the rounding function, Clip(x,min,max) means clipping x to the range [min,max], and D is a hyperparameter representing the maximum integer value that an I-LIF neuron can emit; the membrane potential U[t] decays by a factor of β and is reset by subtracting S[t] after the pulse S[t] is emitted, otherwise H[t] remains unchanged; In the inference phase, the integer pulse values ​​emitted by the I-LIF spiking neuron are converted into binary pulses by the following formula (8), thus ensuring that the inference phase is pulse-driven; Where X l [t] represents the input of the lth layer of neurons, and S l [t,d] represents a pulse sequence, which contains only 0 / 1, and W l represents the coefficient matrix extracted during the expansion process; The entire I-LIF spiking neuron model is expressed as: S=SN(U) (9) Where SN(·) is the spiking neuron layer mentioned above, whose input is the membrane potential tensor U and output is the spiking tensor S; The structure and processing flow of stage-2 and stage-1 are exactly the same. low-stage1 The low-level pulse feature matrix Feature is formed through stage-2 processing low-stage2 .

5. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 4 is characterized in that: The low-level pulse feature matrix Feature low-stage2 The multi-dimensional features of the object are captured from different levels through stage-3 and stage-4. The structures of stage-3 and stage-4 are the same, both of which consist of a downsampling layer and a multi-scale feature fusion module. Finally, the pulse feature matrix Feature output by stage-4 is processed by spatial pyramid pooling SPPF. high-stage4 Perform multi-scale spatial pooling processing; The stage-3 processing flow is as follows: Low-level pulse feature matrix Feature low-stage2 First, the matrix F is obtained by downsampling layer processing D 3 S , and then input into the multi-scale feature fusion module, whose structure is as follows: Feature high-stage3 =F3′+ChannelConv2(F3′) (11) Among them, F3′∈R T×C×H×W Represents the intermediate output matrix after the first residual, Feature high-stage3 ∈R T ×C×H×W Represents the output matrix of the multi-scale feature fusion module; DMSFF(·) is a dilated multi-scale feature fusion module, which is mainly composed of four parallel dilated grouped convolutions with different dilation rates. The four dilated grouped convolution outputs are channel-connected and channel-downsampled using 1×1 convolution blocks. The dilated multi-scale feature fusion module is specifically described as: in, Indicates that the convolution is a dilated convolution of size 3×3, represents the feature matrix obtained by dilated convolution, d represents the dilation rate, g represents the group, c represents the number of channels, SN(·) represents the spike neuron layer, Concat(·) represents the matrix concatenation operation, Conv 1×1 (·) indicates that the convolution is a standard convolution of size 1×1; the result of the dilated multi-scale feature fusion module is input into SepConv(·) to refine the features; ChannelConv2(·) is used as a channel mixer to achieve information fusion between channels. It uses a re-parameterized convolution with a kernel size of 3×3 to minimize the parameter count, which is described as: ChannelConv2(F′3)=BN(RepConv(SN(BN(RepConv(SN(F′3)))))) (14) RepConv(U′)=Conv pw2 (Conv dw1 (Conv pw1 (U′))) (15) Among them, RepConv(·) represents a parameterized convolution with a kernel size of 3×3, U′∈R T×C×H×W RepConv(·) represents the input matrix, which is reparameterized as a standard convolution during inference; Conv pw1 (·) and Conv pw2 (·) is point-wise convolution, Conv dw1 (·) is a depthwise convolution, SN(·) represents a spiking neuron layer, and BN(·) represents a batch normalization operation; The structure and processing flow of stage-4 and stage-3 are exactly the same. high-stage3 The pulse feature matrix Feature is formed through stage-4 processing high-stage4 ; The spatial pyramid pooling is described as: y1=MaxPool(Feature high-stage4 ) y2=MaxPool(y1) y3=MaxPool(y2) (16) SPPF(Feature high-stage4 )=Conv 1×1 (SN(Concat(Feature high-stage4 ,y1,y2,y3))) (17) Among them, MaxPool(·) represents the maximum pooling operation with a kernel size of 5×5, y1, y2, and y3 represent the output matrices after the pooling operation, and Conv 1×1 (·) represents a standard convolution operation with a convolution kernel of 1×1, SN(·) represents a spike neuron layer, and Concat(·) represents a matrix concatenation operation.

6. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 5 is characterized in that: The step 4 is specifically as follows: inputting the output of stage-2 in step 2, the output matrix of stage-3 and SPPF in step 3 into the LCB block for channel alignment, and its structure is as follows: LCB(U in )=BN(Conv(SN(U in ))) (18) Among them U in ∈R T×C×H×W represents the layer input matrix, BN(·) represents batch normalization operation, Conv(·) represents standard convolution operation, and SN(·) represents spiking neuron layer; Conv(·) adopts convolution with kernel 1×1, and LCB block is a spiking convolution with kernel 1×1, which is used to adjust the channel dimension of feature matrix; The feature matrices of stage-2, stage-3 and SPPF after LCB alignment are recorded as S3, S4 and S5 respectively, which are used for intra-scale interaction and feature fusion operations in subsequent processes.

7. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 6 is characterized in that: The intra-scale interaction is accomplished through a Transformer-based SNN block; the Transformer-based SNN block includes a range-adaptive spike attention RASA and a SpikeMLP; the Transformer-based SNN block for intra-scale interaction is expressed as: Attention=S5+RASA(Q,K,V) (19) F5=Attention+SpikeMLP(Attention) (20) SpikeMLP(Attention)=SN(SN(Attention)W1)W2 (21) Among them, Attention∈R T×C×H×W represents the intermediate output matrix after the first residual, F5∈R T×C×H×W Represents the output matrix of Transformer-based SNN Block, W1∈R C×rC and W2∈R C×rC are the learnable parameters of the spiking MLP with expansion ratio r = 4; RASA is expressed as follows: Q s =SN(Conv 1×1 (Q)),K f =Conv 1×1 (K),V f =Conv 1×1 (V) (22) RASA(Q,K,V)=SN(AttnMap V f *c2) (24) Among them, Q, K, V represent the Query, Key, and Value matrices involved in the self-attention operation, which are essentially the input matrix S5∈R T ×C×H×W , Conv 1×1 (·) represents a standard convolution operation with a convolution kernel of 1×1, SN(·) represents a spike neuron layer, · represents matrix multiplication, * represents the multiplication of a matrix and a coefficient, AttnMap is the attention map output in the process, c1 and c2 represent two scaling factors to prevent the gradient vanishing problem in the spike attention process, where: f Qs and f Attn Q s and the average firing rate of the spike attention map Attn, d represents the embedding dimension, H and W represent the spatial height and width of the input, p represents the stride of the convolution used, and D is a hyperparameter representing the maximum integer value of I-LIF spike.

8. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 7 is characterized in that: The cross-scale feature fusion is as follows: first, F5 is upsampled and spatially aligned with the S4 feature matrix, then concatenated and input into MSFF to fuse features of different scales, and then channel-aligned with the S3 feature matrix through LCB: Fusion1=LCB(MSFF(Concat(S4,UpSampling(F5)))) (27) Then the spatial dimension is aligned with the S3 feature matrix through upsampling and input into MSFF: Fusion2=MSFF(Concat(S3,UpSampling(Fusion1))) (28) Among them, UpSampling(·) represents the upsampling operation; Concat(·) represents the matrix concatenation operation; Then, LCB is used to align the dimensions, and then MSFF is used to fuse the features: Fusion3=MSFF(Concat(Fusion1,LCB(Fusion2))) (29) Fusion4=MSFF(Concat(UpSampling(F5),LCB(Fusion3))) (30) Among them, UpSampling(·) represents the upsampling operation; Concat(·) represents the matrix concatenation operation; MSFF(·) is consistent with the description in step 3.

9. The range-adaptive pulse neural network target detection method based on YOLO and Transformer bridging according to claim 8 is characterized in that: Using the detection head and loss function of the YOLOv8 network, Fusion2, Fusion3, and Fusion4 obtained in step 6 are input into the detection head to convert them into specific prediction results, including the coordinates, confidence, and category probability of the bounding box, and then the loss is calculated through the joint loss function; The joint loss function of YOLOv8 is defined as follows: Loss=λ1·L BCE +λ2·L CIoU +λ3·L DFL Concat(·) (31) Among them, L BCE is the binary cross entropy loss used to calculate the classification loss; L CIoU is the complete intersection-union loss, L DFL is the distribution focus loss, and the two together serve as the positioning loss; λ1, λ2, and λ3 are the weights of the loss function.

Citation Information

Cited By

  • Intelligent furniture damage detection method

    CN120388364A

  • Visual perception method and device based on spiking neural network, equipment and medium

    CN120808102A

  • Target detection method and system based on spiking neural network

    CN120912850A

  • Channel buoy detection method based on fusion of multi-mode pulse neural network and visual Transform

    CN121616952A

  • Target detection method and system based on pulse neural network

    CN122550916A