High-efficiency pulse target detection method based on neural dynamic enhancement and dynamic pruning
The SpikeYOLO-NDFEB model enhances SNN performance by integrating NDFEB modules and dynamic pruning, addressing feature extraction and resource constraints, enabling efficient and accurate detection of complex events.
Patent Information
- Application Number
- CN202510497321.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
AI Technical Summary
Existing SNN models face challenges in feature extraction and training stability, leading to subpar performance in complex tasks, while ANN models require high computational resources, making them unsuitable for resource-constrained environments.
A high-efficiency pulse detection method using SpikeYOLO-NDFEB, which integrates NDFEB modules for enhanced feature extraction and dynamic pruning, including separable convolutions, temporal aggregation, adaptive membrane potential thresholding, and sparse connections, optimized through SADAP for resource-efficient performance.
The method significantly improves detection performance and efficiency in complex scenarios, providing a reliable solution for real-time monitoring of events like wildfires, drone interference, and ground collapses, while reducing computational complexity and power consumption.
Smart Images

Figure CN120318654A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image processing, and particularly relates to an efficient spiking target detection method based on neural dynamic enhancement and dynamic pruning. Background Art
[0002] With the rapid development of China's economy, remarkable progress has been made in infrastructure construction, which provides an important guarantee for the continuous improvement of people's livelihood and regional economy. Against this background, the construction of public safety infrastructure such as wildfire monitoring, unmanned aerial vehicle (UAV) safety management, and ground subsidence prevention has gradually accelerated and played an important role in comprehensively promoting China's modernization drive. In recent years, the wide application of UAV technology, intelligent transportation systems, and infrastructure construction has entered a stage of rapid development. However, the accompanying frequent occurrence of emergencies has led to an upward trend in social security risks. This has made the efficient monitoring of wildfires, UAV-bird interference, and ground subsidence an important task in the field of public safety.
[0003] Currently, the detection methods for the above-mentioned emergencies mainly include manual detection, traditional image algorithm detection, and deep learning methods. Manual detection has high safety risks, low efficiency, and is easily interfered by human subjective factors; traditional image processing algorithms have solved the detection problem to a certain extent, but their detection accuracy is limited and it is difficult to meet the requirements of real-time processing of large-scale data; although deep learning methods have significantly improved the detection accuracy, due to high power consumption and high requirements for computing resources, they often rely on high-performance hardware support, which severely restricts their application in resource-constrained scenarios and increases the complexity of dealing with emergencies.
[0004] Spiking Neural Networks (SNNs), as a computational model that simulates the activities of biological neurons, have gradually become an important direction in the research of emergency detection due to their unique advantages in spatio-temporal information processing and energy consumption efficiency. Different from traditional Artificial Neural Networks (ANNs) that rely on continuous-value calculations, SNNs transmit and process information through discrete spike signals, and this mechanism makes them outstanding in the efficient processing of time-domain and space-domain information. At the same time, the low-power consumption characteristics of SNNs make them particularly suitable for resource-constrained real-time response scenarios.
[0005] Although spiking neural networks (SNNs) have significant advantages in energy efficiency, existing SNN-based object detection models still have deficiencies in feature extraction capabilities, resulting in inferior performance in complex tasks compared to artificial neural networks. To improve the performance of SNNs, previous studies have attempted to convert from artificial neural networks (ANNs) to SNNs or adopt alternative training methods, such as using surrogate gradient optimization or converting deep ANN models into equivalent SNN architectures. However, these methods usually come with high computational costs and do not fully exploit the potential advantages of SNNs in temporal dynamic processing. At the same time, although modular enhancement strategies (such as separable convolutions, temporal dynamic aggregation, etc.) have been proposed, while improving detection performance, it is still difficult to achieve an ideal balance between computational efficiency and detection accuracy. Summary of the Invention
[0006] An object of the present invention is to address the above deficiencies in the prior art and provide an energy-efficient spiking object detection method based on neural dynamic enhancement and dynamic pruning, so as to solve the problems in the prior art that although ANN technology has high accuracy, it relies on high computing resources and is difficult to be applied in resource-constrained environments; although SNNs have the advantage of low power consumption, there are bottlenecks in feature expression and training stability, and high model complexity and energy consumption problems further limit the application of these technologies in actual scenarios.
[0007] To achieve the above object, the technical solution adopted by the present invention is:
[0008] An energy-efficient spiking object detection method based on neural dynamic enhancement and dynamic pruning, which includes the following steps:
[0009] S1. Collect images of original wildfires, UAV bird interference, and ground collapses and construct an emergency event dataset;
[0010] S2. Construct an object detection model SpikeYOLO-NDFEB;
[0011] S3. Use the emergency event dataset to train the object detection model SpikeYOLO-NDFEB;
[0012] S4. Use adaptive pruning technology to optimize the trained object detection model SpikeYOLO-NDFEB;
[0013] S5. Evaluate the performance of the optimized object detection model SpikeYOLO-NDFEB.
[0014] Further, S2 specifically includes:
[0015] Integrate the NDFEB module into the baseline model SpikeYOLO to construct the object detection model SpikeYOLO-NDFEB; among them, the NDFEB module is located in the backbone network of the baseline model SpikeYOLO.
[0016] Furthermore, the NDFEB module includes a separable convolution module, a temporal dynamic aggregation module, an adaptive membrane potential threshold adjustment module, an adaptive sparse connection module, and a complexity evaluation module;
[0017] Among them, the image feature map is input into the separable convolution module for spatial feature extraction, and the output is the spatial feature map:
[0018] SepConv(x) = PointwiseConv(DepthwiseConv(x))
[0019] In the formula, SepConv is the separable convolution operation; x is the image feature map; DepthwiseConv is the depth convolution; PointwiseConv is the pointwise convolution;
[0020] The spatial feature map is input into the temporal dynamic aggregation module to capture temporal information, and then the aggregated spatio-temporal feature map is output;
[0021] The aggregated spatio-temporal feature map is input into the adaptive membrane potential threshold adjustment module for adaptive membrane potential threshold adjustment, and the enhanced aggregated spatio-temporal feature map is output;
[0022] The enhanced aggregated spatio-temporal feature map is input into the adaptive sparse connection module to generate a sparse mask, and then the optimized aggregated spatio-temporal feature map is output;
[0023] The optimized aggregated spatio-temporal feature map is input into the complexity evaluation module to output the complexity score, and based on this complexity score, the complexity of the input sample is evaluated;
[0024] Among them, the complexity score is calculated as:
[0025] S = Sigmoid(FC(AvgPool(x)))
[0026] In the formula, S is the complexity score; Sigmoid is the activation function; FC is the fully connected layer; AvgPool is the average pooling.
[0027] Furthermore, the adaptive membrane potential threshold adjustment in the adaptive membrane potential threshold adjustment module includes:
[0028] θ(t) = σ(FC(AvgPool(x(t)))) × θ max
[0029] Wherein, θ(t) is the adaptive threshold; σ is the sigmoid activation function; θ max is the maximum threshold.
[0030] Furthermore, a sparse mask is generated in the adaptive sparse connection module, including:
[0031] M = Sigmoid(Conv(x)) > Threshold
[0032] Wherein, M is the sparse mask; Sigmoid is the activation function; Conv is the convolution; Threshold is the sparsity control value.
[0033] Furthermore, the loss function of the object detection model SpikeYOLO-NDFEB is specifically:
[0034]
[0035] Wherein, is the total loss; is the localization loss; is the classification loss; is the distribution focal loss.
[0036] Furthermore, the localization loss is:
[0037]
[0038] Wherein, IoU refers to the intersection over union; b i is the i-th predicted bounding box; b gt i is the ground truth bounding box corresponding to the predicted bounding box b i ;
[0039] The classification loss is:
[0040]
[0041] Wherein, N represents the total number of samples participating in the loss calculation; y i represents the true class label of the i-th sample; represents the predicted class probability corresponding to the i-th sample;
[0042] The distribution focal loss is:
[0043]
[0044] Wherein:
[0045]
[0046] In the formula, DFL is the distribution focal loss; C is the number of distribution intervals; is the true probability of the j-th distribution interval; p j is the predicted probability of the j-th distribution interval; α t and γ are hyperparameters that control the behavior of the focal loss.
[0047] Furthermore, S4 includes the following sub-steps:
[0048] S41. Evaluate the activity levels of neurons and synapses using spike traces, and calculate the synaptic importance based on the spike traces;
[0049]
[0050] In the formula, I ij is the synaptic importance; T is the total number of time steps for statistical spike traces; and are the spike traces of the presynaptic and postsynaptic neurons at time step t, respectively, and θ is the sliding threshold of the postsynaptic neuron;
[0051] S42. Dynamically adjust the pruning rate according to the recall training progress:
[0052] l% = l0 × exp(-λ × e)
[0053] In the formula, l% is the pruning rate; l0 is the initial pruning rate; λ is the decay constant, and e is the current training epoch;
[0054] S43. Set the pruning threshold according to the synaptic importance:
[0055] θ s = quantile(I ij , p)
[0056]
[0057] θ n = quantile(D i , q)
[0058] In the formula, θ s is the synaptic pruning threshold; θ n is the neuron pruning threshold; quantile refers to the quantile; p is the synaptic pruning target; D i is the total synaptic importance of the neuron; q is the neuron pruning target;
[0059] S44. Perform pruning operations on synapses and neurons according to the synaptic pruning threshold and the neuron pruning threshold, and update the synaptic importance and the total synaptic importance of the neurons.
[0060] Furthermore, in S44, the pruning operations of synapses and neurons are performed, including:
[0061] For each synapse from neuron j to neuron i, if its synapse importance I ij < synaptic pruning threshold θ s , then set the synapse weight w ij to 0 and mark it as the pruned state:
[0062]
[0063] where w ij is the synapse weight;
[0064] For each neuron i, if the total synapse importance D of the neuron i < neuron pruning threshold θ n , then set the weights w ij of all its incoming synapses to 0 and mark them as the pruned state;
[0065]
[0066] The synapse importance and the total synapse importance of the neuron are updated using a progressive decay mechanism:
[0067]
[0068] where is the updated synapse importance; is the updated total synapse importance of the neuron; is the current synapse importance; is the current total synapse importance of the neuron; α is the decay factor.
[0069] Furthermore, in S5, on the burst scenario dataset, the model performance of the optimized object detection model SpikeYOLO-NDFEB is evaluated using the mean average precision metric.
[0070] The high-performance spiking object detection method based on neural dynamic enhancement and dynamic pruning provided by the present invention has the following beneficial effects:
[0071] 1. The present invention constructs an object detection model SpikeYOLO-NDFEB, which integrates a neural dynamic enhancement module and an adaptive pruning technique driven by neural spike activity. Through such module design and efficient pruning strategy, the detection performance and computational efficiency of the model in complex scenarios are significantly improved, providing an efficient and reliable solution for applications such as wildfire monitoring, UAV bird interference recognition, and ground subsidence detection.
[0072] 2. The present invention promotes the practical application of SNN in the field of object detection, provides a solution that can simultaneously meet high performance and high precision, and provides a reliable and highly adaptable technical means for scenarios such as wildfire monitoring, UAV bird interference recognition, and ground subsidence detection.
[0073] 3. The present invention enhances the feature extraction and temporal information processing capabilities in complex dynamic scenarios through NDFEB, and accurately detects wildfire spread, bird trajectories, and ground subsidence changes.
[0074] 4. The present invention reduces the computational complexity and power consumption through the SADAP technology, and realizes efficient deployment in resource-constrained environments.
[0075] 5. The object detection model SpikeYOLO-NDFEB of the present invention has better feature expression and training stability, not only improves the detection performance, but also significantly reduces the computational cost, providing strong support for the real-time detection requirements in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a flowchart of the high-performance pulsed object detection method based on neural dynamic enhancement and dynamic pruning of the present invention.
[0077] Figure 2 is a typical example of the burst dataset collected by the present invention.
[0078] Figure 3 is a network structure diagram of the object detection model SpikeYOLO-NDFEB of the present invention.
[0079] Figure 4 is a network structure diagram of NDFEB of the present invention.
[0080] Figure 5 is a visual comparison of the object detection results of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] The following describes the specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the inventive concept of the present invention are within the scope of protection.
[0082] Example 1
[0083] The high-performance pulsed object detection method based on neural dynamic enhancement and dynamic pruning in this embodiment refers to Figure 1 , and specifically includes the following contents:
[0084] Step S1: Collect images of original wildfires, UAV bird interference, and ground collapses, and construct an emergency event dataset;
[0085] Specifically, in this embodiment, for the three scenarios of wildfire monitoring, UAV bird interference recognition, and ground collapse detection, high-quality images containing multiple categories of targets are collected. The data sources include real-time images captured by UAVs, video frames collected by fixed monitoring devices, and publicly available remote sensing image databases.
[0086] Reference Figure 2 , the collected data covers different scene complexities, lighting conditions, and resolutions to ensure that the model can adapt to diverse environments. Preprocess the obtained images of original wildfires, UAV bird interference, and ground collapses. Use the histogram equalization method to improve the contrast of the defective areas and use bilateral filtering to alleviate noise interference; reduce the image resolution through image augmentation methods such as flipping and cropping to obtain more diverse data samples, and complete data annotation on the professional annotation software LabelImg to construct a high-quality emergency event dataset with multiple scenarios, multiple categories, and diverse samples.
[0087] Step S2: Construct the target detection model SpikeYOLO-NDFEB;
[0088] Reference Figure 3 , in this embodiment, based on the baseline model SpikeYOLO, the NDFEB (NeuroDynamic Feature Enhancement Block) module is integrated into it to construct the target detection model SpikeYOLO-NDFEB; the NDFEB module of this embodiment is located in the backbone network of the baseline model SpikeYOLO.
[0089] Reference Figure 4 , the NDFEB module includes a separable convolution module, a temporal dynamic aggregation module, an adaptive membrane potential threshold adjustment module, an adaptive sparse connection module, and a complexity evaluation module, which are responsible for enhancing and optimizing in the spatial and temporal feature processing of the image feature map. The specific workflow is as follows:
[0090] Separable Convolution (SepConv) module;
[0091] The image feature map is input into the separable convolution module for spatial feature extraction. The input image feature map is processed through depthwise convolution and pointwise convolution operations to effectively capture spatial features at different scales and reduce computational complexity, and then output the spatial feature map;
[0092] Among them, the separable convolution operation is:
[0093] SepConv(x)=PointwiseConv(DepthwiseConv(x))
[0094] Where SepConv is a separable convolution operation; x is an image feature map; DepthwiseConv is a depthwise convolution, which is a single convolution filter applied to each input channel; PointwiseConv is a point-by-point convolution, which uses 1×1 convolution for multi-channel combination; this decomposition not only reduces the number of parameters, but also enables the network to learn more complex spatial features.
[0095] Temporal Dynamic Aggregation Modules (TDAM);
[0096] The spatial feature map is input into the temporal dynamic aggregation module to capture temporal information, and then the aggregated spatiotemporal feature map is output; it can promote the integration of temporal information at multiple time steps. This module uses gated recurrent units (GRUs) to aggregate spatiotemporal features, enabling the network to capture motion dynamics and temporal dependencies, which is crucial for accurate target detection in dynamic environments. The calculation process is as follows:
[0097] The calculation formula of the update gate is:
[0098] z t =σ(W z x t +U z h t-1 )
[0099] In the formula, z t Represents the update gate, which is used to control the current hidden state h t From the previous hidden state h t-1 The amount of information introduced; x t is the input of the current time step, h t-1 is the hidden state of the previous time step, W z and U z are the weight matrices of input and hidden states respectively, and σ is the Sigmoid activation function.
[0100] The calculation of the reset gate is:
[0101] r t =σ(W r x t +U r h t-1 )
[0102] In the formula, r t Represents the reset gate, which is used to determine the hidden state h of the previous time stept-1 which information in it needs to be reset; W r and U r are the weight matrices of the input and the hidden state respectively.
[0103] The calculation formula for the candidate hidden state is:
[0104]
[0105] where is the candidate hidden state, representing the input x at the current time step t and the reset hidden state r t ⊙h t-1 is the combined result; ⊙ represents element-wise multiplication, W h and U h are the weight matrices of the input and the hidden state, and tanh is the hyperbolic tangent activation function, which is used to generate non-linear mapping.
[0106] The calculation formula for the final hidden state is:
[0107]
[0108] where h t is the hidden state at the current time step, which is obtained by weighted combination of the hidden state h t-1 from the previous time step and the candidate hidden state according to the weight of the update gate z t obtained.
[0109] By introducing the update gate and the reset gate, the GRU can flexibly control the information memory and forgetting mechanism, ensuring the effective capture and utilization of time information by the network. The update gate z t determines the contribution ratio from the previous time step in the current hidden state, while the reset gate r t adjusts the influence of the previous hidden state in the calculation of the candidate hidden state. The final hidden state h t is the weighted combination of the hidden state h t-1 from the previous time step and the candidate hidden state so as to achieve the gradual aggregation of time information.
[0110] In this embodiment, through the time dynamic aggregation module TDAM, the network can retain and synthesize the key information of multiple time steps, thereby significantly enhancing the time-dependent detection ability for the target. This mechanism is particularly important in a dynamic environment, such as detecting fast-moving objects or processing tasks with strong time correlation.
[0111] Adaptive Membrane Potential Thresholding Module (AMPT);
[0112] In the adaptive membrane potential threshold adjustment module for the aggregated spatio-temporal feature map, adaptive membrane potential threshold adjustment is performed to output an enhanced aggregated spatio-temporal feature map.
[0113] Specifically, in this embodiment, the aggregated spatio-temporal feature map passes through an Adaptive Membrane Potential Thresholding (AMPT) module to adjust the activation threshold of neurons. The AMPT module uses the feature statistical information generated by the pooling layer to dynamically adjust the activation threshold of neurons, thereby enhancing the sensitivity of neurons to important features.
[0114] Among them, the adaptive threshold θ(t) can be defined as:
[0115] θ(t) = σ(FC(AvgPool(x(t)))) × θ max
[0116] In the formula, θ(t) is the adaptive threshold; σ is the sigmoid activation function; θ max is the maximum threshold.
[0117] This embodiment adopts the dynamic adjustment mechanism of the adaptive membrane potential threshold adjustment module to ensure that neurons are more sensitive to significant features, thereby improving the overall detection accuracy.
[0118] Adaptive Sparse Connections (ASC) module;
[0119] The enhanced aggregated spatio-temporal feature map is input into the adaptive sparse connection module, where a sparse mask is generated according to the importance of the feature map, and then an optimized aggregated spatio-temporal feature map is output. By controlling the sparsity of the calculation path, the ASC module reduces redundant calculations and improves the computational efficiency of the network.
[0120] Among them, the generation method of the sparse mask M is as follows:
[0121] M = Sigmoid(Conv(x)) > Threshold
[0122] In the formula, M is the sparse mask; Sigmoid is the activation function; Conv is the convolution; Threshold is the sparsity control value.
[0123] This module ASC performs selective pruning on the connections. This mechanism enables the network to concentrate computational resources on the most informative paths, thereby improving efficiency and performance.
[0124] Complexity Evaluation Module;
[0125] In the optimized aggregated spatio-temporal feature map input complexity evaluation module, an output complexity score is generated. Based on this complexity score, the complexity of the input sample is evaluated, and the behavior of the network is adjusted accordingly. By evaluating the complexity of the features, this module dynamically allocates computing resources, thereby enhancing the adaptability and generalization ability of the network;
[0126] Among them, the complexity score is calculated as:
[0127] S = Sigmoid(FC(AvgPool(x)))
[0128] In the formula, S is the complexity score; Sigmoid is the activation function; FC is the fully connected layer; AvgPool is the average pooling.
[0129] In this embodiment, based on the complexity score S, the network can adjust its computing path, process simple inputs more efficiently, and allocate more resources for complex scenarios at the same time.
[0130] In addition, this embodiment introduces residual connections. The residual connections are introduced into NDFEB to stabilize the training and promote the gradient flow, thereby improving the convergence speed. By providing a direct path for gradient propagation, the residual connections help to alleviate the vanishing gradient problem, making it possible to train deeper and more complex SNN architectures.
[0131] The loss function of the object detection model SpikeYOLO-NDFEB;
[0132] In this embodiment, in order to complete the object detection task in the SNN, a composite loss function is constructed. The design of this composite loss function can effectively balance the bounding box localization, object classification, and the learning of difficult-to-predict samples, and finally improve the detection performance. This composite loss function includes:
[0133]
[0134] In the formula, is the total loss; is the localization loss; is the classification loss; is the distribution focal loss.
[0135] Among them, the localization loss is used to ensure that the predicted bounding box is as aligned as possible with the actual target position. It uses the intersection over union (IoU) to measure the difference between the predicted bounding box and the true annotation box. Specifically, it is:
[0136]
[0137] In the formula, IoU refers to the Intersection over Union, and its calculation formula is:
[0138]
[0139] Among them, Area(b i ∩b gt i ) represents the intersection area of the predicted bounding box b i and the ground truth bounding box b gt i , and Area(b i ∪b gt i ) represents their union area;
[0140] b i represents the i-th predicted bounding box, which is generated by the object detection model based on the input image and is used to describe the position and size of the predicted object in the image;
[0141] b gt i represents the ground truth bounding box corresponding to the predicted bounding box b i , which is provided by manual annotation or a standard dataset and reflects the actual position and size of the object in the image.
[0142] Classification loss is used to penalize misclassified predictions and encourage the network to achieve accurate object classification. It quantifies the errors in object classification detection using binary cross-entropy (BCE) loss, and specifically:
[0143]
[0144] In the formula, N represents the total number of samples participating in the loss calculation; y i represents the true class label of the i-th sample; represents the predicted class probability corresponding to the i-th sample, which is output by the object detection model and ranges from 0 to 1;
[0145] Distribution focal loss improves the network's ability to accurately predict the bounding box distribution, especially for difficult-to-predict samples, and specifically:
[0146]
[0147] Among them:
[0148]
[0149] In the formula, DFL is the Distribution Focal Loss, which is designed to alleviate the imbalance problem between easy and hard samples by modeling the regression distribution of object detection; C is the number of distribution intervals; is the true probability of the j-th distribution interval; p j is the predicted probability of the h-th distribution interval; α t and γ are hyperparameters that control the behavior of the focal loss
[0150] Step S3: Train the object detection model SpikeYOLO-NDFEB using the emergency event dataset;
[0151] Specifically, the model in this embodiment is trained on an NVIDIA RTX 3090 Graphics Processing Unit (GPU), and the batch size is set to 16. The dataset is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. The training is carried out for a total of 300 rounds. The optimization algorithm uses the Adam optimizer, and the initial learning rate is set to 0.001, and the learning rate is decayed to 1 / 10 of the original at the 200th and 250th rounds respectively.
[0152] Step S4: Optimize the trained object detection model SpikeYOLO-NDFEB using the Spiking Activity-Driven Adaptive Pruning (SADAP) technique;
[0153] In this embodiment, the Spiking Activity-Driven Adaptive Pruning (SADAP) technique is introduced, which can reduce the computational overhead of the network without affecting the detection accuracy. Among them, SADAP optimizes the network structure, reduces parameters and energy consumption, and maintains or improves the model performance by measuring the activity levels of neurons and synapses and dynamically pruning inactive synapses and neurons. Its specific implementation includes the following sub-steps:
[0154] Step S41: Measurement of the activity level before pruning; Use the spike trace to evaluate the activity levels of neurons and synapses, and calculate the synapse importance based on the spike trace. The specific content includes the following:
[0155] After the network training in step S3 is completed, it is necessary to evaluate the activity levels of each neuron and synapse. For this purpose, the spike trace S i (t) is introduced to measure the activity level of neuron i at time step t:
[0156] S i (t + 1) = τ s S i (t) + o i (t + 1)
[0157] In the formula, Si (t + 1) is the spike trace of neuron i at time step t + 1; τ s is the decay constant of the spike trace, o i (t + 1) is the spike output (0 or 1) of neuron i at time step t + 1. The spike trace comprehensively considers the spike sequence of the previous period and the firing state at the current moment, reflecting the overall activity level of the neuron over a period of time.
[0158] In addition, based on the Bienenstock-Cooper-Munro (BCM) theory, the synaptic importance I is defined ij as:
[0159]
[0160] In the formula, I ij is the synaptic importance; T is the total number of time steps for statistical spike traces, reflecting the sampling duration; and are the spike traces of the presynaptic and postsynaptic neurons at time step t respectively, and θ is the sliding threshold of the postsynaptic neuron;
[0161]
[0162] In the formula, Num is the number of historical batches; θ(t - 1) is the sliding threshold of the postsynaptic neuron obtained from the previous batch calculation; is the spike trace of the postsynaptic neuron at the current time step t;
[0163] Through the above calculations in this step S41, the importance of each synapse can be quantified for subsequent pruning decisions.
[0164] Step S42: Dynamically adjust the pruning rate according to the training progress to reflect the "fast first and then slow" pruning process;
[0165] The specific adjustment formula for the pruning rate l% is:
[0166] l% = l0 × exp(-λ × e)
[0167] In the formula, l% is the pruning rate; l0 is the initial pruning rate; λ is the decay constant, and e is the current training epoch; as the training epoch increases, the pruning rate gradually decreases to ensure the smoothness and stability of the pruning process. At the same time, set the lower limit l min % (such as 20%) to avoid the pruning effect being not obvious due to too low pruning rate.
[0168] Step S43: Set the pruning threshold according to the synaptic importance:
[0169] Set the pruning threshold θ according to the importance of synapses and neurons s and θ n to determine the synapses and neurons that need to be pruned. The synaptic pruning threshold θ s is determined by the distribution of all synaptic importances I ij . Select the lowest p% of synaptic importances as the pruning target:
[0170] θ s = quantile(I ij , p)
[0171] For neuron pruning, first calculate the total synaptic importance D i of each neuron i:
[0172]
[0173] Then, according to the distribution of D i of all neurons, set the neuron pruning threshold θ n , and select the q% of neurons with the lowest total importance as the pruning target:
[0174] θ n = quantile(D i , q)
[0175] In the formula, θ s is the synaptic pruning threshold; θ n is the neuron pruning threshold; quantile refers to the quantile, which is a statistical method used to determine the value at a given percentage position in the data distribution; p is the synaptic pruning target; D i is the total synaptic importance of the neuron; q is the neuron pruning target.
[0176] Step S44: Perform pruning; perform synaptic and neuron pruning operations according to the synaptic pruning threshold and the neuron pruning threshold, and update the synaptic importance and the total synaptic importance of the neurons. The specific content is as follows:
[0177] For each synapse (j→i), if the synaptic importance I ij < the synaptic pruning threshold θ s , then set the synaptic weight w ij to 0 and mark it as the pruning state:
[0178]
[0179] In the formula, w ij is the synaptic weight;
[0180] For each neuron i, if the total synaptic importance D i<The neuron pruning threshold θ n , then set the weights w ij of all its incoming synapses to 0 and mark them as pruned;
[0181]
[0182] Through the above operations in this step S44, the low-importance synapses and neurons are removed, thereby optimizing the network structure and reducing parameters and energy consumption. To ensure the stability and continuity of the pruning process, a progressive decay mechanism is introduced to gradually decay the importance metrics of the unpruned synapses and neurons. Specifically, for each synaptic importance I ij and the total synaptic importance D i of each neuron, the following updates are made:
[0183]
[0184] where, is the updated synaptic importance; is the updated total synaptic importance of the neuron; is the current synaptic importance; is the current total synaptic importance of the neuron; α is the decay factor, ensuring that the importance metrics gradually decrease and increasing the likelihood of future pruning. This mechanism makes the pruning process smoother, avoids sudden structural changes, and maintains the overall stability of the network.
[0185] Step S5: Evaluate the performance of the optimized object detection model SpikeYOLO-NDFEB;
[0186] Evaluate the performance of SpikeYOLO-NDFEB on the sudden scenario dataset. The performance evaluation uses the mean average precision (mAP) as the metric, where the mean precision at an intersection over union (IoU) threshold of 50% is denoted as mAP 50 and the mean precision from 50% to 95% IoU thresholds is denoted as mAP 50:95 precision. The power consumption of artificial neural networks and spiking neural networks can be calculated as follows:
[0187] E ANN = O 2 × C in × C out × k 2 × E MAC (15)
[0188] E SNN = (T × D) × f r × O2 ×C in ×C out ×k 2 ×E AC (16)
[0189] Where O is the size of the feature output, C in and C out respectively represent the number of input channels and output channels, k is the kernel size, f r represents the average spike firing rate, T is the time step, and D is the upper limit of integer activation during training. All operations are assumed to be implemented in 32-bit floating point on a 45nm technology, where E MAC = 4.6 pJ and E AC = 0.9 pJ.
[0190] In this embodiment, in order to qualitatively evaluate the performance of the object detection model SpikeYOLO-NDFEB, a visualization experiment was conducted to compare the detection results of different models.
[0191] Figure 5 From left to right are: Ground Truth, SpikeYOLO, SpikeYOLO-NDFEB, SpikeYOLO-NDFEB-20% Pruned. The yellow boxes in the figure represent False Positives, and the green boxes represent True Positives; Figure 5 Further shows the bounding box prediction results in the sample images of the baseline models SpikeYOLO, SpikeYOLO-NDFEB and their pruned versions on the burst scene dataset. The results highlight the significant advantages of the object detection model SpikeYOLO-NDFEB in this embodiment in terms of localization accuracy and object classification ability, especially when dealing with small object and dense object scenarios. For example, the object detection model SpikeYOLO-NDFEB can successfully detect multiple small objects missed or misclassified by the baseline models, demonstrating its enhanced feature representation ability and time processing ability. In addition, Figure 5 The rightmost column shows the detection results of the pruned model, further emphasizing the small degree of performance degradation after pruning. Although the model complexity is reduced, the detection accuracy is hardly affected.
[0192] Table 1 Model Performance Test Comparison (T represents the time step, and D represents the expansion factor in each time step)
[0193]
[0194] Table 1 shows the performance comparison of the target detection model SpikeYOLO-NDFEB of this implementation with the baseline SpikeYOLO and other advanced models.
[0195] The results show that the target detection model SpikeYOLO-NDFEB achieves higher detection accuracy while maintaining competitive computational efficiency. Specifically, the target detection model SpikeYOLO-NDFEB has improved in terms of mAP 50 and mAP 50:95 compared to the baseline SpikeYOLO, indicating enhanced localization and classification capabilities. As shown in Table 1, the target detection model SpikeYOLO-NDFEB reaches 90.6% in mAP 50 and 69.2% in mAP 50:95 , which are 1.1% and 1.1% higher than the baseline SpikeYOLO respectively. This improvement highlights the effectiveness of NDFEB in enhancing the feature representation and temporal processing capabilities of SNNs. Compared with other models (such as EMS-YOLO and traditional artificial neural network models PVT and YOLOv5), SpikeYOLO-NDFEB demonstrates a superior balance between accuracy and computational efficiency. The number of parameters increases from 23.1M to 26.5M, and the floating-point operations (FLOPs) increase from 350G to 420G, while the detection performance is significantly improved, further demonstrating the high efficiency of the architecture optimization of the present invention.
[0196] Table 2 Performance comparison of SpikeYOLO-NDFEB at different pruning ratios
[0197]
[0198] Table 2 shows the pruning research results at different pruning ratios on SpikeYOLO-NDFEB. Experiments show that pruning model parameters can significantly reduce power consumption and model size while incurring minimal performance loss. For example, pruning 10% of the parameters only reduces mAP 50 by 0.2% and mAP 50:95 by 0.1%, while the number of parameters is reduced by approximately 8.3% and the FLOPs are reduced by 1.4%. At a 20% pruning ratio, the number of parameters and FLOPs are reduced by approximately 18.5% and 30.0% respectively, and mAP 50 and mAP 50:95 only decrease by 0.5% and 0.2% respectively. Even at a higher 30% pruning ratio, the model still retains 98.5% of the baseline mAP 50 and 99.3% of the baseline mAP 50:95These results demonstrate the robustness of the SADAP technique, which can maintain a high detection accuracy while significantly improving the computational efficiency and reducing the energy consumption. Selecting a pruning ratio of 20% can achieve a balance between reducing the model size and maintaining the detection performance, reflecting the best trade-off between pruning efficiency and accuracy loss.
[0199] Table 3 Comparison of different pruning methods for SpikeYOLO-NDFEB at a pruning ratio of 20%
[0200]
[0201] As shown in Table 3, the SADAP method performs excellently in terms of reducing the model complexity and power consumption while maintaining the detection accuracy, and has more advantages than the magnitude-based pruning and random pruning. Specifically, SADAP achieves an mAP 50 of 90.1% and an mAP 50:95 of 69.0%, which are 2.1% higher than the magnitude-based pruning and 2.6% higher than the random pruning, respectively. In addition, SADAP is also significantly superior to other pruning methods in reducing FLOPs, further highlighting its efficiency.
[0202] Although the specific implementation manners of the invention have been described in detail with reference to the accompanying drawings, it should not be construed as a limitation on the protection scope of this patent. Within the scope described in the claims, various modifications and deformations that can be made by those skilled in the art without creative efforts still fall within the protection scope of this patent.
Claims
1. An efficient pulsed target detection method based on neural dynamic enhancement and dynamic pruning, characterized in that, It includes the following steps: S1. Collect images of original wildfires, UAV bird interference, and ground collapses, and construct an emergency event dataset; S2. Construct a target detection model SpikeYOLO-NDFEB; S3. Use the emergency event dataset to train the target detection model SpikeYOLO-NDFEB; S4. Optimize the trained target detection model SpikeYOLO-NDFEB using an adaptive pruning technique; S5. Evaluate the performance of the optimized target detection model SpikeYOLO-NDFEB.
2. The high-efficiency pulse target detection method based on neural dynamic enhancement and dynamic pruning according to claim 1, characterized in that The specific content of S2 includes: Integrate the NDFEB module into the baseline model SpikeYOLO to construct the target detection model SpikeYOLO-NDFEB; among them, the NDFEB module is located in the backbone network of the baseline model SpikeYOLO.
3. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 2, wherein The NDFEB module includes a separable convolution module, a temporal dynamic aggregation module, an adaptive membrane potential threshold adjustment module, an adaptive sparse connection module, and a complexity evaluation module; Among them, the image feature map is input into the separable convolution module for spatial feature extraction, and the output is the spatial feature map: SepConv(x) = PointwiseConv(DepthwiseConv(x)) In the formula, SepConv is the separable convolution operation; x is the image feature map; DepthwiseConv is the depth convolution; PointwiseConv is the pointwise convolution; The spatial feature map is input into the temporal dynamic aggregation module to capture temporal information, and then the aggregated spatio-temporal feature map is output; The aggregated spatio-temporal feature map is input into the adaptive membrane potential threshold adjustment module for adaptive membrane potential threshold adjustment, and the enhanced aggregated spatio-temporal feature map is output; The enhanced aggregated spatio-temporal feature map is input into the adaptive sparse connection module to generate a sparse mask, and then the optimized aggregated spatio-temporal feature map is output; The optimized aggregated spatio-temporal feature map is input into the complexity evaluation module to output a complexity score, and based on this complexity score, the complexity of the input sample is evaluated; Among them, the complexity score is calculated as: s = Sigmoid(FC(AvgPool(x))) In the formula, S is the complexity score; Sigmoid is the activation function; FC is the fully connected layer; AvgPool is the average pooling.
4. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 3, wherein The adaptive membrane potential threshold adjustment in the adaptive membrane potential threshold adjustment module includes: θ(t) = σ(FC(AvgPool(x(t)))) × θ max where θ(t) is the adaptive threshold; σ is the sigmoid activation function; θ max is the maximum threshold.
5. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 3, wherein The generation of the sparse mask in the adaptive sparse connection module includes: M = Sigmoid(Conv(x)) > Threshold In the formula, M is the sparse mask; Sigmoid is the activation function; Conv is the convolution; Threshold is the sparsity control value.
6. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 1, wherein The loss function of the target detection model SpikeYOLO-NDFEB is specifically: Wherein, is the total loss; is the localization loss; is the classification loss; is the distribution focal loss.
7. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 6, wherein The positioning loss is as follows: Where, IoU refers to the intersection over union; b i is the i-th predicted bounding box; b gt i is the ground truth bounding box corresponding to the predicted bounding box b i ; Classification loss is as follows: where N represents the total number of samples involved in loss calculation; y i represents the true class label of the i-th sample; represents the predicted class probability corresponding to the i-th sample; Distribution focal loss is as follows: Among them: In the formula, DFL is the distribution focal loss; C is the number of distribution intervals; is the true probability of the j-th distribution interval; p j is the predicted probability of the j-th distribution interval; α t and γ are hyperparameters that control the behavior of the focal loss.
8. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 1, wherein S4 includes the following sub-steps: S41. Use the pulse trace to evaluate the activity levels of neurons and synapses, and calculate the synaptic importance based on the pulse trace; Where I ij is the synaptic importance; T is the total number of time steps for statistical pulse traces; and are the pulse traces of the presynaptic and postsynaptic neurons at time step t, respectively, and θ is the sliding threshold of the postsynaptic neuron; S42. Dynamically adjust the pruning rate according to the recalled training progress: l% = l0 × exp(-λ × e) where l% is the pruning rate; l0 is the initial pruning rate; λ is the decay constant, and e is the current training epoch; S43. Set the pruning threshold according to the synapse importance: θ s = quantile(I ij , p) θ n = quantile(D i , q) where θ s is the synaptic pruning threshold; θ n is the neuron pruning threshold; quantile is the quantile; p is the synaptic pruning target; D i is the total synaptic importance of the neuron; q is the neuron pruning target; S44. Perform synapse and neuron pruning operations according to the synapse pruning threshold and neuron pruning threshold, and update the synapse importance and the total synapse importance of neurons.
9. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 8, characterized in that, In the said S44, performing synapse and neuron pruning operations includes: For each synapse, from neuron j to neuron i, if its synapse importance I ij <is less than the synapse pruning threshold θ s , then set the synapse weight w ij to 0 and mark it as the pruned state: where w ij is the synaptic weight; For each neuron i, if the total synaptic importance D of the neuron i <neuron pruning threshold θ n , then set the weights w of all its incoming synapses ij to 0 and mark them as pruned; Adopt a progressive decay mechanism to update the synapse importance and the total synapse importance of neurons: wherein, is the updated synaptic importance; is the total synaptic importance of the updated neuron; is the current synaptic importance; is the total synaptic importance of the current neuron; α is the decay factor.
10. The high-efficiency pulsed target detection method based on neural dynamic enhancement and dynamic pruning according to claim 1, characterized in that In the said S5, on the burst scenario dataset, use the average precision metric to evaluate the model performance of the optimized object detection model SpikeYOLO-NDFEB.