Unmanned aerial vehicle sudden accident rapid classification method based on spiking neural network
By adopting the pulse neural network classification model SpikeClassifier based on the leak integral-issuance model in the UAV emergency classification task, the calculation delay and energy consumption problems of existing SNNs on edge devices are solved, and efficient and accurate classification of emergencies is achieved.
Patent Information
- Application Number
- CN202510056842.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-27
AI Technical Summary
Existing pulsed neural networks (SNNs) still face the problems of computational delay and excessive energy consumption in practical applications, especially on resource-constrained edge smart devices.
SpikeClassifier, a pulsed neural network classification model based on the leak integral-issuance model, uses a network structure that includes deep convolutional layer, point-by-point convolutional layer and maximum pooling layer, and iterates 25 time steps during the forward propagation process to reduce the number of time steps of the model and improve event driving, reducing calculation complexity and power consumption.
It realizes the accuracy of classification under complex backgrounds and resource constraints, and improves the response speed and efficiency of the model, adapts to resource limitations of edge devices, and improves the real-time and adaptability of classification.
Smart Images

Figure CN120047727A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of UAV image recognition and classification, and particularly relates to a method for rapid classification of UAV emergency accidents based on a spiking neural network. Background Art
[0002] With the rapid development of UAV technology, UAVs have been widely used in fields such as agriculture, environmental monitoring, and logistics distribution. However, during the operation of UAVs, low-flying birds pose a potential threat to their safety. The collision between birds and UAVs may not only cause damage to the UAV and affect the normal execution of flight tasks, but also endanger the safety of ground personnel and property. In addition, in emergency scenarios such as wildfire monitoring and ground defect detection, rapid and accurate classification capabilities are also required. Early detection and timely warning of wildfires can effectively reduce the losses caused by fires; while the detection of ground defects helps to prevent potential safety hazards caused by infrastructure failures. In these situations, as a flexible monitoring tool, UAVs can quickly reach inaccessible locations for real-time monitoring and data collection. However, in order to enable UAVs to effectively respond to these emergencies, the problem of efficient classification under resource-constrained conditions must be solved.
[0003] Traditional computer vision methods rely on manually designed features, such as Support Vector Machine (SVM), Random Forest (RF), etc. Although they perform well in certain specific tasks, their ability to process images in complex backgrounds is limited. In contrast, deep learning, especially Artificial Neural Networks (ANNs), has shown significant advantages in the field of image processing due to its strong feature learning ability. Convolutional Neural Networks (CNNs), as a type of ANN, can automatically learn multi-level feature representations from raw images, greatly improving the accuracy and robustness of image processing. However, with the continuous expansion of the scale of model parameters, ANNs require more computational energy consumption to support the implementation of more complex application tasks, which undoubtedly limits the application and popularization of edge intelligent devices. Although researchers have carried out studies on model compression technologies such as network quantization, structure pruning, and knowledge distillation, the power consumption problem remains very prominent.
[0004] Spiking Neural Networks (SNNs) are the third-generation neural network models, which achieve efficient time-series data processing by simulating the pulse signal transmission of biological neurons. Compared with traditional ANNs, SNNs use single-bit pulses, reducing floating-point multiplication operations and thus lowering energy consumption. Spiking neurons can accumulate temporal information and trigger pulses when the threshold is exceeded. This feature enables SNNs to perform excellently in terms of computational efficiency and energy efficiency and achieve comparable accuracy to ANNs in many tasks. However, existing SNN training methods usually require converting an ANN model into its corresponding SNN model first. This method requires a long number of time steps when dealing with complex tasks, resulting in high energy consumption, which is contrary to the purpose of low energy consumption. Therefore, the existing technology proposes the Spatio-temporal Backpropagation (STBP) method for directly training SNNs, which reduces the required number of time steps through fine-tuning, but there are still problems of delay and energy consumption, especially in low-power applications on edge devices. Summary of the Invention
[0005] The purpose of the present invention is to address the above deficiencies in the prior art and provide a method for rapid classification of UAV emergency accidents based on spiking neural networks, so as to solve the problem that although existing spiking neural networks (SNNs) can theoretically provide high-energy-efficiency computing, they still face problems of high computational delay and energy consumption in practical applications because they require multiple time steps to achieve better computational accuracy.
[0006] To achieve the above object, the technical solution adopted by the present invention is:
[0007] A method for rapid classification of UAV emergency accidents based on spiking neural networks, which includes the following steps:
[0008] S1. Obtain the initial dataset of UAV emergency accidents;
[0009] S2. Preprocess the initial dataset to obtain the UAV emergency accident dataset;
[0010] S3. Select the leaky integrate-and-fire model as the spiking neuron of the spiking neural network;
[0011] S4. Construct a spiking neural network classification model SpikeClassifier based on the leaky integrate-and-fire model;
[0012] S5. Train the spiking neural network classification model SpikeClassifier using the UAV emergency accident dataset to obtain the optimized spiking neural network classification model SpikeClassifier;
[0013] S6. Use the optimized spiking neural network classification model SpikeClassifier to predict and classify UAV emergencies.
[0014] Furthermore, the initial data set in S1 includes positive samples and negative samples;
[0015] The positive samples include dan-bird, qun-bird, fire, crack, stoma, spalling, rust;
[0016] Among them, dan-bird is a single bird image; qun-bird is a flock of bird images; fire is the flame and smoke images during the initial stage to the spreading process of wildfires; crack, stoma, spalling, rust are ground crack images, ground pore images, ground spalling images, and ground rust images respectively;
[0017] The negative samples are ordinary scene images of emergencies, including sky, trees, and buildings.
[0018] Furthermore, S2 specifically includes:
[0019] Clean, augment, adjust the contrast and brightness of the initial data set to obtain a UAV emergency data set, and divide the UAV emergency data set into a training set and a validation set according to a ratio of 8:2.
[0020] Furthermore, in S3, the leaky integrate-and-fire model, as the dynamic equation of the spiking neuron of the spiking neural network, is:
[0021]
[0022] Where is the membrane potential state of the i-th spiking neuron in the n-th layer before generating a spike at time step t, V i n (t - 1) is the membrane potential state of the i-th spiking neuron in the n-th layer at the previous moment (t - 1), is the output current triggered by the presynaptic neuron at the same moment; τ is the membrane time constant;
[0023] When the membrane potential of the spiking neuron reaches the threshold θ, a spike output S(t) is generated, and the membrane potential is reset to V(t):
[0024] S(t) = Θ(H(t) - θ)
[0025]
[0026] V(t) = r(H(t), S(t))
[0027] Among them, Θ() is the Heaviside step function, H(t) is the membrane potential state before the pulse is generated at the time step t; r() is the membrane potential reset function.
[0028] Furthermore, the membrane potential reset function includes hard reset and soft reset;
[0029] The hard reset resets the membrane potential of the spiking neuron to a fixed resting potential V rest ;
[0030] The soft reset reduces the membrane potential by the same magnitude as the firing threshold:
[0031] Furthermore, in the S4, the number of time steps iterated during the forward propagation of the spiking neural network classification model SpikeClassifier is 25.
[0032] Furthermore, in the S4, the network structure of the spiking neural network classification model SpikeClassifier includes:
[0033] The first deep convolutional layer, followed by the first pointwise convolutional layer and the first max pooling layer in sequence after the first deep convolutional layer;
[0034] After the first max pooling layer, it is connected to the second deep convolutional layer, and the output end of the second deep convolutional layer is connected to the second pointwise convolutional layer and the second max pooling layer in sequence;
[0035] After the second max pooling layer, it is connected to the first fully connected layer, the output end of the first fully connected layer is connected to a spiking neuron layer, and after this spiking neuron layer, it is connected to the second fully connected layer; after the second fully connected layer, it is connected to another spiking neuron layer.
[0036] Furthermore, the forward propagation process of the spiking neural network classification model SpikeClassifier includes:
[0037] The input image data is input into the first deep convolutional layer, and after passing through the first pointwise convolutional layer, the first max pooling layer, the second deep convolutional layer, the second pointwise convolutional layer and the second max pooling layer in sequence, the feature map is flattened into a one-dimensional vector;
[0038] The one-dimensional vector is input into the first fully connected layer to compress the flattened one-dimensional vector into a feature vector; the feature vector passes through a spiking neuron layer to output a time-varying dynamic feature vector; the dynamic feature vector is input into the second fully connected layer to map the dynamic feature vector to the classification space and output 8 categories corresponding to seven positive samples and one negative sample; finally, after passing through another spiking neuron layer, it outputs a time-varying static feature.
[0039] The rapid classification method for UAV emergencies based on spiking neural networks provided by the present invention has the following beneficial effects:
[0040] 1. The present invention provides a simple and efficient spiking neural network classification method, which can maintain the classification accuracy and improve the response speed and efficiency of the model under complex backgrounds and resource-constrained conditions.
[0041] 2. The present invention ensures low computational complexity and low power consumption through the event-driven and sparse nature of spiking neural networks, enabling it to adapt to the resource limitations of edge devices.
[0042] 3. By reducing the number of time steps of the model and the event-driven nature of spiking neural networks, the present invention improves the real-time performance of classification, enabling rapid response in the face of emergencies.
[0043] 4. Simpler model structure: Compared with traditional deep learning models, the spiking neural network model SpikeClassifier proposed by the present invention has a simple structure, reduces the number of parameters, thereby reducing the training difficulty and the complexity of the model, and facilitating the maintenance and update of the model.
[0044] 5. Stronger adaptability: The present invention is not only applicable to the classification of flying birds, wildfires, and ground defects, but can also be applied to other scenarios, with strong generalization ability.
[0045] 6. Easy to integrate: Due to the lightweight and low-power characteristics of the model, the present invention is easy to integrate into existing edge devices, and the function can be enhanced without large-scale hardware upgrades. Description of the Drawings
[0046] Figure 1 It is a flowchart of the rapid classification method for UAV emergencies based on spiking neural networks of the present invention.
[0047] Figure 2 It is a typical example diagram of sample data of the present invention; among them, (1) is the single-bird category, (2) is the flock-bird category, (3) is the fire category, (4) is the negative sample (no-bird-and-no-fire) category without birds or fires, (5) is the crack category in ground defects, (6) is the rust category in ground defects, (7) is the spalling category in ground defects, and (8) is the stoma category in ground defects.
[0048] Figure 3 It is the network structure diagram of the spiking neural network classification model SpikeClassifier of the present invention.
[0049] Figure 4 This is the graph of the change in the training accuracy of the SpikeClassifier, the pulse neural network classification model of the present invention.
[0050] Figure 5 This is the confusion matrix graph of the SpikeClassifier, the pulse neural network classification model of the present invention.
[0051] Figure 6 This is the prediction result graph of the SpikeClassifier, the pulse neural network classification model of the present invention. Detailed implementation manners
[0052] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0053] Example 1
[0054] A method for rapid classification of UAV emergency accidents based on a pulse neural network in this embodiment can solve the following technical problems:
[0055] Traditional computer vision methods have limitations in dealing with image classification tasks in complex backgrounds. Although deep learning methods have powerful feature learning capabilities, their large model parameter scale leads to high computational energy consumption, which is particularly obvious on resource-constrained edge intelligent devices. In addition, although existing spiking neural networks (SNNs) can theoretically provide high-energy-efficiency computing, due to the need for multiple time steps to achieve better operation accuracy, they still face problems of high computational latency and energy consumption in practical applications.
[0056] Based on this, referring to Figure 1 , this embodiment specifically includes the following content:
[0057] Step S1: Obtain the initial data set of UAV emergency accidents;
[0058] The data collection in this embodiment specifically includes:
[0059] Bird data, using a high-definition camera installed on the UAV to take bird images at different times and locations, and record their flight postures and backgrounds. These images cover the flight states of different species of birds under various weather conditions, and are divided into single-bird (named dan-bird) category and flock-bird (named qun-bird) category to ensure the diversity and representativeness of the data set.
[0060] For wildfire data, using publicly available datasets, images of flames and smoke during the initial stage to the spreading process of wildfires (named fire) are collected. These datasets contain wildfire cases in different regions and time periods, ensuring the extensiveness and timeliness of the datasets.
[0061] For ground subsidence data, the road surface is photographed by cameras. These images cover different types of ground subsidence, including cracks (named crack), pores (named stoma), spalling (named spalling), and rust (named rust).
[0062] In addition, a large number of ordinary scene images without emergencies are collected as negative samples (named no-bird-and-no-fire). These images include common backgrounds such as the sky, trees, and buildings in daily environments, which are used to improve the classification accuracy of the model in complex backgrounds.
[0063] Based on this, the initial dataset of this embodiment includes positive samples and negative samples;
[0064] The positive samples include dan-bird, qun-bird, fire, crack, stoma, spalling, rust;
[0065] Among them, dan-bird is a single-bird image; qun-bird is a group-bird image; fire is an image of flames and smoke during the initial stage to the spreading process of wildfires; crack, stoma, spalling, rust are ground crack images, ground pore images, ground spalling images, and ground rust images respectively;
[0066] The negative sample no-bird-and-no-fire is an ordinary scene image of no emergency, including the sky, trees, and buildings;
[0067] This embodiment has a total of 7 positive samples and 1 negative sample, with a total of 8000 images.
[0068] Step S2: Preprocess the initial dataset to obtain a UAV emergency accident dataset;
[0069] Since there are certain quality defects in the 7 positive samples and 1 negative sample in the initial dataset, in order to ensure the training effect of the subsequent model, the following processing is performed on the initial dataset:
[0070] First, clean the initial data in the initial dataset, removing blurred, duplicate, and images that do not meet the classification criteria to ensure the quality of the dataset. At the same time, to improve the generalization ability and robustness of the model, data augmentation is performed on the positive sample data. The augmentation methods include image rotation (randomly rotating the image by 0°, 90°, 180°, and 270°), image translation (randomly translating in the horizontal and vertical directions), image scaling (randomly scaling the image while maintaining the ratio), and image flipping (horizontal and vertical flipping of the image). On the basis of data augmentation, image enhancement techniques are further applied, including contrast adjustment, brightness change, etc., to enhance the visual features of the images. These techniques help the model learn more rich features and contribute to improving the classification accuracy.
[0071] After the preprocessing of the dataset is completed, a high-quality drone accident dataset with a total of 12,698 positive and negative samples is obtained, as specifically Figure 2 shown. When training the model, this drone accident dataset is divided into a training set and a validation set. Considering the total scale of the dataset and the training requirements of the model, a ratio of 8:2 is used for the division. The training set is used for the training and optimization of the model, and the validation set is used to evaluate the performance and generalization ability of the model to ensure that each category has sufficient samples in the training set and the validation set, avoiding classification bias caused by sample imbalance. Through the above work, a training set and a validation set with a reasonable structure, balanced categories, and rich features are finally obtained, laying a solid foundation for the subsequent model training.
[0072] Step S3: Select the leaky integrate-and-fire model as the spiking neuron of the spiking neural network;
[0073] In this embodiment, the leaky integrate-and-fire (LIF) model spiking neuron (hereinafter referred to as the LIF neuron model) is selected. The LIF neuron model receives the signals transmitted by the presynaptic neurons, thereby accumulating the membrane potential, that is, integrating; when the membrane potential reaches a specific threshold θ, the LIF neuron fires to generate an output pulse and reset the membrane potential; when no external signal is received, its membrane potential gradually leaks to the resting potential. Its membrane potential V is:
[0074]
[0075] Among them, τ m is the membrane potential time constant, which controls the leakage speed of the LIF neuron membrane potential, V rest is the resting potential, R mis a constant; I is the input current, which is the direct stimulation of the activity of the presynaptic neuron on the current neuron; in computer implementation, usually let V rest = 0 and incorporate R m into the learnable parameters of the neural network. Therefore, Equation (1) is further simplified to:
[0076]
[0077] Using Euler's equation to solve the above first-order differential equation, the dynamic equation of the iterative LIF spiking neuron model for discrete time steps can be obtained:
[0078]
[0079] is the membrane potential state of the i-th spiking neuron in the n-th layer before generating a spike at time step t, and V i n (t - 1) is the membrane potential state of the i-th spiking neuron in the n-th layer at the previous moment (t - 1), is the output current triggered by the presynaptic neuron at the same moment; τ is the membrane time constant, which is a key parameter in the leaky integrate-and-fire model and is used to control the decay rate of the membrane potential, determining the sensitivity of the neuron to time changes.
[0080] When the membrane potential of the LIF neuron reaches the threshold θ, a spike will be generated according to Equation (4). Subsequently, its membrane potential is reset according to Equation (5):
[0081] S(t) = Θ(H(t) - θ) (4)
[0082]
[0083] V(t) = r(H(t), S(t)) (6)
[0084] where S is the generated spike output, Θ() is the Heaviside step function as shown in Equation (5), and r() is the membrane potential reset function, which usually has two ways: hard reset and soft reset. The hard reset resets the membrane potential of the spiking neuron to a fixed value, usually preset to the resting potential V rest , and the soft reset reduces the membrane potential by the same amplitude as the firing threshold. The reset method selected in the present invention is the hard reset.
[0085] Step S4: Construct a spiking neural network classification model SpikeClassifier based on the leaky integrate-and-fire model;
[0086] The time dynamics configuration of the spiking neural network classification model SpikeClassifier in this embodiment is as follows:
[0087] In this embodiment, the number of time steps iterated by the spiking neural network classification model SpikeClassifier during forward propagation is defined as 25, which means that the model will perform 25 time-step iterations on each input sample to fully capture the feature information in the input data; in terms of the settings of the LIF spiking neurons, the decay coefficient of the LIF spiking neurons is set to 0.95. This parameter determines the degree of memory of the neurons between time steps and helps to accumulate input signals when processing temporal information.
[0088] Reference Figure 3 , the network structure of the spiking neural network classification model SpikeClassifier in this embodiment includes:
[0089] The size of the input sample image is 3×224×224;
[0090] The first depth convolutional layer has 3 input channels and 3 output channels, with a convolutional kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 3. The first depth convolutional layer independently extracts features on each input channel.
[0091] After the first depth convolutional layer, the first pointwise convolutional layer and the first max pooling layer are connected in sequence;
[0092] The first pointwise convolutional layer has 3 input channels and 16 output channels, with a convolutional kernel size of 1×1. This layer is responsible for fusing features from different channels. After pointwise convolution, the first max pooling layer is used for spatial dimensionality reduction, with a kernel size of 2×2 and a stride of 2.
[0093] After the first max pooling layer, the second depth convolutional layer is connected. The output end of the second depth convolutional layer is followed by the second pointwise convolutional layer and the second max pooling layer in sequence.
[0094] Among them, the input channels of the depth convolution of the second depth convolutional layer are 16, and the output channels are also 16, with a convolutional kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 16; immediately following is the second pointwise convolution, with 16 input channels and 32 output channels, and a convolutional kernel size of 1×1; then the second max pooling layer is used again for spatial dimensionality reduction, with a kernel size of 2×2 and a stride of 2.
[0095] After the second max pooling layer, the first fully connected layer is connected. The output end of the first fully connected layer is connected to a spiking neuron layer, and after this spiking neuron layer, the second fully connected layer is connected; after the second fully connected layer, another spiking neuron layer is connected.
[0096] After the convolutional layer finishes processing, the feature map is flattened into a one-dimensional vector and input into the first fully connected layer. The input size of this layer is 32×56×56, and the output size is 256. After the first fully connected layer, there is a layer of LIF neurons, which is used to simulate the behavior of spiking neurons. The decay coefficient of this layer is set to 0.95, introducing temporal dynamics to the output of the fully connected layer and allowing the network to accumulate information in the time dimension, thereby enhancing the model's ability to capture temporal features; after the LIF neuron layer is the second fully connected layer, with an input size of 256 and an output size of 8 (the number of classes in the classification task, including seven positive samples and one negative sample). After the second fully connected layer, there is again another layer of LIF neurons, with the same decay coefficient of 0.95. This layer also aims to strengthen the network's temporal processing ability, so that the final output not only considers static features but can also effectively reflect the time evolution of the input sequence.
[0097] The forward propagation process of the spiking neural network classification model SpikeClassifier in this embodiment is as follows:
[0098] The input image data is input into the first depth convolutional layer and sequentially passes through the first pointwise convolutional layer, ReLU activation function ( Figure 3 not shown in the figure), the first max pooling layer, the second depth convolutional layer, the second pointwise convolutional layer, ReLU activation function ( Figure 3 not shown in the figure), and the second max pooling layer. Then, the feature map is flattened into a one-dimensional vector;
[0099] The one-dimensional vector is input into the first fully connected layer to compress the flattened one-dimensional vector into a smaller feature vector; the feature vector passes through a layer of spiking neurons, introducing the time dimension, enabling the network to process temporal information, not only paying attention to static features but also capturing the dynamic features of the input changing over time, and then outputting a dynamic feature vector that changes over time; the dynamic feature vector is input into the second fully connected layer to map the dynamic feature vector to the classification space, outputting 8 classes corresponding to seven positive samples and one negative sample; finally, it passes through another layer of spiking neurons to further strengthen the processing of temporal information, thereby improving the classification accuracy and outputting static features that change over time.
[0100] Step S5: Use the UAV accident dataset to train the spiking neural network classification model SpikeClassifier to obtain the optimized spiking neural network classification model SpikeClassifier. The model training and optimization process in this embodiment is as follows:
[0101] Training configuration;
[0102] The cross-entropy loss function (CrossEntropyLoss) is used as the loss function of the spiking neural network classification model SpikeClassifier. The Adam optimizer is used to optimize the weight parameters of the network. The learning rate is set to 5e-4, and the momentum parameters are 0.9 and 0.999 respectively. The epoch is set to 300, and the batch size is set to 4. The training and testing of the model are run on a device equipped with a GEFORCE RTX 4060 GPU, making full use of the computing power of the GPU to accelerate the training and inference processes of the deep learning model.
[0103] Training data and batch loading;
[0104] In the present invention, the batch loading method is adopted. After dividing the dataset into a training set and a test set through step S2, in each training iteration, a data batch is obtained from the training set, and the input data and the corresponding target labels are passed to the model for forward propagation.
[0105] Forward propagation;
[0106] Each input data sample undergoes 25 time-step iterations. The model calculates the spike recordings and membrane potential recordings of neurons at each time step. The output membrane potential at each time step is processed through the LIF spiking neuron model. The outputs of 25 time steps are accumulated and compared with the target labels to calculate the total loss.
[0107] Backward propagation and optimization;
[0108] After the forward propagation of each batch ends, the model calculates the gradients through backward propagation and updates the weight parameters of the model according to the gradient information of the loss function. Through the Adam optimizer, the model parameters are adaptively updated to gradually reduce the training error.
[0109] Recording of training loss and accuracy;
[0110] During the training process, the training loss of the model is recorded for each epoch, and the training performance of the model is evaluated by calculating the accuracy of the output of each batch compared with the target labels. For this purpose, in the present invention, after each time step, the spike firings of the neurons are accumulated to obtain the final classification result, which is compared with the labels to calculate the classification accuracy.
[0111] Test set evaluation;
[0112] At the end of each epoch, the model is evaluated using the test set, and the loss and classification accuracy of the model on the test set are calculated. During the test process, the spike trains are also accumulated through the output results of 25 time steps to obtain the final classification result, which is then compared with the true label to calculate the accuracy.
[0113] Model saving;
[0114] Whenever the classification accuracy of the model on the test set exceeds the previous best result, the parameters of the current model are saved to ensure that the model with the best performance is finally deployed.
[0115] Visualization of the training process;
[0116] Reference Figure 4 . By recording and visualizing the accuracy curves of training and testing, the present invention demonstrates the performance of the model during the training process. As the number of training iterations increases, the training and testing accuracies of the model gradually improve, indicating that the model effectively learns the features of the input data and has good generalization ability.
[0117] Through this step, the spiking neural network classification model SpikeClassifier has undergone comprehensive training and optimization within 300 training epochs. Through the gradual update of the cross-entropy loss function and the Adam optimizer, the model weights have been effectively adjusted, gradually reducing the training error. Finally, through this training process, the present invention obtains a fully optimized spiking neural network classification model SpikeClassifier, which has the ability to efficiently and accurately classify emergency events.
[0118] Step S6: Use the optimized spiking neural network classification model SpikeClassifier to predict and classify UAV emergency accidents.
[0119] To verify the performance of the spiking neural network classification model SpikeClassifier, it was compared in detail with the current mainstream neural network models. A new and highly competitive spiking neural network classification model Meta-Spikeformer, as well as five mainstream artificial neural network models, including VGG16, PVTv2_B1, FocalNet_S, PatchConvent_S60, and UniFormer_B, were selected as the comparison benchmarks. The comparison metrics covered the classification accuracy, the number of parameters (M), and the inference speed (FPS) of the models, aiming to comprehensively evaluate the performance of each model from multiple dimensions. Table 1 shows the specific performance of each model in these metrics. To analyze more deeply the performance of the spiking neural network classification model SpikeClassifier in actual classification tasks, the confusion matrix of this model was also calculated, presenting in detail its classification performance on different categories. The confusion matrix of SpikeClassifier is as Figure 5 shown. In addition, Figure 6 shows the classification prediction results of this model during the inference process.
[0120] Refer to Figure 5, The confusion matrix of the spiking neural network classification model SpikeClassifier shows the classification performance of the model on 8 categories, generally reflecting the high accuracy of the model. Most samples are correctly classified on the diagonal, demonstrating the excellent classification ability of the model on most categories. Specifically, 98 samples in the "Fire" category are correctly classified, with only a small number of misclassifications, indicating the high accuracy of the model in the sudden fire scenario. The "no-fire-and-no-bird" category performs extremely well, with only 3 samples slightly misclassified out of 111 samples, suggesting a very high precision of the model in identifying no-event scenarios. In the bird recognition task, 111 samples in the "dan-bird" category are correctly classified, with a very good classification effect, and only a small number of samples are misjudged as "qun-bird", which may be due to the relatively similar visual features of the two types of birds. Particularly, all 127 samples in the "qun-bird" category are correctly classified without any misclassifications, reflecting the extremely high accuracy of the model in bird flock recognition. For the "crack" category, 67 samples are correctly classified. Although there are some misclassified samples, the overall accuracy is still relatively high, indicating that the model has a good performance in crack recognition. The performance of the "rust" category is also quite stable, with 72 samples correctly classified. Despite a small number of misclassifications into similar categories, the overall accuracy rate still remains at a relatively high level. In the "spalling" and "stoma" categories, the classification performance of the model is also relatively precise, with 67 and 82 samples correctly classified respectively, demonstrating the powerful ability of the model in different ground defect recognitions. It proves the effectiveness and reliability of the model in multi-category emergency scenarios and is an efficient classification method applicable to complex scenarios.
[0121] Reference Figure 6 , The classification prediction effect diagram shows the high accuracy and extremely fast inference speed of the spiking neural network classification model SpikeClassifier on multiple categories. The classification confidence of all categories is almost 100%, indicating the excellent classification accuracy of the model in different scenarios. Even though the confidence of the `no-bird-and-no-fire` category is slightly lower at 99.9%, the gap is negligible, further reflecting the powerful performance of the model. In terms of inference speed, the processing time of the model for each category is on average about 0.04 seconds (i.e., 25 FPS), indicating that the inference efficiency of the model in different tasks is very consistent and fast.
[0122] Table 1 Evaluation index table of the SNN classification network SpikeClassifier designed in the present invention and mainstream classification networks
[0123]
[0124] As can be seen from Table 1, the spiking neural network classification model SpikeClassifier designed by the present invention has obvious advantages over mainstream classification networks in terms of accuracy, parameter quantity and inference speed. First, in terms of accuracy, the classification accuracy of the spiking neural network classification model SpikeClassifier reached 95.84%, second only to the performance of UniFormer (97.04%), and higher than the competitive Meta-Spikeformer and other mainstream artificial neural networks; secondly, in terms of parameter quantity, the spiking neural network classification model SpikeClassifier only uses 25M parameters, which is much lower than VGG16 (134.2M) and UniFormer (50M), Meta-Spikeformer (30M) and other networks. The smaller parameter quantity not only effectively reduces the demand for computing resources, but also makes the model more lightweight, which is convenient for deployment and application on resource-constrained edge devices. Compared with other mainstream networks, such as PVTv2_B1 (63M), FocalNet_S (89M) and PatchConvent_S60 (48M), the spiking neural network classification model SpikeClassifier significantly reduces the number of model parameters while maintaining high accuracy, indicating that it has achieved a good balance of parameter efficiency in its design. By comparing the inference speed, it can be seen that the spiking neural network classification model SpikeClassifier is far ahead of all the compared models with an inference speed of 25.0FPS, showing extremely high real-time performance. In contrast, the inference speed of Meta-Spikeformer is significantly slower, at 0.71FPS, while the inference speed of PatchConvent_S60 is 0.83FPS, which shows that the spiking neural network classification model SpikeClassifier has a significant advantage in inference speed. Others such as VGG16's 5.88FPS, PVTv2_B1's 3.33FPS, FocalNet_S's 2.85FPS, and UniFormer's 1.72FPS are all significantly slower than SpikeClassifier, which further proves the superiority of the spiking neural network classification model SpikeClassifier of the present invention under real-time requirements.
[0125] Although the specific implementation of the invention is described in detail in conjunction with the drawings, it should not be understood as limiting the scope of protection of this patent. Within the scope described in the claims, various modifications and variations that can be made by those skilled in the art without creative work still fall within the scope of protection of this patent.
Claims
1. A rapid classification method for UAV accidents based on pulse neural network, characterized in that: The following steps are involved: S1. Obtain the initial dataset of UAV accidents; S2, preprocessing the initial data set to obtain a drone accident data set; S3, selecting the leaky integrate-release model as the spiking neuron of the spiking neural network; S4, constructing a spiking neural network classification model SpikeClassifier based on the leakage integration-dispensing model; S5, using the UAV accident data set to train the spiking neural network classification model SpikeClassifier, to obtain an optimized spiking neural network classification model SpikeClassifier; S6. Use the optimized spike neural network classification model SpikeClassifier to predict and classify UAV accidents.
2. The method for rapid classification of unmanned aerial vehicle accidents based on pulse neural network according to claim 1 is characterized in that: The initial data set in S1 includes positive samples and negative samples; The positive samples include dan-bird, qun-bird, fire, crack, stoma, spalling, and rust; Among them, dan-bird is a single bird image; qun-bird is a flock of birds image; fire is the flame and smoke image from the initial stage of wildfire to the spreading process; crack, stoma, spalling, and rust are ground crack images, ground pore images, ground peeling images, and ground rust images respectively; The negative samples are common scene images of sudden accidents, including the sky, trees and buildings.
3. The method for rapid classification of unmanned aerial vehicle accidents based on pulse neural network according to claim 1 is characterized in that: The S2 specifically includes: The initial data set was cleaned, data augmented, contrast and brightness adjusted to obtain the drone accident data set, which was then divided into a training set and a validation set in a ratio of 8:
2.
4. The method for rapid classification of unmanned aerial vehicle accidents based on pulse neural network according to claim 1 is characterized in that: In S3, the dynamic equation of the spiking neuron of the leaky integrate-release model as the spiking neural network is: in, is the membrane potential state of the ith spiking neuron in the nth layer before it generates a pulse at time step t, V i n (t-1) is the membrane potential state of the ith spike neuron in the nth layer at the previous moment (t-1), is the output current induced by the presynaptic neuron at the same time; τ is the membrane time constant; When the membrane potential of the spiking neuron reaches the threshold θ, a pulse output S(t) is generated and the membrane potential is reset to V(t): S(t)=θ(H(t)-θ) V(t)=r(H(t),S(t)) Where Θ() is the Heaviside step function, H(t) is the membrane potential state before the pulse is generated at time step t; r() is the membrane potential reset function.
5. The method for rapid classification of unmanned aerial vehicle accidents based on pulse neural network according to claim 4 is characterized in that: The membrane potential reset function includes hard reset and soft reset; The hard reset resets the membrane potential of the spiking neuron to a fixed resting potential V rest ; The soft reset lowers the membrane potential by the same magnitude as the emission threshold.
6. The method for rapid classification of unmanned aerial vehicle accidents based on pulse neural network according to claim 1 is characterized in that: In S4, the number of iterative time steps in the forward propagation process of the spiking neural network classification model SpikeClassifier is 25.
7. The method for rapid classification of unmanned aerial vehicle accidents based on pulse neural network according to claim 6 is characterized in that: In S4, the network structure of the spike neural network classification model SpikeClassifier includes: A first depthwise convolutional layer, wherein the first depthwise convolutional layer is sequentially followed by a first pointwise convolutional layer and a first maximum pooling layer; The first maximum pooling layer is connected to a second depth convolutional layer, and an output end of the second depth convolutional layer is sequentially connected to a second point-by-point convolutional layer and a second maximum pooling layer; The second maximum pooling layer is connected to the first fully connected layer, the output end of the first fully connected layer is connected to a pulse neuron layer, the pulse neuron layer is connected to the second fully connected layer; the second fully connected layer is connected to another pulse neuron layer.
8. The method for rapid classification of unmanned aerial vehicle accidents based on pulse neural network according to claim 7 is characterized in that: The forward propagation process of the spiking neural network classification model SpikeClassifier includes: The input image data is input to the first deep convolution layer, and then passes through the first point-by-point convolution layer, the first maximum pooling layer, the second deep convolution layer, the second point-by-point convolution layer, and the second maximum pooling layer, and the feature map is flattened into a one-dimensional vector. The one-dimensional vector is input to the first fully connected layer to compress the flattened one-dimensional vector into a feature vector; the feature vector passes through a pulse neuron layer to output a dynamic feature vector that changes with time; the dynamic feature vector is input to the second fully connected layer to map the dynamic feature vector to the classification space and output 8 categories corresponding to seven positive samples and one negative sample; finally, it passes through another pulse neuron layer to output static features that change with time.