Improved Flying Object Target Detection Method and Device Based on Spiking Neural Network

Through the method based on pulse neural network, the image frame sequence is converted into pulse features and multi-scale feature fusion is solved, and the problem of low accuracy in target detection of flying objects in the prior art is achieved, and efficient target recognition under low light and high speed conditions is achieved.

CN119888431BActive Publication Date: 2025-07-18NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361389.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-18
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The existing object detection method has blurred motion images in low-light conditions and fast moving object detection scenarios, resulting in low accuracy in object detection of flying objects.

Method used

Using an improved method based on pulse neural network, the image frame sequence is converted into pulse features through the pulse feature extraction layer, and the features are gradually extracted through the feature extraction subnet, and multi-scale feature fusion is performed in combination with the neck network, and finally the object detection layer is used to identify it, and the event-driven image frame sequence is generated using event data.

Benefits of technology

It improves the accuracy of target recognition in low-brightness environments and high-speed moving object images, reduces data calculations, and effectively captures dynamic targets and reduces static background interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888431B_ABST
    Figure CN119888431B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for improving the target detection of flying objects based on a spiking neural network, which relates to the technical field of target detection and mainly aims to solve the problem of low accuracy of existing target detection for flying objects. It mainly includes obtaining a target detection image and event data, and generating an event-driven image frame sequence based on the target detection image and event data. The image frame sequence is converted into spiking features through a spiking feature extraction layer, and the spiking features are gradually subjected to feature extraction by each feature extraction subnet to obtain image enhancement features corresponding to each feature extraction subnet; multi-scale feature fusion is performed on multiple image enhancement features through a neck network to obtain fusion features corresponding to each feature fusion subnet, and target recognition is respectively performed on the respective corresponding fusion features through each target detection layer to obtain the target detection result of the flying object. It is mainly used for the target detection of flying objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a method and device for detecting flying object targets improved based on a spiking neural network. Background Art

[0002] Object detection is one of the core tasks in the field of computer vision, aiming to automatically identify and locate target objects of interest in images or videos, such as people, vehicles, animals, etc. Its application scope is wide, covering multiple fields such as security monitoring, unmanned aerial vehicle flight control, autonomous driving, medical image analysis, human-computer interaction, intelligent retail, and augmented reality.

[0003] Existing object detection methods mainly rely on image data obtained by frame-based cameras that capture frames at a fixed rate. Although frame-based cameras provide dense intensity information, they often have problems such as limited dynamic range and relatively low frame rate. In application scenarios for detecting low-light conditions and fast-moving objects, the collected images have more motion blur, which in turn affects the richness and accuracy of features in the object detection process, resulting in relatively low accuracy in detecting flying object targets. Summary of the Invention

[0004] In view of this, the present invention provides a method and device for detecting flying object targets improved based on a spiking neural network, mainly aiming to solve the problem of relatively low accuracy in detecting flying object targets by existing object detection methods.

[0005] According to one aspect of the present invention, there is provided a method for detecting flying object targets improved based on a spiking neural network, including:

[0006] Performing object recognition through an object detection model improved based on a spiking neural network, wherein the object detection model includes a backbone network for pulse conversion and pulse feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale object recognition. The backbone network includes a pulse feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism. The neck network includes multiple feature fusion sub-networks of different scales. The head network includes multiple object detection layers of different scales. The method includes:

[0007] Obtaining a target detection image and event data of the target detection image, and generating an event-driven image frame sequence based on the target detection image and the event data, wherein the target detection image includes an object that moves rapidly relative to an image acquisition device;

[0008] The image frame sequence is converted into pulse features through the pulse feature extraction layer, and the pulse features are gradually subjected to feature extraction through each of the feature extraction sub-networks to obtain image enhancement features corresponding to each feature extraction sub-network; the neck network performs multi-scale feature fusion on multiple image enhancement features to obtain fusion features corresponding to each feature fusion sub-network, and each target detection layer respectively performs target recognition on the corresponding fusion features to obtain the flight object target detection result.

[0009] Further, the backbone network sequentially includes a pulse feature extraction layer, a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fast spatial pyramid pooling sub-network;

[0010] The process of gradually extracting features from the pulse features through each of the feature extraction sub-networks to obtain image enhancement features corresponding to each feature extraction sub-network includes:

[0011] The pulse features are input into the first feature extraction sub-network, and feature extraction is gradually performed through each feature extraction sub-network and the fast spatial pyramid pooling sub-network, and the output of the first feature extraction sub-network is taken as the second image enhancement feature, the output of the second feature extraction sub-network is taken as the second image enhancement feature, the output of the third feature extraction sub-network is taken as the third image enhancement feature, and the output of the fast spatial pyramid pooling sub-network is taken as the fourth image enhancement feature.

[0012] Further, the pulse feature extraction layer includes a neuron sub-layer, a convolution sub-layer, and a batch normalization sub-layer. The process of converting the image frame sequence into pulse features through the pulse feature extraction layer includes:

[0013] The spatial features and temporal features of the image frame sequence are extracted through the neuron sub-layer, the spatial features and the temporal features are fused to generate a membrane potential function, and a pulse signal is generated based on the membrane potential threshold and the membrane potential function;

[0014] The convolution sub-layer performs graph feature extraction on the pulse signal to obtain a pulse signal feature map, and the batch normalization sub-layer captures the temporal dependence relationship of the pulse signal feature map to obtain pulse features.

[0015] Further, each of the feature extraction sub-networks includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing a channel and spatial attention mechanism. The feature extraction layer includes a first convolution branch and a second convolution branch. The process of the feature extraction sub-network processing the input features includes:

[0016] Performing cross-channel feature interaction and combination processing on the input features through the first convolution branch to obtain an attention map;

[0017] Capturing local spatial information and feature representations at different scales in the input features through the first convolution branch to obtain a receptive field spatial feature, and fusing the attention map and the receptive field spatial feature to obtain a first image enhancement feature, wherein the number of convolution kernels of the second convolution branch is greater than that of the first convolution branch;

[0018] Performing feature extraction based on channel-space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature.

[0019] Further, each residual connection layer sequentially includes two impulse feature extraction sub-layers and a lightweight attention sub-layer. The lightweight attention sub-layer includes a channel attention branch and a spatial feature branch. The performing feature extraction based on channel-space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature includes:

[0020] Performing feature extraction on the first image enhancement feature step by step through two impulse feature extraction sub-layers to obtain a second image enhancement feature;

[0021] Obtaining a global spatial descriptor of the second image enhancement feature through the channel attention branch, calculating a channel attention map based on the global spatial descriptor, expanding the spatial feature information of the second image enhancement feature through the spatial feature branch to generate a cross-channel average pooling feature and a maximum pooling feature, performing splicing and convolution operations on the average pooling feature and the maximum pooling feature to obtain a cross-channel spatial feature, and fusing the channel attention map and the cross-channel spatial feature to obtain a third image enhancement feature;

[0022] Fusing the first image enhancement feature and the third image enhancement feature to obtain an image enhancement feature.

[0023] Further, the neck network sequentially includes a first feature fusion sub-network, a second feature fusion sub-network, a third feature fusion sub-network, a fourth feature fusion sub-network, a fifth feature fusion sub-network, a sixth feature fusion sub-network, and a seventh feature fusion sub-network. The fusion features include a first fusion feature, a second fusion feature, a third fusion feature, and a fourth fusion feature with gradually decreasing granularity;

[0024] The performing multi-scale feature fusion on multiple image enhancement features through the neck network includes:

[0025] Input multiple image enhancement features into the neck network, extract the output of the first feature fusion sub-network as the first fusion feature, extract the output of the third feature fusion sub-network as the second fusion feature, extract the output of the fifth feature fusion sub-network as the third fusion feature, and extract the output of the seventh fusion sub-network as the fourth fusion feature;

[0026] Among them, each feature fusion sub-network includes a residual connection layer, and the third feature fusion sub-network and the seventh feature fusion sub-network respectively further include a pulsed feature extraction layer for feature fusion.

[0027] Furthermore, the head network includes a small target detection layer, a first target detection layer, a second target detection layer, and a third target detection layer. Each target detection layer respectively performs target recognition on its corresponding fusion feature to obtain the flying object target detection result, including:

[0028] Perform target recognition based on spatio-temporal features on the first fusion feature through the small target detection layer to obtain the small target detection result. Specifically, it includes: obtaining the average pooling feature by performing average pooling operation on the first fusion feature, obtaining the max pooling feature by performing max pooling operation on the first fusion feature, performing weighted fusion on the average pooling operation and the max pooling feature through a shared multi-layer perceptron to obtain the temporal attention map, extracting the global spatial feature of the first fusion feature, and fusing the temporal attention map and the global spatial feature to generate a spatio-temporal fusion feature map, and performing flying object positioning and classification processing based on the spatio-temporal fusion feature map to obtain the small target detection result;

[0029] Perform target recognition on the second fusion feature through the first target detection layer to obtain the first target detection result, perform target recognition on the third fusion feature through the second target detection layer to obtain the second target detection result, and perform target recognition on the fourth fusion feature through the third target detection layer to obtain the third target detection result;

[0030] Determine the bounding box based on the small target detection result, the first target detection result, the second target detection result, and the third target detection result, and generate the flying object target detection result based on the bounding box.

[0031] Furthermore, the obtaining of the target detection image and the event data of the target detection image, and generating an event-driven image frame sequence based on the target detection image and the event data includes:

[0032] Construct event sets corresponding to different timestamps according to the timestamps of each pixel in the event data, where the event sets include the position coordinates, timestamps, and polarities of the pixels corresponding to different timestamps;

[0033] Convert the event data into a frame sequence according to the timestamp-based functional relationship between the pulse pattern tensor and the event set;

[0034] Stitch the frame sequence and the target detection image to obtain an event-driven image frame sequence.

[0035] Further, before performing target recognition through a target detection model improved based on a spiking neural network, the method further includes:

[0036] Obtain a training data set, where the training data set includes samples of event-driven image frame sequences marked with calibration boxes;

[0037] Construct an initial target detection model, where the initial target detection model includes a backbone network, a neck network, and a head network; the backbone network includes a spiking feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism, and each feature extraction sub-network sequentially includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing a channel and spatial attention mechanism; the neck network sequentially includes a first feature fusion sub-network, a second feature fusion sub-network, a third feature fusion sub-network, a fourth feature fusion sub-network, a fifth feature fusion sub-network, a sixth feature fusion sub-network, and a seventh feature fusion sub-network, and the head network includes a small target detection layer, a first target detection layer, a second target detection layer, and a third target detection layer;

[0038] Use the initial target detection model to perform target detection processing on the frame sequence samples obtained by converting the event data to obtain prediction boxes;

[0039] Calculate a first distance between the upper left corner points of the calibration box and the prediction box, and a second distance between the upper right corner points of the calibration box and the prediction box;

[0040] Construct a loss function based on the first distance, the second distance, and an attention weighted loss term;

[0041] Use the loss function to train the initial target detection model to obtain a trained target detection model.

[0042] According to another aspect of the present invention, there is provided an improved flying object target detection device based on a spiking neural network, including: an acquisition module and a target detection module. The target detection module performs target recognition through a target detection model improved based on a spiking neural network. Among them, the target detection model includes a backbone network for pulse conversion and pulse feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale target recognition. Among them, the target detection model includes a backbone network for pulse conversion and pulse feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale target recognition. The backbone network includes a pulse feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism. The neck network includes multiple feature fusion sub-networks of different scales. The head network includes multiple target detection layers of different scales;

[0043] The acquisition module is configured to acquire a target detection image and event data of the target detection image, and generate an event-driven image frame sequence based on the target detection image and the event data. Among them, the target detection image includes an object that moves rapidly relative to the image acquisition device;

[0044] The target detection module is configured to convert the image frame sequence into pulse features through the pulse feature extraction layer, and gradually extract features from the pulse features through each of the feature extraction sub-networks to obtain image enhancement features corresponding to each feature extraction sub-network; perform multi-scale feature fusion on multiple image enhancement features through the neck network to obtain fusion features corresponding to each feature fusion sub-network, and perform target recognition on the respective corresponding fusion features through each target detection layer to obtain a flying object target detection result.

[0045] By means of the above technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:

[0046] The present invention provides a method and device for detecting flying object targets improved based on a spiking neural network. In an embodiment of the present invention, a target detection image and event data of the target detection image are acquired, and an event-driven image frame sequence is generated based on the target detection image and the event data, wherein the target detection image includes an object that moves rapidly relative to an image acquisition device; the image frame sequence is converted into spiking features through a spiking feature extraction layer in a target detection model improved based on a spiking neural network, and the spiking features are gradually subjected to feature extraction through each feature extraction sub-network to obtain an image enhancement feature corresponding to each feature extraction sub-network; multi-scale feature fusion is performed on a plurality of image enhancement features through a neck network to obtain a fusion feature corresponding to each feature fusion sub-network, and target recognition is respectively performed on the respective corresponding fusion features through each target detection layer to obtain a flying object target detection result, realizing target detection based on image data and event data, processing the event data, converting the event-driven image frame sequence into a pulse signal, and performing convolutional feature extraction on the pulse signal, solving the problem that the existing convolutional neural network cannot perform feature extraction on event data, while realizing the processing of event-driven image data, greatly reducing the data calculation amount. In addition, feature extraction is performed based on an image frame sequence combining image static data and event data, which can effectively capture dynamic targets while supplementing high-dimensional features, reducing static background interference, thereby improving the target recognition accuracy of images in low-brightness environments and images of high-speed moving objects.

[0047] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0049] Figure 1 shows a flowchart of a method for detecting flying object targets improved based on a spiking neural network provided by an embodiment of the present invention;

[0050] Figure 2 shows a schematic diagram of the architecture of a target detection model improved based on a spiking neural network provided by an embodiment of the present invention;

[0051] Figure 3Shows a schematic diagram of the architecture of a pulse feature extraction layer provided by an embodiment of the present invention;

[0052] Figure 4 Shows a schematic diagram of the principle of neuron pulse firing provided by an embodiment of the present invention;

[0053] Figure 5 Shows a schematic diagram of the architecture of a residual connection layer based on lightweight attention provided by an embodiment of the present invention;

[0054] Figure 6 Shows a schematic diagram of the processing process of a lightweight attention sublayer provided by an embodiment of the present invention;

[0055] Figure 7 Shows a schematic diagram of the architecture of a small target detection layer provided by an embodiment of the present invention;

[0056] Figure 8 Shows a schematic diagram of the processing process of a spatio-temporal fusion feature extraction sublayer provided by an embodiment of the present invention;

[0057] Figure 9 Shows a block diagram of a flying object target detection device improved based on a spiking neural network provided by an embodiment of the present invention. Detailed implementation manners

[0058] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0059] Aiming at the problem that the existing target detection methods have low accuracy in detecting flying objects. Embodiments of the present invention provide a method for detecting flying object targets improved based on a spiking neural network, as Figure 1 shown, the method includes:

[0060] 101. Obtain a target detection image and event data of the target detection image, and generate an event-driven image frame sequence based on the target detection image and the event data.

[0061] In the embodiments of the present invention, the image to be detected includes an object that moves rapidly relative to the image acquisition device. For example, in application scenarios such as unmanned aerial vehicles, intelligent traffic monitoring, and security patrols, in order to achieve flight safety monitoring and obstacle avoidance, images are collected based on the image acquisition device configured on the flight device. Among them, the image acquisition device is a camera based on a dynamic vision sensor, also known as an event camera. The event camera asynchronously captures the scene through pixel-level light changes, providing lower latency and a higher dynamic range, and can collect images with less motion blur. However, the data generated by the event camera includes two parts: the image to be detected and the event data corresponding to each pixel in the image to be detected. In order to make full use of these two parts of data, it is necessary to convert the event data and integrate the converted event data with the image to be detected to obtain an event-driven image frame sequence. It includes the position coordinates, timestamp, and polarity of the current pixel. The polarity represents an increase or decrease in brightness and can be represented by +1 / -1.

[0062] 102. Perform object recognition through an object detection model improved based on a spiking neural network.

[0063] Among them, the object detection model includes a backbone network for spike conversion and spike feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale object recognition. The backbone network includes a spike feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism. The neck network includes multiple feature fusion sub-networks of different scales. The head network includes multiple object detection layers of different scales.

[0064] In the embodiment of the present invention, based on the pulse neurons in the pulse feature extraction layer, an image frame sequence containing event data is converted into a pulse signal, and convolutional feature extraction is performed on the pulse signal to obtain pulse features. The image frame sequence is converted into pulse features through the pulse feature extraction layer. Furthermore, different feature extraction sub-networks are connected in sequence. After the pulse features are input into the first feature extraction sub-network, each feature extraction sub-network gradually performs feature extraction. The output features of the previous feature extraction sub-network serve as the input of the current feature extraction sub-network, and the output of the current feature extraction sub-network serves as the input of the next feature extraction sub-network. Each of the pulse features is gradually subjected to feature extraction through each of the feature extraction sub-networks to obtain image enhancement features corresponding to each feature extraction sub-network. Among them, the structures of each feature extraction sub-network are the same, and each is a feature extraction sub-network based on a lightweight attention mechanism. This feature extraction sub-network includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing channel and spatial attention mechanisms. Through the neck network, multi-scale feature fusion is performed on multiple image enhancement features to obtain fusion features corresponding to each feature fusion sub-network. Through each object detection layer, object recognition is respectively performed on the corresponding fusion features to obtain the detection results of flying object targets.

[0065] It should be noted that by converting the event-driven image frame sequence into a pulse signal and performing convolutional feature extraction on the pulse signal, the problem that the existing convolutional neural network cannot perform feature extraction on event data is solved. While realizing the processing of event-driven image data, the data calculation amount is also greatly reduced. In addition, based on the image frame sequence combining image static data and event data for feature extraction, while supplementing high-dimensional features, dynamic targets can be effectively captured, static background interference can be reduced, thereby improving the object recognition accuracy of images in low-brightness environments and high-speed moving objects.

[0066] In an embodiment of the present invention, for further illustration and limitation, the step of gradually performing feature extraction on the pulse features through each of the feature extraction sub-networks to obtain image enhancement features corresponding to each feature extraction sub-network includes:

[0067] The pulse features are input into the first feature extraction sub-network, so as to gradually perform feature extraction through each feature extraction sub-network and the fast spatial pyramid pooling sub-network, and obtain the output of the first feature extraction sub-network as the second image enhancement feature, the output of the second feature extraction sub-network as the second image enhancement feature, the output of the third feature extraction sub-network as the third image enhancement feature, and the output of the fast spatial pyramid pooling sub-network as the fourth image enhancement feature.

[0068] In the embodiment of the present invention, the backbone network sequentially includes a pulse feature extraction layer, a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fast spatial pyramid pooling sub-network. As Figure 2 shown, after the event-driven image frame sequence is input into the backbone network, the pulse conversion and pulse feature extraction are first performed by the pulse feature extraction layer, and the pulse features output by the pulse feature extraction layer are used as the input features of the first feature extraction sub-network to start the step-by-step feature extraction by multiple feature extraction sub-networks and the fast spatial pyramid pooling sub-network. Specifically, the output of the first feature extraction sub-network is used as the input feature of the second feature extraction sub-network, the output of the second feature extraction sub-network is used as the input feature of the third feature extraction sub-network, and the output of the third feature extraction sub-network is used as the input feature of the fast spatial pyramid pooling sub-network. While the output of each feature extraction sub-network is used as the input feature of the next feature extraction sub-network, it is also used as the output feature of the backbone network, that is, the first feature extraction sub-network outputs the second image enhancement feature, the second feature extraction sub-network outputs the second image enhancement feature, the third feature extraction sub-network outputs the third image enhancement feature, and the fast spatial pyramid pooling sub-network outputs the fourth image enhancement feature. It should be noted that the fast spatial pyramid pooling sub-network includes a fourth feature extraction sub-network and a fast spatial pyramid pooling layer.

[0069] In an embodiment of the present invention, for further illustration and limitation, the step of converting the image frame sequence into pulse features through the pulse feature extraction layer includes:

[0070] extracting the spatial features and temporal features of the image frame sequence through a neuron sub-layer, fusing the spatial features and the temporal features to generate a membrane potential function, and generating a pulse signal according to a membrane potential threshold and the membrane potential function;

[0071] performing graph feature extraction on the pulse signal through the convolutional sub-layer to obtain a pulse signal feature map, and capturing the temporal dependence relationship of the pulse signal feature map through a batch normalization sub-layer to obtain pulse features.

[0072] In the embodiment of the present invention, as Figure 3 shown, the pulse feature extraction layer includes a neuron sub-layer, a convolutional sub-layer, and a batch normalization sub-layer (Threshold-dependent batch normalization, TDBN). Based on the simulation of the complex spatio-temporal dynamic characteristics of biological neurons and the trade-off of mathematical forms, the neuron can adopt the (leaky integrate-and-fire) neuron model. As Figure 4As shown, the processing of neurons includes: first, gradually extracting features from the input data through a Multi-Layer Perceptron (MLP), convolution, normalization, and average pooling to extract spatial features from the image frame sequence , and integrating the spatial features at the current moment and the time input (the time feature of the previous layer at the current moment) into the membrane potential , generating a spatial pulse tensor for the neurons in the next layer (layer n+1), and generating a new neuron state at the next moment (t+1). When the membrane potential is greater than or equal to the threshold , the output pulse tensor will be output, and the membrane potential gradually decays to the resting potential , obtaining the time feature at the current moment t . At the same time, the pulse tensor will be input into the neurons in layer n+1. If the membrane potential is less than the threshold , then at moment t, the neurons in layer n do not emit pulses , and the time feature gradually decays, waiting for the next input to accumulate. Through the above process, the pulse domain features of the image frame sequence including event data are refined and converted into pulses, and pulse-based convolutional feature extraction is performed, thereby realizing the processing of event data, and at the same time, improving the operation density of the neural network. To solve the problem that pulses cannot be differentiated in backpropagation, the above pulse emission process is represented by a surrogate gradient formula:

[0073] ;

[0074] where i represents the i-th neuron in layer n, represents the pulse tensor output by the i-th neuron, represents the membrane potential output by the i-th neuron, ensures that the integral of the gradient is 1 and determines the steepness of the curve, represents the pulse generation function.

[0075] After performing pulse-based convolutional feature extraction and obtaining the pulse signal feature map, the batch normalization sublayer processes the pulse signal feature map, optimizes the feature representation by modeling the temporal dependence relationship, and considers both the time domain and the spatial domain, thereby improving the accuracy of the extracted features. The processing process of the batch normalization sublayer for the pulse signal feature map can be expressed by the following formula:

[0076] ;

[0077] ;

[0078] Among them, represents the membrane potential generated by the i-th neuron in the (n + 1)-th neuron layer at the (t + 1)-th time step, represents the membrane potential generated by the i-th neuron in the (n + 1)-th neuron layer at the t-th time step, represents the pulse tensor generated by the i-th neuron in the (n + 1)-th neuron layer at the t-th time step, represents the neuron leakage decay factor, represents the total input of the i-th neuron at the (t + 1)-th time step, which is usually the weighted sum of the output pulses of the previous layer of neurons. represents the mean value of each channel under the mini-batch sequence input, represents the variance value of each channel under the mini-batch sequence input, , represents a small constant to avoid division by zero, and represent trainable parameters, represents a hyperparameter dependent on the threshold, represents the synaptic weight between the i-th and j-th neurons.

[0079] In an embodiment of the present invention, for further illustration and limitation, the processing process of the feature extraction sub-network for the input features includes:

[0080] Performing cross-channel feature interaction and combination processing on the input features through the first convolution branch to obtain an attention map;

[0081] Capturing local spatial information and feature representations at different scales in the input features through the first convolution branch to obtain a receptive field spatial feature, and fusing the attention map and the receptive field spatial feature to obtain a first image enhancement feature;

[0082] Performing feature extraction based on channel-space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature.

[0083] In an embodiment of the present invention, each of the feature extraction sub-networks includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing channel and spatial attention mechanisms. The feature extraction layer introducing a receptive field spatial attention mechanism includes a first convolution branch and a second convolution branch. The number of convolution kernels of the second convolution branch is greater than that of the first convolution branch. The first convolution branch can be a convolution branch with a convolution kernel of 1×1, and the second convolution branch is a convolution kernel of The convolutional branch. In the first convolutional branch, to reduce the complexity of the network structure and the number of parameters, first, global feature information is obtained through an average pooling layer, and then channel information interaction is performed through a 1×1 convolution operation. The Softmax function is used to emphasize the importance of different spatial feature information and obtain the corresponding weights, resulting in an attention map. Let the attention map be It can be expressed as:

[0084] ;

[0085] Where, represents average pooling, represents a convolution with a convolution kernel size of 1×1.

[0086] In the second convolutional branch, different receptive field spatial feature structures and sliding strides are formed according to different sizes of convolution kernels. The receptive field space is dynamically adjusted through the convolution kernel slider, and interactively fused with the attention information to generate the corresponding weights, and the receptive field spatial features are preferentially sorted to dynamically adjust the receptive field space. Group convolution is performed with the size to expand the content of the receptive field spatial feature information, and the number of feature information channels is increased to times to ensure that adjacent feature information does not overlap, generating the receptive field spatial feature corresponding to the attention map , which can be expressed as:

[0087] ;

[0088] Where, represents a convolution with a convolution kernel size of , represents normalization, represents the rectified linear unit activation function.

[0089] Combining the attention map with the receptive field spatial feature , that is, × as the first image enhancement feature. By dynamically adjusting the receptive field space, the computational redundancy problems caused by repeated feature extraction and traditional convolution kernel parameter sharing can be solved, thereby reducing the amount of computation and improving the data processing efficiency.

[0090] In an embodiment of the present invention, for further illustration and limitation, the first image enhancement feature is subjected to feature extraction based on channel space fusion attention through the residual connection layer to obtain an image enhancement feature, including:

[0091] The first image enhancement feature is gradually subjected to feature extraction through two pulse feature extraction sub-layers to obtain a second image enhancement feature;

[0092] Obtain the global spatial descriptor of the second image enhancement feature through the channel attention branch, calculate the channel attention map based on the global spatial descriptor, expand the spatial feature information of the second image enhancement feature through the spatial feature branch to generate the average pooling feature and the max pooling feature across channels, perform concatenation and convolution operations on the average pooling feature and the max pooling feature to obtain the cross-channel spatial feature, and fuse the channel attention map and the cross-channel spatial feature to obtain the third image enhancement feature;

[0093] Fuse the first image enhancement feature and the third image enhancement feature to obtain the image enhancement feature.

[0094] In the embodiment of the present invention, as Figure 5 shown, each residual connection layer sequentially includes two pulse feature extraction sub-layers and a lightweight attention sub-layer. Among them, the structure combination of the pulse feature extraction sub-layer is the same as that of the pulse feature extraction layer in the backbone network. The first image enhancement feature After passing through two layers of pulse feature extraction layers, the second image enhancement feature is obtained. Furthermore, it needs to pass through the lightweight attention sub-layer to obtain the third image enhancement feature denoted as , and then link to the previous layer (layer n-1) at the moment of to obtain the image enhancement feature , this process is expressed as . Among them, the lightweight attention sub-layer includes a channel attention branch and a spatial feature branch. As Figure 6 shown. The extraction process of the channel attention branch includes: the input feature passes through the global average pooling layer with a size of (avgpool) to obtain the global spatial descriptor, and jointly performs linear processing, the Rectified Linear Unit (ReLU) function, and the Sigmoid activation function with the max pooling result of the input feature with a size of , and performs two-layer shared weight weighting to obtain a one-dimensional channel attention map which can be expressed as:

[0095] ;

[0096] Among them, , represents the number of channels after operations such as convolution, represents the average pooling result, represents the max pooling result, represents the sigmoid function, Denote the weights of the $i$-th layer under the current number of channels, Denote the weights of the $j$-th layer under the current number of channels, i.e., and are the shared MLP weights for different time steps, is the channel reduction factor. As Figure 6 shown, the feature extraction process of the spatial feature branch includes content expansion of the spatial feature information through grouped convolution (feature size is ), followed by normalization and linear correction (feature size is ), then tensor shape adjustment (feature size is ). After the above normalization operations, average pooling and max pooling are performed to aggregate the channel information of the feature map, generating two 2D feature maps denoted as . Among them, each feature map represents the average pooling feature and max pooling feature across channels. Then the two feature maps are concatenated and cross-channel spatial features are generated through convolution ( ). Among them, in the above representation of the feature size is the number of channels, is the number of convolution kernels, is the width × height of the feature map size, and the convolution kernel of the convolution is . The cross-channel spatial features can be expressed as:

[0097] ;

[0098] where , , represents the sigmoid function, represents the number of image heights after max pooling operation, represents the number of image widths after max pooling operation.

[0099] In an embodiment of the present invention, for further illustration and limitation, multi-scale feature fusion of multiple image enhancement features is performed through the neck network, including:

[0100] Input multiple image enhancement features into the neck network, and extract the output of the first feature fusion sub-network as the first fusion feature, extract the output of the third feature fusion sub-network as the second fusion feature, extract the output of the fifth feature fusion sub-network as the third fusion feature, and extract the output of the seventh feature fusion sub-network as the fourth fusion feature.

[0101] In an embodiment of the present invention, as Figure 2As shown in the figure, the neck network sequentially includes a first feature fusion sub-network, a second feature fusion sub-network, a third feature fusion sub-network, a fourth feature fusion sub-network, a fifth feature fusion sub-network, a sixth feature fusion sub-network, and a seventh feature fusion sub-network. The fused features include a first fused feature, a second fused feature, a third fused feature, and a fourth fused feature with gradually decreasing granularity. Among them, each feature fusion sub-network includes a residual connection layer, and the third feature fusion sub-network and the seventh feature fusion sub-network respectively further include a pulse feature extraction layer for feature fusion. In addition, each feature fusion sub-network further includes a splicing sub-layer and an up-sampling sub-layer. Among them, the larger the granularity, the larger the image scale of the corresponding feature and the larger the spatial receptive field. On the contrary, the smaller the granularity, the smaller the image scale of the corresponding feature and the smaller the spatial receptive field.

[0102] In an embodiment of the present invention, for further illustration and limitation, the small target detection layer performs target recognition on the first fused feature based on spatio-temporal features to obtain a small target detection result;

[0103] The first target detection layer performs target recognition on the second fused feature to obtain a first target detection result, the second target detection layer performs target recognition on the third fused feature to obtain a second target detection result, and the third target detection layer performs target recognition on the fourth fused feature to obtain a third target detection result;

[0104] Based on the small target detection result, the first target detection result, the second target detection result, and the third target detection result, a positioning frame is determined, and a flying object target detection result is generated according to the positioning frame.

[0105] In an embodiment of the present invention, as Figure 2 shown, the head network includes a small target detection layer, a first target detection layer, a second target detection layer, and a third target detection layer. Among them, the architecture of the small target detection layer, as Figure 7 shown, includes a classification branch and a bounding box branch. The structures of the two branches are the same. From the input end, they are sequentially a pulse feature extraction sub-layer, a spatio-temporal fusion feature sub-layer, and a convolutional sub-layer, so as to construct their respective corresponding loss functions based on the output of the convolutional sub-layer during the training process. The process of the spatio-temporal fusion feature sub-layer, as Figure 8 shown, specifically includes: performing an average pooling operation on the first fused feature to obtain an average pooling feature (feature size is ), performing a max pooling operation on the first fused feature to obtain a max pooling feature (feature size is ), performing weighted fusion on the average pooling operation and the max pooling feature through a shared multi-layer perceptron to obtain a temporal attention map, and extracting the global spatial feature of the first fused feature (feature size is ),(the time attention map and the global spatial features are fused to generate a spatio-temporal fusion feature map, and object detection results of small targets are obtained by performing flying object localization and classification processing based on the spatio-temporal fusion feature map. To enhance small target detection, the internal features of the network in the time dimension are refined by exploring the cross-time-step relationships of the feature blocks, thereby achieving both performance gain and energy consumption reduction. First, the spatial channel information of the feature blocks is aggregated at each time step through average pooling and max pooling operations to generate two different temporal context descriptors, representing the average pooling feature and the max pooling feature respectively; then, the average pooling and max pooling features are converted into a time attention map through a shared MLP) , detecting the time characteristics of flying objects. The time attention map can be expressed as:

[0106] ;

[0107] where 、 , represents average pooling, represents the result of the max pooling layer, represents the number of time steps after operations such as convolution, represents the sigmoid function, and are the weights of the i-th and j-th linear layers in the shared MLP, represents a time reduction factor for controlling the computational burden of the multi-layer perceptron.

[0108] Then, global spatial enhanced feature extraction is performed on the first fusion feature, and a global spatial feature map is generated based on grouped convolution, normalization, and activation functions, and then the time attention map is fused to generate a spatio-temporal fusion feature map.

[0109] In an embodiment of the present invention, for further illustration and limitation, the steps of obtaining the target detection image and the event data of the target detection image, and generating an event-driven image frame sequence based on the target detection image and the event data include:

[0110] According to the timestamps corresponding to each pixel in the event data, event sets corresponding to different timestamps are constructed;

[0111] According to the function relationship based on timestamps between the pulse pattern tensor and the event set, the event data is converted into a frame sequence;

[0112] The frame sequence is spliced with the target detection image to obtain an event-driven image frame sequence.

[0113] In an embodiment of the present invention, the event set includes the position coordinates, timestamps, and polarities of pixels corresponding to different timestamps. Assume that the spatial resolution of the image is and the pulse mode tensor is equal to the event set at the timestamp , where represents the th pixel, and ([[]] , ) represents the position coordinates of the th pixel, and represents the polarity of the event. Based on the timestamp-based functional relationship between the pulse mode tensor and the event set, the event data is converted into a frame form to obtain a frame sequence. The functional relationship is expressed as follows:

[0114] ;

[0115] represents the frame form into which the input event data is converted at time t, where t ∈ {1, 2, …, T} is the time step, is the element addition function, and represents the time resolution factor, which determines how many consecutive events are included in each frame.

[0116] In an embodiment of the present invention, for further illustration and limitation, before performing object recognition through an object detection model improved based on a spiking neural network, the method further includes:

[0117] Obtain a training data set;

[0118] Construct an initial object detection model;

[0119] Use the initial object detection model to perform object detection processing on the frame sequence samples obtained by converting the event data to obtain prediction boxes;

[0120] Calculate the first distance between the calibration box and the upper left corner point of the prediction box, and the second distance between the calibration box and the upper right corner point of the prediction box;

[0121] Construct a loss function based on the first distance, the second distance, and the attention weighted loss term;

[0122] Use the loss function to train the initial object detection model to obtain a trained object detection model.

[0123] In the embodiments of the present invention, the training data set includes event-driven image frame sequence samples marked with calibration boxes. The initial object detection model includes a backbone network, a neck network, and a head network; the backbone network includes a pulse feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism, and each feature extraction sub-network successively includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing a channel and spatial attention mechanism; the neck network successively includes a first feature fusion sub-network, a second feature fusion sub-network, a third feature fusion sub-network, a fourth feature fusion sub-network, a fifth feature fusion sub-network, a sixth feature fusion sub-network, and a seventh feature fusion sub-network, and the head network includes a small object detection layer, a first object detection layer, a second object detection layer, and a third object detection layer. Set as the distance between the upper left corner points of the prediction box and the calibration box; as the distance between the lower right corner points of the prediction box and the calibration box, then the loss can be expressed as:

[0124] ;

[0125] Among them, , is the width of the prediction box, is the height of the prediction box, is the width of the calibration box, is the height of the calibration box, represents the pixel width, represents the pixel height.

[0126] Based on the above loss, attention weighting is performed to obtain a loss function, expressed as:

[0127] ; Among them, represents the weight coefficient, which is a hyperparameter, represents the th attention weight of the prediction box, calculated by the attention mechanism, represents the total number of target boxes detected in the current image (or feature map). Training the initial object detection layer based on the above loss function can dynamically adjust the geometric error weight, alleviate the positioning deviation in complex scenarios, and at the same time, improve the bounding box regression accuracy, solve the positioning deviation in complex scenarios, thereby improving the robustness of the model. In a specific application example, the model training parameter configuration can be that the optimizer selects Adam, the number of training rounds is 150, the batch size is 16, the input image size / pixel: 640×640, and the initial learning rate is 0.01.

[0128] Verified through actual ablation experiments and comparative experiments, and taking the original YOLOv8 network model as the experimental comparison: The above-mentioned object detection model improved based on pulses has the most obvious performance improvement. Compared with the original YOLOv8 network model, when the threshold of the average precision mean of the algorithm is 0.5, it has increased by 5.3 percentage points, the precision has increased by 4.8 percentage points, the recall rate has increased by 4.9 percentage points, and the total floating-point operation volume has decreased by 0.75×109. It not only improves the detection accuracy but also realizes the lightweight of the model.

[0129] The present invention provides a method for detecting flying object targets improved based on a spiking neural network. In an embodiment of the present invention, by obtaining a target detection image and event data of the target detection image, and generating an event-driven image frame sequence based on the target detection image and the event data, wherein the target detection image includes an object moving rapidly relative to an image acquisition device; converting the image frame sequence into spiking features through a spiking feature extraction layer in the object detection model improved based on a spiking neural network, and gradually extracting features of the spiking features through each feature extraction sub-network to obtain image enhancement features corresponding to each feature extraction sub-network; performing multi-scale feature fusion on multiple image enhancement features through a neck network to obtain fusion features corresponding to each feature fusion sub-network, and respectively performing target recognition on the respective corresponding fusion features through each target detection layer to obtain a flying object target detection result, realizing target detection based on image data and event data, processing the event data, converting the event-driven image frame sequence into a spiking signal, and performing convolutional feature extraction on the spiking signal, solving the problem that the existing convolutional neural network cannot perform feature extraction on event data. While realizing the processing of event-driven image data, it also greatly reduces the data calculation amount. In addition, based on the image frame sequence combining image static data and event data for feature extraction, it can effectively capture dynamic targets while supplementing high-dimensional features, reduce static background interference, and thus improve the target recognition accuracy of images in low-light environments and images of high-speed moving objects.

[0130] Further, as an implementation of the above Figure 1 shown method, an embodiment of the present invention provides a device for detecting flying object targets improved based on a spiking neural network, as Figure 9As shown in the figure, the device includes: an acquisition module 21 and a target detection module 22. The target detection module 22 performs target recognition through a target detection model improved based on a spiking neural network. Among them, the target detection model includes a backbone network for pulse conversion and pulse feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale target recognition. Among them, the target detection model includes a backbone network for pulse conversion and pulse feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale target recognition. The backbone network includes a pulse feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism. The neck network includes multiple feature fusion sub-networks of different scales. The head network includes multiple target detection layers of different scales;

[0131] The acquisition module 21 is configured to acquire a target detection image and event data of the target detection image, and generate an event-driven image frame sequence based on the target detection image and the event data. Among them, the target detection image includes an object that moves rapidly relative to an image acquisition device;

[0132] The target detection module 22 is configured to convert the image frame sequence into pulse features through the pulse feature extraction layer, and gradually perform feature extraction on the pulse features through each of the feature extraction sub-networks to obtain image enhancement features corresponding to each feature extraction sub-network; perform multi-scale feature fusion on multiple image enhancement features through the neck network to obtain fusion features corresponding to each feature fusion sub-network, and perform target recognition on the respective corresponding fusion features through each target detection layer to obtain a flying object target detection result.

[0133] Further, the backbone network sequentially includes a pulse feature extraction layer, a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fast spatial pyramid pooling sub-network; the target detection module includes:

[0134] The first feature extraction unit is configured to input the pulse features into the first feature extraction sub-network through the second feature extraction sub-network, so as to gradually perform feature extraction through each feature extraction sub-network and the fast spatial pyramid pooling sub-network, and obtain the output of the first feature extraction sub-network as the second image enhancement feature, the output of the second feature extraction sub-network as the second image enhancement feature, the output of the third feature extraction sub-network as the third image enhancement feature, and the output of the fast spatial pyramid pooling sub-network as the fourth image enhancement feature.

[0135] Further, the pulse feature extraction layer includes a neuron sub-layer, a convolution sub-layer, and a batch normalization sub-layer. The image frame sequence is converted into pulse features through the pulse feature extraction layer. The target detection module includes:

[0136] A generation unit for extracting the spatial and temporal features of the image frame sequence through the neuron sub-layer, fusing the spatial and temporal features to generate a membrane potential function, and generating a pulse signal based on a membrane potential threshold and the membrane potential function;

[0137] A third feature extraction unit for performing graph feature extraction on the pulse signal through the convolution sub-layer to obtain a pulse signal feature map, and capturing the temporal sequence dependence of the pulse signal feature map through the batch normalization sub-layer to obtain pulse features.

[0138] Further, the target detection module 22 includes:

[0139] A first processing unit for performing cross-channel feature interaction and combination processing on the input features through the first convolution branch to obtain an attention map;

[0140] A first fusion unit for capturing the local spatial information and feature representations of different scales in the input features through the first convolution branch to obtain a receptive field spatial feature, and fusing the attention map and the receptive field spatial feature to obtain a first image enhancement feature, where the number of convolution kernels of the second convolution branch is greater than that of the first convolution branch;

[0141] A fourth feature extraction unit for performing feature extraction based on channel space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature.

[0142] Further, in a specific application scenario, the fourth feature extraction unit is specifically configured to gradually perform feature extraction on the first image enhancement feature through two pulse feature extraction sub-layers to obtain a second image enhancement feature; obtain a global spatial descriptor of the second image enhancement feature through the channel attention branch, and calculate a channel attention map based on the global spatial descriptor; expand the spatial feature information of the second image enhancement feature through the spatial feature branch to generate a cross-channel average pooling feature and a maximum pooling feature; perform splicing and convolution operations on the average pooling feature and the maximum pooling feature to obtain a cross-channel spatial feature, and fuse the channel attention map and the cross-channel spatial feature to obtain a third image enhancement feature; and fuse the first image enhancement feature and the third image enhancement feature to obtain an image enhancement feature.

[0143] Further, the target detection module 22 includes:

[0144] The fifth feature extraction unit is configured to input multiple image enhancement features into the neck network, and extract the output of the first feature fusion sub-network as the first fusion feature, extract the output of the third feature fusion sub-network as the second feature fusion feature, extract the output of the fifth feature fusion sub-network as the third fusion feature, and extract the output of the seventh fusion sub-network as the fourth fusion feature; wherein each feature fusion sub-network includes a residual connection layer, and the third feature fusion sub-network and the seventh feature fusion sub-network respectively further include a pulsed feature extraction layer for feature fusion.

[0145] Further, the object detection module 22 includes:

[0146] The sixth feature extraction unit is configured to perform object recognition based on spatio-temporal features on the first fusion feature through the small object detection layer to obtain a small object detection result, specifically including: performing an average pooling operation on the first fusion feature to obtain an average pooling feature, performing a max pooling operation on the first fusion feature to obtain a max pooling feature, performing weighted fusion on the average pooling operation and the max pooling feature through a shared multi-layer perceptron to obtain a temporal attention map, extracting the global spatial features of the first fusion feature, and fusing the temporal attention map and the global spatial features to generate a spatio-temporal fusion feature map, and performing flying object positioning and classification processing based on the spatio-temporal fusion feature map to obtain a small object detection result;

[0147] The second fusion unit is configured to perform object recognition on the second fusion feature through the first object detection layer to obtain a first object detection result, perform object recognition on the third fusion feature through the second object detection layer to obtain a second object detection result, and perform object recognition on the fourth fusion feature through the third object detection layer to obtain a third object detection result;

[0148] The determination unit is configured to determine a bounding box based on the small object detection result, the first object detection result, the second object detection result, and the third object detection result, and generate a flying object target detection result based on the bounding box.

[0149] Further, the acquisition module 21 includes:

[0150] The construction unit is configured to construct an event set corresponding to different timestamps according to the timestamps corresponding to each pixel in the event data, where the event set includes the position coordinates, timestamps, and polarities of the pixels corresponding to different timestamps;

[0151] The conversion unit is configured to convert the event data into a frame sequence according to the functional relationship based on timestamps between the pulsed mode tensor and the event set;

[0152] The splicing unit is used for the third fusion to splice the frame sequence and the target detection image to obtain an event-driven image frame sequence.

[0153] Furthermore, the device further includes:

[0154] The obtaining module 21 is further used to obtain a training data set, wherein the training data set includes event-driven image frame sequence samples marked with calibration boxes;

[0155] The first construction module is used to construct an initial target detection model, wherein the initial target detection model includes a backbone network, a neck network and a head network; the backbone network includes a pulse feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism, and each feature extraction sub-network sequentially includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing a channel and spatial attention mechanism; the neck network sequentially includes a first feature fusion sub-network, a second feature fusion sub-network, a third feature fusion sub-network, a fourth feature fusion sub-network, a fifth feature fusion sub-network, a sixth feature fusion sub-network and a seventh feature fusion sub-network, and the head network includes a small target detection layer, a first target detection layer, a second target detection layer and a third target detection layer;

[0156] The second construction module is used to perform target detection processing on the event data converted into frame sequence samples by using the initial target detection model to obtain prediction boxes; calculate a first distance between the upper left corner points of the calibration box and the prediction box, and a second distance between the upper right corner points of the calibration box and the prediction box; construct a loss function according to the first distance, the second distance and an attention weighted loss term;

[0157] The training module is used to train the initial target detection model by using the loss function to obtain a trained target detection model.

[0158] The present invention provides a flight object target detection device improved based on a spiking neural network. In an embodiment of the present invention, a target detection image and event data of the target detection image are obtained, and an event-driven image frame sequence is generated based on the target detection image and the event data, wherein the target detection image includes an object moving rapidly relative to an image acquisition device; the image frame sequence is converted into spiking features through a spiking feature extraction layer in a target detection model improved based on a spiking neural network, and the spiking features are gradually subjected to feature extraction by each feature extraction sub-network to obtain an image enhancement feature corresponding to each feature extraction sub-network; multi-scale feature fusion is performed on multiple image enhancement features through a neck network to obtain a fusion feature corresponding to each feature fusion sub-network, and each target detection layer respectively performs target recognition on the corresponding fusion feature to obtain a flight object target detection result, realizing target detection based on image data and event data, processing the event data, converting the event-driven image frame sequence into a spiking signal, and performing convolutional feature extraction on the spiking signal, solving the problem that an existing convolutional neural network cannot perform feature extraction on event data. While realizing the processing of event-driven image data, the data calculation amount is also greatly reduced. In addition, feature extraction is performed based on an image frame sequence combining image static data and event data, which can effectively capture dynamic targets while supplementing high-dimensional features, reduce static background interference, and thus improve the target recognition accuracy of images in low-light environments and images of high-speed moving objects.

[0159] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0160] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An improved method for detecting flying object targets based on spiking neural networks, characterized in that Target recognition is performed through an improved object detection model based on a spiking neural network. The object detection model includes a backbone network for spike conversion and spike feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale object recognition. The backbone network includes a spike feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism. The neck network includes multiple feature fusion sub-networks of different scales. The head network includes multiple object detection layers of different scales. The method includes: Obtain an object detection image and event data of the object detection image, and generate an event-driven image frame sequence based on the object detection image and the event data, where the object detection image includes an object moving rapidly relative to an image acquisition device; Convert the image frame sequence into spike features through the spike feature extraction layer, and gradually perform feature extraction on the spike features through each of the feature extraction sub-networks to obtain an image enhancement feature corresponding to each feature extraction sub-network; perform multi-scale feature fusion on multiple image enhancement features through the neck network to obtain a fusion feature corresponding to each feature fusion sub-network, and perform object recognition on the respective corresponding fusion features through each object detection layer to obtain a flying object target detection result; Each of the feature extraction sub-networks includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing a channel and spatial attention mechanism. The feature extraction layer includes a first convolutional branch and a second convolutional branch. The process of the feature extraction sub-network processing the input features includes: Perform cross-channel feature interaction and combination processing on the input features through the first convolutional branch to obtain an attention map; Capture local spatial information and feature representations of different scales in the input features through the first convolutional branch to obtain a receptive field spatial feature, and fuse the attention map and the receptive field spatial feature to obtain a first image enhancement feature, where the number of convolutional kernels of the second convolutional branch is greater than that of the first convolutional branch; Perform feature extraction based on channel-space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature. Each residual connection layer sequentially includes two spike feature extraction sub-layers and a lightweight attention sub-layer. The lightweight attention sub-layer includes a channel attention branch and a spatial feature branch; The performing feature extraction based on channel-space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature specifically includes: Gradually perform feature extraction on the first image enhancement feature through two spike feature extraction sub-layers to obtain a second image enhancement feature; Obtain the global spatial descriptor of the second image enhancement feature through the channel attention branch, calculate the channel attention map based on the global spatial descriptor, expand the spatial feature information of the second image enhancement feature through the spatial feature branch to generate the average pooling feature and the maximum pooling feature across channels, obtain the cross-channel spatial feature through concatenating and convolving the average pooling feature and the maximum pooling feature, and fuse the channel attention map and the cross-channel spatial feature to obtain the third image enhancement feature; Fuse the first image enhancement feature and the third image enhancement feature to obtain the image enhancement feature.

2. The method according to claim 1, characterized in that, The backbone network sequentially includes a pulse feature extraction layer, a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fast spatial pyramid pooling sub-network. Gradually extracting features from the pulse feature through each of the feature extraction sub-networks to obtain the image enhancement feature corresponding to each feature extraction sub-network includes: Input the pulse feature into the first feature extraction sub-network, and gradually perform feature extraction through each feature extraction sub-network and the fast spatial pyramid pooling sub-network to obtain the output of the first feature extraction sub-network as the second image enhancement feature, the output of the second feature extraction sub-network as the second image enhancement feature, the output of the third feature extraction sub-network as the third image enhancement feature, and the output of the fast spatial pyramid pooling sub-network as the fourth image enhancement feature.

3. The method according to claim 1, characterized in that, The pulse feature extraction layer includes a neuron sub-layer, a convolution sub-layer, and a batch normalization sub-layer. Converting the image frame sequence into a pulse feature through the pulse feature extraction layer includes: Extract the spatial feature and the temporal feature of the image frame sequence through the neuron sub-layer, fuse and generate a membrane potential function based on the spatial feature and the temporal feature, and generate a pulse signal based on the membrane potential threshold and the membrane potential function; Perform graph feature extraction on the pulse signal through the convolution sub-layer to obtain a pulse signal feature map, and capture the temporal dependence relationship of the pulse signal feature map through the batch normalization sub-layer to obtain a pulse feature.

4. The method according to claim 1, wherein The neck network sequentially includes a first feature fusion sub-network, a second feature fusion sub-network, a third feature fusion sub-network, a fourth feature fusion sub-network, a fifth feature fusion sub-network, a sixth feature fusion sub-network, and a seventh feature fusion sub-network. The fusion features include a first fusion feature, a second fusion feature, a third fusion feature, and a fourth fusion feature with gradually decreasing granularity; Performing multi-scale feature fusion on multiple image enhancement features through the neck network includes: Input multiple image enhancement features into the neck network, and extract the output of the first feature fusion sub-network as the first fusion feature, the output of the third feature fusion sub-network as the second fusion feature, the output of the fifth feature fusion sub-network as the third fusion feature, and the output of the seventh fusion sub-network as the fourth fusion feature; Among them, each feature fusion sub-network includes a residual connection layer introducing channel and spatial attention mechanisms, and the third feature fusion sub-network and the seventh feature fusion sub-network further include impulse feature extraction layers for feature fusion respectively.

5. The method according to claim 4, wherein The head network includes a small target detection layer, a first target detection layer, a second target detection layer, and a third target detection layer. Each target detection layer performs target recognition on its corresponding fused feature to obtain a flying object target detection result, including: Performing target recognition based on spatio-temporal features on the first fused feature through the small target detection layer to obtain a small target detection result, which specifically includes: obtaining an average pooling feature by performing an average pooling operation on the first fused feature, obtaining a maximum pooling feature by performing a maximum pooling operation on the first fused feature, performing weighted fusion on the average pooling operation and the maximum pooling feature through a shared multi-layer perceptron to obtain a temporal attention map, extracting the global spatial feature of the first fused feature, and fusing the temporal attention map and the global spatial feature to generate a spatio-temporal fusion feature map, and performing flying object positioning and classification processing based on the spatio-temporal fusion feature map to obtain a small target detection result; Performing target recognition on the second fused feature through the first target detection layer to obtain a first target detection result, performing target recognition on the third fused feature through the second target detection layer to obtain a second target detection result, and performing target recognition on the fourth fused feature through the third target detection layer to obtain a third target detection result; Determining a bounding box based on the small target detection result, the first target detection result, the second target detection result, and the third target detection result, and generating a flying object target detection result based on the bounding box.

6. The method according to claim 1, wherein Obtaining the target detection image and the event data of the target detection image, and generating an event-driven image frame sequence based on the target detection image and the event data, including: Constructing event sets corresponding to different timestamps according to the timestamps of each pixel in the event data, where the event sets include the position coordinates, timestamps, and polarities of the pixels corresponding to different timestamps; Converting the event data into a frame sequence according to the functional relationship based on timestamps between the impulse mode tensor and the event sets; Splicing the frame sequence and the target detection image to obtain an event-driven image frame sequence.

7. The method according to claim 1, characterized in that, Before performing target recognition through the target detection model improved based on the spiking neural network, the method further includes: Obtaining a training data set, where the training data set includes event-driven image frame sequence samples marked with calibration boxes; Construct an initial object detection model, where the initial object detection model includes a backbone network, a neck network, and a head network; the backbone network includes a pulsed feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism, and each feature extraction sub-network sequentially includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing a channel and spatial attention mechanism; the neck network sequentially includes a first feature fusion sub-network, a second feature fusion sub-network, a third feature fusion sub-network, a fourth feature fusion sub-network, a fifth feature fusion sub-network, a sixth feature fusion sub-network, and a seventh feature fusion sub-network, and the head network includes a small object detection layer, a first object detection layer, a second object detection layer, and a third object detection layer; Use the initial object detection model to perform object detection processing on the frame sequence samples obtained by converting the event data, and obtain prediction boxes; Calculate a first distance between the calibration box and the upper left corner point of the prediction box, and a second distance between the calibration box and the upper right corner point of the prediction box; Construct a loss function based on the first distance, the second distance, and the attention weighted loss term; Use the loss function to train the initial object detection model to obtain a trained object detection model.

8. An improved flight object target detection device based on a spiking neural network, characterized in that, The device includes an acquisition module and an object detection module. The object detection module performs object recognition through an object detection model improved based on a spiking neural network. The object detection model includes a backbone network for pulse conversion and pulsed feature extraction, a neck network for multi-scale feature fusion, and a head network for multi-scale object recognition. The backbone network includes a pulsed feature extraction layer and multiple feature extraction sub-networks based on a lightweight attention mechanism. The neck network includes multiple feature fusion sub-networks of different scales. The head network includes multiple object detection layers of different scales; The acquisition module is used to acquire an object detection image and event data of the object detection image, and generate an event-driven image frame sequence based on the object detection image and the event data, where the object detection image includes an object moving rapidly relative to an image acquisition device; The object detection module is used to convert the image frame sequence into pulsed features through the pulsed feature extraction layer, and gradually perform feature extraction on the pulsed features through each of the feature extraction sub-networks to obtain image enhancement features corresponding to each feature extraction sub-network; perform multi-scale feature fusion on multiple image enhancement features through the neck network to obtain fusion features corresponding to each feature fusion sub-network, and perform object recognition on the respective corresponding fusion features through each object detection layer to obtain a flying object object detection result; Wherein, each of the feature extraction sub-networks includes a feature extraction layer introducing a receptive field spatial attention mechanism and a residual connection layer introducing a channel and spatial attention mechanism. The feature extraction layer includes a first convolution branch and a second convolution branch. The processing process of the feature extraction sub-network for the input features includes: Performing cross-channel feature interaction and combination processing on the input features through the first convolution branch to obtain an attention map; Capturing local spatial information and feature representations at different scales in the input features through the first convolution branch to obtain receptive field spatial features, and fusing the attention map and the receptive field spatial features to obtain a first image enhancement feature, wherein the number of convolution kernels of the second convolution branch is greater than that of the first convolution branch; Performing feature extraction based on channel-space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature, wherein each residual connection layer sequentially includes two pulse feature extraction sub-layers and a lightweight attention sub-layer, and the lightweight attention sub-layer includes a channel attention branch and a spatial feature branch; The performing feature extraction based on channel-space fusion attention on the first image enhancement feature through the residual connection layer to obtain an image enhancement feature specifically includes: Performing feature extraction on the first image enhancement feature step by step through two pulse feature extraction sub-layers to obtain a second image enhancement feature; Obtaining a global spatial descriptor of the second image enhancement feature through the channel attention branch, calculating a channel attention map based on the global spatial descriptor, expanding the spatial feature information of the second image enhancement feature through the spatial feature branch to generate a cross-channel average pooling feature and a maximum pooling feature, performing splicing and convolution operations on the average pooling feature and the maximum pooling feature to obtain a cross-channel spatial feature, and fusing the channel attention map and the cross-channel spatial feature to obtain a third image enhancement feature; Fusing the first image enhancement feature and the third image enhancement feature to obtain an image enhancement feature.

Citation Information

Patent Citations

  • Target identification method, device and equipment based on residual pulse neural network

    CN118397295A

  • SNN target tracking method and system fusing event and RGB image

    CN119477976A