Pulse neural network tracking method and system based on event camera

By constructing a spiking neural network tracking method based on an event camera, multi-time-step event data and low-resolution grayscale images are acquired to form a multi-channel event tensor. Dynamic and static features are extracted and fused to generate a high-resolution image for target tracking. This solves the problems of blurred target contours and loss of details in existing technologies, and achieves high-quality reconstruction and accurate tracking under high dynamic range.

CN120997257AActive Publication Date: 2025-11-21ACADEMY OF MILITARY MEDICAL SCIENCES

Patent Information

Application Number
CN202511526372.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2025-11-21
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Existing event camera visual tracking methods are difficult to meet the application requirements of mobile platforms such as UAVs under low power consumption and high dynamic range conditions. They suffer from problems such as blurred target outlines, loss of details, tracking drift and high failure rate. Furthermore, the spatiotemporal characteristics of sparse event data are difficult to utilize effectively.

Method used

By constructing a spiking neural network tracking method based on an event camera, multi-time-step event data and low-resolution grayscale images are acquired to form a multi-channel event tensor. Dynamic and static features are extracted and fused, a high-resolution image is generated using a super-resolution decoder, and target tracking is performed using a spiking neural network with LIF neurons.

Benefits of technology

It achieves high-quality target reconstruction and accurate tracking under low-resolution conditions, improves robustness and real-time performance, and is suitable for UAV platforms in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997257A_ABST
    Figure CN120997257A_ABST
Patent Text Reader

Abstract

The invention provides a pulse neural network tracking method and system based on an event camera. Acquiring multi-time-step event data and a low-resolution grayscale image output by an event camera, and combining the event data into a multi-channel event tensor according to a time sequence and polarity; dynamic event features in the event tensor and static image features in the grayscale image are extracted respectively, and feature fusion is carried out through channel splicing and convolution operation; the fusion features are input into a super-resolution decoder, up-sampling of the image is achieved through transposition convolution operation, and a high-resolution moving target image is generated; and inputting the high-resolution image into a pulse neural network feature extraction module based on LIF neurons, extracting features of a target template and a search area, and determining a target position through cross-correlation operation to realize tracking. According to the invention, the spatial resolution and target positioning precision of the image output by the event camera can be effectively improved, and robust target tracking in a complex dynamic environment is realized on a low-power-consumption edge device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and brain-like perception, and belongs to the cross-technology field of image processing, event vision and target tracking, in particular to a pulse neural network tracking method and system based on an event camera. BACKGROUND

[0002] With the rapid development of artificial intelligence and brain-like computing technology, event vision perception, as an important part of high dynamic vision systems, has shown significant advantages in unmanned aerial vehicles, autonomous driving and industrial detection, etc. The event camera outputs an asynchronous pixel-level brightness change event stream, has the characteristics of microsecond-level response, high dynamic range, low power consumption and low data redundancy, etc., and provides a new technical path for real-time vision tasks in complex environments. In related technologies, a multi-modal perception and processing system is constructed through the collaborative work of event data and grayscale images. Specifically, this technology covers the whole process from event stream acquisition, multi-time step event frame construction, multi-modal feature extraction to target reconstruction and tracking, including event feature extraction, image feature enhancement, feature fusion, super-resolution decoding and pulse neural network (SNN) driven tracking mechanism, etc.

[0003] However, in the existing event camera vision tracking method, low-resolution images and sparse event streams are directly used for target perception, and the dynamic information in the time dimension is not fully exploited, which may cause the target contour to be blurred, details to be lost, and the positioning accuracy to be reduced, or tracking drift and a significant increase in failure rate in complex scenes such as severe changes in illumination and multiple target overlaps, thereby affecting the tracking stability and robustness of the unmanned aerial vehicle platform under the constraints of high dynamic range and low power consumption. In addition, the traditional method lacks effective modeling of the spatio-temporal consistency and edge preservation capability in the feature fusion and reconstruction process, making it difficult to meet the dual demands of real-time performance and computational efficiency of edge computing devices. The existing technical problems are as follows: (1) Under the constraints of low power consumption and high dynamic range, traditional target tracking methods based on frame images are difficult to adapt to the application requirements of mobile platforms such as unmanned aerial vehicles, and the tracking robustness and real-time performance are insufficient.

[0004] (2) The event stream and grayscale image frame spatial resolution output by the existing event camera are low, resulting in image detail loss, which makes it difficult to meet the requirements of image quality and perception accuracy in high dynamic vision tasks such as target tracking; (3) The sparse event data contains rich dynamic information in the time dimension, but the existing methods are difficult to effectively extract and utilize its spatio-temporal features, limiting the further improvement of target perception and tracking accuracy. Therefore, it is urgent to study how to deeply exploit the spatio-temporal dynamic characteristics of event data. SUMMARY

[0005] The present application aims to at least partially solve one of the technical problems in the related art.

[0006] To this end, a first object of the present application is to propose an event camera-based pulse neural network tracking method, to construct an information fusion mechanism with timing consistency modeling capability, to realize high-quality reconstruction of low-resolution images, and to effectively improve the accuracy and robustness of target tracking.

[0007] A second object of the present application is to propose an event camera-based pulse neural network tracking system.

[0008] To achieve the above object, the first aspect of the present application proposes an event camera-based pulse neural network tracking method, comprising: S1, acquiring multi-time step event data and low-resolution grayscale images output by an event camera, and combining the event data in time sequence and polarity into a multi-channel event tensor; S2, extracting dynamic event features in the event tensor and static image features in the grayscale image, respectively, and performing feature fusion through channel splicing and convolution operation to form a spatiotemporal consistency feature representation; S3, inputting the fused features into a super-resolution decoder to realize up-sampling of the image through transposed convolution operation, and generating a high-resolution moving target image; S4, inputting the high-resolution image into a pulse neural network feature extraction module based on LIF neurons to extract features of the target template and the search area, and determining the target position through cross-correlation operation to realize tracking.

[0009] In one embodiment of the present application, the acquisition of multi-time step event data and low-resolution grayscale images output by the event camera, and the combination of the event data in time sequence and polarity into a multi-channel event tensor, comprises: S11, reading event images corresponding to time steps from ON and OFF subdirectories, a total of 8 images, splicing them in time sequence into a tensor with a shape of [8, H, W], wherein 4 time steps correspond to positive and negative polarity event images, respectively; S12, performing normalization processing on all image data, and uniformly normalizing the intensity values of the event images and the grayscale images to the range of [0, 1] to enhance the convergence stability and generalization ability of the model.

[0010] In one embodiment of the present application, the extraction of dynamic event features in the event tensor and static image features in the grayscale image, and the feature fusion through channel splicing and convolution operation to form a spatiotemporal consistency feature representation, further comprises: S21, dynamic event features are extracted through three-layer continuous two-dimensional convolution operation, each layer is followed by batch normalization and ReLU activation function, and finally a feature map with 64 output channels is output; S22, static image features are extracted through a parallel multi-scale convolution structure, including three groups of convolution channels with kernel sizes of 3*3, 5*5 and 7*7, and the responses of key regions are enhanced through an attention mechanism after splicing.

[0011] In an embodiment of the present application, the fused features are input into the super-resolution decoder, the up-sampling of the image is realized through the transposed convolution operation, and a high-resolution moving target image is generated, and the present application further comprises: S31, the transposed convolution operation is used, the kernel size is configured as 4, the step is configured as 2, and the padding is configured as 1, so that the spatial size of the image is enlarged to 2 times of the original size; S32, the batch normalization and ReLU activation function are used to enhance the non-linear modeling capability, and a 1-channel grayscale image is output through the final convolution layer.

[0012] In an embodiment of the present application, the present application further comprises: S5, the high-resolution moving target image is subjected to image quality evaluation, the Sobel operator is used to calculate the image gradient amplitude graph, and the gradient sparsity is constrained through the L1 norm, so as to further optimize the image edge retention capability.

[0013] To achieve the above purpose, a second embodiment of the present application proposes an event camera-based pulse neural network tracking system, which comprises: An event data acquisition module is configured to acquire multi-time step event data and low-resolution grayscale images output by an event camera, and combine the event data in time sequence and polarity into a multi-channel event tensor; A feature extraction and fusion module is configured to extract dynamic event features in the event tensor and static image features in the grayscale image respectively, and perform feature fusion through channel splicing and convolution operation, so as to form a spatio-temporal consistent feature representation; An image super-resolution reconstruction module is configured to input the fused features into a super-resolution decoder, realize the up-sampling of the image through the transposed convolution operation, and generate a high-resolution moving target image; A target tracking feature processing module is configured to input the high-resolution image into a pulse neural network feature extraction module based on LIF neurons, extract features of a target template and a search area, and determine the target position through cross-correlation operation to realize tracking.

[0014] The method and system of the embodiments of the present application realize high-quality super-resolution reconstruction of a moving target under the constraint of low spatial resolution event camera data, significantly improve the accuracy and robustness of target tracking, and maintain low power consumption and real-time performance.

[0015] Additional aspects and advantages of the present application will be made apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0016] The above and / or additional aspects and advantages of the present application will become apparent and be more readily understood from the following description, by reference to which: Figure 1 is a flow chart of an event camera based spiking neural network tracking method according to an embodiment of the present application; Figure 2 is a system architecture diagram according to an embodiment of the present application; Figure 3 is a data processing flow chart of a system according to an embodiment of the present application; Figure 4 is a structure diagram of an event camera based spiking neural network tracking system according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] It should be noted that the embodiments and features of the embodiments in the present application can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0018] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0019] A kind of event camera based spiking neural network tracking method and system according to the embodiments of the present application will be described below with reference to the accompanying drawings.

[0020] Embodiment 1 Figure 1 is a flow chart of an event camera based spiking neural network tracking method according to an embodiment of the present application, as shown in Figure 1 , comprising: S1, obtaining the multi-time step event data and low-resolution grayscale image output by the event camera, and combining the event data into a multi-channel event tensor in time sequence and polarity.

[0021] Specifically, this step aims to obtain multi-time step event data and low-resolution grayscale images output by an event camera, and combine the event data in time sequence and polarity into a multi-channel event tensor, which is a key link in data preprocessing and feature input construction in the entire algorithm system.

[0022] In an embodiment of the application, the event camera outputs an event stream in an asynchronous manner, and each event contains a timestamp, a pixel coordinate, and a polarity (ON / OFF), reflecting the dynamic process of the change in pixel brightness. In order to facilitate subsequent spiking neural network (SNN) processing, the event data needs to be aggregated in fixed time steps (T=4) to form an event frame. In each time step, the event data is divided into an ON event map and an OFF event map, respectively representing the event distribution of brightness increase and decrease. By reading and splicing the ON and OFF event images of the four time steps in time sequence, an event tensor with a shape of [8, H, W] is finally formed, where the eight channels correspond to the positive and negative polarity event maps of the four time steps, and H and W are the height and width of the image. This tensor structure can effectively retain the temporal dynamic characteristics of the event, providing a unified input format for subsequent feature extraction and information fusion.

[0023] Further, the construction of the event tensor depends on the setting of the time step T. In the present application, T=4, i.e., every 4 time steps are taken as a processing unit, ensuring the continuity and representativeness of event information in the time dimension. The normalization of the event image adopts the standard image intensity normalization method, mapping the pixel values of all event images and grayscale images to the range of [0, 1] to unify the intensity distribution and improve the stability and generalization ability of model training. The number of channels of the event tensor is 8, and the spatial size is consistent with the original image, facilitating channel splicing and fusion with image features in subsequent modules.

[0024] This step is suitable for real-time target tracking tasks on a UAV platform. Due to its high dynamic range and low latency, the event camera is particularly suitable for use in scenes with severe changes in illumination and fast target motion. By synchronously processing event data and grayscale images and constructing a multi-channel event tensor, high-quality, structured input data are provided for subsequent motion target super-resolution reconstruction and SNN tracking algorithms, thereby improving the perception ability and tracking robustness of the system in complex environments.

[0025] This step achieves alignment of event streams and image frames in time and spatial dimensions through structured processing of event data, laying a foundation for multi-modal feature fusion. The multi-channel event tensor not only retains the temporal dynamic information of the event, but also enhances the discrimination ability of the motion direction through polarity separation, which helps to improve the modeling accuracy and stability of the subsequent network for the target motion trajectory.

[0026] Further, S1 comprises: S11, read the event images of corresponding time steps from the ON and OFF sub-directories respectively, a total of 8 images, spliced in time sequence into a tensor with shape [8, H, W], wherein 4 time steps correspond to positive and negative polarity event images respectively.

[0027] Specifically, this step involves reading event images of corresponding time steps from the ON and OFF sub-directories of the event camera, and splicing them in time sequence into a tensor with shape [8, H, W], which contains 4 time steps, each corresponding to a positive and negative polarity event image. This step is a key data preprocessing link in the entire motion target super-resolution reconstruction and target tracking algorithm, and its technical implementation is based on the asynchronous output mechanism of the event camera and the construction strategy of multi-time-step event frames.

[0028] In some implementations, the event camera outputs event streams in an asynchronous manner, and each event records the pixel position, polarity (ON or OFF), and occurrence timestamp. To facilitate neural network processing, the event stream needs to be aggregated in fixed time steps (T=4) to generate event frames. Each time step contains two polarity event images, representing the event density graph of brightness increase (ON) and brightness decrease (OFF) in the time window. Therefore, each time step corresponds to two event images, and a total of 4 time steps form 8 event images.

[0029] The specific operation is as follows: first, according to the timestamp information of the event camera, the event data is divided into 4 continuous time windows, each with a length of T=4 milliseconds (which can be adjusted according to actual task requirements). In each time window, the pixel-level count of ON and OFF events is counted respectively to generate the corresponding event image. The event image is usually represented in the form of a single-channel grayscale image, and its pixel value reflects the frequency or intensity of event occurrence, with a value range of [0, 255]. To ensure the consistency of model input, all event images are normalized to the [0, 1] interval, conforming to the input standard of deep learning models.

[0030] Further, the 8 event images are channel-spliced in time sequence to form a four-dimensional tensor with shape [8, H, W], where 8 represents the number of time-polarity channels, H and W are the height and width of the image respectively. This tensor is used as the input of the event feature extractor for subsequent feature extraction and fusion processing.

[0031] The technical value of this step lies in that, by structuring the event data of the organization, a multi-time step event frame tensor is constructed to provide a time-consistent input representation for the neural network, thereby effectively capturing the motion trajectory and change trend of the target in the time dimension. In the target tracking task, this method can enhance the perception ability of the model to fast-moving targets, improve the robustness and real-time performance of tracking, and is especially suitable for complex scenes with large changes in light and fast target motion.

[0032] S12, normalize all image data, and normalize the intensity values of the event image and the grayscale image to the range of [0, 1] to enhance the convergence stability and generalization ability of the model.

[0033] Specifically, this step aims to normalize the intensity values of the event image and the grayscale image, and map them to the standardized numerical interval of [0, 1], thereby improving the convergence stability and generalization ability of the subsequent spiking neural network (SNN) model. In some implementations, the normalization operation is based on the Min-Max Normalization method, In specific implementations, considering that the event image output by the event camera is usually sparse binary or counting data (such as ON / OFF event count), its intensity value range may be much smaller than the 0-255 range of the grayscale image. Therefore, the normalization operation needs to be adapted to the numerical characteristics of different modal images. For example, the intensity value of the event image may be 0 to several event counts (such as 0-100), while the grayscale image is an 8-bit integer of 0-255. In the normalization process, a channel-level or batch-level normalization strategy can be optionally used to ensure consistent perception ability of the model under different input conditions.

[0034] Further, the normalization process plays a key role in this application. Due to the significant differences in time resolution, dynamic response mechanism and intensity distribution between event images and grayscale images, if normalization is not performed, the model may experience gradient shock and slow convergence during training due to uneven input data distribution. By normalizing all image data to the [0, 1] interval, the sensitivity of the model to input scale is reduced, and the fusion efficiency of the multi-modal data by the feature extraction module (such as the event feature extractor and the image feature extractor) is improved.

[0035] In practical applications, this step is usually completed in the image preprocessing stage and is suitable for edge computing environments on unmanned aerial vehicle platforms. The normalized image data serves as input for the subsequent convolutional neural network, which helps to improve the robustness of the model in complex lighting and dynamic scenes, especially in low-resolution and high-noise environments, effectively enhancing the recognizability of target outlines and providing a stable data foundation for super-resolution reconstruction and tracking of moving targets.

[0036] S2, respectively extracts the dynamic event features in the event tensor and the static image features in the grayscale image, and performs feature fusion through channel splicing and convolution operation to form a spatiotemporal consistent feature representation.

[0037] Specifically, this step involves the fusion processing of event tensor and grayscale image features, which is a key link to realize the spatiotemporal consistent modeling of moving targets. In some implementations, the dynamic event features output by the event feature extractor and the static image features extracted by the image feature extractor have dimensions of [B, 64, H, W], where 64 is the number of feature channels, and H and W are the height and width of the image. To effectively integrate multi-modal information, the event features and image features are first spliced in the channel dimension to form a fusion feature tensor with a channel number of 128, i.e., [B, 128, H, W]. This splicing operation preserves the temporal dynamic characteristics of event data and the spatial structure information of image data, providing a basis for subsequent feature interaction.

[0038] Further, to reduce the feature dimension and improve the fusion efficiency, a 1x1 convolution operation is used to compress the channels of the spliced features. The 1x1 convolution has the characteristics of small parameter quantity and high computational efficiency, and can realize cross-channel information integration and feature mapping. In this invention, the output channel number of the 1x1 convolution is set to 64, consistent with the original channel number of the event features and image features, thereby realizing the alignment of the feature space and the efficient fusion of information. The convolution operation is followed by batch normalization (Batch Normalization) and ReLU activation function to enhance the non-linear expression ability of the model and speed up the training process.

[0039] Optionally, to improve the training stability and feature expression ability of the model, a residual connection mechanism is introduced to add the original spliced features and the features processed by the 1x1 convolution element by element to form the final fusion features. This fusion feature not only preserves the time sensitivity of event data, but also integrates the structural information of image data, thereby constructing a feature representation with spatiotemporal consistency, providing high-quality input for subsequent super-resolution reconstruction and target tracking.

[0040] In practical applications, this step is applicable to real-time target tracking systems on unmanned aerial vehicle platforms, especially in high dynamic, low light or fast motion scenarios, which can effectively improve the robustness and accuracy of target perception. Through this fusion mechanism, the model can more accurately capture the motion trajectory and contour information of the target when processing low-resolution images and sparse event streams, thereby significantly improving the tracking performance.

[0041] Further, S2 includes: S21, dynamic event features are extracted through three-layer continuous two-dimensional convolution operation, each layer is followed by batch normalization and ReLU activation function, and finally a feature map with 64 output channels is output.

[0042] Specifically, this step is a key processing link in the event feature extractor, and the technical implementation principle is based on the deep convolutional neural network for layer-by-layer extraction and enhancement of the spatiotemporal features of event data. The event data output by the event camera records the change of pixel brightness in the form of asynchronous pulses, has high time resolution and low data redundancy characteristics. In order to effectively model the dynamic information in the event stream, this step adopts three-layer continuous two-dimensional convolution operation, each layer is followed by batch normalization (Batch Normalization, BN) and ReLU activation function to enhance the nonlinear expression ability of the model and speed up the training process. The dimension of the input event tensor is [B, 8, H, W], where the 8 channels correspond to the ON / OFF event map of 4 time steps, B is the batch size, and H and W are the height and width of the image.

[0043] In specific implementation, the parameter configuration of each two-dimensional convolution layer can be adjusted according to actual task requirements. In some implementations, the first convolution kernel size is 3x3, the step is 1, the input channel number is 8, and the output channel number is 16; the second convolution kernel size is still 3x3, the input channel number is 16, and the output channel number is 32; the third convolution kernel size is 3x3, the input channel number is 32, and the output channel number is 64. All convolution layers use zero padding (padding=1) to keep the spatial size of the feature map unchanged. The batch normalization operation is performed after each convolution, and the standardization parameters (mean and variance) are dynamically updated during the training process, which helps to alleviate the gradient vanishing problem and improve the model stability. The ReLU activation function introduces nonlinearity and enhances the model's ability to fit complex event patterns.

[0044] The feature map output by this step has a size of [B, 64, H, W] and a channel number of 64, which can fully express the motion trajectory, edge change and time dynamic features in the event data. In practical applications, this module is deployed on a UAV edge computing platform to process multi-time step event data collected by an event camera in real time, providing high-dimensional and high-discriminative dynamic feature representation for subsequent image reconstruction and target tracking. Through this structure design, the event data feature extraction efficiency and representation quality are effectively improved, laying a solid foundation for realizing high-precision and low-power motion target tracking.

[0045] S22, static image features are extracted through a parallel multi-scale convolution structure, including three groups of convolution channels with kernel sizes of 3x3, 5x5 and 7x7, and the responses of key regions are enhanced through attention mechanism after splicing.

[0046] In the present application, the image feature extractor adopts a parallel multi-scale convolution structure to extract static texture features and structural information in low-resolution grayscale images. This structure captures different scales of local features in the image by setting different receptive field convolution kernels (3x3, 5x5, and 7x7) in parallel, thereby enhancing the model's perception of target contours, edges, and detailed information. In some implementations, each convolution channel uses standard convolution operations with an input of a single-channel grayscale image tensor with dimensions [B, 1, H, W], where B represents the batch size, and H and W are the height and width of the image. The number of output channels of each convolution layer is consistent, all being 64, to ensure that the channel dimensions are aligned for subsequent feature fusion.

[0047] Specifically, the 3x3 convolution kernel has a small receptive field and is suitable for extracting local edges and high-frequency texture features in the image; the 5x5 convolution kernel can capture medium-scale structural information such as the contours and local shapes of the target; and the 7x7 convolution kernel has a larger receptive field and is used to extract global context information, which helps to enhance the model's understanding of the overall structure of the target. The convolution operations in each channel use the same stride (stride=1) and padding method (padding=same) to maintain the spatial dimensions of the output feature maps consistent. After the convolution operation, an attention mechanism (such as the SE attention module or CBAM) is introduced to weight and fuse the features of different scales, highlighting the response of key regions and improving the discriminability of feature expression.

[0048] This step plays a key role in the entire technical solution, and the multi-scale image features and event features output by this step are spliced and dimensionally reduced in the feature fusion module to provide complementary information of structure and texture for subsequent super-resolution reconstruction. By introducing a multi-scale convolution structure, the model can effectively retain the structural details of the target in low-resolution images, thereby improving the quality of the reconstructed image and the robustness of target tracking. This design meets the comprehensive needs of the unmanned aerial vehicle platform for lightweight, low power consumption, and high precision, and has significant engineering practical value.

[0049] S3, inputting the fused features into a super-resolution decoder to realize image up-sampling through transposed convolution operation and generating a high-resolution moving target image.

[0050] Specifically, in the technical solution of the present application, the step of "inputting the fused features into a super-resolution decoder to realize image up-sampling through transposed convolution operation and generating a high-resolution moving target image" is one of the core steps for implementing moving target super-resolution reconstruction. This step expands the spatial dimensions of the fused multi-modal feature map through transposed convolution (Transposed Convolution) operation in deep learning, thereby restoring a high-resolution moving target image and providing high-quality visual input for subsequent target tracking.

[0051] In this step, the input dimension of the fused features is [B, 64, H, W], where B represents the batch size, and H and W represent the height and width of the feature map. The super-resolution decoder uses a single-layer transposed convolution structure, with the configuration parameters being: a convolution kernel size of 4x4, a stride of 2, and a padding of 1. This configuration causes the spatial size of the output image to be expanded to twice the original size in each dimension, i.e., the output size is [B, 1, Hx2, Wx2]. The transposed convolution operation achieves upsampling through a backpropagation mechanism, which essentially interpolates and expands the information in the low-resolution feature map in the spatial domain, thereby generating more detailed image details.

[0052] The output channel number of the transposed convolution layer is 1, corresponding to a grayscale image. After this layer, a batch normalization (Batch Normalization) and a ReLU activation function are usually connected to enhance the non-linear expression ability of the model and speed up the training process. In addition, the up-sampling scale of the super-resolution decoder is 2, i.e., the output image is twice the resolution of the input image. In actual deployment, the computational complexity of this module is relatively low, making it suitable for running on edge devices such as drones, meeting the real-time and low-power requirements.

[0053] This step is mainly applied to the target tracking task of small targets in high dynamic, low light, or fast motion scenes. Due to the low resolution of the image output by the event camera, the super-resolution reconstruction through this step can significantly improve the target contour clarity and detail retention capability, thereby enhancing the robustness of target detection and tracking. In particular, in complex scenes such as small target recognition and occlusion recovery, high-resolution images can provide more rich edge and texture information, which helps to improve tracking accuracy.

[0054] This step achieves image up-sampling through transposed convolution, effectively recovering the lost detail information in the low-resolution image while preserving the dynamic information contained in the event data. Combined with the multi-modal features output by the feature fusion module, the super-resolution decoder can generate high-resolution images with spatio-temporal consistency, providing high-quality input for subsequent SNN (Spiking Neural Network) feature extraction and target positioning, thereby improving the performance and stability of the entire tracking system.

[0055] Further, S3 comprises: S31, using a transposed convolution operation with a kernel size of 4, a stride of 2, and a padding of 1, expands the image spatial size to twice the original size.

[0056] S32, enhances the non-linear modeling capability through batch normalization and ReLU activation function, and outputs a 1-channel grayscale image through the final convolution layer.

[0057] Specifically, in some implementations, the transpose convolution operation in the super-resolution encoder is a key step to realize the up-sampling of low-resolution images, and its technical principle is based on the mathematical mechanism of deconvolution or transposed convolution, which is used to enlarge the feature map in the spatial dimension, so as to restore the high-resolution details of the image. This step receives the fused feature map from the feature fusion module, and the input feature dimension is [B, 64, H, W], where B represents the batch size, H and W are the height and width of the current feature map. The transpose convolution layer is configured with a standard parameter combination of kernel size 4x4, stride 2, and padding 1, which can realize 2 times up-sampling of the image size, that is, the output size is [B, 64, 2H, 2W], which provides a basis for generating high-resolution grayscale images subsequently.

[0058] In terms of specific operation, the transpose convolution realizes spatial expansion through the convolution operation of back propagation, and its essence is to map each feature point in the low-resolution feature map to multiple positions in the high-resolution space, so as to interpolate and enhance details in the spatial dimension. In the present application, the layer adopts a standard convolution kernel initialization method, and the weight parameters are set by He initialization to alleviate the gradient vanishing problem and improve the training efficiency. The number of channels of the convolution kernel is 64, and the number of output channels is still 64, so as to keep the consistency of the feature dimension and facilitate the access of subsequent processing modules.

[0059] The stride (stride=2) of the transpose convolution determines the magnification of the image in the width and height directions, and the padding (padding=1) ensures that the output size is twice the input, not twice minus 2. This configuration conforms to the common up-sampling strategy in the image super-resolution task (such as the 4x4 convolution kernel and 2 times up-sampling structure widely used in models such as SRGAN, ESRGAN, etc.), which can effectively balance the computational complexity and image quality. In addition, the layer is followed by batch normalization (Batch Normalization) and ReLU activation function to enhance the non-linear expression ability and improve the recovery effect of image details.

[0060] The transpose convolution operation is deployed on the edge computing platform to process the low-resolution images and event stream data collected by the event camera in real time, generate high-resolution motion target images, and thus improve the accuracy and robustness of target tracking. Especially in high dynamic scenes, such as fast moving targets, severe changes in light or low light environments, this step can effectively restore the target edge and texture information, providing high-quality input for subsequent SNN feature extraction and related matching.

[0061] Further, this step plays a connecting role between the previous step and the subsequent step, and is a key link between feature fusion and image reconstruction. Through the up-sampling operation of transposed convolution, the model can realize high-resolution reconstruction of the moving target while maintaining computational efficiency, thereby significantly improving the accuracy and stability of target perception and meeting the strict requirements of the UAV platform for real-time performance and low power consumption.

[0062] S4, inputting the high-resolution image into a LIF neuron-based spiking neural network feature extraction module, extracting features of the target template and the search area, and determining the target position through cross-correlation operation to realize tracking.

[0063] Specifically, this step involves inputting a high-resolution image into a LIF (Leaky Integrate-and-Fire) neuron-based spiking neural network (SNN) feature extraction module to extract features of the target template and the search area, and determining the target position through cross-correlation operation to realize tracking of the moving target. This step is the core perception and decision-making link in the entire tracking algorithm, and its technical implementation is based on a brain-like computing architecture, which has the characteristics of low power consumption, high timing sensitivity, and strong robustness.

[0064] In some implementations, the SNN feature extraction module adopts a convolution-based spiking neural network structure, which is based on the topological structure of AlexNet and includes five convolution layers with convolution kernel sizes of 11x11, 5x5, 3x3, 3x3, and 3x3, and output channel numbers of 96, 256, 384, 384, and 256, respectively. After each layer of convolution operation, a LIF neuron model is connected for simulating the pulse firing mechanism of biological neurons, thereby realizing temporal coding and dynamic response of image features.

[0065] Further, the feature extraction module adopts a twin network structure to extract features of the target template image and the search area image of the current frame, respectively, and the two branches share convolution layer weights to ensure consistency of the feature space. The extracted feature map has a size of [B, 256, H', W'], where B is the batch size, and H' and W' are the spatial dimensions of the feature map. Subsequently, a response score map is generated through cross-channel cross-correlation operation, In practical applications, this step is suitable for real-time target tracking tasks of small targets in complex dynamic environments, especially in scenes with severe changes in light and fast target motion, which can effectively improve the tracking robustness and stability. By introducing the pulse neuron model, not only the computational energy consumption of the model is reduced, but also the temporal perception ability of event data is enhanced, thereby realizing efficient deployment on low-power edge devices.

[0066] In summary, this step extracts the spatio-temporal features in the high-resolution image through the LIF neuron-driven spiking neural network, and realizes the accurate positioning of the target position in combination with the cross-correlation operation, which is one of the key innovations of the application in the fusion application of event vision and target tracking.

[0067] The target tracking method based on the spiking neural network and the event camera in the embodiment of the application effectively improves the reconstruction quality of the low-resolution event camera image and the robustness of the target tracking, and realizes the accurate perception and stable tracking of the moving target in a high-dynamic and low-power consumption scene.

[0068] Further comprising: S5, image quality evaluation is performed on the high-resolution moving target image, a Sobel operator is used to calculate an image gradient amplitude map, and L1 norm is used to constrain the gradient sparsity, so as to further optimize the image edge retention capability.

[0069] Specifically, in the application, in the step of performing image quality evaluation on the high-resolution moving target image, a Sobel operator is used to calculate an image gradient amplitude map, and L1 norm is used to constrain the gradient sparsity, so as to further optimize the image edge retention capability. The core of this step is to quantitatively evaluate the edge structure of the reconstructed image by using the image gradient information, and to improve the structure fidelity of the image in the super-resolution reconstruction process by means of the sparsity constraint.

[0070] The Sobel operator is a classical edge detection operator, which extracts edge information by calculating the gradient amplitude of the image in the horizontal and vertical directions. Specifically, the Sobel operator is composed of two 3x3 convolution kernels, which are used to calculate the gradient response in the x and y directions respectively. In the application, the Sobel operator is applied to the reconstructed high-resolution image to generate a gradient amplitude map.

[0071] Further, in order to enhance the clarity of the image edge and suppress noise, the application introduces L1 norm to constrain the sparsity of the gradient amplitude map. L1 norm has the sparsity characteristic, which can make the response of the non-edge area in the gradient amplitude map tend to zero, so as to highlight the true edge information.

[0072] This step is mainly used for evaluating and optimizing the quality of the super-resolution reconstructed image in the unmanned aerial vehicle tracking system. The image output by the event camera has low resolution, and the edge information is easy to be lost. However, by jointly applying the Sobel gradient and the L1 sparsity constraint, the target edge structure can be effectively recovered, and the distinguishability of the image in the complex background can be improved. Especially in the low-light, high-speed motion or occlusion scene, this method can significantly enhance the boundary clarity of the target, thereby improving the positioning accuracy and stability of the tracking algorithm.

[0073] The technical effect of this step is reflected in: on the one hand, through the calculation of the gradient amplitude map, the edge quality of the image can be quantitatively evaluated; on the other hand, the L1 norm constraint helps to suppress artifacts and noise in the reconstructed image, enhancing the sparsity and saliency of the edge. Overall, this step provides a structure-aware auxiliary supervision signal for the model under the unsupervised learning framework, improving the reconstruction quality of the moving target image and the performance of the subsequent tracking algorithm.

[0074] Embodiment 2 To implement the pulse neural network tracking method based on the event camera proposed in the application, first, small target video data needs to be collected by the event camera. The event camera can accurately capture microsecond-level visual changes and can output traditional image frames and ON (light intensity increase) and OFF (light intensity decrease) event data at the same time. To ensure the consistency and training efficiency of the neural network model input, the event data and grayscale images need to be standardized and preprocessed. For the event data output by the event camera, the event data is matched in time with the image frames, and is divided into multiple time step event frames with a fixed time step (T=4). The specific processing procedure is as follows: First, S101, event data and image data are obtained based on the event camera, as follows: Read the event images of the corresponding time steps from the ON and OFF subdirectories respectively, a total of 8 images, and splice them in time order into a tensor with a shape of [8, H, W], where the channel dimension is 4 time steps x 2 positive and negative polarity event images.

[0075] All image data (event images and grayscale images) are normalized to the range [0, 1] to unify the image intensity distribution and enhance the convergence stability and generalization ability of the model.

[0076] Secondly, S102, motion target super-resolution reconstruction based on multi-time step event data is performed, as follows: It can be understood that the application aims to use the sparse visual information captured by the event camera at high temporal resolution to perform super-resolution reconstruction of the moving target in the target time period. The overall framework mainly consists of an event feature extractor, an image feature extractor, a feature fusion module, a super-resolution decoder, and a self-encoder module. This architecture can effectively fuse multi-time step event stream and low-resolution image frame information, realize spatiotemporal consistency modeling and detail recovery of the moving target. The overall framework process is as shown in Figure 2 .

[0077] Among them, the event feature extractor. The event feature extractor is used for extracting discriminative spatial features from event data output by a brain-like visual perception chip. The input of the module is an event tensor with a dimension of [B, 8, H, W], wherein B represents a batch size, H and W represent the height and width of the image respectively, and the 8 channels correspond to positive and negative polarity event maps extracted at 4 time steps respectively.

[0078] The module includes three layers of continuous two-dimensional convolution operations, each layer being connected with a batch normalization (Batch Normalization) and a ReLU activation function after convolution, so as to enhance the nonlinear expression ability of the network and accelerate the convergence speed. Finally, the module outputs a feature map with a channel number of 64, and the size is [B, 64, H, W], which provides a rich semantic information basis for subsequent image reconstruction and tracking.

[0079] Among them, the image feature extractor. The module is used for extracting static texture features and structural information in a low-resolution gray image. The input is a single-channel image with a size of [B, 1, H, W]. In order to obtain multi-scale context information, the module introduces a parallel multi-scale convolution structure, including three groups of convolution channels with kernel sizes of 3x3, 5x5 and 7x7. After the features are spliced, the key area response is further enhanced through the attention mechanism, so that the network can more accurately capture the edges and target areas in the image. Finally, the output dimension is [B, 64, H, W], which is aligned with the event feature in the spatial size and channel dimension.

[0080] Among them, the feature fusion module. The module is used for fusion processing of event features and image features, realizing deep integration of multi-source information. The specific process is as follows: The event feature map and the image feature map are spliced in the channel dimension to form a fusion feature with a dimension of [B, 128, H, W]; The channel dimension is compressed to 64 through 1x1 convolution operation, realizing efficient information fusion and dimension reduction; The residual connection mechanism is introduced to add the input and the processed features, so as to improve the training stability and strengthen the key feature expression.

[0081] The module effectively fuses dynamic event information and static image texture features, improves the overall feature representation ability, and enhances the perception ability of the model to moving targets.

[0082] Among them, the super-resolution encoder. In order to realize the enlargement and detail recovery of the low-resolution image, the application designs a structure efficient up-sampling module. The module receives the fusion feature as the input, and generates a high-resolution output through the following structure, specifically as follows: One layer transpose convolution operation, with kernel size 4, stride 2, padding 1, expands the image spatial size to 2 times of the original size; Batch normalization and ReLU activation function are used to enhance the non-linear modeling ability. The final convolutional layer outputs a 1-channel grayscale image with a size of [B, 1, Hxscale, Wxscale], where scale represents the up-sampling ratio.

[0083] The decoder takes into account both image quality and computational efficiency, and is suitable for edge platforms such as real-time deployment of unmanned aerial vehicles.

[0084] Among them, the auto-encoder-decoder structure. In order to enhance the learning ability of the network in the unsupervised scene, an auto-encoder structure is specially introduced as an auxiliary training module. This module mainly includes: Encoder: composed of two layers of convolution, the input image is compressed into a 64-channel low-dimensional feature by the encoder; Decoder: symmetric structure, also composed of two layers of convolution, used to restore the feature to the original size of the image; Output: the reconstructed low-resolution image, with a size of [B, 1, H, W], consistent with the original image.

[0085] The training target of this module is to minimize the reconstruction error between the input image and the reconstructed image, guiding the main network to learn more robust feature representation, and still having good generalization performance in the absence of high-quality labeled data.

[0086] Among them, the loss function design. In order to realize the training optimization process of the motion target super-resolution reconstruction method proposed in the present application, a set of composite loss function system composed of multiple components working together is designed. Under the unsupervised learning framework, the model is constrained and guided by using multiple source supervision signals, which improves the quality and structural fidelity of the reconstructed motion target, ensures that the model has strong edge preservation ability and feature consistency, and mainly includes the following four items: Auto-encoder reconstruction loss. This loss term is used to measure the reconstruction ability of the auto-encoder module for the input low-resolution image, which guides the encoder-decoder structure to learn the main texture and structure information of the image, and enhances the representation ability of the model under unsupervised conditions. This item is composed of two standard error forms.

[0087] L1 loss. Pixel-level absolute error is used, which has strong edge preservation ability, and is defined as:

[0088] Among them, indicates the reconstructed image of the auto-encoder, is the input low-resolution image.

[0089] Mean Squared Error Loss (MSE). Used to measure the average square value of overall pixel difference, defined as:

[0090] Combining these two losses can ensure the accuracy of the overall image structure and the ability to retain details.

[0091] Consistency Loss. To ensure that the super-resolution image generated by the main path is consistent with the structure of the autoencoder path, a consistency loss term is introduced. Under this constraint, the down-sampled super-resolution image should be highly similar to the autoencoder reconstructed image, defined as:

[0092] where, denotes the super-resolution reconstructed image, denotes the 2x down-sampling operation using bilinear interpolation or other methods. This loss term promotes consistency between the two information flow paths in the feature space, strengthening the model's ability to model image structure.

[0093] Smoothness Loss. To reduce the artifacts and noise that may appear in the super-resolution image, an image smoothness loss is introduced to encourage the continuity of the generated image in spatial distribution. This term is achieved by calculating the gradient difference between adjacent pixels in the image, defined as:

[0094] This term helps to eliminate high-frequency noise in the image, improving the naturalness and visual quality of the reconstructed image.

[0095] Gradient Loss. To further preserve key edge information in the image, a gradient loss is introduced to strengthen the sparsity of the image gradient magnitude map. This term can calculate the edge intensity of the image through the Sobel operator or difference operator, and impose sparsity constraints, defined as:

[0096] where, ▽ denotes the gradient calculation operation of the image, and L1 norm is used to enhance the response of the edge region and suppress meaningless detail noise.

[0097] Finally, the total loss function is represented in the form of weighted sum:

[0098] where, to are the weight parameters of each loss term, which can be adjusted according to task requirements and training performance. This multi-task loss design can consider the content accuracy, structural consistency, edge preservation and visual quality of the image, and is one of the core optimization methods in the model training phase of the invention.

[0099] The network is trained and optimized. For training the whole image reconstruction model, this part adopts Adam optimizer for end-to-end training, and the following parameters are set: Initial learning rate: 0.0001; Weight decay: 0.0001; Momentum parameter: .

[0100] During the training process, the following strategies are used to improve the performance of the model: Reduce the learning rate to 0.1 times of the original every 20 epochs.

[0101] Finally, S103, the implementation process of the tracking algorithm based on the reconstructed moving target is as follows: The embodiment proposes a target tracking algorithm based on reconstructed moving target and spiking neural network (SNN). The model uses reconstructed high-resolution moving target images to improve the adaptability and robustness of the model to small targets in complex environments, while maintaining temporal dynamic information and significantly reducing computational overhead, suitable for low-power edge devices.

[0102] The overall framework of the model is shown in FIG. 3. The input includes two modalities of image and corresponding event data stream. The two modalities of data are input into the super-resolution moving target reconstruction module mentioned above to obtain high-resolution moving target images. Then the SNN feature extraction module is used to obtain the spatial features of the template and search branch. Then the spatial features of the template and search frame are correlated to find the maximum correlation position, which is the position of the target.

[0103] Specifically, the SNN feature extraction module, the loss function and the training strategy are as follows: The feature extraction module adopts a convolution-based spiking neural network structure to extract features from the target template image and the search region image. The network structure is based on a simplified AlexNet and consists of multiple convolution layers, spiking neuron models, local response normalization layers (LRN) and pooling layers. The parameters of the convolution layer are shared between the two branches of the twin network to ensure that the extracted features have consistency in the same space.

[0104] The network structure is as follows, including five convolution layers, which are configured as follows: The first layer of convolution kernel size is 11x11, the step is 2, and the output channel is 96; The second layer of convolution kernel size is 5x5, and the output channel is 256; The third, fourth and fifth convolution kernel sizes are 3x3, and the output channels are 384, 384 and 256, respectively. All convolutional layers are followed by an impulse neuron model, and some layers are followed by a pooling or LRN operation.

[0105] The extracted target template feature map and the search region feature map are subjected to a cross-channel cross-correlation operation to generate a response score map, and the maximum value position in the response score map is the predicted position of the target in the current frame.

[0106] The impulse neurons in the network use the classic LIF neuron model, and the dynamic model of LIF can be described by three equations of charging, discharging and resetting. The charging equation is:

[0107] wherein is the membrane time constant. represents the membrane voltage of the neuron at the previous time. is the external input at the current time.

[0108] When the instantaneous membrane potential exceeds the threshold , a pulse is fired, otherwise not. The neuron firing process is as follows:

[0109] wherein is a step function:

[0110] When the impulse neuron fires a pulse, the voltage is reset to . When the impulse neuron does not fire a pulse, the voltage remains unchanged. The reset equation of the impulse neuron is as follows:

[0111] wherein is the membrane potential at the current time.

[0112] Loss function and training strategy. The model uses a logistic regression loss function to optimize the accuracy of the response map. The loss function is defined as follows:

[0113] wherein: represents the predicted value of the iii position in the response map; is the label corresponding to the position, the central region is set as a positive sample (+1), and the region away from the central region is set as a negative sample (-1); N represents the number of all pixel points in the response map.

[0114] The event camera based pulse neural network tracking method of the embodiment effectively improves the reconstruction quality of a low-resolution event camera image and the target tracking robustness, and realizes accurate perception and stable tracking of a moving target in a high-dynamic and low-power consumption scene.

[0115] Embodiment 3 To achieve the above-mentioned embodiments, as shown in the figure, the embodiment further provides an event camera based pulse neural network tracking system 10, comprising: Figure 4 An event data acquisition module 100 is configured to acquire multi-time step event data and low-resolution grayscale images output by an event camera, and combine the event data in time sequence and polarity into a multi-channel event tensor; A feature extraction and fusion module 200 is configured to extract dynamic event features in the event tensor and static image features in the grayscale image respectively, and perform feature fusion through channel splicing and convolution operation to form a spatio-temporal consistent feature representation; An image super-resolution reconstruction module 300 is configured to input the fused features into a super-resolution decoder, and realize image up-sampling through transposed convolution operation to generate a high-resolution moving target image; A target tracking feature processing module 400 is configured to input the high-resolution image into a pulse neural network feature extraction module based on LIF neurons, extract features of a target template and a search area, and determine a target position through cross-correlation operation to realize tracking. Further, the event data acquisition module is further configured to:

[0116] read event images of corresponding time steps from ON and OFF subdirectories, a total of 8 images, and splice them in time sequence into a tensor with a shape of [8, H, W], wherein 4 time steps correspond to positive and negative polarity event images respectively; perform normalization processing on all image data, and normalize intensity values of the event images and the grayscale images to the range of [0, 1] to enhance the convergence stability and generalization ability of the model. Further, the feature extraction and fusion module is further configured to:

[0117] extract dynamic event features through three-layer continuous two-dimensional convolution operation, and after each layer of convolution, connect batch normalization and ReLU activation function, and finally output a feature map with a channel number of 64; extract static image features through a parallel multi-scale convolution structure, including three groups of convolution channels with kernel sizes of 3x3, 5x5 and 7x7, splice them and enhance the response of key regions through an attention mechanism. Further, the image super-resolution reconstruction module is further configured to:

[0118] ​The image space size is enlarged to 2 times of the original size by using the transpose convolution operation, and the kernel size is 4, the step is 2, and the padding is 1. The non-linear modeling capability is enhanced by batch normalization and ReLU activation function, and a 1-channel grayscale image is output through the final convolution layer.

[0119] Further, it also comprises: An image quality evaluation module is configured to evaluate the image quality of the high-resolution moving target image, calculate an image gradient amplitude graph by using a Sobel operator, and constrain the gradient sparsity by using an L1 norm to further optimize the image edge retention capability.

[0120] The event camera-based pulse neural network tracking system of the embodiment of the present application effectively improves the reconstruction quality and target tracking robustness of a low-resolution event camera image, and realizes accurate perception and stable tracking of a moving target in a high-dynamic and low-power consumption scene.

[0121] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0122] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise specifically limited.

Claims

1. A spiking neural network tracking method based on an event camera, characterized in that, include: S1, acquire multi-timestep event data and low-resolution grayscale image output by the event camera, and combine the event data into a multi-channel event tensor according to time order and polarity; S2 extracts dynamic event features from the event tensor and static image features from the grayscale image, and fuses the features through channel concatenation and convolution operations to form a spatiotemporally consistent feature representation. S3 inputs the fused features into the super-resolution decoder, and upsamples the image through transposed convolution to generate a high-resolution moving target image; S4 inputs the high-resolution image into the LIF neuron-based spiking neural network feature extraction module to extract features of the target template and the search area, and determines the target position through cross-correlation operation to achieve tracking.

2. The method as described in claim 1, characterized in that, The process of acquiring multi-timestep event data and low-resolution grayscale images output by the event camera, and combining the event data into a multi-channel event tensor according to time order and polarity, includes: S11, read the event images corresponding to the time steps from the ON and OFF subdirectories respectively, a total of 8 images, and stitch them together in chronological order into a tensor of shape [8, H, W], where 4 time steps correspond to positive and negative polarity event images respectively; S12, normalize all image data, and normalize the intensity values ​​of event images and grayscale images to the range of [0, 1] to enhance the convergence stability and generalization ability of the model.

3. The method as described in claim 1, characterized in that, The step of extracting dynamic event features from the event tensor and static image features from the grayscale image, and fusing these features through channel concatenation and convolution operations to form a spatiotemporally consistent feature representation, further includes: S21 extracts dynamic event features through three consecutive two-dimensional convolution operations. Each convolution is followed by batch normalization and ReLU activation function, and the final output is a feature map with 64 channels. S22 extracts static image features through a parallel multi-scale convolutional structure, including three sets of convolutional channels with kernel sizes of 3×3, 5×5 and 7×7. After concatenation, the key region response is enhanced through an attention mechanism.

4. The method as described in claim 1, characterized in that, The step of inputting the fused features into the super-resolution decoder and upsampling the image through transposed convolution to generate a high-resolution moving target image further includes: S31 uses transposed convolution with a kernel size of 4, a stride of 2, and padding of 1 to enlarge the image space size to twice its original size. S32 enhances nonlinear modeling capabilities through batch normalization and ReLU activation functions, and outputs a 1-channel grayscale image through the final convolutional layer.

5. The method as described in claim 1, characterized in that, Also includes: S5 performs image quality assessment on high-resolution moving target images, uses the Sobel operator to calculate the image gradient magnitude map, and constrains gradient sparsity through the L1 norm to further optimize the image edge preservation capability.

6. A spiking neural network tracking system based on an event camera, characterized in that, include: The event data acquisition module is used to acquire multi-time step event data and low-resolution grayscale images output by the event camera, and combine the event data into a multi-channel event tensor according to time order and polarity. The feature extraction and fusion module is used to extract dynamic event features from the event tensor and static image features from the grayscale image, and to fuse the features through channel concatenation and convolution operations to form a spatiotemporally consistent feature representation. The image super-resolution reconstruction module is used to input fused features into the super-resolution decoder, and to achieve image upsampling through transposed convolution operation to generate high-resolution moving target images; The target tracking feature processing module is used to input high-resolution images into the LIF neuron-based spiking neural network feature extraction module, extract the features of the target template and the search area, and determine the target position through cross-correlation calculation to achieve tracking.

7. The system as described in claim 6, characterized in that, The event data acquisition module is also used for: The event images corresponding to the time steps are read from the ON and OFF subdirectories, totaling 8 images. They are then stitched together in chronological order into a tensor of shape [8, H, W], where 4 time steps correspond to positive and negative polarity event images respectively. All image data are normalized to unify the intensity values ​​of event images and grayscale images to the range of [0,1], thereby enhancing the convergence stability and generalization ability of the model.

8. The system as described in claim 6, characterized in that, The feature extraction and fusion module is also used for: Dynamic event features are extracted by three consecutive two-dimensional convolution operations. Each convolution is followed by batch normalization and ReLU activation function, and the final output is a feature map with 64 channels. Static image features are extracted using a parallel multi-scale convolutional structure, including three sets of convolutional channels with kernel sizes of 3×3, 5×5, and 7×7. After concatenation, the response of key regions is enhanced through an attention mechanism.

9. The system as described in claim 6, characterized in that, The image super-resolution reconstruction module is also used for: By using transposed convolution with a kernel size of 4, a stride of 2, and padding of 1, the image spatial size is enlarged to twice its original size. The nonlinear modeling capability is enhanced by batch normalization and ReLU activation function, and a 1-channel grayscale image is output through the final convolutional layer.

10. The system as described in claim 6, characterized in that, Also includes: The image quality assessment module is used to assess the image quality of high-resolution moving target images. It uses the Sobel operator to calculate the image gradient magnitude map and constrains gradient sparsity through the L1 norm to further optimize the image edge preservation capability.

Citation Information

Patent Citations

  • Pulse neural network target tracking method and system based on event camera

    CN114429491A

  • Target tracking method of spiking neural network based on multiple attention mechanisms

    CN117314972A

  • SNN target tracking method and system fusing event and RGB image

    CN119477976A

  • Neuromorphic visual target tracking method and system based on image processing

    CN120411188A

  • Dynamic scene perception method and system based on event camera and pulse neural network, terminal and medium

    CN120510544A

Cited By

  • Image restoration method and system based on spiking neural network

    CN121190334A

  • An image restoration method and system based on pulse neural network

    CN121190334B