Low-illumination target detection method and system based on spiking neural network

By combining the pre-trained YOLO model with the pulse neural network, and using feature coding and alternative gradient technology to optimize the pulse sequence, the energy efficiency and detection accuracy problems of the pulse neural network in low-illumination environments are solved, and efficient low-illumination object detection is achieved.

CN120298762APending Publication Date: 2025-07-11SICHUAN INFORMATION TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510347072.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing pulse neural network has a balance between energy efficiency and detection accuracy in object detection in low-illumination environments, and the existing methods are susceptible to noise interference in low-light environments, resulting in a decrease in detection accuracy and an increase in energy consumption.

Method used

The pre-trained YOLO model is used as the backbone network, and the continuous feature map is converted into a spatiotemporal pulse sequence through a feature encoding adaptive method, and the pulse sequence is optimized in combination with the Poisson process and time dynamic encoding. The end-to-end pulse network training method is used to jointly optimize based on alternative gradient technology to minimize the total loss function.

Benefits of technology

Achieve high-precision object detection in low-illumination environments, improving detection performance, maintaining low power consumption while improving the robustness and detection accuracy of the network in low-light environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298762A_ABST
    Figure CN120298762A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination target detection method and system based on a pulse neural network, and the method comprises the steps: taking a pre-trained YOLO model as a backbone network, carrying out the feature extraction of an input image through a feature coding adaptive method, converting a continuous feature graph into a space-time pulse sequence, and carrying out the coding optimization of the space-time pulse sequence. A space-time pulse sequence meeting the pulse neural network processing requirement is generated, space-time characteristics of the pulse sequence are optimized through characteristic value normalization, Poisson process pulse generation and time dynamic coding, and the pulse neural network is used for carrying out target positioning and classification on the space-time pulse sequence meeting the pulse neural network processing requirement. A target bounding box and a target category label of a target are generated, the detection performance in a low-illumination environment is improved through an end-to-end training framework, a YOLO backbone network and an SNN detection head are jointly optimized through a substitution gradient technology, a total loss function is minimized, the problem that an SNN pulse mechanism cannot be differentiated is solved, and high-precision detection is achieved under the condition of low power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of spiking neural networks, and more specifically, to a low-light target detection method and system based on spiking neural networks. Background Art

[0002] Target detection is an important research direction in the field of computer vision and is widely used in fields such as autonomous driving, security monitoring, and robot navigation. Traditional artificial neural network (ANN) methods, especially models based on convolutional neural networks (CNNs) and their variants (such as YOLO), have achieved remarkable detection results under normal lighting conditions. However, in low-light environments, due to low signal-to-noise ratio (SNR) and severe noise interference, traditional methods often struggle to maintain high detection accuracy and consume a large amount of computing resources, limiting their application in practical scenarios. In recent years, spiking neural networks (SNNs), as a computational model inspired by the biological nervous system, have gradually attracted attention due to their unique advantages in energy efficiency and temporal information processing. SNNs transmit information through discrete spike signals and adopt an event-driven computational mode, which can significantly reduce power consumption during information processing. In addition, the sparse activation characteristic of SNNs enables them to effectively process temporal patterns and is suitable for real-time processing tasks in dynamic environments. Although SNNs have demonstrated good performance in theory and experiments, their popularization in practical applications such as low-light target detection still faces many challenges.

[0003] Currently, some studies have explored target detection technologies based on deep learning models, including traditional ANN methods and SNN methods. In low-light environments, CNN-based methods usually improve detection performance through image enhancement techniques, but these methods require high computing resources and are prone to failure in high-noise environments. At the same time, the advantages of SNNs in low energy consumption, event-driven computing, and temporal pattern recognition make them a potential solution for low-light target detection. However, existing SNN research mainly focuses on well-lit environments, and there is still relatively limited research on target detection under low-light conditions. In addition, the technology of converting traditional ANN models into SNNs, as a means of leveraging the advantages of SNNs, has received extensive attention. These conversion methods usually rely on surrogate gradients for backpropagation training, but in low-light environments, due to severe noise interference, problems such as time misalignment and membrane potential quantization are likely to occur, resulting in a decline in system performance. Although surrogate gradient technology provides crucial support for the gradient optimization of SNNs and has made some progress in the training of deep SNNs, its application in low-light target detection still has obvious deficiencies.

[0004] In summary, although the existing technologies have made many advances in the fields of object detection and computer vision, there are still obvious deficiencies in the application of SNNs for object detection in low-light environments. There is an urgent need to develop a new solution that can fully leverage the energy efficiency advantages of SNNs and meet the requirements of high-precision detection in low-light environments, so as to promote the wide application of SNNs in practical scenarios. Summary of the Invention

[0005] The purpose of this application is to provide a low-light object detection method and system based on spiking neural networks, which jointly optimize the YOLO backbone network and the SNN detection head through surrogate gradient techniques to minimize the total loss function, effectively solve the problem of non-differentiability of the SNN spiking mechanism, and achieve high-precision detection under low-power conditions, in order to overcome the existing technical deficiencies.

[0006] The purpose of this application is achieved through the following technical solutions:

[0007] In the first aspect, this application proposes a low-light object detection method based on spiking neural networks, and the method includes:

[0008] Using a pre-trained YOLO model as the backbone network, and extracting continuous feature maps from the input image using a feature encoding adaptive method;

[0009] Converting the continuous feature maps into spatio-temporal spike sequences and performing encoding optimization to generate spatio-temporal spike sequences that meet the processing requirements of spiking neural networks, where the encoding optimization includes eigenvalue normalization, spike generation based on the Poisson process, and time dynamic encoding;

[0010] Using a spiking neural network to perform object localization and classification on the spatio-temporal spike sequences that meet the processing requirements of spiking neural networks, and generating object bounding boxes and object category labels;

[0011] Adopting an end-to-end spiking network training method, and performing joint optimization based on surrogate gradient techniques to minimize the total loss function.

[0012] In a possible embodiment, the step of converting the continuous feature maps into spatio-temporal spike sequences and performing encoding optimization to generate spatio-temporal spike sequences that meet the processing requirements of spiking neural networks includes:

[0013] Performing normalization processing on the continuous feature maps to obtain normalized feature maps;

[0014] Mapping the normalized feature values in the normalized feature maps through the Poisson process to obtain spatio-temporal spike sequences that meet the processing requirements of spiking neural networks, where the frequency of the spike sequences is proportional to the intensity of the normalized feature values.

[0015] In a possible embodiment, during the process of obtaining the pulse sequence, the pulse generation probability is: where is the normalized eigenvalue, λ i,j is the firing rate of the normalized eigenvalue, and Δt is the time step.

[0016] In a possible embodiment, time dynamic encoding dynamically adjusts the pulse time interval according to the change of the feature map over time, and the change of the feature map over time is: where represents the eigenvalue at time t, and the pulse time interval is: ∈ is a constant to prevent the denominator from being zero.

[0017] In a possible embodiment, using a spiking neural network to perform object localization and classification on a spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network, and generating the steps of the target bounding box and class label include:

[0018] Using a spiking neural network to capture the target spatial features in the input image according to the spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network;

[0019] Mapping the time series of the spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network to the spatial coordinates of the target position;

[0020] Generating a target bounding box B = f bbox S sail , f bbox represents the bounding box prediction function generated by the pulse sequence, and S sail is the pulse sequence after spatial information encoding;

[0021] Mapping each detected target feature to the activation state of the spiking neuron;

[0022] Using the spiking neural network to process the activation state and generate output pulses related to the category;

[0023] Decoding the output pulses to obtain the target class label f class is the mapping function for the spiking neural network to classify the spatio-temporal pulse sequence, is the category prediction of the target.

[0024] In a possible embodiment, adopting an end-to-end spiking network training method, performing joint optimization based on the surrogate gradient technique, and minimizing the total loss function steps include:

[0025] Adopting a surrogate function to perform continuous approximation on the non-differentiable pulse function of the spiking neuron, and performing backpropagation by calculating the surrogate gradient;

[0026] Construct the total loss function and optimize the network parameters through the total loss function. The total loss function where is the mean squared error loss for bounding box prediction, is the cross-entropy loss for class prediction;

[0027] Jointly optimize the weights of the pulsed backbone network and the pulsed detection head through the gradient descent algorithm, so that the spatial feature extraction ability of the YOLO backbone network and the temporal dynamic processing ability of the pulsed neural network are improved synchronously in end-to-end training.

[0028] In one possible embodiment, the surrogate function is: α is a parameter that controls the approximate steepness, θ is the threshold, and the surrogate gradient is:

[0029] In one possible embodiment, the pre-trained YOLO model includes multiple convolutional layers for extracting hierarchical features from the input image.

[0030] In a second aspect, the present application proposes a low-light target detection system based on a pulsed neural network. The system includes:

[0031] A pulsed backbone network module for using a pre-trained YOLO model as the backbone network and extracting continuous feature maps from the input image using the feature encoding adaptation method;

[0032] A pulsed feature encoding adaptation module for converting the continuous feature map into a spatio-temporal pulse sequence and performing encoding optimization to generate a spatio-temporal pulse sequence that meets the processing requirements of the pulsed neural network. The encoding optimization includes eigenvalue normalization, pulse generation based on the Poisson process, and time dynamic encoding;

[0033] A pulsed detection head module for using the pulsed neural network to perform target localization and classification on the spatio-temporal pulse sequence that meets the processing requirements of the pulsed neural network, and generating the bounding box and class label of the target;

[0034] An end-to-end pulsed network training module for using the end-to-end pulsed network training method, performing joint optimization based on the surrogate gradient technique, and minimizing the total loss function.

[0035] The main solution of the present application and its various further alternative solutions can be freely combined to form multiple solutions, all of which are solutions that can be adopted and claimed by the present application; and in the present application, (each non-conflicting selection) can be freely combined between selections and with other selections. Those skilled in the art can understand that there are multiple combinations according to the prior art and common general knowledge after understanding the solution of the present application, all of which are the technical solutions to be protected by the present application, and are not enumerated here.

[0036] The present application discloses a low - illumination target detection method and system based on a spiking neural network. The pre - trained YOLO model is used as the backbone network. The input image is feature - extracted by using a feature encoding adaptive method, and the continuous feature map is converted into a spatio - temporal spike sequence. The spatio - temporal spike sequence is encoded and optimized to generate a spatio - temporal spike sequence that meets the processing requirements of the spiking neural network. The spatio - temporal characteristics of the pulse sequence are optimized through eigenvalue normalization, Poisson process pulse generation, and time - dynamic encoding. The spiking neural network is used to perform target localization and classification on the spatio - temporal spike sequence that meets the processing requirements of the spiking neural network, generating the target's bounding box and target class label. The detection performance in low - illumination environments is improved through an end - to - end training framework. The YOLO backbone network and the SNN detection head are jointly optimized through surrogate gradient techniques to minimize the total loss function, effectively solving the problem that the spike mechanism of the SNN is non - differentiable, and achieving high - precision detection under low - power conditions. Brief Description of the Drawings

[0037] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 Fig. shows the schematic flowchart of a low - illumination target detection method based on a spiking neural network proposed in an embodiment of the present application.

[0039] Figure 2 Fig. shows the model architecture diagram proposed in an embodiment of the present application.

[0040] Figure 3 Fig. shows the experimental result diagram of the model proposed in an embodiment of the present application on the Dark Face dataset.

[0041] Figure 4 Fig. shows the experimental result diagram of the model proposed in an embodiment of the present application on the ExDark dataset. Detailed Embodiments

[0042] The following illustrates the embodiments of the present application through specific examples. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0043] Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of this application.

[0044] Spiking neural networks (SNNs) are inspired by biological neural mechanisms and have significant advantages over traditional artificial neural networks (ANNs) in terms of energy efficiency and temporal information processing. SNNs mimic the behavior of biological neurons and transmit information through discrete pulses (spikes) rather than continuous signals. This event-driven computing mode enables SNNs to process information more efficiently, with greatly reduced energy consumption compared to ANNs that rely on continuous activation. In addition, the sparse activation of neurons in SNNs also enables them to process temporal patterns in real time, making them particularly suitable for dynamic environments. However, despite the many advantages of SNNs, their application in practical tasks is still relatively limited, especially in the field of low-light object detection. Currently, SNNs for low-light object detection still face the following challenges: how to achieve a balance between energy efficiency and detection accuracy of SNNs in low-light object detection.

[0045] In low-light environments, the energy efficiency advantage of SNNs is challenged by complex processing requirements to compensate for the degraded input quality. Low-light environments introduce severe noise (e.g., thermal and shot noise) that interferes with spike encoding and reduces the signal-to-noise ratio (SNR). For example, random noise spikes may drown out meaningful signals, while faint target features may fail to trigger neuronal activation. In addition, the low contrast between objects and backgrounds requires enhanced spatiotemporal sensitivity to capture subtle visual cues, but existing SNNs often have difficulty in amplifying these features while maintaining energy consumption. In addition, in low-light scenes, sparse spike activations limit the network's ability to extract discriminative features, especially for small or occluded objects. Improving detection accuracy usually requires deeper network architectures or additional preprocessing, which weakens the energy efficiency advantage of SNNs.

[0046] Currently, most methods rely on ANN-to-SNN conversion, which maps continuous ANN activations to discrete spike frequencies. However, this conversion process introduces significant quantization errors, especially in low-light scenarios, where accurate spatiotemporal feature representation is critical. For example, during inference, the conversion of membrane potential to spikes loses fine-grained temporal information, resulting in up to 30% accuracy drop compared to ANNs. Although there have been some breakthroughs in directly training SNNs using surrogate gradients, it still faces the problem of gradient instability in deep networks and a lack of architectures optimized for low-light feature extraction. In addition, hybrid methods that combine SNN backbone networks with ANN detection heads introduce non-spiking computations, weakening the energy efficiency advantage of SNNs. These limitations highlight the urgent need for an end-to-end, fully spiked framework to avoid conversion errors and achieve efficient low-light object detection.

[0047] The present application discloses a method and system for low-light target detection based on a spiking neural network. By introducing a dedicated end-to-end SNN training framework and combining it with an improved feature encoding technology - converting the high-quality spatial features of the pre-trained YOLO model into a spatiotemporal spike sequence, the detection performance of the network under low signal-to-noise ratio conditions is effectively improved, thereby providing a competitive solution for low-light target detection that takes into account both energy efficiency and detection accuracy.

[0048] Please refer to Figure 1 , Figure 1 A schematic diagram of a low-light target detection method based on a pulse neural network proposed in an embodiment of the present application is shown. Figure 2 The model architecture diagram proposed in the embodiment of the present application is shown, wherein the input is a low-light image (Low-Light Image), and feature signals (SpikesSignal) and feature maps (Feature) are extracted from the low-light image through feature extraction (FEA), and these feature signals and feature maps constitute the input representation stage.

[0049] The extracted feature signals and feature maps then enter the backbone network (Backbone), which consists of multiple processing layers to gradually extract and enhance the features of the image. The feature maps processed by the backbone network enter the LIF (Leaky Integrate-and-Fire) layer for further processing and optimization of the feature maps.

[0050] The feature map output by the LIF layer enters multiple multi-layer perceptron (MLP) modules. These MLP modules process different parts or features at different levels of the feature map respectively, and finally generate an output image (Output). The output image is enhanced, with higher brightness and clearer details, thereby improving the quality of low-light images.

[0051] The low-light target detection method includes the following steps:

[0052] Step S1: Use the pre-trained YOLO model as the backbone network, and use the feature encoding adaptive method to extract features from the input image to obtain a continuous feature map.

[0053] The low-light image to be detected is input into the pre-trained YOLO model. The image enters the model as raw data and begins to be processed by a series of convolutional layers. Each convolutional layer of the YOLO model performs a convolution operation on the input image. The convolution kernel slides on the image, and mathematical operations are performed on the convolution kernel and the pixel values ​​of the local area of ​​the image to extract features at different levels of the image. As data is passed between convolutional layers, the feature information in the image is continuously mined and integrated.

[0054] After being processed by all the convolutional layers in the YOLO backbone network, one or more feature maps are finally output. These feature maps are data structures in the form of two-dimensional matrices, where each element f i,j represents the feature vector extracted from the image at specific positions i, j. Since these features are real values calculated by the convolutional neural network, the elements in the feature map are all continuous values, so it is called a continuous feature map. This continuous feature map contains rich spatial information of the input image and the object feature representation learned by the model, but due to its continuity, it cannot be directly applied to the spiking neural network (SNN), and further processing by the feature encoding adaptation method is required later.

[0055] The pre-trained YOLO model includes multiple convolutional layers for extracting hierarchical features from the input image.

[0056] The YOLO model contains multiple convolutional layers, which are the core components for effectively extracting image features. The convolutional layer slides a convolutional kernel over the input image to perform a convolution operation. During each sliding process, the convolutional kernel performs an element-wise multiplication operation with the corresponding local region of the image and sums the results to generate the feature value of that region. As the convolutional layer continuously slides over the image, it can traverse the entire image and extract rich local feature information.

[0057] Step S2: Convert the continuous feature map into a spatio-temporal spike train and perform encoding optimization to generate a spatio-temporal spike train that meets the processing requirements of the spiking neural network. The encoding optimization includes eigenvalue normalization, Poisson process-based spike generation, and temporal dynamic encoding.

[0058] To convert the continuous feature map output by the YOLO model into a spatio-temporal spike train suitable for processing by the spiking neural network (SNN), encoding optimization is required, which mainly includes three key steps: eigenvalue normalization, Poisson process-based spike generation, and temporal dynamic encoding. First, perform eigenvalue normalization. Since the eigenvalue range and distribution in the continuous feature map output by YOLO may not be suitable for the SNN, it is mapped to a suitable interval through methods such as linear normalization to obtain a normalized feature map. Then, generate spikes based on the Poisson process. For each eigenvalue in the normalized feature map, define the spike occurrence probability according to a specific formula, and determine whether to generate a spike at each time step to form the corresponding spatio-temporal spike train, enabling the SNN to effectively process spatial feature information. Finally, perform temporal dynamic encoding. Considering the temporal dynamic characteristics of the scene, calculate the spike time interval according to the change of the feature map over time, and dynamically adjust the spike generation time, allowing the SNN to capture the change characteristics of the target in the time dimension and improve the accuracy of target detection in low-light environments.

[0059] Step S2 includes:

[0060] Normalize the continuous feature map to obtain a normalized feature map;

[0061] Map the normalized feature values in the normalized feature map through a Poisson process to obtain a spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network. The frequency of the pulse sequence is proportional to the intensity of the normalized feature value.

[0062] After normalizing the continuous feature map, the feature values f i,j in the feature map F output by the pre-trained YOLO model are continuous real numbers, while the spiking neural network (SNN) usually requires discrete input data suitable for the pulse generation mechanism. The normalization process can map the feature values to a suitable range to meet the conditions for pulse generation, facilitating subsequent conversion into a spatio-temporal pulse sequence that can be processed by the SNN. The specific operation is to perform a normalization operation on each feature value f i,j in the feature map F to obtain the normalized feature map i,j where is the normalized feature value. is the normalized feature value.

[0063] After that, map the normalized feature values in the normalized feature map through a Poisson process to obtain a spatio-temporal pulse sequence. By using the Poisson process to generate the pulse sequence, a connection is established between the normalized feature value and the probability of pulse occurrence. The mapping process is as follows: for each normalized feature value in the normalized feature map map it through the Poisson process to a spatio-temporal pulse sequence S i,j (t), and the pulse frequency of this pulse sequence is proportional to the intensity of the normalized feature value.

[0064] Time dynamic encoding dynamically adjusts the pulse time interval according to the change of the feature map over time. The change of the feature map over time is: where represents the feature value at time t, and the pulse time interval is: ∈ is a constant to prevent the denominator from being zero.

[0065] is a set of feature maps that change over time t. At each moment t, there is a corresponding feature map, and each element in the feature map represents the feature value at the grid position i, j at time t. This feature value is obtained after previous feature extraction, normalization, etc., and it contains an abstract representation of the image information at that position.

[0066] The pulse occurrence probability is: Δt, where is the normalized eigenvalue, λ i,j is the firing rate of the normalized eigenvalue, and Δt is the time step. Within each time step Δt, according to the above probability formula, it is determined whether to generate a pulse, thus forming a pulse sequence S i,j (t). In this way, each normalized eigenvalue corresponds to a spatio-temporal pulse sequence, enabling the SNN to receive and process the spatial feature information originally extracted by YOLO in a manner suitable for its processing. Finally, the continuous feature map output by YOLO can be converted into a spatio-temporal pulse sequence suitable for processing by the spiking neural network, and the frequency characteristics of this pulse sequence are associated with the intensity of the original eigenvalue. On the basis of retaining the spatial feature information, the SNN can effectively process this information, laying a foundation for subsequent object detection in low-light environments.

[0067] Step S3: Use the spiking neural network to perform object localization and classification on the spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network, and generate the object bounding box and the object class label.

[0068] The spiking neural network uses the spatio-temporal pulse sequence that meets its processing requirements to perform object localization and classification. In terms of object localization, the network first processes these spatio-temporal pulse sequences, extracts the key feature information reflecting the object position from them, and converts it into spatial coordinate information, thereby determining the range of the object in the image, and finally generating the object bounding box. In the object classification link, the network maps the object features to the activation state of the spiking neurons, generates output pulses related to different categories based on this, and then identifies the category to which the object belongs and outputs the category label through a decoding operation.

[0069] The steps of step S3 include:

[0070] Use the spiking neural network to capture the object spatial features in the input image according to the spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network;

[0071] Map the time sequence of the spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network to the spatial coordinates of the object position;

[0072] Generate the object bounding box B = f bbox S sail f bbox represents the bounding box prediction function generated by the pulse sequence, and S sail is the pulse sequence after spatial information encoding;

[0073] Map each detected object feature to the activation state of the spiking neurons;

[0074] Use the spiking neural network to process the activation state and generate output pulses related to the category;

[0075] Decode the output pulse to obtain the target class label f class is the mapping function for the spiking neural network to classify spatio-temporal pulse sequences, which is the class prediction of the target.

[0076] The spiking neural network captures the target spatial features in the input image based on the spatio-temporal pulse sequence that meets the processing requirements of the spiking neural network. This process uses spiking neurons to process the input spatio-temporal pulse sequence and extract the information related to the target spatial position from it. Then, the temporal information of the spatio-temporal pulse sequence is mapped to the spatial coordinates of the target position. Through specific mapping rules, the time characteristics carried by the pulse sequence are converted into the coordinate information of the position where the target is located in the image, so as to determine the spatial position of the target in the image. Based on the captured target spatial features and the mapped target position spatial coordinates, the bounding box of the target is generated. This process is achieved through the bounding box prediction function f bbox which acts on the pulse sequence S after spatial information encoding sail , and finally obtains the coordinates B=(x min ,y min ,x max ,y max ) representing the bounding box of the target, thereby clarifying the position and range of the target in the image.

[0077] For object classification, each detected target feature is mapped to the activation state of the spiking neuron. The various feature information of the target is converted into the activation state that the spiking neuron can process, and the spiking neural network is used to process the neurons in the activated state to generate output pulses related to the class. Based on its unique computational mechanism, the spiking neural network performs operations and processing on these activation states and outputs pulse signals corresponding to different target classes. Decoding operations are performed on the generated output pulses related to the class, so as to obtain the class label of the target. Through specific decoding algorithms and rules, the information that can represent the class to which the target belongs is extracted from the output pulses, and the classification and recognition of the target are realized.

[0078] Finally, the pulse detection head module outputs the bounding box coordinates (calculated by B = f bbox S sail ) and class labels (calculated by ) of each target, enabling Dark-SNN to efficiently complete the target localization and classification tasks in low-light environments.

[0079] Step S4: Adopt an end-to-end spiking network training method, perform joint optimization based on the surrogate gradient technique, and minimize the total loss function.

[0080] When dealing with complex tasks such as object detection in low-light environments, traditional training methods may require training different network modules separately and then undergoing a complex integration process. The end-to-end spiking neural network training method abandons this cumbersome approach. It regards the entire network architecture as a unified whole and conducts training from the input data (such as images) to the final output results (such as the location and category of objects) throughout the process. The advantage of this method is that it can better capture the mutual relationships and dependencies among various parts of the network, enabling the performance of the entire network to be more fully optimized.

[0081] When dealing with complex tasks (such as object detection in low-light environments), the end-to-end spiking neural network training method regards the entire network architecture as a unified whole for training, optimizes the whole process from input data to final output, can better capture the mutual relationships among various parts of the network, and improves the network performance.

[0082] In the joint optimization process based on surrogate gradient techniques, due to the non-differentiability caused by the discrete spike mechanism of the spiking neural network (SNN), the traditional gradient backpropagation algorithm cannot be directly applied. The surrogate gradient technique designs an approximate differentiable function to replace the non-differentiable spike function and calculates the approximate gradient in backpropagation to update the weights. In the end-to-end spiking neural network training, considering the collaborative work of various parts of the network comprehensively, it uses the surrogate gradient to calculate the parameter gradient and update the weights, and through multiple iterations, makes the network converge to a better state.

[0083] The total loss function measures the difference between the network prediction result and the true label by considering multiple factors. In the object detection task, it usually includes sub-loss functions such as localization loss and classification loss. Minimizing the total loss function is the core goal of training, which can make the network prediction closer to the true label and improve the accuracy and robustness of object detection in low-light environments.

[0084] Step S4 includes:

[0085] Use a surrogate function to make a continuous approximation of the non-differentiable spike function of the spiking neuron and perform backpropagation by calculating the surrogate gradient;

[0086] Construct the total loss function and optimize the network parameters through the total loss function. The total loss function where is the mean squared error loss for bounding box prediction, is the cross-entropy loss for class prediction;

[0087] Jointly optimize the weights of the spiking backbone network and the spiking detection head through the gradient descent algorithm, so that the spatial feature extraction ability of the YOLO backbone network and the temporal dynamic processing ability of the spiking neural network are improved synchronously in the end-to-end training.

[0088] The surrogate function is: α is a parameter that controls the approximation steepness, θ is the threshold, and the surrogate gradient is:

[0089] The end-to-end spiking neural network training framework uses the surrogate gradient to approximate the gradient of the non-differentiable spiking function. In a traditional SNN, when the membrane potential of a neuron exceeds the threshold, it emits a spike, i.e., u(t) > θ is satisfied, where θ is the threshold. The spike train S(t) is: Since S(t) is non-differentiable, the surrogate function is used for approximation, such as the following continuous approximation: where α is a parameter that controls the approximation steepness. The surrogate gradient is used for backpropagation to calculate the gradient of the weight update in the SNN:

[0090] In the end-to-end spiking neural network training framework, the entire Dark-SNN architecture is trained end-to-end from the input image to the final output prediction without manual conversion from ANN to SNN. The network uses the surrogate gradient to train the backbone network (YOLO feature extraction) and the spiking detection head (object detection and classification), ensuring that the temporal dynamics of the SNN can directly optimize object detection in low-light environments. Let be the total loss function, which consists of the localization loss and the classification loss : where is the loss related to object localization, usually measured by the mean squared error (MSE) between the predicted bounding box and the ground truth bounding box; is the loss related to object classification, usually calculated using the cross-entropy loss. The goal of the E2EST framework is to minimize the total loss function and optimize the weights of the SNN using the surrogate gradient through backpropagation. For each spike train generated by the SNN, the gradient with respect to the loss function is calculated and the weights are updated accordingly.

[0091] During training, the surrogate gradient propagates backward through the network, and the weights are updated using a gradient descent-based optimization algorithm. The end-to-end nature of the E2EST framework ensures that the spatial and temporal aspects of the Dark-SNN architecture can be jointly optimized. This makes the model more effective in object detection in low-light environments because it can simultaneously learn the feature extraction ability of YOLO and the temporal processing ability of the SNN in a single unified training process.

[0092] Figure 3 Shows the experimental result graph of the model proposed in the embodiment of the present application on the Dark Face dataset, Figure 4The figure shows the experimental results of the model proposed in the embodiments of the present application on the ExDark dataset. A large number of experiments have been carried out on the ExDark and Dark Face datasets, demonstrating that Dark-SNN performs excellently among 14 state-of-the-art methods. To evaluate the effectiveness of Dark-SNN, we conducted extensive experiments on two widely used datasets, ExDark and Dark Face, which are specifically designed for low-light object detection. These datasets contain images with different noise and lighting challenges. The experimental results show that Dark-SNN outperforms multiple existing methods in terms of detection accuracy and localization accuracy, demonstrating the strong capabilities of the proposed method in real low-light scenarios.

[0093] In summary, although spiking neural networks (SNNs) have advantages in energy efficiency and temporal information processing, their practical applications are limited. Although the sparse activation of their neurons is beneficial for real-time processing of temporal patterns, the ability of the network to extract discriminative features is limited, and it is difficult to detect small or occluded objects; improving accuracy often requires a deeper network structure or additional preprocessing, which will weaken the energy efficiency advantage. Moreover, the low-light environment introduces severe noise (thermal noise and scattering noise), which interferes with pulse coding, reduces the signal-to-noise ratio, and makes it difficult for weak target features to trigger neuron activation; the low contrast between objects and the background requires enhanced spatio-temporal sensitivity, and existing SNNs are difficult to amplify features without increasing energy consumption; the sparse pulse activation limits the network's ability to extract discriminative features. At the same time, most existing methods rely on the conversion from ANN to SNN, introducing significant quantization errors, which affect the accuracy in low-light scenarios; directly training SNNs using surrogate gradients shows potential, but faces challenges such as unstable gradients in deep networks and the lack of an architecture customized for low-light feature extraction; the hybrid method combining the SNN backbone and the ANN detection head eliminates the energy efficiency advantage of SNNs.

[0094] The Dark-SNN architecture proposed in this application is a spiking neural network (SNN) architecture specifically designed to address the challenges of low-light object detection. It consists of a backbone network (YOLO) for feature extraction and an SNN head for object detection and classification. An adaptive method based on pulse-based feature encoding is used to convert the output of the pre-trained YOLO model into a spatio-temporal pulse sequence. It combines the strong spatial feature extraction ability of YOLO with the energy-saving temporal encoding ability of SNN to form a collaborative mechanism. By converting the high-quality spatial features of YOLO into a temporal pulse pattern, Dark-SNN can not only utilize the spatial understanding ability of YOLO but also give full play to the inherent temporal dynamic processing ability of SNN to achieve efficient detection with the lowest energy consumption.

[0095] The end-to-end SNN training framework uses surrogate gradient training to eliminate problems in the process of converting traditional ANNs to SNNs, such as time misalignment and membrane potential quantization errors. Traditional SNN training methods for converting pre-trained ANNs to SNNs are full of inaccuracies, leading to performance degradation. The EEST framework can directly optimize the dynamic characteristics of SNNs, learn more effectively without conversion, enhance the robustness of Dark-SNN, and make it particularly suitable for low-light object detection tasks.

[0096] A large number of experiments were conducted on two widely used datasets, ExDark and Dark Face, which are designed for low-light object detection. These datasets contain images with different noise levels and lighting challenges. The experimental results show that Dark-SNN outperforms many existing methods in terms of both detection accuracy and object localization accuracy, and performs excellently among 14 state-of-the-art methods, fully verifying the excellent performance of the proposed method in real low-light scenarios.

[0097] The following gives a possible implementation of a low-light object detection system based on a spiking neural network, which is used to execute each execution step and corresponding technical effects of the low-light object detection method shown in the above embodiments and possible implementations. The system includes:

[0098] A spiking backbone network module, which uses a pre-trained YOLO model as the backbone network and uses a feature encoding adaptation method to extract features from the input image to obtain a continuous feature map;

[0099] A spiking feature encoding adaptation module, which is used to convert the continuous feature map into a spatio-temporal spike sequence and perform encoding optimization to generate a spatio-temporal spike sequence that meets the processing requirements of the spiking neural network. The encoding optimization includes eigenvalue normalization, spike generation based on the Poisson process, and time dynamic encoding;

[0100] A spiking detection head module, which uses a spiking neural network to perform object localization and classification on the spatio-temporal spike sequence that meets the processing requirements of the spiking neural network, and generates the bounding box and class label of the object;

[0101] An end-to-end spiking network training module, which uses an end-to-end spiking network training method to perform joint optimization based on surrogate gradient technology to minimize the total loss function.

[0102] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A low-light target detection method based on spiking neural networks, characterized in that, The method includes: Using a pre-trained YOLO model as the backbone network, and extracting continuous feature maps from the input image by using a feature encoding adaptation method; Converting the continuous feature maps into spatio-temporal pulse sequences and performing encoding optimization to generate spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network, where the encoding optimization includes eigenvalue normalization, Poisson process-based pulse generation, and time dynamic encoding; Using the spiking neural network to perform object localization and classification on the spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network, and generating object bounding boxes and object class labels; Adopting an end-to-end spiking network training method, and performing joint optimization based on the surrogate gradient technique to minimize the total loss function.

2. The low-light target detection method according to claim 1, wherein The step of converting the continuous feature maps into spatio-temporal pulse sequences and performing encoding optimization to generate spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network includes: Performing normalization processing on the continuous feature maps to obtain normalized feature maps; Mapping the normalized eigenvalues in the normalized feature maps through a Poisson process to obtain spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network, where the frequency of the pulse sequences is proportional to the intensity of the normalized eigenvalues.

3. The low-light target detection method according to claim 2, wherein, During the process of obtaining the pulse sequence, the pulse generation probability is: where is the normalized eigenvalue, λ i,j is the pulse firing rate of the normalized eigenvalue, and Δt is the time step.

4. The low-light target detection method according to claim 1, characterized in that Temporal dynamic encoding dynamically adjusts the pulse time interval according to the change of the feature map over time, and the change of the feature map over time is as follows: where represents the feature value at time t, and the pulse time interval is: ∈ is a constant to prevent the denominator from being zero.

5. The low-light target detection method according to claim 1, characterized in that The step of using the spiking neural network to perform object localization and classification on the spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network and generating object bounding boxes and class labels includes: Using the spiking neural network to capture the target spatial features in the input image according to the spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network; Mapping the time sequence of the spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network to the spatial coordinates of the target positions; Generate the target bounding box B = f based on the target space features and the spatial coordinates of the target position bbox S sail , f bbox represents the bounding box prediction function generated by the pulse sequence, and S sail is the pulse sequence after spatial information encoding; Mapping each detected target feature to the activation state of the spiking neuron; Using the spiking neural network to process the activation state and generate output pulses related to the class; Decode the output pulse to obtain the target class label is the mapping function for the spiking neural network to classify spatio-temporal pulse sequences, which is the class prediction of the target.

6. The low-light target detection method according to claim 4, characterized in that The step of adopting an end-to-end spiking network training method and performing joint optimization based on the surrogate gradient technique to minimize the total loss function includes: Adopting a surrogate function to perform continuous approximation on the non-differentiable pulse function of the spiking neuron, and performing backpropagation by calculating the surrogate gradient; Construct the total loss function and optimize the network parameters through the total loss function. The total loss function is where is the mean squared error loss for bounding box prediction, is the cross-entropy loss for class prediction; Jointly optimizing the weights of the spiking backbone network and the spiking detection head through the gradient descent algorithm, so that the spatial feature extraction ability of the YOLO backbone network and the temporal dynamic processing ability of the spiking neural network are synchronously improved in the end-to-end training.

7. The low-light target detection method according to claim 6, characterized in that, The substitution function is: α is a parameter that controls the approximate steepness, θ is the threshold, and the substitution gradient is:

8. The low-light target detection method according to claim 1, characterized in that The pre-trained YOLO model includes multiple convolutional layers for extracting hierarchical features from the input image.

9. A low-light target detection system based on a spiking neural network, characterized in that, The system includes: A spiking backbone network module for using a pre-trained YOLO model as the backbone network and extracting continuous feature maps from the input image by using a feature encoding adaptation method; A spiking feature encoding adaptation module for converting the continuous feature maps into spatio-temporal pulse sequences and performing encoding optimization to generate spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network, where the encoding optimization includes eigenvalue normalization, Poisson process-based pulse generation, and time dynamic encoding; A spiking detection head module for using the spiking neural network to perform object localization and classification on the spatio-temporal pulse sequences that meet the processing requirements of the spiking neural network and generating the bounding boxes and class labels of the objects; An end-to-end spiking network training module, which is used to perform joint optimization based on surrogate gradient techniques by adopting an end-to-end spiking network training method to minimize the total loss function.