A binocular event depth estimation method based on spiking neural networks

By introducing feature supervision module and edge regularization module into deep pulse neural networks, the problem of performance degradation of SNN on complex tasks is solved, and feature extraction capabilities and visual effects of edge regions are improved.

CN114926517BActive Publication Date: 2025-06-10NANJING UNIV

Patent Information

Application Number
CN202210548361.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-06-10
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Deep pulse neural networks (SNNs) deteriorate their performance on complex tasks such as stereo event matching, resulting in a decline in feature extraction capabilities and hindering their application on complex tasks.

Method used

The feature supervision module is designed using knowledge distillation technology to transfer the information of the convolutional neural network (CNN) to the SNN, helping SNN to generalize better, and using the binary characteristics of edge maps and pulses in the edge regularization module to improve the visual effect of the edge area of ​​the depth map.

Benefits of technology

Through the combination of knowledge distillation technology and edge regularization module, the feature extraction and generalization performance of SNN in depth map estimation is improved, especially in edge regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926517B_ABST
    Figure CN114926517B_ABST
Patent Text Reader

Abstract

The present invention discloses a binocular event depth estimation method based on a spiking neural network. The steps are as follows: (1) Select a binocular event depth dataset; (2) Construct a spiking neural network based on UNet, including an encoder, a bottleneck module, a decoder, a feature supervision module, and an edge regularization module; (3) Input the binocular events into the neural network, extract multi-scale encoded spikes and encoded features, multi-stage decoded spikes and edge prediction spikes, and input the multi-stage decoded spikes into an infinite threshold non-spiking neuron in sequence. The potential value of the non-spiking neuron is added to the potential value of the edge prediction neuron to obtain a depth map; (4) Construct a loss function for the neural network, including a depth map regression loss, a feature absolute value loss, and an edge binary cross-entropy loss, and train the neural network according to the loss function; (5) Input the binocular events of the test set into the trained neural network to obtain a depth map. The present invention can solve the problem of performance degradation of deep SNNs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bio-inspired vision and deep learning, and particularly to a binocular event depth estimation method based on a spiking neural network. Background Art

[0002] In high-speed applications, traditional cameras capture two-dimensional images, which usually have problems such as blurring and low dynamic range. Inspired by biological vision, event cameras have been proposed. These are cameras that can detect changes in the logarithmic illumination of the environment and output an asynchronous binary event stream, with the advantages of high dynamic range, low latency, and low power consumption. With the wide application of event cameras in fields such as image reconstruction and optical flow estimation, binocular event camera systems have been proposed for stereo depth estimation. At the same time, many algorithms based on Convolutional Neural Network (CNN) have achieved good results. However, algorithms based on convolutional neural networks require a large amount of computational cost, greatly weakening the main advantages of event cameras.

[0003] Recently, various studies have shown that spiking neural networks (SNNs), due to their characteristic of only outputting binary spikes in each layer, can be very conveniently directly applied to low-power neuromorphic vision systems. Therefore, compared with convolutional neural networks, they can perfectly match binary event-based visual scenes. However, the phenomenon of spike disappearance in deep SNNs causes the network to be unable to effectively learn the features of the input, resulting in performance degradation and hindering applications in complex tasks such as stereo event matching. Therefore, how to improve the feature extraction ability of SNNs and better generalize is an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a binocular event depth estimation method based on a spiking neural network. To solve the performance degradation problem of deep SNNs, the present invention first uses knowledge distillation technology and designs a feature supervision module to transfer the rich information of CNNs into SNNs, thereby helping SNNs better generalize. Secondly, to achieve better visual effects in the edge region, an edge regularization module is designed.

[0005] The technical solution adopted by the method of the present invention is as follows:

[0006] A binocular event depth estimation method based on a spiking neural network, comprising the following steps:

[0007] Step 1, select a binocular event depth data set and divide it into a training set and a test set; the data set includes binocular events, APS images, and depth maps;

[0008] Step 2: Construct a spiking neural network based on UNet, including an encoder, a bottleneck module, a decoder, a feature supervision module, and an edge regularization module; the encoder, the bottleneck module, the decoder, and the edge regularization module are connected in sequence, and the feature supervision module is connected to the encoder;

[0009] Step 3: Input binocular events into the spiking neural network. The binocular events pass through the encoder to output multi-scale encoded spikes, and then the encoded spikes are input into the feature supervision module to obtain spike features; the encoded spikes pass through the bottleneck module and the decoder to obtain multi-stage decoded spikes, and the multi-stage decoded spikes are input into the edge regularization module to obtain multi-stage edge prediction spikes. Then, the multi-stage decoded spikes are respectively input into non-spiking neurons with an infinite threshold and the multi-stage edge prediction spikes are input into spiking neurons, and the potential values of the two types of neurons are added to obtain a depth map;

[0010] Step 4: Construct a loss function for the spiking neural network, including a depth map regression loss, a feature absolute value loss, and an edge binary cross-entropy loss, and train the spiking neural network according to the loss function;

[0011] Step 5: Input the binocular events of the test set into the trained spiking neural network, and pass through the encoder, the bottleneck module, the decoder, and the edge regularization module to obtain a depth map.

[0012] Further, in Step 1, the dataset also includes a binary edge map obtained by inputting a depth map and an APS image into an edge detection network for processing.

[0013] Further, in Step 2, the encoder uses convolution with a stride of 2 for downsampling; the bottleneck module adopts a general residual structure; the decoder uses nearest neighbor 2D upsampling and convolution with a stride of 1 to upsample the encoded spikes passing through the encoder and the bottleneck module to the size of the depth map.

[0014] Further, in Step 2, the feature supervision module includes convolution with a stride of 1 and ReLU operations to learn the floating-point features of the multi-scale encoded spikes output by the encoder.

[0015] Further, in Step 2, the edge regularization module upsamples the multi-stage decoded spikes of the decoder to predict a binary edge map.

[0016] The present invention has the following advantages compared with the prior art:

[0017] (1) The present invention is an end-to-end model that can directly output a depth map according to the input binocular events.

[0018] (2) When training the SNN, the knowledge distillation technique is used to guide the SNN to learn by using the existing general CNN. During inference, it remains a pure SNN structure.

[0019] (3) The proposed edge regularization module effectively utilizes the characteristic that both the edge map and the pulse are binary, improving the visual effect of the edge region of the depth map. Brief Description of the Drawings

[0020] Figure 1 is a schematic flowchart of the method of the present invention;

[0021] Figure 2 is a flowchart of obtaining a binary edge map in the network of the present invention;

[0022] Figure 3 is an overall network structure diagram of the method of the present invention;

[0023] Figure 4 is a flowchart of the neuron accumulation potential in the network of the present invention;

[0024] Figure 5 is the depth map estimation result of the method of the present invention on the dataset MVSEC. Detailed Embodiments

[0025] The present invention will be described in detail below in conjunction with the drawings and embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0026] This embodiment provides a binocular event depth estimation method based on a spiking neural network. The specific steps are as Figure 1 shown and include the following steps:

[0027] (1) Select a binocular event depth dataset and divide it into a training set and a test set. The dataset contains binocular events, APS images, and depth maps. At the same time, a binary edge map is constructed according to the APS images and depth maps.

[0028] (2) Construct a spiking neural network based on UNet, including an encoder, a bottleneck module, and a decoder, and add a feature supervision module and an edge regularization module.

[0029] (3) Input binocular events into the neural network. The events pass through the encoder to output multi-scale encoded pulses, and the encoded pulses are input into the feature supervision module to obtain pulse features. The encoded pulses pass through the bottleneck module and the decoder to obtain multi-stage decoded pulses. At the same time, the intermediate results of the multi-stage decoded pulses are input into the edge regularization module to obtain multi-stage edge prediction pulses. The multi-stage decoded pulses are respectively input into the infinite threshold non-spiking neurons and the multi-stage edge prediction pulses are input into the spiking neurons, and then the potential values of the two types of neurons are added to obtain the depth map.

[0030] (4) Construct the loss function of the neural network, including depth map regression loss, feature absolute value loss, and edge binary cross-entropy loss, and train the neural network according to the loss function. Input the pulse features obtained by the feature supervision module in step (3) and the features output by the APS image through the general convolutional neural network into the feature absolute value loss to train the network and obtain better results.

[0031] (5) Input the binocular events of the test set into the trained neural network, and obtain the depth map through the encoder, bottleneck module, decoder, and edge regularization module.

[0032] 1. Describe in detail the part of the database construction involved in step (1):

[0033] The dataset used in this embodiment is MVSEC, and the dataset provides binocular events, APS images, and depth maps. According to the requirements of the edge regularization module in the neural network model, a binary edge map is constructed for subsequent supervision of the edge regularization module. Specifically, as Figure 2 shown, use the edge detection network HEDNet to detect the edges of the real depth map. Since the edge values detected by HEDNet range from [0, 255], a threshold of 25 is manually selected here, and the part with an edge value greater than the threshold is set to 1, and the part less than the threshold is set to 0. Since there are some invalid points in the depth map, invalid edges will be detected. To solve this problem, the APS image is also edge-detected and used as a mask to remove the invalid points in the depth map edges. In addition, to further remove the edges around the invalid points in the depth map, the invalid points of the depth map are extracted and dilated, and used as a mask to further remove the invalid points in the edge map. This process can be represented by the following formula

[0034]

[0035] where S d is the result obtained by using HEDNet to detect the edges of the real depth map, S img is the result obtained by using HEDNet to detect the edges of the APS image, M final is the finally used edge map, Bin τ(·) is the threshold-based binarization factor, τ 1 and τ 2 are thresholds, ⊙ is the dot product operation, M false is the mask of invalid points in the depth map, and Dilation(·) is the dilation operation. Here, τ 1 is selected as 0, and τ 2 is 25.

[0036] 2. The overall network structure of this embodiment is as shown in Figure 3 . First, the working mode of the SNN will be introduced below, and then the encoder, bottleneck module, decoder, feature supervision module, and edge regularization module involved in step (2) will be described in detail respectively.

[0037] The event output format of the event camera is (t, x, y, p), where t is the timestamp, (x, y) is the event coordinate, and p is the polarity of the event. For convenient network processing, events within a period of time in a single view are accumulated to form an event stack with a size of [2, H, W], where H and W are the resolutions of the event camera, and the value of each pixel point is the total number of events occurring. The binocular event stacks are stacked to obtain the input of the network as [4, H, W].

[0038] In this embodiment, the basic component neuron of the spiking neural network is the IF neuron. The IF neuron can receive weighted spikes from the previous layer neurons, update the potential. If the potential is greater than the threshold, it outputs a spike and resets the potential to 0, otherwise it retains the original spike. The working mode of the neuron in the l + 1 layer can be expressed as

[0039]

[0040]

[0041] where t is the current time, V is the potential value, o(·) is the output spike, W(i, j) is the weight value between the j-th neuron in the l-th layer and the i-th neuron in the l + 1-th layer, f(·) is the Dirac function, and V th is the threshold, which is set to 1 in this example. Since the Dirac function is not differentiable, in this embodiment, the arctan function is used to replace the Dirac function

[0042]

[0043] and its gradient is

[0044]

[0045] 2.1 Encoder, used to extract high-dimensional features of input events. The event stack passes through a convolutional layer and IF neurons to extract initial features, and then through a combination of 4 convolutional layers with a stride of 2 and IF neurons to gradually extract high-dimensional pulse features.

[0046] 2.2 Bottleneck module, containing 2 basic residual connection blocks, and each residual connection block contains a combination of 2 convolutions and IF neurons. This module further extracts pulse features.

[0047] 2.3 Decoder, using 2D nearest neighbor upsampling, convolution with a stride of 1 and IF neurons to gradually upsample the pulses to the size of the depth map. The result obtained after each upsampling by the decoder will be directly upsampled to the size of the depth map through upsampling and convolution, and the obtained floating-point values will be sequentially input into the final non-pulse neurons to obtain the potential values of 4 stages. The last potential value is used as the initial depth map result.

[0048] 2.4 Feature supervision module, accumulating the pulses obtained in 4 stages of the encoder in the time dimension respectively, and inputting them into a combination of convolution with a stride of 1 and ReLU, called the conversion module, to obtain the features of the pulses

[0049]

[0050] where t 0 is the timestamp corresponding to the depth map during training, T is the accumulated time, and the output feature f and the 4-dimensional feature maps obtained from the APS image through VGGNet are input into the feature absolute value loss to train the parameters of the network.

[0051] 2.5 Edge regularization module, directly upsampling the result obtained after each upsampling by the decoder to the size of the depth map through upsampling and convolution, and sequentially inputting it into the final pulse neurons, as Figure 4 shown. For example, after the value of stage 1 is input into the pulse neuron, the neuron state changes, obtaining the output pulse and potential value of stage 1; the value of stage 2 continues to be input into the pulse neuron, and the neuron state changes, obtaining the output pulse and potential value of stage 2. The output pulses of 4 stages will all be compared with the edge map.

[0052] 3. Provide a detailed description of the part of training the neural network involved in step (4):

[0053] In this embodiment, three loss functions are used, namely depth map regression loss, feature absolute value loss, and edge binary cross-entropy loss.

[0054] The depth map regression loss adopts scale-invariant loss and multi-scale-scale-invariant gradient matching loss, that is, given the output depth map and the label depth map d, let the residual The scale - invariant loss is

[0055]

[0056] where N is the number of valid pixels u in the label depth map. The gradient - matching loss is

[0057]

[0058] where is the gradient calculated using the Sobel operator. Therefore, the depth - map regression loss is

[0059]

[0060] The feature absolute - value loss selects the SmoothL1 loss, that is, for the feature f extracted by VGGNet and the impulse feature obtained from formula (5) its loss is

[0061]

[0062] where M is the total number of pixels in the three - dimensional feature map, and (i, j, k) are the pixel coordinates.

[0063] The edge binary cross - entropy loss adopts the weighted cross - entropy loss, that is, for the output impulse of the final impulse neuron and the binary edge map y, its loss is

[0064]

[0065] where y + is the total number of pixels with edges, y - = N - y + . To balance positive and negative samples, a weight - balancing factor α = y - / y + is added. Since the output of the edge regularization module is already binary, in order to reduce the number of training weights, in this embodiment, the common practice of adding a convolutional layer and a Sigmoid layer is not adopted, but and are directly set, where σ(·) represents Sigmoid.

[0066] The total loss is l total = l si + λ 0 l grad + λ 1 l FSM + λ 2 l MRM . In addition, since the feature supervision module, the decoder, and the edge regularization module all have four - stage sequential outputs, therefore, at each stage, the corresponding l needs to be calculatedsi , l grad , l FSM and l MRM , are added together to obtain the final loss. The parameter is selected as λ 0 = 0.5, λ 1 = 1.0, λ 2 = 0.75.

[0067] The entire network is trained in an end-to-end manner without step-by-step training. The PyTorch and SpikingJelly frameworks are used for training, the Adam optimizer is used, the total number of epochs is set to 70, the initial learning rate is set to 0.0002, and a 0.5-fold decay is performed at the 8th, 42nd, and 60th epochs.

[0068] 4. Describe the test part involved in step (5) in detail:

[0069] In the test phase, first, the neural network model loads the trained network weights, inputs the binocular events of the test set into the network, and finally obtains the depth map through the encoder, bottleneck module, decoder, and edge regularization module in step (2).

[0070] As Figure 5 shown, the present invention is trained and tested on the MVSEC dataset to obtain the depth map. Among them, the first row is the result without using the feature supervision module and the edge regularization module, and the second row is the result using the method of the present invention. It can be seen that the present invention has better color consistency and edge effects.

Claims

1. A binocular event depth estimation method based on a spiking neural network, characterized in that, it includes the following steps: Step 1, select a binocular event depth dataset, which is divided into a training set and a test set; the dataset includes binocular events, APS images, and depth maps; Step 2, construct a spiking neural network based on UNet, including an encoder, a bottleneck module, a decoder, a feature supervision module, and an edge regularization module; the encoder, the bottleneck module, the decoder, and the edge regularization module are connected in sequence, and the feature supervision module is connected to the encoder; the encoder uses a convolution with a stride of 2 for downsampling; the bottleneck module adopts a general residual structure; the decoder uses nearest neighbor 2D upsampling and a convolution with a stride of 1 to upsample the encoded spikes passing through the encoder and the bottleneck module to the size of the depth map; the feature supervision module includes a convolution with a stride of 1 and a ReLU operation to learn the floating-point features of the multi-scale encoded spikes output by the encoder; the edge regularization module upsamples the multi-stage decoded spikes of the decoder to predict a binary edge map; Step 3, input the binocular events into the spiking neural network, the binocular events pass through the encoder to output multi-scale encoded spikes, and then input the encoded spikes into the feature supervision module to obtain spike features; The encoded spikes pass through the bottleneck module and the decoder to obtain multi-stage decoded spikes, input the multi-stage decoded spikes into the edge regularization module to obtain multi-stage edge prediction spikes, then input the multi-stage decoded spikes into an infinite threshold non-spiking neuron and input the multi-stage edge prediction spikes into a spiking neuron respectively, and then add the potential values of the two neurons to obtain a depth map; Step 4, construct a loss function for the spiking neural network, including a depth map regression loss, a feature absolute value loss, and an edge binary cross-entropy loss, and train the spiking neural network according to the loss function; Step 5, input the binocular events of the test set into the trained spiking neural network, and pass through the encoder, the bottleneck module, the decoder, and the edge regularization module to obtain a depth map.

2. The binocular event depth estimation method based on a spiking neural network according to claim 1, characterized in that, in Step 1, the dataset further includes a binary edge map obtained by inputting the depth map and the APS image into an edge detection network for processing.

Citation Information

Patent Citations

  • A binocular depth estimation method based on depth neural network

    CN109377530A

  • Monocular depth estimation method based on deep learning

    CN110738697A

Cited By

  • Monocular self-supervision depth estimation method fusing multi-resolution features and global context

    CN120976282A