Event camera monocular depth estimation method based on state space model

By introducing state space models and event prior knowledge in monocular depth estimation, a depth estimation network that can efficiently mine event data characteristics is designed, which solves the problem of insufficient depth map quality in complex scenarios by existing methods, and achieves higher quality and robust depth estimation effects.

CN120125632AActive Publication Date: 2025-06-10NANJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510586473.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-10
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing monocular depth estimation method is difficult to generate high-quality depth maps in complex scenarios such as high-speed motion and low-light, and the deep learning-based algorithms do not mine the feature of event data.

Method used

Using the event camera monocular depth estimation method based on the state space model, a network framework that can fully mine event data characteristics and efficient depth estimation is designed by introducing event prior knowledge and time-frequency encoder, spatial encoder, decoder and other components.

Benefits of technology

This method significantly improves the quality and robustness of the depth map in complex scenarios, can estimate and generate dense depth information of fine-grained structures during daytime and nighttime, reduces computational complexity, and enables low-latency event camera applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125632A_ABST
    Figure CN120125632A_ABST
Patent Text Reader

Abstract

The invention relates to the field of event camera and artificial intelligence computer vision, in particular to an event camera monocular depth estimation method based on a state space model, which comprises the following steps: preparing a data set containing event data and a depth map corresponding to the event data; preprocessing the event data into a space-time voxel form; calculating an event distribution map as priori knowledge; constructing an event state space module which comprises a multi-frequency state space module and an event distribution gating unit; constructing a vision-based selective state space model for encoding spatial features of the event data; and constructing a dense residual block as a decoder to generate a required dense depth map. The invention aims to comprehensively integrate time, frequency and space information in a state space model under the guidance of event prior, enhance the modeling capability of the model for asynchronous sparse event data, and improve the accuracy of depth estimation. The method is helpful for improving the robustness of robot navigation, automatic driving, virtual reality and other practical applications for sensing a complex three-dimensional reality scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of event cameras and artificial intelligence computer vision, and specifically to a method for monocular depth estimation of event cameras based on a state space model. Background Art

[0002] Monocular depth estimation (MDE) plays an important role in the field of computer vision. It improves the perception and understanding of the three-dimensional environment at a relatively low cost and has been widely used in applications such as robot navigation, autonomous driving, and virtual reality. However, due to the limitations of low dynamic range and easy motion blur of traditional frame-based cameras, it is difficult to generate high-quality depth maps relying solely on RGB images in complex scenarios such as high-speed motion and low light.

[0003] In recent years, event camera technology based on biological perception has attracted wide attention. Different from traditional cameras that capture scenes at a fixed frame rate, event cameras record changes in pixel intensity at microsecond-level resolution, generating an asynchronous event stream containing timestamp, position, and polarity information. Event cameras have significant advantages in hardware, such as high dynamic range (HDR) and high temporal resolution, while avoiding the acquisition of redundant data. However, event data is sparse in space and dynamically rich in time, which poses new challenges to deep learning-based algorithms.

[0004] Existing network models mostly migrate CNN networks in the traditional image field to extract spatial features and simply model temporal cues with recurrent neural networks, resulting in poor event information extraction capabilities. Recently, some works use Transformer to solve this problem, but it brings greater computational complexity, which is contrary to the low-latency nature of events. Generally speaking, previous methods have insufficiently mined the characteristics of event data. Specifically, it is not enough to only consider the spatio-temporal characteristics of events. These methods tend to rely on complex models to achieve feature learning. Therefore, how to design an efficient depth estimation network framework that can comprehensively mine the characteristics of event data and focus on events is an important challenge in current research. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a monocular depth estimation model guided by event priors. On the one hand, the computational complexity is reduced by introducing a state space model into the problem of monocular depth estimation of event cameras. On the other hand, event prior knowledge is introduced into the state space architecture to further improve the depth estimation performance of event data. The present invention will help to improve the robustness of the perception of complex three-dimensional real-world scenarios in practical applications such as robot navigation, autonomous driving, and virtual reality.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: A monocular depth estimation method for event cameras based on a state space model, characterized in that the method comprises: Step 1, prepare event data, including simulated event data and real-scene event data, and depth labels corresponding to the two event data respectively; Step 2, preprocess the original event data respectively by spatio-temporal voxel grids and calculate event distribution maps, which are respectively represented as dense spatio-temporal voxels and calculate event distribution maps ; Step 3, design a monocular depth estimation network based on an event camera under a state space model, including a time-frequency encoder, a spatial encoder and a decoder with a state space structure; Step 4, divide the preprocessed event data into a training set, a validation set and a test set; and use the training set to train the model, continuously adjust the parameters of the model through the backpropagation algorithm to minimize the loss function; meanwhile, during the training process, use the validation set to monitor the performance of the model; Step 5, use the trained model to estimate the relative depth information of the test event data within a certain time interval to obtain a depth map.

[0007] Further, the format of the event data in step S1 is , where is the pixel position, represents the timestamp of the triggered event, represents the polarity of the event. The depth label represents the relative depth value of each pixel position.

[0008] Further, the step of preprocessing the original event data into spatio-temporal voxels in step 2 is to convert the collected M events into a spatio-temporal voxel grid similar to a tensor representation , where H, W represents the height and width of the event frame to fully retain the four-dimensional information of the event. Specifically, after discretizing the time dimension of the event stream , the event stream within the time window is projected onto C time windows according to the timestamps. The formula is as follows: ; where is the unit impulse function, is the bilinear sampling kernel function, is the normalized event timestamp. The present invention samples event data within ΔT = 50 ms, and the number of time windows C = 5.

[0009] Further, the event distribution map in step 2 The calculation formula is as follows: ; where is an indicator function, recording the triggering of M events in the current time window.

[0010] Furthermore, the depth estimation network in step 3 consists of three parts: The first part is an event state space model designed for event data as a time-frequency encoder; The second part is a vision-based selective state space model as a spatial encoder; The third part is a decoder constructed by dense residual blocks.

[0011] Furthermore, the event state space model in the first part consists of two branches, namely the recurrent branch and the activation branch. The recurrent branch consists of multiple multi-frequency state space modules and skip connections, encoding the event spatio-temporal voxel tensor in the time domain and frequency domain to obtain the update of , and the activation branch calculates the event distribution and calculates the event distribution attention map through the event gating activation module and the gating weight , and the activation encoder outputs to obtain . The overall architecture is shown in the following formula: .

[0012] Furthermore, the multi-frequency state space module in the recurrent branch is a state space module guided by wavelet transform, extracting the time-frequency information of events in different frequency bands, including the following steps: Using the high-frequency and low-frequency filters generated by the Haar wavelet transform to form 4 different filters, and using wavelet convolution to decompose the input event spatio-temporal voxel grid to obtain three high-frequency bands ( , , ) and one low-frequency band ( ), and the resolution in each spatial dimension is half of ; ; where represents the event of the i th channel, Conv represents convolution, T represents transpose, BDenotes the batch size. Integrate the frequency components corresponding to the three high-frequency bands and one low-frequency band calculated above to obtain the event frequency feature map .

[0013] Adopt a three-path multi-modal modeling strategy to map event data to three state representations: multi-frequency expansion features , time expansion features and compressed space features .

[0014] Among them, the three-path multi-modal modeling strategy includes: S31-1. The first path is obtained by expanding in the time dimension , and a frequency domain attention module is introduced. This module uses a 2D convolution and a shared multi-layer perceptron ( MLP ) model to generate a channel frequency matrix . The formula is as follows: ; Where flatten represents the current expansion along the time dimension, MLP represents the multi-layer perceptron with shared weights, represents the sigmoid activation function.

[0015] S31-2. The second path is spatial compression. Perform spatial compression on the data of multiple time steps of the original event spatio-temporal voxels, perform attention processing on the input time dimension, calculate the time weights through parallel average pooling and max pooling operations , and then generate a time weight vector through a shared multi-layer perceptron to capture the temporal correlation. The formula is as follows: ; Where represents activation function, and are two multi-layer perceptrons with shared weights.

[0016] S31-3. The third path is spatial expansion. Further expand the multi-channel event frequency domain map after frequency domain decomposition in space to obtain . The formula is as follows: ; Where flatten represents expansion in the spatial dimension.

[0017] The expanded event feature map is input into a state space model (SSM) block, which is designed to consist of an S5 block and a feed-forward network module FFN of a standard Transformer with normalization, activation, and residual connection modules. After the output of the SSM block undergoes an inverse wavelet transform, it is multiplied by the temporal attention weights and the frequency-domain channel matrix to obtain the final time-frequency encoding output . The formula is as follows: ; ; ; Among them, IWT represents the inverse wavelet transform, represents the dot product, represents the output feature map of S5, represents the hidden layer at the current moment, represents the feed-forward network output.

[0018] Stacking multiple multi-frequency state space modules results in a multi-frequency state space group, and finally, the output of the multi-frequency state space group is obtained after passing through a feed-forward network with residuals .

[0019] Furthermore, the activation branch includes an event distribution gating module that generates an event distribution map and gating weights, including the following steps: The module extracts multi-scale spatial features of the event distribution through depthwise separable convolutions and pointwise convolutions with different kernel sizes (3×3, 5×5, 7×7), smooths the event distribution, and integrates distribution information from multiple scales. The multi-scale distribution map passes through a module CRM composed of two convolutional layers plus function, and then applies the Sigmoid activation function ( ) to generate a pixel-level attention map . As shown in the following formula: ; where and respectively represent the depthwise separable convolution and pointwise convolution with the i th convolutional kernel size, and LN represents layer normalization.

[0020] Applying the attention map to the multi-scale event distribution map results in . After calculating its global statistical characteristics through average pooling and max pooling, and then passing through a module composed of two linear layers and LRM module, the sigmoid activation function is applied to obtain the gating weight . As shown in the following formula: ; ; Furthermore, the vision-based selective state space model adopts Vision Mamba (Vim). Specifically, the event data after time-frequency encoding is first converted into event blocks; then these event blocks are linearly projected onto a higher-dimensional vector, and positional embeddings are added to these projected vectors . The first branch feeds the sequence of event blocks into a bi-directional Vim encoder layer for event data spatial encoding, which is the default setting of Vim, . In the second branch, the event blocks pass through a linear layer to adjust the encoder output via a gated activation function, and finally the final spatial encoding output is obtained through a residual connection.

[0021] Furthermore, the gradient descent algorithm is used for backpropagation training, and the loss function will include traditional L 1 loss and multi-scale gradient matching loss L grad , expressed as: ; ; where is the logarithmic depth difference, is the predicted value, is the ground truth value, N is the number of pixels, is the depth difference corresponding to the scale s , and represent the gradients calculated using the Sobel operator in the x and y directions respectively, , .

[0022] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention proposes a monocular depth estimation method for event cameras based on a state space model, introduces the state space model into the field of monocular depth estimation for event cameras, focuses on events, makes full use of event prior knowledge, realizes the efficient unification of event data in the frequency domain, time domain, and spatial domain under the guidance of event prior information, fully excavates the non-uniform frequency characteristics, continuous time characteristics, and sparse spatial characteristics of event data, and this model can estimate dense depth information with fine-grained structures during both day and night. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings: Figure 1 is a flowchart of an embodiment of the present invention; Figure 2 is the overall network architecture of an embodiment of the present invention; Figure 3 is the multi-frequency state space model structure (MFSSB) of an embodiment of the present invention; Figure 4 is the event gating unit structure (EDGU) of an embodiment of the present invention; Figure 5 is the qualitative comparison result of an embodiment of the present invention on MVSEC; Figure 6 is the qualitative comparison result of an embodiment of the present invention on DENSE; Figure 7 is the ablation study and evaluation diagram of an embodiment of the present invention on MVSEC; Figure 8 is the evaluation diagram of an embodiment of the present invention on DENSE; Figure 9 is the ablation experiment comparison diagram of each module of an embodiment of the present invention. Specific embodiments

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] Please refer to Figures 1 - 9 , the present invention provides a technical solution: Embodiment 1: As Figure 1 shown, this embodiment provides a method for monocular depth estimation of an event camera based on a state space model, including the following steps: (1) Prepare a data set; The prepared dataset includes simulated event data, real-scenario event data, and their corresponding depth labels. In this example, publicly available datasets are used, namely the MVSEC dataset in real scenarios and the DENSE simulation dataset. Among them, MVESC includes four real scenarios (Outdoor day1, Outdoor night1, Outdoor night2, Outdoor night3), and the data format is hdf5. The DENSE dataset contains eight scenarios (Town 01, Town 02, Town 03, Town 04, Town06, Town 07, Town 10), and the data format is numpy.

[0026] (2) Preprocess the original event data; Represent it as a dense spatio-temporal voxel , and calculate the event distribution map . Specifically, the step of preprocessing the original event data into spatio-temporal voxels is to convert the M collected events into a spatio-temporal voxel grid similar to a tensor representation , where H, W represents the height and width of the event frame to fully retain the four-dimensional information of the event. Specifically, after discretizing the time dimension of the event stream , the event stream within the time window is projected onto C time windows according to the event timestamps. The formula is as follows: ; where is the unit impulse function, is the bilinear sampling kernel function, is the normalized event timestamp. In the present invention, event data within ΔT = 50ms is sampled, and the number of time windows C = 5; Further, the calculation formula for the event distribution map is: ; where is the indicator function, records the triggering of M events within the current time window.

[0027] (3) Model design; The depth estimation network model is designed in three parts. The first part is an event state space model designed for event data as a time-frequency encoder. The second part is a vision-based selective state space model as a spatial encoder. The third part is a decoder constructed by dense residual blocks.

[0028] Among them, the event state space model consists of two branches, namely the recurrent branch and the activation branch. The recurrent branch is composed of multiple multi-frequency state space modules and skip connections, and encodes the event spatio-temporal voxel tensor in the time domain and frequency domain to obtain updates , and the activation branch calculates the event distribution and calculates the event distribution attention map through the event gated activation module and the gated weight , and the activation encoder outputs . The overall architecture is shown in the following formula: ; The multi-frequency state space module is a state space module guided by wavelet transform, which extracts the time-frequency information of events in different frequency bands, including the following steps: combining the high-frequency and low-frequency filters generated by Haar wavelet transform into 4 different filters, and using wavelet convolution to decompose the input event spatio-temporal voxel grid to obtain three high-frequency bands ( , , ) and one low-frequency band ( ), and the resolution in each spatial dimension is half of that of ; ; Among them, represents the event of the i th channel, Conv represents convolution, T represents transpose, B represents the batch size.

[0029] Integrate the frequency components corresponding to the three high-frequency bands and one low-frequency band calculated above to obtain the event frequency feature map . Adopt a three-path multi-modal modeling strategy to map event data to three state representations: multi-frequency expansion feature , time expansion feature and compressed space feature .

[0030] Specifically, the first path is to expand in the time dimension to obtain , and a frequency-domain attention module is introduced. This module uses 2D convolution and a shared multi-layer perceptron ( MLP ) model to generate a channel frequency matrix . The formula is as follows: ; where flatten represents the current unfolding along the time dimension, MLP represents the multi-layer perceptron with shared weights, represents the sigmoid activation function.

[0031] The second path is spatial compression. The data of multiple time steps of the original event spatio-temporal voxels is spatially compressed, and attention processing is performed on the input time dimension. The temporal weights are calculated through parallel average pooling and max pooling operations , and then a temporal weight vector is generated through a shared multi-layer perceptron to capture the temporal correlation. The formula is as follows: ; where represents the activation function, and are two multi-layer perceptrons with shared weights.

[0032] The third path is spatial unfolding. The multi-channel event frequency-domain map after frequency-domain decomposition is further unfolded in space to obtain . The formula is as follows: ; where flatten represents unfolding in the spatial dimension.

[0033] The unfolded event feature map is input into a state space model (SSM) block, which is designed to consist of an S5 block and a feed-forward network module FFN of a standard Transformer with normalization, activation, and residual connection modules. After the output of the SSM block undergoes an inverse wavelet transform, it is multiplied by the temporal attention weights and the frequency-domain channel matrix to obtain the final time-frequency coding output . The formula is as follows: ; ; ; where IWT represents the inverse wavelet transform, represents the dot product, represents the output feature map of S5, represents the hidden layer at the current moment, Represents the output of the feedforward network.

[0034] Stacking multiple multi-frequency state space modules results in a multi-frequency state space group, and finally the output of the multi-frequency state space group is obtained after passing through a feedforward network with residuals. 。

[0035] The activation branch contains an event distribution gating module that generates an event distribution map and gating weights, including the following steps: The module extracts multi-scale spatial features of the event distribution through depthwise separable convolutions and pointwise convolutions with different kernel sizes (3×3, 5×5, 7×7), smooths the event distribution, and integrates distribution information from multiple scales. The multi-scale distribution map After passing through a module CRM composed of two convolutional layers plus function, the Sigmoid activation function ( ) is applied to generate a pixel-level attention map . As shown in the following formula: ; where and represent the depthwise separable convolution and pointwise convolution of the i th convolutional kernel size respectively, and LN represents layer normalization.

[0036] Applying the attention map to the multi-scale event distribution map results in . After average pooling and max pooling, its global statistical characteristics are further calculated, and then after passing through a module composed of two linear layers and LRM module, the sigmoid activation function is applied to obtain the gating weights . As shown in the following formula: ; ; The vision-based selective state space model uses Vision Mamba (Vim). Specifically, the event data after time-frequency encoding is first converted into event blocks; then these event blocks are linearly projected into a higher-dimensional vector, and positional embeddings are added to these projected vectors . The first branch feeds the sequence of event blocks into a layer bidirectional Vim encoder for spatial encoding of event data, which is the default setting of Vim, . In the second branch, the event blocks pass through a linear layer to adjust the encoder output through a gating activation function, and finally the final spatial encoding output is obtained through a residual connection.

[0037] (4) Model training; Backpropagation training is performed using the gradient descent algorithm. The AdamW optimizer is used, and the training data is randomly cropped to 224×224 for data augmentation. The learning rate is set to 0.0001 and trained for 200 epochs. The loss function will include the traditional L 1 loss and the multi-scale gradient matching loss L grad , expressed as: ; ; where is the logarithmic depth difference, is the predicted value, is the ground truth value, N is the number of pixels, is the depth difference corresponding to the scale s , and represent the gradients calculated using the Sobel operator in the x and y directions respectively, , .

[0038] (5) Experimental testing: To verify the effectiveness of the proposed method, we compare the method of the present invention with a variety of existing state-of-the-art event-based methods (E2Depth (Hidalgo-Carrio et al., 2020, Int. Conf. 3DV), E2Depth+ (Hidalgo-Carrio et al., 2020, Int. Conf. 3DV), EReFormer (Liu et al., 2024, IEEE TCSVT), DTL- (Wang et al., 2021, Int. Conf. CVPR)).

[0039] E2Depth+ was pre-trained on the first 1000 samples in the DENSE dataset and then re-trained on two datasets that share the same architecture as the method of the present invention.

[0040] Quantitative evaluation on MVSEC is as shown in Figure 7 (the evaluation metrics are the same as those of the existing methods, including: AbsRel, RMSElog, SIlog, δ<1.25, δ<1.25 2 , δ<1.25 3 etc.), and qualitative comparison is as shown in Figure 5 ; Quantitative evaluation on DENSE is asFigure 8 As shown, the qualitative comparison is as follows Figure 6 As shown, the results prove the effectiveness of the prior knowledge extraction of the method of the present invention and the applicability of the state space modeling to the event data modeling.

[0041] Figure 9 For the ablation experiment comparison of each module, Experiment 1 is compared with the experimental baseline. Especially in the unseen night scenes, the method of the present invention shows significant advantages, which verifies that the frequency differences between datasets are caused by factors such as vehicle speed and lighting conditions, and these differences need to be considered in multiple frequencies to achieve better generalization ability. In addition, the method of the present invention achieves a 26% reduction in latency. In Experiment 2, the activation branch is replaced with a linear module for comparison, and the results show that the method of the present invention has superior performance. We believe that using a unified MLP gating architecture cannot effectively solve each specific problem, which highlights the advantage of the more customized design of the method of the present invention, making it more suitable as a gating mechanism for events. In addition, the prior of the event distribution in EDGU provides stronger guidance for the encoder, enabling it to learn more effectively.

[0042] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for monocular depth estimation of an event camera based on a state-space model, characterized in that: Methods include: Step 1: Prepare event data, including simulated event data and real scene event data, as well as the deep labels corresponding to the two event data respectively; Step 2: Preprocess the original event data into space-time voxel grids and calculate event distribution maps, respectively, and represent them as dense space-time voxels and calculate the event distribution graph ; Step 3: Design a monocular depth estimation network based on an event camera under a state-space model, including a time-frequency encoder, a spatial encoder, and a decoder of a state-space structure; Step 4: Divide the preprocessed event data into a training set, a validation set, and a test set; use the training set to train the model, and continuously adjust the model parameters through the back propagation algorithm to minimize the loss function; at the same time, during the training process, use the validation set to monitor the performance of the model; Step 5: Use the trained model to estimate the relative depth information of the test event data within a certain time interval to obtain a depth map.

2. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 1, characterized in that: The event data format in step 1 is: ; in is the pixel position, Indicates the timestamp of the triggering event. It indicates the polarity of the event; The depth label indicates the relative depth value of each pixel location.

3. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 1, characterized in that: The step 2 is to pre-process the spatiotemporal voxel grid, including: The collected M Events Convert to tensor-like representation The space-time voxel grid of in, H.W. Represents the height and width of the event frame to fully preserve the four-dimensional information of the event; Specifically, the event stream After the time dimension is discretized, the time window is divided into The event stream in is projected to C In a time window, according to the formula: ; in represents the unit pulse function, represents the bilinear sampling kernel function, is the normalized event timestamp.

4. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 1, characterized in that: The preprocessing of calculating the event distribution graph in step 2 includes: According to the formula, we get the event distribution diagram : ; in, represents the indicator function, Indicates that the current time window is recorded M Events The trigger.

5. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 1, characterized in that: The depth estimation network in step 3 includes: three parts; The first part is to use the event state space model designed for event data as a time-frequency encoder; The second part is to use the vision-based selective state space model as a spatial encoder; The third part is the decoder built with dense residual blocks.

6. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 5, characterized in that: The event state space model in the first part includes: a loop branch and an activation branch; The loop branch consists of multiple multi-frequency state-space modules and jump connections, which encodes the event spatiotemporal voxel tensor in time domain and frequency domain to obtain a voxel tensor Update ; The activation branch calculates the event distribution and calculates the event distribution attention map through the event gated activation module and gate weights , thereby activating the encoder output to obtain ; Based on this, the specific architecture of the event state space model is shown in the following formula: 。 7. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 6, characterized in that: The multi-frequency state space module in the loop branch is a state space module guided by wavelet transform, which extracts time-frequency information of events in different frequency bands, including the following steps: The high-frequency and low-frequency filters generated by Haar wavelet transform are combined into 4 different filters, and the wavelet convolution kernel is used F The input event space-time voxel grid is decomposed to obtain three high-frequency bands: , , , and a low-frequency band: ; The resolution in each spatial dimension is Then, the low-frequency components are recursively decomposed and cascaded with wavelet convolution to improve the expression ability of static information. According to the formula: ; ; in, Indicates i events for each channel, Conv represents convolution, T represents transpose, B Indicates the batch size; By integrating the frequency components corresponding to three high-frequency bands and one low-frequency band, we can obtain the event frequency characteristic map ; A three-path multi-modal modeling strategy is used to map event data into three state representations: multi-frequency unfolding features , time expansion characteristics and compressed space features ; Among them, the three-path multi-morphological modeling strategy includes: S31-1. The first path is to expand using the time dimension , and introduces a frequency domain attention module. According to the formula, the frequency domain attention module uses 2D convolution and shared multi-layer perceptron model to generate a channel frequency matrix : ; in, flatten Indicates that the current expansion is along the time dimension. MLP represents a multilayer perceptron with shared weights, Represents the sigmoid activation function; S31-2, the second path is to compress the multiple time step data of the original event space-time voxels in space: according to the formula, the input time dimension is processed with attention, and the time weight is calculated by parallel average pooling and maximum pooling operations , and then generate the time weight vector through a shared multilayer perceptron , to capture temporal correlations: ; in, express Activation function, and Two multi-layer perceptrons with shared weights; S31-3. The third path is to further expand the multi-channel event frequency domain diagram after frequency domain decomposition in space to obtain , according to the formula: ; in, flatten It means expansion in the spatial dimension; Inputting the expanded event feature graph into a state-space model block, wherein the state-space model block includes an S5 block and a feed-forward network module FFN of a standard Transformer with normalization, activation and residual connection modules; According to the formula, the output of the state space model block is inversely wavelet transformed and multiplied with the time attention weight and the frequency domain channel matrix to obtain the final time-frequency coding output : ; ; ; Where IWT stands for inverse wavelet transform, represents dot product, represents the output feature map of S5, represents the hidden layer at the current moment, represents the output of the feedforward network; Multiple multi-frequency state space modules are stacked to obtain a multi-frequency state space group, and finally the output of the multi-frequency state space group is obtained after passing through a feed-forward network with residuals. .

8. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 6, characterized in that: The activation branch includes an event distribution gating module, which generates an event distribution map and a gating weight, including the following steps: S32-1, the module extracts multi-scale spatial features of event distribution through deep separation convolution and point convolution with different kernel sizes, smoothes event distribution and integrates distribution information from multiple scales: According to the formula, the multi-scale distribution map After two convolutional layers plus After the module CRM composed of functions, the Sigmoid activation function is applied ( ) Generate pixel-level attention map : ; in and Respectively represent i Depthwise separable convolution and pointwise convolution with kernel size of , LN represents layer normalization; S32-2. According to the formula, the attention map is applied to the multi-scale event distribution map After getting , further calculate its global statistical characteristics through average pooling and maximum pooling, and then pass through two linear layers and Modules LRM The sigmoid activation function is applied after the module to obtain the gate weight : ; 。 9. The method for monocular depth estimation of an event camera based on a state space model as claimed in claim 5, characterized in that: The vision-based selective state space model in the second part is used as a spatial encoder, including: Event data after time-frequency encoding First converted into The event block; Linearly project the event block to a higher dimensional vector and add position embedding to the projected vector ; The first branch feeds the event block sequence into Layer bidirectional Vim encoder for event data space encoding; in, is the default setting of Vim. ; In the second branch, the event block passes through a linear layer and adjusts the encoder output through a gated activation function, and finally obtains the final spatial encoding output through a residual connection.

10. The method for monocular depth estimation of an event camera based on a state space model according to claim 1, characterized in that: In step 4, the parameters of the model are continuously adjusted through the back propagation algorithm, including: Using the gradient descent algorithm for back propagation training, the loss function will include the traditional L 1 loss and multi-scale gradient matching loss L grad , expressed as: ; ; in, is the logarithmic depth difference, is the predicted value, is the true value, N is the number of pixels, The corresponding scale s The depth difference, and Represents the use of Sobel operator to calculate x and y The gradient in direction, , .

Citation Information

Patent Citations

  • Unmanned underwater vehicle autonomous decision control method based on visual depth estimation

    CN111340868A

  • Depth estimation method based on laser radar and event camera fusion

    CN114359744A

  • Monocular depth estimation method and device for pulse camera

    CN114998402A