A Monocular Depth Estimation Method for Event Cameras Based on State Space Model
Through the combination of state space model and event prior knowledge, an event camera monocular depth estimation network was designed, which solved the problem of insufficient mining of event camera data characteristics, and achieved efficient depth estimation and robust complex scene perception.
Patent Information
- Application Number
- CN202510586473.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing monocular depth estimation method is insufficient to mine the data characteristics of event cameras, and it is difficult to generate high-quality depth maps in complex scenarios. The traditional method has high computational complexity and depends on complex models.
The state space model is used to introduce the event camera monocular depth estimation, and the network is designed through time-frequency encoder, spatial encoder and decoder. Combined with event prior knowledge, the gradient descent algorithm is used to optimize the model parameters, and the multi-frequency state space module and event distribution gate module are used for feature extraction.
It improves the depth estimation performance of event cameras in complex scenarios, reduces the computational complexity, and can estimate dense depth information of fine-grained structures during the day and at night, improving the robustness of applications such as robot navigation and autonomous driving.
Smart Images

Figure CN120125632B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of event cameras and artificial intelligence computer vision, and specifically to a method for monocular depth estimation of event cameras based on a state space model. Background Art
[0002] Monocular depth estimation (MDE) plays an important role in the field of computer vision. It improves the perception and understanding of the three-dimensional environment at a low cost and has been widely used in applications such as robot navigation, autonomous driving, and virtual reality. However, due to the limitations of low dynamic range and easy motion blur of traditional frame-based cameras, it is difficult to generate high-quality depth maps relying solely on RGB images in complex scenarios such as high-speed motion and low light.
[0003] In recent years, event camera technology based on biological perception has attracted wide attention. Different from traditional cameras that capture scenes at a fixed frame rate, event cameras record the changes in pixel intensity at a microsecond-level resolution, generating an asynchronous event stream containing timestamp, position, and polarity information. Event cameras have significant advantages in hardware, such as high dynamic range (HDR) and high temporal resolution, while avoiding the acquisition of redundant data. However, event data is sparse in space and dynamically rich in time, which poses new challenges to deep learning-based algorithms.
[0004] Existing network models mostly migrate CNN networks in the traditional image field to extract spatial features, and at the same time use recurrent neural networks to simply model temporal cues, resulting in poor event information extraction ability. Recently, some works use Transformer to solve this problem, but it brings greater computational complexity, which is contrary to the low-latency events. Generally speaking, previous methods have insufficient mining of event data characteristics. Specifically, it is not enough to only consider the spatio-temporal characteristics of events. These methods tend to rely on complex models to achieve feature learning. Therefore, how to design an efficient depth estimation network framework that can comprehensively mine event data characteristics and focus on events is an important challenge in current research. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a monocular depth estimation model guided by event priors. On the one hand, the computational complexity is reduced by introducing a state space model into the monocular depth estimation problem of event cameras. On the other hand, event prior knowledge is introduced into the state space architecture to further improve the depth estimation performance of event data. The present invention will help improve the robustness of the perception of complex three-dimensional real-world scenes in practical applications such as robot navigation, autonomous driving, and virtual reality.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] A monocular depth estimation method for event cameras based on a state space model, characterized in that the method comprises:
[0008] Step 1, prepare event data, including simulated event data and real-scene event data, and depth labels corresponding to the two event data respectively;
[0009] Step 2, preprocess the original event data respectively by spatio-temporal voxel grid and calculate the event distribution map, which are respectively represented as dense spatio-temporal voxels and calculate the event distribution map ;
[0010] Step 3, design a monocular depth estimation network based on the event camera under the state space model, including a time-frequency encoder, a spatial encoder and a decoder with a state space structure;
[0011] Step 4, divide the preprocessed event data into a training set, a validation set and a test set; and use the training set to train the model, continuously adjust the parameters of the model through the backpropagation algorithm to minimize the loss function; at the same time, during the training process, use the validation set to monitor the performance of the model;
[0012] Step 5, use the trained model to estimate the relative depth information of the test event data within a certain time interval to obtain a depth map.
[0013] Further, the format of the event data in step S1 is , where is the pixel position, represents the timestamp of the triggered event, then represents the polarity of the event. The depth label represents the relative depth value of each pixel position.
[0014] Further, the step of preprocessing the original event data into spatio-temporal voxels in step 2 is to convert the collected M events into a spatio-temporal voxel grid similar to tensor representation , where H, W represents the height and width of the event frame to fully retain the four-dimensional information of the event. Specifically, after discretizing the time dimension of the event stream , project the event stream within the time window onto C time windows according to the timestamps. The formula is as follows:
[0015] ;
[0016] where is the unit impulse function, is the bilinear sampling kernel function, It is a standardized event timestamp. The present invention samples event data within ΔT = 50 ms, and the number of time windows C = 5.
[0017] Furthermore, in step 2, the event distribution diagram The calculation formula is:
[0018] ;
[0019] where is an indicator function, records the triggering of M events in the current time window.
[0020] Furthermore, the depth estimation network in step 3 consists of three parts:
[0021] The first part is an event state space model designed for event data as a time-frequency encoder;
[0022] The second part is a vision-based selective state space model as a spatial encoder;
[0023] The third part is a decoder constructed by dense residual blocks.
[0024] Furthermore, the event state space model in the first part consists of two branches, namely the recurrent branch and the activation branch. The recurrent branch is composed of multiple multi-frequency state space modules and skip connections, and performs time-domain and frequency-domain encoding on the event spatio-temporal voxel tensor to obtain the update of , the activation branch calculates the event distribution and calculates the event distribution attention map through the event gating activation module and the gating weight , and the activation encoder outputs . The overall architecture is shown in the following formula:
[0025] .
[0026] Furthermore, the multi-frequency state space module in the recurrent branch is a state space module guided by wavelet transform, which extracts the time-frequency information of events in different frequency bands, including the following steps:
[0027] Use the high-frequency and low-frequency filters generated by the Haar wavelet transform to form 4 different filters, and use wavelet convolution to decompose the input event spatio-temporal voxel grid to obtain three high-frequency bands ( , , ) and one low-frequency band ( ), and the resolution in each spatial dimension is Half of it, and then the low-frequency components are recursively decomposed and cascaded wavelet convolutions are used to enhance the expression ability of static information. The formula is as follows:
[0028] ;
[0029] ;
[0030] Among them, represents the event of the i th channel, Conv represents convolution, T represents transpose, B represents the batch size. Integrate the frequency components corresponding to the three high-frequency bands and one low-frequency band calculated above to obtain the event frequency feature map .
[0031] Adopt a three-path multi-modal modeling strategy to map event data to three state representations: multi-frequency expansion features , time expansion features and compressed space features .
[0032] Among them, the three-path multi-modal modeling strategy includes:
[0033] S31-1. The first path is to expand in the time dimension to obtain , and a frequency domain attention module is introduced. This module uses a 2D convolution and a shared multi-layer perceptron ( MLP ) model to generate a channel frequency matrix . The formula is as follows:
[0034] ;
[0035] Among them flatten represents the current expansion along the time dimension, MLP represents the multi-layer perceptron with shared weights, represents the sigmoid activation function.
[0036] S31-2. The second path is spatial compression. The data of multiple time steps of the original event spatio-temporal voxels are spatially compressed, and attention processing is performed on the input time dimension. The time weights are calculated through parallel average pooling and max pooling operations, and then a time weight vector is generated through a shared multi-layer perceptron to capture the temporal correlation. The formula is as follows:
[0037] ;
[0038] Among them represents activation function, and are two multi-layer perceptrons sharing weights.
[0039] S31-3. The third path is spatial unfolding. The multi-channel event frequency-domain map after frequency-domain decomposition is further unfolded in space to obtain . The formula is as follows:
[0040] ;
[0041] where flatten represents unfolding in the spatial dimension.
[0042] The unfolded event feature map is input into a state space model (SSM) block, which is designed to consist of an S5 block and a feed-forward network module FFN of a standard Transformer with normalization, activation, and residual connection modules. The output of the SSM block is subjected to an inverse wavelet transform and then multiplied by the temporal attention weights and the frequency-domain channel matrix to obtain the final time-frequency coding output . The formula is as follows:
[0043] ;
[0044] ;
[0045] ;
[0046] where IWT represents the inverse wavelet transform, represents element-wise multiplication, represents the output feature map of S5, represents the hidden layer at the current moment, represents the feed-forward network output.
[0047] Stacking multiple multi-frequency state space modules gives a multi-frequency state space group, and finally the output of the multi-frequency state space group is obtained after passing through a feed-forward network with residuals .
[0048] Furthermore, the activation branch includes an event distribution gating module that generates an event distribution map and gating weights, including the following steps:
[0049] The module extracts multi-scale spatial features of the event distribution through depthwise separable convolutions and pointwise convolutions with different kernel sizes (3×3, 5×5, 7×7), smooths the event distribution, and integrates distribution information from multiple scales. The multi-scale distribution map passes through a module CRM composed of two convolutional layers plus function, and then applies the Sigmoid activation function ( ) to generate a pixel-level attention map As shown in the following formula:
[0050] ;
[0051] where and respectively represent the depthwise separable convolution and point convolution of the i th convolutional kernel size, and LN represents layer normalization.
[0052] Applying the attention map to the multi-scale event distribution map results in , and its global statistical characteristics are further calculated through average pooling and max pooling, and then passed through a module composed of two linear layers and LRM module, and the sigmoid activation function is applied to obtain the gating weight . As shown in the following formula:
[0053] ;
[0054] ;
[0055] Furthermore, the vision-based selective state space model adopts Vision Mamba (Vim). Specifically, the event data after time-frequency encoding is first converted into event blocks; then these event blocks are linearly projected into a higher-dimensional vector, and positional embeddings are added to these projected vectors. The first branch sends the sequence of event blocks into layers of bidirectional Vim encoders for event data spatial encoding, is the default setting of Vim, . In the second branch, the event blocks pass through a linear layer to adjust the encoder output through a gating activation function, and finally the final spatial encoding output is obtained through a residual connection.
[0056] Furthermore, the gradient descent algorithm is used for backpropagation training, and the loss function will include the traditional L 1 loss and the multi-scale gradient matching loss L grad , expressed as:
[0057] ;
[0058] ;
[0059] where is the logarithmic depth difference, is the predicted value, is the true value, N is the number of pixels, is the corresponding scale s of the depth difference, and represent the gradients calculated using the Sobel operator in the x and y directions respectively, , .
[0060] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0061] The present invention proposes a monocular depth estimation method for event cameras based on a state space model, introducing the state space model into the field of monocular depth estimation for event cameras, focusing on events, making full use of event prior knowledge, achieving efficient unification of event data in the frequency domain, time domain, and spatial domain under the guidance of event prior information, fully exploiting the non-uniform frequency characteristics, continuous time characteristics, and sparse spatial characteristics of event data. This model can estimate dense depth information with fine-grained structures during both day and night. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0063] Figure 1 is the flowchart of the embodiment of the present invention;
[0064] Figure 2 is the overall network architecture of the embodiment of the present invention;
[0065] Figure 3 is the multi-frequency state space model structure (MFSSB) of the embodiment of the present invention;
[0066] Figure 4 is the event gating unit structure (EDGU) of the embodiment of the present invention;
[0067] Figure 5 is the qualitative comparison result of the embodiment of the present invention on MVSEC;
[0068] Figure 6 is the qualitative comparison result of the embodiment of the present invention on DENSE;
[0069] Figure 7 is the ablation study and evaluation diagram of the embodiment of the present invention on MVSEC;
[0070] Figure 8 is the evaluation diagram of the embodiment of the present invention on DENSE;
[0071] Figure 9 This is the ablation experiment comparison chart of each module in the embodiments of the present invention. Specific implementation manners
[0072] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0073] Please refer to Figures 1 - 9 , the present invention provides a technical solution:
[0074] Embodiment 1: As Figure 1 shown, this embodiment provides a monocular depth estimation method for an event camera based on a state space model, including the following steps:
[0075] (1) Prepare a data set;
[0076] Preparing the data set includes simulated event data and real-scene event data, as well as their corresponding depth labels; in this example, publicly available data sets are used, namely the MVSEC data set in the real scene and the simulated data set DENSE. Among them, MVESC includes four real scenes (Outdoor day1, Outdoor night1, Outdoor night2, Outdoor night3), and the data format is hdf5 format. The DENSE data set contains eight scenes (Town 01, Town 02, Town 03, Town 04, Town06, Town 07, Town 10), and the data format is numpy format.
[0077] (2) Preprocess the original event data;
[0078] Represent it as a dense spatio-temporal voxel , and calculate the event distribution map . Specifically, the step of preprocessing the original event data into spatio-temporal voxels is to convert the M events collected into a spatio-temporal voxel grid similar to a tensor representation , where H, W represents the height and width of the event frame to fully retain the four-dimensional information of the event. Specifically, after discretizing the time dimension of the event stream , the event stream within the time window is projected onto C time windows according to the event timestamps. The formula is as follows:
[0079] ;
[0080] wherein is the unit impulse function, is the bilinear sampling kernel function, is the normalized event timestamp. The present invention samples the event data within ΔT = 50 ms, and the number of time windows C = 5;
[0081] Furthermore, the event distribution map is calculated by the formula:
[0082] ;
[0083] wherein is the indicator function, records the triggering of M events within the current time window .
[0084] (3) Model design;
[0085] The depth estimation network model is designed into three parts. The first part is the event state space model designed for event data as the time-frequency encoder, the second part is the vision-based selective state space model as the spatial encoder, and the third part is the decoder constructed by dense residual blocks.
[0086] Among them, the event state space model consists of two branches, namely the recurrent branch and the activation branch. The recurrent branch is composed of multiple multi-frequency state space modules and skip connections, and encodes the event spatio-temporal voxel tensor in the time domain and frequency domain to obtain the update of , and the activation branch calculates the event distribution and calculates the event distribution attention map through the event gating activation module and the gating weight , and the activation encoder outputs to obtain . The overall architecture is shown in the following formula:
[0087] ;
[0088] The multi-frequency state space module is a state space module guided by wavelet transform, which extracts the time-frequency information of events in different frequency bands, including the following steps: combining the high-frequency and low-frequency filters generated by Haar wavelet transform into 4 different filters, and using wavelet convolution to decompose the input event spatio-temporal voxel grid to obtain three high-frequency bands ( , , ) and one low-frequency band ( ), the resolution in each spatial dimension is half of that, and then the low-frequency components are recursively decomposed and cascaded wavelet convolutions are used to enhance the expression ability of static information. The formula is as follows:
[0089] ;
[0090] ;
[0091] Among them, represents the event of the i th channel, Conv represents convolution, T represents transpose, B represents the batch size.
[0092] Integrate the frequency components corresponding to the three high-frequency bands and one low-frequency band calculated above to obtain the event frequency feature map . Adopt a three-path multi-modal modeling strategy to map the event data to three state representations: multi-frequency expansion feature , time expansion feature and compressed spatial feature .
[0093] Specifically, the first path is to expand in the time dimension to obtain , and a frequency domain attention module is introduced. This module uses a 2D convolution and a shared multi-layer perceptron ( MLP ) model to generate the channel frequency matrix . The formula is as follows:
[0094] ;
[0095] Among them flatten represents the current expansion along the time dimension, MLP represents the multi-layer perceptron with shared weights, represents the sigmoid activation function.
[0096] The second path is spatial compression. The data of multiple time steps of the original event spatio-temporal voxels are compressed in space, and attention processing is performed on the input time dimension. The time weights are calculated through parallel average pooling and max pooling operations, and then the time weight vector is generated through a shared multi-layer perceptron to capture the temporal correlation. The formula is as follows:
[0097] ;
[0098] Among them represents activation function, and It is two multi-layer perceptrons sharing weights.
[0099] The third path is spatial unfolding. The multi-channel event frequency-domain map after frequency-domain decomposition is further unfolded in space to obtain . The formula is as follows:
[0100] ;
[0101] where flatten represents unfolding in the spatial dimension.
[0102] The unfolded event feature map is input into a state space model (SSM) block, which is designed to consist of an S5 block and a feed-forward network module FFN of a standard Transformer with normalization, activation, and residual connection modules. The output of the SSM block is multiplied by the temporal attention weight and the frequency-domain channel matrix after inverse wavelet transform to obtain the final time-frequency encoding output . The formula is as follows:
[0103] ;
[0104] ;
[0105] ;
[0106] where, IWT represents inverse wavelet transform, represents dot product, represents the output feature map of S5, represents the hidden layer at the current moment, represents the feed-forward network output.
[0107] Stacking multiple multi-frequency state space modules results in a multi-frequency state space group, and finally, the output of the multi-frequency state space group is obtained after a feed-forward network with residuals .
[0108] The activation branch contains an event distribution gating module that generates an event distribution map and gating weights, including the following steps: The module extracts multi-scale spatial features of the event distribution through depthwise separable convolutions and point convolutions with different kernel sizes (3×3, 5×5, 7×7), smooths the event distribution, and integrates distribution information from multiple scales. The multi-scale distribution map passes through a module CRM composed of two convolutional layers plus function, and then applies the Sigmoid activation function ( ) to generate a pixel-level attention map . As shown in the following formula:
[0109] ;
[0110] wherein and respectively represent the depthwise separable convolution and point convolution of the i -th convolutional kernel size, and LN represents layer normalization.
[0111] Applying the attention map to the multi-scale event distribution map results in , and its global statistical features are further calculated through average pooling and max pooling, and then passed through a module composed of two linear layers and LRM module, and then the sigmoid activation function is applied to obtain the gating weight . As shown in the following formula:
[0112] ;
[0113] ;
[0114] The vision-based selective state space model adopts Vision Mamba (Vim). Specifically, the event data after time-frequency encoding is first converted into event blocks; then these event blocks are linearly projected into a higher-dimensional vector, and positional embeddings are added to these projected vectors. The first branch sends the sequence of event blocks into a -layer bidirectional Vim encoder for spatial encoding of event data, being the default setting of Vim, . In the second branch, the event blocks pass through a linear layer to adjust the encoder output through a gating activation function, and finally the final spatial encoding output is obtained through a residual connection.
[0115] (4) Model training;
[0116] Backpropagation training is performed using the gradient descent algorithm, the AdamW optimizer is used, the training data is randomly cropped to 224×224 for data augmentation, and the learning rate is set to 0.0001 for 200 epochs. The loss function will include the traditional L 1 loss and the multi-scale gradient matching loss L grad , expressed as:
[0117] ;
[0118] ;
[0119] wherein is the logarithmic depth difference, is the predicted value, is the true value, N is the number of pixels, is the corresponding scale s of the depth difference, and represent calculating respectively using the Sobel operator x and y the gradients in the directions, , .
[0120] (5) Experimental tests:
[0121] To verify the effectiveness of the proposed method, we compared the method of the present invention with a variety of existing state-of-the-art event-based methods (E2Depth (Hidalgo-Carrio et al., 2020, Int. Conf. 3DV), E2Depth+ (Hidalgo-Carrio et al., 2020, Int. Conf. 3DV), EReFormer (Liu et al., 2024, IEEE TCSVT), DTL- (Wang et al., 2021, Int. Conf. CVPR)).
[0122] E2Depth+ was pre-trained on the first 1000 samples in the DENSE dataset and then re-trained on two datasets that share the same architecture as the method of the present invention.
[0123] Quantitative evaluation on MVSEC is as Figure 7 shown (the evaluation metrics are the same as those of the existing methods, including: AbsRel, RMSElog, SIlog, δ<1.25, δ<1.25 2 , δ<1.25 3 etc.), and qualitative comparison is as Figure 5 shown; quantitative evaluation on DENSE is as Figure 8 shown, and qualitative comparison is as Figure 6 shown. The results prove the effectiveness of the extraction of event prior knowledge and the applicability of state space modeling to event data modeling of the method of the present invention.
[0124] Figure 9For the ablation experiment comparison of each module, Experiment 1 was compared with the experimental baseline. Especially in unseen night scenes, the method of the present invention demonstrated significant advantages, which verified that the frequency differences between datasets were caused by factors such as vehicle speed and lighting conditions, and these differences needed to be considered in multiple frequencies to achieve better generalization ability. In addition, the method of the present invention achieved a 26% reduction in latency. In Experiment 2, the activation branch was replaced with a linear module for comparison, and the results showed that the method of the present invention had more superior performance. We believe that using a unified MLP gating architecture cannot effectively solve each specific problem, which highlights the advantage of the more customized design of the method of the present invention, making it more suitable as a gating mechanism for events. In addition, the prior of the event distribution in EDGU provides stronger guidance for the encoder, enabling it to learn more effectively.
[0125] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A monocular depth estimation method for event cameras based on the state space model, characterized in that: The method includes: Step 1: Prepare event data, including simulated event data and real-scene event data, as well as the depth labels corresponding to the two event data respectively; Step 2: Preprocess the original event data into a spatio-temporal voxel grid and a calculated event distribution map, respectively represented as a dense spatio-temporal voxel and a calculated event distribution map ; Step 3: Design a monocular depth estimation network based on an event camera under a state space model, including a time-frequency encoder, a spatial encoder, and a decoder of the state space structure; Step 4: Divide the preprocessed event data into a training set, a validation set, and a test set; and use the training set to train the model, continuously adjusting the parameters of the model through the backpropagation algorithm to minimize the loss function; meanwhile, during the training process, use the validation set to monitor the performance of the model; Step 5: Use the trained model to estimate the relative depth information of the test event data within a certain time interval to obtain a depth map; Among them, the depth estimation network in step 3 includes three parts; The first part is to use the event state space model designed for event data as the time-frequency encoder; The second part is to use the vision-based selective state space model as the spatial encoder; The third part is a decoder constructed by dense residual blocks; The event state space model in the first part includes a cyclic branch and an activation branch; The loop branch consists of multiple multi-frequency state space modules and skip connections, and obtains the updated voxel tensor by performing time-domain and frequency-domain encoding on the event spatio-temporal voxel tensor. update ; The activation branch calculates the event distribution and calculates the event distribution attention map through the event gating activation module and the gating weight , thereby activating the encoder output to obtain ; Based on this, the specific architecture of the event state space model is shown by the following formula: ; The activation branch contains an event distribution gating module that generates an event distribution map and gating weights, including the following steps: S32-1: The module extracts multi-scale spatial features of the event distribution through depthwise separable convolutions and point convolutions with different kernel sizes, smooths the event distribution, and integrates distribution information from multiple scales: According to the formula, the multi-scale distribution map After passing through the module CRM composed of two convolutional layers plus function, the Sigmoid activation function ( ) is applied to generate a pixel-level attention map : ; Among them and respectively represent depthwise separable convolution and point convolution with the size of the i-th convolution kernel, and LN represents layer normalization; S32-2. Apply the attention map to the multi-scale event distribution map according to the formula to obtain , further calculate its global statistical characteristics through average pooling and max pooling, and then apply the sigmoid activation function after passing through the LRM module composed of two linear layers and module to obtain the gating weight : ; 。 2. The monocular depth estimation method for an event camera based on a state space model according to claim 1, characterized in that, The event data format in the said step 1 is ; wherein is the pixel position, represents the timestamp of the triggering event, and represents the polarity of the event; The depth label represents the relative depth value of each pixel position.
3. The monocular depth estimation method for an event camera based on a state space model according to claim 1, wherein, The preprocessing of the spatio-temporal voxel grid in step 2 includes: Convert the M events collected into a spatio-temporal voxel grid in a tensor-like representation; Among them, H and W represent the height and width of the event frame to fully retain the four-dimensional information of the event; Specifically, after discretizing the time dimension of the event stream and projecting the event stream within the time window onto C time windows according to the timestamps, according to the formula: ; wherein represents the unit impulse function, represents the bilinear sampling kernel function, is the normalized event timestamp.
4. The monocular depth estimation method for an event camera based on a state space model according to claim 1, wherein, The preprocessing of calculating the event distribution map in step 2 includes: According to the formula, the event distribution diagram is obtained : ; Among them, represents an indicator function, indicating the triggering of M events within the current time window .
5. The monocular depth estimation method for an event camera based on a state space model according to claim 1, characterized in that, The multi-frequency state space module in the cyclic branch is a state space module guided by wavelet transform, which extracts the time-frequency information of events in different frequency bands, including the following steps: The high-frequency and low-frequency filters generated by the Haar wavelet transform are combined into four different filters, and the input event spatio-temporal voxel grid is decomposed using the wavelet convolution kernel F to obtain three high-frequency bands: , , , and a low-frequency band: ; the resolution in each spatial dimension is half of that, and then the low-frequency components are recursively decomposed and cascaded wavelet convolutions are used to enhance the expression ability of static information, according to the formula: ; ; Among them, represents the event of the i-th channel, Conv represents convolution, T represents transpose, and B represents the batch size; By integrating the frequency components corresponding to three high-frequency bands and one low-frequency band, an event frequency feature map is obtained ; Adopt a three-path multi-modal modeling strategy to map event data to three state representations: multi-frequency expansion features , time expansion features and compressed space features ; Among them, the three-path multi-modal modeling strategy includes: S31-1. The first path is obtained by unfolding in the time dimension , and a frequency domain attention module is introduced. According to the formula, the frequency domain attention module uses 2D convolution and a shared multi-layer perceptron model to generate a channel frequency matrix : ; Among them, "flatten" means to expand along the time dimension currently, and "MLP" means a multi-layer perceptron with shared weights. represents the sigmoid activation function; S31-2. The second path is to perform spatial compression on the data of multiple time steps of the original event spatio-temporal voxel: According to the formula, perform attention processing on the input time dimension, and calculate the time weights through parallel average pooling and max pooling operations , and then generate a time weight vector through a shared multi-layer perceptron , to capture the temporal correlation: ; Among them, denotes the activation function, and are two multi-layer perceptrons sharing weights; S31-3. The third path is to further expand the multi-channel event frequency-domain graph after frequency-domain decomposition in space to obtain , according to the formula: ; Among them, flatten means to expand in the spatial dimension; Input the unfolded event feature map into the state space model block, where the state space model block includes an S5 block and a feed-forward network module FFN of a standard Transformer with normalization, activation, and residual connection modules; According to the formula, the output of the state space model block is multiplied by the time attention weight and the frequency domain channel matrix after inverse wavelet transform to obtain the final time-frequency coding output : ; ; ; Among them, IWT represents the inverse wavelet transform, represents dot product, represents the output feature map of S5, represents the hidden layer at the current moment, represents the output of the feedforward network; Stacking multiple multi-frequency state space modules results in a multi-frequency state space group, and finally the output of the multi-frequency state space group is obtained after passing through a feed-forward network with residuals .
6. The monocular depth estimation method for an event camera based on a state space model according to claim 1, characterized in that The vision-based selective state space model in the second part is used as the spatial encoder, including: The event data after time-frequency encoding is first converted into event blocks; Linearly project the event block into a higher-dimensional vector and add positional embeddings to the projected vector ; The first branch sends the event block sequence into the two-way Vim encoder in the layer for event data space encoding; Among them, is the default setting of Vim, ; In the second branch, the event block passes through a linear layer to adjust the encoder output through a gating activation function, and finally obtains the final spatial encoding output through a residual connection.
7. The monocular depth estimation method for an event camera based on a state space model according to claim 1, characterized in that, In step 4, continuously adjusting the parameters of the model through the backpropagation algorithm includes: Backpropagation training using the gradient descent algorithm, the loss function will include the traditional L1 loss and the multi-scale gradient matching loss L grad , expressed as: ; ; Among them, is the logarithmic depth difference, is the predicted value, is the true value, N is the number of pixels, is the depth difference corresponding to scale s, and represent calculating the gradients in the x and y directions respectively using the Sobel operator, , .
Citation Information
Patent Citations
Unmanned underwater vehicle autonomous decision control method based on visual depth estimation
CN111340868A
Monocular depth estimation method and device for pulse camera
CN114998402A