Event camera target detection method based on sparse Mama
Through the event camera object detection method based on sparse Mamba, the problems of image quality damage and redundant calculation in the prior art in complex scenarios are solved, and efficient and accurate object detection effect is achieved.
Patent Information
- Application Number
- CN202510093649.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-13
AI Technical Summary
The existing object detection method based on event cameras is damaged in complex scenarios such as high-speed motion, low light and overexposure, resulting in difficulty in extracting key features, and ignores the sparsity and low signal-to-noise ratio of event data, resulting in redundant calculations and suboptimal performance.
Using the event camera object detection method based on sparse Mamba, by converting the event stream into voxel tensors and dividing it into fragments, a time continuity score graph is generated, average pooling and Gaussian smoothing are performed, and spatiotemporal continuity score graph is obtained, and a sparse graph is obtained by binarization. Then, feature extraction is performed through the multi-layer sparse space Mamba module and the space-channel hybrid Mamba module, multi-scale feature fusion is performed by combining the feature pyramid network, and finally the YOLOX detection head is input for object detection.
This method adaptively discards event-free and noise markers through space-time continuity evaluation, improves computing efficiency and accuracy, maintains stable performance in complex scenarios, and achieves the best balance between accuracy and efficiency.
Smart Images

Figure CN120147596A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of object detection, and specifically, to an object detection method for event cameras based on sparse Mamba. Background Art
[0002] Robust object detection is crucial for intelligent systems such as autonomous driving and robotics. However, traditional frame-based cameras have inherent limitations in frame rate and dynamic range, which can lead to impaired image quality in complex scenarios such as high-speed motion, low light, and overexposure, seriously affecting the effective extraction of key features. In contrast, event cameras, with their ability to detect pixel-level light intensity changes asynchronously, exhibit high temporal resolution and high dynamic range characteristics, and can maintain stable and robust performance even in extremely challenging environments.
[0003] To fully leverage the advantages of event cameras in object detection tasks, researchers have developed various methods based on advanced neural network architectures, covering methods based on spiking neural networks (SNNs), graph neural networks (GNNs), convolutional neural networks (CNNs), and Transformers. In theory, methods based on SNNs and GNNs can achieve low-latency inference, but this usually requires specific hardware support and performs poorly in practical application scenarios. Recent research progress has shown that methods based on Transformers, due to their larger receptive fields, significantly outperform CNN methods based on local receptive fields in terms of performance. However, these methods usually convert the event stream into discrete tokens and uniformly process the event-free and noisy regions, ignoring the impact of the spatial sparsity and low signal-to-noise ratio of event data, resulting in a large amount of redundant computation and suboptimal performance. To improve the computational efficiency on sparse event data, some researchers have proposed a sparsification strategy based on window attention to discard unimportant regions, but at the expense of global modeling ability. Summary of the Invention
[0004] To overcome at least one deficiency in the prior art, this application provides an object detection method for event cameras based on sparse Mamba.
[0005] In a first aspect, there is provided an object detection method for event cameras based on sparse Mamba, including:
[0006] Converting the event stream acquired by the event camera into a voxel tensor, and dividing the voxel tensor into multiple segments;
[0007] Based on the event stream, accumulating the timestamps of all events at each pixel point to generate a temporal continuity score map; performing an average pooling operation on the temporal continuity score map, and smoothing and denoising based on a Gaussian function to obtain a spatio-temporal continuity score map; binarizing the spatio-temporal continuity score map to obtain a sparsification map;
[0008] Input multiple segments and a sparsified graph into the first sparse spatial Mamba module for feature extraction to obtain the first feature;
[0009] Perform max pooling on the sparsified graph to obtain the first pooled feature; perform downsampling on the first feature to obtain the first downsampled feature; input the first pooled feature and the first downsampled feature into the second sparse spatial Mamba module for feature extraction to obtain the second feature;
[0010] Perform max pooling on the first pooled feature to obtain the second pooled feature; perform downsampling on the second feature to obtain the second downsampled feature; input the second pooled feature and the second downsampled feature into the first spatial-channel hybrid Mamba module for feature extraction to obtain the third feature;
[0011] Perform max pooling on the second pooled feature to obtain the third pooled feature; perform downsampling on the third feature to obtain the third downsampled feature; input the third pooled feature and the third downsampled feature into the second spatial-channel hybrid Mamba module for feature extraction to obtain the fourth feature;
[0012] Input the fourth feature, the third feature, and the second feature into the feature pyramid network for multi-scale feature fusion to obtain the fused feature;
[0013] Input the fused feature into the YOLOX detection head to obtain the object detection result.
[0014] In one embodiment, the first sparse spatial Mamba module includes 2 sparsified spatial Mamba layers with the same structure and connected in sequence;
[0015] Each sparsified spatial Mamba layer includes a sparse SS2D module, a sparse MLP module, and a ConvLSTM;
[0016] Input multiple segments and a sparsified graph into the sparse SS2D module for two-dimensional selective scanning, add the output of the sparse SS2D module to the multiple segments to obtain the first addition result; input the first addition result into the sparse MLP module, add the output of the sparse MLP module to the first addition result to obtain the second addition result; input the second addition result into the ConvLSTM;
[0017] The sparse SS2D module includes LayerNorm, Linear, DWConv, Sparse SS2D unit, Linear, and Spatial Attention connected in sequence;
[0018] The sparse MLP module includes LayerNorm, Sparse MLP unit, and channel Attention connected in sequence.
[0019] In one embodiment, tokens are generated based on the sparsified map from the output feature map of DWConv; the tokens include the features of the pixels with pixel values of 1.
[0020] The Sparse SS2D unit is used to implement the following functions:
[0021] The tokens are input into the Bidi-Scan module, and the features of the pixels with pixel values of 1 are sorted in the horizontal forward and horizontal reverse directions respectively to obtain the first sorting result and the second sorting result.
[0022] The tokens are divided into image patches, and the maximum value of the spatio-temporal continuity scores of each pixel in each image patch is determined, and the maximum value is used as the score of the image patch; the features of the pixels with pixel values of 1 in all image patches are sorted according to the scores to obtain the third sorting result.
[0023] The first sorting result, the second sorting result and the third sorting result are respectively input into the S6 module to extract the context relationship between the features, and the first sorting result, the second sorting result and the third sorting result containing the feature relationship are obtained.
[0024] The features of each pixel in the first sorting result, the second sorting result and the third sorting result containing the feature relationship are mapped back to the original positions in the tokens, and the mean value of the three features corresponding to each pixel is calculated as the new feature of the pixel.
[0025] In one embodiment, the first spatial-channel hybrid Mamba module includes 2 spatially-channel hybrid Mamba layers that are identical in structure and connected in sequence.
[0026] Each spatially-channel hybrid Mamba layer includes a Sparse SS2D module, a GCI module, channel Attention, and ConvLSTM.
[0027] The second pooled feature and the second downsampled feature are input into the Sparse SS2D module. The output of the Sparse SS2D module is added to the second downsampled feature to obtain a third addition result. The third addition result is input into the GCI module, and the output of the GCI module passes through channel Attention to obtain the channel attention feature.
[0028] The channel attention feature is added to the third addition result to obtain a fourth addition result; the fourth addition result is input into ConvLSTM.
[0029] In one embodiment, the GCI module includes two branches. The first branch includes a 1×1 convolutional layer, and the second branch includes Linear, DWConv, Bidi-channel Scan, and Linear connected in sequence;
[0030] After the outputs of the first branch and the second branch are added together, the result is used as the output of the GCI module.
[0031] In one embodiment, Bidi-channel Scan is used to implement the following functions:
[0032] After flattening and transposing the input along different dimensions, a first sequence is obtained; the first sequence is flipped to obtain a second sequence;
[0033] The first sequence is input into the S6 module to obtain a first output; the second sequence is input into the S6 module and flipped to obtain a second output;
[0034] The addition result obtained by adding the first output and the second output is then transposed and flattened to obtain the output of Bidi-channel Scan.
[0035] In a second aspect, a sparse Mamba-based event camera target detection device is provided, including:
[0036] An event stream conversion module, configured to convert the event stream acquired by the event camera into a voxel tensor and divide the voxel tensor into multiple segments;
[0037] A sparsification graph acquisition module, configured to accumulate the timestamps of all events at each pixel point based on the event stream to generate a temporal continuity score map; perform average pooling on the temporal continuity score map and perform smoothing denoising based on a Gaussian function to obtain a spatio-temporal continuity score map; binarize the spatio-temporal continuity score map to obtain a sparsification graph;
[0038] A first feature extraction module, configured to input the multiple segments and the sparsification graph into a first sparse spatial Mamba module for feature extraction to obtain a first feature;
[0039] A second feature extraction module, configured to perform max pooling on the sparsification graph to obtain a first pooled feature; perform downsampling on the first feature to obtain a first downsampled feature; input the first pooled feature and the first downsampled feature into a second sparse spatial Mamba module for feature extraction to obtain a second feature;
[0040] The third feature extraction module is used to perform max pooling on the first pooled feature to obtain the second pooled feature; downsample the second feature to obtain the second downsampled feature; input the second pooled feature and the second downsampled feature into the first spatio-channel hybrid Mamba module for feature extraction to obtain the third feature;
[0041] The fourth feature extraction module is used to perform max pooling on the second pooled feature to obtain the third pooled feature; downsample the third feature to obtain the third downsampled feature; input the third pooled feature and the third downsampled feature into the second spatio-channel hybrid Mamba module for feature extraction to obtain the fourth feature;
[0042] The feature fusion module is used to input the fourth feature, the third feature, and the second feature into the feature pyramid network for multi-scale feature fusion to obtain the fused feature;
[0043] The detection module is used to input the fused feature into the YOLOX detection head to obtain the object detection result.
[0044] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned event camera object detection method based on sparse Mamba.
[0045] In a fourth aspect, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the above-mentioned event camera object detection method based on sparse Mamba.
[0046] Compared with the prior art, the present application has the following beneficial effects: The event camera object detection method based on sparse Mamba of the present application adaptively discards eventless and noise labels based on spatio-temporal continuity evaluation, and simultaneously captures global dependencies in both spatial and channel dimensions, showing the best balance between accuracy and efficiency; proposes an information-first local scanning strategy (IPL-Scan) to guide the model to focus on high-information labels during the scanning process, thereby improving the spatial context modeling ability; designs a global channel interaction module (GCI) aimed at aggregating channel information from a global spatial perspective and extending global interaction to the 3D feature space to further improve the global modeling ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The present application can be better understood by referring to the description given below in conjunction with the accompanying drawings. The drawings, together with the following detailed description, are included in this specification and form a part of this specification. In the drawings:
[0048] Figure 1Shows the schematic diagram of the event camera target detection method based on Sparse Mamba;
[0049] Figure 2 Shows the schematic diagram of the Spatiotemporal Continuity Assessment (STCA) module;
[0050] Figure 3 Shows the structural schematic diagram of the Sparse Spatial Mamba layer;
[0051] Figure 4 Shows the schematic diagram of the Sparse SS2D unit;
[0052] Figure 5 Shows the structural schematic diagram of the Spatial-Channel Hybrid Mamba layer;
[0053] Figure 6 Shows the structural schematic diagram of the Bidi-channel Scan;
[0054] Figure 7 Shows the visualization diagrams of the original events, score maps, sparsification maps, and sparsification results on the eTram and 1Mpx datasets, where (a) is the original events on the eTram and 1Mpx datasets, (b) is the score maps on the eTram and 1Mpx datasets, (c) is the sparsification maps on the eTram and 1Mpx datasets, and (d) is the sparsification results on the eTram and 1Mpx datasets. Detailed implementation manners
[0055] In the following, exemplary embodiments of the present application will be described with reference to the accompanying drawings. For clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many specific decisions specific to the embodiments can be made during the development of any such actual embodiment to achieve the specific goals of the developer, and these decisions may vary with different embodiments.
[0056] Here, it should also be noted that, in order to avoid obscuring the present application due to unnecessary details, only the device structures closely related to the solution of the present application are shown in the drawings, while other details less related to the present application are omitted.
[0057] It should be understood that the present application is not limited to the described implementation forms only due to the following description with reference to the accompanying drawings. In this document, where feasible, embodiments can be combined with each other, features can be replaced or borrowed between different embodiments, and one or more features can be omitted in one embodiment.
[0058] The embodiments of the present application provide an event camera target detection method based on Sparse Mamba, Figure 1The schematic diagram of the event camera target detection method based on sparse Mamba is shown. See Figure 1 , and the method mainly includes the following steps:
[0059] Step S1: Convert the event stream obtained by the event camera into a voxel tensor, and divide the voxel tensor into multiple segments.
[0060] Given the event stream where (x i , y i ) are spatial coordinates, t i is the timestamp, p i is the event polarity, i is the event label, and N is the number of events.
[0061] Step S2: Based on the event stream, accumulate the timestamps of all events at each pixel point to generate a temporal continuity score map; perform average pooling operation on the temporal continuity score map, and perform smoothing denoising based on the Gaussian function to obtain a spatio-temporal continuity score map; binarize the spatio-temporal continuity score map to obtain a sparsification map.
[0062] Here, accumulating the timestamps of all times at each pixel point, the following formula is used:
[0063]
[0064] where is the temporal continuity score of the pixel point (x, y).
[0065] The temporal continuity scores of all pixel points form a temporal continuity score map.
[0066] Then, use average pooling with a kernel size and stride of P to extract the temporal information content corresponding to each label, where P is the block size used in the event tokenization process. Subsequently, neighborhood information is effectively aggregated to evaluate spatial continuity. For active events, nearby neighbors are more likely to be triggered by the same moving edge, while distant neighbors are more likely to be noise. Therefore, to reduce the impact of noise on the information quantity evaluation, a Gaussian function is used for distance-based weighted aggregation within the neighborhood, thereby smoothing the noise and maintaining a more complete target structure. In the obtained spatio-temporal continuity score map S st , each pixel value represents the information content of the active event of the label. The larger the value, the more important the label.
[0067] Then, calculate the average value α of the spatio-temporal continuity scores of all pixel points in the spatio-temporal continuity score map to represent the sparsity of the scene, and use it as the threshold for discarding uninformative labels. Binarize the spatio-temporal continuity score map according to α to generate a sparsification map D:
[0068]
[0069] D x,y To sparsify the pixel value of the pixel point (x, y) in graph D.
[0070] In this step, the event camera only triggers events asynchronously at positions where the brightness change exceeds the threshold, resulting in high spatial sparsity, especially when the camera is stationary. In addition, the inherent circuit characteristics of the event camera will inevitably generate a large amount of noise. These blank and noise regions do not contain valid information, which will lead to unnecessary calculations and potential interference. Through observation, there are significant differences in the spatio-temporal distribution between active events and noise events. Specifically, noise events are spatially isolated and temporally discontinuous, while active events are usually concentrated on the edges of moving objects, showing spatial proximity and temporal continuity. Based on this prior, a spatio-temporal continuity assessment (STCA) module is adopted here. Figure 2 Shows the schematic diagram of the spatio-temporal continuity assessment (STCA) module, which evaluates the importance of labels by assessing the spatio-temporal continuity of events and selectively discards uninformative labels, thereby reducing the computational overhead.
[0071] Step S3: Input multiple segments and the sparsified graph into the first sparse space Mamba module for feature extraction to obtain the first feature.
[0072] Step S4: Perform max pooling on the sparsified graph to obtain the first pooled feature; downsample the first feature to obtain the first downsampled feature; input the first pooled feature and the first downsampled feature into the second sparse space Mamba module for feature extraction to obtain the second feature.
[0073] Step S5: Perform max pooling on the first pooled feature to obtain the second pooled feature; downsample the second feature to obtain the second downsampled feature; input the second pooled feature and the second downsampled feature into the first spatial-channel hybrid Mamba module for feature extraction to obtain the third feature.
[0074] Step S6: Perform max pooling on the second pooled feature to obtain the third pooled feature; downsample the third feature to obtain the third downsampled feature; input the third pooled feature and the third downsampled feature into the second spatial-channel hybrid Mamba module for feature extraction to obtain the fourth feature.
[0075] Step S7: Input the fourth feature, the third feature, and the second feature into the feature pyramid network for multi-scale feature fusion to obtain the fused feature.
[0076] Step S8: Input the fused feature into the YOLOX detection head to obtain the target detection result.
[0077] In this embodiment, first, the event stream is input into the Spatiotemporal Continuity Assessment (STCA) module, which generates a sparsification graph to guide the sparsification operation. At the same time, the event stream is converted into a voxel tensor and divided into segments for tokenization. Then, these tokens undergo multi-scale feature extraction through four stages. The first two stages use the Sparse Spatial Mamba (SSM) module, which is used to enhance the global spatial interaction between the retained tokens, further reduce the computational overhead, and transmit spatiotemporal information across time steps, and its output is transmitted to the subsequent layer. The last two stages use the Spatial-Channel Mixed Mamba (SCMM) module, which promotes channel interaction from a global perspective, thereby extending global modeling to three-dimensional representations. The features generated by the last three stages are input into the Feature Pyramid Network (FPN) for multi-scale feature fusion. Finally, the YOLOX detection head is used to output the detection results. This embodiment can adaptively discard eventless and noisy tokens based on spatiotemporal continuity assessment, and at the same time capture global dependencies in both spatial and channel dimensions, showing the best balance between accuracy and efficiency.
[0078] In one embodiment, the first Sparse Spatial Mamba module includes 2 Sparse Spatial Mamba layers with the same structure and connected in sequence; Figure 3 The structural schematic diagram of the Sparse Spatial Mamba layer is shown, see Figure 3 , each Sparse Spatial Mamba layer includes a Sparse SS2D module, a Sparse MLP module, and a ConvLSTM; here, ConvLSTM is the Convolutional Long Short-Term Memory network, which is a deep learning model for processing sequence data.
[0079] Multiple segments and the sparsification graph are input into the Sparse SS2D module for two-dimensional selective scanning. The output of the Sparse SS2D module is added to the multiple segments to obtain a first addition result; the first addition result is input into the Sparse MLP module, and the output of the Sparse MLP module is added to the first addition result to obtain a second addition result; the second addition result is input into the ConvLSTM;
[0080] The Sparse SS2D module includes LayerNorm (Layer Normalization), Linear (Linear Layer), DWConv (Depthwise Separable Convolution), Sparse SS2D unit, Linear (Linear Layer), and Spatial Attention connected in sequence;
[0081] The Sparse MLP module includes LayerNorm, Sparse MLP unit, and channel Attention connected in sequence.
[0082] In this embodiment, the sparse SS2D module is used to enhance the global spatial interaction between the retained tokens, the sparse MLP module is used to further reduce the computational overhead, and the ConvLSTM is used to transmit spatio-temporal information across time steps and transmit its output to the subsequent layers.
[0083] It should be noted that the structure of the second sparse spatial Mamba module is the same as that of the first sparse spatial Mamba module, and reference can be made to the above embodiment.
[0084] Specifically, Figure 4 shows the schematic diagram of the Sparse SS2D unit. Refer to Figure 4 , generate Tokens based on the sparsified graph for the feature map output by DWConv, where the pixels in the feature map correspond one-to-one with the pixels in the sparsified graph; the Tokens include the features of the pixels with a pixel value of 1 in the feature map; the Sparse SS2D unit is used to implement the following functions:
[0085] The Tokens are input into the Bidi-Scan module (bidirectional scanning module), and the features of the pixels with a pixel value of 1 are sorted in the horizontal forward and horizontal reverse directions to obtain the first sorting result and the second sorting result respectively;
[0086] Divide the Tokens into image patches, and determine the maximum value of the spatio-temporal continuity scores of each pixel in each image patch, and use the maximum value as the score of the image patch; sort the features of the pixels with a pixel value of 1 in all image patches according to the scores to obtain the third sorting result;
[0087] Input the first sorting result, the second sorting result, and the third sorting result into the S6 module to extract the context relationship between the features, and obtain the first sorting result, the second sorting result, and the third sorting result containing the feature relationship; here, the S6 module is a complex component in the Mamba architecture, which is mainly composed of a series of linear transformations and discretization processes and is used to process the input feature sequence. It plays a crucial role in capturing the temporal dynamic features of the sequence, and the temporal dynamic features are a key aspect of sequence modeling tasks such as language modeling.
[0088] Map the features of each pixel in the first sorting result, the second sorting result, and the third sorting result containing the feature relationship back to the original positions in the Tokens, and calculate the mean value of the three features corresponding to each pixel as the new feature of the pixel.
[0089] In the above embodiments, the spatio-temporal continuity score map quantifies the information content of the markers. The higher the score of a marker, the greater the likelihood of representing a foreground target. Based on this map, the markers are re-ordered, and the markers with higher information content are processed earlier, thereby shortening the scanning distance between important markers and promoting the interaction between them. In addition, the markers with lower scores are processed later, thereby reducing the interference caused by noise.
[0090] Considering that direct re-ordering may destroy local information, local constraints are introduced during the sorting process. When a marker is processed, its k×k neighborhood (image patch) is also processed immediately. Specifically, the maximum value is extracted from each k×k local window using max pooling with a kernel and stride of k. These maximum values representing the local windows are initially sorted. Subsequently, the sorted result is upsampled by a factor of k to obtain the window-level sorting result. This strategy effectively promotes the interaction between potential target regions while retaining important local information.
[0091] In one embodiment, the first spatio-channel hybrid Mamba module includes two spatio-channel hybrid Mamba layers with the same structure and connected in sequence; Figure 5 The structural schematic diagram of the spatio-channel hybrid Mamba layer is shown. Refer to Figure 5 , each spatio-channel hybrid Mamba layer includes a sparse SS2D module, a GCI module (global channel interaction module), channel Attention, and ConvLSTM;
[0092] The second pooled feature and the second downsampled feature are input into the sparse SS2D module. The output of the sparse SS2D module is added to the second downsampled feature to obtain a third addition result. The third addition result is input into the GCI module, and the output of the GCI module passes through channel Attention to obtain the channel attention feature;
[0093] The channel attention feature is added to the third addition result to obtain a fourth addition result; the fourth addition result is input into ConvLSTM.
[0094] In this embodiment, the GCI module promotes channel interaction from a global perspective, thereby extending global modeling to three-dimensional representations. In this embodiment, two-dimensional spatial scanning strategies, such as Bidi-Scan, may disperse the markers related to the same target in the scanning sequence, resulting in a relatively large scanning interval and weakening the interaction between them. Therefore, this embodiment proposes Information-Priority Local Scan (IPL-Scan) to alleviate the limitations of two-dimensional scanning methods, and designs a Sparse SS2D unit to combine IPL-Scan and Bidi-Scan to promote global interaction.
[0095] It should be noted that the second spatial-channel hybrid Mamba module has the same structure as the first spatial-channel hybrid Mamba module, and reference can be made to this embodiment.
[0096] Specifically, referring to Figure 5 , the GCI module includes two branches. The first branch includes a 1×1 convolutional layer, and the second branch includes Linear, DWConv, Bidi-channel Scan, and Linear connected in sequence;
[0097] After the outputs of the first branch and the second branch are added, the result is used as the output of the GCI module.
[0098] Specifically, Figure 6 shows a schematic structural diagram of Bidi-channel Scan. Referring to Figure 5 , Bidi-channelScan is used to implement the following functions:
[0099] After flattening and transposing the input along different dimensions, a first sequence is obtained; the first sequence is flipped to obtain a second sequence;
[0100] The first sequence is input into the S6 module to obtain a first output; the second sequence is input into the S6 module and flipped to obtain a second output;
[0101] The addition result obtained by adding the first output and the second output is then transposed and flattened to obtain the output of Bidi-channel Scan.
[0102] In this embodiment, in the first branch, 1×1 convolution is used to capture pixel-level dependencies between channels, thereby realizing local adaptive interaction. In the second branch, after linear and DWConv preprocessing to capture local context, it is then fed into the bidirectional channel scan Bidi-channel Scan. Through flattening and transposing operations, the global spatial information of each channel is used as the basic unit of interaction. Subsequently, a reverse sequence is generated by flipping, and it is input into S6 together with the original sequence to achieve adaptive interaction from a global perspective. Selective scanning based on the global spatial context enables each channel to selectively focus on other channels from a more comprehensive global perspective, accurately capture the dependencies between channels, and further enhance the global modeling ability. Finally, the results of the two branches are integrated to achieve comprehensive channel interaction.
[0103] To further verify the effectiveness of the method of this application, the following experimental analysis was carried out.
[0104] Comparative experiments were conducted on two autonomous driving datasets, Gen1 and 1Mpx, and a traffic monitoring dataset, eTram. The COCO mAP (mean average precision) was used to evaluate the accuracy of object detection. The size of the model was measured by the total number of parameters. In addition, the average FLOPs (floating point operations per second) of the first 1000 samples in the test set were calculated to evaluate the computational complexity.
[0105] On the Gen1 and 1Mpx datasets, the method of this application was compared and analyzed with two CNN-based methods: RED and ASTMNet; and five Transformer-based methods: ERGO-12, RVT, GET, SAST, and S5-ViT. On the eTram dataset, the method of this application was compared with three Transformer-based methods: RVT, SAST, and S5-ViT. To compare with the SSM-based method, the VSS block in VMamba was used to construct a detection framework named VSS. In addition, a baseline model without using the sparsification strategy was constructed to evaluate the effectiveness of the method of this application.
[0106] Table 1 Comparison results of detection performance with state-of-the-art methods on the autonomous driving datasets Gen1 and 1Mpx
[0107]
[0108] The results are shown in Table 1 and Table 2. On the Gen1 dataset, the SMamba of this application outperformed all other methods with the lowest FLOPs and the number of parameters. Compared with ERGO-12, SMamba achieved the same mAP with only 5% of the FLOPs and 27% of the number of parameters. On the 1Mpx and eTram datasets, the baseline method of this application was better than all SOTA methods. By integrating the sparsification strategy of this application into the baseline, the FLOPs of SMamba were reduced by 22% and 31% respectively, and the mAP was increased by 0.1% and 0.3% respectively, exceeding all other methods while maintaining the lowest number of parameters. The sparsification operation of this application enables the network to focus on important regions, alleviates the interference of blank and noise regions, thereby reducing the computational overhead and improving the accuracy. The consistent performance improvement on the autonomous driving and traffic monitoring datasets indicates that the method of this application can be generalized to different sparse levels and achieve an excellent balance between accuracy and efficiency. Event sparse and event dense
[0109] Table 2 Comparison results of detection performance with state-of-the-art methods on the intelligent monitoring dataset eTram
[0110]
[0111] Figure 7Visualization diagrams of the original events, score maps, sparsification maps, and sparsification results on the eTram and 1Mpx datasets are shown. Among them, (a) shows the original events on the eTram and 1Mpx datasets, (b) shows the score maps on the eTram and 1Mpx datasets, (c) shows the sparsification maps on the eTram and 1Mpx datasets, and (d) shows the sparsification results on the eTram and 1Mpx datasets. The scene complexity of the two diagrams in the same dataset increases sequentially. The eTram dataset collected by a stationary camera shows higher sparsity than the 1Mpx dataset collected by a moving camera. As the event density increases, the STCA module retains more markers. This indicates that the STCA of this application has a strong scene adaptation ability. While selecting important markers, it can effectively reduce the interference of blank areas and noise.
[0112] Adopting the same inventive concept as the event camera target detection method based on sparse Mamba, this embodiment also provides a corresponding event camera target detection device based on sparse Mamba, including:
[0113] An event stream conversion module, configured to convert the event stream obtained by the event camera into a voxel tensor and divide the voxel tensor into multiple segments;
[0114] A sparsification map acquisition module, configured to accumulate the timestamps of all events at each pixel point based on the event stream to generate a temporal continuity score map; perform average pooling operation on the temporal continuity score map, and perform smoothing and denoising based on the Gaussian function to obtain a spatio-temporal continuity score map; perform binarization on the spatio-temporal continuity score map to obtain a sparsification map;
[0115] A first feature extraction module, configured to input the multiple segments and the sparsification map into a first sparse spatial Mamba module for feature extraction to obtain a first feature;
[0116] A second feature extraction module, configured to perform max pooling on the sparsification map to obtain a first pooled feature; perform downsampling on the first feature to obtain a first downsampled feature; input the first pooled feature and the first downsampled feature into a second sparse spatial Mamba module for feature extraction to obtain a second feature;
[0117] A third feature extraction module, configured to perform max pooling on the first pooled feature to obtain a second pooled feature; perform downsampling on the second feature to obtain a second downsampled feature; input the second pooled feature and the second downsampled feature into a first spatial-channel hybrid Mamba module for feature extraction to obtain a third feature;
[0118] The fourth feature extraction module is used to perform max pooling on the second pooled feature to obtain the third pooled feature; downsample the third feature to obtain the third downsampled feature; the third pooled feature and the third downsampled feature are input into the second spatial-channel hybrid Mamba module for feature extraction to obtain the fourth feature;
[0119] The feature fusion module is used to input the fourth feature, the third feature, and the second feature into the feature pyramid network for multi-scale feature fusion to obtain the fused feature;
[0120] The detection module is used to input the fused feature into the YOLOX detection head to obtain the object detection result.
[0121] The event camera object detection device based on sparse Mamba in this embodiment has the same inventive concept as the above-mentioned event camera object detection method based on sparse Mamba. Therefore, the specific implementation of this device can be seen in the embodiment part of the event camera object detection method based on sparse Mamba in the previous text, and its technical effects correspond to those of the above method, which will not be elaborated here.
[0122] This application embodiment provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the above-mentioned event camera object detection method based on sparse Mamba.
[0123] This application embodiment provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, it is used to implement the above-mentioned event camera object detection method based on sparse Mamba.
[0124] As mentioned above, these are only various implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for event camera target detection based on sparse Mamba, characterized in that: include: Convert an event stream acquired by an event camera into a voxel tensor, and divide the voxel tensor into a plurality of segments; Based on the event stream, the timestamps of all events at each pixel are accumulated to generate a time continuity score map; Performing an average pooling operation on the temporal continuity score map, and performing smoothing and denoising based on a Gaussian function to obtain a spatiotemporal continuity score map; binarizing the spatiotemporal continuity score map to obtain a sparse map; Inputting the plurality of segments and the sparse graph into a first sparse space Mamba module for feature extraction to obtain a first feature; Performing maximum pooling on the sparse graph to obtain a first pooling feature; downsampling the first feature to obtain a first downsampling feature; inputting the first pooling feature and the first downsampling feature into a second sparse space Mamba module for feature extraction to obtain a second feature; Performing maximum pooling on the first pooled feature to obtain a second pooled feature; downsampling the second feature to obtain a second downsampled feature; inputting the second pooled feature and the second downsampled feature into a first space-channel hybrid Mamba module for feature extraction to obtain a third feature; Performing maximum pooling on the second pooling feature to obtain a third pooling feature; downsampling the third feature to obtain a third downsampled feature; inputting the third pooling feature and the third downsampled feature into a second space-channel hybrid Mamba module for feature extraction to obtain a fourth feature; The fourth feature, the third feature, and the second feature are all input into a feature pyramid network for multi-scale feature fusion to obtain a fused feature; The fused features are input into the YOLOX detection head to obtain the target detection result.
2. The method according to claim 1, characterized in that The first sparse space Mamba module includes two sparse space Mamba layers having the same structure and connected in sequence; Each of the sparse spatial Mamba layers includes a sparse SS2D module, a sparse MLP module, and a ConvLSTM; The multiple segments and the sparse graph are input into the sparse SS2D module for two-dimensional selective scanning, the output of the sparse SS2D module is added to the multiple segments to obtain a first addition result; the first addition result is input into the sparse MLP module, the output of the sparse MLP module is added to the first addition result to obtain a second addition result; the second addition result is input into the ConvLSTM; The sparse SS2D module includes LayerNorm, Linear, DWConv, Sparse SS2D unit, Linear, and SpatialAttention connected in sequence; The sparse MLP module includes LayerNorm, Sparse MLP unit, and channel Attention which are connected in sequence.
3. The method according to claim 2, characterized in that Generate Tokens based on the sparse graph from the feature graph output by the DWConv; the Tokens include features of pixels whose pixel values are 1; The Sparse SS2D unit is used to implement the following functions: The tokens are input into the Bidi-Scan module, and the features of pixels with a pixel value of 1 are sorted in a horizontal forward direction and a horizontal reverse direction to obtain a first sorting result and a second sorting result respectively; Divide the Tokens into image blocks, and determine the maximum value of the spatiotemporal continuity scores of each pixel in each image block, and use the maximum value as the score of the image block; sort the features of pixels with pixel values of 1 in all image blocks according to the scores to obtain a third sorting result; Input the first sorting result, the second sorting result and the third sorting result into the S6 module respectively to extract the contextual relationship between the features, so as to obtain the first sorting result, the second sorting result and the third sorting result including the feature relationship; The features of each pixel in the first sorting result, the second sorting result and the third sorting result containing the feature relationship are mapped back to the original position in Tokens, and the average of the three features corresponding to each pixel is calculated as the new feature of the pixel.
4. The method according to claim 1, characterized in that The first space-channel hybrid Mamba module includes two space-channel hybrid Mamba layers having the same structure and connected in sequence; Each of the spatial-channel hybrid Mamba layers includes a sparse SS2D module, a GCI module, a channelAttention and a ConvLSTM; The second pooling feature and the second down-sampled feature are input into the sparse SS2D module, the output of the sparse SS2D module is added to the second down-sampled feature to obtain a third addition result, the third addition result is input into the GCI module, and the output of the GCI module is passed through channelAttention to obtain a channel attention feature; The channel attention feature is added to the third addition result to obtain a fourth addition result; the fourth addition result is input into the ConvLSTM.
5. The method according to claim 4, characterized in that The GCI module includes two branches, the first branch includes a 1×1 convolutional layer, and the second branch includes a Linear, a DWConv, a Bidi-channel Scan, and a Linear connected in sequence; The outputs of the first branch and the second branch are added together to serve as the output of the GCI module.
6. The method according to claim 5, characterized in that The Bidi-channel Scan is used to achieve the following functions: After flattening and transposing the input along different dimensions, a first sequence is obtained; the first sequence is flipped to obtain a second sequence; The first sequence is input into the S6 module to obtain a first output; The second sequence is input into the S6 module and flipped to obtain a second output; The addition result obtained by adding the first output and the second output is transposed and flattened to obtain the output of the Bidi-channel Scan.
7. An event camera target detection device based on sparse Mamba, characterized in that: include: An event stream conversion module, used for converting the event stream acquired by the event camera into a voxel tensor, and dividing the voxel tensor into a plurality of segments; A sparse graph acquisition module, used to accumulate the timestamps of all events at each pixel point based on the event stream to generate a time continuity score graph; Performing an average pooling operation on the temporal continuity score map, and performing smoothing and denoising based on a Gaussian function to obtain a spatiotemporal continuity score map; Binarizing the spatiotemporal continuity score map to obtain a sparse map; A first feature extraction module, configured to input the plurality of segments and the sparse graph into a first sparse space Mamba module for feature extraction to obtain a first feature; A second feature extraction module is used to perform maximum pooling on the sparse graph to obtain a first pooling feature; downsample the first feature to obtain a first downsampled feature; the first pooling feature and the first downsampled feature are input into a second sparse space Mamba module for feature extraction to obtain a second feature; A third feature extraction module is used to perform maximum pooling on the first pooled feature to obtain a second pooled feature; downsample the second feature to obtain a second downsampled feature; the second pooled feature and the second downsampled feature are input into the first space-channel hybrid Mamba module for feature extraction to obtain a third feature; A fourth feature extraction module is used to perform maximum pooling on the second pooling feature to obtain a third pooling feature; downsample the third feature to obtain a third downsampled feature; the third pooling feature and the third downsampled feature are input into a second space-channel hybrid Mamba module for feature extraction to obtain a fourth feature; A feature fusion module, used for inputting the fourth feature, the third feature, and the second feature into a feature pyramid network for multi-scale feature fusion to obtain a fused feature; The detection module is used to input the fusion features into the YOLOX detection head to obtain the target detection result.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the event camera target detection method based on sparse Mamba is implemented as described in any one of claims 1 to 6.
9. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the event camera target detection method based on sparse Mamba as described in any one of claims 1 to 6.
Citation Information
Cited By
Sea temperature complementation method and system based on recursive double-current Mama
CN120543373A
Sea temperature completion method and system based on recursive dual-stream Mamba
CN120543373B
Hotel room type matching method and device
CN120687466A