Mamba-based polarization-event-radar multi-modal collaborative perception method
By employing Mamba's polarization-event-radar multimodal collaborative sensing method, and utilizing bi-branch feature extraction and a state-space model to fuse biomimetic polarization images, event images, and radar point cloud data, the accuracy and robustness issues of sensing technology in underground scenes are addressed, achieving high-precision environmental perception and target detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTH CHINA UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2026-01-15
- Publication Date
- 2026-06-26
AI Technical Summary
Existing sensing technologies struggle to provide high accuracy and robustness in underground scenarios, especially in low-light or dynamically changing environments. Traditional single-sensor or simple data fusion technologies are insufficient to meet the requirements, and existing multimodal collaborative sensing systems are not yet mature.
A Mamba-based polarization-event-radar multimodal collaborative sensing method is adopted. The biomimetic polarization image features and event image features are extracted by a dual-branch feature extraction network, and the Mamba sequential deep learning architecture driven by the state space model (SSM) is used for feature fusion. Combined with the bird's-eye view feature encoding of radar point cloud data, multimodal feature fusion is achieved.
It improves the accuracy and environmental adaptability of target detection, enabling high-precision perception in complex underground environments and supporting applications in fields such as autonomous driving, robotics, underground space operations, and mineral resource exploration.
Smart Images

Figure CN121883822B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sensor fusion technology, and in particular to a polarization-event-radar multimodal cooperative sensing method based on Mamba. Background Technology
[0002] Underground scenarios, including mineral exploration, subway tunnels, and underground space operations, present unique environmental challenges that place even more stringent requirements on perception systems.
[0003] First, the lighting conditions in underground environments are often unstable, with flickering or complete absence of light, directly impacting the performance of light-dependent sensors, such as traditional cameras. Second, the confined spaces and complex, variable structures make it difficult for traditional single-sensor or simple data fusion techniques to provide the required high accuracy and robustness. Furthermore, the dynamic nature of underground environments demands high adaptability and flexibility from the sensing system. Movement of construction equipment, changes in geological structures, or other temporary environmental alterations require the sensing system to respond quickly and update its understanding of the scene in real time.
[0004] Therefore, the application of existing sensing technologies in underground scenarios is significantly limited, and there is an urgent need for a solution that can integrate data from multiple sensors to provide more comprehensive, accurate and stable sensing capabilities. Summary of the Invention
[0005] In view of this, embodiments of this application provide a polarization-event-radar multimodal collaborative sensing method based on Mamba to solve the problem in the prior art of how to fully integrate multi-sensor data in highly dynamic underground scenarios to improve the accuracy and robustness of environmental perception.
[0006] A first aspect of this application provides a polarization-event-radar multimodal cooperative sensing method based on Mamba, comprising:
[0007] Acquire images from a biomimetic polarization camera, event camera images, and radar point cloud data;
[0008] A dual-branch feature extraction network is used to extract biomimetic polarization image features and event image features, respectively; the event image features are temporal features obtained by a multi-scale event aggregation method.
[0009] The Mamba sequence deep learning architecture, driven by the State Space Model (SSM), fuses features from biomimetic polarization images and event images to obtain fused image features. Here, the event image features are the state variables of the Mamba decoder layer, and the biomimetic polarization image features are updated through multiple decoder layers to obtain the fused image features.
[0010] Bird's-Eye-View (BEV) feature encoding is performed on radar point cloud data to obtain the BEV encoded features of radar point cloud data;
[0011] Multimodal fusion features are obtained by fusing the coded features of fused image data and radar point cloud data based on Mamba.
[0012] The set of detection targets should be determined based on at least multimodal fusion features.
[0013] A second aspect of this application provides a Mamba-based polarization-event-radar multimodal cooperative sensing device, comprising:
[0014] The acquisition module is configured to acquire images from a biomimetic polarization camera, images from an event camera, and radar point cloud data.
[0015] The feature extraction module is configured to extract biomimetic polarization image features and event image features using a dual-branch feature extraction network; the event image features are temporal features obtained using a multi-scale event aggregation method.
[0016] The fusion module is configured to fuse features from biomimetic polarization image features and event image features using the Mamba sequence deep learning architecture driven by the state-space model (SSM) to obtain fused image features. The event image features are the state variables of the Mamba decoder layer, and the biomimetic polarization image features are updated through multiple decoder layers to obtain the fused image features.
[0017] The encoding module is configured to perform bird's-eye view BEV feature encoding on radar point cloud data to obtain the BEV encoded features of radar point cloud data.
[0018] The fusion module is also configured to fuse the coded features of the fused image and the radar point cloud data based on Mamba to obtain multimodal fused features;
[0019] The perception module is configured to determine the set of detection targets based at least on multimodal fusion features.
[0020] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0021] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0022] The beneficial effects of the embodiments in this application compared with the prior art are:
[0023] This application uses a dual-branch feature extraction network to extract biomimetic polarization image features and event image features respectively. It then uses SSM-driven Mamba to fuse the biomimetic polarization image features and event image features to obtain fused image features. Additionally, it acquires the BEV feature encoding of radar point cloud data. Using Mamba, it fuses the fused image features and the encoded features of the radar point cloud data to obtain multimodal fused features, thereby determining the target set. This fully utilizes the data characteristics of different sensors to improve the accuracy of target detection and enhances the ability to adapt to complex environments. It can provide strong technical support for fields such as autonomous driving, embodied intelligence, robotics, underground space operations, and mineral resource exploration, and can achieve high-precision perception of targets and the environment, thus realizing excellent environmental perception capabilities and demonstrating broad application potential and significant practical value. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating a polarization-event-radar multimodal cooperative sensing method based on Mamba, provided in an embodiment of this application.
[0026] Figure 2 This is a flowchart illustrating the method for obtaining event image features using a multi-scale event aggregation method provided in this application embodiment.
[0027] Figure 3 This is a flowchart illustrating the method for extracting biomimetic polarization image features provided in an embodiment of this application.
[0028] Figure 4 This is a flowchart illustrating the method for BEV feature encoding of radar point cloud data provided in this application embodiment.
[0029] Figure 5 This is a flowchart illustrating a method for fusing encoded features of fused image features and radar point cloud data based on Mamba, as provided in an embodiment of this application.
[0030] Figure 6 This is a flowchart illustrating a method for determining a set of detection targets based on at least multimodal fusion features, as provided in an embodiment of this application.
[0031] Figure 7 This is a schematic diagram of a polarization-event-radar multimodal cooperative sensing device based on Mamba, provided in an embodiment of this application.
[0032] Figure 8 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0033] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0034] The following will describe in detail, with reference to the accompanying drawings, a polarization-event-radar multimodal cooperative sensing method and apparatus based on Mamba according to embodiments of this application.
[0035] As mentioned above, the application of existing sensing technologies in underground scenarios is significantly limited. Specifically:
[0036] 1) Traditional sensing technologies generally rely on single sensors such as radar and RGB cameras. Such solutions are poorly adaptable in low light or dynamic and complex environments. For example, radar performance is significantly degraded in strong reflection and multipath interference scenarios, and RGB cameras cannot collect effective and clear environmental images when there is insufficient light.
[0037] 2) Existing data fusion methods are mostly simple methods such as direct superposition and mean weighting. Not only is it difficult to fully integrate the technical advantages of different sensors, but it is also easy to introduce additional errors due to data compatibility issues, reducing the credibility of the perception results.
[0038] 3) In highly dynamic underground scenarios, the low radar scanning frequency can cause untimely tracking of dynamic targets, while the conventional exposure mode of the polarization camera will cause motion blur, affecting the accurate identification of the target outline.
[0039] 4) Under extreme lighting conditions such as strong light glare or dim light, traditional visual perception solutions are prone to failure, exhibiting obvious technical "long tail shortcomings".
[0040] Meanwhile, although some solutions attempt to eliminate motion blur through the microsecond-level response of event cameras, filter out glare by utilizing the physical characteristics of polarization cameras, and improve the above problems by combining radar active detection capabilities, an integrated multimodal collaborative sensing system has not yet been formed.
[0041] In view of this, this application provides a polarization-event-radar multimodal collaborative sensing method based on Mamba. By using multi-sensor fusion technology, it overcomes the above-mentioned limitations and improves the object perception accuracy and environmental perception capability, thereby providing richer and more reliable environmental information for highly dynamic underground scenarios to support applications such as safe navigation, accurate mapping, efficient construction and real-time monitoring.
[0042] Figure 1 This is a flowchart illustrating a polarization-event-radar multimodal cooperative sensing method based on Mamba, provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps:
[0043] In step S101, images from a biomimetic polarization camera, images from an event camera, and radar point cloud data are acquired.
[0044] In step S102, a dual-branch feature extraction network is used to extract biomimetic polarization image features and event image features, respectively.
[0045] Among them, the event image features are temporal features obtained using a multi-scale event aggregation method.
[0046] In step S103, the biomimetic polarization image features and event image features are fused using Mamba driven by SSM to obtain fused image features.
[0047] Among them, the event image features are the state variables of the Mamba decoder layer, and the biomimetic polarization image features are updated through multiple decoder layers to obtain the fused image features.
[0048] In step S104, BEV feature encoding is performed on the radar point cloud data to obtain the BEV encoded features of the radar point cloud data.
[0049] In step S105, the encoded features of the fused image and the radar point cloud data are fused based on Mamba to obtain multimodal fused features.
[0050] In step S106, the set of detection targets is determined at least based on multimodal fusion features.
[0051] In some embodiments of this application, the method can be executed by a server or by a terminal device with certain processing capabilities, for collaborative sensing of multimodal data such as bionic polarization cameras, event cameras and radar, thereby achieving high-precision target detection in specific scenarios such as highly dynamic underground environments.
[0052] In some embodiments of this application, images from a biomimetic polarization camera, images from an event camera, and radar point cloud data can be acquired first.
[0053] On one hand, a dual-branch feature extraction network can be used to extract biomimetic polarization image features and event image features separately. Then, based on SSM-driven Mamba, the biomimetic polarization image features and event image features are fused to obtain fused image features. When fusing the biomimetic polarization image features and event image features, the event image features are the state variables of the Mamba decoder layer, and the biomimetic polarization image features are updated through multiple decoder layers to obtain the fused image features.
[0054] On the other hand, bird's-eye view BEV feature encoding can also be performed on radar point cloud data to obtain the BEV encoded features of radar point cloud data.
[0055] Next, based on Mamba, the encoded features of the fused image and radar point cloud data can be fused to obtain multimodal fused features, and then the set of detection targets can be determined.
[0056] According to the technical solution provided in the embodiments of this application, a dual-branch feature extraction network is used to extract biomimetic polarization image features and event image features respectively. Mamba driven by SSM is used to fuse the biomimetic polarization image features and event image features to obtain fused image features. BEV feature encoding of radar point cloud data is also obtained. Mamba is then used to fuse the fused image features and the encoded features of the radar point cloud data to obtain multimodal fused features, thereby determining the set of detection targets. This fully utilizes the data characteristics of different sensors to improve the accuracy of target detection and enhances the ability to adapt to complex environments. It can provide strong technical support for fields such as autonomous driving, embodied intelligence, robotics, underground space operations, and mineral resource exploration, and can achieve high-precision perception of targets and the environment, thus achieving excellent environmental perception capabilities, demonstrating broad application potential and significant practical value.
[0057] In some embodiments of this application, before acquiring images from the bionic polarization camera, event camera, and radar point cloud data, each sensor can be calibrated separately to obtain the parameters of each sensor, as well as the rotation matrix and translation vector between different sensors; wherein the sensors include a bionic polarization camera, an event camera, and radar; then, timestamps are assigned to the measurement data of each sensor, and data synchronization of each sensor is achieved through timestamp interpolation and clock drift compensation algorithms. After completing the above preparations, each sensor can be put into operation, thereby acquiring images from the bionic polarization camera, event camera, and radar point cloud data.
[0058] In some implementations, calibrating each sensor separately may include: first, selecting a standard calibration reference target, which may include, but is not limited to, a calibration plate with a geometric pattern; then, placing the calibration reference target at multiple positions and angles within the sensor's field of view to ensure that the calibration reference target is clearly captured by all sensors; selecting the radar coordinate system as the reference coordinate system, establishing a data transformation relationship between the reference coordinate system and the coordinate systems of the bionic polarization camera and the event camera; calculating the optimal rotation matrix and translation vector based on the same captured calibration reference target using an optimization algorithm to transform the bionic polarization camera coordinate system, the event camera coordinate system, and the reference coordinate system; finally, uniformly transforming the data from the bionic polarization camera and the event camera to the radar coordinate system based on the calculated rotation matrix and translation vector to achieve multi-sensor data synchronization.
[0059] Taking the transformation of vector P from coordinate system A of one sensor to coordinate system B of another sensor as an example, the transformation formula between different sensor coordinate systems is as follows: ;in, R is a vector in the original coordinate system A, R is the rotation matrix, and t is the translation vector. It is a vector in the transformed coordinate system B. The calibration process can involve sampling the same vector P multiple times in different coordinate systems and simultaneously calculating the rotation matrix R and the translation vector t.
[0060] The data synchronization mechanism achieves time alignment of multi-sensor data by establishing a unified time benchmark. This includes, but is not limited to, selecting the event camera data timestamp as the synchronization benchmark. First, the original timestamp sequences of the event camera, bionic polarization camera, and radar are extracted. By comparing the timestamp data of the three types of sensors, the sensor with the lowest output frequency is selected as the anchor sensor. Its data acquisition time is used as the global synchronization anchor point. The time offset of the bionic polarization camera and radar relative to the selected benchmark (such as the event camera) is calculated. Based on this offset, the original timestamps of the two types of sensors are calibrated and adjusted. An alignment window is set in combination with the acquisition period of the anchor sensor. Through data matching, weighted fusion, or linear interpolation completion strategies, the image data of the bionic polarization camera, the point cloud data of the radar, and the event camera data are made to correspond precisely on the same time axis, ultimately achieving time synchronization of multi-sensor data.
[0061] Figure 2 This is a flowchart illustrating the method for obtaining event image features using a multi-scale event aggregation method provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps:
[0062] In step S201, an asynchronous event stream is generated by measuring the pixel brightness changes of the event camera image.
[0063] In step S202, the time is obtained from the asynchronous event stream. The corresponding n scale events.
[0064] Among them, events of different scales correspond to events in the asynchronous event stream within different time ranges, with the largest scale event corresponding to the largest time range of this aggregation, and the remaining time ranges being... ; The exposure time for a single frame of an image captured by a bionic polarization camera. It is a positive integer less than n. The aggregation time range is defined, where n is a positive integer greater than or equal to 3.
[0065] In step S203, the n scale events are projected onto the feature space, and the projected event features are pooled.
[0066] Specifically, the pooling size for pooling small-scale events is smaller than that for pooling large-scale events, and the time range corresponding to small-scale events is smaller than that corresponding to large-scale events.
[0067] In step S204, the pooled event features of events other than the largest scale event are upsampled, and the upsampled event features are concatenated and convolved with the pooled event features of the largest scale event to obtain the event features within the time range. Inner Time Event image features.
[0068] Event cameras record brightness changes in a scene asynchronously, without capturing images at a fixed frame rate. Therefore, asynchronous event streams can be generated by measuring pixel brightness changes in the image.
[0069] Multi-scale event aggregation can be applied to this asynchronous event stream to obtain event image features within the current aggregation time range. For example, it can be applied to any time range. Aggregate asynchronous event streams within that time frame. At this point, events can be aggregated within that time frame. Further, multiple time ranges of different scales can be constructed. Taking three scales as an example, it can provide... , and All events occurring within each time frame are called scale events within that time frame. The three scale events corresponding to the three time frames mentioned above can be denoted as follows: , and .
[0070] Events at different scales can be projected into the feature space separately, with the projection weights shared for each scale event input. In this case, the projection step size can be expressed as... , ;in This indicates the projection module.
[0071] Pooling can be applied to the projected event features separately. Coarse pooling is used for event features with larger projection steps, while fine pooling is used for event features with smaller projection steps. Formally, pooling with different window sizes can be implemented. Features of multi-scale events at the highest level Applying max pooling .
[0072] To combine them, features of smaller shapes are upsampled to adjust the resolution. Finally, all features are concatenated and fed into a convolution to generate event image features. , where up represents upsampling for matching resolution.
[0073] Figure 3 This is a flowchart illustrating the method for extracting biomimetic polarization image features provided in an embodiment of this application. For example... Figure 3 As shown, the method includes the following steps:
[0074] In step S301, the time... Image degradation processing is performed on the biomimetic polarization image to obtain the time... Degraded biomimetic polarization image.
[0075] Among them, time This is the timestamp of the current image frame from the bionic polarization camera.
[0076] In step S301, the time... Feature encoding is performed on the degraded biomimetic polarization image to obtain the time step. The biomimetic polarization image features.
[0077] In some embodiments of this application, the time can be determined first. Image degradation processing is performed on the biomimetic polarization image to obtain the time... The degraded biomimetic polarization image. Then, at time... Feature encoding is performed on the degraded biomimetic polarization image to obtain the time step. The biomimetic polarization image features. At this point, the obtained time... The biomimetic polarization image features can be represented as ,in Indicates time The captured biomimetic polarization image, where δ(•) is the image degradation function. It is a biomimetic polarization image feature encoder.
[0078] Next, feature fusion can be performed on the biomimetic polarization image features and the event image features. During feature fusion, the biomimetic polarization image can be used as a reference. That is, on the one hand, the image frame timestamp of the biomimetic polarization image needs to be used as a reference to align the event image with it; on the other hand, the features of the biomimetic polarization image should be the primary focus, prioritizing the preservation of physical property information such as surface material, depth, and illumination direction rich in the polarization image during the fusion process, while simultaneously enhancing target details in the event image.
[0079] In some implementations, when performing feature fusion of biomimetic polarization image features and event image features using SSM-driven Mamba, the time frame can be determined first. Biomimetic polarization image features and time Information constraints are performed on the event image features to obtain the time. The initial fusion characteristics at this moment; The initial fusion features can be used for reverse matching of time points. A biomimetic polarization image.
[0080] Among them, time Event image features by time range It is obtained by multi-scale event aggregation of the asynchronous event stream within a given time. In other words, at a given moment... When selecting event image features, you can choose , … This serves as a time window for multi-scale events, which are then aggregated to obtain the time interval. Event image features.
[0081] Time The asynchronous event stream is recorded as Then at time The event image features can be represented as ,in This is an event image feature encoder. At this point, for time... Biomimetic polarization image features and time The event image features can be constrained and associated with information as follows: ,in, It is an information constraint association and fusion module. express Constraints are imposed on the fusion information of polarization image and event image based on information from polarization image, so that the fusion features of event image and polarization image features are kept consistent.
[0082] Furthermore, moments The initial fusion characteristics can be represented as .
[0083] Then from Beginning, regarding the time Preliminary fusion characteristics and timing The event image features are iteratively correlated with information constraints until... ;in, This refers to the event segment period of the event camera. This processing method ensures the temporal consistency between the event camera and the biomimetic polarization camera. This process can be called sequence constraint.
[0084] Similarly, time Event image features by time range It is obtained by multi-scale event aggregation of the asynchronous event stream within a given time. In other words, at a given moment... When selecting event image features, you can choose , … This serves as a time window for multi-scale events, which are then aggregated to obtain the time interval. Event image features.
[0085] Finally, it was determined that each moment... The corresponding initial fusion features constitute the fusion feature sequence at time [time]. The fused image features.
[0086] This can be denoted as the frame period of the biomimetic polarization image. , and All of these can be preset. In one example, if we set... Milliseconds (ms) If ms, then when performing sequence constraints, you can first... and time range Multi-scale event aggregation is performed on the asynchronous event stream within the event to obtain event image features. Information constraints are then used for correlation to obtain the time-series data. Preliminary fusion characteristics Then on and time range Multi-scale event aggregation is performed on the asynchronous event stream within the event to obtain event image features. Information constraints are then used for correlation to obtain the time-series data. Preliminary fusion characteristics And so on, until the moment is obtained. Preliminary fusion characteristics At this point, the initial fusion features can be... , , … The fused feature sequence is used as time step The fused image features.
[0087] Figure 4 This is a flowchart illustrating the method for BEV feature encoding of radar point cloud data provided in an embodiment of this application. Figure 4 As shown, the method includes the following steps:
[0088] In step S401, local radar features are extracted using a point-based backbone network.
[0089] In step S402, global radar features are obtained using a Transformer-based backbone network.
[0090] In step S403, a cross-attention-based injection and extraction network is used to interact with the radar local features and radar global features to obtain optimized radar local features and optimized radar global features.
[0091] In step S404, the BEV pixel coordinates of the target radar point are obtained.
[0092] The target radar point can be any radar point.
[0093] In step S405, the optimized local radar features and optimized global radar features of the target radar point are scattered to the BEV pixel of the radar point and its nearby pixels to obtain the initial BEV features of the target radar point.
[0094] Among them, the neighboring pixels of each BEV pixel are pixels whose distance from this BEV pixel is less than a preset distance threshold;
[0095] In step S406, the initial BEV features of each radar point are aggregated to obtain radar BEV features.
[0096] In step S407, a Gaussian-like BEV weight map is determined based on the radar cross-section RCS value of the target radar point, and the Gaussian-like BEV weight map of all radar points is maximized to obtain the radar Gaussian-like BEV weight map.
[0097] In step S408, the radar BEV features and the radar-like Gaussian BEV weight map are connected, and the connected features are processed by a multilayer perceptron (MLP) to obtain the BEV encoded features.
[0098] In some embodiments of this application, when performing BEV feature encoding on radar point cloud data, radar local features can be extracted using a point-based backbone network and radar global features can be obtained using a Transformer-based backbone network. Then, an injection and extraction network based on cross-attention is used to interact with the radar local features and radar global features to obtain optimized radar local features and optimized radar global features.
[0099] Then, the BEV pixel coordinates of the target radar point are obtained, and the optimized radar local features and optimized radar global features of the target radar point are scattered to the BEV pixels of the radar point and its nearby pixels to obtain the initial BEV features of the target radar point.
[0100] Next, the initial BEV features of each radar point are aggregated to obtain radar BEV features, and a Gaussian-like BEV weight map is determined based on the radar cross-section RCS value of the target radar point. The Gaussian-like BEV weight map of all radar points is maximized to obtain the radar Gaussian-like BEV weight map.
[0101] Finally, the radar BEV features and the radar-like Gaussian BEV weight map are connected, and the connected features are processed by a multi-layer perceptron (MLP) to obtain the BEV encoded features.
[0102] Among them, using a point-based backbone network to extract local radar features may include: inputting radar point cloud data into an MLP to improve its feature dimension; performing max pooling on the radar point cloud data with improved feature dimension to extract global information and fuse it into high-dimensional radar features.
[0103] Interacting with radar local and global features using a cross-attention-based injection and extraction network can include: calculating pairwise distances between all radar points; generating a Gaussian-like weight map based on the pairwise distances; using the Gaussian-like weight map to adjust the weights of the target radar point relative to different spatial locations, with the Gaussian-like weight map being positively correlated with the location distance, where the location distance is the distance between the spatial location and the target radar point; adjusting the attention mechanism using the Gaussian-like weight map; and optimizing the interaction between radar local and global features based on the attention mechanism to obtain the optimized radar global features: The optimized radar local features are as follows: .
[0104] in, For the first generation of Transformer-based backbone networks Each feature block For the first point-based backbone network Each feature block, and in the injection and extraction network, For query, For keys and values; LN is the layer normalization function, and γ is the learnable scaling parameter. It uses a cross-attention mechanism, and FFN is a feedforward network.
[0105] In other words, for target detection in radar point cloud data, the point cloud data is processed directly to determine the target detection information in the radar point cloud data, and the radar point cloud data is encoded into BEV features using the Point Cloud Bird's-eye View Feature Encoding Network (LidarBEVNet).
[0106] Efficient BEV feature extraction from radar data is achieved using a point cloud bird's-eye view feature encoding network, including but not limited to LidarBEV Net. This network employs a dual-stream backbone architecture, consisting of a point-based backbone and a Transformer-based backbone. The point-based backbone learns local radar features, while the Transformer-based backbone acquires global information. For the point-based backbone, a planar structure, including but not limited to PointNet, is used. This backbone consists of multiple modules, each including an MLP and a max-pooling layer. First, the input radar point cloud features are fed into the MLP to increase their feature dimensionality. Then, max-pooling is performed on all radar point clouds to extract global information and fuse it into the high-dimensional radar features. The entire extraction process can be summarized as follows: .
[0107] The Transformer-based backbone network consists of multiple standard Transformer blocks with attention mechanisms, a feedforward network, and a normalization layer. This embodiment employs Distance Modulated Self-Attention (DMSA) for feature extraction, aiming to effectively aggregate neighboring information during the initial training iterations of the model to accelerate its convergence.
[0108] Specifically, given the coordinates of N radar points, we can first calculate the pairwise distances D∈R between all points. N×N Then, a Gaussian-like weighted graph G is generated based on the pairwise distance D. ,in It is a learnable parameter used to control the bandwidth of the Gaussian-like distribution. Essentially, the Gaussian-like weighted graph G assigns high weights to spatial locations near points and low weights to locations far from points.
[0109] The attention mechanism is adjusted using the generated weights G as follows: .
[0110] A cross-attention-based injection and extraction module is introduced to optimize the interaction between features from two different backbone radars, and this module is applied to each block of the two backbone networks. Assume that the i-th feature block of the point-based backbone network and the Transformer-based backbone network are respectively... and During the injection operation, As a query As keys and values.
[0111] Multi-head cross-attention mechanism integrates Transformer features Encoding converted into point features , can be represented as
[0112] Similarly, the extraction operation can extract point features with cross-interest from the Transformer-based backbone network. The extraction operation is defined as follows: Updated features and It is sent to the next block of the corresponding trunk.
[0113] Radar Cross-Section (RCS) is an indicator of an object's detectability by radar. Generally, larger objects produce stronger radar wave reflections, resulting in higher RCS values. Therefore, RCS can be used as a rough estimate of object size. This paper proposes a BEV encoder capable of sensing RCS to detect RCS scattering characteristics. This involves distributing the radar point's features across multiple pixels, rather than limiting it to a single pixel in the BEV space, thereby using RCS as a reference for target size.
[0114] Given a specific radar point and its RCS value 3D coordinates BEV pixel coordinates And feature f. The feature f is scattered onto pixel p and its neighboring pixels, where the pixel distance from p is less than... If a pixel in the BEV feature is covered by multiple radar features, sum-pooling is used to aggregate these features. This method yields the radar BEV feature. .
[0115] In addition, a Gaussian-like BEV weight map can be introduced for each point based on the RCS value. , where x and y are pixel coordinates.
[0116] Final Gaussian-like BEV weight map It is obtained by maximizing the weight maps of all Gaussian-like BEV classes. and Connect them and send them to the MLP to obtain the final RCS-aware BEV features. .
[0117] Figure 5 This is a flowchart illustrating a method for fusing coded features of fused image features and radar point cloud data based on Mamba, as provided in an embodiment of this application. Figure 5 As shown, the method includes the following steps:
[0118] In step S501, the fused image features and BEV coding features are projected onto the hidden state space of the SSM to obtain hidden fused image features and hidden BEV coding features.
[0119] In step S502, the fused image features and BEV coded features are projected to obtain image gating parameters and radar gating parameters, respectively.
[0120] In step S503, image gating parameters and radar gating parameters are used to modulate the hidden fused image features and the hidden BEV coded features to obtain the hidden fused image features after multi-image fusion and the hidden BEV coded features after feature interaction.
[0121] In step S504, the hidden fused image features after multi-image fusion and the hidden BEV encoded features after feature interaction are projected back into the original space and passed through residual connections to obtain complementary features.
[0122] In step S504, the multimodal fusion feature is determined based on the complementary feature by the Mamba decoder.
[0123] In some embodiments of this application, when fusing fused image features and coded features of radar point cloud data based on Mamba, the fused image features and the BEV coded features can be first projected onto the hidden state space of the SSM to obtain hidden fused image features. and hidden BEV coding features Then, the fused image features and the BEV encoded features are projected separately to obtain image gating parameters. and radar gating parameters .
[0124] Next, use the above... and stated right and Modulation is performed to obtain the hidden fused image features after multi-image fusion. Hidden BEV encoded features after interaction with features and the and Projecting back into the original space and passing it through residual connections yields complementary features. and .in, , , Represents the projection operation of a linear transformation.
[0125] Finally, the complementary features are used by the Mamba decoder. and Determine the multimodal fusion features.
[0126] In other words, the fused image features and radar point cloud BEV features can be fused using the Mamba Fusion Block, and then the final detection result can be obtained by outputting the Mamba decoder.
[0127] SSM is used to represent linear time-invariant systems by passing a one-dimensional input sequence x(t) ∈ R to an intermediate implicit state h(t) ∈ R. N To produce the output y(t) ∈ R, where N is a positive integer, it can be represented by the following linear ordinary differential equation: ; .in, It is a state vector. State vector The first derivative with respect to time, It is the input vector. It is the output vector, A∈R N×N Let B represent the state transition matrix, where B∈R N×1 , C∈R N×1 Denotes the projection parameters, D∈R N×1 This indicates a skip connection.
[0128] After discretization, the equation with step size Δ can be expressed as: ; ;
[0129] in, Δ represents the time scale parameter. I is an identity matrix.
[0130] The state-space model is achieved by using structured convolution kernels. Obtained by global convolution calculation .
[0131] The features of two modalities can be projected into the hidden state space through the Vision State Space (VSS) block, and the hidden state transition can be constructed using a gating mechanism to achieve cross-modal deep feature fusion.
[0132] In some implementations, they can first be projected into the hidden state space using ungated VSS blocks to obtain... ,in This represents the operation of projecting features into the hidden state space. It encodes features for BEV. and fusion of image features By projecting the parameters separately, the gating parameters can be obtained. and : , ;in and These represent the parameters in the two streams respectively. and Gating operation.
[0133] Using z and Gated output pairs and Modulation is performed to fuse hidden state features into , ;in, and These represent the radar point cloud BEV features after feature interaction and the hidden state features after multi-image fusion, respectively.
[0134] Then and Projecting back into the original space and passing it through residual connections yields complementary features. and : , ;in Represents the projection operation of a linear transformation.
[0135] Finally, the Mamba decoder outputs multimodal fusion features, and based on these features, a preliminary set of detected targets (including target category, coordinates, confidence level, etc.) is generated, thus obtaining the set of detected targets after multi-sensor data fusion.
[0136] Figure 6 This is a flowchart illustrating a method for determining a set of detection targets based at least on multimodal fusion features, as provided in an embodiment of this application. Figure 6 As shown, the method includes the following steps:
[0137] In step S601, a preliminary set of detection targets is determined based on multimodal fusion features.
[0138] In step S602, based on the event camera time, time synchronization processing is performed on each detection target in the preliminary detection target set, and detection targets exceeding the preset threshold are discarded.
[0139] In step S603, the cross-union ratio (CUI) calculation method based on the detection frame is used to merge the detection targets in the preliminary detection target set after time synchronization to obtain the detection target set.
[0140] In some embodiments of this application, a preliminary set of detection targets can be determined based on multimodal fusion features, and then time synchronization processing can be performed on each detection target in the preliminary set of detection targets based on the event camera time, discarding detection targets that exceed a preset threshold.
[0141] For example, if the event camera captures a reference timestamp Meanwhile, the maximum time deviation threshold Δtmax = 50ms allowed for multi-sensor data fusion was pre-calibrated. The target set obtained from the initial detection included target 1, target 2 and target 3. The three targets were collected by the bionic polarization camera and radar respectively and carried the timestamps of their respective sensors.
[0142] First, extract the sensor timestamps for each target, where the timestamp for target 1 (polarization camera) is... +30ms, Target 2 (radar) timestamp is -60ms, Target 3 (polarization camera) timestamp is +45ms.
[0143] Then, the timestamps of each target and the base time are calculated. The absolute deviation values are calculated: target 1 has a deviation of 30ms, target 2 has a deviation of 60ms, and target 3 has a deviation of 45ms. Each deviation value is then compared with a preset threshold Δtmax = 50ms. Targets 1 and 3 with deviations less than or equal to the threshold are determined to be valid synchronization targets, while target 2 with a deviation exceeding the threshold is deemed an invalid target and discarded.
[0144] Finally, the timestamps of Target 1 and Target 3 were uniformly calibrated to the base time. The complete detection information for these two targets is then stored in the dataset, completing this time synchronization and filtering process.
[0145] Then, the intersection-union ratio (IUU) method based on the detection box is used to merge the detection targets in the initial detection target set after time synchronization to obtain the detection target set.
[0146] in, Let i be the set of the i-th detection targets. , This is the initial set of detected targets after the i-th time synchronization. , These represent the target bounding boxes of the biomimetic polarization camera, the event camera, and the radar, respectively. This indicates the calculation of intersection-union ratio. This represents the intersection of the sizes of the bounding boxes of the i-th detected target. This represents the union of the sizes of the bounding boxes of the i-th detected target. This indicates the set threshold.
[0147] If the number of targets in the merged detection target set is greater than If the same target is detected multiple times, the confidence level of the repeated targets can be calculated. The target with the higher confidence level is retained. If the confidence levels are close, the target with the more stable position is selected.
[0148] The technical solution provided in this application integrates sensors such as biomimetic polarization cameras, event cameras, and radar, fully utilizing the characteristics of data from different sensors. This method not only significantly improves the accuracy of target detection but also enhances the ability to adapt to complex environments, providing strong technical support for fields such as underground space operations, autonomous driving, embodied intelligence, and mineral resource exploration.
[0149] Meanwhile, the technical solution provided in this application can also achieve high-precision perception of targets and the environment, thereby achieving excellent environmental perception capabilities and demonstrating broad application potential and significant practical value. Since this method relies almost entirely on the results of a single sensor, its robustness in detecting targets is also improved. Overall, it achieves superior environmental perception capabilities compared to processing multiple sensor channels independently.
[0150] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0151] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0152] Figure 7 This is a schematic diagram of a polarization-event-radar multimodal cooperative sensing device based on Mamba, provided in an embodiment of this application. Figure 7 As shown, the device includes:
[0153] The acquisition module 701 is configured to acquire images from a biomimetic polarization camera, images from an event camera, and radar point cloud data.
[0154] The feature extraction module 702 is configured to extract biomimetic polarization image features and event image features using a dual-branch feature extraction network; the event image features are temporal features obtained using a multi-scale event aggregation method.
[0155] The fusion module 703 is configured to perform feature fusion on biomimetic polarization image features and event image features based on the Mamba sequence deep learning architecture driven by the state-space model SSM, so as to obtain fused image features; wherein, the event image features are the state variables of the Mamba decoder layer, and the biomimetic polarization image features are updated through multiple decoder layers to obtain fused image features.
[0156] Encoding module 704 is configured to perform bird's-eye view BEV feature encoding on radar point cloud data to obtain BEV encoded features of radar point cloud data.
[0157] The fusion module 703 is also configured to fuse the coded features of the fused image and the radar point cloud data based on Mamba to obtain multimodal fused features.
[0158] The perception module 705 is configured to determine the set of detection targets based at least on multimodal fusion features.
[0159] According to the technical solution provided in the embodiments of this application, a dual-branch feature extraction network is used to extract biomimetic polarization image features and event image features respectively. Mamba driven by SSM is used to fuse the biomimetic polarization image features and event image features to obtain fused image features. BEV feature encoding of radar point cloud data is also obtained. Mamba is then used to fuse the fused image features and the encoded features of the radar point cloud data to obtain multimodal fused features, thereby determining the set of detection targets. This fully utilizes the data characteristics of different sensors to improve the accuracy of target detection and enhances the ability to adapt to complex environments. It can provide strong technical support for fields such as autonomous driving, embodied intelligence, robotics, underground space operations, and mineral resource exploration, and can achieve high-precision perception of targets and the environment, thus achieving excellent environmental perception capabilities, demonstrating broad application potential and significant practical value.
[0160] In some implementations, event image features are obtained using a multi-scale event aggregation method, including: generating an asynchronous event stream by measuring pixel brightness changes in event camera images; and obtaining time points from the asynchronous event stream. The event corresponds to n scale events; where different scale events correspond to events in the asynchronous event stream within different time ranges, the largest scale event corresponds to the largest time range of this aggregation, and the remaining time ranges are... ; The exposure time for a single frame of an image captured by a bionic polarization camera. It is a positive integer less than n. To define the aggregation time range, n is a positive integer greater than or equal to 3. The n scale events are projected onto the feature space, and the projected event features are pooled. The pooling size for small-scale events is smaller than that for large-scale events, and the time range corresponding to small-scale events is smaller than that corresponding to large-scale events. The pooled event features of all scale events except the largest scale event are upsampled, and the upsampled event features are concatenated and convolved with the pooled event features of the largest scale event to obtain the aggregation time range. Inner Time Event image features.
[0161] In some implementations, biomimetic polarization image features are extracted in the following manner: for time... Image degradation processing is performed on the biomimetic polarization image to obtain the time... Degraded biomimetic polarization image; for time Feature encoding is performed on the degraded biomimetic polarization image to obtain the time step. The biomimetic polarization image features.
[0162] In some implementations, Mamba driven by SSM performs feature fusion on biomimetic polarization image features and event image features, including: time... Biomimetic polarization image features and time Information constraints are performed on the event image features to obtain the time. The initial fusion characteristics; among which, time Event image features by time range The asynchronous event stream within is obtained by multi-scale event aggregation; time. The initial fusion features can be used for reverse matching of time points. Bionic polarization images; self Beginning, regarding the time Preliminary fusion characteristics and timing The event image features are iteratively correlated with information constraints until... ;in, The event segment period of the event camera, time. Event image features by time range The asynchronous event stream within is obtained by multi-scale event aggregation; the determination is based on the events at each time point. The corresponding initial fusion features constitute the fusion feature sequence at time [time]. The fused image features.
[0163] In some implementations, BEV feature encoding is performed on radar point cloud data, including: extracting local radar features using a point-based backbone network; obtaining global radar features using a Transformer-based backbone network; interacting the local and global radar features using a cross-attention-based injection and extraction network to obtain optimized local and global radar features; obtaining the BEV pixel coordinates of a target radar point; the target radar point is any radar point; scattering the optimized local and global radar features of the target radar point to the BEV pixels of the radar point and its neighboring pixels to obtain the initial BEV features of the target radar point; wherein, the neighboring pixels of each BEV pixel are pixels whose distance from the BEV pixel is less than a preset distance threshold; aggregating the initial BEV features of each radar point to obtain radar BEV features; determining a Gaussian-like BEV weight map based on the RCS value of the radar cross section of the target radar point, maximizing the Gaussian-like BEV weight map of all radar points to obtain a radar Gaussian-like BEV weight map; concatenating the radar BEV features and the radar Gaussian-like BEV weight map, and performing multilayer perceptron (MLP) processing on the concatenated features to obtain BEV encoded features.
[0164] In some implementations, a point-based backbone network is used to extract local radar features, including: inputting radar point cloud data into an MLP to increase its feature dimension; performing max pooling on the radar point cloud data with increased feature dimension to extract global information and fuse it into high-dimensional radar features.
[0165] In some implementations, a cross-attention-based injection and extraction network is used to interact with local and global radar features, including: calculating pairwise distances between all radar points; generating a Gaussian-like weight map based on the pairwise distances; using the Gaussian-like weight map to adjust the weights of the target radar point relative to different spatial locations, with the Gaussian-like weight map being positively correlated with the location distance, where the location distance is the distance between the spatial location and the target radar point; adjusting the attention mechanism using the Gaussian-like weight map; and optimizing the interaction between local and global radar features based on the attention mechanism to obtain the optimized global radar features: The optimized radar local features are as follows: ;in, For the first generation of Transformer-based backbone networks Each feature block For the first point-based backbone network Each feature block, and in the injection and extraction network, For query, For keys and values; LN is the layer normalization function, and γ is the learnable scaling parameter. It uses a cross-attention mechanism, and FFN is a feedforward network.
[0166] In some implementations, the fused image features and the encoded features of radar point cloud data are fused based on Mamba, including: projecting the fused image features and BEV encoded features onto the hidden state space of the SSM respectively to obtain hidden fused image features. and hidden BEV coding features The image gating parameters are obtained by projecting the fused image features and BEV-coded features separately. and radar gating parameters ;use and right and Modulation is performed to obtain the hidden fused image features after multi-image fusion. Hidden BEV encoded features after interaction with features ;Will and Projecting back into the original space and passing it through residual connections yields complementary features. and ;in, , , Represents the projection operation of a linear transformation; based on complementary features using the Mamba decoder. and Determine the multimodal fusion features.
[0167] In some implementations, before acquiring images from the bionic polarization camera, event camera, and radar point cloud data, the method further includes: calibrating each sensor separately, acquiring the parameters of each sensor, and the rotation matrix and translation vector between different sensors; wherein the sensors include a bionic polarization camera, an event camera, and radar; assigning timestamps to the measurement data of each sensor, and achieving data synchronization of each sensor through timestamp interpolation and clock drift compensation algorithms.
[0168] In some implementations, the set of detection targets is determined at least based on multimodal fusion features, including: determining a preliminary set of detection targets based on multimodal fusion features; performing time synchronization processing on each detection target in the preliminary set of detection targets using the event camera time as a reference, discarding detection targets exceeding a preset threshold; merging each detection target in the time-synchronized preliminary set of detection targets using an intersection-union ratio (IUU) calculation method based on detection boxes to obtain a final set of detection targets; wherein, Let i be the set of the i-th detection targets. , This is the initial set of detected targets after the i-th time synchronization. , These represent the target bounding boxes of the biomimetic polarization camera, the event camera, and the radar, respectively. This indicates the calculation of intersection-union ratio. This represents the intersection of the sizes of the bounding boxes of the i-th detected target. This represents the union of the sizes of the bounding boxes of the i-th detected target. This indicates the set threshold.
[0169] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0170] Figure 8 This is a schematic diagram of the electronic device provided in an embodiment of this application. For example... Figure 8 As shown, the electronic device 8 of this embodiment includes a processor 801, a memory 802, and a computer program 803 stored in the memory 802 and executable on the processor 801. When the processor 801 executes the computer program 803, it implements the steps in the various method embodiments described above. Alternatively, when the processor 801 executes the computer program 803, it implements the functions of each module / unit in the various device embodiments described above.
[0171] Electronic device 8 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 8 may include, but is not limited to, processor 801 and memory 802. Those skilled in the art will understand that... Figure 8 This is merely an example of electronic device 8 and does not constitute a limitation on electronic device 8. It may include more or fewer components than shown, or different components.
[0172] The processor 801 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0173] The memory 802 can be an internal storage unit of the electronic device 8, such as a hard disk or RAM of the electronic device 8. The memory 802 can also be an external storage device of the electronic device 8, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 8. The memory 802 can also include both internal and external storage units of the electronic device 8. The memory 802 is used to store computer programs and other programs and data required by the electronic device.
[0174] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0175] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0176] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A polarization-event-radar multimodal cooperative sensing method based on Mamba, characterized in that, include: Acquire images from a biomimetic polarization camera, event camera images, and radar point cloud data; A dual-branch feature extraction network is used to extract biomimetic polarization image features and event image features, respectively; the event image features are temporal features obtained using a multi-scale event aggregation method. The Mamba sequence deep learning architecture, driven by the State-Space Model (SSM), fuses the biomimetic polarization image features and the event image features to obtain fused image features. The event image features are the state variables of the Mamba decoder layer, and the biomimetic polarization image features are updated through multiple decoder layers to obtain the fused image features. Bird’s-eye view BEV feature encoding is performed on the radar point cloud data to obtain the BEV encoded features of the radar point cloud data. Based on the Mamba, the fused image features and the encoded features of the radar point cloud data are fused to obtain multimodal fused features; The set of detection targets is determined at least based on the multimodal fusion features; Among them, the event image features are obtained using a multi-scale event aggregation method, including: An asynchronous event stream is generated by measuring the pixel brightness changes in the event camera image; Obtaining time from the asynchronous event stream The event corresponds to n scale events; where different scale events correspond to events in the asynchronous event stream within different time ranges, the largest scale event corresponds to the largest time range of this aggregation, and the remaining time ranges are... ; The exposure time for a single frame of an image captured by a bionic polarization camera. It is a positive integer less than n. The aggregation time range is defined, where n is a positive integer greater than or equal to 3. The n scale events are projected into the feature space respectively, and the projected event features are pooled respectively; wherein, the pooling size for pooling small-scale events is smaller than the pooling size for pooling large-scale events, and the time range corresponding to the small-scale events is smaller than the time range corresponding to the large-scale events. The event features after pooling for events of all scales except the largest scale event are upsampled. Then, the upsampled event features are concatenated and convolved with the event features after pooling for the largest scale event to obtain the event features within the time range. Inner Time Event image features.
2. The method according to claim 1, characterized in that, The biomimetic polarization image features are extracted using the following method: For time Image degradation processing is performed on the biomimetic polarization image to obtain the time... Degraded biomimetic polarization images; where, at time This is the timestamp of the current image frame from the biomimetic polarization camera. For time Feature encoding is performed on the degraded biomimetic polarization image to obtain the time step. The biomimetic polarization image features.
3. The method according to claim 2, characterized in that, Based on SSM-driven Mamba, feature fusion is performed on the biomimetic polarization image features and the event image features, including: For time Biomimetic polarization image features and time Information constraints are performed on the event image features to obtain the time. The initial fusion characteristics; among which, time Event image features by time range The asynchronous event stream within is obtained by multi-scale event aggregation; the time... The initial fusion features can be used to reverse match the time. Bionic polarization images; since Beginning, for the time Preliminary fusion characteristics and timing The event image features are iteratively correlated with information constraints until... ;in, The event segment period of the event camera, time. Event image features by time range The asynchronous event stream within is obtained by multi-scale event aggregation. Determined by each time point The corresponding initial fusion features constitute the fusion feature sequence at time [time]. The fused image features.
4. The method according to claim 1, characterized in that, BEV feature encoding is performed on the radar point cloud data, including: Local radar features are extracted using a point-based backbone network. Global radar features are obtained using a Transformer-based backbone network. The radar local features and radar global features are interacted using a cross-attention-based injection and extraction network to obtain optimized radar local features and optimized radar global features. Obtain the BEV pixel coordinates of the target radar point; the target radar point can be any radar point. The optimized local radar features and optimized global radar features of the target radar point are scattered to the BEV pixel of the radar point and its neighboring pixels to obtain the initial BEV features of the target radar point; wherein, the neighboring pixels of each BEV pixel are pixels whose distance from the BEV pixel is less than a preset distance threshold. The initial BEV features of each radar point are aggregated to obtain the radar BEV features; The Gaussian-like BEV weight map is determined based on the RCS value of the target radar point. The Gaussian-like BEV weight map of all radar points is maximized to obtain the radar Gaussian-like BEV weight map. The radar BEV features and the radar-like Gaussian BEV weight map are connected, and the connected features are processed by a multilayer perceptron (MLP) to obtain the BEV encoded features.
5. The method according to claim 4, characterized in that, Local radar features are extracted using a point-based backbone network, including: The radar point cloud data is input into the MLP to enhance its feature dimensions; Max pooling is performed on the radar point cloud data after feature dimension enhancement to extract global information and fuse it into high-dimensional radar features.
6. The method according to claim 4, characterized in that, The radar local features and the radar global features are interacted using a cross-attention-based injection and extraction network, including: Calculate the pairwise distances between all radar points; A Gaussian-like weighted map is generated based on the pairwise distances; the Gaussian-like weighted map is used to adjust the weights of the target radar point relative to different spatial locations, and the Gaussian-like weighted map is positively correlated with the location distance, where the location distance is the distance between the spatial location and the target radar point; The attention mechanism is adjusted using the Gaussian-like weight graph. Based on the attention mechanism, the interaction between the radar local features and the radar global features is optimized to obtain the optimized radar global features as follows: ; The optimized radar local features are as follows: ; in, For the first generation of Transformer-based backbone networks Each feature block For the first point-based backbone network Each feature block, and in the injection and extraction network, For query, For keys and values; LN is the layer normalization function, and γ is the learnable scaling parameter. It uses a cross-attention mechanism, and FFN is a feedforward network.
7. The method according to claim 1, characterized in that, The fusion of the fused image features and the encoded features of the radar point cloud data based on the Mamba is performed, including: The fused image features and the BEV encoded features are projected onto the hidden state space of the SSM to obtain the hidden fused image features. and hidden BEV coding features ; The fused image features and the BEV-coded features are projected separately to obtain image gating parameters. and radar gating parameters ; Using the above and stated right and Modulation is performed to obtain the hidden fused image features after multi-image fusion. Hidden BEV encoded features after interaction with features ; Will and Projecting back into the original space and passing it through residual connections yields complementary features. and ;in, , , Represents the projection operation of a linear transformation; Based on the complementary features via the Mamba decoder and Determine the multimodal fusion features.
8. The method according to claim 1, characterized in that, Before acquiring images from the biomimetic polarization camera, event camera images, and radar point cloud data, the method further includes: Each sensor is calibrated separately to obtain the parameters of each sensor, as well as the rotation matrix and translation vector between different sensors; wherein, the sensors include a bionic polarization camera, an event camera, and radar; Timestamps are assigned to the measurement data of each sensor, and data synchronization of each sensor is achieved through timestamp interpolation and clock drift compensation algorithms.
9. The method according to claim 1, characterized in that, Determining the set of detection targets based at least on the multimodal fusion features includes: A preliminary set of detection targets is determined based on the aforementioned multimodal fusion features; Based on the event camera time, time synchronization processing is performed on each detection target in the preliminary detection target set, and detection targets exceeding the preset threshold are discarded; The cross-union ratio (CUNR) method based on detection boxes is used to merge the detection targets in the initial detection target set after time synchronization to obtain the detection target set. in, Let i be the set of the i-th detection targets. , This is the initial set of detected targets after the i-th time synchronization. , These represent the target bounding boxes of the biomimetic polarization camera, the event camera, and the radar, respectively. This indicates the calculation of intersection-union ratio. This represents the intersection of the sizes of the bounding boxes of the i-th detected target. This represents the union of the sizes of the bounding boxes of the i-th detected target. This indicates the set threshold.
10. A polarization-event-radar multimodal cooperative sensing device based on Mamba, characterized in that, include: The acquisition module is configured to acquire images from a biomimetic polarization camera, images from an event camera, and radar point cloud data. The feature extraction module is configured to extract biomimetic polarization image features and event image features using a dual-branch feature extraction network; the event image features are temporal features obtained using a multi-scale event aggregation method. The fusion module is configured to perform feature fusion on the biomimetic polarization image features and the event image features based on the Mamba sequence deep learning architecture driven by the state-space model (SSM) to obtain fused image features; wherein, the event image features are the state variables of the Mamba decoder layer, and the biomimetic polarization image features are updated through multiple decoder layers to obtain the fused image features; The encoding module is configured to perform bird's-eye view BEV feature encoding on the radar point cloud data to obtain the BEV encoded features of the radar point cloud data. The fusion module is also configured to fuse the fused image features and the encoded features of the radar point cloud data based on the Mamba to obtain multimodal fusion features; The perception module is configured to determine a set of detection targets based at least on the multimodal fusion features; Among them, the event image features are obtained using a multi-scale event aggregation method, including: An asynchronous event stream is generated by measuring the pixel brightness changes in the event camera image; Obtaining time from the asynchronous event stream The event corresponds to n scale events; where different scale events correspond to events in the asynchronous event stream within different time ranges, the largest scale event corresponds to the largest time range of this aggregation, and the remaining time ranges are... ; The exposure time for a single frame of an image captured by a bionic polarization camera. It is a positive integer less than n. The aggregation time range is defined, where n is a positive integer greater than or equal to 3. The n scale events are projected into the feature space respectively, and the projected event features are pooled respectively; wherein, the pooling size for pooling small-scale events is smaller than the pooling size for pooling large-scale events, and the time range corresponding to the small-scale events is smaller than the time range corresponding to the large-scale events. The event features after pooling for events of all scales except the largest scale event are upsampled. Then, the upsampled event features are concatenated and convolved with the event features after pooling for the largest scale event to obtain the event features within the time range. Inner Time Event image features.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 9.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Radar video intelligent fusion and early warning method and system
CN112562405A
Radar / camera fusion 3D target detection method based on bidirectional cross-modal attention
CN119992065A