A high-level semantic guided low-level controlled alignment multi-modal three-dimensional target detection method, system and vehicle

By performing multi-scale feature fusion and state space model interaction in bird's-eye view space, the problem of insufficient fusion in multimodal 3D target detection is solved, achieving efficient detection in complex traffic scenarios and adverse weather conditions, and improving detection robustness and accuracy.

CN122415989APending Publication Date: 2026-07-17JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU UNIV
Filing Date
2026-04-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies in multimodal 3D target detection suffer from problems such as early fusion leading to deep ambiguity and loss of geometric accuracy, insufficient intermodal interaction in late fusion, and difficulty in simultaneously ensuring high-level global semantic consistency and low-level multi-scale fine-grained alignment in deep fusion. These issues result in insufficient robustness and accuracy in complex traffic scenarios and adverse weather conditions.

Method used

The method involves extracting multi-scale features from camera images and LiDAR point clouds and projecting them onto a unified Bird's-Eye View (BEV) space. A high-level cross-modal interaction is performed through a hybrid Mamba fusion module to generate high-level fused BEV features with semantic consensus. A prediction module is used to generate multi-scale offset fields and fusion weight maps. The method is then combined with the Cross-Mamba structure of the state space model to perform cross-modal information exchange and controlled alignment.

Benefits of technology

It achieves robustness and accuracy improvement in 3D target detection under complex traffic scenarios and severe weather conditions. By guiding low-level controlled alignment and dynamic fusion with high-level semantics, it reduces computational complexity and improves global semantic consistency and local geometric stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415989A_ABST
    Figure CN122415989A_ABST
Patent Text Reader

Abstract

The application discloses a kind of high-level semantic guide low-level controlled alignment multi-modal three-dimensional target detection method, system and vehicle, introduce Cross Mamba cross-modal global fusion mechanism on high-level BEV feature, realize continuous stable cross-modal semantic propagation in BEV grid by sharing hidden state.The scene-level consistent semantics of high-level fusion feature is used as a constraint condition, and small-range controllable spatial correction is performed on low-level multi-scale features, and neighborhood constraint is applied to the offset amplitude, and the laser radar BEV feature is used as a geometric anchor to keep the sampling position unchanged or only perform limited fine-tuning.After controlled alignment, scale guide features and aligned modal features are used to predict dynamic fusion weights, so that the image modality occupies a higher weight in the semantic reliable area, and the laser radar modality occupies a higher weight in the geometric stable area;At the same time, cross-scale state space modeling is performed on multi-scale fusion features, so that global information and local information between different scales can be bidirectionally propagated and cooperatively integrated.
Need to check novelty before this filing date? Find Prior Art