一种基于多模态信息融合的3D开放场景目标检测方法
By fusing multimodal information from LiDAR and visible light cameras, 3D point cloud and visual image features are extracted and fused, solving the problem of detecting unknown targets in open scenes. This achieves high-accuracy and highly adaptable target detection, suitable for autonomous driving and robot navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA ACAD OF LAUNCH VEHICLE TECH
- Filing Date
- 2025-11-14
- Publication Date
- 2026-07-17
AI Technical Summary
Existing 3D target detection methods cannot effectively identify unknown targets in open scenes, and single visual information and multimodal fusion methods are insufficient in target scale estimation, localization accuracy and occlusion handling, which cannot meet the detection needs of complex environments.
Data is collected simultaneously by LiDAR and visible light cameras. After preprocessing, point cloud and visual image features are extracted. Combined with a multimodal feature fusion network, an improved YOLOv7 and BLIP pre-trained model is used for feature extraction and fusion. Target detection is performed using an attention mechanism and an open set learning strategy.
It achieves accurate detection and localization of multiple categories and unknown targets in open scenes, improves detection accuracy and environmental adaptability, can identify unseen categories and unknown targets, reduces the impact of lighting changes and occlusion, and is suitable for autonomous driving and robot navigation.
Smart Images

Figure CN121661446B_ABST