一种基于多模态信息融合的3D开放场景目标检测方法

By fusing multimodal information from LiDAR and visible light cameras, 3D point cloud and visual image features are extracted and fused, solving the problem of detecting unknown targets in open scenes. This achieves high-accuracy and highly adaptable target detection, suitable for autonomous driving and robot navigation.

CN121661446BActive Publication Date: 2026-07-17CHINA ACAD OF LAUNCH VEHICLE TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA ACAD OF LAUNCH VEHICLE TECH
Filing Date
2025-11-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing 3D target detection methods cannot effectively identify unknown targets in open scenes, and single visual information and multimodal fusion methods are insufficient in target scale estimation, localization accuracy and occlusion handling, which cannot meet the detection needs of complex environments.

Method used

Data is collected simultaneously by LiDAR and visible light cameras. After preprocessing, point cloud and visual image features are extracted. Combined with a multimodal feature fusion network, an improved YOLOv7 and BLIP pre-trained model is used for feature extraction and fusion. Target detection is performed using an attention mechanism and an open set learning strategy.

Benefits of technology

It achieves accurate detection and localization of multiple categories and unknown targets in open scenes, improves detection accuracy and environmental adaptability, can identify unseen categories and unknown targets, reduces the impact of lighting changes and occlusion, and is suitable for autonomous driving and robot navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661446B_ABST
    Figure CN121661446B_ABST
Patent Text Reader

Abstract

本发明一种基于多模态信息融合的3D开放场景目标检测方法,利用激光雷达采集的3D点云数据、可见光相机采集的视觉图像数据、输入或基于视觉图像生成的包含待检测目标语义信息的文本指令,分别进行点云数据、图像数据和语义数据的特征提取与特征融合,利用融合后的特征实现3D开放场景目标检测。本发明通过对激光雷达点云和视觉图像的深度特征提取、精准对齐及融合,结合开放集目标检测策略,解决了因单一视觉信息局限性和多模态融合场景适用性差导致的目标检测效果不佳等问题,实现对开放场景中多类别、未知目标的准确检测和定位,显著提升了3D目标检测在开放环境中的准确性和适应性。
Need to check novelty before this filing date? Find Prior Art