A multi-modal three-dimensional target detection and task-driven reconstruction method and system based on modal independence

By constructing a shared geometric hub based on point cloud representation, unified feature extraction and consistent processing of RGB images, depth maps, and point cloud modalities are achieved, solving the modality dependency and detection instability problems of multimodal 3D perception methods and improving the robustness and applicability of the system.

CN122415894APending Publication Date: 2026-07-17赵嫣然

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
赵嫣然
Filing Date
2026-06-01
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multimodal 3D perception methods are highly dependent on input modalities and sensor configurations. The geometric representations between modalities are not uniform, and the performance degrades significantly when a modality is missing. It is difficult to achieve cross-scene reuse. Furthermore, existing methods lack stable geometric central representations, resulting in inconsistent detection results and poor robustness.

Method used

A unified geometric processing hub is constructed, with point cloud representation as the core. RGB images, depth maps, and point cloud modalities are uniformly incorporated into a shared geometric feature space for consistent processing and fusion, forming a stable multimodal geometric feature representation. Encoder parameters are optimized in the shared space to achieve modality-independent perception.

Benefits of technology

It improves the system's generalization ability and robustness in complex deployment environments, reduces the dependence on fixed modal combinations and sensor configurations, and ensures that stable 3D detection results can still be output under single modal input or modal combination changes, thereby enhancing the applicability and flexibility of the multimodal sensing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415894A_ABST
    Figure CN122415894A_ABST
Patent Text Reader

Abstract

一种基于模态独立的多模态三维目标检测与任务驱动重建方法及系统,涉及多模态感知、三维目标检测和自动驾驶技术领域。技术要点:所述方法包括特征重建和模态独立感知两个部分:在特征重建部分,构建点云、深度图、RGB图像等多模态的重建链路,形成以点云为核心的几何中枢,并保存经评估较优的网络参数;在模态独立检测部分,将RGB图像、深度图和点云作为并行输入,提取几何特征并进行一致性约束学习,最终输出三维检测框进行目标检测。本发明通过建立模态与统一几何表示的映射关系,使模型在不同模态输入或模态缺失变化下,依然能保持稳定的三维检测性能,减少对不同传感器配置下专用模型的重复训练需求,提高系统的泛化能力、鲁棒性及部署灵活性。该方法适用于多模态感知、自动驾驶、智能巡检及复杂环境下的三维目标检测任务。
Need to check novelty before this filing date? Find Prior Art