A multi-modal three-dimensional target detection and task-driven reconstruction method and system based on modal independence
By constructing a shared geometric hub based on point cloud representation, unified feature extraction and consistent processing of RGB images, depth maps, and point cloud modalities are achieved, solving the modality dependency and detection instability problems of multimodal 3D perception methods and improving the robustness and applicability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 赵嫣然
- Filing Date
- 2026-06-01
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multimodal 3D perception methods are highly dependent on input modalities and sensor configurations. The geometric representations between modalities are not uniform, and the performance degrades significantly when a modality is missing. It is difficult to achieve cross-scene reuse. Furthermore, existing methods lack stable geometric central representations, resulting in inconsistent detection results and poor robustness.
A unified geometric processing hub is constructed, with point cloud representation as the core. RGB images, depth maps, and point cloud modalities are uniformly incorporated into a shared geometric feature space for consistent processing and fusion, forming a stable multimodal geometric feature representation. Encoder parameters are optimized in the shared space to achieve modality-independent perception.
It improves the system's generalization ability and robustness in complex deployment environments, reduces the dependence on fixed modal combinations and sensor configurations, and ensures that stable 3D detection results can still be output under single modal input or modal combination changes, thereby enhancing the applicability and flexibility of the multimodal sensing system.
Smart Images

Figure CN122415894A_ABST