一种基于多模态大模型推理的视觉问答与目标定位方法
By employing a visual question answering and target localization method based on multimodal large model reasoning, the semantic alignment problem in multimodal fusion is solved, achieving high-precision target localization and verifiable pixel-level evidence, thereby improving the accuracy and robustness of scene understanding and target localization in autonomous driving systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF SCI & TECH
- Filing Date
- 2026-06-18
- Publication Date
- 2026-07-17
AI Technical Summary
Multimodal large language models and segmentation models suffer from semantic alignment issues in scene understanding and target localization of images, leading to a disconnect between inference conclusions and segmentation results, which affects the accuracy and robustness of scene understanding and target localization in autonomous driving.
A visual question answering and target localization method based on multimodal large model reasoning is adopted. By combining a multimodal large language model, an anchored inference pooling module, an image encoder, a fusion module and a mask decoder, end-to-end joint training is performed to achieve deep fusion and semantic alignment of inference semantics and visual features, and to generate pixel-level target masks.
It improves the accuracy and interpretability of target localization, ensures that the localization results are strictly guided by semantic reasoning, enhances the decision-making transparency and reliability of autonomous driving systems, and improves the segmentation robustness under complex road conditions.
Smart Images

Figure CN122416461A_ABST