一种基于多模态大模型推理的视觉问答与目标定位方法

By employing a visual question answering and target localization method based on multimodal large model reasoning, the semantic alignment problem in multimodal fusion is solved, achieving high-precision target localization and verifiable pixel-level evidence, thereby improving the accuracy and robustness of scene understanding and target localization in autonomous driving systems.

CN122416461APending Publication Date: 2026-07-17NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2026-06-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Multimodal large language models and segmentation models suffer from semantic alignment issues in scene understanding and target localization of images, leading to a disconnect between inference conclusions and segmentation results, which affects the accuracy and robustness of scene understanding and target localization in autonomous driving.

Method used

A visual question answering and target localization method based on multimodal large model reasoning is adopted. By combining a multimodal large language model, an anchored inference pooling module, an image encoder, a fusion module and a mask decoder, end-to-end joint training is performed to achieve deep fusion and semantic alignment of inference semantics and visual features, and to generate pixel-level target masks.

Benefits of technology

It improves the accuracy and interpretability of target localization, ensures that the localization results are strictly guided by semantic reasoning, enhances the decision-making transparency and reliability of autonomous driving systems, and improves the segmentation robustness under complex road conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416461A_ABST
    Figure CN122416461A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于多模态大模型推理的视觉问答与目标定位方法,属于图像识别的技术领域;包括以下步骤:步骤S1:获取图像信息和文字信息;步骤S2:设置问答处理模型,处理图像信息和文字信息,步骤S3:对应生产掩码,对掩码进行解码形成掩码图像,结合文字形成问答;步骤S4:对问答处理模型进行训练,直到问答处理模型能对图像和问题生成符合场景实际的推理答案。本发明采用上述方法,将图片和语言进行融合处理,降低因推理与分割脱节、语义对齐不足导致的定位误差,为场景推理结论提供可验证的像素级视觉证据,通过联合优化实现场景推理与目标定位的语义对齐,形成完整的“推理‑融合‑定位‑优化”闭环。
Need to check novelty before this filing date? Find Prior Art