基于深度信息场景图生成的多尺度目标检测方法

By constructing a multimodal scene graph generation network that integrates deep information, the problem of inaccurate predicate prediction in target detection under complex scenes is solved, and more accurate multi-scale target detection is achieved, especially the prediction of the relationship between subject and object targets.

CN118470301BActive Publication Date: 2026-07-17XIDIAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2024-05-16
Publication Date
2026-07-17

Smart Images

  • Figure CN118470301B_ABST
    Figure CN118470301B_ABST
Patent Text Reader

Abstract

基于深度信息场景图生成的多尺度目标检测方法,包括以下步骤;步骤1,获取Visual Genome数据集;步骤2,构建融合深度信息的多模态场景图生成网络S,将深度信息引入场景图的生成过程,以避免视觉错位效应对谓词预测的干扰;步骤3,对融合深度信息的多模态场景图生成网络S进行场景图生成实验,得到该网络在Visual Genome数据集下的场景图生成任务性能指标。本发明使得目标间依赖关系更好的辅助多尺度目标检测任务,用于解决复杂场景下对于主题目标和客体目标的谓词预测不准确的问题。
Need to check novelty before this filing date? Find Prior Art