多模态AI人机协同标注研究方法及系统

By using multimodal association weight matrix and semantic inertia modeling, combined with cross-modal confidence fusion and topological temporal correction, the boundary drift problem of multimodal annotation systems in dynamic scenes is solved, achieving highly accurate semantic recognition and annotation.

CN122065027BActive Publication Date: 2026-07-17GUIZHOU CARTHAGE INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU CARTHAGE INFORMATION TECH CO LTD
Filing Date
2026-01-16
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multimodal AI annotation systems often fail to detect boundaries accurately in dynamic scenes, leading to errors in identifying object boundaries that dynamically drift or overlap. This results in semantic misalignment of labels across different time periods, affecting the accuracy and consistency of the annotation results.

Method used

By performing time synchronization processing through a multimodal correlation weight matrix, generating a boundary dynamic response map using a temporal gradient difference model, constructing a semantic inertia tensor model, and performing cross-modal confidence fusion and graph convolutional temporal network prediction, the system can automatically detect and dynamically correct potential semantic misalignment regions.

Benefits of technology

In complex lighting and occlusion environments, it accurately captures the dynamic contours and semantic change trends of targets, eliminates semantic drift between consecutive frames, improves the accuracy of semantic recognition and annotation in dynamic scenes, and ensures the continuity and stability of label information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065027B_ABST
    Figure CN122065027B_ABST
Patent Text Reader

Abstract

本发明公开了多模态AI人机协同标注研究方法及系统,涉及人机协同技术领域,包括以下步骤:获取动态场景下的多模态原始数据流,对视觉模态、语音模态与点云模态进行时间同步处理,并基于各模态特征分布的相关性构建多模态关联权重矩阵;基于多模态关联权重矩阵,利用时序梯度差分模型提取连续帧间的边界变化特征,生成边界动态响应图;依据边界动态响应图,构建语义惯性张量模型。本发明通过多模态关联权重与语义惯性建模实现多源信息对齐与时序连续,精确捕获目标动态变化,消除语义漂移并提高标注准确性;同时通过跨模态置信融合与拓扑时序校正自动修正语义错位,确保语义一致性与稳定性,提升标注的自适应性与连贯性。
Need to check novelty before this filing date? Find Prior Art