多模态AI人机协同标注研究方法及系统
By using multimodal association weight matrix and semantic inertia modeling, combined with cross-modal confidence fusion and topological temporal correction, the boundary drift problem of multimodal annotation systems in dynamic scenes is solved, achieving highly accurate semantic recognition and annotation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU CARTHAGE INFORMATION TECH CO LTD
- Filing Date
- 2026-01-16
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multimodal AI annotation systems often fail to detect boundaries accurately in dynamic scenes, leading to errors in identifying object boundaries that dynamically drift or overlap. This results in semantic misalignment of labels across different time periods, affecting the accuracy and consistency of the annotation results.
By performing time synchronization processing through a multimodal correlation weight matrix, generating a boundary dynamic response map using a temporal gradient difference model, constructing a semantic inertia tensor model, and performing cross-modal confidence fusion and graph convolutional temporal network prediction, the system can automatically detect and dynamically correct potential semantic misalignment regions.
In complex lighting and occlusion environments, it accurately captures the dynamic contours and semantic change trends of targets, eliminates semantic drift between consecutive frames, improves the accuracy of semantic recognition and annotation in dynamic scenes, and ensures the continuity and stability of label information.
Smart Images

Figure CN122065027B_ABST