Reference multi-target tracking method based on spatial text-visual interaction
By using a spatial cross-modal feature alignment method, fine-grained semantic information in text descriptions is explicitly split and aligned, which solves the problem of insufficient utilization of cross-modal information in existing methods and improves the target selection and tracking accuracy of multi-target tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF ELECTRONICS SCI & TECH OF CHINA
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-21
AI Technical Summary
Existing multi-target tracking methods do not fully utilize cross-modal information, especially the fine-grained spatial positioning information in text, which limits the target selection capability and tracking accuracy in complex semantic scenarios.
A method based on spatial cross-modal feature alignment is adopted. The text decoupling module explicitly splits the target object, spatial location and appearance color into three semantic components. The fine-grained semantic feature alignment module performs multi-level alignment at the pixel level, attribute level and channel level to improve the matching accuracy between visual features and text semantics.
It significantly improves the target selection capability and tracking accuracy in complex semantic scenarios, achieves accurate matching and continuous tracking of text descriptions, and has good versatility and scalability.
Smart Images

Figure CN122434973A_ABST