Reference multi-target tracking method based on spatial text-visual interaction

By using a spatial cross-modal feature alignment method, fine-grained semantic information in text descriptions is explicitly split and aligned, which solves the problem of insufficient utilization of cross-modal information in existing methods and improves the target selection and tracking accuracy of multi-target tracking.

CN122434973APending Publication Date: 2026-07-21UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2026-04-08
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing multi-target tracking methods do not fully utilize cross-modal information, especially the fine-grained spatial positioning information in text, which limits the target selection capability and tracking accuracy in complex semantic scenarios.

Method used

A method based on spatial cross-modal feature alignment is adopted. The text decoupling module explicitly splits the target object, spatial location and appearance color into three semantic components. The fine-grained semantic feature alignment module performs multi-level alignment at the pixel level, attribute level and channel level to improve the matching accuracy between visual features and text semantics.

Benefits of technology

It significantly improves the target selection capability and tracking accuracy in complex semantic scenarios, achieves accurate matching and continuous tracking of text descriptions, and has good versatility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434973A_ABST
    Figure CN122434973A_ABST
Patent Text Reader

Abstract

The application discloses a kind of referring multi-target tracking methods based on space text-visual interaction, preliminary interaction of visual feature and text feature is realized by bidirectional cross-modal feature fusion encoder;Second, text description is explicitly split into three kinds of fine-grained semantic components of target object, spatial position and appearance color by text decoupling module, and the corresponding semantic mask is generated;Then, fine-grained semantic feature alignment module is sequentially executed pixel-level feature alignment, attribute feature alignment and channel modulation, wherein attribute feature alignment sets independent branch for target object, spatial position and appearance color respectively to carry out targeted alignment;Finally, through timing enhancement module, continuous and consistent target tracking trajectory is output.The application improves the target screening ability and tracking accuracy in complex semantic reference scene by text explicit decoupling and multi-granularity progressive alignment, and can be used in the fields of automatic driving, video monitoring and human-computer interaction.
Need to check novelty before this filing date? Find Prior Art