A visual target tracking method, system and device based on multi-modal fusion and a storage medium

By decoupling the visual target tracking task into label optical flow tracking and target tracking subtasks, and introducing high-precision visual label detection, a multi-hypothesis fusion strategy and a closed-loop feedback mechanism are adopted to solve the accuracy and stability problems of visual target tracking in complex environments, thus achieving high-precision and stable target tracking.

CN122416338APending Publication Date: 2026-07-17SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-04-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing visual target tracking methods are not accurate enough or prone to drift when faced with target deformation, occlusion, and changes in lighting. Furthermore, the multi-sensor fusion strategy is simple and lacks an effective feedback mechanism, which leads to a decline in system performance in complex environments.

Method used

The visual target tracking task is decoupled into two parallel subtasks: tag optical flow tracking and target tracking. High-precision visual tag detection is introduced as an anchor point, and dynamic correction is performed through a fusion filter. A multi-hypothesis fusion strategy and a closed-loop feedback mechanism triggered by visual tag detection are adopted to achieve reliability assessment and intelligent fusion of the tracking source.

Benefits of technology

It improves the stability and adaptability of visual target tracking, ensuring high-precision target tracking in complex environments. It suppresses tracking drift through visual marker detection and provides intuitive feedback on position reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416338A_ABST
    Figure CN122416338A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于多模态融合的视觉目标跟踪方法、系统、设备及存储介质,其中方法包括:对视频序列初始化,标注静态或准静态参考标签与待跟踪目标,提取标签稀疏特征点并建立像素与真实空间坐标转换关系;主循环对每一后续帧基于光流法跟踪参考标签以获取稳定位置,并行获取目标视觉标记检测离散位置观测值与连续跟踪观测值,检测成功时重置跟踪器;采用多假设融合策略更新融合滤波器状态,输出目标位置估计,结合参考标签稳定位置完成坐标转换输出真实空间坐标,还可依据状态协方差矩阵迹反馈定位置信度。本发明克服了单一跟踪方法精度不足或漂移严重的缺陷,解决多源观测融合难点,兼顾跟踪连续性与定位精度,抑制累计误差,可实现稳定可靠的真实空间定位,适配复杂工况,鲁棒性与实时性更佳。
Need to check novelty before this filing date? Find Prior Art