一种基于深度学习的端到端的多飞行器跟踪方法

By employing an end-to-end deep learning approach, combined with a dynamic multi-scale spatiotemporal network and a global-local extraction module, the problems of occlusion and recognition errors in multi-target tracking models during airport surface surveillance were solved, achieving higher tracking accuracy and stability.

CN117765027BActive Publication Date: 2026-07-17UNIV OF ELECTRONICS SCI & TECH OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2023-12-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multi-target tracking models cannot effectively utilize the temporal information between video frames in airport surface surveillance, leading to aircraft target occlusion and identification errors. Furthermore, traditional methods cannot extract global and local features simultaneously, resulting in decreased tracking stability.

Method used

We adopt an end-to-end approach based on deep learning, extracting the spatiotemporal fusion features of the current frame and keyframes through a dynamic multi-scale spatiotemporal network and a global-local extraction module. Combining the representational capabilities of the Transformer model, we use a deep similarity network for target association and matching, abandoning traditional constraints and directly using detection embeddings and tracking embedding vectors for similarity calculation.

Benefits of technology

It effectively solves the problem of aircraft target obstruction, improves the robustness and stability of multi-target tracking, and enhances tracking accuracy in airport surface surveillance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117765027B_ABST
    Figure CN117765027B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于深度学习的端到端的多飞行器跟踪方法,包括步骤:使用结构相似度算法对机场场面视频中当前帧之前的T帧进行关键帧选择,选择其中与当前帧低层次语义特征差距最大的帧,作为关键帧;将当前帧和关键帧输入动态多尺度空时网络,获取当前帧和关键帧之间的多尺度空时融合特征;将多尺度空时融合特征作为空时相似度估计网络的输入,利用Transformer模型的表示能力获得检测嵌入和跟踪嵌入;并使用深度相似度估计网络生成相似度矩阵;根据相似度矩阵对当前帧中的飞行器目标与关键帧中的飞行器目标进行关联和匹配,计算飞行器轨迹,完成端到端的多飞行器跟踪,本发明充分利用了整个视频帧之间的时序信息。
Need to check novelty before this filing date? Find Prior Art