The invention discloses a real-time open vocabulary video target detection method based on efficient
time sequence aggregation and knowledge migration, and relates to a
computer vision technology. The problems of poor real-time performance, weak generalization ability of novel categories, insufficient stability of cross-frame detection and the like in the prior art are solved. Firstly, a sparse sampling strategy is adopted to select key frames to reduce calculation overhead; secondly, an enhanced supervision set is constructed, and a basic category real
label and a novel category pseudo
label generated by a teacher model are fused; extracting multi-scale anchor-level features through a
backbone network, and aggregating
time sequence features through an anchor-level memory attention module; and finally, reasoning by adopting a
label space expert mixed strategy, and optimizing the model in combination with a multi-task loss. Compared with the prior art, the method has the advantages that excellent performance-efficiency balance is achieved, the method is outstanding in performance on data sets such as LV-VIS and BURST, the reasoning speed is increased by 70-120 times compared with an existing method, meanwhile, the robustness to novel categories and specific video challenges is enhanced, and the method is suitable for real-time application scenes such as intelligent monitoring and automatic driving.