一种基于多模态特征融合与特征增强的无人机航拍图像长尾目标检测方法
By employing multimodal feature fusion and feature enhancement methods, the problems of insufficient feature extraction and recognition confusion in long-tail target detection in UAV aerial images are solved, improving detection accuracy and robustness, and achieving efficient recognition of long-tail targets.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
- Filing Date
- 2026-03-04
- Publication Date
- 2026-07-17
AI Technical Summary
Drone aerial images exhibit a long-tail distribution, which leads to a decline in target detection performance. In particular, the recognition accuracy and robustness of long-tail targets are insufficient, and existing methods struggle to effectively address the problems of insufficient feature extraction and recognition confusion in complex backgrounds.
A multimodal feature fusion and feature enhancement method is adopted. Semantic and visual features are extracted through text encoder and visual encoder, and feature enhancement is performed by combining position encoding and self-attention mechanism. A target query set is constructed to optimize the classification performance of the decoding module. A classification loss function is constructed by using class instance adaptation mechanism and semantic visual boundary constraint mechanism.
It improves the detection accuracy and robustness of long-tail targets in complex backgrounds, reduces the risk of identification confusion and false detection, and realizes end-to-end detection of long-tail targets.
Smart Images

Figure CN121767894B_ABST