一种基于多模态特征融合与特征增强的无人机航拍图像长尾目标检测方法

By employing multimodal feature fusion and feature enhancement methods, the problems of insufficient feature extraction and recognition confusion in long-tail target detection in UAV aerial images are solved, improving detection accuracy and robustness, and achieving efficient recognition of long-tail targets.

CN121767894BActive Publication Date: 2026-07-17NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
Filing Date
2026-03-04
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Drone aerial images exhibit a long-tail distribution, which leads to a decline in target detection performance. In particular, the recognition accuracy and robustness of long-tail targets are insufficient, and existing methods struggle to effectively address the problems of insufficient feature extraction and recognition confusion in complex backgrounds.

Method used

A multimodal feature fusion and feature enhancement method is adopted. Semantic and visual features are extracted through text encoder and visual encoder, and feature enhancement is performed by combining position encoding and self-attention mechanism. A target query set is constructed to optimize the classification performance of the decoding module. A classification loss function is constructed by using class instance adaptation mechanism and semantic visual boundary constraint mechanism.

Benefits of technology

It improves the detection accuracy and robustness of long-tail targets in complex backgrounds, reduces the risk of identification confusion and false detection, and realizes end-to-end detection of long-tail targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767894B_ABST
    Figure CN121767894B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于多模态特征融合与特征增强的无人机航拍图像长尾目标检测方法,涉及目标检测技术领域,包括:通过特征提取模块对无人机航拍图像进行特征提取,基于多模态特征融合模块在语义维度、视觉维度进行特征提取与加权融合,获取融合增强特征向量;在通过位置编码模块进行编码后,与融合增强特征向量相加进行全局特征分析,获取编码增强特征向量;通过密度增强特征向量、目标查询数量和引导向量,构建目标查询集合;解码模块基于多层交互注意力机制根据目标查询集合对编码增强特征向量进行逐步解码,获取目标检测结果;并通过分类损失函数模块优化解码模块的分类性能。本发明实现了在复杂背景和噪声干扰下对长尾目标的检测。
Need to check novelty before this filing date? Find Prior Art