A multi-purpose pedestrian re-identification method based on instruction guidance under the perspective of a drone

By employing a command-guided method from the perspective of drones, utilizing improved PSMNet and Transformer feature extraction, and combining multimodal feature fusion, the challenge of pedestrian re-identification in complex environments was solved, achieving versatile and robust pedestrian recognition performance.

CN122435485APending Publication Date: 2026-07-21SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610544016.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing pedestrian re-identification technologies face challenges in complex environments, such as changes in clothing, modal differences, and perspective, making it difficult to achieve flexible and robust recognition for multiple purposes.

Method used

We adopt a command-guided approach from the perspective of UAVs, perform disparity estimation through an improved PSMNet, combine Transformer feature extraction and multimodal feature fusion, and construct an end-to-end joint loss function using a command-aware Q-Former module and a deep interactive multimodal feature fusion module to achieve unified modeling for multiple tasks.

Benefits of technology

It reduces system deployment complexity, improves retrieval flexibility, alleviates perspective bias under UAV viewpoint, enhances discrimination capabilities in scenarios such as occlusion and lighting changes, and achieves robustness and accuracy of multi-purpose pedestrian re-identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435485A_ABST
    Figure CN122435485A_ABST
Patent Text Reader

Abstract

The application discloses a multi-purpose pedestrian re-identification method based on instruction guidance under the visual angle of a UAV. In view of the appearance deviation caused by the overhead visual angle of the UAV and the problem that the existing model is difficult to adapt to various retrieval requirements, the application obtains pedestrian spatial angle information by binocular disparity estimation and embeds the feature sequence; a Transformer network containing spatial marker selection and frequency marker selection is constructed to extract pedestrian features; a Q-Former module is used to fine-tune the pedestrian features according to the instruction; the features are deeply fused through a multi-modal feature fusion module; and semantic weight triplet loss and identity loss are combined for optimization. The application introduces a text instruction as high-level semantic prior, so that a single model can adapt to various tasks such as traditional re-identification, clothing change re-identification, cross-modal re-identification and text-image re-identification, and the spatial angle information is used to relieve the overhead visual angle deviation and improve the recognition accuracy and robustness in a complex scene.
Need to check novelty before this filing date? Find Prior Art