A Deep Learning-Based Method and Apparatus for Overlaying Unmanned Aerial Vehicle Video Geographic Information

By fusing UAV visual and map features using deep learning technology, the deviation problem caused by GNSS and attitude drift in UAV video geographic information overlay was solved, achieving accurate alignment and stable overlay of geographic information and improving the accuracy of UAV video geographic information overlay.

CN122089585BActive Publication Date: 2026-07-17NORTHEASTERN UNIV CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHEASTERN UNIV CHINA
Filing Date
2026-04-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for overlaying geographic information in UAV video suffer from the instability of GNSS signals and gimbal attitude drift, resulting in the geographic information drifting or misalignment in the video footage and large deviations in the overlay results.

Method used

By employing a deep learning-based approach, a current field-of-view model is constructed by fusing visual semantic features from UAV detection data with map structure features from cloud-based vector map data. A spatial detection model is then used for cross-modal fusion, implicitly inferring the relationship between visual content and geographic elements to achieve precise alignment and stable overlay of geographic information.

Benefits of technology

It reduces the reliance on high-precision GNSS and gimbal attitude, improves the accuracy and robustness of UAV video geographic information overlay, achieves pixel-level precise alignment, and reduces the deviation of the overlay results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089585B_ABST
    Figure CN122089585B_ABST
Patent Text Reader

Abstract

This application provides a method and apparatus for overlaying UAV video geographic information based on deep learning, relating to the field of UAV visual enhancement and geographic information fusion technology. After acquiring UAV detection data, the method constructs a current field-of-view model based on positioning and flight status data, and retrieves vector map data from a map database based on the field-of-view model. Then, a deep learning-based spatial detection model is used to perform cross-modal fusion of the visual semantic features of the video frame image and the map structure features of the vector map data, and spatial correspondence information between visual content and geographic features is established based on implicit inference. Finally, geographic information is overlaid onto the video frame image according to the spatial correspondence information to generate an overlaid frame image. This method can utilize the cross-modal fusion capability of visual semantics and map priors to achieve pixel-level precise alignment of geographic information in the video frame image, reducing the deviation in the video geographic information overlay result.
Need to check novelty before this filing date? Find Prior Art