Unmanned aerial vehicle aerial image analysis method and system

By combining multi-source remote sensing data synchronization and dual-branch network feature extraction with UAV aerial image analysis based on LiDAR point cloud data, the problems of target detection accuracy and energy efficiency of UAVs in complex environments were solved, achieving high-precision detection and optimized endurance.

CN121170629APending Publication Date: 2025-12-19BEIJING ZHONGSHI RONGCHUANG TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511227709.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing UAV aerial image analysis technology is limited in target detection accuracy, decision-making intelligence, and energy efficiency in complex and dynamic environments. The efficiency of multi-source data collaboration is low, and the fragmentation of functional modules leads to insufficient system practicality.

Method used

Multi-source remote sensing data is processed by time synchronization and spatial alignment. Multi-scale fusion features are extracted by combining the dual-branch Blur-PANet network to conduct threat risk assessment and generate obstacle avoidance instructions. Transfer learning is then performed using LiDAR point cloud data for verification.

Benefits of technology

It achieves high-precision target detection in complex environments, reduces collision risk, optimizes energy consumption and endurance, and improves the reliability of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170629A_ABST
    Figure CN121170629A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle aerial image analysis method and system, and the method comprises the steps: collecting multi-source remote sensing data, and carrying out the time synchronization and space alignment processing, and obtaining the preprocessing data; inputting the preprocessed data into a feature extraction model, and extracting multi-scale fusion features; performing threat risk assessment on a threat target in the environment where the unmanned aerial vehicle is located; when the threat risk meets a preset avoidance condition, aligning the ground data and the aerial photography data by using a feature matching algorithm, and cooperatively generating an obstacle avoidance instruction carrying an unmanned aerial vehicle obstacle avoidance path; and after the unmanned aerial vehicle executes the obstacle avoidance instruction, aligning the current ground data with the aerial photography data by using the feature matching algorithm again, projecting the target motion track of the unmanned aerial vehicle to the geographic space to generate a motion mask in combination with the LiDAR point cloud data, and carrying out transfer learning verification on the static region to obtain an unmanned aerial vehicle aerial photography image analysis result. According to the invention, the autonomous operation performance, robustness and cruising ability of the unmanned aerial vehicle in a complex dynamic environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned aerial vehicle remote sensing vision, in particular to an unmanned aerial vehicle aerial image analysis method and system. BACKGROUND

[0002] The unmanned aerial vehicle aerial image analysis technology is widely used in the fields of power inspection and disaster search and rescue. The current unmanned aerial vehicle aerial image analysis technology mainly relies on a deep learning framework, optimizes a region proposal network (RPN) to improve the quality of a candidate frame, and the detection precision (AP50) of a medium-sized target such as a vehicle or a building can reach 78.5% under ideal lighting conditions. Meanwhile, a mainstream three-dimensional reconstruction scheme adopts a motion recovery structure (SfM) technology, generates a sparse point cloud by using a sequence of aerial images, and constructs a dense point cloud model by using multi-view stereo matching (MVS). In terms of dynamic obstacle avoidance, a traditional method combines a YOLO series model to output a detection frame in real time, combines an artificially set safety distance threshold to trigger an avoidance action, and a multi-sensor fusion field generally adopts an offline calibration and timestamp alignment strategy, fuses visible light and infrared data streams by using Kalman filtering, and on the hardware level, a commercial unmanned aerial vehicle generally adopts a rolling shutter sensor in combination with a mechanical shutter to balance the cost and imaging speed requirements.

[0003] However, although the existing technology has made certain progress, there are still the following systematic bottlenecks:

[0004] 1) Low coordination efficiency of multi-source data, different sensors cause time and space mismatch due to hardware synchronization error, and the detection frame of a moving target is shifted by up to 10 pixels; air-ground image registration relies on manual selection of control points, and the DEM reconstruction height error is more than 2 meters, which cannot meet the high-precision requirements of geological disaster monitoring.

[0005] 2) Dynamic decision and energy efficiency imbalance: the obstacle avoidance strategy is only based on a static safety distance, without modeling the motion state of the target, and the collision probability increases when encountering a high-speed moving obstacle; meanwhile, the flight parameters are fixedly set, and the unmanned aerial vehicle maintains low-altitude flight in steep slope areas, which consumes additional energy; on the hardware level, the inference delay of an 80MB-level model on an edge device is more than 200ms, and the battery state is not associated with the task scheduling, and high-load tasks such as three-dimensional reconstruction aggravate the decline of the endurance.

[0006] 3) Cross-technology module fragmentation: the detection, obstacle avoidance and reconstruction functions are independently run, for example, the change detection does not use the motion trajectory data generated by the real-time obstacle avoidance to cause a high false detection rate, and the SLAM module does not feed back the terrain curvature to the flight control system, causing the flight path planning to be out of touch with the terrain matching, and these defects jointly restrict the practicability of the system in the scenes of power inspection and disaster search and rescue. SUMMARY

[0007] To this end, the unmanned aerial vehicle aerial image analysis method and system are provided, aiming at solving the technical problems of limited target detection accuracy, decision intelligence and energy efficiency of the unmanned aerial vehicle aerial image analysis in the prior art in a complex dynamic environment.

[0008] To achieve the above object, the application adopts the following technical solutions:

[0009] According to the first aspect of the application, the application provides an unmanned aerial vehicle aerial image analysis method, which comprises:

[0010] Collecting multi-source remote sensing data, and performing time synchronization and spatial alignment processing on the multi-source remote sensing data to obtain preprocessed data; the multi-source remote sensing data at least includes visible light images and infrared images;

[0011] Inputting the preprocessed data into a feature extraction model based on a double-branch Blur-PANet network to extract multi-scale fusion features in the preprocessed data;

[0012] Based on the multi-scale fusion features, threat risk assessment is performed on the threat targets in the environment where the unmanned aerial vehicle is located;

[0013] When the threat risk meets the preset avoidance condition, the feature matching algorithm is used to align the ground data and aerial data, and the obstacle avoidance instruction carrying the unmanned aerial vehicle obstacle avoidance path is generated in collaboration;

[0014] After the unmanned aerial vehicle executes the obstacle avoidance instruction, the feature matching algorithm is used again to align the current ground data and aerial data, and the LiDAR point cloud data is combined to project the target motion trajectory of the unmanned aerial vehicle to the geographic space to generate a motion mask, and the static area is verified by transfer learning to obtain the unmanned aerial vehicle aerial image analysis result.

[0015] Further, the multi-source remote sensing data is collected, comprising:

[0016] The multi-source remote sensing data is collected by using a global exposure CMOS sensor, specifically comprising:

[0017] The exposure parameters for collecting the multi-source remote sensing data are dynamically adjusted based on an environment perception function; the mathematical expression of the environment perception dynamic function is:

[0018] E(x,y,z,t)=λ1·T+λ2·O+λ3·W

[0019] Wherein, E(x,y,z,t) represents the environment perception evaluation value for adjusting the exposure parameter; T represents the terrain feature; O represents the obstacle distribution; W represents the weather condition; λ1, λ2, λ3 represent the dynamic weight coefficient; (x,y,z), t represent the coordinate and time parameter respectively.

[0020] Further, the time synchronization and spatial alignment processing of the multi-source remote sensing data comprises:

[0021] adopting a timestamp interpolation method to compress the time synchronization deviation of the multi-source remote sensing data to within a preset time length threshold; and

[0022] controlling the moving target detection frame offset within a preset pixel threshold by fusing the edge features of the visible light image and the thermal radiation features of the infrared image in the multi-source remote sensing data through a phase consistency algorithm, specifically comprising:

[0023] establishing a pixel coordinate mapping function by using pre-acquired lens distortion parameters;

[0024] extracting phase consistency features of the visible light image and the infrared image based on the pixel coordinate mapping function, and eliminating edge misregistration caused by space-time offset by fusing gradient amplitudes of the visible light image and the infrared image.

[0025] Further, the backbone network of the feature extraction model adopts a depth separable convolution, and the neck network adopts a weighted bidirectional cyclic path.

[0026] the multi-scale fusion features in the preprocessed data comprise:

[0027] applying adaptive histogram equalization to the visible light image, and performing non-uniformity correction on the infrared image to generate a training sample set;

[0028] the backbone network extracts shallow high-resolution features and deep semantic features in the training sample set based on the depth separable convolution;

[0029] the neck network fuses the shallow high-resolution features through a forward path, injects the deep semantic features through a reverse path, adjusts feature contribution degrees according to learnable weights, and fuses to generate double-branch Blur-PANet network features as multi-scale fusion features.

[0030] Further, the adjusting feature contribution degrees according to learnable weights and fusing to generate double-branch Blur-PANet network features comprise:

[0031] a fusion variable is calculated by using a learnable weight fusion formula, and the mathematical expression is:

[0032] w = σ (α·LBP + β·OFM)

[0033] wherein, w represents the fusion variable; σ represents an activation function; LBP represents an LBP texture feature; OFM represents an optical flow motion feature; and α and β represent corresponding dynamic adjustment coefficients, respectively.

[0034] In the feature extraction process, the angular velocity value of the unmanned aerial vehicle is multiplied by the inscribed circle radius of the detected target, the displacement vector is calculated in combination with the inter-frame time interval, the sampling position of the deformable convolution kernel is dynamically adjusted according to the displacement vector, the perception field of the deformable convolution kernel is offset along the target motion direction, and the edge texture of the corresponding region in the feature map is locally reconstructed.

[0035] Further, the threat risk assessment of the threat target in the environment where the unmanned aerial vehicle is located based on the multi-scale fusion feature includes:

[0036] A direction perception loss function is introduced in the regression task of the multi-scale fusion feature, an energy focusing loss is used to balance positive and negative samples, a threat function value is calculated in real time, and specifically includes:

[0037] The sine difference of the direction angle of the predicted box and the real box is calculated to quantify the angle deviation, and the loss weight is dynamically adjusted in combination with the rotated box intersection over union ratio; at the same time, the energy focusing loss is used to reconstruct the classification task, an exponential penalty weight is applied to the difficult classification samples, and the loss contribution of the easy classification samples is reduced by adjusting the focusing factor, so as to optimize the weight ratio of positive and negative samples;

[0038] Based on the positive and negative samples, a threat assessment model is established, which fuses the relative speed, collision time and angular velocity module length of the threat target and the unmanned aerial vehicle, and contains a dynamic risk term and a safety margin, the threat risk assessment model is used to assess the threat target in the environment where the unmanned aerial vehicle is located, and a threat assessment value is obtained.

[0039] Further, when the threat risk meets the preset avoidance condition, the feature matching algorithm is used to align the ground data and the aerial data, and the obstacle avoidance instruction carrying the unmanned aerial vehicle obstacle avoidance path is generated in cooperation, including:

[0040] When the threat assessment value is lower than a preset safety threshold, it is considered that the preset avoidance condition is met;

[0041] The coordinates of the safety envelope boundary points are calculated as the obstacle avoidance maneuver points, the yaw angle and pitch angle control amounts are output, and are sent to the ground control point;

[0042] The ground control point aligns the ground data and the aerial data through the feature matching algorithm, calculates the flight curvature according to the digital elevation model, and dynamically adjusts the flight parameters of the unmanned aerial vehicle;

[0043] The flight parameters include flight height and flight speed.

[0044] Further, the method further includes:

[0045] When the unmanned aerial vehicle deviates from the target route, a segmented regression path is dynamically generated through a route regression function, specifically including:

[0046] based on the horizontal error distance and the azimuth deviation of the current position and the target waypoint, the regression heading angle is calculated in real time by using the route regression function;

[0047] when the horizontal error distance exceeds the preset error threshold, the segmented regression path is generated;

[0048] wherein, the segmented regression path includes a first segment of the regression path for fast approaching the target route projection point with a large inclination angle, a second segment of the regression path for calibrating the attitude along the route tangent direction, and a third segment of the regression path for regression of the target route with constant height;

[0049] the flight parameters and the segmented regression path adjusted dynamically are sent to the UAV as the obstacle avoidance instruction.

[0050] Further, the current ground data and aerial data are aligned by using the feature matching algorithm, the target motion trajectory of the UAV is projected to the geographic space to generate a motion mask, the static region is verified by transfer learning, and the UAV aerial image analysis result is obtained, including:

[0051] the control points of the ground data and the aerial data are extracted by the feature matching algorithm, and the geographic coordinate system of the ground data, the aerial data and the LiDAR point cloud data is unified by using the homography matrix to reconstruct the geographic space;

[0052] the target motion trajectory output by the multi-target tracking is projected to the geographic space to generate a motion mask, and the change detection result of the static region is retained based on the motion mask filtering temporary change interference;

[0053] the change detection result is input into the pre-trained ShuffleNet fine-tuning learning network model, the true and false changes of the static region are verified by RGB-LBP fusion features, and the UAV aerial image analysis result is obtained.

[0054] According to the second aspect of the present application, the present application provides a UAV aerial image analysis system, the system comprising:

[0055] a multi-source acquisition module for acquiring multi-source remote sensing data, and performing time synchronization and spatial alignment processing on the multi-source remote sensing data to obtain preprocessed data; the multi-source remote sensing data includes visible light images, infrared images and / or depth images;

[0056] a feature fusion module for inputting the preprocessed data into a feature extraction model based on a double-branch Blur-PANet network to extract multi-scale fusion features in the preprocessed data;

[0057] a threat assessment module configured to perform threat risk assessment on a threat target in an environment where the UAV is located based on the multi-scale fusion features;

[0058] a collaborative decision-making module configured to, when the threat risk meets a preset avoidance condition, align ground data and aerial data by using a feature matching algorithm, and collaboratively generate an obstacle avoidance instruction carrying an obstacle avoidance path of the UAV;

[0059] a result verification module configured to, after the UAV executes the obstacle avoidance instruction, align current ground data and aerial data by using the feature matching algorithm again, and, in combination with LiDAR point cloud data, project a target motion trajectory of the UAV to a geographic space to generate a motion mask, verify a static region by transfer learning, and obtain an analysis result of a UAV aerial image.

[0060] The technical scheme of the present application has at least the following beneficial effects:

[0061] According to the technical scheme of the present application, multi-source remote sensing data is collected, and the multi-source remote sensing data is subjected to time synchronization and spatial alignment processing to obtain preprocessed data; the preprocessed data is input into a feature extraction model based on a double-branch Blur-PANet network to extract multi-scale fusion features in the preprocessed data; based on the multi-scale fusion features, threat risk assessment is performed on a threat target in an environment where the UAV is located; when the threat risk meets a preset avoidance condition, ground data and aerial data are aligned by using a feature matching algorithm, and an obstacle avoidance instruction carrying an obstacle avoidance path of the UAV is collaboratively generated; after the UAV executes the obstacle avoidance instruction, the feature matching algorithm is used again to align current ground data and aerial data, and, in combination with LiDAR point cloud data, a target motion trajectory of the UAV is projected to a geographic space to generate a motion mask, a static region is verified by transfer learning, and an analysis result of a UAV aerial image is obtained. At least the following technical effects are achieved:

[0062] 1) By constructing a collaborative closed-loop architecture from data collection, feature analysis to dynamic decision-making and closed-loop verification, the functional modules are integrated, the defects of fragmented functions in the prior art are overcome, the triple optimization of energy consumption, precision and real-time performance is achieved, and the autonomous operation performance, robustness and endurance of the UAV in a complex dynamic environment are comprehensively improved;

[0063] 2) By high-precision time synchronization and spatial alignment processing, the target detection error caused by time-space mismatch of multi-source data is effectively reduced, and in combination with a network capable of extracting multi-scale fusion features, the detection accuracy of small targets in a complex environment is significantly improved;

[0064] 3) Through real-time risk assessment and dynamic avoidance mechanism based on deep features, it can effectively deal with high-speed moving obstacles and reduce the risk of collision; at the same time, based on the energy efficiency adaptive optimization strategy of environment and self-state, the energy consumption of unmanned aerial vehicle in complex terrain is significantly reduced, and the endurance time is effectively prolonged;

[0065] 4) Through the closed-loop feedback and verification mechanism after executing the instruction, it can effectively filter dynamic interference, distinguish real changes from temporary occlusion, significantly reduce the false detection rate of advanced tasks such as change detection, and ensure the reliability of the final analysis result.

[0066] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, a brief introduction will be given below to the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0068] Figure 1 A system architecture schematic diagram applied to the unmanned aerial vehicle aerial image analysis method provided by an embodiment of the present application is shown;

[0069] Figure 2 A flowchart of the unmanned aerial vehicle aerial image analysis method provided by an embodiment of the present application is shown;

[0070] Figure 3 A structure schematic diagram of the unmanned aerial vehicle aerial image analysis system provided by an embodiment of the present application is shown;

[0071] Figure 4 An entity structure schematic diagram of the computer device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0072] The exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0073] It is to be noted that the relative terms such as first and second and the like are used herein solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... " does not, without more limitations, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0074] Referring to Figure 1 , a system architecture diagram applied by a UAV aerial image analysis method is shown. The system architecture mainly includes a UAV 10 deployed in the air and a ground control point 20 arranged on the ground, and the two can communicate with each other through a high-bandwidth, low-latency wireless data link. The UAV 10 serves as a task execution and front-end perception platform, and can integrate a multi-modal sensor group 11 for collecting multi-source remote sensing data, a high-performance on-board processing unit 12 responsible for computing power, and a flight control unit 13 responsible for flight attitude and trajectory control. Correspondingly, the ground control point 20 serves as the decision center of the whole system, and the core is a collaborative decision center 21 responsible for processing massive data, running complex algorithms, and performing global task planning and energy efficiency optimization.

[0075] A UAV aerial image analysis method provided by an embodiment of the present application will be described below. As shown in Figure 2 , the method can at least include the following steps S201-S205:

[0076] In step S201, multi-source remote sensing data is collected, and time synchronization and spatial alignment processing are performed on the multi-source remote sensing data to obtain preprocessed data.

[0077] The multi-source remote sensing data in the embodiment of the present application can include visible light images, infrared images, and depth images. In actual application, a global exposure CMOS sensor can be used for collection, and an electronic shutter can be used to eliminate the rolling shutter effect. Specifically, when collecting multi-source remote sensing data, the exposure parameters for collecting multi-source remote sensing data can be dynamically adjusted based on an environment perception function; the mathematical expression of the environment perception dynamic function is:

[0078] E(x, y, z, t) = λ1·T + λ2·O + λ3·W

[0079] Wherein, E(x,y,z,t) represents an environment perception evaluation value for adjusting exposure parameters;T represents a terrain feature, which can be extracted by preloading a digital elevation model;O represents an obstacle distribution, which is generated by real-time point cloud segmentation;W represents weather conditions, which integrates temperature and humidity, wind speed sensor data, etc.;λ1, λ2, λ3 represent dynamic weight coefficients;(x,y,z), t represent coordinates and time parameters respectively.

[0080] Further, when time synchronization and spatial alignment processing is performed on the multi-source remote sensing data, the time synchronization deviation of the multi-source remote sensing data can be compressed to within a preset time length threshold by using a timestamp interpolation method;And by using a phase consistency algorithm to fuse the edge features of the visible light image and the thermal radiation features of the infrared image in the multi-source remote sensing data, the shift of the moving target detection frame is controlled within a preset pixel threshold.

[0081] That is, the embodiment of the application can use the lens distortion parameters obtained in advance by the six-degree-of-freedom calibration robot arm to establish a pixel coordinate mapping function;Based on the pixel coordinate mapping function, the phase consistency features of the visible light image and the infrared image are extracted, and the edge misregistration caused by the spatio-temporal shift is eliminated by fusing the gradient amplitudes of the visible light image and the infrared image.

[0082] In actual application, a six-degree-of-freedom robot arm is used to hold a multi-sensor group, move along a preset spiral trajectory around a high-precision checkerboard calibration board, collect calibration images at 72 poses, solve the intrinsic parameter matrix K and the distortion coefficients [k1, k2, p1, p2, k3] by Zhang Zhengyou calibration method, establish a pixel coordinate mapping function, and generate a pixel coordinate mapping table containing 1280*720 pixels:

[0083]

[0084] Therefore, the phase consistency features of the visible light image and the infrared image are extracted in 5 scale spaces, the gradient direction of the visible light image is calculated, the isotherm normal vector of the infrared image is extracted, and the edge misregistration caused by the spatio-temporal shift of the sensor is eliminated by fusing the gradient amplitudes through directional weighting.

[0085] In step S202, the preprocessed data is input into a feature extraction model based on a double-branch Blur-PANet network to extract multi-scale fusion features in the preprocessed data.

[0086] The feature extraction model proposed in the embodiment of the application uses a depth separable convolution to replace a standard convolution layer in the backbone network, and uses a weighted bidirectional cyclic path in the neck network to fuse shallow high-resolution features and deep semantic features according to a learnable weight to generate double-branch Blur-PANet network data.

[0087] Further, the multi-scale fusion features in the pre-processed data are extracted, specifically, adaptive histogram equalization can be applied to the visible light image, and non-uniformity correction can be performed on the infrared image to generate a training sample set; the backbone network extracts shallow high-resolution features and deep semantic features in the training sample set based on a depth separable convolution; the neck network fuses the shallow high-resolution features through a forward path, injects the deep semantic features through a reverse path, adjusts the feature contribution degree according to a learnable weight, and fuses to generate a double-branch Blur-PANet network feature as the multi-scale fusion feature.

[0088] In actual application, the transmission delay needs to be compensated by using a hardware time stamp interpolation first to ensure that the multi-modal data time synchronization deviation is less than or equal to 1 ms, and then the radiation correction is performed, adaptive histogram equalization is applied to the visible light image, non-uniformity correction is performed on the infrared image, and a training sample set is generated by using random rotation and Mosaic splicing technology; the C3 module of YOLOv5 is replaced by a depth separable convolution in the backbone network, the standard 3x3 convolution is decomposed into a channel-independent depth convolution and a 1x1 point convolution, and the neck network is innovatively designed with a weighted bidirectional cyclic path, which fuses the shallow high-resolution features through a forward path and injects the deep semantic features through a reverse path to output, and dynamically adjusts the feature contribution degree through a learnable gating weight.

[0089] Further, when the double-branch Blur-PANet network feature is fused by adjusting the feature contribution degree according to the learnable weight, the learnable weight fusion formula can be used to calculate the fusion variable as follows:

[0090] w = σ (a LBP + b OFM)

[0091] wherein w represents the fusion variable; sigma represents an activation function; LBP represents an LBP texture feature; OFM represents an optical flow motion feature; a and b represent corresponding dynamic adjustment coefficients, i.e., contribution weights. The initial values can be set as a = 0.6 and b = 0.4, which are automatically optimized through back propagation.

[0092] It should be noted that in the feature extraction process, the angular velocity value of the unmanned aerial vehicle is multiplied by the inscribed circle radius of the detected target, the displacement vector is calculated in combination with the inter-frame time interval, the sampling position of the deformable convolution kernel is dynamically adjusted according to the displacement vector, the perception field of the deformable convolution kernel is shifted along the target motion direction, and the edge texture of the corresponding area in the feature map is locally reconstructed.

[0093] Step S203, based on the multi-scale fusion features, threat risk assessment is performed on the threat target in the environment where the unmanned aerial vehicle is located.

[0094] By introducing a direction perception loss function in the regression task of multi-scale fusion features, an energy focus loss is used to balance positive and negative samples, and a threat risk assessment is performed in real time by calculating the value of the threat function.

[0095] Specifically, the angle deviation is quantified by calculating the sine difference of the direction angle between the predicted box and the real box, and the loss weight is dynamically adjusted by combining the rotated box intersection over union. Meanwhile, the energy focus loss is used to reconstruct the classification task, and the exponential penalty weight is applied to the difficult classification samples, and the loss contribution of the easy classification samples is reduced by adjusting the focus factor, so as to optimize the weight ratio of positive and negative samples. For example, the weight ratio of positive and negative samples is optimized from 1:100 to 1:3. Further, based on the positive and negative samples, a threat assessment model is established, which integrates the relative speed of the threat target and the unmanned aerial vehicle, the collision time and the angular velocity module length, and contains a dynamic risk term and a safety margin. The threat assessment model is used to assess the threat risk of the threat target in the environment of the unmanned aerial vehicle, and a threat assessment value is obtained.

[0096] The threat assessment value is used to compare with a preset safety threshold, and when it is lower than the preset safety threshold, an obstacle avoidance instruction is immediately generated, and the coordinates of the safety envelope boundary points are output to generate the yaw angle and pitch angle control amount, so as to realize the full-link response from threat detection to avoidance action.

[0097] In step S204, when the threat risk meets the preset avoidance condition, the feature matching algorithm is used to align the ground data and the aerial data, and the obstacle avoidance instruction carrying the obstacle avoidance path of the unmanned aerial vehicle is generated.

[0098] As described above, the embodiment of the present application judges whether the threat risk meets the preset avoidance condition by comparing the threat assessment value with the preset safety threshold. When the threat assessment value is lower than the preset safety threshold, a semi-elliptical-semicircular combined envelope is generated based on the target angular velocity, the nearest point of the envelope boundary is calculated as the avoidance maneuver point, and the visual servo control amount including the yaw angle and pitch angle control amount is output to the ground control point; the ground control point aligns the ground data and the aerial data by using the feature matching algorithm, calculates the flight curvature according to the digital elevation model, and dynamically adjusts the flight parameters of the unmanned aerial vehicle, including the flight height and the flight speed.

[0099] The flight height can be calculated as: flight height = reference height + γ x curvature; and the flight speed can be calculated as: flight speed = maximum flight x e^(-β x curvature). Wherein, γ is a curvature correction coefficient, and β is a decay coefficient.

[0100] It should be noted that in actual application, when the battery power of the unmanned aerial vehicle is lower than a preset power threshold (for example, 20%), the three-dimensional reconstruction thread is closed, and the dynamically adjusted flight parameters and the segmented regression path are sent to the unmanned aerial vehicle as the obstacle avoidance instruction by using the route regression function.

[0101] Specifically, when the UAV deviates from the target flight path, a regression heading angle can be calculated in real time based on the horizontal error distance and the azimuth deviation of the current position from the target waypoint using a flight path regression function; when the horizontal error distance exceeds a preset error threshold (e.g., 50 meters), a segmented regression path is generated. The segmented regression path includes a first segment of the regression path that rapidly approaches the target flight path projection point at a large inclination angle, a second segment of the regression path that calibrates the attitude along the tangent direction of the flight path, and a third segment of the regression path that returns to the target flight path at a constant height.

[0102] When sending dynamic adjustment data to the UAV, the flight parameters and segmented path point coordinates can be packaged as a binary instruction stream, which is sent through a low-delay data transmission link. The on-board navigation module can parse the instructions in real time and adjust the motor speed and rudder deflection through a PID controller. In steep slope segments with high curvature (e.g., > 0.1 m -1 ), the height is automatically increased to ensure that the trajectory tracking error is within the preset error threshold.

[0103] Step S205: After the UAV executes the obstacle avoidance instruction, the current ground data and aerial data are aligned again using the feature matching algorithm, and the target motion trajectory of the UAV is projected into the geographic space to generate a motion mask. The static region is verified by transfer learning to obtain the UAV aerial image analysis result.

[0104] Specifically, the control points of the ground data and the aerial data are extracted by the feature matching algorithm, and the geographic coordinate system of the ground data, the aerial data, and the LiDAR point cloud data is unified by the homography matrix to reconstruct the geographic space. The target motion trajectory output by the multi-target tracking is projected into the geographic space to generate a motion mask. Based on the motion mask, temporary changes are filtered, and the change detection results of the static region are retained. The change detection results are input into the pre-trained ShuffleNet fine-tuning learning network model to verify the true and false changes of the static region through RGB-LBP fusion features, and the UAV aerial image analysis result is obtained.

[0105] In practical applications, the homography matrix can be solved by the random sample consensus algorithm. When reconstructing the geographic space, the aerial UAV view, the ground view, and the LiDAR point cloud can be unified to the geographic coordinate system. The k-nearest neighbor average distance of each point is calculated, and the outliers exceeding 3 times the global average distance standard deviation are removed. The moving least squares method is used to reconstruct the terrain surface. Then, the motion trajectory output by the multi-target tracking is projected into the geographic space grid to generate a binary motion mask. Based on the mask, temporary changes are filtered, and only the change detection results of the static region are retained. Subsequently, the pre-trained ShuffleNet fine-tuning learning network is input, and the true and false changes are verified through RGB-LBP fusion features to achieve image analysis result verification.

[0106] It should be noted that the preset time threshold, the preset pixel threshold, the preset error threshold and the preset safety threshold mentioned in the above description can be set according to the needs in actual application, and the present application does not limit this.

[0107] The embodiment of the present application provides a UAV aerial image analysis method, which comprises the following steps: collecting multi-source remote sensing data, and performing time synchronization and space alignment processing on the multi-source remote sensing data to obtain preprocessed data; inputting the preprocessed data into a feature extraction model based on a double-branch Blur-PANet network to extract multi-scale fusion features in the preprocessed data; based on the multi-scale fusion features, performing threat risk assessment on threat targets in an environment where the UAV is located; when the threat risk meets a preset avoidance condition, using a feature matching algorithm to align ground data and aerial data, and collaboratively generating an obstacle avoidance instruction carrying an obstacle avoidance path of the UAV; after the UAV executes the obstacle avoidance instruction, again using the feature matching algorithm to align the current ground data and aerial data, and combining LiDAR point cloud data, projecting a target motion trajectory of the UAV into a geographic space to generate a motion mask, verifying a static region through transfer learning, and obtaining a UAV aerial image analysis result. The present application at least has the following technical effects:

[0108] 1) Through the multi-modal sensor collaborative mechanism and the lightweight network architecture innovation, the performance of small target detection in complex environments is significantly improved, and the global exposure CMOS sensor cooperates with the time stamp interpolation method to compress the time difference of multi-source data to within 1 millisecond;

[0109] 2) Combined with the phase consistency fusion algorithm to eliminate the edge misregistration of infrared / visible light, the detection box of the moving target is offset by ≤0.5 pixels, the double-branch Blur-PANet network reduces the calculation amount through depth separable convolution, and the weighted bidirectional cyclic path adaptively fuses multi-scale features through learnable gating weights, the small target detection accuracy is improved on the COCO-Air dataset, in addition, hardware perception quantization and dynamic calculation scheduling make the model volume compressed, and the real-time obstacle avoidance demand is met;

[0110] 3) Constructing a cross-module collaborative decision chain and an energy efficiency optimization system, a direction perception loss function is combined with a threat function to trigger a semi-ellipse-semicircle envelope avoidance mechanism, so that the collision risk of dynamic obstacles is reduced, the flight height and speed are dynamically adjusted based on the curvature of the digital elevation model, the energy consumption in steep slope areas is reduced, and when the battery power is less than 20%, the three-dimensional reconstruction task is turned off, a segmented path is generated combined with the air route regression function, and the endurance is prolonged;

[0111] 4) LiDAR point cloud statistical analysis filtering combined with transfer learning network, change detection false detection rate compression, sensor-network-decision closed loop architecture breaks through small target detection and dynamic obstacle avoidance collaborative bottleneck, realizes energy consumption-precision-real-time triple optimization, edge device computing efficiency is improved, and the endurance of the aerial unmanned aerial vehicle is greatly increased.

[0112] Further, as a specific implementation of Figure 2 , the embodiment of the present application provides an unmanned aerial vehicle aerial image analysis system, as shown in Figure 3 , the system can include: a multi-source acquisition module 310, a feature fusion module 320, a threat assessment module 330, a collaborative decision module 340, and a result verification module 350.

[0113] The multi-source acquisition module 310 can be used to acquire multi-source remote sensing data, and perform time synchronization and spatial alignment processing on the multi-source remote sensing data to obtain preprocessed data; the multi-source remote sensing data includes visible light images, infrared images and / or depth images;

[0114] The feature fusion module 320 can be used to input the preprocessed data into a feature extraction model based on a double-branch Blur-PANet network to extract multi-scale fusion features in the preprocessed data;

[0115] The threat assessment module 330 can be used to perform threat risk assessment on threat targets in an environment where the unmanned aerial vehicle is located based on the multi-scale fusion features;

[0116] The collaborative decision module 340 can be used to align ground data and aerial data using a feature matching algorithm when the threat risk meets a preset avoidance condition, and collaboratively generate an obstacle avoidance instruction carrying an unmanned aerial vehicle obstacle avoidance path;

[0117] The result verification module 350 can be used to, after the unmanned aerial vehicle executes the obstacle avoidance instruction, align the current ground data and aerial data using the feature matching algorithm again, and combine LiDAR point cloud data to project a target motion trajectory of the unmanned aerial vehicle to a geographic space to generate a motion mask, perform transfer learning verification on a static region, and obtain an unmanned aerial vehicle aerial image analysis result.

[0118] It should be noted that other corresponding descriptions of the various functional modules involved in the unmanned aerial vehicle aerial image analysis system provided by the embodiment of the present application can refer to the corresponding descriptions of the method shown in Figure 2 , which will not be described here.

[0119] Based on the above method shown in Figure 2 , accordingly, the embodiment of the present application also provides a computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the steps of the unmanned aerial vehicle aerial image analysis method described in any of the above embodiments.

[0120] Based on the above method as shown in Figure 2 and the system as shown in Figure 3 Embodiments of the present application also provide a physical structure diagram of a computer device, as shown in Figure 4 which can include a communication bus, a processor, a memory and a communication interface, and can further include an input / output interface and a display device, wherein the communication between the functional units can be completed through the bus. The memory stores a computer program, and the processor is configured to execute the program stored in the memory to execute the steps of the UAV aerial image analysis method described in the above embodiments.

[0121] Those skilled in the art can clearly understand the specific working process of the above-described system, device, module and unit, and can refer to the corresponding process in the foregoing method embodiments. For brevity, no further description is given here.

[0122] In addition, the functional units in each of the embodiments of the present application can be physically independent of each other, or two or more functional units can be integrated together, and all functional units can be integrated in one processing unit. The integrated functional units can be realized in the form of hardware or software or firmware.

[0123] Those skilled in the art can understand that the integrated functional units, if realized in the form of software and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be essentially embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computing device (such as a personal computer, a server, or a network device) to execute all or part of the steps of the method described in the embodiments of the present application when the instructions are executed. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0124] Alternatively, all or part of the steps of the foregoing method embodiments can be completed by program instruction related hardware (such as a computing device of a personal computer, a server, or a network device), which can be stored in a computer readable storage medium, and when the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the method described in the embodiments of the present application.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that, within the spirit and principle of the present application, the technical solutions recorded in the foregoing embodiments can still be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of the present application.

Claims

1. A method for analyzing aerial images captured by unmanned aerial vehicles, characterized in that, The method includes: Multi-source remote sensing data is acquired, and the multi-source remote sensing data is subjected to time synchronization and spatial alignment processing to obtain preprocessed data; the multi-source remote sensing data includes at least visible light images and infrared images; The preprocessed data is input into a feature extraction model based on a dual-branch Blur-PANet network to extract multi-scale fusion features from the preprocessed data. Based on the multi-scale fusion features, a threat risk assessment is performed on the threat targets in the environment where the UAV is located; When the threat risk meets the preset avoidance conditions, the feature matching algorithm is used to align ground data and aerial data, and obstacle avoidance instructions carrying the UAV obstacle avoidance path are generated collaboratively. After the UAV executes the obstacle avoidance command, the current ground data and aerial data are aligned again using the feature matching algorithm, and combined with LiDAR point cloud data, the target motion trajectory of the UAV is projected onto the geospatial space to generate a motion mask. Transfer learning verification is performed on the static area to obtain the UAV aerial image analysis results.

2. The method according to claim 1, characterized in that, The collection of multi-source remote sensing data includes: The multi-source remote sensing data is acquired using a global exposure CMOS sensor, specifically including: The exposure parameters of the acquired multi-source remote sensing data are dynamically adjusted based on an environmental perception function; the mathematical expression of the environmental perception dynamic function is: E(x,y,z,t)=λ1·T+λ2·O+λ3·W Where E(x,y,z,t) represents the environmental perception assessment value used to adjust the exposure parameters; T represents the terrain features; O represents the obstacle distribution; W represents the meteorological conditions; λ1, λ2, and λ3 represent the dynamic weighting coefficients; and (x,y,z) and t represent the coordinate and time parameters, respectively.

3. The method according to claim 1, characterized in that, The time synchronization and spatial alignment processing of the multi-source remote sensing data includes: The time synchronization deviation of the multi-source remote sensing data is compressed to within a preset duration threshold using timestamp interpolation; and, By fusing edge features of visible light images and thermal radiation features of infrared images from the multi-source remote sensing data using a phase consistency algorithm, the offset of the moving target detection box is controlled within a preset pixel threshold. Specifically, this includes: A pixel coordinate mapping function is established using pre-acquired lens distortion parameters; Based on the pixel coordinate mapping function, the phase consistency features of the visible light image and the infrared image are extracted. By fusing the gradient magnitudes of the visible light image and the infrared image, edge misalignment caused by spatiotemporal offset is eliminated.

4. The method according to claim 1, characterized in that, The backbone network of the feature extraction model uses depthwise separable convolution, and the neck network uses a weighted bidirectional recurrent path. The extraction of multi-scale fusion features from the preprocessed data includes: Adaptive histogram equalization is applied to the visible light image, and non-uniformity correction is performed on the infrared image to generate a training sample set; The backbone network extracts shallow high-resolution features and deep semantic features from the training sample set based on the depthwise separable convolution; The neck network fuses the shallow high-resolution features through the forward path and injects the deep semantic features through the reverse path. The feature contribution is adjusted according to learnable weights, and the features are fused to generate a dual-branch Blur-PANet network feature as a multi-scale fusion feature.

5. The method according to claim 4, characterized in that, The step of adjusting feature contribution according to learnable weights and fusing to generate dual-branch Blur-PANet network features includes: The fusion variables are calculated using a learnable weight fusion formula, the mathematical expression of which is: w = σ(α·LBP + β·OFM) Where w represents the fusion variable; σ represents the activation function; LBP represents the LBP texture feature; OFM represents the optical flow motion feature; α and β represent the corresponding dynamic adjustment coefficients, respectively; During feature extraction, the angular velocity value of the UAV is multiplied by the radius of the inscribed circle of the target, and the displacement vector is calculated in combination with the inter-frame time interval. The sampling position of the deformable convolution kernel is dynamically adjusted according to the displacement vector, so that the receptive field of the deformable convolution kernel is shifted along the target motion direction, and the edge texture of the corresponding region in the feature map is locally reconstructed.

6. The method according to claim 1, characterized in that, The threat risk assessment of threat targets in the environment where the UAV is located, based on the multi-scale fusion features, includes: In the regression task of the multi-scale fused features, a direction-aware loss function is introduced, and energy focusing loss is used to balance positive and negative samples. The threat function value is calculated in real time, specifically including: The sinusoidal difference between the orientation angles of the predicted box and the ground truth box is calculated to quantify the angular deviation. The loss weight is dynamically adjusted in conjunction with the intersection-union ratio of the rotated boxes. At the same time, the energy focusing loss is used to reconstruct the classification task, and an exponential penalty weight is applied to the hard-to-classify samples. The loss contribution of the easy-to-classify samples is reduced by adjusting the focusing factor, so as to optimize the weight ratio of positive and negative samples. Based on the positive and negative samples, a threat assessment model is established that integrates the relative speed, collision time, and angular velocity magnitude of the threat target and the UAV, and includes dynamic risk items and safety margins. The threat assessment model is used to assess the threat risk of the threat target in the environment where the UAV is located, and obtain the threat assessment value.

7. The method according to claim 6, characterized in that, When the threat risk meets the preset avoidance conditions, a feature matching algorithm is used to align ground data and aerial data, and obstacle avoidance instructions carrying the UAV obstacle avoidance path are collaboratively generated, including: When the threat assessment value is lower than the preset security threshold, it is considered that the preset avoidance conditions are met; Calculate the coordinates of the safety envelope boundary point as the obstacle avoidance maneuver point, output the yaw angle and pitch angle control values, and send them to the ground control point; The ground control point aligns ground data and aerial data using the feature matching algorithm, calculates flight curvature based on the digital elevation model, and dynamically adjusts the flight parameters of the UAV. The flight parameters include flight altitude and flight speed.

8. The method according to claim 7, characterized in that, The method further includes: When the UAV deviates from the target flight path, a segmented regression path is dynamically generated using a flight path regression function, specifically including: Based on the horizontal error distance and azimuth deviation between the current position and the target waypoint, the regression heading angle is calculated in real time using the waypoint regression function. When the horizontal error distance exceeds a preset error threshold, the segmented regression path is generated; The segmented regression path includes a first regression path that rapidly approaches the target route projection point at a large tilt angle, a second regression path that calibrates the attitude along the route tangent, and a third regression path that returns to the target route at a constant altitude. The dynamically adjusted flight parameters and the segmented regression path are sent to the UAV as the obstacle avoidance command.

9. The method according to any one of claims 1 to 8, characterized in that, The process involves aligning current ground and aerial data using the feature matching algorithm, combining LiDAR point cloud data, projecting the target trajectory of the UAV onto geospatial data to generate a motion mask, performing transfer learning verification on static areas, and obtaining UAV aerial image analysis results, including: The control points of the ground data and aerial data are extracted by the feature matching algorithm, and the geographic coordinate system of the ground data, the aerial data and the LiDAR point cloud data is unified by the homography matrix to reconstruct the geospatial data. The target motion trajectory output from multi-target tracking is projected onto the geographic space to generate a motion mask. Based on the motion mask, temporary change interference is filtered out, while the change detection results of the static area are retained. The change detection results are input into a pre-trained ShuffleNet fine-tuning learning network model, and the true and false changes in the static region are verified by RGB-LBP fusion features to obtain the analysis results of the UAV aerial image.

10. A drone aerial image analysis system, characterized in that, The system includes: A multi-source acquisition module is used to acquire multi-source remote sensing data and perform time synchronization and spatial alignment processing on the multi-source remote sensing data to obtain pre-processed data; the multi-source remote sensing data includes visible light images, infrared images and / or depth images; The feature fusion module is used to input the preprocessed data into a feature extraction model based on a dual-branch Blur-PANet network to extract multi-scale fusion features from the preprocessed data. The threat assessment module is used to assess the threat risk of targets in the environment where the UAV is located based on the multi-scale fusion features. The collaborative decision-making module is used to align ground data and aerial data using a feature matching algorithm when the threat risk meets the preset avoidance conditions, and collaboratively generate obstacle avoidance instructions carrying the UAV obstacle avoidance path. The result verification module is used to, after the UAV executes the obstacle avoidance command, again use the feature matching algorithm to align the current ground data and aerial data, and combine LiDAR point cloud data to project the target motion trajectory of the UAV onto the geospatial space to generate a motion mask, perform transfer learning verification on the static area, and obtain the UAV aerial image analysis results.

Citation Information

Patent Citations

  • Changed building rapid detection method based on new and old time phase images of unmanned aerial vehicle

    CN111898477A

  • Unmanned aerial vehicle intelligent inspection method for energy facility inspection

    CN119693824A

  • Unmanned aerial vehicle multi-source sensing fusion AI real-time intelligent guidance and adaptive obstacle avoidance method

    CN120178906A

  • Guidance optimization method for AI self-adaptive unmanned aerial vehicle for high-speed long voyage

    CN120523017A

  • Methods for target detection based on visible cameras, infrared cameras, and lidars

    US20240355105A1