Unmanned aerial vehicle autonomous inspection ground-imitated flight method
By incorporating multi-source sensors and a lightweight PSPNet-MobileNetV2 model for pixel-level semantic segmentation and laser point cloud fusion, the problem of inaccurate obstacle avoidance judgment by UAVs in complex terrain is solved, achieving highly safe and stable autonomous inspection.
Patent Information
- Application Number
- CN202511218414.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing drone terrain-following flight technology cannot effectively integrate terrain material semantic information and structural features in complex terrain scenarios, resulting in inaccurate obstacle avoidance judgments and affecting the safety and reliability of inspection operations.
Equipped with multi-source sensors to collect terrain data in real time, the system uses a lightweight PSPNet-MobileNetV2 hybrid architecture model for pixel-level semantic segmentation, outputs terrain material classification labels, and integrates laser point cloud spatial topology data to generate dynamic control commands. Combined with dual-end queue constraints, it achieves smooth cross-modal switching and closed-loop control.
It enables precise perception and adaptive flight control of UAVs in complex terrain, improves the safety and stability of inspections, solves the problem of inaccurate obstacle avoidance judgment, and meets the requirements of high-precision and high-reliability inspections.
Smart Images

Figure CN121165769A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of unmanned aerial vehicle autonomous control, in particular to a method for unmanned aerial vehicle autonomous inspection and ground-following flight. BACKGROUND
[0002] Currently, the method for unmanned aerial vehicle autonomous inspection and ground-following flight is mainly applied to complex terrain scenes such as power inspection, photovoltaic power station inspection, and mountain environment monitoring. In the above application scenarios, the ground-following flight technology is one of the key technologies to ensure the quality and safety of unmanned aerial vehicle inspection operations. The core principle is to dynamically adjust the flight height of the unmanned aerial vehicle to adapt to the changes in the terrain, ensuring that the unmanned aerial vehicle maintains a relatively constant height difference from the ground or the target inspection object, thereby ensuring the collection accuracy and effectiveness of inspection data (such as laser point cloud and image data) and avoiding collision risks of the unmanned aerial vehicle due to sudden changes in the terrain. With the large-scale application of unmanned aerial vehicle inspection technology in the fields of energy and geological exploration, complex terrain scenes (such as mountain vegetation areas, water periphery, and dense power transmission line areas) have higher requirements for the terrain adaptability, scene compatibility, and decision-making intelligence of the ground-following flight technology.
[0003] In the prior art, the unmanned aerial vehicle ground-following flight scheme mainly relies on terrain elevation data to achieve height control. There are two main technical paths: one is based on real-time collected elevation data, such as directly obtaining terrain elevation information through a laser radar or constructing a digital elevation model (DEM) in real time to adjust the flight height. For example, Chinese patent application No. CN116627164B “Unmanned aerial vehicle ground-following flight control method and system based on terrain height” constructs a DEM model in real time to calculate the terrain elevation difference and dynamically adjusts the flight height of the unmanned aerial vehicle. The other is based on a pre-constructed elevation model, such as pre-generating variable-height waypoints based on a digital surface model (DSM). For example, Chinese patent application No. CN111966129B “Photovoltaic inspection unmanned aerial vehicle and its ground-following flight method” pre-plans the flight trajectory and height of the photovoltaic inspection area based on the DSM model to achieve global unified ground-following flight control. In addition, in the power line inspection scene, existing technologies mostly plan the line path based on the geometric elevation of the line; in the photovoltaic inspection scene, they mostly pre-generate waypoints with fixed height intervals based on the DSM. The above schemes can achieve basic ground-following functions in scenes with relatively simple terrain conditions and single terrain features.
[0004] However, the existing hover technology has a core defect: relying only on terrain geometric elevation data for flight height adjustment, without incorporating terrain material semantic information (such as water, sand, tree canopy, bare rock, power lines, buildings, and other different material categories) and terrain structure features (such as slope, vegetation NDVI value, bare rock spectral reflectance, power line sag gradient, etc.) into the flight decision logic, it cannot dynamically switch flight modes according to real-time perception of terrain properties. This core defect leads to the problem that in complex inspection scenarios, the UAV cannot adapt the optimal flight strategy to the terrain characteristics, thereby causing the key problem of obstacle avoidance judgment error - for example, in water scenarios, relying only on geometric elevation data makes it difficult to distinguish between water surface specular reflection and real ground, leading to distortion of laser radar point cloud, and the UAV may misjudge the height and approach the water surface; in vegetation canopy scenarios, it is difficult to identify canopy thickness and NDVI characteristics, and it is easy to set too low a flight height by mistakenly considering the top of the canopy as the ground, causing the UAV to collide with the canopy; in power line dense areas, it is difficult to perceive the sag gradient of the power line, and only relying on the preset safety distance to avoid obstacles may miss the lowest point of the sag and cause collision risk. The above obstacle avoidance judgment error problem directly threatens the safety of UAV inspection operations, and is the core bottleneck restricting the reliable application of existing hover technology in complex terrain scenarios.
[0005] In summary, although the existing UAV hover technology can meet the basic height tracking needs in simple terrain scenarios, it has significant deficiencies in terrain semantic understanding and dynamic decision-making, especially lacking the fusion of terrain material semantic and structure features, leading to the prominent problem of insufficient obstacle avoidance reliability in complex inspection scenarios. Therefore, there is an urgent need in the current field for an adaptive hover technology that can integrate terrain material semantic recognition and structure feature analysis to solve the obstacle avoidance judgment error problem caused by relying only on geometric elevation data in existing technology, thereby improving the safety, stability, and intelligence level of UAV autonomous inspection in complex terrain scenarios, and meeting the actual needs of high-precision, high-reliability inspection operations in the fields of power, photovoltaic, and mountain monitoring. SUMMARY
[0006] The unmanned aerial vehicle autonomous inspection ground-imitating flight method aims to make up for the shortage of the prior art, and can realize real-time collection of terrain data through a multi-source sensor carried on the unmanned aerial vehicle, perform pixel-level semantic segmentation on the data by using a light-weight PSPNet-MobileNetV2 hybrid architecture model to output terrain material classification labels, fuse the semantic segmentation result and laser point cloud spatial topology data to extract terrain structure features, match preset flight modes according to the combination of the terrain material classification labels and the terrain structure features, and generate dynamic control instructions, realize cross-mode smooth switching through double-ended queue constraint, form a closed-loop control in combination with sensor feedback, and adaptively adjust the flight strategy of the unmanned aerial vehicle according to different complex terrains, accurately adapt to terrain properties, effectively avoid inaccurate obstacle avoidance judgment, and improve the safety, stability and intelligent level of the unmanned aerial vehicle in autonomous inspection ground-imitating flight in a complex scene.
[0007] The unmanned aerial vehicle autonomous inspection ground-imitating flight method aims to make up for the shortage of the prior art, and can realize real-time collection of terrain data through a multi-source sensor carried on the unmanned aerial vehicle, perform pixel-level semantic segmentation on the data by using a light-weight PSPNet-MobileNetV2 hybrid architecture model to output terrain material classification labels, fuse the semantic segmentation result and laser point cloud spatial topology data to extract terrain structure features, match preset flight modes according to the combination of the terrain material classification labels and the terrain structure features, and generate dynamic control instructions, realize cross-mode smooth switching through double-ended queue constraint, form a closed-loop control in combination with sensor feedback, and adaptively adjust the flight strategy of the unmanned aerial vehicle according to different complex terrains, accurately adapt to terrain properties, effectively avoid inaccurate obstacle avoidance judgment, and improve the safety, stability and intelligent level of the unmanned aerial vehicle in autonomous inspection ground-imitating flight in a complex scene. S100, real-time collection of terrain data through a multi-source sensor, the multi-source sensor including a light-weight laser radar, a multi-spectral camera and an ultrasonic radar, wherein the ranging accuracy of the light-weight laser radar is ±3cm, the multi-spectral camera covers 5 spectral bands, and the obstacle avoidance accuracy of the ultrasonic radar is ±5cm; S200, pixel-level semantic segmentation of the real-time collected terrain data based on a light-weight PSPNet-MobileNetV2 hybrid architecture model, to output terrain material classification labels, wherein the terrain material classification labels include water surface, sand, tree crown, bare rock, power transmission line and building; S300, fusion of the semantic segmentation result and laser point cloud spatial topology data to extract terrain structure features, wherein the terrain structure features include cliffs with a slope greater than 60°, valleys with continuous negative terrain, vegetation canopy with NDVI greater than 0.7, and bare rock with spectral reflectance greater than 0.4; S400, combination matching of preset flight modes according to the terrain material classification labels and the terrain structure features to generate dynamic control instructions; the flight modes include: when the water surface is identified and the reflectivity is higher than a threshold value, a low disturbance mode is enabled, the height of the unmanned aerial vehicle is controlled to be reduced to 3m, and the downward-looking laser is turned off; when the tree crown is identified and the NDVI is greater than 0.7, a penetrating scanning mode is enabled, the height is increased to 2m above the crown layer, and multi-spectral penetrating imaging is started; when the bare rock is identified and the slope is greater than 30°, a high-precision reconstruction mode is enabled, the tilt angle of the gimbal is adjusted to-45°, and the laser scanning frequency is increased to 100Hz; when a power transmission line dense area is identified and there is an arc drop gradient, a line-imitating obstacle avoidance mode is enabled, flight is performed along the electromagnetic field strength gradient, and the lateral obstacle avoidance distance is set to 5m. S500, constrain the height deviation of adjacent frames by a double-ended queue to be less than 20%, to realize smooth switching across flight modes; S600, evaluate the execution effect of the control instruction in real time during flight, correct the decision deviation according to the sensor feedback data, and form a closed-loop control.
[0008] Further, the semantic segmentation model in step S200 removes 20% redundant layers through channel pruning, and adopts INT8 quantization to compress the model volume to 0.8MB; The model introduces a FocalLoss function for dynamic interference, and its expression is: ; Wherein, represents the predicted probability value of the current pixel point belonging to a specific terrain material category, represents the weight adjustment coefficient of the current terrain material category, represents the adjustment factor of the focal difficult sample.
[0009] Further, the extraction of terrain structure features in step S300 includes: The point cloud sampling density in the valley and cliff area is forced to be increased by 50%, and the three-dimensional obstacle avoidance path is generated; Based on the semantic distribution of the power transmission line, the sag minimum point is predicted, and a counterclockwise path with a radius of 10m is generated to avoid the electromagnetic blind area.
[0010] Further, the dynamic switching logic of the flight mode in step S400 includes: When the tree crown ratio exceeds 40%, automatically switch from the ground mode to the penetration scanning mode, and simultaneously increase the vegetation filtering strength of the laser radar by 2 times; When a sudden increase in wind speed is detected, the distance of the wire-like inspection is shortened to 10km and the standby barometric height mode is switched.
[0011] Further, the smooth switching algorithm across flight modes in step S500 constrains the height change by the following formula: ; Wherein, represents the height change rate of adjacent frames, represents the current flight height measurement value of the UAV, represents the flight height measurement value of the UAV at the previous time.
[0012] Further, the asynchronous fusion method of the multi-source sensor includes: The sampling frequency of the laser radar data is 10Hz, and the sampling frequency of the visual data is 30Hz; Through timestamp alignment, the calculation overhead is reduced by 35%.
[0013] Further, the closed-loop control feedback mechanism in step S600 includes: Real-time monitoring of the actual distance to the obstacle by ultrasonic radar, and when the deviation from the preset obstacle avoidance threshold exceeds ±10%, the flight mode is re-matched; In the power line dense area, dynamically calibrate the flight heading angle based on the direction of the electromagnetic field strength gradient, and control the horizontal error within ±0.5m.
[0014] Further, the joint output of terrain material classification and structural features includes: Synchronous generation of 6 types of material labels and 4 types of structural feature labels; The model inference time is less than 50ms, and it is deployed on the NVIDIA Jetson TX2 embedded platform.
[0015] Further, the optimization deployment in resource-constrained scenarios includes: Laser radar and visual data are aligned by timestamp to realize asynchronous fusion; Under the memory limit of embedded devices, model pruning and quantization compression are used to occupy computing resources.
[0016] Further, the flight mode decision matrix is constructed according to the following logic: When the recognized terrain semantic type is water surface, and its structural feature is flat and reflectivity is higher than the preset threshold, the low disturbance flight mode is enabled; When the recognized terrain semantic type is tree crown, and its structural feature satisfies NDVI value greater than 0.7 and crown thickness exceeds the set standard, the penetrating scanning flight mode is enabled; When the recognized terrain semantic type is bare rock, and its structural feature is slope greater than 30°, the high-precision reconstruction flight mode is enabled; When the recognized terrain semantic type is a power line dense area, and its structural feature has sag gradient, the line-imitating obstacle avoidance flight mode is enabled.
[0017] Compared with the prior art, the unmanned aerial vehicle autonomous inspection and ground flight method has the following beneficial effects: The present application can realize accurate perception and adaptive flight control of complex terrain, so that the unmanned aerial vehicle can adapt to the optimal flight strategy for different terrain properties during the inspection process, effectively solve the problem of inaccurate obstacle avoidance caused by the fact that the prior art only relies on terrain geometric elevation data for flight height adjustment and does not include terrain material semantic information and terrain structure features in flight decision logic, improve the safety and reliability of autonomous inspection of the unmanned aerial vehicle in complex terrain scenes, and meet the actual needs of high-precision and high-reliability ground simulation flight in the fields of power inspection and mountain environment monitoring.
[0018] The present application can realize accurate perception and adaptive flight control of complex terrain, so that the unmanned aerial vehicle can adapt to the optimal flight strategy for different terrain properties during the inspection process, effectively solve the problem of inaccurate obstacle avoidance caused by the fact that the prior art only relies on terrain geometric elevation data for flight height adjustment and does not include terrain material semantic information and terrain structure features in flight decision logic, improve the safety and reliability of autonomous inspection of the unmanned aerial vehicle in complex terrain scenes, and meet the actual needs of high-precision and high-reliability ground simulation flight in the fields of power inspection and mountain environment monitoring.
[0019] Other advantages, objects, and features of the present application will be apparent to those skilled in the art from the following specification, in some degree of respect, based on the study of the following, or can be taught from the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0021] Figure 1 The operation flowchart of the present application; Figure 2A step block diagram of the present application; Figure 3 A multi-sensor fusion decision control system architecture diagram of the present application. DETAILED DESCRIPTION
[0022] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined application purposes, the specific embodiments, structures, features and effects according to the present application are described in detail below in combination with the drawings and preferred embodiments.
[0023] Embodiment one: as shown in Figure 2 and Figure 3 , this embodiment aims at the high safety and high adaptability requirements of unmanned aerial vehicle (UAV) ground effect flight in the complex terrain (including water area, vegetation canopy, bare rock cliff and dense power transmission line area) of mountain 220kV power transmission line inspection scene, and specifically implements a method of UAV autonomous inspection ground effect flight. The method realizes real-time terrain data collection by UAV carrying multi-source sensors, completes terrain material classification through a lightweight semantic segmentation model, extracts terrain structure features by fusing laser point cloud topology data, matches dynamic flight modes according to the combination of material and structure, and guarantees flight safety and inspection accuracy through smooth switching algorithm and closed-loop control mechanism. This embodiment elaborates the specific operation process, parameter design basis and working principle of each technical link, verifies the feasibility and superiority of the method in the complex mountain power transmission line inspection scene, effectively solves the obstacle avoidance judgment error problem caused by relying only on terrain elevation data in the prior art, and improves the intelligence and reliability of UAV inspection.
[0024] Scene and system deployment: The application scene of this embodiment is selected as the mountain 220kV power transmission line inspection area of a provincial power company, which has a span of about 20km and covers various complex terrains: 3 intermountain water areas with an area of more than 0.5km², a dense vegetation canopy area (mainly pine forest) with a length of about 5km, 2 sections of bare rock cliff with a slope of more than 60°, and a mountain valley span section with obvious sag gradient of power transmission line (the maximum sag difference is up to 8m). In this scene, power transmission line inspection needs to meet three core requirements: "avoiding collision with terrain / obstacles", "ensuring line detail collection accuracy" and "adapting to complex weather interference (such as sudden gust) ". The traditional ground effect flight scheme based on pre-generated DSM is prone to problems such as "water surface misjudgment leading to low altitude approach", "canopy collision risk" and "sag minimum point missing", which need to be realized by the dynamic adaptive control method of the present application.
[0025] The UAV system deployment of this embodiment: 1. UAV platform selection: Select a multi-rotor industrial UAV with a maximum load of 1.5 kg and a continuous time of 40 minutes. It has a multi-sensor integrated interface and an open-source flight control system (supports custom control instruction access). The load capacity and endurance characteristics of this platform can meet the long-term mounting requirements of multi-source sensors, and the open-source flight control system provides hardware support for real-time issuance of dynamic control instructions.
[0026] 2. Multi-source sensor installation and parameter design: Lightweight laser radar: installed at the central position below the UAV body, using TOF (Time of Flight) ranging principle, ranging range 0.5-100m, design ranging accuracy ±3cm. The accuracy design basis is: in power line inspection, the safety distance between UAV and line needs to be controlled within 5m, and the ranging error of ±3cm can ensure that the calculation deviation of laser point cloud to line position and terrain elevation does not affect the safety decision; at the same time, lightweight design (weight less than 300g) can avoid excessive occupation of UAV load, and ensure the endurance time.
[0027] Multispectral camera: installed below the UAV gimbal, at a 15° angle with the laser radar (to avoid mutual obstruction of the field of view), covering 5 spectral bands, including 450-520nm (blue light), 520-600nm (green light), 630-690nm (red light), 760-900nm (near-infrared), and 900-1700nm (short-wave infrared). The core purpose of this band selection is: red and near-infrared bands are used to calculate the NDVI (Normalized Difference Vegetation Index) of vegetation, to judge the density of vegetation canopy; the short-wave infrared band can penetrate part of the vegetation leaves to obtain the terrain information under the canopy; the blue and green bands are used to improve the classification accuracy of water surface, bare rock and other materials.
[0028] Ultrasonic radar: a total of 4 are installed, respectively deployed around the UAV body (front, back, left, right), detection range 0.1-5m, avoidance accuracy design ±5cm. The deployment scenario of this sensor focuses on "short-distance sudden obstacle detection", such as protruding rocks and low shrubs encountered during low-altitude flight, and the ±5cm accuracy can ensure that the deviation of feedback data does not result in avoidance decision errors when the distance to the obstacle is within 1m.
[0029] 3. Embedded computing platform deployment: NVIDIA Jetson TX2 embedded platform is used, installed inside the UAV body, connected with multi-source sensors through USB3.0 interface, realizing real-time data reception and processing; at the same time, it communicates with the flight control system through UART interface, completing the issuance of control instructions. The basis for choosing this platform is that its integrated GPU (Pascal architecture, 256 cores) can meet the real-time inference requirements of lightweight semantic segmentation models, and the power consumption is only 15W, which will not excessively consume the battery power of the UAV, meeting the energy constraints of long-term inspection.
[0030] Multi-source sensor data acquisition and asynchronous fusion (corresponding to step S100): The core goal of this link is to obtain comprehensive and real-time terrain data through multi-source sensors, and to solve the problem of data redundancy and processing delay caused by the difference in sampling frequency of different sensors through asynchronous fusion method, to provide high-quality input data for subsequent semantic segmentation and feature extraction.
[0031] 3.1 Data acquisition process of each sensor 1. Lightweight laser radar data acquisition: The laser radar emits laser beams at a sampling frequency of 10 Hz, generating about 100,000 laser points per frame, each containing three-dimensional coordinates (x, y, z, based on the UAV coordinate system) and reflection intensity values. The basis for setting the sampling frequency to 10 Hz is that the time scale of mountain terrain fluctuation is relatively slow (the flight speed of the UAV is about 5 m / s, and 10 Hz sampling can ensure that the terrain elevation data is obtained once every 0.1 s, covering a flight distance of 50 cm, which is sufficient to capture the macro fluctuations of the terrain); If the sampling frequency is too high (such as 20 Hz), it will double the amount of point cloud data per frame, increasing the computational burden of the embedded platform, and the improvement of terrain details is limited; If it is too low (such as 5 Hz), it may miss the elevation changes of local steep terrain (such as cliff edges), causing collision risks. During the acquisition process, the laser radar transmits the point cloud data in real time to the embedded platform through the USB3.0 interface, and each frame of data is attached with a timestamp (accurate to the millisecond level) for subsequent data fusion.
[0032] 2. Multi-spectral camera data acquisition: The multi-spectral camera takes terrain images at a sampling frequency of 30 Hz, with a resolution of 1280x720 pixels per frame, and each pixel corresponds to the reflectivity value of 5 spectral bands (output after internal calibration). The design logic of 30 Hz sampling frequency is: visual data needs to capture more detailed terrain material texture (such as the fine structure of power lines, the reflection texture of water surface), a higher frame rate can avoid image blur caused by UAV movement; At the same time, the frequency ratio of 30 Hz to laser radar 10 Hz is 3:1, which is convenient for subsequent alignment of time stamps to realize "3 frames of visual data matching 1 frame of laser point cloud data", which not only ensures the integrity of visual details, but also does not produce too much redundant data. During acquisition, the multi-spectral camera synchronously records the timestamp of each frame of image (calibrated based on the same clock source as the laser radar timestamp), and packs the image data and reflectivity data for transmission to the embedded platform.
[0033] 3. Ultrasonic radar data collection: Four ultrasonic radars detect the distance of surrounding obstacles in real time at a sampling frequency of 20 Hz. When the detected distance is less than a preset threshold (initially set to 3 m), a high-frequency warning is triggered (the sampling frequency is increased to 50 Hz), ensuring a quick response to close-range obstacles. The base sampling frequency of 20 Hz can meet the close-range monitoring needs of the UAV during normal flight, and the warning frequency of 50 Hz is aimed at sudden obstacles (such as protruding shrubs encountered during low-altitude flight), leaving enough reaction time for obstacle avoidance decisions (when the UAV flies at a speed of 5 m / s, the 50 Hz sampling can update the obstacle distance every 0.02 s, ensuring that the obstacle avoidance process of "detection-decision-execution" is completed within 0.1 s).
[0034] 3.2 Multi-source sensor asynchronous fusion implementation: Due to the differences in sampling frequencies of the laser radar (10 Hz), multispectral camera (30 Hz), and ultrasonic radar (20 Hz / 50 Hz), if the "synchronous collection-synchronous processing" mode is adopted, the high-frequency sensor data will wait for the low-frequency sensor data, resulting in data accumulation and processing delay. Therefore, this embodiment adopts the "timestamp alignment + asynchronous fusion" strategy, and the specific process is as follows: 1. Timestamp calibration: During the system startup phase, the timestamps of the three sensors are uniformly calibrated through the clock module of the embedded platform, ensuring that the timestamp deviation of all sensors is less than 1 ms (to avoid data matching errors caused by different timestamps). The calibration method is as follows: taking the system clock of the embedded platform as the reference, a synchronous trigger signal is sent to each sensor, and the difference between the timestamp feedback of each sensor and the system clock is recorded. In the subsequent collection process, the timestamps are unified through difference compensation.
[0035] 2. Data alignment and selection: taking the sampling timestamp of the laser radar as the "reference time axis" (since the laser radar data is the core source of terrain elevation and spatial topology, it needs to be used as the fusion reference), the multispectral camera data and ultrasonic radar data are aligned: For the 30 Hz data of the multispectral camera, every 3 frames of images correspond to 1 frame of laser radar data (the 3 frames of images with the smallest timestamp difference), and the reflectivity data of the 3 frames of images is fused through the bilinear interpolation algorithm to generate 1 frame of "fusion vision data" matching the laser radar timestamp, avoiding the noise influence of single-frame vision data.
[0036] For the ultrasonic radar data, all radar data with a timestamp difference of less than 50 ms from the laser radar is taken, and the average value is calculated as the "effective obstacle distance data" at that moment, which not only ensures the real-time nature of the data, but also reduces the measurement noise through averaging filtering.
[0037] 3. Fusion efficiency optimization: Through the above timestamp alignment strategy, only the "effective data on the reference timeline" is processed subsequently, avoiding the calculation of redundant data that is not aligned (such as images in a multispectral camera that do not match the lidar timestamp). Ultimately, this results in a 35% reduction in computational overhead. The principle of this optimization is that unaligned data, if involved in processing, will result in repeated calculations (such as the same terrain area being repeatedly analyzed by multiple frames of visual data), while aligned data only retains visual and obstacle information corresponding to terrain elevation, significantly reducing the computational load on the embedded platform and ensuring the real-time performance of subsequent semantic segmentation and feature extraction.
[0038] Pixel-level semantic segmentation processing of terrain data (corresponding to step S200): This step uses a lightweight PSPNet-MobileNetV2 hybrid architecture model to perform pixel-level semantic segmentation on the asynchronously fused terrain data (lidar reflectance intensity data + multispectral fused visual data), outputting 6-class terrain material classification labels (water, sand, tree canopy, bare rock, power line, and building). This provides a material category basis for subsequent terrain structure feature extraction. This step requires a detailed explanation of the model architecture design, optimization process, loss function application, and semantic segmentation implementation process.
[0039] 4.1 Selection and advantages of PSPNet-MobileNetV2 hybrid architecture; The PSPNet-MobileNetV2 hybrid architecture is chosen as the semantic segmentation model. The core basis for this choice is that this architecture balances "multi-scale terrain feature capture capability" and "lightweight deployment characteristics". The specific design logic is as follows: 1. Role of PSPNet's pyramid pooling module: In the mountain power line inspection scene, the scale of terrain materials varies significantly (such as large-area water, small-scale power lines, and medium-scale tree canopies). Traditional semantic segmentation models (such as U-Net) are prone to "incomplete segmentation of large-scale targets" and "missing small-scale targets" when processing multi-scale targets. The pyramid pooling module of PSPNet uses 4 different scale pooling operations (such as 1x1, 2x2, 3x3, and 6x6 pooling) to capture global terrain features (such as the distribution of the entire water area), medium-scale features (such as the canopy of a single tree), and small-scale features (such as a single power line). Through feature fusion, multi-scale information is integrated, ensuring the segmentation accuracy of the 6-class material labels, especially the recognition rate of small-scale targets such as power lines, which is improved to over 95% (verified by previous sample testing).
[0040] 2. Lightweight support of MobileNetV2: The memory (8 GB) and computing power (1.3 TFLOPS) of the embedded platform (NVIDIA Jetson TX2) carried by the UAV are limited. The traditional PSPNet uses ResNet50 as the backbone network, and the model size is more than 100 MB, and the inference time is more than 200 ms, which cannot meet the real-time processing requirements (when the UAV flies at a speed of 5 m / s, it needs to process a frame of data every 50 ms to avoid lag). MobileNetV2 uses "depth separable convolution + linear bottleneck structure", which decomposes the standard convolution into depth convolution (channel-wise convolution) and point convolution (1x1 convolution), reducing the number of parameters by more than 90% while ensuring feature extraction capability; at the same time, the linear bottleneck structure solves the feature degradation problem caused by depth separable convolution, ensuring model accuracy. Using MobileNetV2 as the backbone network of PSPNet can achieve model lightweight while retaining multi-scale segmentation capability.
[0041] 4.2 Optimization process of semantic segmentation model (channel pruning and INT8 quantization): This embodiment performs "channel pruning + INT8 quantization" double optimization on the PSPNet-MobileNetV2 hybrid architecture, and the specific process is as follows: 1. Channel pruning: remove 20% redundant layers: Pruning premise: Through "channel contribution analysis" in the model training process, the contribution value of each channel in each convolution layer to the final semantic segmentation accuracy is calculated (the contribution value is calculated by the mutual information between the channel output feature map and the label, and the lower the mutual information, the smaller the contribution).
[0042] Pruning process: First, freeze the core channels of the backbone network (MobileNetV2) and the pyramid pooling module (the top 80% of the contribution degree), to avoid affecting the core feature extraction capability of the model; then remove the channels with a contribution degree of less than 20% in the transition convolution layer (the convolution layer connecting different modules) of the model, the mutual information between the output feature map of these channels and the label is less than 0.1 (after testing, the model accuracy only decreases by 0.5% after removal, which is within the acceptable range); finally, fine-tune the pruned model to update the weight parameters of the remaining channels through back propagation, to make up for the slight accuracy loss caused by pruning.
[0043] Pruning effect: The number of parameters of the model is reduced from 2.5M before pruning to 2.0M, and the calculation amount is reduced by 20%, laying a foundation for subsequent quantization compression.
[0044] 2. INT8 quantization: compress the model size to 0.8 MB: Quantization principle: The weight parameters and activation values of traditional deep learning models are stored in 32-bit floating-point numbers (FP32), which occupies a large amount of memory. INT8 quantization maps FP32 data to 8-bit integers (INT8) through a linear quantization formula (INT8 value = (FP32 value - mean value) / scaling factor) to achieve data conversion, while ensuring that the accuracy loss of the quantized model is controllable through "calibration set calibration".
[0045] Quantization process: First, select 1000 frames of terrain data of the mountain power line inspection scene as the calibration set, covering all 6 types of terrain materials. Then run the FP32 model on the calibration set, calculate the distribution range (maximum value, minimum value, mean value) of each layer weight and activation value, and calculate the optimal scaling factor (to ensure that INT8 data can cover the effective range of FP32 data, and reduce truncation error). Finally, convert the FP32 model to the INT8 model based on the scaling factor, and verify the accuracy of the quantized model (the mIOU (mean intersection over union) of semantic segmentation decreases from 89.2% of FP32 to 88.7% of INT8, with an accuracy loss of only 0.5%).
[0046] Quantization effect: The model size is compressed from 8MB (FP32) to 0.8MB (INT8) after pruning, reducing memory usage by 90%. The inference time is reduced from 80ms after pruning to 45ms, meeting the real-time processing requirements of embedded platforms (inference time < 50ms).
[0047] 4.3 Application of FocalLoss function and implementation of semantic segmentation: In the mountain power line inspection scene, the pixel distribution of terrain materials is significantly unbalanced: the pixel ratio of water and power lines is only 5%-8%, while the pixel ratio of bare rock and ground is more than 60%. If the traditional cross-entropy loss function is used, the model will tend to predict the majority class (bare rock, ground), resulting in low segmentation accuracy of the minority class (power line, water). Therefore, this embodiment introduces the FocalLoss function to solve the class imbalance problem, and the specific application process is as follows: 1. Parameter design and significance of FocalLoss function: The expression of FocalLoss function is: , is the predicted probability value of the current pixel belonging to a specific terrain material class, which is calculated by the output layer (Softmax layer) of the semantic segmentation model, with a value range of [0, 1]. For example, the probability of a pixel being predicted as "power line" by the model is , the probability of being predicted as "bare rock" is , then the of the pixel is
[0048] is the weight adjustment coefficient of the current terrain material category, which is set according to the distribution proportion of pixels in each category. The purpose is to balance the loss contribution of the majority and the minority. In this embodiment, the value of is: transmission line (0.7), water surface (0.6), tree crown (0.5), bare rock (0.3), sandy land (0.4), and building (0.5). The basis for the value is: the transmission line and the water surface are minority categories, and a higher value is given to increase their loss weight; the bare rock is a majority category, and a lower value is given to reduce its loss weight, so as to avoid submerging the loss signal of the minority.
[0049] is the adjustment factor for focusing on difficult-to-classify samples, and the value is 2 (after preliminary testing, the focusing effect of the model on difficult-to-classify samples such as fuzzy transmission line pixels and bare rock pixels similar to water surface reflection is optimal). For easy-to-classify samples (close to 1, such as clear bare rock pixels), close to 0, the loss value is suppressed; for difficult-to-classify samples (close to 0.5, such as fuzzy transmission line pixels), close to 0.25, the loss value is amplified, and the model will preferentially learn the features of the difficult-to-classify samples.
[0050] 2. Calculation of FocalLoss function and model training: The loss calculation process: taking a "fuzzy transmission line pixel" in a certain frame of terrain data as an example, the true label of the pixel is "transmission line", and the model predicts that the "transmission line" is , , . Substituting the formula to calculate: If the pixel is mispredicted as "bare rock" , the loss value is: A higher loss value will guide the model to adjust the parameters and reduce such misjudgments.
[0051] Model training process: the stochastic gradient descent (SGD) optimizer is used, the learning rate is initially set to 0.001, and is attenuated to 0.1 of the original every 10 epochs; the training data set contains 5000 frames of labeled data of mountain transmission line inspection scenes (each frame of data is labeled with 6 material labels), the training round is 100 epochs, and the model converges stably on the validation set (fluctuation less than 0.1%) until the mIOU is stable.
[0052] 3. Semantic segmentation result output: After the INT8 quantization model is deployed on the NVIDIA Jetson TX2 embedded platform, the terrain data fused asynchronously (each frame of data contains a laser radar reflectivity map and a multispectral fusion map) is processed in real time: the model first normalizes the input data (scales the pixel value to [0, 1]), then extracts the basic features through the MobileNetV2 backbone network, fuses the multi-scale features through the PSPNet pyramid pooling module, and finally outputs the material category probability of each pixel through the Softmax layer. The class with the highest probability is selected as the final material label of the pixel to generate a "terrain material semantic segmentation map" (each pixel corresponds to one of the six labels) with the same resolution as the input data, and the semantic segmentation result is stored in the cache of the embedded platform for subsequent terrain structure feature extraction.
[0053] Terrain structure feature extraction (corresponding to step S300): The core of this step is to fuse the "terrain material semantic segmentation result" and the "laser point cloud spatial topology data" to extract four key terrain structure features (cliff with a slope > 60°, valley with continuous negative terrain, vegetation canopy with NDVI > 0.7, and bare rock with spectral reflectance > 0.4), providing a structural dimension judgment basis for subsequent flight mode matching. This step needs to elaborate the data fusion method, feature extraction principle, and enhancement processing of special areas.
[0054] 5.1 Fusion method of semantic segmentation result and laser point cloud: The laser point cloud spatial topology data contains the three-dimensional coordinates (x, y, z) and reflectivity of each laser point, which can reflect the elevation change and spatial distribution of the terrain; the semantic segmentation result provides the material category of each pixel, and the fusion of the two needs to achieve "spatial correspondence between point cloud and pixel", the specific process is as follows: 1. Coordinate mapping: Obtain the real-time position (latitude, longitude, elevation) and attitude (roll angle, pitch angle, heading angle) of the UAV through the POS system (Position and Attitude System) of the UAV, establish the conversion relationship between the laser radar coordinate system and the multispectral camera coordinate system (based on the calibration parameters during sensor installation, such as relative position offset, angle offset), and then convert the three-dimensional coordinates of each laser point (laser radar coordinate system) to the image coordinate system of the multispectral camera (pixel coordinates), achieving one-to-one correspondence between "laser point-pixel", i.e. each laser point can be matched to a pixel in the semantic segmentation map, thereby obtaining the terrain material category corresponding to the laser point.
[0055] 2. Data association: The matched "laser point (three-dimensional coordinates, reflection intensity) + material label" data is stored in a structured data format (such as PCD format), and each laser point contains 7 attributes: x (east coordinate), y (north coordinate), z (elevation), intensity (reflection intensity), material (material label), ndvi (NDVI value, only for vegetation-related material calculation), reflectivity (spectral reflectance, only for bare rock material calculation). This structured data not only retains the spatial topological information of the laser point cloud, but also integrates the material category and key spectral parameters, providing a complete data basis for structural feature extraction.
[0056] 5.2 Key topographic structural feature extraction principle: Based on the above fused structured data, the following method is used to extract 4 types of topographic structural features: 1. Cliff extraction with slope > 60°: Slope calculation principle: Slope is the degree of inclination of the terrain surface at any point, defined as the elevation rate of change at that point, and the calculation formula is: Where is the elevation difference between adjacent laser points, is the horizontal distance between adjacent laser points (calculated from x, y coordinates: .
[0057] Extraction process: First, divide the fused laser point cloud into 5m x 5m grids (grid size is set according to terrain resolution requirements, 5m grid can balance macro slope and calculation efficiency), calculate the average elevation of all laser points in each grid; then select adjacent grids (such as the 4 adjacent grids above and below), calculate the average elevation difference and horizontal distance between the current grid and the adjacent grid, to get the slope value of the grid; finally, mark the grid with a slope value > 60° as a "cliff area" and record the laser point cloud data of the area (for subsequent obstacle avoidance path generation).
[0058] Extraction basis: Areas with a slope > 60° belong to steep terrain, and if the UAV flies at a regular height, it is easy to collide due to terrain changes, so special flight modes (such as increasing height or generating an obstacle avoidance path) are needed to deal with it, so this feature needs to be accurately extracted.
[0059] 2. Valley extraction of continuous negative terrain: Negative terrain definition: Negative terrain refers to areas with lower elevations than surrounding terrain, and valleys are typical continuous negative terrain (characterized by extending in a certain direction, with continuous multiple grids having lower elevations than the two adjacent grids).
[0060] Extraction process: Firstly, calculate the elevation difference between each 5m x 5m grid and the average elevation of its 8 adjacent grids If , the grid is a negative terrain grid; then use the "region growing algorithm" to merge the adjacent (up, down, left, right) and grids into a negative terrain region; finally, filter out the continuous negative terrain region with a length > 50m (along the extension direction) and a width > 10m, and mark it as a "valley region".
[0061] Extraction basis: The terrain of the valley region is complex, and there are low shrubs, rocks and other obstacles, and the sag of the power transmission line in the valley region is usually large, so it is necessary to increase the sampling density of the point cloud to obtain more detailed terrain information and avoid collision risks.
[0062] 3. Extraction of vegetation canopy with NDVI > 0.7: NDVI calculation principle: NDVI (Normalized Difference Vegetation Index) is a core indicator reflecting the growth and density of vegetation, and the calculation formula is: , where is the spectral reflectance of the near-infrared band, is the spectral reflectance of the red light band (both collected and calibrated by the multispectral camera). The NDVI value ranges from -1 to 1, and the larger the value, the denser the vegetation.
[0063] Extraction process: Firstly, filter out the laser points with material label "tree canopy" from the fused data, and obtain the and corresponding to these points; then calculate the average NDVI value of the tree canopy points in each 5m x 5m grid; finally, mark the grid with an average NDVI value > 0.7 as a "dense vegetation canopy region".
[0064] Extraction basis: NDVI > 0.7 indicates that the vegetation canopy is dense (such as the pine forest in this embodiment, with NDVI up to 0.8-0.9 in summer), and if the flight height of the UAV is too low, it is easy to collide with the top of the canopy; at the same time, the dense canopy needs to obtain the underlying terrain information through the penetration scanning mode, so it is necessary to accurately extract this feature.
[0065] 4. Extraction of bare rock with spectral reflectance > 0.4: Spectral reflectance acquisition: The reflectance data of bare rock regions collected by the multispectral camera (calibrated by radiation, eliminating the effects of light, atmosphere, etc.), and the average reflectance of the visible light band (450-690nm) is selected as the spectral reflectance index of bare rock (the reflectance of bare rock in the visible light band is significantly higher than that of vegetation and water surface).
[0066] Extraction process: Filter out the laser points with the material label of "bare rock" from the fused data, and calculate the average spectral reflectance of the bare rock points in each 5m x 5m grid.
[0067] Extraction basis: Bare rock with spectral reflectance > 0.4 is usually a bare rock surface without vegetation coverage, and the slope of such bare rock is often larger (such as the cliff section in this embodiment), which requires high-precision reconstruction mode to obtain its terrain details to ensure the safety of UAV obstacle avoidance.
[0068] 5.3 Enhancement processing of special areas; For high-risk areas such as valleys, cliffs, and sagging sections of power lines, enhanced processing of laser point cloud data is required to improve feature extraction accuracy and the reliability of obstacle avoidance path generation. The specific processing methods are as follows: 1. Point cloud sampling density improvement in valley and cliff areas: Processing principle: The terrain of valley and cliff areas is complex, and the conventional 10Hz sampling frequency of laser point cloud may have data sparse areas (such as low point cloud density at the edge of the cliff, which cannot accurately reflect the terrain profile), resulting in slope calculation deviation or obstacle detection failure.
[0069] Processing process: When a valley or cliff area is extracted, the embedded platform sends a "density improvement instruction" to the laser radar, which increases the laser point cloud sampling density in this area by 50% (i.e. from 100,000 to 150,000 points per frame); At the same time, extend the point cloud data collection time of this area (collect one frame every 0.2s instead of the conventional 0.1s), to ensure that enough dense point cloud data is obtained.
[0070] Subsequent processing: Based on the point cloud data after density improvement, the continuous range of the valley and the slope value of the cliff are recalculated to ensure the accuracy of feature extraction; At the same time, trigger the three-dimensional obstacle avoidance path generation (use RRT algorithm to search for collision-free path in point cloud data, the safety distance between path and terrain is set to 3m), to avoid the UAV flying directly through the complex area.
[0071] 2. Prediction of sagging lowest point of power line and generation of surrounding path: Sagging lowest point prediction principle: The sagging of the power line is affected by temperature and load, showing a "low in the middle and high at both ends" curve along the line direction, and the sagging lowest point is the area with the highest collision risk (closest to the ground). Based on the semantic distribution of the power line (the "power line" label in the semantic segmentation map) and the elevation data of the laser point cloud, the position of the sagging lowest point can be predicted.
[0072] Prediction process: First, extract all laser points with the material label "power transmission line" from the fused data. Divide these points into multiple segments (each segment corresponds to a 50m long power transmission line) according to the route (based on the UAV heading angle and the spatial distribution of the power transmission line). Then, perform curve fitting on the elevation data of the laser points of each segment (using quadratic function fitting, with the sag of the power transmission line approximating a parabola). Calculate the three-dimensional coordinates of the lowest point of the fitted curve (i.e., the lowest point of the sag).
[0073] Path generation: For the predicted lowest point of the sag, a counter-clockwise loop path with a radius of 10m is generated. The reason for choosing counter-clockwise looping is that the electromagnetic field generated by the power line is distributed in a clockwise rotation. Counter-clockwise looping can avoid the UAV entering the electromagnetic field blind zone (the electromagnetic field blind zone will cause signal interference of ultrasonic radar and lidar, affecting data acquisition). The design basis of 10m radius is that the electromagnetic influence range of the power line is about 8m. A 10m radius can ensure that the UAV loops within a safe distance, while clearly collecting detailed data of the lowest point of the sag (such as line wear and insulator condition).
[0074] The implementation of flight mode matching, smooth switching and closed-loop control (corresponding to steps S400-S600) is as follows; This stage is the core decision-making and control stage for UAV terrain-following flight. It matches preset flight modes based on a combination of "terrain material classification labels + terrain structure features" to generate dynamic control commands; it achieves smooth switching between modes through dual-ended queue constraints; and it constructs closed-loop control based on sensor feedback data to correct decision-making deviations and ensure flight safety and inspection accuracy. This stage requires a detailed explanation of the specific implementation of the mode matching logic, smooth switching algorithm, and closed-loop control mechanism.
[0075] 6.1 Flight mode matching logic and control command generation (S400): This embodiment pre-defines four core flight modes (low-disturbance mode, penetration scan mode, high-precision reconstruction mode, and line-following obstacle avoidance mode). Each mode corresponds to a specific "material-structure" combination. The matching logic is designed based on the principle that "terrain attributes determine the optimal flight strategy," as detailed below: 1. Matching and control of low-disturbance modes: Matching conditions: The terrain material label is "water surface," and the structural features meet the requirements of "flat water surface area (slope < 5°) and spectral reflectance higher than the preset threshold." The reflectance threshold is set to 0.6, based on the following design principle: the reflectance of water surface in the visible light band is typically 0.5-0.8. When the reflectance > 0.6, the TOF signal of the lidar is easily interfered with by specular reflection, leading to point cloud data distortion (misjudging the water surface as "high-elevation ground"). A low-disturbance mode is needed to avoid data distortion and water surface disturbance.
[0076] Control instruction generation: When the matching condition is met, the embedded platform sends the following control instructions to the flight control system: Height control: Reduce the UAV flight height to 3m (above water surface), the design basis of 3m height is: to avoid too low to cause propeller airflow disturbance to water surface (affect water quality inspection data, if water quality monitoring is carried out in this area at the same time), and to avoid too high to cause the decline of laser radar point cloud collection accuracy of water surface details (such as floating objects).
[0077] Sensor control: Turn off the downward-looking laser radar (only keep the forward-looking laser radar for forward terrain detection), the reason is that the data of downward-looking laser radar under water surface reflection is invalid, after turning off, it can reduce data redundancy processing and reduce computing load.
[0078] Speed control: Reduce the flight speed from the conventional 5m / s to 3m / s, slowing down the flight speed can increase the collection time of multispectral camera for water surface details, and improve the identification accuracy of water quality monitoring (if equipped with water quality sensor) or water surface obstacles (such as floating branches).
[0079] 2. Matching and control of penetration scanning mode: Matching condition: The terrain material label is "crown", and the structural characteristics meet "NDVI>0.7 (dense canopy) and canopy thickness exceeds the set standard" (canopy thickness is calculated by laser radar point cloud, that is, the height difference between the top and bottom of the canopy is >5m). The design basis of this condition is: the canopy with NDVI>0.7 is dense, and the UAV is easy to collide with the top of the canopy in the conventional ground mode; the canopy thickness >5m indicates that there may be terrain undulation under the canopy, which needs to be obtained through penetration scanning.
[0080] Control instruction generation: Height control: Increase the UAV flight height to 2m above the canopy, the design basis of 2m height is: the short-wave infrared band of multispectral camera can penetrate the canopy leaves within 2m to obtain the lower terrain elevation data, while avoiding the decline of collection accuracy of canopy top details (such as tree diseases) caused by too high height.
[0081] Sensor control: Start the penetration imaging mode of multispectral camera (increase the gain of short-wave infrared band by 2 times to enhance the penetration ability), and at the same time, increase the vegetation filtering strength of laser radar by 2 times (enhance the filtering of vegetation point cloud by algorithm to retain the lower terrain point cloud).
[0082] Gimbal control: Adjust the pitch angle of multispectral camera gimbal to -30° (tilt down by 30°) to ensure that the camera lens is aimed at the lower layer of the canopy to obtain clearer penetration images.
[0083] 3. Matching and control of high-precision reconstruction mode: Matching condition: The terrain material label is "bare rock", and the structural feature meets "slope > 30°". The design basis of this condition is: the bare rock area with a slope > 30° is steep, and obstacles such as protruding rocks are likely to exist, so high-precision terrain reconstruction is needed to obtain details and ensure obstacle avoidance safety; at the same time, the foundation of the transmission line in the bare rock area (such as the iron tower) is easily affected by weathering, and high-precision reconstruction can be used for foundation disease detection.
[0084] Control instruction generation: Gimbal control: adjust the inclination angle of the laser radar gimbal to -45° (tilt 45° downward), which can make the scanning range of the laser radar more consistent with the surface of steep bare rock, obtain more dense vertical direction point cloud data, and improve the terrain reconstruction accuracy.
[0085] Sensor control: increase the scanning frequency of the laser radar from the regular 10Hz to 100Hz, and the 100Hz scanning frequency can increase the number of point clouds per frame to 1 million, and the point cloud density is increased by 10 times, ensuring that the resolution of the reconstructed terrain model reaches 0.1m (meeting the detection requirements of protruding rocks).
[0086] Speed control: reduce the flight speed to 2m / s, slow flight can ensure that the laser radar obtains enough point cloud data at each position, avoiding sparse point clouds due to fast flight.
[0087] 4. Matching and control of the line-following obstacle avoidance mode: Matching condition: The terrain material label is "dense transmission line area" (the pixel ratio of the transmission line is > 20%), and the structural feature meets "there is a sag gradient" (the sag difference of adjacent 50m transmission lines is > 2m). The design basis of this condition is: when there are dense transmission lines and a large sag gradient, the conventional flight path parallel to the line is likely to collide with the lowest point of the sag, so it is necessary to fly along the electromagnetic field strength gradient to dynamically track the line direction.
[0088] Control instruction generation: Heading control: adjust the flight heading angle along the electromagnetic field strength gradient direction, which is detected by the electromagnetic sensor carried by the UAV (the electromagnetic sensor is installed at the front of the fuselage, which feedbacks the change rate of the electromagnetic field strength in real time, and the gradient direction is the tangent direction of the line direction), ensuring that the UAV always flies along the tangent direction of the line, avoiding deviation.
[0089] Distance control: set the lateral obstacle avoidance distance to 5m (the horizontal distance between the UAV and the transmission line is kept at 5m), and the design basis of 5m distance is: the safety distance standard of 220kV transmission line is 3m, and 5m can reserve 2m of error redundancy to avoid collision due to heading angle deviation.
[0090] Height control: Based on the predicted height of the sag low point, the UAV flight height is set to "sag low point height + 3m", ensuring a vertical safety distance from the line.
[0091] 6.2 Dynamic switching logic of flight mode: In addition to the above basic matching logic, this embodiment also designs dynamic switching logic based on environmental changes to deal with sudden scenarios (such as changes in vegetation proportion, sudden gusts), as follows: 1. Switch triggered by tree crown proportion: When the embedded platform counts the tree crown pixel proportion of the current area through the semantic segmentation map to exceed 40% (counted every 0.5s), it automatically switches from "hugging mode" (the default mode for regular terrain, with a height set to 5m above ground) to "penetration scanning mode". At the same time, the vegetation filtering strength of the laser radar is increased by 2 times. The reason is that when the tree crown proportion is high, the number of vegetation points in the point cloud increases significantly, and enhancing filtering can effectively retain the point cloud data of the ground and the power line, avoiding interference from vegetation points to terrain feature extraction.
[0092] 2. Switch triggered by sudden wind speed: The wind speed sensor carried by the UAV monitors the flight environment wind speed in real time (sampling frequency 10Hz), and when a sudden wind speed exceeding 8m / s (sudden gust) is detected, the embedded platform immediately sends instructions: Inspection range adjustment: The current line inspection distance is shortened from the regular 20km to 10km, reducing flight time and reducing the risk of sustained impact of gusts.
[0093] Fixed height mode switching: Switch from "laser radar height fixing" (adjust height based on laser radar height data) to "backup barometric height fixing" (adjust height based on barometric sensor height data), because strong winds can cause the UAV attitude to be unstable, and the height data of the laser radar is easily affected by the body sway, resulting in deviation, while barometric height fixing is more stable (although the accuracy is slightly lower, but in strong wind scenarios, priority is given to stability).
[0094] 6.3 Smooth switching algorithm across modes (S500): The height difference between different flight modes is large (such as low disturbance mode height 3m, penetration scanning mode height 15m), if directly switching modes, it will cause the UAV height to mutate, causing the body attitude to fluctuate sharply, and even lose control. Therefore, this embodiment adopts a smooth switching algorithm of "double-ended queue constraint + height change rate control", which constrains the height deviation of adjacent frames through the following formula: Where: The constraint threshold is set to 20% for the adjacent frame height change rate, that is, the height change of each frame (0.1s) should not exceed 20% of the height of the previous frame. The basis for designing this threshold is that the maximum height change rate of an industrial-grade unmanned aerial vehicle is usually 1m / s. If the height of the previous frame is 10m, a change rate of 20% corresponds to a height change of 2m / s, which is within the controllable range of the unmanned aerial vehicle and can avoid attitude fluctuations.
[0095] is the flight height measurement value of the unmanned aerial vehicle at the current time (t ).
[0096] is the flight height measurement value of the unmanned aerial vehicle at the previous time (t ) and the time interval (t ) is 0.1s.
[0097] The specific implementation process of smooth switching is as follows: 1. Double-ended queue initialization: create a double-ended queue with a length of 5 to store the flight height measurement values of the last 5 frames (h ). The role of the queue is to smooth the height measurement noise (such as height jumps caused by sensor fluctuations).
[0098] 2. Height change rate calculation and judgment: when switching flight modes is needed (such as switching from low disturbance mode to penetration scanning mode, and the target height is raised from 3m to 15m), first calculate the difference between the current height (h ) and the target height (h ). Then, based on the constraint of h , calculate the maximum height adjustment allowed for each frame (h ). If h , adjust the height of the current frame by h , and adjust the remaining height difference in subsequent frames. If h , adjust directly to the target height.
[0099] 3. Queue update and feedback: after adjusting the height of each frame, add the new height value (h ) to the double-ended queue and remove the old value (h ) at the head of the queue. Calculate the average of the 5 height values in the queue as the reference height (h ) for the next frame. If the deviation between h and h exceeds 5%, adjust the height adjustment amount of the next frame to further smooth the height change.
[0100] For example, the process of switching from h to h : 1st frame (t , after adjustment , meet the constraints.
[0101] Frame 2 (( , after adjustment .
[0102] Similarly, each frame is raised by a maximum change rate of 20%, until the target height of 15m is reached around frame 10, the height change is smooth throughout the process, and the body attitude fluctuation is less than ±2° (verified by actual test). 6.4 Closed-loop control feedback mechanism (S600): To ensure that the control command execution effect of the flight mode meets the expectations, and to avoid decision bias caused by sensor errors or environmental interference, the embodiment constructs a closed-loop control mechanism based on sensor feedback, which includes the following two core links: 1. Obstacle avoidance deviation correction of ultrasonic radar: Feedback data collection: Four ultrasonic radars monitor the actual distance between the unmanned aerial vehicle and the surrounding obstacles in real time (sampling frequency 20Hz), and transmit the distance data to the embedded platform.
[0103] Deviation judgment: The embedded platform compares the actual distance with the preset obstacle avoidance threshold of the current flight mode (such as 3m for low disturbance mode and 5m for penetration scanning mode), and calculates the deviation rate: , where is the actual distance, is the preset threshold.
[0104] Correction measures: when (the deviation of actual distance and threshold exceeds 10%), it is determined as "obstacle avoidance deviation", and the embedded platform re-matches the flight mode: If (actual distance is too close), the flight height is increased by 50%, and the "temporary obstacle avoidance mode" is switched (flight speed is reduced to 2m / s, and laser radar scanning frequency is increased to 50Hz), until the actual distance returns to the threshold range.
[0105] If (actual distance is too far), the flight height is reduced by 30% to ensure the accuracy of data collection (too far will cause details to be blurred).
[0106] 2. Dynamic calibration of heading angle in power line dense area: Feedback data collection: The electromagnetic sensor carried by the UAV detects the electromagnetic field strength gradient direction of the power transmission line in real time (sampling frequency 10 Hz), and the laser radar collects the three-dimensional coordinates of the power transmission line in real time, calculates the deviation (horizontal error ).
[0107] Deviation judgment: The allowable range of horizontal error is set to ±0.5 m, and the design basis of this range is that the power transmission line inspection needs to ensure that the UAV flies along the tangent direction of the line, and a horizontal error exceeding 0.5 m may cause collision with adjacent lines (if the line spacing is small).
[0108] Calibration measures: When the horizontal error or , the embedded platform sends a heading angle calibration instruction to the flight control system: If (the UAV deviates to the right side of the line), the heading angle is adjusted to the left by ( is a proportional coefficient, set to 0.1° / m, that is, adjust 0.1° for every 1 m of deviation), to avoid excessive adjustment leading to heading fluctuations.
[0109] If (the UAV deviates to the left side of the line), the heading angle is adjusted to the right, and the adjustment amplitude is also calculated according to .
[0110] After calibration, the horizontal error is recalculated through the power transmission line coordinates fed back by the laser radar until , forming a closed-loop calibration.
[0111] In summary, the embodiment verifies the feasibility and effectiveness of the autonomous inspection flight method of the UAV by deploying the complete technical process of "multi-source sensor fusion-semantic segmentation-structural feature extraction-dynamic mode matching-smooth switching-closed-loop control" in the mountain 220kV power transmission line inspection scene. In the implementation process, the parameter design of each technical link (such as sensor accuracy, model optimization index, and flight mode height) is determined based on actual inspection requirements and hardware constraints to ensure the practicality of the technical scheme; the FocalLoss function is used to solve the class imbalance problem and improve the segmentation accuracy of the minority class materials (such as power transmission lines); the smooth switching algorithm with double-ended queue constraints is used to avoid body fluctuations caused by mode switching; and the closed-loop control mechanism is used to correct the decision deviation caused by sensor errors and environmental interference. The embodiment fully embodies the superiority of the method in complex terrain inspection scenes through specific technical implementation and parameter design, effectively solves the obstacle avoidance judgment error problem caused by relying only on elevation data in the prior art, provides a feasible technical solution for autonomous inspection flight of UAVs, and has the value of large-scale application in the fields of electric power, photovoltaic, and mountain monitoring.
[0112] Embodiment Two: As shown in Embodiment One, on the basis of Embodiment One, this embodiment details the specific steps of a method for unmanned aerial vehicle autonomous inspection of terrain flying in operation, and the specific steps are: Figure 1 1. Multi-source sensor data acquisition: The unmanned aerial vehicle synchronously acquires terrain data through the light-weight laser radar, multi-spectral camera and ultrasonic radar carried thereon. The laser radar obtains three-dimensional coordinates and reflection intensity of the terrain, the multi-spectral camera captures 5-band spectral reflectance, and the ultrasonic radar monitors real-time near-distance obstacle information. 2. Data asynchronous fusion and alignment:
[0113] The sensor data of different sampling frequencies (laser radar 10 Hz, multi-spectral camera 30 Hz, ultrasonic radar 20 Hz) are time-stamped aligned, and a spatiotemporally consistent data frame is generated through linear interpolation and mean filtering, reducing 35% of calculation redundancy. 3. Terrain material semantic segmentation:
[0114] A light-weight PSPNet-MobileNetV2 model is used to perform pixel-level semantic segmentation on the fused data, outputting 6 categories of terrain material labels (water surface, sand, tree canopy, bare rock, power line, building); the model is optimized by channel pruning and INT8 quantization, with a volume of 0.8 MB and a single-frame inference time of <50 ms. 4. Terrain structure feature extraction:
[0115] Fusion of semantic labels and laser point cloud topology data, calculation of four types of key features; Cliff with slope > 60° (based on grid elevation difference calculation); Valley of continuous negative terrain (identified by region growing algorithm); Vegetation canopy with NDVI > 0.7 (combined with red / near-infrared reflectance); Bare rock with spectral reflectance > 0.4 (visible light band mean analysis). 5. Flight mode dynamic decision:
[0116] According to the material-feature combination matching, the preset flight mode is enabled; Water surface + high reflectivity → enable low disturbance mode (height 3 m, turn off downward-looking laser); Tree canopy + NDVI > 0.7 → enable penetration scanning mode (2 m above the canopy, start multi-spectral imaging); Bare rock + slope > 30° → enable high-precision reconstruction mode (gimbal -45°, laser frequency 100 Hz); Power line dense area + sag gradient → enable line-imitating obstacle avoidance mode (lateral obstacle avoidance 5 m, fly along the electromagnetic gradient).
[0117] 6. Cross-modal smooth switching: Store the height data of the last 5 frames in a double-ended queue To Constrain the rate of change of adjacent frames to be less than or equal to 20%, and use a gradual height adjustment strategy to avoid sudden changes in posture.
[0118] 7. Closed-loop control and real-time correction: Based on the ultrasonic radar to monitor the actual obstacle avoidance distance, when the deviation is more than ±10%, re-match the flight mode; In the dense area of transmission lines, dynamically calibrate the heading angle according to the electromagnetic field strength gradient, and control the horizontal error within ±0.5m.
[0119] The above is only the preferred embodiment of the present application, not any form of limitation on the present application, although the present application has been disclosed as above, however, not to limit the present application, any person skilled in the art, without departing from the scope of the present application, when the above disclosed technical content can make some changes or modifications for equivalent embodiments, but as long as it does not deviate from the technical solution of the present application, according to the technical essence of the present application, any modification, equivalent change and modification of the above embodiment, all still belong to the scope of the present application technical solution.
Claims
1. A method for autonomous inspection of a UAV flying in a ground-like manner, characterized in that, The method comprises the following steps: S100, collecting terrain data in real time through a plurality of source sensors, the plurality of source sensors comprising a lightweight laser radar, a multispectral camera and an ultrasonic radar, wherein the ranging accuracy of the lightweight laser radar is ±3 cm, the multispectral camera covers 5 spectral bands, and the obstacle avoidance accuracy of the ultrasonic radar is ±5 cm; S200, performing pixel-level semantic segmentation on the terrain data collected in real time based on a lightweight PSPNet-MobileNetV2 hybrid architecture model, and outputting a terrain material classification label, the terrain material classification label comprising water surface, sand, tree crown, bare rock, power transmission line and building; S300, fusing the semantic segmentation result and laser point cloud spatial topology data to extract terrain structure features, the terrain structure features comprising cliff with a slope greater than 60°, valley with continuous negative terrain, vegetation canopy with NDVI greater than 0.7, and bare rock with spectral reflectance greater than 0.4; S400, matching a preset flight mode according to the combination of the terrain material classification label and the terrain structure features, and generating a dynamic control instruction; the flight mode comprises: when the water surface is identified and the reflectivity is higher than a threshold value, a low disturbance mode is enabled, the height of the unmanned aerial vehicle is controlled to be reduced to 3 m, and the downward-looking laser is turned off; when the tree crown is identified and the NDVI is greater than 0.7, a penetration scanning mode is enabled, the height is increased to 2 m above the crown layer, and multispectral penetration imaging is started; when the bare rock is identified and the slope is greater than 30°, a high-precision reconstruction mode is enabled, the tilt angle of the holder is adjusted to -45°, and the laser scanning frequency is increased to 100 Hz; when a power transmission line dense area is identified and there is an arc drop gradient, a line-imitating obstacle avoidance mode is enabled, the unmanned aerial vehicle is flown along the electromagnetic field strength gradient, and the lateral obstacle avoidance distance is set to 5 m; S500, the height deviation of adjacent frames is constrained to be less than 20% by using a double-ended queue, and smooth switching across flight modes is achieved; S600, the execution effect of the control instruction is evaluated in real time during flight, the decision deviation is corrected according to the sensor feedback data, and a closed-loop control is formed. 2.The method of claim 1, wherein, The semantic segmentation model in the step S200 removes 20% redundant layers by channel pruning, and the model volume is compressed to 0.8 MB by using INT8 quantization; The model introduces a FocalLoss function for dynamic interference, and its expression is: ; wherein, represents a predicted probability value that the current pixel point belongs to a specific terrain material category, represents a weight adjustment coefficient of the current terrain material category, represents an adjustment factor of the focused hard-to-classify samples. 3.The method of claim 1, wherein, the extraction of the terrain structure features in the step S300 comprises: the point cloud sampling density in the valley and cliff area is forcibly increased by 50%, and a three-dimensional obstacle avoidance path is generated; based on the semantic distribution of the power transmission line, the lowest point of the arc drop is predicted, and a counterclockwise ring around path with a radius of 10 m is generated to avoid the electromagnetic blind area. 4.The method of claim 1, wherein, the dynamic switching logic of the flight mode in the step S400 comprises: when the proportion of the tree crown is more than 40%, the ground-hugging mode is automatically switched to the penetration scanning mode, and the vegetation filtering strength of the laser radar is increased by 2 times synchronously; when a sudden increase in wind speed is detected and exceeds 8 m / s, the line-imitating inspection distance is shortened to 10 km, and the standby barometric height-holding mode is switched.
5. The method of claim 1, wherein, The smooth switching algorithm across flight modes in the step S500 constrains the height variation by the following equation: ; wherein, represents a change rate of adjacent frame height, represents a current time UAV flight height measurement value, represents a previous time UAV flight height measurement value. 6.The method of claim 1, wherein, the asynchronous fusion method of the plurality of source sensors comprises: the sampling frequency of the laser radar data is 10 Hz, and the sampling frequency of the visual data is 30 Hz; 35% of the calculation overhead is reduced by time stamp alignment.
7. The method of claim 1, wherein, the closed-loop control feedback mechanism in the step S600 comprises: Real-time monitoring the actual distance to obstacles by ultrasonic radar, when the deviation from the preset obstacle avoidance threshold exceeds ± 10%, re-match the flight mode; In the dense area of power transmission lines, dynamically calibrate the flight heading angle based on the direction of the electromagnetic field strength gradient, and control the horizontal error within ± 0.5m. 8.The method of claim 1, wherein, The joint output of terrain material classification and structural characteristics includes: Synchronously generate 6 categories of material labels and 4 categories of structural feature labels; The model inference time is less than 50ms, and it is deployed on the NVIDIA Jetson TX2 embedded platform. 9.The method of claim 1, wherein, Optimized deployment in resource-constrained scenarios includes: Laser radar and visual data are aligned by timestamp to achieve asynchronous fusion; Under the memory limit of embedded devices, model pruning and quantization are used to compress computing resources. 10.The method of claim 1, wherein, The flight mode decision matrix is constructed based on the following logic: When the recognized terrain semantic type is water surface, and its structural characteristics are flat and reflectivity is higher than the preset threshold, enable low disturbance flight mode; When the recognized terrain semantic type is tree crown, and its structural characteristics meet the conditions of NDVI value greater than 0.7 and crown thickness exceeding the set standard, enable the penetrating scanning flight mode; When the recognized terrain semantic type is bare rock, and its structural characteristics are slope greater than 30°, enable high-precision reconstruction flight mode; When the recognized terrain semantic type is dense area of power transmission lines, and its structural characteristics have sag gradient, enable the line-imitating obstacle avoidance flight mode.
Citation Information
Patent Citations
A photovoltaic inspection drone and its terrain-following flight method
CN111966129B
A method and system for controlling unmanned aerial vehicle terrain-imitating flight based on terrain height
CN116627164B
Unmanned aerial vehicle track smooth prediction method and system based on terrain height
CN116643580A
Unmanned aerial vehicle laser radar imitation flight method based on geographic space data
CN117406778A
Power transmission line unmanned aerial vehicle auxiliary inspection method and system
CN118707988A