Intelligent power distribution line unmanned aerial vehicle inspection method and device based on light and visual integration

CN122551216APending Publication Date: 2026-08-11HEBEI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0007]本发明提供了一种基于光视一体化的智能配电线路无人机巡检方法及设备,解决了配电线路巡检存在的无人机识别困难,可靠性差的问题

Benefits of technology

[0012] This invention provides a method and equipment for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration. The invention employs an optical-visual integrated inspection architecture, simultaneously acquiring 3D point cloud data from a lidar radar and 2D image data from a visible light camera. After fusion preprocessing, key point cloud data of critical inspection targets are extracted, overcoming the inherent limitations of single visible light imaging, such as poor environmental adaptability and lack of 3D spatial information. The key point cloud data is precisely semantically segmented using the RandLA-CGNet network to obtain a 3D semantic point cloud with multi-category semantic information. This leverages the 3D characteristics of the lidar point cloud to accurately quantify the spatial distance and net spatial gap of the inspection target, while also improving target recognition accuracy under complex lighting and adverse weather conditions through optical-visual fusion features. A unified semantic cost map is constructed based on the 3D semantic point cloud to complete path planning, achieving highly reliable autonomous UAV inspection and reducing inspection miss rates and operational safety risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551216A_ABST
    Figure CN122551216A_ABST
Patent Text Reader

Abstract

This invention provides a method and equipment for intelligent power distribution line inspection using an integrated optical-visual system, relating to the field of power distribution line inspection technology. The invention employs an integrated optical-visual inspection architecture, simultaneously acquiring 3D point cloud data from a LiDAR and 2D image data from a visible light camera. After fusion preprocessing, key point cloud data of critical inspection targets is extracted, overcoming the inherent limitations of single visible light imaging, such as poor environmental adaptability and lack of 3D spatial information. The key point cloud data is precisely semantically segmented using the RandLA-CGNet network, resulting in a 3D semantic point cloud with multi-category semantic information. This achieves accurate quantification of spatial distance and net spatial gap of inspection targets, while also improving target recognition accuracy under complex lighting and adverse weather conditions. A unified semantic cost map is constructed based on the 3D semantic point cloud to complete path planning, enabling highly reliable autonomous UAV inspection and reducing inspection miss rates and operational safety risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power distribution line inspection technology, and in particular to a method and equipment for intelligent power distribution line unmanned aerial vehicle (UAV) inspection based on optical-visual integration. Background Technology

[0002] As the core terminal link of the power system, the distribution network directly supplies power to end users, and its operational safety and reliability are directly related to the overall stability of the power system. Distribution overhead lines are characterized by their wide distribution range, complex laying environment, and long-term outdoor exposure, making them susceptible to factors such as weather conditions, external damage, and vegetation growth. This can lead to safety hazards such as conductor damage, insulator aging, tree intrusion, and insufficient clearance at crossings. Regular inspections are the core technical means to promptly identify potential hazards and prevent line faults.

[0003] With its advantages of flexible operation, wide coverage, and no high-altitude operation risks, drone inspection is gradually replacing the traditional manual pole climbing and ground inspection mode, becoming the mainstream technical solution for power distribution line inspection. Current mainstream drone power distribution line inspection solutions mostly use a monocular visible light camera as the core data acquisition device, relying on visible light images to achieve line target identification and hidden danger investigation.

[0004] On the one hand, visible light imaging has extremely poor environmental adaptability, and the accuracy of target recognition cannot be guaranteed under complex lighting and adverse weather conditions. The imaging quality of visible light cameras is highly dependent on ambient light and atmospheric transparency. In strong light and backlight scenes, problems such as overexposure and halo occlusion are prone to occur, resulting in the loss of the outline and features of key inspection targets such as conductors, insulators, and hardware. In low-light scenes such as cloudy days and nights, the image signal-to-noise ratio is greatly reduced, and target features are blurred, making effective recognition impossible. In adverse weather conditions such as foggy days, hazy days, and rainy days, the atmospheric penetration ability of visible light is weak, and the image is prone to problems such as fogging, blurring, and loss of details, further increasing the difficulty of recognition. The above problems result in a high rate of missed detection and false detection in existing visible light inspection solutions, which cannot meet the requirements of all-weather, all-scenario inspection of power distribution lines. Often, manual secondary verification is required, which greatly reduces the efficiency of inspection operations.

[0005] On the other hand, visible light cameras can only acquire two-dimensional planar images and cannot reconstruct the true three-dimensional spatial topology of the inspection scene, thus failing to meet the core safety requirements of power distribution line inspection. Based on two-dimensional visible light images, only planar position recognition of targets can be achieved, but it is impossible to accurately measure key three-dimensional parameters such as the true spatial distance, vertical height, and clear clearance between conductors and towers, trees, buildings, and crossing lines. The core control requirement of power distribution line inspection is the quantitative investigation of hidden dangers such as tree intrusion and insufficient crossings. At the same time, obstacle avoidance safety control of drone autonomous inspection also highly depends on the true three-dimensional spatial information of the scene. Existing two-dimensional imaging solutions can only qualitatively determine the presence of hidden dangers through image pixel ratios, but cannot achieve quantitative assessment of the level of hidden dangers, nor can they provide accurate safety distance control basis for drone autonomous inspection. This not only makes it difficult to achieve refined control of hidden dangers, but also easily leads to safety accidents such as collisions with lines and drone crashes during drone inspections, making it unsuitable for the engineering application requirements of autonomous power distribution line inspection.

[0006] Therefore, current power distribution line inspections suffer from difficulties in drone identification and poor reliability. Summary of the Invention

[0007] This invention provides a method and equipment for intelligent power distribution line drone inspection based on optical-visual integration, which solves the problems of difficult drone identification and poor reliability in power distribution line inspection.

[0008] In a first aspect, this invention provides a method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration. The method includes: collecting 3D point cloud data from a lidar and 2D image data from a visible light camera during the inspection process; fusing and preprocessing the 3D point cloud data and 2D image data to extract multiple key inspection targets, including conductors, towers, insulators, trees, and buildings, to obtain key point cloud data; performing semantic segmentation on the key point cloud data based on the RandLA-CGNet network to obtain a 3D semantic point cloud, which includes categories such as conductors, towers, passage obstacles, and safe areas; projecting the 3D semantic point cloud onto a 2D grid plane to construct a unified semantic cost map, which includes an occupancy layer, a semantic cost layer, and a safety expansion cost layer; and performing path planning based on the unified semantic cost map to obtain the optimal inspection route, thereby enabling autonomous inspection control of the UAV.

[0009] Secondly, embodiments of the present invention provide an intelligent power distribution line UAV inspection device based on optical-visual integration. This inspection device includes a communication module and a processing module. The communication module is used to collect three-dimensional point cloud data from a lidar and two-dimensional image data from a visible light camera during the inspection process. The processing module is used to fuse and preprocess the three-dimensional point cloud data and the two-dimensional image data, extracting multiple key inspection targets such as conductors, towers, insulators, trees, and buildings to obtain key point cloud data. Based on the RandLA-CGNet network, the key point cloud data is semantically segmented to obtain a three-dimensional semantic point cloud, which includes categories such as conductors, towers, passage obstacles, and safe areas. The three-dimensional semantic point cloud is projected onto a two-dimensional grid plane to construct a unified semantic cost map, which includes an occupancy layer, a semantic cost layer, and a safety expansion cost layer. Based on the unified semantic cost map, path planning is performed to obtain the optimal inspection route, enabling the UAV to perform autonomous inspection control.

[0010] Thirdly, embodiments of the present invention provide an electronic device including a memory and a processor. The memory stores a computer program, and the processor is configured to call and run the computer program stored in the memory to perform the steps of the method as described in the first aspect and any possible implementation thereof.

[0011] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method as described in the first aspect and any possible implementation thereof.

[0012] This invention provides a method and equipment for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration. The invention employs an optical-visual integrated inspection architecture, simultaneously acquiring 3D point cloud data from a lidar radar and 2D image data from a visible light camera. After fusion preprocessing, key point cloud data of critical inspection targets are extracted, overcoming the inherent limitations of single visible light imaging, such as poor environmental adaptability and lack of 3D spatial information. The key point cloud data is precisely semantically segmented using the RandLA-CGNet network to obtain a 3D semantic point cloud with multi-category semantic information. This leverages the 3D characteristics of the lidar point cloud to accurately quantify the spatial distance and net spatial gap of the inspection target, while also improving target recognition accuracy under complex lighting and adverse weather conditions through optical-visual fusion features. A unified semantic cost map is constructed based on the 3D semantic point cloud to complete path planning, achieving highly reliable autonomous UAV inspection and reducing inspection miss rates and operational safety risks. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating an intelligent power distribution line unmanned aerial vehicle (UAV) inspection method based on optical-visual integration, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of a mapping from world coordinates to grid coordinates provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the construction of a two-dimensional occupied grid map from a three-dimensional point cloud, as provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of a semantic cost layer and a security expansion layer provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a grid neighborhood connection provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the overall network structure of RandLA-Net provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of a RandLA-CGNet network structure provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of a local feature aggregation module provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of a local-global context fusion module provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a dynamic ratio contextual attention module provided in an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of a norm-gated channel feature module provided in an embodiment of the present invention; Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the present invention clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0016] like Figure 1 As shown, this invention provides a method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration. The method includes steps S101-S105.

[0017] S101. Collect 3D point cloud data from the lidar and 2D image data from the visible light camera during the inspection process.

[0018] In some embodiments, the three-dimensional point cloud data is a discrete set of points obtained by lidar through a three-dimensional scanning of the inspection scene using a time-of-flight ranging mechanism. Each point contains three-dimensional spatial coordinates (x, y, z) and laser reflection intensity information.

[0019] In some embodiments, the two-dimensional image data is a digital image of an inspection scene captured by a visible light camera, and each pixel contains texture, color, and grayscale features to supplement the missing visual semantic information in the point cloud.

[0020] For example, in embodiments of the present invention, the joint calibration of the lidar and visible light camera carried by the UAV can be completed in advance to determine the spatial mapping relationship between the intrinsic and extrinsic parameters of the two types of sensors; the UAV is controlled to fly along the preset initial route of the power distribution line, and the inspection line is scanned and image acquired in 360° by the lidar and visible light camera that are triggered simultaneously, and three-dimensional point cloud data and two-dimensional image data are acquired simultaneously.

[0021] S102. The three-dimensional point cloud data and two-dimensional image data are fused and preprocessed to extract multiple key inspection targets such as conductors, towers, insulators, trees and buildings to obtain key point cloud data.

[0022] As one possible implementation, step S102 can be specifically implemented as steps S1021-S1024.

[0023] S1021. Based on the 3D point cloud data and 2D image data, as well as the pre-calibrated spatial mapping relationship between the UAV's lidar and visible light camera, timestamp synchronization and spatial coordinate alignment are performed to obtain spatiotemporally registered 3D point cloud data and 2D image data.

[0024] In some embodiments, the spatial mapping relationship is obtained by joint calibration of the lidar and the visible light camera, including the camera intrinsic parameter matrix, distortion coefficient, and the extrinsic parameter rotation matrix and translation vector between the lidar and the camera, which is used to realize the coordinate mapping from the three-dimensional point cloud to the two-dimensional image pixel plane.

[0025] For example, in embodiments of the present invention, each spatial point of the three-dimensional point cloud can be projected onto the pixel plane of the two-dimensional image based on the calibrated spatial mapping relationship, the pixel coordinates corresponding to each point can be calculated, a one-to-one correspondence between the spatial positions of the three-dimensional point cloud and the two-dimensional image can be established, and spatial dimension alignment can be completed; the validity of the point cloud and image data that have completed time synchronization and spatial alignment can be verified to obtain the spatiotemporally registered three-dimensional point cloud data and two-dimensional image data.

[0026] S1022. Perform optical and visual feature fusion on the spatiotemporally registered 3D point cloud and 2D image, mapping the texture, color, and grayscale features of the 2D image to each spatial point of the 3D point cloud, adding visual features to the point cloud, and obtaining the fused feature point cloud.

[0027] In some embodiments, optical-visual feature fusion, also known as laser-visual feature fusion, deeply fuses the visual semantic features of a two-dimensional image with the geometric features of a three-dimensional laser point cloud, enabling the point cloud to simultaneously possess three-dimensional spatial geometric information and two-dimensional visual texture information, thus solving the problem of insufficient information from a single sensor.

[0028] For example, in this embodiment of the invention, each spatial point of the spatiotemporally registered 3D point cloud can be projected onto the pixel plane of a 2D image based on a spatial mapping relationship to obtain the pixel coordinates corresponding to each spatial point; the texture, color, and grayscale visual features of the pixel position corresponding to the pixel coordinates of each spatial point are extracted, the extracted visual features are assigned to the spatial points corresponding to the 3D point cloud and feature dimension stitching is performed, and the geometric features such as the 3D coordinates and laser reflection intensity of the point cloud are stitched together with the newly added visual features in the channel dimension to obtain a fused feature point cloud that simultaneously possesses geometric and visual features.

[0029] S1023. Perform two-stage filtering on the fused feature point cloud to obtain filtered point cloud data.

[0030] In some embodiments, two-stage filtering refers to a two-stage filtering process: first, statistical filtering for noise reduction, followed by direct-pass filtering for pruning, balancing noise removal and effective region selection. Statistical filtering is a point cloud outlier removal algorithm that calculates the average distance between each point and its k nearest neighbors, using a Gaussian distribution to remove outlier noise points whose distance exceeds the mean ± standard deviation. Direct-pass filtering is a point cloud pruning algorithm that retains effective point cloud data within a specified dimension (such as the Z-axis elevation) while removing invalid data.

[0031] S1024. The filtered point cloud data is subjected to dimensionality reduction processing by random downsampling to extract key point cloud data containing key inspection targets such as conductors, towers, insulators, trees, and buildings.

[0032] In some embodiments, random downsampling is a point cloud dimensionality reduction algorithm. According to a preset sampling ratio, a corresponding number of points are randomly extracted from the original point cloud. While preserving the overall geometric structure and spatial distribution of the point cloud, the point cloud density is reduced, thereby reducing the computational load of subsequent algorithms.

[0033] For example, in this embodiment of the invention, a sampling ratio for random downsampling can be preset based on the total number of points in the filtered point cloud data and the minimum size of the target for power line inspection. The sampling ratio must ensure that the geometric structure of small targets such as thin wires and small insulators is not lost. According to the preset sampling ratio, random sampling is performed on the filtered point cloud data, and each point is assigned a random sampling weight. After sorting by weight, the corresponding number of valid points are retained to obtain key point cloud data.

[0034] S103. Semantic segmentation of key point cloud data is performed based on the RandLA-CGNet network to obtain a three-dimensional semantic point cloud.

[0035] In this application, the three-dimensional semantic point cloud includes categories such as conductors, towers, passageway obstacles, and safe zones.

[0036] As one possible implementation, step S103 can be specifically implemented as steps S1031-S1034.

[0037] S1031. Input the key point cloud data into the encoder of the RandLA-CGNet network. Through the local feature aggregation module of the encoder, extract the local geometric structure features of the power distribution line point cloud and generate shallow multi-scale features containing details of thin conductors and small insulators.

[0038] In some embodiments, the Local Feature Aggregation (LFA) module is a lightweight feature extraction unit native to RandLA-Net and a shallow core module of the encoder. Through a three-level process of local spatial encoding, shared fully connected mapping, and attention pooling, it preserves the geometric details of tiny targets such as thin wires and small insulators during random downsampling.

[0039] For example, in embodiments of the present invention, visual features (color, texture, grayscale) attached to a point cloud can be concatenated with spatial coordinates to form an input feature vector. Let the input keypoint cloud be... , where p i ∈R 3 Let be the 3D spatial coordinates of the i-th point, and N be the total number of points in the input point cloud. The keypoint cloud is then input into the encoder of the RandLA-CGNet network. First, it enters the encoder's first-layer downsampling module, where k-nearest neighbor search (k=16) is used to find the k-nearest neighbor for each center point p. i Matching neighborhood point sets Construct local neighborhood relationships. Perform local spatial encoding on each center point and its set of neighboring points to explicitly encode local geometric relationships: ; where r ij Center point p i With neighboring point p j The local space encoding vector, The encoding, which represents the Euclidean distance between two points, fully preserves the relative position and geometric structure information within the neighborhood. A shared fully connected layer is applied to the local spatial encoding vector to generate high-dimensional point features; then, attention pooling is used to weight and aggregate the neighborhood features, enhancing the representation of key detail features. ;in, For the neighborhood point p j Features mapped by the fully connected layer, g( ) is the attention weight generation function, α ij The attention weights for neighboring points, The local features are aggregated from the center points. Random downsampling is performed on the aggregated local features, compressing the number of points to 1 / 4 of the original. Then, the local features are input into the second-layer downsampling module of the encoder, and the above local feature aggregation process is repeated to generate a second-layer shallow feature containing multi-scale detailed information. The shallow feature is simultaneously sent to the deep module of the encoder and the corresponding layer of the decoder to complete the shallow feature extraction.

[0040] S1032. Through the local-global context fusion module of the encoder, local feature aggregation and dynamic ratio context attention calculation are performed in parallel to extract local fine structure information and global line channel semantic information simultaneously, generating deep coding features that take into account both small target details and large scene consistency.

[0041] In some embodiments, the Local-Global Context Fusion (LGCF) module has a dual-branch parallel structure. The local branch retains the ability to capture details, while the global branch injects global semantic information through Dynamic Ratio Context Attention (DRCA), which solves the problem of poor semantic consistency in large-scale scenarios in traditional networks.

[0042] In some embodiments, Dynamic Ratio Context Attention (DRCA) is the core of the LGCF global branch. It adaptively adjusts the channel compression dimension through a trainable compression ratio vector, replacing the traditional fixed compression ratio, and balancing detail preservation and computational efficiency.

[0043] In some embodiments, deep coding features are high-dimensional features output from the last three layers of the encoder, which include local fine structure and global line channel semantic information.

[0044] For example, the features output from the shallow layer of the encoder are randomly downsampled and then input into the LGCF module of the third layer of the encoder. The module is divided into local branches and global branches for parallel computation. Local branch computation: The local feature aggregation process is repeated, performing k-nearest neighbor search, local spatial encoding, and attention pooling on the input features, and outputting the local aggregated features F. local It preserves detailed information about fine wires and small hardware.

[0045] ; where s ij For attention weights, h ij For the output features of the fully connected layer, f j These are the original features of the neighboring points.

[0046] Global Branching (DRCA) Calculation: Performing Dynamic Ratio Contextual Attention Calculation: First, perform global average pooling along the point dimension on the input features to obtain global contextual features: Where C is the number of feature channels, Let be the global feature value of the c-th channel. The second step involves adaptively calculating the channel compression dimension using a trainable compression ratio vector r, performing channel compression and decompression on the global features, and generating dynamic channel weights 'a'. c Then, the dynamic weights are multiplied channel by channel with the local features, injecting global semantic information to output the global enhanced feature F. global : Adaptive fusion is performed on the output features of the local and global branches using a trainable fusion coefficient vector. Where α is the trainable channel fusion coefficient, which is automatically optimized during network training to balance the contribution ratio of local and global features. Batch normalization and LeakyReLU activation are performed on the fused features to generate the third layer of deep features; then these are input into the fourth and fifth layers of the encoder, and the above LGCF calculation process is repeated to finally generate deep encoded features that take into account both the details of small targets and the global semantic consistency of large scenes.

[0047] S1033. Input the deep coding features into the decoder, combine them with the shallow features of the corresponding layer of the skip connection fusion encoder, and after feature fusion of each layer of the decoder, perform adaptive recalibration of the feature channels through the norm-gated channel feature module to generate high-resolution point-by-point features.

[0048] In some embodiments, the Norm-Gated Channel Feature Module (NGCF) is inserted after feature fusion at each layer of the decoder. Through L2 norm aggregation, gated recalibration, and residual fusion, it adaptively amplifies key feature channels and suppresses noise channels, improving boundary segmentation accuracy. Skip connections are feature fusion channels at corresponding layers of the encoder and decoder, directly passing shallow detail features to the decoder to compensate for detail loss during downsampling. High-resolution point-by-point features are the high-dimensional features at the final output of the decoder that correspond to each point in the input point cloud, serving as the direct basis for semantic classification.

[0049] For example, in this embodiment of the invention, the deep encoded features finally output by the encoder can be input into the first layer of the decoder. The point cloud resolution is restored through trilinear interpolation upsampling, ensuring that the number of points in the upsampled features matches the number of points in the shallow features of the fourth layer of the encoder. Skip connection fusion is then performed: the upsampled features are aligned with the channel dimensions of the corresponding shallow features of the encoder layer, and then fused by channel concatenation to obtain the fused feature F. raw ∈R B×N×C Where B is the batch size, N is the number of points, and C is the number of channels. The fused features are input into the NGCF module for channel adaptive recalibration: Step 1, L2 norm aggregation: L2 norm aggregation is performed on the fused features along the point dimension to obtain the global response value for each channel. ; where f c This is the global response value for the c-th channel. The first step is to calculate the L2 norm. The second step is to normalize the channel response values ​​to eliminate scaling differences. The third step is gate recalibration: A gate function is constructed using trainable gate parameters γ and β to generate adaptive channel weights. Among them, g c Here, represents the gating weight for channel c, and 1+ is the design feature to ensure that the original characteristic response is preserved even if the gating fails. Fourth step: Perform channel recalibration. In the formula, F ng These are the features after gated recalibration. Fifth step, residual fusion: Using the optimal fixed weight λ=0.5 from the ablation experiment, residual fusion is performed. Step 6: Perform batch normalization and LeakyReLU activation on the fused features: Repeat the above upsampling-skip connection fusion-NGCF recalibration process to complete the five-layer calculation of the decoder, gradually restoring the point cloud resolution to the original number of input points, and finally outputting high-resolution point-by-point features.

[0050] S1034. Based on high-resolution point-by-point features, the semantic category prediction results of each point are output through the point-by-point classification head to generate a three-dimensional semantic point cloud containing categories such as conductors, towers, passage obstacles, and safe areas.

[0051] In some embodiments, the point-by-point classification head is the network output layer, consisting of 1×1 shared convolutions and Softmax activation, which maps high-dimensional point-by-point features to the class probability distribution of each point and outputs the semantic prediction result. The 3D semantic point cloud is the point cloud after semantic segmentation, with each point having an attached semantic class label.

[0052] For example, embodiments of the present invention can convert the high-resolution point-by-point feature F output by the decoder into... final ∈R N×CfeatThe input point-by-point classification header is processed by a 1×1 shared convolutional layer to map the feature dimension to a preset number of semantic categories K (K=4 in this scheme, corresponding to wires, towers, passage obstacles, and safe areas), thus obtaining the category prediction feature F. cls ∈R N×K The probability distribution for each category is calculated using the Softmax activation function, as shown in the following formula: Among them, p i,k Let i be the predicted probability that the i-th point belongs to the k-th class, satisfying The probability values ​​range from [0,1]. A maximum probability voting strategy is used to determine the final semantic category of each point, as shown in the following formula: Among them, l i This is the final semantic category label for the i-th point. The semantic category label is then concatenated with the spatial coordinates and visual features of the original keypoint cloud to generate a 3D semantic point cloud with point-by-point semantic information. The semantic categories cover four types: conductors, towers, passageway obstacles, and safe areas.

[0053] S104. Project the 3D semantic point cloud onto the 2D raster plane to construct a unified semantic cost map.

[0054] In this embodiment, the unified semantic cost map includes an occupancy layer, a semantic cost layer, and a security inflation cost layer.

[0055] As one possible implementation, step S104 can be specifically implemented as steps S1041-S1045.

[0056] S1041. Based on the resolution and planar coordinate boundaries of the three-dimensional semantic point cloud and the two-dimensional raster map, vertical projection is performed to realize the mapping transformation from three-dimensional spatial coordinates to two-dimensional raster index, thereby obtaining two-dimensional semantic raster data.

[0057] In some embodiments, the resolution of the two-dimensional grid map is the physical space size corresponding to a single grid cell, in meters (m). For power distribution inspection scenarios, the resolution is set to 0.1m to 0.5m. The planar coordinate boundary is the range of X-axis and Y-axis values ​​of the two-dimensional grid map in the world coordinate system, corresponding to the overall spatial range of the power distribution line inspection corridor, and determining the overall size and number of rows and columns of the grid map.

[0058] In some embodiments, vertical projection is the operation of projecting a three-dimensional spatial point cloud along the elevation Z-axis onto a two-dimensional XOY plane, converting the three-dimensional semantic point cloud into planar data that can be processed by a two-dimensional raster.

[0059] In some embodiments, the two-dimensional semantic raster data is a discretized dataset in which each raster cell contains information such as the semantic category, number of points, and spatial coordinates of the corresponding three-dimensional point cloud after projection mapping is completed.

[0060] For example, Figure 2 This diagram illustrates the mapping from world coordinates to raster coordinates, demonstrating the principle of the transformation from world space coordinates to two-dimensional raster indexes in a 3D semantic point cloud. This is fundamental to the rasterization of 3D point clouds and the construction of a unified semantic cost map. The left side represents the world coordinate system, corresponding to the real physical space of a power distribution line inspection scenario. The right side represents the raster division coordinate system, corresponding to the discrete index space of a two-dimensional raster map. Using the origin of the world coordinate system as the origin of the raster coordinates, the continuous physical space is discretized into a regular grid at a preset resolution. P(i,j) represents a physical space point, and P(x,y) corresponds to the raster row and column indices after mapping, achieving a one-to-one correspondence between the continuous physical space and the discrete raster space.

[0061] For example, embodiments of the present invention can perform elevation filtering on three-dimensional semantic point clouds, retaining point clouds within the effective elevation range of power distribution line inspections. The filtered three-dimensional semantic point clouds are then subjected to vertical projection onto the XOY plane, mapping the spatial coordinates (x, y, z) of each three-dimensional point to a two-dimensional raster index. Where (i,j) are the raster row and column indices corresponding to point (x,y), and floor( The function is a floor function that completes a one-to-one mapping from 3D spatial coordinates to 2D raster indices. All projected point clouds are grouped by raster index, and the number of point clouds contained within each raster cell and the semantic category label of each point are counted to generate 2D semantic raster data containing semantic information.

[0062] S1042. Based on the two-dimensional semantic raster data and the semantic categories corresponding to the three-dimensional semantic point cloud, mark the categories of impassable hard obstacles, count the number of hard obstacle points in each raster, and determine the occupancy status of the raster by combining the preset occupancy threshold, and generate a binary occupancy layer.

[0063] In some embodiments, the category of impassable hard obstacles is: in power distribution inspection scenarios, the semantic category that drones absolutely cannot cross includes wires, towers, and buildings. The point cloud of this category is the core basis for grid occupancy determination.

[0064] In some embodiments, the binary occupancy layer, also known as the accessibility occupancy layer, is the hard constraint basis of the unified semantic cost map. Each grid cell has only two states: "occupied (unaccessible)" and "idle (accessible)".

[0065] In some embodiments, a preset occupancy threshold is defined as the number of hard obstacle points used to determine whether a grid is occupied, which is used to suppress the interference of point cloud noise on the occupancy determination.

[0066] Figure 3This diagram illustrates the construction of a 3D point cloud into a 2D occupancy grid map. It demonstrates the complete transformation process of a 3D point cloud with semantic information, after vertical projection and occupancy status determination, to generate a binary occupancy layer. The left side shows the original 3D point cloud: a 3D semantic point cloud of a power line inspection scene after semantic segmentation, with different colors corresponding to different semantic categories. The right side shows the 2D occupancy grid representation: a binary occupancy layer generated from the 3D point cloud after vertical projection, hard obstacle point statistics, and occupancy status determination. Black grids are marked as "occupied," representing impassable hard obstacles, while white grids are marked as "free," representing passable grids.

[0067] For example, embodiments of the present invention can perform traversability attribute mapping based on the semantic categories of three-dimensional semantic point clouds, marking wires, towers, and buildings as impassable hard obstacles, trees and passageways as traversable soft obstacles, and safe areas as freely traversable. Each grid cell of the two-dimensional semantic raster data is traversed, and the number of point clouds representing hard obstacles within each grid cell is counted. Where x is the current raster cell, and n hard (x) represents the number of hard obstacle points within the grid, l k Let I( be the semantic category label of the k-th point within the raster). The function is an indicator function that outputs 1 if the condition is met, and 0 otherwise. A threshold τ is set for the occupancy of hard obstacle points. hard Based on the number of hard obstacle points within the grid, the grid occupancy status is determined: Where Occ(x) is the final occupancy state of grid x, with an output of 1 indicating occupancy (impassable) and an output of 0 indicating vacancy (passable). The majority vote semantic category for the grid is determined. Based on the occupancy status determination results of all grids, a binary occupancy layer is generated, where occupied grids are marked in black and idle grids are marked in white.

[0068] S1043. Based on two-dimensional semantic raster data and the semantic categories corresponding to three-dimensional semantic point clouds, different levels of access risk are divided, and a corresponding level of semantic cost value is set for each raster to generate a semantic cost layer.

[0069] In some embodiments, the semantic cost layer is the core semantic constraint layer of the unified semantic cost map. It converts discrete semantic category labels into continuous generation values ​​that can be accumulated and calculated for path planning, and distinguishes between high-risk areas that are "passable but not recommended to pass" and low-risk areas that are "preferred to pass".

[0070] In some embodiments, the access risk level is a gradient level for semantic categories based on the accessibility and inspection security risk of semantic categories. The higher the level, the higher the semantic value of the corresponding grid.

[0071] In some embodiments, semantic cost value is a quantified value of the passage risk corresponding to each grid cell. The higher the cost value, the higher the passage risk of that grid cell.

[0072] Figure 4 Figure (a) shows the semantic cost layer and the safety inflation layer, which are used to illustrate the rasterized representation of the two core cost layers in the unified semantic cost map. Figure (b) shows the semantic cost layer: based on the passage risk level of the semantic category within the raster, a gradient-increasing semantic cost value is set for different rasters; Figure (c) shows the safety inflation cost layer: using the hard obstacle rasters of the binary occupancy layer as the inflation source, the gradient cost value is generated through distance transformation and exponential decay function; the black rasters are the core hard obstacle rasters, with a cost value of the maximum value of 1; the yellow rasters are the safety inflation buffer areas around the obstacles, the closer to the obstacle rasters, the higher the cost value, and the cost value decays to 0 after the distance exceeds the safety inflation radius.

[0073] For example, step S1043 can be specifically implemented as steps A1-A5.

[0074] A1. Based on the semantic categories of 3D semantic point clouds, safe areas are classified into low-risk levels, trees and passage obstacles into medium-risk levels, and wires, towers and buildings into high-risk levels.

[0075] In some embodiments, the following risk levels are defined: Low Risk (Free): No obstacles or safety risks; a priority area for route planning, representing a safe flight zone for power line inspection corridors. Medium Risk (Soft): Geometrically passable, but with risks of collisions and scrapes; an area to be avoided as much as possible in route planning, including trees and passageway obstacles. High Risk (Hard): Absolutely impassable; poses a fatal risk of drone crashes and electrocution; an area to be completely avoided in route planning, including power lines, poles, and buildings.

[0076] A2. Set fixed semantic values ​​with increasing gradients for low-risk, medium-risk, and high-risk levels respectively.

[0077] In some embodiments, a fixed semantic cost value is defined as the standardized cost value corresponding to each risk level, with a numerical range of [0,1]. The higher the risk level, the greater the cost value. A gradient-increasing rule is applied: cost value for high-risk levels > cost value for medium-risk levels > cost value for low-risk levels. This gradient difference in cost values ​​guides path planning to prioritize low-risk areas and avoid high-risk areas.

[0078] For example, low-risk level (security zone): fixed semantic value set to C. Free =0, no semantic penalty, preferred choice in path planning; Medium risk level (trees, obstacles): fixed semantic cost set to C. Soft=0.5, resulting in a moderate semantic penalty during path planning, which the algorithm tries to avoid; High-risk level (conductors, towers, buildings): fixed semantic cost is set to C. Hard =1.0, which will generate the maximum semantic penalty in path planning, forming a double guarantee with the hard obstacle occupancy constraint.

[0079] A3. Traverse each grid cell of the two-dimensional semantic raster data and count the percentage of point clouds of different semantic categories within the grid cell.

[0080] In some embodiments, the semantic category point cloud quantity ratio is the proportion of the number of point clouds of a certain semantic category within a single raster to the total number of point clouds within that raster, used to determine the dominant semantic category of the raster.

[0081] For example, embodiments of the present invention can traverse each grid cell of the two-dimensional semantic raster data and perform point cloud statistics on each grid cell: First, count the total number of point clouds N in the grid cell. total The second step is to count the number N of semantic category point clouds corresponding to high-risk, medium-risk, and low-risk levels within the raster. Hard N Soft N Free The percentage of point cloud data corresponding to each risk level is calculated using the following formula: Among them, R Hard R Soft R Free The percentages of point cloud numbers representing high, medium, and low risk levels are respectively, and the sum of the three is 1.

[0082] A4. Based on the proportion of point cloud quantity of different semantic categories in each raster and the fixed semantic cost of each risk level, the semantic cost of each raster is calculated.

[0083] For example, in embodiments of the present invention, a hard obstacle ratio threshold τ can be preset. occ When the proportion of high-risk point clouds within the grid is R Hard ≥τ occ When this condition is met, the semantic cost of the raster is directly forced to be set to 1.0, consistent with the hard barrier constraint of the binary occupancy layer. For rasters that do not meet the hard barrier threshold, a weighted summation is performed based on the proportion of point cloud points and the fixed semantic cost to obtain the final semantic cost of the raster, as shown in the following formula: Among them, C sem (x) represents the final semantic value of raster x, with a value range of [0,1].

[0084] A5. Generate a semantic cost layer based on the semantic cost of each grid.

[0085] In some embodiments, the semantic cost layer is the final generated single-channel two-dimensional raster matrix, where each element of the matrix corresponds to the semantic cost of a raster, and its size, resolution, and coordinate boundaries are completely consistent with those of the binary occupancy layer and the security inflation cost layer.

[0086] For example, embodiments of the present invention can create a two-dimensional matrix with the exact same size as the binary occupancy layer based on the number of rows and columns of the two-dimensional semantic raster data. The final semantic cost value calculated for each raster is filled into the corresponding index position of the two-dimensional matrix to form a semantic cost matrix. Based on the semantic cost matrix, a semantic cost layer is generated.

[0087] S1044. Based on the binary occupancy layer, using the obstacle grid as the expansion source, the distance from each grid to the nearest obstacle grid is calculated through distance transformation, and an exponential decay function is used to generate a safe expansion cost layer.

[0088] In some embodiments, the safety inflation cost layer is a safety constraint layer of the unified semantic cost map. It uses obstacle grids as inflation sources and sets gradient-increasing cost values ​​for grids around obstacles through distance transformation and exponential decay functions, so that path planning can automatically maintain a safe distance from obstacles.

[0089] In some embodiments, the distance transformation is an algorithm that calculates the Euclidean distance from each free grid cell to the nearest occupied obstacle grid cell in a two-dimensional grid map. The exponential decay function is a safety cost mapping function with the core characteristic that the closer to the obstacle grid cell, the higher the cost; once the distance to the obstacle exceeds the safety expansion radius, the cost decays to 0, thus quantifying the safety margin around the obstacle.

[0090] For example, in this embodiment of the invention, the obstacle grid marked as "occupied" in the binary occupancy layer can be used as the sole expansion source. An Euclidean distance transformation is performed on the entire grid map to calculate the shortest Euclidean distance d(x) from each grid x to the nearest obstacle grid, in units of grid numbers. A preset safe expansion radius d for power distribution inspection is also provided. max (Corresponding to electrical safety distance), expansion attenuation coefficient α, using an exponential attenuation function, calculate the safety expansion cost of each grid cell: Among them, C infl Occ(x) represents the safety inflation cost of grid x, with a value range of [0,1]. Occ(x)=1 represents an obstacle grid, and its cost is directly set to the maximum value of 1. When the distance exceeds the safety inflation radius dmax, the cost decays to 0, eliminating the impact on distant grids. Based on the safety inflation cost calculation results of all grids, a safety inflation cost layer is generated.

[0091] S1045. Based on the binary occupancy layer, semantic cost layer, and security expansion cost layer, and with preset fusion weights and linear weighted fusion, the value of the fused full map is normalized, and maximum cost obstacle placement is performed on the grids corresponding to hard obstacles to generate a unified semantic cost map.

[0092] In some embodiments, a unified semantic cost map is obtained by linearly weighted fusion of a binary occupancy layer, a semantic cost layer, and a security expansion cost layer, thereby unifying geometric obstacle constraints, semantic risk constraints, and security distance constraints into a single cost field.

[0093] In some embodiments, linear weighted fusion is a multi-cost layer fusion method that sets corresponding fusion weights for different cost layers to balance the influence weights of each constraint in path planning.

[0094] In some embodiments, maximum cost obstacle setting: hard obstacle grids are forced to be set to the maximum cost of the entire graph, ensuring that obstacle grids are never selected in path planning, and avoiding the weakening of the impassable constraints of hard obstacles during the fusion and normalization process.

[0095] For example, in this embodiment of the invention, preset fixed fusion weights ω1, ω2, and ω3 can be set for the binary occupancy layer, semantic cost layer, and security inflation cost layer, respectively, with the sum of the weights being 1. Min-max normalization is performed on the grid cost values ​​of the three cost layers to unify the numerical range of the cost values ​​to [0,1], eliminating the numerical scale differences between different cost layers. Based on the preset fusion weights, linear weighted fusion is performed on the normalized three cost layers to calculate the initial fusion cost value of each grid. ; among which, J init (x) represents the initial fusion cost of raster x. The initial fusion cost of the entire map is normalized again to unify the total map cost to the [0,1] interval; simultaneously, maximum cost obstacle placement is performed on the hard obstacle raster marked as impassable in the binary occupied layer. Where J(x) is the final unified cost value of raster x, and J(x) = 1 is the maximum cost value of the entire map, representing absolutely impassable. A unified semantic cost map is generated based on the final unified cost values ​​of all rasters.

[0096] S105. Based on the unified semantic cost map, path planning is performed to obtain the optimal inspection route, enabling the UAV to perform autonomous inspection control.

[0097] As one possible implementation, step S105 can be specifically implemented as steps S1051-S1057.

[0098] S1051. Based on the unified semantic cost map, determine the starting point, ending point, and necessary nodes for fixed-point detection of key components along the route of the UAV inspection.

[0099] In some embodiments, the necessary nodes for the fixed-point inspection of key components of the line are: the drone waypoints corresponding to key components such as poles, insulators, conductor joints, and fittings in the power distribution line that must be photographed and inspected at close range.

[0100] For example, in this embodiment of the invention, based on the operational scope of the current power distribution line inspection task, the takeoff point of the UAV is determined as the inspection start point and the landing and recovery point as the inspection end point, and the three-dimensional coordinates of the two nodes in the world coordinate system are recorded. According to the actual route of the power distribution line and the tower numbering order, all necessary nodes are sorted sequentially to generate an ordered sequence of necessary nodes. The coordinates of the start point, ordered necessary nodes, and end point are mapped and transformed with the raster coordinate system of the unified semantic cost map to obtain the raster index corresponding to each node.

[0101] S1052. Based on the starting point, ending point, and necessary nodes of UAV inspection, and the eight-neighbor grid search mode, the path search space is constructed by decomposing the necessary nodes into continuous sub-path search units.

[0102] In some embodiments, such as Figure 5 As shown, the eight-neighbor grid search mode is a standard neighborhood model used in path planning. This means that each grid node can establish connections with adjacent grids in four orthogonal directions (up, down, left, and right) and four diagonal directions (upper left, upper right, lower left, and lower right). Compared with the four-neighbor model, it can better approximate a straight path.

[0103] In some embodiments, the sub-path search unit: takes two adjacent necessary nodes as the starting and ending points of an independent path search interval, decomposes the long-distance global path into multiple short-distance sub-paths, reduces the computational load of a single search, and ensures that all necessary nodes are completely covered.

[0104] In some embodiments, the path search space, consisting of all traversable grid nodes of the unified semantic cost map, the eight-neighborhood connections between nodes, and segmented sub-path units, is the carrier for the SE-A* algorithm to perform path search.

[0105] For example, embodiments of the present invention can use an eight-neighbor search mode, with mandatory nodes as segment nodes, to perform global path decomposition: First, the inspection starting point is taken as the first segment starting point, and the first node of the ordered mandatory node sequence is taken as the first segment ending point, generating the first sub-path search unit; Second, the ending point of the previous segment is taken as the starting point of the next segment, and the next ordered mandatory node is taken as the ending point of the next segment, generating consecutive sub-path search units; Third, the last ordered mandatory node is taken as the segment starting point, and the inspection ending point is taken as the segment ending point, generating the last sub-path search unit. Based on the two-dimensional grid plane of the unified semantic cost map, and combined with the eight-neighbor connection rules and segmented sub-path search units, a global path search space is constructed.

[0106] S1053. Based on the path search space and unified semantic cost map, within the framework of graph search-based heuristic path planning algorithms, semantic penalty terms and turning penalty terms are introduced to construct the cumulative cost function of the SE-A ​​algorithm.

[0107] In some embodiments, the cumulative cost function is the actual cumulative passage cost from the starting point to the current node, the semantic penalty term is related to the semantic cost layer and the security expansion cost layer of the unified semantic cost map, and the turning penalty term is related to the angle between the movement directions of adjacent nodes.

[0108] In some embodiments, the SE-A ​​algorithm, the core path planning algorithm, also known as Semantic-Enhanced A-star, introduces semantic penalty terms and turning penalty terms on the basis of the traditional A* algorithm framework, which can simultaneously take into account the semantic compliance, smoothness and global optimality of the path.

[0109] In some embodiments, the cumulative cost function g(n) is a component of the core evaluation function of the SE-A ​​algorithm, referring to the actual cumulative travel cost from the path start point to the current node n, and is the core object of the SE-A ​​algorithm optimization. Semantic penalty term: Incorporating the semantic cost value and safety inflation cost value of the unified semantic cost map into the path's cumulative cost, guiding the algorithm to prioritize low-risk safe areas and avoid high-risk areas, achieving semantically aware path planning. Turning penalty term: Calculating the turning cost based on the angle between the movement directions of adjacent nodes, suppressing frequent path turns, improving path smoothness, adapting to the kinematic characteristics of the UAV, and reducing acceleration, deceleration, and attitude adjustments during flight.

[0110] For example, embodiments of the present invention can be based on the classic A algorithm framework, constructing the core evaluation function of the SE-A ​​algorithm: f(n) = g(n) + h(n); where f(n) is the comprehensive evaluation function of node n, g(n) is the cumulative cost function from the starting point to the current node, and h(n) is the heuristic estimated cost from the current node to the destination. The cumulative cost function g(n) is constructed by introducing semantic penalty terms and turning penalty terms on the basis of traditional geometric movement costs: Among them, g geo (n): The basic geometric cost of moving from the starting point to the current node. Orthogonal movement costs 1, and diagonal movement costs... C sem (n): Semantic penalty term, whose value is the grid fusion cost of the unified semantic cost map corresponding to the current node, and is directly bound to the semantic cost layer and the security expansion cost layer; C turn (n): The turning penalty term, which is directly related to the angle between the moving directions of adjacent nodes; λ s , λ t These represent the semantic penalty weight and the turning penalty weight, respectively. The semantic penalty term is directly taken from the raster cost value of the unified semantic cost map, and the formula is as follows: C sem (n) = J(n); where J(n) is the final fusion cost of the raster corresponding to the current node n in the unified semantic cost map, realizing the direct binding of semantic cost and path cumulative cost, guiding the algorithm to avoid high-risk areas. The specific calculation steps of the turning penalty term are as follows: First, calculate the movement direction vector of the current node relative to its parent node. The movement direction vector of the parent node relative to its parent node. The second step is to calculate the angle Δθ between the two direction vectors using the dot product, as shown in the following formula: The third step is to calculate the steering penalty value based on the included angle, using the following formula: The larger the angle, the higher the turning penalty value; the penalty value is 0 when moving in the same direction. The cumulative cost function is constructed, clarifying the binding relationship between the semantic penalty term and the unified semantic cost map, and the binding relationship between the turning penalty term and the angle between the movement directions, thus forming the core cost calculation rules of the SE-A* algorithm.

[0111] S1054. Based on the cumulative cost function and the diagonal distance heuristic function that satisfies the adoptability condition, the node comprehensive evaluation function is obtained by summing them.

[0112] In some embodiments, the heuristic function h(n) is a core component of the A* algorithm, referring to the estimated cost from the current node n to the target node. It is used to guide the search direction, reduce unnecessary node expansion, and improve search efficiency.

[0113] In some embodiments, the admissibility condition is the core mathematical premise for Algorithm A to guarantee global optimality. It means that the estimated cost of the heuristic function is never greater than the actual minimum cost from the current node to the target node. As long as the admissibility condition is met, Algorithm A will definitely find the globally optimal path.

[0114] In some embodiments, the diagonal distance heuristic function adapts to the eight-neighbor grid search pattern, better fits the eight-neighbor movement rules than the Manhattan distance, and strictly satisfies the acceptability condition.

[0115] For example, embodiments of the present invention can set a diagonal distance heuristic function adapted to the eight-neighbor search pattern, strictly satisfying the admissibility condition, as shown in the following formula: Where dx and dy are the absolute distances between the current node and the target node on the X and Y axes of the grid, respectively; D is the unit cost of moving in the orthogonal direction, with a value of 1. The cost per unit of movement in the diagonal direction, taking a value of Verifying the admissibility of the heuristic function: The estimated value of the diagonal distance heuristic function is always less than or equal to the actual minimum movement cost from the current node to the target node, fully satisfying the admissibility condition of the A* algorithm, mathematically guaranteeing that the algorithm can find the globally optimal path. Summing the cumulative cost function g(n) with the heuristic function h(n) yields the node comprehensive evaluation function f(n), which serves as the sole criterion for determining the node expansion priority in subsequent path searches. The smaller the value of f(n), the higher the node search priority.

[0116] S1055. Within the global path search space, perform a global grid path search with the goal of minimizing the node comprehensive evaluation function to generate an initial inspection path.

[0117] In some embodiments, global grid path search is a graph search process based on the A algorithm. Prioritizing the minimization of the comprehensive cost function, it traverses all walkable grid nodes in the global path search space to find the path with the minimum cumulative cost from the starting point to the ending point. After completing the global search, the original path generated by backtracking through parent nodes is formed. This path consists of a continuous sequence of grid nodes and contains the complete spatial orientation of the path.

[0118] For example, step S1055 can be implemented as steps one through five.

[0119] Step 1: Initialize the open list and the closed list. Add the inspection starting node to the open list, and at the same time initialize the cumulative cost, heuristic cost, and node comprehensive evaluation function value of all grid nodes in the path search space.

[0120] In some embodiments, the OpenList is a core data structure of the A* algorithm, used to store raster nodes that have been discovered but not yet expanded, serving as a candidate pool for selecting the optimal expanded node in each round. The ClosedList is also a core data structure of the A* algorithm, used to store raster nodes that have already been expanded. Nodes added to the ClosedList do not participate in subsequent expansion calculations, avoiding redundant searches.

[0121] In some embodiments, the cost value is initialized: the cost value of all nodes in the search space is initialized, the cumulative cost value g(n) of the starting node is initialized to 0, and the g(n), h(n), and f(n) of other nodes are initialized to infinity, ensuring that only the starting node has search priority in the initial state.

[0122] For example, embodiments of the present invention can initialize two empty list structures: an open list and a closed list. Each element of the list stores the index, cost value, and parent node index information of a raster node. Global cost value initialization is performed: for all raster nodes in the path search space, the cumulative cost value g(n), heuristic cost value h(n), and comprehensive evaluation function value f(n) are all initialized to infinity (∞); for the inspection start node, its cumulative cost value g(start) is initialized to 0, its heuristic cost value h(start) is calculated based on the diagonal distance heuristic function, and the comprehensive evaluation function value f(start) = g(start) + h(start). The initialized inspection start node is added to the open list, completing the initialization preparation work before the search.

[0123] Step 2: Traverse the open list, select the grid node with the smallest node comprehensive evaluation function value as the current expansion node, remove the current expansion node from the open list, and add it to the closed list.

[0124] In some embodiments, the current expanding node is the node with the smallest comprehensive evaluation function value f(n) selected from the open list in each round of search. This node is the core node that needs to traverse its neighborhood and complete the expansion calculation in this round. Node expansion is the process of performing traversability verification, cost calculation, and parent node update on all eight neighboring nodes of the current node. This is the core action of the A* algorithm's search path.

[0125] For example, in this embodiment of the invention, all nodes in the current open list can be traversed, and the comprehensive evaluation function value f(n) of each node can be compared. The node with the smallest f(n) value is selected as the current expansion node for this round. If there are multiple nodes with the same f(n) value in the open list, the node with the smaller heuristic cost h(n) is selected as the current expansion node, prioritizing the search towards the direction closer to the endpoint to improve search efficiency. The selected current expansion node is removed from the open list and added to the closed list, marked as having completed expansion, and will not be processed repeatedly in subsequent searches.

[0126] Step 3: Traverse the eight neighboring grid nodes of the current extended node, perform a passability check on each neighboring node, calculate the cost value and update the parent node for the neighboring nodes that pass the check, and add the qualified neighboring nodes to the open list. In some embodiments, traversability verification is a compliance check to determine whether neighboring nodes can participate in path search. Parent node: records the node preceding the current node in the path; it is the core basis for backtracking and generating the complete path after the search is completed. The parent node of each node corresponds to the node preceding the optimal path leading to that node.

[0127] For example, in this embodiment of the invention, the eight neighboring grid nodes of the current extended node can be traversed, and a traversability check is performed on each neighboring node in sequence. The check rules are as follows: Rule 1: Whether the neighboring node is within the grid boundary of the path search space; if it exceeds the boundary, the check fails. Rule 2: Whether the neighboring node is in the closed list; if it is in the closed list, the check fails. Rule 3: Whether the neighboring node is an inaccessible hard obstacle grid marked by the unified semantic cost map; if it is a hard obstacle, the check fails. Only neighboring nodes that simultaneously satisfy the above three rules can pass the traversability check and enter the subsequent cost value calculation stage. For neighboring nodes that pass the check, the new cumulative cost value g from the starting point through the current extended node to that neighboring node is calculated. new The calculation rule adopts the cumulative cost function of the SE-A* algorithm, which includes geometric movement cost, semantic penalty term, and turning penalty term.

[0128] Perform cost value comparison and parent node update: If the new cumulative cost value gnew is less than the current cumulative cost value g(n) of the neighboring node, it means that the path to the neighboring node via the current expanding node is better; update the cumulative cost value g(n) = gnew of the neighboring node, recalculate the comprehensive evaluation function value f(n) = g(n) + h(n); set the parent node of the neighboring node as the current expanding node, and record the path association relationship. If the neighboring node is not currently in the open list, add it to the open list to participate in the next round of optimal node selection.

[0129] Step 4: Repeat steps 2 and 3 until the inspection endpoint node is added to the closed list or there are no nodes to be expanded in the open list.

[0130] For example, after completing the neighborhood traversal and update of each currently expanded node, the termination condition is immediately checked: First priority check: Check if the endpoint node has been added to the closed list. If it has, the search is immediately terminated, and the search is considered successful. Second priority check: Check if the open list is empty. If it is empty and the endpoint has not been added to the closed list, the search is immediately terminated, and the search is considered failed, as there is no feasible path to inspect. If the termination condition is not triggered, steps two and three are repeated to continue node expansion and searching until the termination condition is triggered.

[0131] Step 5: When the inspection endpoint node is added to the closed list, start from the endpoint node and backtrack along the parent node relationship to the starting node to generate the initial inspection path.

[0132] For example, when the search successfully terminates, the inspection endpoint node is used as the backtracking starting point. The parent node index of the endpoint node is read to obtain the previous path node. The backtracking continues along the parent node relationships, and each node is added to the path node sequence until the inspection starting point node is reached. The path node sequence obtained by backtracking is reversed to obtain a complete node sequence arranged in order from the starting point to the endpoint, which is the initial inspection path.

[0133] S1056. Perform three-point collinear redundant node removal processing on the initial inspection path to generate the optimal inspection route.

[0134] For example, in this embodiment of the invention, a continuous node sequence of the initial inspection path can be extracted, denoted as P=[p0,p1,p2,...,pn], where p0 is the starting point and pn is the ending point. The optimized path sequence is initialized by first adding the starting point p0 to the optimization sequence, setting the current reference node as p0, and the next node to be verified as p2. The continuous nodes of the initial path are traversed, and a three-point collinearity check is performed: for three consecutive nodes p1... 1. Using pi and pi+1, determine if the three nodes are on the same straight line. If the three points are collinear, the middle node pi is considered redundant and is removed. If the three points are not collinear, pi is considered a critical inflection point, retained, and the baseline node is updated to pi. After removing redundant nodes along the entire path, the retained critical inflection points and all necessary nodes are sequentially concatenated to generate the final optimal inspection route. The removal process does not change the cumulative cost and constraint compliance of the path, and fully inherits the global optimality of the initial path.

[0135] S1057. Convert the optimal inspection route into waypoint control commands that the UAV can execute, and control the UAV to perform autonomous inspection of power distribution lines.

[0136] In some embodiments, waypoint control commands are flight commands that the UAV flight control system can directly recognize and execute. These commands include parameters such as waypoint coordinates, flight altitude, cruise speed, gimbal attitude, and dwell time, and serve as the engineered execution carrier for the optimal inspection route.

[0137] For example, step S1057 can be specifically implemented as steps B1-B4.

[0138] B1. Analyze the path nodes of the optimal inspection route, extract the coordinates of the key inflection points of the route, match the necessary nodes for the fixed-point detection of key components of the route, and generate a basic waypoint sequence.

[0139] In some embodiments, the key inflection point, the node where the flight direction changes in the optimal inspection route, is the core control point for the UAV's route turning and determines the overall direction of the route.

[0140] In some embodiments, the basic waypoint sequence is a set of waypoints formed by sequentially piecing together key inflection points and necessary nodes. It is the core framework of the UAV flight path, with each waypoint corresponding to a three-dimensional spatial coordinate in the world coordinate system.

[0141] For example, this embodiment of the invention can analyze the complete path nodes of the optimal inspection route, separating two types of core waypoints: one type is key inflection points where the route direction changes, and the other type is essential nodes for the fixed-point detection of key components along the route. Coordinate deduplication is performed on both types of waypoints to eliminate redundant waypoints with overlapping spatial locations, preventing the UAV from repeatedly hovering in the same position. According to the flight sequence of the inspection route, the deduplicated waypoints are sorted sequentially, and a unique waypoint number is assigned to each waypoint, generating an ordered basic waypoint sequence. This ensures that the waypoint sequence progresses continuously along the route without turning back or abrupt changes. For each waypoint in the basic waypoint sequence, the corresponding waypoint type is marked: ordinary cruise waypoint (key inflection point) or fixed-point detection waypoint (essential node), providing a basis for subsequent flight control parameter configuration.

[0142] B2 configures flight control parameters for each waypoint in the base waypoint sequence.

[0143] In some embodiments, flight control parameters include flight altitude, cruise speed, stationary dwell time, and inspection gimbal attitude parameters.

[0144] In some embodiments, flight altitude: the relative ground / absolute altitude of the UAV at the waypoint. In power distribution inspection scenarios, an electrical safety distance must be maintained from power lines and towers. Cruise speed: the forward speed of the UAV between two waypoints, divided into cruise speed and deceleration speed near the waypoint, adapted to the UAV's kinematic characteristics. Stationary dwell time: the time the UAV hovers after reaching the stationary inspection waypoint, used to ensure that the gimbal camera can clearly capture key components, corresponding to the shooting requirements of the inspection task. Inspection gimbal attitude parameters: the pitch angle, roll angle, and yaw angle parameters of the UAV's onboard gimbal, used to control the camera's shooting angle and ensure that key components appear completely in the captured image.

[0145] B3. Encode the basic waypoint sequence and flight control parameters according to the preset communication protocol format to generate waypoint control commands for the UAV.

[0146] In some embodiments, a preset communication protocol format is used: a standard data communication protocol between the UAV flight control system and the ground station, commonly including the MAVLink protocol and the DJI Onboard SDK protocol, which is a standard format that the flight control system can recognize and execute commands. Command encoding: The waypoint sequence and flight control parameters are converted into binary / hexadecimal command data packets according to the frame structure, data format, and verification rules specified in the communication protocol, ensuring that the flight control system can correctly parse and execute them. Command verification: CRC cyclic redundancy check is performed on the generated waypoint control commands to ensure that no data errors occur during transmission, guaranteeing flight safety.

[0147] For example, in this embodiment of the invention, the three-dimensional coordinates, flight control parameters, and gimbal attitude parameters of each waypoint can be converted into numerical types and unit formats specified in the protocol, according to the protocol format requirements. The parameter data of all waypoints are encapsulated into a waypoint task data package as specified in the protocol, in order of waypoint number, and mandatory protocol fields such as frame header, frame trailer, device address, and task number are added.

[0148] B4. Based on waypoint control commands, control the UAV to perform autonomous cruise, conduct fixed-point detection of key components of the line and collect inspection data synchronously, and complete the autonomous inspection of power distribution lines.

[0149] For example, in this embodiment of the invention, the generated waypoint control commands can be sent to the UAV flight control system via a wireless data transmission link between the ground station and the UAV. After receiving the commands, the flight control system performs compliance and security checks, and sends back a confirmation signal after the checks are passed, loading the waypoint task. A task start command is sent to the UAV, controlling it to take off from the takeoff point, climb to a preset safe altitude, and then perform autonomous cruise flight along the waypoint sequence, strictly following preset speed, altitude, and transition modes to track the route. When the UAV reaches a designated detection waypoint, it automatically switches to hovering mode, adjusts the camera shooting angle according to preset gimbal attitude parameters, and completes multi-angle shooting of key components of the line for a preset dwell time, simultaneously recording the waypoint position and UAV attitude information corresponding to the shooting data. During the entire flight path, the UAV simultaneously triggers the lidar and visible light camera to collect real-time 3D point cloud data and 2D image data of the inspection scene at preset frequencies, completing the synchronous collection and storage of inspection data. When the UAV completes the flight and detection tasks of all waypoints and reaches the inspection endpoint, it automatically lands and recovers, completing the entire autonomous inspection operation of the power distribution line.

[0150] This invention provides a method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration. It employs an integrated optical-visual inspection architecture, simultaneously acquiring 3D point cloud data from a lidar radar and 2D image data from a visible light camera. After fusion preprocessing, key point cloud data of critical inspection targets are extracted, overcoming the inherent limitations of single visible light imaging, such as poor environmental adaptability and lack of 3D spatial information. The key point cloud data is precisely semantically segmented using a RandLA-CGNet network, resulting in a 3D semantic point cloud with multi-category semantic information. This leverages the 3D characteristics of the lidar point cloud to accurately quantify the spatial distance and net spatial gap of the inspection target, while also improving target recognition accuracy under complex lighting and adverse weather conditions through optical-visual fusion features. A unified semantic cost map is constructed based on the 3D semantic point cloud to complete path planning, achieving highly reliable autonomous UAV inspection and reducing inspection miss rates and operational safety risks.

[0151] Optionally, the UAV inspection method for intelligent power distribution lines based on optical-visual integration provided in this embodiment of the invention further includes steps S201-S207 before step S103.

[0152] S201. Construct the overall encoder-decoder architecture of the RandLA-CGNet network.

[0153] Figure 6 This invention provides a schematic diagram of the overall network structure of RandLA-Net. Figure 6 This is the basic prototype network of the RandLA-CGNet network of this invention. It adopts a classic symmetric encoder-decoder end-to-end architecture, with the original point cloud model as the input and the segmented point cloud model as the output. Figure 7 This invention provides a RandLA-CGNet network structure diagram. Figure 7 The network is a symmetrical encoder-decoder end-to-end architecture. The input is the original point cloud model of the power distribution line inspection scenario, and the output is the segmented point cloud model with point-by-point semantic labels.

[0154] In some embodiments, the encoder consists of five stacked downsampling modules for extracting multi-scale semantic features. Each layer contains a random downsampling unit (RS) and a feature aggregation unit. The first two downsampling modules are configured with a local feature aggregation module (LFA) to extract local geometric details of small targets such as thin wires and small insulators. The last three downsampling modules are configured with a local-global context fusion module (LGCF) to simultaneously extract local fine structure information and global line channel semantic information. After each downsampling layer, the number of point cloud points is compressed to 1 / 4 of the original, and the feature dimension is increased step by step. Finally, the deep encoded features that take into account both the details of small targets and the consistency of large scenes are output.

[0155] In some embodiments, the decoder consists of five stacked upsampling modules connected in skip connections to restore the spatial resolution of the point cloud and output point-by-point semantic predictions. Corresponding one-to-one with the five downsampling modules of the encoder, skip connections are set between corresponding levels of the encoder and decoder to fuse the shallow features of the encoder with the upsampling features of the decoder. After feature fusion at each layer of the decoder, a norm-gated channel feature module (NGCF) is inserted to adaptively recalibrate the feature channels, enhance key detail features, and suppress noise channels. After each layer of upsampling, the number of points in the point cloud is restored to four times the original, the feature dimension is reduced step by step, and finally the high-resolution point-by-point features with the same resolution as the input point cloud are output.

[0156] In some embodiments, the classification output branch: the high-resolution point-by-point features output by the decoder are passed through two fully connected layers (FC), a dropout layer (DP), and a classifier to finally output the semantic category prediction result for each point, thus completing the point cloud semantic segmentation.

[0157] In some embodiments, the encoder-decoder architecture is a classic symmetric architecture for semantic segmentation, where the encoder is responsible for downsampling to extract multi-scale semantic features, and the decoder is responsible for upsampling to restore the point cloud resolution.

[0158] In some embodiments, the downsampling module is a basic unit of the encoder. Each layer consists of a random downsampling module and a feature aggregation module. After each layer, the number of points is compressed to 1 / 4 of the original, and the feature dimension is increased by 1 time.

[0159] In some embodiments, the upsampling module is a basic unit of the decoder. Each layer consists of a trilinear interpolation upsampling module and a feature fusion module. After each layer, the number of points is restored to 4 times the original, and the feature dimension is reduced by 1 time.

[0160] For example, embodiments of the present invention can construct a symmetric encoder-decoder end-to-end architecture of the RandLA-CGNet network. The input is a 3D point cloud with visual features, and the output is a point-by-point semantic category prediction result. Encoder branch construction: A 5-layer stacked downsampling module is set, with a fixed random sampling rate of 1 / 4 for each layer. When the number of input points is N, the number of output points and feature dimensions of each layer satisfy the following formula: ;in, This represents the number of output points of the encoder's l-th layer. To correspond to the feature dimensions, multi-scale feature extraction from local details to global semantics is achieved through 5 layers of stacking.

[0161] Construct the decoder branch: Set up 5 stacked upsampling modules, corresponding one-to-one with the encoder's 5 downsampling layers. The number of output points and feature dimensions of each layer satisfy the following formula: ;in, This represents the number of output points of the decoder's layer l. To correspond to the feature dimension, the final output number of points is consistent with the original input number of points N.

[0162] To establish a skip connection channel between corresponding layers of the encoder and decoder, the output features of the encoder's layer l and the decoder's layer 6 are connected. The upsampled features from layer l are concatenated and fused using the following formula: ;in, For the features fused by skip connections, Concat( This involves concatenating channels to form a complete encoder-decoder architecture.

[0163] S202, the core module for building the encoder.

[0164] In some embodiments, such as Figure 8As shown, the encoder's first two layers include a local feature aggregation module for extracting local geometric details from the point cloud. This module employs a two-level cascaded feature aggregation architecture, taking point cloud features with visual characteristics as input and ultimately outputting aggregated high-dimensional local features. The specific structure is as follows: First-level aggregation link: The input features first pass through a shared fully connected layer (Shared MLP) to complete initial feature mapping, then the relative geometric relationship between the center point and neighboring points is explicitly encoded through a Local Spatial Encoding Unit (LocSE), and then the neighborhood features are weighted and aggregated through an Attentive Pooling Unit to enhance the feature representation of key geometric details; Second-level aggregation link: The features output from the first-level aggregation are again processed through the LocSE Unit and the Attentive Pooling Unit to perform secondary feature aggregation, further expanding the local receptive field and improving the ability to capture complex geometric structures; Residual fusion output: The features after the two-level cascaded aggregation are fused with the original input features of the module through residual connections, mapped through a shared fully connected layer, and the final local aggregated features are output, alleviating the gradient vanishing problem in deep network training while preserving the core geometric information of the original input.

[0165] In some embodiments, such as Figure 9 and Figure 10 The encoder shown has a local-global context fusion module in its last three layers. The local-global context fusion module has a built-in dynamic ratio context attention branch. The dynamic ratio context attention branch is used to inject global semantic information to generate encoded features that take into account both the details of small targets and the consistency of the large scene.

[0166] The Local-Global Context Fusion (LGCF) module includes local branches: employing Figure 8 The Local Feature Aggregation (LFA) module performs k-nearest neighbor search, local spatial encoding, and attention pooling on the input features, outputting local aggregated features that retain detailed information about fine wires and small hardware fittings. The global branch employs a Dynamic Ratio Contextual Attention (DRCA) module to perform global context modeling and adaptive channel reweighting on the input features, outputting globally enhanced features that inject overall semantic information across the scene. The adaptive fusion output combines the output features from the local and global branches through adaptive weighted fusion using a trainable channel fusion coefficient α. The fused features are then subjected to batch normalization (BN) and ReLU activation to output the final local-global fused features, balancing the contribution ratio of local details to global semantics.

[0167] The Dynamic Ratio Contextual Attention (DRCA) module includes the following stages: A context modeling stage: Global average pooling is performed on the input features along the point dimension to extract global contextual features for each channel, compressing the spatial dimension features into a channel-dimensional global descriptive vector, thus achieving aggregation of global semantic information. A feature reweighting stage: The aggregated global contextual features are processed through two 1×1 convolutions and a ReLU activation function to complete channel compression and restoration. The channel compression dimension is adaptively adjusted using a trainable dynamic ratio, replacing the traditional fixed compression ratio and balancing detail preservation and computational efficiency. Finally, an adaptive weight for each channel is generated using a Sigmoid activation function, multiplied channel-by-channel with the original input features to inject global semantic information and reweight the features, outputting the globally enhanced features.

[0168] For example, in the embodiments of the present invention, the first two layers of the encoder downsampling module are configured with a Local Feature Aggregation (LFA) module: each LFA module has a built-in two-level cascaded local spatial encoding and attention pooling unit, the neighborhood search k value is fixed at 16, and the formulas for local spatial encoding and attention pooling are as shown in step S103, ensuring the shallow network's ability to extract local geometric details. For the last three layers of the encoder downsampling module, a Local-Global Context Fusion (LGCF) module is configured, and each LGCF module has parallel local and global branches: the local branches have the same structure as the LFA modules and perform local feature aggregation using the formula in step S103; the global branches have a built-in Dynamic Ratio Context Attention (DRCA) branch, defining a trainable compression ratio vector r∈R. C The dimension is consistent with the number of input feature channels, and the compression ratio of each channel is adaptively controlled by the Sigmoid function, as shown in the following formula: ; where r compress For each channel's dynamic compression dimension, the Sigmoid function restricts the output to the [0,1] interval, ensuring the reasonableness of the compression dimension. A trainable channel fusion coefficient vector α∈R is set for the LGCF module. C The outputs of the local and global branches are adaptively weighted and fused using the formula in step S103. After fusion, the outputs are batch normalized and activated by LeakyReLU. The first two layers of the fixed encoder only perform local feature aggregation, while the last three layers only perform local-global context fusion, forming a hierarchical feature extraction architecture.

[0169] S203, the core module for building the decoder.

[0170] In some embodiments, after upsampling and skip connection feature fusion at each layer of the decoder, a norm-gated channel feature module is inserted. The norm-gated channel feature module is used to perform norm aggregation, gated recalibration, and residual fusion on the feature channels.

[0171] like Figure 11 As shown, the Norm-Gated Channel Feature Module (NGCF) is used to adaptively recalibrate the feature channels, dynamically amplify key feature channels, suppress noise channels, and improve the segmentation accuracy of detailed structures such as thin conductors and insulator boundaries. This module sequentially sets up three core units: global context embedding, channel normalization, and gating adaptation. The specific structure is as follows: Global Context Embedding Unit: Performs L1 / L2 norm aggregation on the input feature Finv along the point dimension to obtain the global response value of each channel, generating a channel-dimensional global context embedding vector to characterize the contribution of each channel to semantic segmentation; Channel Normalization Unit: Performs L2 norm normalization on the global context embedding vector to eliminate numerical scale differences between different channels, obtaining a normalized channel response vector; Gating Adaptive Unit: Constructs a gating function using trainable gating parameters γ and β, generates channel adaptive weights, and performs channel-by-channel recalibration on the input features; The recalibrated features and the original input features undergo residual fusion with fixed fusion weights λ, ultimately outputting the recalibrated high-resolution feature Fout, which both strengthens key detail features and avoids excessive perturbation to the original feature distribution, ensuring the stability of network training.

[0172] For example, in the embodiments of the present invention, a feature fusion unit can be set up for each layer of the decoder upsampling module: first, the shallow features and upsampled features of the corresponding layer of the encoder are aligned in channel dimension, and then feature fusion is completed by channel splicing.

[0173] After the feature fusion unit of each layer of the decoder, a norm-gated channel feature module (NGCF) is inserted. Each NGCF module is configured with a norm aggregation unit, a gated recalibration unit, and a residual fusion unit in sequence. The complete process adopts the formula in step S103.

[0174] The norm aggregation unit uses the L2 norm aggregation method and performs channel response value calculation and normalization according to the formula in step S103; the gating recalibration unit is set with two sets of trainable gating parameters γ∈R. C , β∈R C The gating function is constructed using the formula in step S103, and the channel weights are adaptively adjusted.

[0175] The residual fusion unit uses the optimal fusion weight λ=0.5 from the ablation experiment to perform residual fusion. After fusion, the output is activated by batch normalization and LeakyReLU, thus completing the construction of the core module of the decoder.

[0176] S204. Construct the loss function for the network.

[0177] In some embodiments, the loss function is a fusion-type FCE loss function obtained by linearly weighting and fusing the cross-entropy loss and the focus loss.

[0178] In some embodiments, the fusion-type FCE (Focused Cross Entropy Loss) is obtained by linearly weighting and fusing the cross-entropy loss (CE) and the focus loss (FL), mitigating the class imbalance problem. Cross-entropy loss (CE): a classic loss function for classification tasks, ensuring the overall convergence stability of the model. Focus loss (FL): an improvement on the cross-entropy loss, reducing the weights of easily classified samples and focusing on difficult-to-classify samples and minority class samples.

[0179] For example, embodiments of the present invention can construct a basic cross-entropy loss function Lce to calculate the overall classification loss: ; where t i,k Let ti,k be the true category label, where ti,k = 1 if the true category of the i-th point is k, and 0 otherwise; pi,k is the predicted category probability by the model, and K is the total number of semantic categories. Construct the focus loss function L. fl Strengthen the learning of difficult-to-distinguish samples and minority class samples: ; where α k For class balancing weights, γ is the focusing parameter, set γ=2, α k The loss weights for easily classified samples are adaptively adjusted based on the sample size by category, and the weights of these samples are reduced. A fusion-based FCE loss function is constructed using a linear weighted fusion method. The fusion weight λ adopts the optimal value of λ=0.5 from the ablation experiment to balance overall convergence and the learning effect of difficult samples. The fusion-type FCE loss function is set as the total loss function of the network and is used for gradient backpropagation and weight parameter updates during the training process.

[0180] S205. Based on the public point cloud dataset and the power distribution line-specific point cloud dataset, the RandLA-CGNet network is pre-trained and fine-tuned to obtain the network.

[0181] In some embodiments, pre-training uses a large-scale public dataset to initially train the network, learning its general ability to extract point cloud features. Fine-tuning: Based on the pre-trained weights, a secondary training is performed using a power distribution line-specific dataset with a small learning rate to adapt to power distribution inspection scenarios. Accuracy evaluation metrics: Overall classification accuracy (OA), average accuracy (mAcc), and average intersection-over-union ratio (mIoU) are used.

[0182] For example, in this embodiment of the invention, data augmentation and annotation can be performed on the publicly available S3DIS dataset and the power distribution line-specific point cloud dataset, divided into training set, validation set, and test set in a 7:2:1 ratio; semantic category ground truth is labeled for each point to construct training samples. A training environment is set up using the Adam optimizer, with an initial learning rate of 0.01, a learning rate decay rate of 0.95, a batch size of 6, and an epoch of 100. Pre-training is performed based on the S3DIS dataset, and the accuracy is verified using the validation set after each training round, saving the pre-training weights. The pre-training weights are loaded, the initial learning rate is adjusted to 0.001, the weights of the first two layers of the encoder are frozen, and transfer learning fine-tuning is performed using the power distribution line-specific dataset, with an epoch of 50. Training is terminated early when the validation set accuracy shows no improvement for 10 consecutive rounds. Accuracy verification: The model is tested using three core metrics, with the corresponding formulas as follows: Overall classification accuracy (OA): Average accuracy (mAcc): Mean Intersection over Union (mIoU): Where TPi is the number of true positive samples in class i, FPi is the number of false positive samples, FNi is the number of false negative samples, and K is the total number of semantic categories. Network weights that meet the accuracy requirements of power distribution inspection are stored, and a RandLA-CGNet network adapted to power distribution line inspection scenarios is generated.

[0183] Thus, by constructing a RandLA-CGNet network adapted for power distribution inspection, designing a dedicated feature extraction module and fusion loss function, and fine-tuning it through pre-training on a dataset, this invention significantly improves the segmentation accuracy of small targets and global semantic consistency, providing accurate semantic input for subsequent map construction and path planning.

[0184] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0185] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 500 includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program 503, it implements the steps in the above-described method embodiments. Alternatively, when the processor 501 executes the computer program 503, it implements the functions of each module / unit in the above-described device embodiments.

[0186] For example, the computer program 503 may be divided into one or more modules / units, which are stored in the memory 502 and executed by the processor 501 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 503 in the electronic device 500.

[0187] The processor 501 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0188] The memory 502 can be an internal storage unit of the electronic device 500, such as a hard disk or memory of the electronic device 500. The memory 502 can also be an external storage device of the electronic device 500, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 500. Furthermore, the memory 502 can include both internal and external storage units of the electronic device 500. The memory 502 is used to store the computer program and other programs and data required by the terminal. The memory 502 can also be used to temporarily store data that has been output or will be output.

[0189] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration, characterized in that, include: Collect 3D point cloud data from lidar and 2D image data from visible light cameras during the inspection process; The three-dimensional point cloud data and two-dimensional image data are fused and preprocessed to extract multiple key inspection targets such as conductors, towers, insulators, trees and buildings to obtain key point cloud data; The key point cloud data is semantically segmented based on the RandLA-CGNet network to obtain a three-dimensional semantic point cloud, which includes categories such as wires, towers, passage obstacles, and safe areas. The three-dimensional semantic point cloud is projected onto a two-dimensional grid plane to construct a unified semantic cost map, which includes an occupancy layer, a semantic cost layer, and a security expansion cost layer. Based on the unified semantic cost map, path planning is performed to obtain the optimal inspection route, enabling the UAV to perform autonomous inspection control.

2. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 1, characterized in that, The process involves fusing and preprocessing the 3D point cloud data and 2D image data to extract multiple key inspection targets, including conductors, towers, insulators, trees, and buildings, resulting in key point cloud data, including: Based on the aforementioned 3D point cloud data and 2D image data, as well as the pre-calibrated spatial mapping relationship between the UAV's lidar and visible light camera, timestamp synchronization and spatial coordinate alignment are performed to obtain spatiotemporally registered 3D point cloud data and 2D image data. The spatiotemporally registered 3D point cloud and 2D image are fused by optical and visual feature fusion. The texture, color, and grayscale features of the 2D image are mapped to each spatial point of the 3D point cloud, adding visual features to the point cloud to obtain the fused feature point cloud. The fused feature point cloud is subjected to two-stage filtering to obtain filtered point cloud data. The filtered point cloud data is subjected to dimensionality reduction processing using random downsampling to extract key point cloud data containing key inspection targets such as conductors, towers, insulators, trees, and buildings.

3. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 1, characterized in that, The semantic segmentation of the key point cloud data based on the RandLA-CGNet network to obtain a three-dimensional semantic point cloud includes: The key point cloud data is input into the encoder of the RandLA-CGNet network. Through the local feature aggregation module of the encoder, the local geometric structure features of the power distribution line point cloud are extracted to generate shallow multi-scale features containing details of fine conductors and small insulators. Through the local-global context fusion module of the encoder, local feature aggregation and dynamic ratio context attention calculation are performed in parallel, and local fine structure information and global line channel semantic information are extracted simultaneously to generate deep coding features that take into account both small target details and large scene consistency. The deep coding features are input into the decoder and combined with the shallow features of the corresponding layer of the skip connection fusion encoder. After feature fusion at each layer of the decoder, the feature channels are adaptively recalibrated through the norm-gated channel feature module to generate high-resolution pointwise features. Based on the high-resolution point-by-point features, the semantic category prediction results of each point are output through the point-by-point classification head, generating a three-dimensional semantic point cloud containing categories such as conductors, towers, passage obstacles, and safe areas.

4. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 1, characterized in that, Before performing semantic segmentation on the key point cloud data based on the RandLA-CGNet network to obtain the 3D semantic point cloud, the process also includes: The overall encoder-decoder architecture of the RandLA-CGNet network is constructed. The encoder consists of 5 stacked downsampling modules for extracting multi-scale semantic features. The decoder consists of 5 stacked upsampling modules with skip connections for restoring the spatial resolution of the point cloud and outputting point-by-point semantic predictions. The core module for constructing the encoder is set up as follows: the first two layers of the encoder are set up with a local feature aggregation module to extract local geometric details of the point cloud; the last three layers of the encoder are set up with a local-global context fusion module, which has a built-in dynamic ratio context attention branch; the dynamic ratio context attention branch is used to inject global semantic information to generate encoded features that take into account both small target details and large scene consistency. The core module of the decoder is constructed by inserting a norm-gated channel feature module after upsampling and skip connection feature fusion at each layer of the decoder. The norm-gated channel feature module is used to perform norm aggregation, gated recalibration and residual fusion of the feature channels. Construct a loss function for the network, wherein the loss function is a fusion-type FCE loss function obtained by linearly weighting and fusing the cross-entropy loss and the focus loss; The RandLA-CGNet network was obtained by pre-training and fine-tuning based on public point cloud datasets and power distribution line-specific point cloud datasets.

5. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 1, characterized in that, The step of projecting the three-dimensional semantic point cloud onto a two-dimensional grid plane to construct a unified semantic cost map includes: Based on the three-dimensional semantic point cloud and the resolution and planar coordinate boundary of the two-dimensional raster map, vertical projection is performed to realize the mapping transformation from three-dimensional spatial coordinates to two-dimensional raster index, and obtain two-dimensional semantic raster data. Based on the two-dimensional semantic raster data and the semantic categories corresponding to the three-dimensional semantic point cloud, the categories of impassable hard obstacles are marked, the number of hard obstacle points in each raster is counted, and the occupancy status of the raster is determined by combining the preset occupancy threshold to generate a binary occupancy layer. Based on the two-dimensional semantic raster data and the semantic categories corresponding to the three-dimensional semantic point cloud, different access risk levels are divided, and a corresponding semantic cost value is set for each raster to generate a semantic cost layer. Based on the binary occupancy layer, with the obstacle grid as the expansion source, the distance from each grid to the nearest obstacle grid is calculated through distance transformation, and a safe expansion cost layer is generated using an exponential decay function. Based on the binary occupancy layer, the semantic cost layer, the security expansion cost layer, and the preset fusion weights and linear weighted fusion, the fused full-map cost is normalized, and the grid corresponding to hard obstacles is subject to maximum cost obstacle placement to generate a unified semantic cost map.

6. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 5, characterized in that, Based on the two-dimensional semantic raster data and the semantic categories corresponding to the three-dimensional semantic point cloud, different traffic risk levels are divided, and a corresponding semantic cost value is set for each raster to generate a semantic cost layer, including: Based on the semantic categories of 3D semantic point clouds, safe areas are classified into low-risk levels, trees and passage obstacles into medium-risk levels, and wires, towers and buildings into high-risk levels. Set fixed semantic values ​​with progressively increasing gradients for low-risk, medium-risk, and high-risk levels; Traverse each grid cell of the two-dimensional semantic raster data and count the percentage of point clouds of different semantic categories within the grid cell; The semantic value of each grid is calculated based on the proportion of point cloud numbers of different semantic categories within each grid and the fixed semantic value of each risk level. A semantic cost layer is generated based on the semantic cost of each grid.

7. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 1, characterized in that, The process of performing path planning based on the unified semantic cost map to obtain the optimal inspection route and enabling the UAV to perform autonomous inspection control includes: Based on the unified semantic cost map, the starting point, ending point, and necessary nodes for fixed-point detection of key components along the route of the UAV inspection are determined. Based on the starting point, ending point, and necessary nodes of the UAV inspection, and the eight-neighbor grid search mode, the path search space is constructed by decomposing the necessary nodes into continuous sub-path search units. Based on the path search space and the unified semantic cost map, within the framework of the graph search-based heuristic path planning algorithm, a semantic penalty term and a turning penalty term are introduced to construct the cumulative cost function of the SE-A ​​algorithm. The cumulative cost function is the actual cumulative travel cost from the starting point to the current node. The semantic penalty term is related to the semantic cost layer and the safety expansion cost layer of the unified semantic cost map. The turning penalty term is related to the angle between the movement directions of adjacent nodes. The node comprehensive evaluation function is obtained by summing the cumulative cost function and the diagonal distance heuristic function that satisfies the adoptability condition. Within the global path search space, a global grid path search is performed with the goal of minimizing the node comprehensive evaluation function to generate an initial inspection path. The initial inspection path is processed to remove three collinear redundant nodes, and the optimal inspection route is generated. The optimal inspection route is converted into waypoint control commands that can be executed by the UAV, enabling the UAV to perform autonomous inspections of power distribution lines.

8. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 7, characterized in that, Within the global path search space, a global grid path search is performed with the objective of minimizing the node comprehensive evaluation function to generate an initial inspection path, including: Step 1: Initialize the open list and the closed list. Add the inspection starting node to the open list, and at the same time initialize the cumulative cost, heuristic cost, and node comprehensive evaluation function value of all grid nodes in the path search space. Step 2: Traverse the open list, select the grid node with the smallest comprehensive evaluation function value as the current expansion node, remove the current expansion node from the open list, and add it to the closed list; Step 3: Traverse the eight neighboring grid nodes of the current extended node, perform a passability check on each neighboring node, calculate the cost value and update the parent node for the neighboring nodes that pass the check, and add the qualified neighboring nodes to the open list. Step 4: Repeat steps 2 and 3 until the inspection endpoint node is added to the closed list or there are no nodes to be expanded in the open list. Step 5: When the inspection endpoint node is added to the closed list, start from the endpoint node and backtrack along the parent node relationship to the starting node to generate the initial inspection path.

9. The method for unmanned aerial vehicle (UAV) inspection of intelligent power distribution lines based on optical-visual integration according to claim 1, characterized in that, The process of converting the optimal inspection route into waypoint control commands executable by the UAV, and controlling the UAV to perform autonomous inspection of power distribution lines, includes: The path nodes of the optimal inspection route are analyzed, the coordinates of the key inflection points of the route are extracted, the necessary nodes for the fixed-point detection of key components of the route are matched, and a basic waypoint sequence is generated. Configure flight control parameters for each waypoint in the basic waypoint sequence. The flight control parameters include flight altitude, cruise speed, stationary dwell time, and inspection gimbal attitude parameters. According to the preset communication protocol format, the basic waypoint sequence and flight control parameters are encoded to generate waypoint control commands for the UAV; Based on the waypoint control commands, the UAV is controlled to perform autonomous cruise, fixed-point detection of key components of the line, and synchronous collection of inspection data to complete the autonomous inspection of the power distribution line.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor being configured to invoke and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 9.