Unmanned aerial vehicle-oriented in-line autonomous patrol flight control method and system

By collecting monocular time-series image sequences and attitude-motion data, region of interest extraction and inter-frame relative pose estimation are performed. A pseudo-multi-view image group is constructed, and a lightweight three-dimensional Gaussian representation model is established. This solves the problems of insufficient accuracy and heavy equipment in UAV inspection, and enables efficient inspection of UAVs autonomously flying along cables.

CN122632877APending Publication Date: 2026-08-25STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611131528.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing drone inspection methods cannot dynamically adjust according to the actual spatial shape of the cable, resulting in deviations between the inspection route and the cable position. Furthermore, LiDAR-based solutions are costly and heavy, making them difficult to deploy on small drone platforms, and their depth estimation accuracy is insufficient in long-distance cable observation scenarios.

Method used

By acquiring monocular time-series image sequences and attitude-motion data based on a preset observation window, region of interest extraction and inter-frame relative pose estimation are performed. A pseudo-multi-view image group is constructed, a lightweight three-dimensional Gaussian representation model is established, and flight control commands are generated by combining inertial measurement information to enable the UAV to fly autonomously along the cable path.

Benefits of technology

It enables drones to conduct stable and autonomous inspections in low-texture cable scenarios, reducing computational burden, adapting to onboard real-time processing, improving inspection accuracy and efficiency, and eliminating the need for human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122632877A_ABST
    Figure CN122632877A_ABST
Patent Text Reader

Abstract

The application discloses a UAV-oriented along-line autonomous patrol flight control method and system, relates to the technical field of UAV inspection, and comprises the following steps: collecting monocular time sequence images and attitude-motion data of a UAV based on a preset observation window; extracting a region of interest from the monocular time sequence images to obtain a mask time sequence image sequence, performing interframe relative pose estimation based on the attitude-motion data, and obtaining a relative pose sequence; combining the mask time sequence image sequence and the relative pose sequence to construct a pseudo multi-view image group, establishing a lightweight three-dimensional Gaussian representation model according to the pseudo multi-view image group, and obtaining cable observation position information; and generating flight control instructions according to the cable observation position information and real-time pose information provided by an inertial measurement unit, combining constraint conditions of a target inspection task, and controlling the UAV to autonomously fly along a cable path. The application solves the technical problems of low inspection efficiency, insufficient inspection accuracy and difficulty in adapting to an inspection environment in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology inspection, specifically to a method and system for autonomous patrol flight control of UAVs along a designated route. Background Technology

[0002] Using drones for routine inspections of cable infrastructure such as power transmission lines and optical fiber cables has become an important operation and maintenance method in the power and communications industries. Current drone inspection methods mainly rely on manually pre-setting waypoints to generate fixed inspection routes, with the drone flying along the predetermined route and collecting image data of the cables and the surrounding environment.

[0003] However, existing drone inspection methods rely on manually preset waypoints for inspection, which cannot be dynamically adjusted according to the actual spatial shape of the cable, and are prone to deviation between the inspection route and the actual position of the cable; inspection solutions based on lidar have high equipment costs and weight, making them difficult to deploy on small drone platforms; and inspection solutions based on pre-planned routes are limited by baseline length, resulting in insufficient depth estimation accuracy in long-distance cable observation scenarios. Summary of the Invention

[0004] This invention provides a method and system for autonomous inspection flight control along a route for unmanned aerial vehicles (UAVs), which solves the technical problems of low inspection efficiency, insufficient inspection accuracy, and difficulty in adapting to the inspection environment in the prior art.

[0005] In a first aspect, the present invention provides a method for autonomous patrol flight control along a route for unmanned aerial vehicles (UAVs), the method comprising: Based on a preset observation window, a monocular time-series image sequence of the UAV and the attitude-motion data at the corresponding time are collected; Region of interest extraction is performed on the monocular temporal image sequence to obtain a masked temporal image sequence, and inter-frame relative pose estimation is performed based on the pose-motion data to obtain a relative pose sequence; By combining the masked temporal image sequence with the relative pose sequence, a pseudo-multi-view image group is constructed, and a lightweight three-dimensional Gaussian representation model for the target cable is established based on the pseudo-multi-view image group to obtain the cable observation position information of the target cable. Based on the cable observation location information and the real-time pose information provided by inertial measurement, combined with the constraints of the target inspection task, flight control commands are generated to control the UAV to fly autonomously along the cable path.

[0006] Secondly, the present invention provides an autonomous patrol flight control system for unmanned aerial vehicles (UAVs) along a designated route, the system comprising: The attitude-motion data acquisition module is used to acquire monocular time-series image sequences of the UAV and attitude-motion data at corresponding times based on a preset observation window; The region of interest extraction and relative pose estimation module is used to extract the region of interest from the monocular temporal image sequence, obtain the masked temporal image sequence, and perform inter-frame relative pose estimation based on the pose-motion data to obtain the relative pose sequence. The cable observation location information acquisition module is used to combine the mask time sequence image sequence with the relative pose sequence to construct a pseudo multi-view image group, and based on the pseudo multi-view image group, to establish a lightweight three-dimensional Gaussian representation model for the target cable and obtain the cable observation location information of the target cable. The flight control command generation module is used to generate flight control commands based on the cable observation position information and the real-time pose information provided by inertial measurement, combined with the constraints of the target inspection task, to control the UAV to fly autonomously along the cable path.

[0007] One or more technical solutions provided in this invention have at least the following technical effects or advantages: This invention provides a method and system for autonomous patrol flight control of unmanned aerial vehicles (UAVs) along a cable route. First, through a preset observation window, the system synchronously acquires and aligns monocular temporal image sequences with attitude-motion data, providing a spatiotemporally consistent data foundation for subsequent inter-frame pose estimation and 3D reconstruction. Second, a lightweight semantic segmentation network is used to detect cable pixels and extract regions of interest, significantly reducing the subsequent processing scope to a local area of ​​the cable, lowering computational burden and eliminating background interference. Inter-frame relative pose sequences are obtained based on the integration and concatenation of inertial measurement data, exhibiting strong stability in low-texture cable scenes. Third, scene depth is recovered through depth estimation using mask consistency matching, transforming the monocular temporal image sequence into a pseudo-multi-view image group, providing structured input for 3D modeling. A lightweight 3D Gaussian representation model and cable-specific constraints are used to accurately represent the slender structure of the cable with fewer Gaussian units, achieving model lightweighting for airborne real-time processing. Finally, the cable observation location information is fused with real-time pose, and flight control commands for heading, altitude, and speed are generated through a hierarchical control strategy, driving the UAV to autonomously fly along the cable path, achieving autonomous patrol along the route without human intervention. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating the autonomous patrol flight control method for unmanned aerial vehicles (UAVs) provided by the present invention. Figure 2 This is a logical schematic diagram of the autonomous patrol flight control method for unmanned aerial vehicles (UAVs) provided by the present invention. Figure 3 This is a schematic diagram of the structure of the autonomous patrol flight control system for unmanned aerial vehicles (UAVs) provided by the present invention.

[0009] In the attached diagram, the components represented by each number are as follows: Attitude-motion data acquisition module 11; Region of interest extraction and relative pose estimation module 12; Cable observation position information acquisition module 13; Flight control command generation module 14. Detailed Implementation

[0010] This invention provides a method and system for autonomous inspection flight control along a route for unmanned aerial vehicles (UAVs), which solves the technical problems of low inspection efficiency, insufficient inspection accuracy, and difficulty in adapting to the inspection environment in the prior art.

[0011] The present invention will now be described in detail with reference to the accompanying drawings.

[0012] Example 1, as Figure 1 , Figure 2 As shown, this invention provides a method for autonomous patrol flight control along a route for unmanned aerial vehicles (UAVs), the method comprising: S100: Based on a preset observation window, collect monocular time-series image sequences of the UAV and attitude-motion data at corresponding times; In this embodiment of the invention, the preset observation window is a sliding window length pre-set in the time dimension, used to limit the number of historical frames participating in the current processing stage; the monocular time-series image sequence is a set of image frames with a clear temporal order obtained by continuously capturing images at a preset acquisition frequency using a monocular vision sensor mounted on the UAV. Compared with binocular or lidar solutions, monocular vision sensors have the advantages of being lightweight, having low power consumption, and being low cost, making them suitable for small rotary-wing UAV platforms that are sensitive to load; attitude-motion data refers to the three-axis angular velocity data and three-axis acceleration data synchronously output by the UAV's inertial measurement unit at the acquisition time of each image frame, used to describe the rotational and translational motion states of the UAV in three-dimensional space.

[0013] Specifically, according to a preset observation window length and acquisition frequency, the monocular time-series image sequence output by the monocular vision sensor and the three-axis angular velocity data and three-axis acceleration data output by the inertial measurement unit at the corresponding time are synchronously acquired at a fixed frame rate. After acquisition, the acquired monocular time-series image sequence is subjected to distortion correction, downsampling, and quantization compression processing in sequence. Finally, the processed monocular time-series image sequence is timestamped with the three-axis angular velocity data and three-axis acceleration data, so that each frame image is associated with the attitude-motion data at the same time, forming a time-series image sequence with attitude-motion data labels.

[0014] Step S100 in the method of this embodiment of the invention includes: The observation window length and acquisition frequency are preset, wherein the observation window length is defined as the number of historical frames retained within the sliding window; According to the acquisition frequency, the monocular time-series image sequence output by the monocular vision sensor and the triaxial angular velocity data and triaxial acceleration data output by the inertial measurement unit at the corresponding time are synchronously acquired at a preset frame rate. The monocular temporal image sequence is sequentially subjected to distortion correction, downsampling, and quantization compression processing; The monocular time-series image sequence, the three-axis angular velocity data, and the three-axis acceleration data are time-stamped to obtain the monocular time-series image sequence associated with the attitude-motion data.

[0015] In this embodiment of the invention, firstly, based on the UAV's flight speed, the computing power limitations of the onboard computing platform, and the spatial complexity of the target cable, the length of the observation window and the image acquisition frequency are preset. The observation window length is the number of historical frames retained within the sliding window. The preset observation window length is typically set in units of frames to ensure a constant amount of data within the window, facilitating time complexity analysis of the algorithm. The acquisition frequency determines the number of image frames acquired per unit time; a higher frequency results in higher temporal resolution, but also a larger data volume.

[0016] Secondly, according to the preset acquisition frequency, data acquisition is simultaneously triggered between the monocular vision sensor and the inertial measurement unit via hardware synchronization signals or software timestamp synchronization. The vision sensor outputs a monocular time-series image sequence, and the inertial measurement unit outputs triaxial angular velocity data and triaxial acceleration data at the same time.

[0017] Next, the monocular time-series image sequence is subjected to distortion correction, downsampling, and quantization compression processing in sequence.

[0018] Specifically, distortion correction is performed by mapping the original image to a distortion-free image space using the intrinsic parameter matrix and distortion coefficients obtained from camera calibration. Distortion correction eliminates radial and tangential distortion based on pre-calibrated camera distortion coefficients. Next, downsampling is performed to reduce the image resolution to a preset working resolution. Finally, quantization compression is performed to compress pixel precision from floating-point to integer precision, reducing the storage volume of image data. By executing these steps sequentially, distortion correction ensures geometric precision is prioritized, while downsampling and quantization compression reduce data volume while preserving effective information.

[0019] Finally, the preprocessed monocular temporal image sequence is aligned with the corresponding three-axis angular velocity data and three-axis acceleration data according to timestamps. Then, for each frame, the data frame with the closest timestamp is found and used as the attitude-motion data label for that frame. The final output is a temporal image sequence with precise attitude-motion data associated with each frame.

[0020] In this embodiment of the invention, by pre-setting an observation window and synchronously acquiring data, spatiotemporal alignment of images and inertial data is achieved, providing a reliable input foundation for subsequent processing. Distortion correction eliminates lens geometric distortion, ensuring pixel spatial mapping accuracy. Downsampling and quantization compression reduce data volume and computational burden, laying a data foundation for cable identification, 3D modeling, and autonomous flight control.

[0021] S200: Extract the region of interest from the monocular temporal image sequence to obtain the masked temporal image sequence, and perform inter-frame relative pose estimation based on the pose-motion data to obtain the relative pose sequence; In this embodiment of the invention, the region of interest is a local area in the image that contains the target cable. Since the cable typically occupies a small percentage of pixels in the overall image, the masked temporal image sequence is an image sequence generated by semantically segmenting each frame of the original monocular temporal image sequence, retaining only the pixels in the cable region while masking the background region. A mask is a binary image where the value of each pixel identifies whether the pixel belongs to the target category or the background category.

[0022] Specifically, based on a pre-trained lightweight semantic segmentation network, each frame in the monocular temporal image sequence is processed to predict the probability that each pixel belongs to the cable category. Then, the pixel-by-pixel classification probability map is thresholded to obtain a binary mask for the cable in the current frame. Based on this binary mask, the minimum bounding rectangle of the cable region is extracted, and the rectangle is expanded outwards by a preset margin along each side to form the region of interest bounding box. The region of interest image is obtained through spatial cropping, and the output is a masked temporal image. Simultaneously, the triaxial angular velocity data is integrated to obtain the inter-frame relative rotation matrix; the triaxial acceleration data after removing the gravity component is integrated twice to obtain the inter-frame relative translation vector. The inter-frame relative rotation matrix and the inter-frame relative translation vector are combined to form the inter-frame relative rigid body transformation matrix. Multiple inter-frame relative rigid body transformation matrices within a preset observation window are concatenated and accumulated to obtain the complete relative pose of each frame relative to a reference frame within the window, and the output is a relative pose sequence.

[0023] Step S200 in the method of this embodiment of the invention includes: Based on a pre-trained lightweight semantic segmentation network, each frame of the monocular temporal image sequence is processed to obtain a pixel-wise classification probability map. The pixel-by-pixel classification probability map is thresholded to obtain the cable binary mask of the current frame. In the cable binary mask, the region with a pixel value of a first preset value represents the cable region, and the region with a pixel value of a second preset value represents the background region. Based on the binary mask of the cable, the minimum bounding rectangle of the cable region is extracted, and along each side of the minimum bounding rectangle, a preset margin is extended outward to form the bounding box of the region of interest. Spatial cropping is performed on each frame of the image to obtain the region of interest image, and the output is a masked temporal image. Iteratively obtain the masked temporal image sequence corresponding to the monocular temporal image sequence; The inter-frame relative rotation matrix is ​​obtained by integrating the three-axis angular velocity data, and the inter-frame relative translation vector is obtained by double integration of the three-axis acceleration data after removing the gravity component. The inter-frame relative rigid body transformation matrix is ​​formed by the inter-frame relative rotation matrix and the inter-frame relative translation vector, and the relative pose sequence is obtained by cascading and accumulating multiple inter-frame relative rigid body transformation matrices within the preset observation window.

[0024] In this embodiment of the invention, firstly, based on a pre-trained lightweight semantic segmentation network, each frame of the monocular temporal image sequence is input into the pre-trained lightweight semantic segmentation network. The lightweight semantic segmentation network performs forward inference computation, outputting the probability value of each pixel in the monocular temporal image belonging to the cable category. The probability values ​​of all pixels constitute a pixel-by-pixel classification probability map of the same size as the input monocular temporal image, thus obtaining the pixel-by-pixel classification probability map. This probability map reflects the confidence distribution of the lightweight semantic segmentation network regarding which regions in the monocular temporal image belong to cables.

[0025] This lightweight semantic segmentation network employs an asymmetric encoder-decoder architecture. The encoder, based on a lightweight feature extraction backbone network, significantly reduces the number of parameters and computational cost by replacing standard convolutions with depthwise separable convolutions, progressively extracting multi-scale semantic features from the input image. An asymmetric convolution enhancement module is placed at the encoder's end, using multi-branch asymmetric convolution kernels to expand the kernel size in a single direction, enhancing the response capability to features along the direction of linear targets. Simultaneously, it combines dilated convolutions to expand the receptive field and introduces a channel attention mechanism to strengthen the response of key feature channels. The decoder performs upsampling through transposed convolutions or bilinear interpolation combined with depthwise separable convolutions, gradually restoring the spatial resolution of the feature map. It also improves segmentation accuracy by fusing shallow detail features and deep semantic features from the encoder through skip connections. After network training, the network is quantized and compressed to integer precision, significantly reducing inference latency and storage overhead while maintaining segmentation accuracy, meeting the real-time requirements of edge deployment on UAVs.

[0026] In one feasible embodiment, the lightweight semantic segmentation network employs an encoder-decoder architecture. The encoder uses a lightweight MobileNetV2 as its backbone feature extraction network, with its core building blocks being depthwise separable convolutional blocks with inverse residual structures. The encoder undergoes 3×3 convolutions, followed by a bottleneck layer for multi-level downsampling and feature extraction. The bottleneck layer uses 1×1 convolutions for dimensionality upsampling, 3×3 convolutions for spatial feature mixing, and 1×1 convolutions to reduce dimensionality to the bottleneck channel layer by layer, gradually reducing the spatial resolution of the feature map while gradually increasing the number of channels. An asymmetric convolution enhancement module is set at the end of the encoder. The asymmetric convolution kernel expands the receptive field in both the horizontal and vertical directions, and further increases the effective receptive field in conjunction with dilated convolutions. Channel attention mechanisms are used to enhance the response of channels related to linear targets, thereby enhancing the feature extraction capability for slender cable targets. The decoder performs upsampling through transposed convolutions or bilinear interpolation combined with depthwise separable convolutions, gradually restoring the spatial resolution to the original input size. Skip connections are used to fuse shallow detail features and deep semantic features from the encoder to improve the segmentation edge accuracy. After network training, the data is quantized and compressed to integer precision, significantly reducing inference latency and storage overhead while maintaining segmentation accuracy. Inputting a monocular temporal image, the encoder's bottleneck layer performs step-by-step downsampling and feature extraction. At the encoder's end, an asymmetric convolution enhancement module strengthens the feature response of the cable target. Finally, the decoder upsamples the data to the preset resolution, outputting a pixel-by-pixel classification probability map.

[0027] Secondly, the pixel-by-pixel classification probability values ​​are converted into a cable binary mask using a preset threshold. Thresholding the pixel-by-pixel classification probability map yields the cable binary mask for the current frame. Pixels with pixel-by-pixel classification probability values ​​higher than the preset threshold are classified as target categories and marked in the binary mask as regions with a first preset value, representing the cable region. Pixels with pixel-by-pixel classification probability values ​​lower than the preset threshold are classified as background categories and marked in the binary mask as regions with a second preset value, representing the background region.

[0028] Next, the cable region is extracted based on the binary mask of the cable. The pixel coordinates of all pixels marked as cable regions are extracted from the binary mask of the cable. The minimum and maximum values ​​of the pixels in the horizontal and vertical directions are calculated to determine the minimum bounding rectangle that can contain all cable pixels. The minimum bounding rectangle is then extended outward along the four sides of the minimum bounding rectangle by a preset margin to form the region of interest bounding box. Each frame of the image is spatially cropped according to the region of interest bounding box to obtain the region of interest image and obtain the mask temporal image.

[0029] Specifically, the preset margin refers to the number of pixels or relative proportion by which the minimum bounding rectangle expands outwards. The purpose is to preserve the cable area while incorporating adequate surrounding context information, avoiding insufficient spatial context in subsequent processing due to excessively close cropping to the cable. Offline experiments were conducted to statistically analyze the extreme values ​​of inter-frame displacement of the cable at different flight speeds, verifying that the margin is sufficient to cover fluctuations in cable position without introducing excessive background redundancy. Finally, the expansion amount is determined based on the selected fixed value or proportion, such as expanding by 50 pixels on each side or expanding the rectangle size by 10%.

[0030] Furthermore, the masked temporal image sequence corresponding to the monocular temporal image sequence is obtained iteratively.

[0031] Specifically, using the relative rotation matrix and translation vector obtained from inter-frame relative pose estimation, motion compensation is applied to the binary cable mask of the previous frame, transforming it to the current frame's image coordinate system to obtain a registered historical mask. Based on the pixel-by-pixel classification probability map output by the semantic segmentation network of the current frame, the average or maximum classification probability within the cable region of the current frame is calculated. This statistic is used as the detection confidence of the current frame, and the fusion weights of the historical mask and the current frame mask are adaptively determined based on this confidence. Finally, the motion-compensated binary cable mask of the previous frame is weighted and fused with the binary cable mask of the current frame to obtain the fused binary cable mask. The binary cable mask is calculated for each frame in the monocular temporal image sequence to obtain a mask temporal image sequence.

[0032] Furthermore, the triaxial angular velocity data is integrated to obtain the inter-frame relative rotation matrix, yielding the triaxial angular changes between adjacent frames. Simultaneously, the gravity component is subtracted from the triaxial acceleration data using a priori gravity direction data to obtain the linear acceleration generated by the UAV's own motion. This linear acceleration is then integrated twice: the first integration calculates the velocity change, and the second integration calculates the displacement change, obtaining the inter-frame relative translation vector. The priori gravity direction is determined by the accelerometer output when stationary or in uniform motion. Specifically, during the uniform linear motion segment before takeoff or during flight, triaxial acceleration data from the accelerometer is collected, filtered using a moving average to suppress measurement noise, and the direction of the filtered acceleration vector is determined as the priori gravity direction.

[0033] Finally, an inter-frame relative rigid body transformation matrix is ​​formed based on the inter-frame relative rotation matrix and the inter-frame relative translation vector. Then, taking a frame within the observation window as a reference frame, the transformation matrices between adjacent frames are multiplied in chronological order. Multiple inter-frame relative rigid body transformation matrices within the preset observation window are accumulated in cascade to obtain the complete relative pose of each frame within the window relative to the reference frame. The relative poses of all frames are arranged in chronological order to obtain the relative pose sequence.

[0034] Preferably, a baseline quality assessment is performed on the relative pose sequence. The relative pose sequence is a set of pose transformations of all other frames within the observation window relative to the first frame, using the first frame as a reference. For each frame in the relative pose sequence, the baseline length (i.e., the magnitude of the translation vector of that frame relative to the reference frame) is calculated, and the estimated distance from the cable to the UAV is obtained. The ratio of the baseline length to the estimated distance is calculated and used as the baseline quality evaluation index for that frame. When the ratio of a frame is lower than a preset baseline ratio threshold, the frame is determined to be too short to provide effective disparity information and cannot contribute sufficient geometric constraints to 3D reconstruction, and is therefore discarded. After traversing all frames in the relative pose sequence and discarding all frames that do not meet the baseline quality requirements, a valid relative pose sequence is obtained. The estimated distance is determined by the optimal depth estimate.

[0035] In step S200 of the method of this embodiment of the invention, a pseudo-multi-view image group is constructed by combining the masked temporal image sequence and the relative pose sequence, including: Based on historical patrol records, define a preset number of discrete depth assumptions; Under each discrete depth assumption, the relative pose sequence is traversed, and the corresponding pixels in the mask temporal image sequence are subjected to 3D backprojection and cross-frame reprojection by combining the camera intrinsic parameter matrix and the inter-frame relative rigid body transformation matrix to obtain discrete projection results. Based on the discrete projection results, the proportion of the projected pixels corresponding to each discrete depth assumption value falling within the mask area belonging to the cable region in the current frame is statistically calculated, and the mask consistency matching score of each discrete depth assumption value is calculated. The depth hypothesis with the highest matching score is selected as the optimal depth estimate for the current frame. The mask temporal image sequence, the relative pose sequence, and the optimal depth estimate are then combined to construct a pseudo-multi-view image group.

[0036] In this embodiment of the invention, firstly, a preset number of discrete depth assumptions are defined based on historical patrol records. These historical patrol records are flight and observation data accumulated by the UAV during previous patrol missions, including flight trajectories, actual spatial locations of cables, and distances between cameras and cables. The discrete depth assumptions are several candidate depth values ​​sampled uniformly or non-uniformly within a possible depth range, with each assumption representing a possible scene depth.

[0037] Specifically, based on the typical distance range from the camera to the cable statistically recorded in historical patrol records, such as the minimum to maximum distance of [15, 50] meters, samples are uniformly sampled at certain intervals and discretized into L levels to generate a preset number of discrete depth assumption values, where the value of L ranges from L∈[3, 5].

[0038] Secondly, for each discrete depth hypothesis, each frame in the relative pose sequence is traversed. For each cable region pixel in the temporal image of the reference frame mask, the pixel is back-projected from the image coordinate system to 3D space using the inverse of the camera intrinsic matrix and the discrete depth hypothesis, obtaining its 3D spatial coordinates under the discrete depth hypothesis. Then, using the inter-frame relative rigid body transformation matrix from the reference frame to the target frame, the 3D spatial coordinates are transformed to the camera coordinate system of the target frame, and then projected onto the image plane of the target frame using the camera intrinsic matrix of the target frame, obtaining the projected pixel coordinates. The above operation is performed on all cable region pixels to obtain the set of all projected pixels under the depth hypothesis, i.e., the discrete projection result.

[0039] Next, based on the discrete projection results, for each discrete depth hypothesis, the number of similar pixels falling within the area marked as the cable region in the binary mask of the target frame's cable is counted among all the projected pixels corresponding to that value. This number of pixels is divided by the total number of projected pixels to obtain the mask consistency matching score for that depth hypothesis. The scores from multiple target frames are averaged or the minimum value is taken as the final matching score for that depth hypothesis.

[0040] Specifically, for each discrete depth assumption value, the final matching score is calculated as follows: All pixels within the cable mask region of the reference frame are obtained. For each pixel, the assumed depth value is used to backproject the pixel from the image plane to 3D space. Then, based on the relative pose relationship between frames, the pixel is reprojected onto the image plane of the current frame to obtain the projected pixel position in the current frame. Subsequently, the value of the projected pixel position in the cable mask of the current frame is determined. A value of one indicates that the pixel falls within the mask region, while a value of zero indicates that the pixel falls within the background region. The spatial confidence weight of the projected pixel position relative to the center pixel of the cable mask region of the current frame is determined. This weight is quantized using a Gaussian kernel function, so that projected pixels closer to the mask center are assigned higher confidence, while those farther away have decreasing weights.

[0041] The product of the mask value and the confidence weight is then used as the matching contribution of the projected pixel to the current depth hypothesis. The matching contributions of all pixels within the cable mask region of the reference frame are summed, and then normalized by dividing by the total number of pixels within the cable mask region of the reference frame. The result is the matching score under this depth hypothesis. This score is essentially the average proportion of spatially confidence-weighted projected pixels falling within the cable mask region of the current frame. Specifically, the formula for calculating the final matching score of the depth hypothesis value is: Where K is the set of pixels within the cable mask region, S is the final matching score of the depth hypothesis, M is the value of the cable mask at the specified pixel position in the current frame, C is the confidence weight of the reprojection, and d is the typical distance range. These are the projected pixels of the cable mask area.

[0042] Finally, the highest score among all discrete depth hypothesis values ​​is selected, and its corresponding depth value is used as the global optimal depth estimate for the cable region in the current frame. The masked temporal image sequence, relative pose sequence, and optimal depth estimate are then combined to construct a pseudo-multi-view image group. This pseudo-multi-view image group is equivalent to a set of acquisition positions for each frame of an image with known relative pose relationships and each frame associated with depth information.

[0043] In this embodiment of the invention, a lightweight semantic segmentation network and extraction of the region of interest (ROI) from monocular temporal images reduce the monocular temporal images to be processed to a local area of ​​a cable, reducing computational burden and effectively eliminating background interference. Simultaneously, based on the integration and concatenation of inertial measurement data, an accurate inter-frame relative pose sequence is obtained. Subsequently, mask consistency matching is performed to search for the optimal depth estimate within the discrete depth assumption, effectively solving the problem of missing depth information in monocular vision. The temporal image sequence is then transformed into a pseudo-multi-view image set with pose and depth annotations, providing input data for the subsequent lightweight 3D Gaussian representation model.

[0044] S300: Combine the masked temporal image sequence with the relative pose sequence to construct a pseudo-multi-view image group, and based on the pseudo-multi-view image group, establish a lightweight three-dimensional Gaussian representation model for the target cable to obtain the cable observation position information of the target cable. In this embodiment of the invention, the pseudo-multi-view image group is a set of images that simulates multi-view stereoscopic visual effects, constructed from a temporal image sequence acquired by a monocular camera by means of inter-frame relative pose information and through three-dimensional back projection and cross-frame reprojection techniques; the lightweight three-dimensional Gaussian representation model is a three-dimensional scene representation method based on differentiable Gaussian splashing, which represents the geometric structure and appearance information of the three-dimensional scene by weighted superposition of a large number of parameterized three-dimensional Gaussian primitives.

[0045] Specifically, a pseudo-multi-view image set is constructed. A preset number of discrete depth assumptions are defined based on historical inspection records. The relative pose sequence is traversed, and the corresponding pixels in the masked temporal image sequence are subjected to 3D backprojection and cross-frame reprojection using the camera intrinsic matrix and inter-frame relative rigid body transformation matrix. Discrete projection results are obtained, and the proportion of projected pixels corresponding to each discrete depth assumption value falling within the masked area belonging to the cable region in the current frame is calculated. The mask consistency matching score is then calculated. The optimal depth estimate for the current frame is obtained. Finally, the masked temporal image sequence, the relative pose sequence, and the optimal depth estimate are combined to construct the pseudo-multi-view image set.

[0046] A lightweight 3D Gaussian representation model is established based on a pseudo-multi-view image set. The model utilizes the multi-view observation information provided by the pseudo-multi-view image set and performs iterative optimization through a differentiable Gaussian splashing framework to generate a Gaussian element set and extract the cable observation location information of the target cable.

[0047] Step S300 in the method of this embodiment of the invention includes: Establish cable-specific constraints and apply geometric restrictions to the initialization parameters of Gaussian elements based on the cable-specific constraints to perform sparse Gaussian element initialization; Combining the cable-specific constraints, a multidimensional loss function is constructed, and the initialization results of sparse Gaussian primitives are iteratively rendered and optimized to obtain a set of Gaussian primitives that meet the preset constraints. The output is a lightweight three-dimensional Gaussian representation model. Based on the lightweight three-dimensional Gaussian characterization model, adaptive spatial curve fitting with dynamic resolution is performed to obtain the cable observation location information of the target cable.

[0048] In this embodiment of the invention, firstly, cable-specific constraints are established, which include at least: anisotropic constraints, bounded curvature constraints, spatial continuity constraints, and sparse regularization constraints. Then, geometric constraints are applied to the initialization parameters of the Gaussian elements according to the cable-specific constraints, and sparse Gaussian element initialization is performed.

[0049] Specifically, based on the optimal depth estimate in the pseudo-multi-view image set, pixels in the masked temporal image sequence are back-projected into 3D space to obtain an initial 3D point set. This point set is then uniformly downsampled to obtain a sparse sampling point set. Combining cable-specific constraints, a Gaussian element is initialized centered on each sparse sampling point. The covariance matrix of each Gaussian element determines the local tangent direction based on the difference direction of adjacent sparse sampling points. This direction is used as the major axis to initialize the anisotropic ellipsoid shape. The opacity parameter is set to a preset initial opacity value, and the color representation uses the coefficients of a first-order spherical harmonic function that retains only the diffuse reflection component, thus completing the sparse Gaussian element initialization.

[0050] Secondly, by combining cable-specific constraints, a multidimensional loss function is constructed that includes a rendering consistency loss term and additional loss terms. The rendering consistency loss term measures the difference between the rendered image and the real mask temporal image; the additional loss terms include at least one or more of the following: elongation regularization loss term, curvature regularization loss term, spatial continuity regularization loss term, and sparse regularization loss term.

[0051] Subsequently, starting with the initialization results of sparse Gaussian primitives and using a multidimensional loss function as the optimization objective, iterative rendering optimization is performed on the position parameters, covariance matrix parameters, opacity parameters, and color representation of each Gaussian primitive for a preset number of iterations. After the iterative rendering optimization is completed, Gaussian primitives with opacity parameters below a preset opacity threshold are removed, and a set of Gaussian primitives satisfying preset constraints is obtained. The position parameters, covariance matrix parameters, and opacity parameters of each Gaussian primitive are then quantized and compressed to output a lightweight 3D Gaussian representation model. Here, Gaussian primitives are the basic building blocks of the 3D Gaussian representation model, and each primitive is defined by parameters such as position parameters, covariance matrix, opacity parameters, and color representation.

[0052] Finally, based on a lightweight 3D Gaussian characterization model, adaptive spatial curve fitting with dynamic resolution is performed to obtain the cable observation location information of the target cable.

[0053] Specifically, the 3D center positions of each Gaussian element are extracted from the lightweight 3D Gaussian representation model to form a 3D point set for the cable centerline. The fitting residual is calculated for this 3D point set, representing the distance error from the point set to the fitted curve. Based on the fitting residual, an adaptive spatial curve model is selected for fitting: when the fitting residual is below a first preset threshold, a linear fitting model along the principal direction is used; when the fitting residual is below a second preset threshold but above the first preset threshold, a quadratic polynomial curve fitting model is used; and when the fitting residual is above the second preset threshold, a B-spline curve fitting model is used to obtain the 3D spatial centerline. The first preset threshold is less than the second preset threshold.

[0054] Subsequently, based on the fitted 3D spatial centerline, the current tracking point and forward-looking tracking point of the target cable are extracted, and the corresponding cable observation position information is determined. The current tracking point is the closest projection point of the UAV's current position on the 3D spatial centerline, offset by a preset flight altitude along the normal direction; the forward-looking tracking point is the point at a preset forward-looking distance along the 3D spatial centerline, offset by a preset flight altitude along the normal direction.

[0055] In the method of this embodiment of the invention, step S300 involves establishing cable-specific constraints and applying geometric restrictions to the initialization parameters of Gaussian elements based on the cable-specific constraints, thereby performing sparse Gaussian element initialization, including: The cable-specific constraints are established, and the cable-specific constraints include at least the following: Anisotropic constraints are used to constrain the spatial shape of Gaussian elements to be an ellipsoid stretched along a preset direction; The curvature bounded constraint is used to ensure that the spatial discrete curvature formed by the center positions of adjacent Gaussian elements does not exceed a preset curvature upper limit. Spatial continuity constraint is used to ensure that the distance between the center positions of adjacent Gaussian elements does not exceed a preset maximum distance; Sparse regularization constraints are used to constrain the Gaussian element opacity parameter to tend to zero in order to achieve adaptive sparsity. Based on the optimal depth estimate in the pseudo-multi-view image group, the pixels in the masked temporal image sequence are back-projected into three-dimensional space to obtain an initial three-dimensional point set, and the initial three-dimensional point set is uniformly downsampled to obtain a sparse sampling point set. Based on the cable-specific constraints, a Gaussian primitive is initialized with each sparse sampling point as the center, and the sparse Gaussian primitive initialization result is obtained. In this process, the covariance matrix of each Gaussian element determines the local tangent direction based on the difference direction of adjacent sparse sampling points, and the anisotropic ellipsoid shape is initialized with the local tangent direction as the major axis direction. The opacity parameter of each Gaussian element is set to a preset initial opacity value, and the color representation of each Gaussian element adopts the first-order spherical harmonic function coefficients that retain only the diffuse reflection component.

[0056] In this embodiment of the invention, firstly, cable-specific constraints are established, which include at least: Anisotropy constraints constrain the spatial shape of Gaussian primitives to be an ellipsoid stretched along a predetermined direction. In a lightweight 3D Gaussian representation model, each Gaussian primitive contains a covariance matrix. The eigenvalues ​​and eigenvectors of this matrix determine the degree of stretching and orientation of the Gaussian primitive in 3D space. Therefore, anisotropy constraints require the covariance matrix to have significantly unequal eigenvalues, causing the primitive to be stretched significantly along the cable's direction and compressed along the cross-sectional direction, precisely matching the slender shape of the cable.

[0057] Bounded curvature constraints are used to ensure that the spatial discrete curvature formed by the centers of adjacent Gaussian elements does not exceed a preset upper limit. In three-dimensional space, the direction of a cable changes gradually, and its spatial curvature is limited by the material's mechanical properties and the suspension method, thus having a physical upper limit. Bounded curvature constraints calculate the discrete curvature of the spatial curve formed by the centers of three or more adjacent Gaussian elements and limit it to not exceeding a preset threshold, ensuring that the reconstructed cable model remains geometrically smooth and free of abnormal bends.

[0058] Spatial continuity constraints are used to ensure that the distance between the centers of adjacent Gaussian elements does not exceed a preset maximum distance. Cables are continuous, unbroken solid structures in physical space, and their three-dimensional centerlines should not have gaps exceeding a certain distance. Spatial continuity constraints ensure that Gaussian elements are continuously distributed along the entire cable by setting an upper limit on the Euclidean distance between the centers of adjacent Gaussian elements, avoiding model breaks caused by data sparsity or optimization instability.

[0059] Sparse regularization constraints are used to constrain the opacity parameter of Gaussian primitives to tend towards zero to achieve adaptive sparsity. In the Gaussian splashing framework, the opacity parameter controls the contribution of primitives to the final rendering result; Gaussian primitives with opacity approaching zero have no effect on scene representation. Sparse regularization constraints, by introducing an additional penalty term into the optimization objective of the opacity parameter, encourage unnecessary Gaussian primitives to reduce their opacity to near zero during the optimization process, thereby reducing the number of primitives.

[0060] Specifically, the cable-specific constraint Z = λ1 × Z 各向异性 +λ2×Z 曲率有界 +λ3×Z 空间连续 +λ4×Z 稀疏 λ1, λ2, λ3, and λ4 are the weight coefficients corresponding to the four constraints. Based on prior knowledge of typical cable inspection scenarios, combined with historical flight data and statistical characteristics of cable geometric parameters, the baseline values ​​of each constraint can be initially set. A grid search can be performed using a simulation environment or offline dataset, with accuracy as the optimization objective, to iteratively fine-tune the weight coefficients, ensuring that the regularization constraints reach a balance during the optimization process.

[0061] Secondly, based on the optimal depth estimate in the pseudo-multi-view image group, the pixels in the masked temporal image sequence are back-projected into three-dimensional space to obtain an initial three-dimensional point set, and the initial three-dimensional point set is uniformly downsampled to obtain a sparse sampling point set.

[0062] Specifically, based on the optimal depth estimate, each pixel within the cable mask region of the current frame undergoes 3D backprojection through the inverse transformation of the camera intrinsic matrix. Combined with the extrinsic transformation matrix from the camera coordinate system to the body coordinate system, the 3D points obtained from the backprojection are transformed from the camera coordinate system to the body coordinate system. For all pixels marked as cable regions in each frame, the 2D coordinates of that pixel on the image plane and the corresponding optimal depth estimate for that frame are obtained. Using the camera intrinsic matrix, the backprojection formula is applied... Where (u,v) are two-dimensional coordinates, d is the optimal depth estimate, and T k→cam The extrinsic transformation from the camera coordinate system to the body coordinate system is defined, and K is the camera intrinsic parameter matrix. Subsequently, an initial set of 3D points corresponding to the pixels within the masked region is obtained. This initial 3D point set is derived from the extrinsic transformation from the camera coordinate system to the body coordinate system and the pixels within the cable masked region, 1 / The product of the three.

[0063] To reduce the computational complexity and number of model parameters in subsequent Gaussian modeling, the initial 3D point set is spatially uniformly downsampled. Specifically, according to a preset pixel interval, every few pixels within the mask region, a corresponding 3D point is selected. This extracts a significantly reduced number of sparsely sampled points from the dense initial 3D point set, while maintaining a uniform spatial distribution. The number of Gaussian elements in this sparsely sampled point set is controlled within a preset range based on the actual scene complexity. Where [X] is the sparse sampling point set, J is the ratio of the total number of pixels in the cable mask area to the pixel spacing parameter, and n p s represents the total number of pixels within the cable mask area. p The pixel interval is [number].

[0064] Next, taking into account the cable-specific constraints, a Gaussian primitive is initialized with each sparse sampling point as the center to obtain the sparse Gaussian primitive initialization result.

[0065] For position parameter initialization, the spatial coordinates of each 3D point in the sparse sampling point set are used as the initial values ​​of the corresponding Gaussian elements. For covariance matrix initialization, for each sparse sampling point, the spatial difference vector between it and its neighboring sparse sampling points is first calculated. The direction of this difference vector is determined as the local tangent direction at that point, and the covariance matrix is ​​initialized using this local tangent direction as the major axis direction. Specifically, the unit vector of the local tangent direction is used as the first principal component direction of the covariance matrix. Two orthogonal directions are arbitrarily selected in the plane perpendicular to the tangent direction as the second and third principal component directions, and the eigenvalues ​​in the three directions are set as preset major axis radius and minor axis radius values, respectively.

[0066] For opacity parameter initialization, the opacity parameters of all Gaussian elements are set to preset initial opacity values, which are then adaptively adjusted through iterative optimization. For color representation initialization, the original image pixel color value corresponding to each sparse sampling point is obtained. This color value is used as the initial value of the zeroth-order component in the coefficients of the first-order spherical harmonic function. The first-order component is initialized to zero, meaning only the diffuse component is retained, ignoring higher-order illumination variations related to the viewing angle. The original image pixel color value is the average of multiple frames in the pseudo-multi-view image group. After completing the above configuration, the parameter set of all Gaussian elements is the sparse Gaussian element initialization result.

[0067] In this process, the covariance matrix of each Gaussian element determines the local tangent direction based on the difference direction of adjacent sparse sampling points. The anisotropic ellipsoid shape is initialized with the local tangent direction as the major axis direction. The opacity parameter of each Gaussian element is set to a preset initial opacity value, and the color representation of each Gaussian element uses the coefficients of a first-order spherical harmonic function that retains only the diffuse reflection component. The standard deviation along the major axis direction in the initialized anisotropic ellipsoid shape is set to a preset larger value, corresponding to the major axis radius; the standard deviation perpendicular to the major axis direction is set to a preset smaller value, corresponding to the minor axis radius.

[0068] In one embodiment, assume the spatial coordinates of the i-th point in the sparse sampling point set are (10, 5, 3), its preceding neighbor is (9, 9, 5, 3), and its following neighbor is (10, 1, 5, 3). The difference direction is the positive x-axis direction, which is taken as the local tangent direction. Set the major axis radius to 0.3m, the minor axis radius to 0.03m, and the eigenvalue of the covariance matrix in the x-direction to 0.3. 2 =0.09, eigenvalues ​​in the y and z directions are 0.03. 2 =0.0009. The initial opacity value is set to 0.5. The average color of the multi-frame corresponding to this point is (R=128,G=120,B=110). (128,120,110) is used as the zeroth order spherical harmonic coefficients, and a total of 400 Gaussian elements are initialized to form the sparse Gaussian element initialization result.

[0069] In step S300 of the method in this embodiment of the invention, a multidimensional loss function is constructed in conjunction with the cable-specific constraints. Iterative rendering optimization is performed on the sparse Gaussian primitive initialization results to obtain a set of Gaussian primitives that satisfy preset constraints. The output is a lightweight three-dimensional Gaussian representation model, including: Construct a multidimensional loss function that includes a rendering consistency loss term and an additional loss term, wherein the additional loss term includes at least one of the following: a slenderness regularization loss term, a curvature regularization loss term, a spatial continuity regularization loss term, and a sparse regularization loss term. Starting with the Gaussian initialization parameters in the sparse Gaussian initialization result, and with the multidimensional loss function as the optimization objective, the position parameters, covariance matrix parameters, opacity parameters, and color representation of each Gaussian are iteratively optimized for a preset number of iterations. After optimization and iteration, Gaussian elements with opacity parameters lower than the preset opacity threshold are removed, the Gaussian element set is obtained, and the position parameters, covariance matrix parameters, and opacity parameters of each Gaussian element are quantized and compressed to output the lightweight three-dimensional Gaussian representation model.

[0070] In this embodiment of the invention, firstly, a multidimensional loss function is constructed that includes a rendering consistency loss term and an additional loss term, wherein the additional loss term includes at least one of a slenderness regularization loss term, a curvature regularization loss term, a spatial continuity regularization loss term, and a sparse regularization loss term.

[0071] Specifically, rendering consistency loss characterizes the difference between the cable masks rendered from each pseudo-multi-view direction and the actual cable masks of the corresponding historical frames, as well as the pixel-level difference between the rendered image and the actual image. Its specific form can be L1 loss, L2 loss, or structural similarity loss, etc.; slenderness regularization loss characterizes the degree to which the ratio of the minimum to the maximum eigenvalue of each Gaussian covariance matrix exceeds a preset slenderness threshold. Its specific calculation method can be the reciprocal of the condition number of the covariance matrix or the square of the difference between the ratio of the major and minor axes and a preset target value; curvature regularization loss... The spatial continuity regularization loss is used to characterize the degree to which the second derivative norm of the discrete space formed by the center positions of adjacent Gaussian elements exceeds a preset upper limit of curvature. Specifically, it can be calculated by applying an additional quadratic penalty to the discrete curvature values ​​that exceed the upper limit of curvature. The spatial continuity regularization loss is used to characterize the degree to which the distance between the center positions of adjacent Gaussian elements exceeds a preset maximum distance. Specifically, it can be calculated by applying an additional quadratic penalty to the adjacent distances that exceed the maximum distance. The sparsity regularization loss is calculated by summing the opacities of each element in the form of the L1 norm, which characterizes the L1 norm of the opacity parameter of each Gaussian element.

[0072] Multidimensional loss function L t =λ 渲染一致性 ×L 渲染一致性 +λ 细长性正则化 ×L 细长性正则化 +λ 曲率正则化 ×L 曲率正则化 +λ 空间连续正则化 ×L 空间连续正则化 +λ 稀疏正则化 ×L 稀疏正则化 Wherein, λ 渲染一致性 , λ 细长性正则化 , λ 曲率正则化 , λ 空间连续正则化 , λ 稀疏正则化 These are the weighting coefficients for each loss, used to adjust the relative importance of each loss in the overall optimization objective. The sum of the weighting coefficients is 1. The values ​​of each weighting coefficient can be optimized according to the actual scenario through methods such as grid search or Bayesian optimization, or determined based on historical inspection data.

[0073] Secondly, starting with the Gaussian initialization parameters in the sparse Gaussian initialization results, and with the multidimensional loss function as the optimization objective, the position parameters, covariance matrix parameters, opacity parameters, and color representation of each Gaussian are iteratively optimized for a preset number of iterations.

[0074] Specifically, in each iteration, the current Gaussian set is input into the differentiable renderer. Starting from the camera poses of each frame in the pseudo-multi-view image set, a 2D image at the corresponding viewpoint is rendered. Then, the rendering consistency loss between the rendered image and the real mask temporal image is calculated, along with various regularization losses. All loss terms are weighted and summed to obtain the multidimensional loss function value. Next, the gradient of the loss function with respect to each Gaussian parameter is calculated through backpropagation. Finally, the optimizer updates each parameter based on the gradient information, causing the loss function value to decrease. The above process is repeated for a preset number of iterations, such as 100 to 200 times, which can be dynamically adjusted according to the convergence results. During the optimization process, the covariance matrix must always maintain positive semi-definiteness, which can be ensured by parameterizing it as a combination of rotation matrix and scaling vector; the opacity parameter must be kept within the range of 0 to 1, which can be constrained by parameterization using the Sigmoid function.

[0075] Preferably, the differentiable rendering iterative optimization further includes: adaptively classifying the rendering precision level based on the spatial distance between the center position of each Gaussian primitive and the optical center of the camera from the current rendering viewpoint. Higher rendering precision is assigned to Gaussian primitives closer to the camera, and lower rendering precision is assigned to Gaussian primitives farther away, in order to reduce redundant computation in distant areas while ensuring rendering quality in the near areas.

[0076] The region of interest (ROI) images corresponding to each frame in the pseudo-multi-view image group are divided into multiple parallel rendering sub-regions. During rendering calculation, only Gaussian primitives whose center positions are projected to fall within the range of each parallel rendering sub-region are associated with that sub-region. Then, a frustum visibility determination is performed on the associated Gaussian primitives, and Gaussian primitives located outside the frustum range of the current rendering view are removed, thereby avoiding the rendering overhead of invalid primitives. The remaining Gaussian primitives are simplified and sorted according to the depth direction, and the effective depth layer within the same parallel rendering sub-region is limited to a preset maximum occlusion layer. That is, Gaussian primitives with depth values ​​exceeding this layer threshold are directly discarded and do not participate in the rendering calculation of the current view. Finally, the Gaussian primitive sets of each sub-region after the above processing are composited and rendered to output a rendering mask and a depth map.

[0077] Finally, after optimization iterations, all Gaussian primitives are traversed, and the opacity parameter value of each primitive is checked. Primitives with opacity below a preset opacity threshold are removed from the model, resulting in a Gaussian primitive set. The position parameters, covariance matrix parameters, and opacity parameters of each Gaussian primitive are then quantized and compressed to output a lightweight 3D Gaussian representation model. The opacity threshold is a pre-defined lower limit for the opacity parameter, used to determine which Gaussian primitives contribute negligibly to the scene representation. After optimization iterations, Gaussian primitives with opacity parameters below this threshold have minimal weighted contributions during rendering and can be safely removed without affecting model accuracy. The opacity threshold can be determined by averaging, for example, 0.05 or 0.1, based on a labeled dataset of typical cable inspection scenes.

[0078] Specifically, during quantization compression, the output of the lightweight 3D Gaussian representation model also includes parameter simplification and storage optimization of the Gaussian element set: the position parameters of each Gaussian element are compressed from 32-bit floating-point numbers to 16-bit floating-point numbers to achieve simplified positional accuracy; the opacity parameters of each Gaussian element are compressed from 32-bit floating-point numbers to 8-bit fixed-point numbers to reduce storage volume; the covariance matrix of each Gaussian element is decomposed into rotation quaternions and scaling vectors, where rotation quaternions and scaling vectors are stored using 16-bit floating-point numbers, replacing the original nine-component storage method of the three-by-three covariance matrix, thereby significantly reducing storage overhead while retaining the complete geometric information of the covariance matrix; after the above compression and optimization processing, the lightweight 3D Gaussian representation model is output.

[0079] Step S300 in the method of this embodiment of the invention involves performing adaptive spatial curve fitting based on the lightweight three-dimensional Gaussian representation model to obtain the cable observation location information of the target cable, including: The three-dimensional center positions of each Gaussian element are extracted from the lightweight three-dimensional Gaussian representation model to form a three-dimensional point set of the cable centerline; Calculate the fitting residual for the three-dimensional point set of the cable centerline, and adaptively select a spatial curve model for fitting based on the fitting residual to obtain the three-dimensional spatial centerline, including: When the fitting residual is lower than the first preset threshold, the main direction linear fitting model is adopted; When the fitting residual is lower than the second preset threshold and higher than the first preset threshold, a quadratic polynomial curve fitting model is used. When the fitting residual is higher than the second preset threshold, a B-spline curve fitting model is used, wherein the first preset threshold is less than the second preset threshold; Based on the three-dimensional space centerline, the current tracking point and the forward tracking point are extracted and the corresponding cable observation position information is determined. The current tracking point is the nearest projection point of the UAV's current position on the three-dimensional space centerline, which is offset by a preset flight height along the normal direction. The forward tracking point is the point that is offset by a preset flight height along the normal direction at a preset forward distance along the three-dimensional space centerline.

[0080] In this embodiment of the invention, firstly, in the lightweight three-dimensional Gaussian representation model, the position parameter of each Gaussian primitive is the mean vector of the probability density function of its spatial distribution, representing the spatial center position of the local area of ​​the scene represented by the primitive. The three-dimensional center positions of each Gaussian primitive are extracted from the lightweight three-dimensional Gaussian representation model to form a three-dimensional point set of the cable centerline.

[0081] Secondly, the fitting residuals are calculated for the three-dimensional point set of the cable centerline. The fitting residuals are statistical values ​​of the perpendicular distances between points in the three-dimensional point set of the cable centerline and the fitted curve. The smaller the fitting residuals, the closer the geometry of the three-dimensional point set of the cable centerline is to the adopted spatial curve model. Subsequently, a spatial curve model is adaptively selected based on the fitting residuals for fitting to obtain the three-dimensional spatial centerline, including: Furthermore, when the fitting residual is lower than the first preset threshold, the principal direction straight line fitting model is adopted. The principal direction straight line fitting model is a mathematical model that fits a three-dimensional point set into a spatial straight line, determining the direction vector and the point through which the line passes by minimizing the perpendicular distance from the point to the line.

[0082] When the fitting residual is lower than the second preset threshold and higher than the first preset threshold, a quadratic polynomial curve fitting model is used. The quadratic polynomial curve fitting model is a mathematical model that fits a three-dimensional point set into a quadratic polynomial curve in three-dimensional space. Its parameter form is as follows: .

[0083] When the fitting residual exceeds a second preset threshold, a B-spline curve fitting model is adopted, where the first preset threshold is less than the second preset threshold. The B-spline curve fitting model is a mathematical model that fits a 3D point set into a B-spline curve. B-spline curves are defined by control points and node vectors, exhibiting good local support and smoothness. After adaptive selection and fitting, the resulting 3D spatial centerline describes the geometric orientation of the cable in 3D space as a continuous parametric curve.

[0084] The first preset threshold is used to determine whether the cable segment is approximately a straight line. Its setting needs to consider the degree of catenary curvature formed by the cable's own weight in the mid-span between towers. A typical value for this threshold can be set to 2 to 4 times the cable diameter, for example, 0.02m to 0.05m, to ensure that the lateral deviation introduced by the straight-line fitting is much smaller than the safe distance between the UAV and the cable. The second preset threshold is used to determine whether a higher-order curve model is needed for the cable segment. Its setting needs to comprehensively consider the geometric characteristics of the significantly increased curvature near the tower suspension point and the requirements for viewing angle stability in inspection image acquisition. A typical value for this threshold can be set to the upper limit of the UAV's expected flight path tracking accuracy, for example, 0.10m to 0.20m.

[0085] Finally, based on the three-dimensional space centerline, the current tracking point and the forward tracking point are extracted and the corresponding cable observation position information is determined. The current tracking point is the closest projection point of the UAV's current position on the three-dimensional space centerline, which is offset by a preset flight height along the normal direction. The forward tracking point is the point that is offset by a preset flight height along the normal direction at a preset forward distance along the three-dimensional space centerline.

[0086] In this embodiment of the invention, 3D geometric modeling of cables is achieved through sparse Gaussian primitive initialization guided by cable-specific constraints, collaborative optimization of multidimensional loss functions, and adaptive curve fitting. Anisotropy, bounded curvature, spatial continuity, and sparse regularization constraints suppress geometric distortion of slender structures; quantization compression reduces the model volume to adapt to real-time airborne processing; subsequently, adaptive curve fitting dynamically selects the optimal model based on the local curvature of the cable, ensuring the accuracy of centerline extraction and the smoothness of flight control, providing a reliable spatial reference and navigation path for autonomous inspection along the cable route.

[0087] S400: Based on the cable observation position information and the real-time pose information provided by inertial measurement, combined with the constraints of the target inspection task, generate flight control commands to control the UAV to fly autonomously along the cable path.

[0088] In this embodiment of the invention, the real-time pose information is the position and attitude data of the UAV in three-dimensional space at the current moment provided by the UAV inertial measurement unit, which usually includes three-axis coordinates and three-axis attitude angles; the constraints of the target inspection task are the flight restrictions and requirements set for a specific inspection task; the flight control command is the execution command sent to the UAV flight control system, which usually includes control quantities such as the desired heading angle, desired vertical speed, and desired forward speed, used to drive the UAV to fly according to the planned path.

[0089] Specifically, the cable observation position information provided by the lightweight 3D Gaussian representation model is mapped to the navigation coordinate system to obtain the cable tracking position. Then, the horizontal azimuth angle of the forward-looking tracking point relative to the UAV's current position is calculated as the desired heading angle command. Simultaneously, the lateral offset distance from the UAV's current position to the cable tracking position is calculated, and a heading fine-tuning amount is generated via a PID controller, which is then superimposed on the desired heading angle command to form the final heading angle command. Based on the difference between the vertical coordinates of the current tracking point and the UAV's current vertical coordinates, a vertical velocity command is generated via a proportional controller. The curvature value of the cable tracking position at the current tracking point is obtained, the velocity attenuation coefficient is calculated, and the forward velocity command is obtained. The final flight control command is sent to the flight control system to drive the UAV to fly autonomously along the cable tracking position.

[0090] Step S400 in the method of this embodiment of the invention includes: The cable observation position information is mapped to the navigation coordinate system to obtain the cable tracking position in the navigation coordinate system; Calculate the horizontal azimuth angle of the forward-looking tracking point relative to the current position of the UAV in the cable tracking position, and use it as the desired heading angle command; Calculate the lateral offset distance from the current position of the UAV to the cable tracking position, generate a heading fine-tuning amount based on the PID controller, and superimpose the heading fine-tuning amount onto the desired heading angle command to obtain the final heading angle command; Based on the difference between the vertical coordinates of the current tracking point in the cable tracking position and the current vertical coordinates of the UAV, a vertical speed command is generated through a proportional controller. Obtain the curvature value of the cable tracking position at the current tracking point, calculate the speed attenuation coefficient based on the curvature value, and multiply the maximum flight speed by the speed attenuation coefficient to obtain the forward speed command; The final heading angle command, the vertical speed command, and the forward speed command are combined into the flight control command, which drives the UAV to fly autonomously along the cable tracking position.

[0091] In this embodiment of the invention, firstly, the cable observation position information is mapped to the navigation coordinate system to obtain the cable tracking position in the navigation coordinate system. The navigation coordinate system is a global reference coordinate system describing the position and attitude of the UAV, typically using a northeast-central coordinate system or a geocentric coordinate system. In the navigation coordinate system, the UAV's position is represented by three-dimensional coordinates (x, y, z), where x points north, y points east, and z points towards the geocentric direction or altitude.

[0092] Specifically, using a pre-calibrated extrinsic transformation matrix from the camera coordinate system to the body coordinate system, the cable observation position information is transformed from the camera coordinate system to the body coordinate system. Real-time UAV pose information provided by the inertial measurement unit and navigation system is acquired, and a rotation and translation transformation matrix from the body coordinate system to the navigation coordinate system is constructed. The cable position in the body coordinate system is then transformed to the navigation coordinate system using the rotation and translation transformation matrix, yielding the cable tracking position that the UAV should track in the global reference coordinate system.

[0093] Secondly, the horizontal azimuth angle of the forward-looking tracking point relative to the current position of the UAV in the cable tracking position is calculated as the desired heading angle command. The horizontal azimuth angle is the angle between the direction from the current position of the UAV to the forward-looking tracking point and true north in the horizontal plane, usually expressed in degrees or radians.

[0094] Specifically, the 3D coordinates of the forward-looking tracking point are extracted from the cable tracking position in the navigation coordinate system, and the 3D coordinates of the UAV's current position are also obtained. The relative vectors of the forward-looking tracking point with respect to the UAV's current position on the horizontal plane are calculated, i.e., the north and east coordinates of the forward-looking tracking point are subtracted from the north and east coordinates of the UAV's current position, respectively. The clockwise angle of this relative vector with respect to true north is calculated to obtain the azimuth angle, which is then output as the desired heading angle command to the yaw channel of the flight control system.

[0095] Next, the lateral offset distance from the UAV's current position to the cable tracking position is calculated, a heading fine-tuning amount based on the PID controller is generated, and this heading fine-tuning amount is superimposed on the desired heading angle command to obtain the final heading angle command. Here, the lateral offset distance is the vertical distance on the horizontal plane from the UAV's current position to the current tracking point within the cable tracking position.

[0096] Specifically, the direction of the line connecting the foreseeable tracking point and the current tracking point is taken as the desired flight path direction. The vertical distance from the UAV's current position to this desired flight path is calculated, which is the lateral offset distance. This lateral offset distance is input as an error signal to the PID controller. The PID controller outputs a heading fine-tuning amount based on the comprehensive calculations of the proportional, integral, and derivative components. The proportional component generates an immediate correction based on the current offset; the integral component eliminates steady-state errors; and the derivative component predicts the offset change trend and suppresses oscillations. Subsequently, the heading fine-tuning amount is superimposed on the desired heading angle command to obtain the final heading angle command, which is sent to the yaw channel of the flight control system.

[0097] Furthermore, the vertical coordinates of the current tracking point are extracted from the cable tracking position in the navigation coordinate system, and the vertical coordinates of the UAV's current position are also obtained. The difference between the two is calculated to obtain the vertical deviation. This vertical deviation is multiplied by a preset scaling factor to obtain the vertical speed command. If the UAV's current altitude is higher than the desired altitude, the vertical deviation is negative, and the vertical speed command is negative; otherwise, it is positive. The vertical speed command is then sent to the vertical channel of the flight control system.

[0098] Furthermore, the curvature value is calculated from the second derivative information of the centerline in three-dimensional space at the current tracking point. Then, according to the preset curvature-velocity mapping relationship, the curvature value is mapped to a velocity attenuation coefficient. Specifically, when the curvature is 0, the attenuation coefficient is 1.0, and the UAV flies at its maximum flight speed; as the curvature increases, the attenuation coefficient decreases monotonically; when the curvature reaches the preset maximum allowable curvature, the attenuation coefficient drops to the preset minimum value. Subsequently, the maximum flight speed is multiplied by the velocity attenuation coefficient to obtain the forward speed command, which is sent to the horizontal velocity channel of the flight control system.

[0099] Finally, the final heading angle command, vertical speed command, and forward speed command are combined to form a complete flight control command, which is then sent to the flight control system via the data link between the UAV and the ground station or the onboard communication bus. Based on the received commands, the flight control system drives the UAV's motors and control surface actuators through the underlying attitude and speed control loop, enabling the UAV to autonomously fly along the cable tracking position. This control process is repeated in each control cycle, forming a continuous closed-loop control that allows the UAV to continuously track the cable path until the inspection task is completed or a stop command is received.

[0100] In this embodiment of the invention, the cable observation position is mapped to the navigation coordinate system, and the desired heading angle command is generated by combining the forward tracking point. A PID controller then generates a heading fine-tuning amount based on the lateral offset, achieving both predictive and closed-loop dual-mode control of the heading. Subsequently, a proportional controller based on the vertical deviation adjusts the flight altitude in real time, and the flight speed is adaptively adjusted based on the curvature of the current tracking point, enabling the UAV to automatically decelerate on curves to ensure flight safety and inspection image quality. Finally, the heading angle command, vertical speed, and forward speed command are fused into a control command to drive the UAV to fly autonomously along the cable path, ensuring the safety and stability of the inspection process.

[0101] Through the above specific implementation methods, the embodiments of the present invention achieve the following technical effects: In this embodiment of the invention, the spatiotemporal alignment of images and inertial data is first achieved through a preset observation window and synchronous acquisition, providing a reliable input foundation for subsequent processing. Distortion correction eliminates lens geometric distortion, ensuring pixel spatial mapping accuracy. Downsampling and quantization compression reduce data volume and computational burden, laying the data foundation for cable identification, 3D modeling, and autonomous flight control.

[0102] Secondly, by using a lightweight semantic segmentation network and extracting regions of interest from monocular temporal images, the monocular temporal images to be processed are reduced to local areas of cables, reducing computational burden and effectively eliminating background interference. Simultaneously, based on the integration and concatenation of inertial measurement data, accurate inter-frame relative pose sequences are obtained. Subsequently, mask consistency matching is performed to search for the optimal depth estimate under the discrete depth assumption, effectively solving the problem of missing depth information in monocular vision. The temporal image sequence is then transformed into a pseudo-multi-view image set with pose and depth annotations, providing input data for the subsequent lightweight 3D Gaussian representation model.

[0103] Furthermore, 3D geometric modeling of the cable is achieved through cable-specific constraint-guided sparse Gaussian primitive initialization, multidimensional loss function co-optimization, and adaptive curve fitting. Anisotropy, bounded curvature, spatial continuity, and sparse regularization constraints suppress geometric distortion of slender structures; quantization compression reduces model volume to adapt to real-time airborne processing; subsequently, adaptive curve fitting dynamically selects the optimal model based on the local curvature of the cable, ensuring centerline extraction accuracy and flight control smoothness, providing a reliable spatial reference and navigation path for autonomous inspection along the cable route.

[0104] Finally, the cable observation position is mapped to the navigation coordinate system, and the desired heading angle command is generated by combining the forward tracking point. Based on the lateral offset, a PID controller generates a fine-tuning amount for the heading, realizing predictive and closed-loop dual-mode control of the heading. Subsequently, a proportional controller based on the vertical deviation adjusts the flight altitude in real time, and the flight speed is adaptively adjusted based on the curvature of the current tracking point, enabling the UAV to automatically decelerate on curves to ensure flight safety and inspection image quality. Finally, the heading angle command, vertical speed, and forward speed command are fused into a control command to drive the UAV to fly autonomously along the cable path, ensuring the safety and stability of the inspection process.

[0105] Example 2, as Figure 3 As shown, based on the same inventive concept as the unmanned aerial vehicle (UAV) autonomous patrol flight control method provided in Embodiment 1, this embodiment of the invention also provides an unmanned aerial vehicle (UAV) autonomous patrol flight control system, the system comprising: The attitude-motion data acquisition module 11 is used to acquire monocular time-series image sequences of the UAV and attitude-motion data at corresponding times based on a preset observation window; The region of interest extraction and relative pose estimation module 12 is used to extract the region of interest from the monocular temporal image sequence, obtain the masked temporal image sequence, and perform inter-frame relative pose estimation based on the pose-motion data to obtain the relative pose sequence. The cable observation location information acquisition module 13 is used to combine the mask time sequence image sequence with the relative pose sequence to construct a pseudo multi-view image group, and based on the pseudo multi-view image group, to establish a lightweight three-dimensional Gaussian representation model for the target cable and acquire the cable observation location information of the target cable. The flight control command generation module 14 is used to generate flight control commands based on the cable observation position information and the real-time pose information provided by inertial measurement, combined with the constraints of the target inspection task, to control the UAV to fly autonomously along the cable path.

[0106] In one embodiment, the attitude-motion data acquisition module 11 is used for: The observation window length and acquisition frequency are preset, wherein the observation window length is defined as the number of historical frames retained within the sliding window; According to the acquisition frequency, the monocular time-series image sequence output by the monocular vision sensor and the triaxial angular velocity data and triaxial acceleration data output by the inertial measurement unit at the corresponding time are synchronously acquired at a preset frame rate. The monocular temporal image sequence is sequentially subjected to distortion correction, downsampling, and quantization compression processing; The monocular time-series image sequence, the three-axis angular velocity data, and the three-axis acceleration data are time-stamped to obtain the monocular time-series image sequence associated with the attitude-motion data.

[0107] In one embodiment, the region of interest extraction and relative pose estimation module 12 is used for: Based on a pre-trained lightweight semantic segmentation network, each frame of the monocular temporal image sequence is processed to obtain a pixel-wise classification probability map. The pixel-by-pixel classification probability map is thresholded to obtain the cable binary mask of the current frame. In the cable binary mask, the region with a pixel value of a first preset value represents the cable region, and the region with a pixel value of a second preset value represents the background region. Based on the binary mask of the cable, the minimum bounding rectangle of the cable region is extracted, and along each side of the minimum bounding rectangle, a preset margin is extended outward to form the bounding box of the region of interest. Spatial cropping is performed on each frame of the image to obtain the region of interest image, and the output is a masked temporal image. Iteratively obtain the masked temporal image sequence corresponding to the monocular temporal image sequence; The inter-frame relative rotation matrix is ​​obtained by integrating the three-axis angular velocity data, and the inter-frame relative translation vector is obtained by double integration of the three-axis acceleration data after removing the gravity component. The inter-frame relative rigid body transformation matrix is ​​formed by the inter-frame relative rotation matrix and the inter-frame relative translation vector, and the relative pose sequence is obtained by cascading and accumulating multiple inter-frame relative rigid body transformation matrices within the preset observation window.

[0108] Specifically, by combining the masked temporal image sequence with the relative pose sequence, a pseudo-multi-view image group is constructed, including: Based on historical patrol records, define a preset number of discrete depth assumptions; Under each discrete depth assumption, the relative pose sequence is traversed, and the corresponding pixels in the mask temporal image sequence are subjected to 3D backprojection and cross-frame reprojection by combining the camera intrinsic parameter matrix and the inter-frame relative rigid body transformation matrix to obtain discrete projection results. Based on the discrete projection results, the proportion of the projected pixels corresponding to each discrete depth assumption value falling within the mask area belonging to the cable region in the current frame is statistically calculated, and the mask consistency matching score of each discrete depth assumption value is calculated. The depth hypothesis with the highest matching score is selected as the optimal depth estimate for the current frame. The mask temporal image sequence, the relative pose sequence, and the optimal depth estimate are then combined to construct a pseudo-multi-view image group.

[0109] In one embodiment, the cable observation location information acquisition module 13 is used for: Establish cable-specific constraints and apply geometric restrictions to the initialization parameters of Gaussian elements based on the cable-specific constraints to perform sparse Gaussian element initialization; Combining the cable-specific constraints, a multidimensional loss function is constructed, and the initialization results of sparse Gaussian primitives are iteratively rendered and optimized to obtain a set of Gaussian primitives that meet the preset constraints. The output is a lightweight three-dimensional Gaussian representation model. Based on the lightweight three-dimensional Gaussian characterization model, adaptive spatial curve fitting with dynamic resolution is performed to obtain the cable observation location information of the target cable.

[0110] Specifically, establishing cable-specific constraints and imposing geometric restrictions on the initialization parameters of Gaussian elements based on these constraints, and performing sparse Gaussian element initialization, includes: The cable-specific constraints are established, and the cable-specific constraints include at least the following: Anisotropic constraints are used to constrain the spatial shape of Gaussian elements to be an ellipsoid stretched along a preset direction; The curvature bounded constraint is used to ensure that the spatial discrete curvature formed by the center positions of adjacent Gaussian elements does not exceed a preset curvature upper limit. Spatial continuity constraint is used to ensure that the distance between the center positions of adjacent Gaussian elements does not exceed a preset maximum distance; Sparse regularization constraints are used to constrain the Gaussian element opacity parameter to tend to zero in order to achieve adaptive sparsity. Based on the optimal depth estimate in the pseudo-multi-view image group, the pixels in the masked temporal image sequence are back-projected into three-dimensional space to obtain an initial three-dimensional point set, and the initial three-dimensional point set is uniformly downsampled to obtain a sparse sampling point set. Based on the cable-specific constraints, a Gaussian primitive is initialized with each sparse sampling point as the center, and the sparse Gaussian primitive initialization result is obtained. In this process, the covariance matrix of each Gaussian element determines the local tangent direction based on the difference direction of adjacent sparse sampling points, and the anisotropic ellipsoid shape is initialized with the local tangent direction as the major axis direction. The opacity parameter of each Gaussian element is set to a preset initial opacity value, and the color representation of each Gaussian element adopts the first-order spherical harmonic function coefficients that retain only the diffuse reflection component.

[0111] Specifically, a multidimensional loss function is constructed based on the cable-specific constraints. The initialization results of sparse Gaussian primitives are iteratively rendered and optimized to obtain a set of Gaussian primitives that satisfy preset constraints. The output is a lightweight three-dimensional Gaussian representation model, including: Construct a multidimensional loss function that includes a rendering consistency loss term and an additional loss term, wherein the additional loss term includes at least one of the following: a slenderness regularization loss term, a curvature regularization loss term, a spatial continuity regularization loss term, and a sparse regularization loss term. Starting with the Gaussian initialization parameters in the sparse Gaussian initialization result, and with the multidimensional loss function as the optimization objective, the position parameters, covariance matrix parameters, opacity parameters, and color representation of each Gaussian are iteratively optimized for a preset number of iterations. After optimization and iteration, Gaussian elements with opacity parameters lower than the preset opacity threshold are removed, the Gaussian element set is obtained, and the position parameters, covariance matrix parameters, and opacity parameters of each Gaussian element are quantized and compressed to output the lightweight three-dimensional Gaussian representation model.

[0112] Specifically, based on the lightweight three-dimensional Gaussian representation model, adaptive spatial curve fitting is performed to obtain the cable observation location information of the target cable, including: The three-dimensional center positions of each Gaussian element are extracted from the lightweight three-dimensional Gaussian representation model to form a three-dimensional point set of the cable centerline; Calculate the fitting residual for the three-dimensional point set of the cable centerline, and adaptively select a spatial curve model for fitting based on the fitting residual to obtain the three-dimensional spatial centerline, including: When the fitting residual is lower than the first preset threshold, the main direction linear fitting model is adopted; When the fitting residual is lower than the second preset threshold and higher than the first preset threshold, a quadratic polynomial curve fitting model is used. When the fitting residual is higher than the second preset threshold, a B-spline curve fitting model is used, wherein the first preset threshold is less than the second preset threshold; Based on the three-dimensional space centerline, the current tracking point and the forward tracking point are extracted and the corresponding cable observation position information is determined. The current tracking point is the nearest projection point of the UAV's current position on the three-dimensional space centerline, which is offset by a preset flight height along the normal direction. The forward tracking point is the point that is offset by a preset flight height along the normal direction at a preset forward distance along the three-dimensional space centerline.

[0113] In one embodiment, the flight control command generation module 14 is used for: The cable observation position information is mapped to the navigation coordinate system to obtain the cable tracking position in the navigation coordinate system; Calculate the horizontal azimuth angle of the forward-looking tracking point relative to the current position of the UAV in the cable tracking position, and use it as the desired heading angle command; Calculate the lateral offset distance from the current position of the UAV to the cable tracking position, generate a heading fine-tuning amount based on the PID controller, and superimpose the heading fine-tuning amount onto the desired heading angle command to obtain the final heading angle command; Based on the difference between the vertical coordinates of the current tracking point in the cable tracking position and the current vertical coordinates of the UAV, a vertical speed command is generated through a proportional controller. Obtain the curvature value of the cable tracking position at the current tracking point, calculate the speed attenuation coefficient based on the curvature value, and multiply the maximum flight speed by the speed attenuation coefficient to obtain the forward speed command; The final heading angle command, the vertical speed command, and the forward speed command are combined into the flight control command, which drives the UAV to fly autonomously along the cable tracking position.

[0114] Compared to existing technologies, this invention firstly uses a preset observation window to synchronously acquire and time-align monocular temporal image sequences and attitude-motion data, providing a spatiotemporally consistent data foundation for subsequent inter-frame pose estimation and 3D reconstruction. Secondly, a lightweight semantic segmentation network is used to achieve cable pixel detection and region of interest extraction, significantly reducing the subsequent processing scope to the local area of ​​the cable, reducing computational burden and eliminating background interference. Inter-frame relative pose sequences are obtained based on the integration and cascading of inertial measurement data, exhibiting strong stability in low-texture cable scenes. Thirdly, scene depth is recovered through depth estimation using mask consistency matching, transforming the monocular temporal image sequence into a pseudo-multi-view image group, providing structured input for 3D modeling. A lightweight 3D Gaussian representation model and cable-specific constraints are used to accurately represent the slender structure of the cable with fewer Gaussian units, achieving model lightweighting to adapt to airborne real-time processing. Finally, the cable observation position information is fused with real-time pose, and flight control commands for heading, altitude, and speed are generated through a hierarchical control strategy, driving the UAV to autonomously fly along the cable path, achieving autonomous inspection along the line without human intervention.

Claims

1. A method for autonomous patrol flight control along a designated route for unmanned aerial vehicles (UAVs), characterized in that, include: Based on a preset observation window, a monocular time-series image sequence of the UAV and the attitude-motion data at the corresponding time are collected; Region of interest extraction is performed on the monocular temporal image sequence to obtain a masked temporal image sequence, and inter-frame relative pose estimation is performed based on the pose-motion data to obtain a relative pose sequence; By combining the masked temporal image sequence with the relative pose sequence, a pseudo-multi-view image group is constructed, and a lightweight three-dimensional Gaussian representation model for the target cable is established based on the pseudo-multi-view image group to obtain the cable observation position information of the target cable. Based on the cable observation location information and the real-time pose information provided by inertial measurement, combined with the constraints of the target inspection task, flight control commands are generated to control the UAV to fly autonomously along the cable path.

2. The method for autonomous patrol flight control along a route for unmanned aerial vehicles (UAVs) as described in claim 1, characterized in that, Based on a preset observation window, monocular time-series image sequences of the UAV and corresponding attitude-motion data at corresponding times are collected, including: The observation window length and acquisition frequency are preset, wherein the observation window length is defined as the number of historical frames retained within the sliding window; According to the acquisition frequency, the monocular time-series image sequence output by the monocular vision sensor and the triaxial angular velocity data and triaxial acceleration data output by the inertial measurement unit at the corresponding time are synchronously acquired at a preset frame rate. The monocular temporal image sequence is sequentially subjected to distortion correction, downsampling, and quantization compression processing; The monocular temporal image sequence, the three-axis angular velocity data, and the three-axis acceleration data are time-stamped to obtain the monocular temporal image sequence associated with the attitude-motion data.

3. The method for autonomous patrol flight control along a route for unmanned aerial vehicles (UAVs) as described in claim 1, characterized in that, The monocular temporal image sequence is subjected to region of interest extraction to obtain a masked temporal image sequence, and inter-frame relative pose estimation is performed based on the pose-motion data to obtain a relative pose sequence, including: Based on a pre-trained lightweight semantic segmentation network, each frame of the monocular temporal image sequence is processed to obtain a pixel-wise classification probability map. The pixel-by-pixel classification probability map is thresholded to obtain the cable binary mask of the current frame. In the cable binary mask, the region with a pixel value of a first preset value represents the cable region, and the region with a pixel value of a second preset value represents the background region. Based on the binary mask of the cable, the minimum bounding rectangle of the cable region is extracted, and along each side of the minimum bounding rectangle, a preset margin is extended outward to form the bounding box of the region of interest. Spatial cropping is performed on each frame of the image to obtain the region of interest image, and the output is a masked temporal image. Iteratively obtain the masked temporal image sequence corresponding to the monocular temporal image sequence; The inter-frame relative rotation matrix is ​​obtained by integrating the three-axis angular velocity data, and the inter-frame relative translation vector is obtained by double integration of the three-axis acceleration data after removing the gravity component. The inter-frame relative rigid body transformation matrix is ​​formed by the inter-frame relative rotation matrix and the inter-frame relative translation vector, and the relative pose sequence is obtained by cascading and accumulating multiple inter-frame relative rigid body transformation matrices within the preset observation window.

4. The method for autonomous patrol flight control along a route for unmanned aerial vehicles (UAVs) as described in claim 3, characterized in that, By combining the masked temporal image sequence with the relative pose sequence, a pseudo-multi-view image group is constructed, including: Based on historical patrol records, define a preset number of discrete depth assumptions; Under each discrete depth assumption, the relative pose sequence is traversed, and the corresponding pixels in the mask temporal image sequence are subjected to 3D backprojection and cross-frame reprojection by combining the camera intrinsic parameter matrix and the inter-frame relative rigid body transformation matrix to obtain discrete projection results. Based on the discrete projection results, the proportion of the projected pixels corresponding to each discrete depth assumption value falling within the mask area belonging to the cable region in the current frame is statistically calculated, and the mask consistency matching score of each discrete depth assumption value is calculated. The depth hypothesis with the highest matching score is selected as the optimal depth estimate for the current frame. The mask temporal image sequence, the relative pose sequence, and the optimal depth estimate are then combined to construct a pseudo-multi-view image group.

5. The method for autonomous patrol flight control along a route for unmanned aerial vehicles (UAVs) as described in claim 1, characterized in that, Based on the pseudo-multi-view image set, a lightweight three-dimensional Gaussian representation model for the target cable is established to obtain the cable observation location information of the target cable, including: Establish cable-specific constraints and apply geometric restrictions to the initialization parameters of Gaussian elements based on the cable-specific constraints to perform sparse Gaussian element initialization; Combining the cable-specific constraints, a multidimensional loss function is constructed, and the initialization results of sparse Gaussian primitives are iteratively rendered and optimized to obtain a set of Gaussian primitives that meet the preset constraints. The output is a lightweight three-dimensional Gaussian representation model. Based on the lightweight three-dimensional Gaussian characterization model, adaptive spatial curve fitting with dynamic resolution is performed to obtain the cable observation location information of the target cable.

6. The method for autonomous patrol flight control along a route for unmanned aerial vehicles (UAVs) as described in claim 5, characterized in that, Establish cable-specific constraints and apply geometric restrictions to the initialization parameters of Gaussian elements based on these constraints, performing sparse Gaussian element initialization, including: The cable-specific constraints are established, and the cable-specific constraints include at least the following: Anisotropic constraints are used to constrain the spatial shape of Gaussian elements to be an ellipsoid stretched along a preset direction; The curvature bounded constraint is used to ensure that the spatial discrete curvature formed by the center positions of adjacent Gaussian elements does not exceed a preset curvature upper limit. Spatial continuity constraint is used to ensure that the distance between the center positions of adjacent Gaussian elements does not exceed a preset maximum distance; Sparse regularization constraints are used to constrain the Gaussian element opacity parameter to tend to zero in order to achieve adaptive sparsity. Based on the optimal depth estimate in the pseudo-multi-view image group, the pixels in the masked temporal image sequence are back-projected into three-dimensional space to obtain an initial three-dimensional point set, and the initial three-dimensional point set is uniformly downsampled to obtain a sparse sampling point set. Based on the cable-specific constraints, a Gaussian primitive is initialized with each sparse sampling point as the center, and the sparse Gaussian primitive initialization result is obtained. In this process, the covariance matrix of each Gaussian element determines the local tangent direction based on the difference direction of adjacent sparse sampling points, and the anisotropic ellipsoid shape is initialized with the local tangent direction as the major axis direction. The opacity parameter of each Gaussian element is set to a preset initial opacity value, and the color representation of each Gaussian element adopts the first-order spherical harmonic function coefficients that retain only the diffuse reflection component.

7. The method for autonomous patrol flight control along a route for unmanned aerial vehicles as described in claim 5, characterized in that, Combining the cable-specific constraints, a multidimensional loss function is constructed, and the initialization results of sparse Gaussian primitives are iteratively rendered and optimized to obtain a set of Gaussian primitives that satisfy preset constraints. The output is a lightweight three-dimensional Gaussian representation model, including: Construct a multidimensional loss function that includes a rendering consistency loss term and an additional loss term, wherein the additional loss term includes at least one of the following: a slenderness regularization loss term, a curvature regularization loss term, a spatial continuity regularization loss term, and a sparse regularization loss term. Starting with the Gaussian initialization parameters in the sparse Gaussian initialization result, and with the multidimensional loss function as the optimization objective, the position parameters, covariance matrix parameters, opacity parameters, and color representation of each Gaussian are iteratively optimized for a preset number of iterations. After optimization and iteration, Gaussian elements with opacity parameters lower than the preset opacity threshold are removed, the Gaussian element set is obtained, and the position parameters, covariance matrix parameters, and opacity parameters of each Gaussian element are quantized and compressed to output the lightweight three-dimensional Gaussian representation model.

8. The method for autonomous patrol flight control along a route for unmanned aerial vehicles as described in claim 5, characterized in that, Based on the aforementioned lightweight 3D Gaussian representation model, adaptive spatial curve fitting is performed to obtain the cable observation location information of the target cable, including: The three-dimensional center positions of each Gaussian element are extracted from the lightweight three-dimensional Gaussian representation model to form a three-dimensional point set of the cable centerline; Calculate the fitting residual for the three-dimensional point set of the cable centerline, and adaptively select a spatial curve model for fitting based on the fitting residual to obtain the three-dimensional spatial centerline, including: When the fitting residual is lower than the first preset threshold, the main direction linear fitting model is adopted; When the fitting residual is lower than the second preset threshold and higher than the first preset threshold, a quadratic polynomial curve fitting model is used. When the fitting residual is higher than the second preset threshold, a B-spline curve fitting model is used, wherein the first preset threshold is less than the second preset threshold; Based on the three-dimensional space centerline, the current tracking point and the forward tracking point are extracted and the corresponding cable observation position information is determined. The current tracking point is the nearest projection point of the UAV's current position on the three-dimensional space centerline, which is offset by a preset flight height along the normal direction. The forward tracking point is the point that is offset by a preset flight height along the normal direction at a preset forward distance along the three-dimensional space centerline.

9. The method for autonomous patrol flight control along a route for unmanned aerial vehicles as described in claim 1, characterized in that, Based on the cable observation location information and the real-time pose information provided by inertial measurement, combined with the constraints of the target inspection task, flight control commands are generated to control the UAV to fly autonomously along the cable path, including: The cable observation position information is mapped to the navigation coordinate system to obtain the cable tracking position in the navigation coordinate system; Calculate the horizontal azimuth angle of the forward-looking tracking point relative to the current position of the UAV in the cable tracking position, and use it as the desired heading angle command; Calculate the lateral offset distance from the current position of the UAV to the cable tracking position, generate a heading fine-tuning amount based on the PID controller, and superimpose the heading fine-tuning amount onto the desired heading angle command to obtain the final heading angle command; Based on the difference between the vertical coordinates of the current tracking point in the cable tracking position and the current vertical coordinates of the UAV, a vertical speed command is generated through a proportional controller. Obtain the curvature value of the cable tracking position at the current tracking point, calculate the speed attenuation coefficient based on the curvature value, and multiply the maximum flight speed by the speed attenuation coefficient to obtain the forward speed command; The final heading angle command, the vertical speed command, and the forward speed command are combined into the flight control command, which drives the UAV to fly autonomously along the cable tracking position.

10. A flight control system for autonomous patrol along a route for unmanned aerial vehicles (UAVs), characterized in that: For implementing the autonomous patrol flight control method for unmanned aerial vehicles (UAVs) according to any one of claims 1-9, the system comprises: The attitude-motion data acquisition module is used to acquire monocular time-series image sequences of the UAV and attitude-motion data at corresponding times based on a preset observation window; The region of interest extraction and relative pose estimation module is used to extract the region of interest from the monocular temporal image sequence, obtain the masked temporal image sequence, and perform inter-frame relative pose estimation based on the pose-motion data to obtain the relative pose sequence. The cable observation location information acquisition module is used to combine the mask time sequence image sequence with the relative pose sequence to construct a pseudo multi-view image group, and based on the pseudo multi-view image group, to establish a lightweight three-dimensional Gaussian representation model for the target cable and obtain the cable observation location information of the target cable. The flight control command generation module is used to generate flight control commands based on the cable observation position information and the real-time pose information provided by inertial measurement, combined with the constraints of the target inspection task, to control the UAV to fly autonomously along the cable path.