Unmanned aerial vehicle line simulation flight method and system based on sub-pixel segmentation optimization three-dimensional reconstruction
By employing subpixel-level image segmentation and 3D reconstruction technologies, the tracking accuracy problem of UAVs in long-distance power transmission scenarios has been solved, enabling high-precision wire identification and equidistant flight of UAVs, thereby improving inspection efficiency and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-31
AI Technical Summary
Existing UAV line-following flight technology struggles to maintain equidistant tracking in long-distance power transmission scenarios, and its insufficient 3D reconstruction accuracy leads to significant errors in calculation results, making it difficult to meet the needs of refined operation and maintenance.
Subpixel-level image segmentation technology is introduced. Image data is acquired through UAV onboard sensors, filtered and preprocessed, and then segmented at the subpixel level. Triangulation and bundle adjustment are combined to perform 3D reconstruction, and the geometric feature extraction and sag calculation of conductors and towers are optimized.
It improves the identification accuracy of conductors and towers, enables centimeter-level dynamic positioning and drone tracking flight at equal distances, and enhances the perception accuracy and operational reliability of drone inspections.
Smart Images

Figure CN121764147A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electrical automation, and specifically relates to a method and system for UAV line-following flight based on sub-pixel segmentation optimized 3D reconstruction. Background Technology
[0002] In the power transmission system, high-voltage overhead conductors serve as the core artery, and their safe and stable operation directly affects the reliability of energy supply. Traditional manual inspection methods suffer from inherent drawbacks such as low efficiency, high risk, and numerous blind spots—making them costly, slow, difficult, and dangerous—and unsuitable for the development needs of modern power grids. To improve the efficiency and safety of high-voltage overhead transmission line inspections, drone technology is gradually replacing the traditional operation mode that relies on manual foot patrols, telescopes, and infrared thermal imagers. Against this backdrop, line-following flight technology has emerged as an advanced form of drone inspection.
[0003] The emergence of line-following flight technology has brought significant changes to the field of power line inspection. During inspection, sensors such as lidar and visual cameras on drones can perceive and track the position of power lines in real time. Combined with control algorithms, this ensures that a stable detection distance is maintained between the drone and the power line. However, current drone line-following flight technology still has some unresolved issues: in long-distance power transmission scenarios, the sag phenomenon caused by the weight of the power line can make it difficult for the drone to maintain equidistant tracking. Although 3D reconstruction methods can calculate the sag, the insufficient accuracy of 3D modeling leads to significant errors in the calculation results, making it difficult to meet the needs of refined operation and maintenance. There is an urgent need to introduce sub-pixel segmentation technology to improve the accuracy of power line contour extraction and fundamentally enhance the accuracy of geometric reconstruction. Summary of the Invention
[0004] To address the shortcomings of existing technologies, one of the objectives of this invention is to provide a method for UAV line-following flight based on subpixel segmentation-optimized 3D reconstruction. By introducing subpixel-level image segmentation technology, the accuracy of 3D reconstruction input data is improved at the source, accurately restoring the spatial morphology and sag distribution of power lines and towers, providing more accurate position information for UAV tracking and control, thereby achieving high-precision identification of power transmission lines, centimeter-level dynamic positioning, and equidistant line-following flight of UAVs in complex scenarios.
[0005] The second objective of this invention is to provide a system for implementing the UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction.
[0006] This invention provides a method for UAV line-following flight based on sub-pixel segmentation optimized 3D reconstruction, comprising the following steps:
[0007] S1: Acquire image data based on the onboard sensors of the UAV;
[0008] S2: Based on the image data obtained in step S1, preprocess the image data to obtain filtered image data;
[0009] S3: Based on the filtered image data obtained in step S2, perform sub-pixel segmentation on the conductor and tower to obtain a sub-pixel segmentation image of the conductor and tower.
[0010] S4: Based on the sub-pixel segmentation map of the conductor and tower obtained in step S3, high-precision reconstruction from two-dimensional image to three-dimensional space is achieved to obtain a three-dimensional point cloud map.
[0011] S5: Based on the spatial geometric information of the conductor and tower in the 3D point cloud obtained in step S4, output the sag data with the catenary equation as the core basis.
[0012] S6. Based on the sag data obtained in step S5, complete the drone's line-following flight.
[0013] In step S1, the onboard sensors of the UAV include lidar and a visual camera.
[0014] In step S2, the preprocessing is a nonlocal filtering method based on structural similarity measurement to filter the image data and achieve directional protection of the geometric features of the conductor through a feature coupling mechanism.
[0015] The preprocessing specifically includes:
[0016] First, a nonlocal search window is set up with each pixel as the center, and the pixels within the window are used as candidate matching objects. Then, a normalized composite weight function is constructed and normalized within this window. The weights are jointly determined by the similarity of brightness distribution, gradient field consistency, and local structure similarity. Furthermore, an anisotropic constraint factor is introduced to selectively modulate the pixel association, thereby strengthening the high-frequency information matching along the dominant direction of the conductor and weakening the erroneous associations dominated by lateral and noise. Finally, the pixel association model establishes nonlocal association relationships within the nonlocal search window with pixels as nodes and normalized composite weights as edge weights, achieving denoising while suppressing the interference of random noise on the shape of slender targets.
[0017] An adaptive smoothing strategy based on regional saliency discrimination constructs a saliency map by comprehensively analyzing the characteristics of different regions and their neighbors, including differences in brightness distribution, gradient response intensity, and consistency of line orientation. The filtering intensity of each region is then dynamically adjusted based on this saliency map: smoothing is enhanced in low-saliency regions to suppress noise, while smoothing is weakened in high-saliency regions to preserve wire edges and details. Continuous adjustment is used in transition regions to achieve a smooth transition, thus establishing a balance between noise suppression and target edge preservation. This leads to more accurate extraction of wire geometric features and improved wire recognition performance.
[0018] In step S3, based on the filtered image data obtained in step S2, a two-stage segmentation strategy from coarse to fine is applied to the conductors and towers, specifically as follows:
[0019] In the first stage, pixel-level coarse segmentation is performed to quickly locate the main areas of the conductor and tower. In the second stage, the sub-pixel-level fine segmentation stage is entered. Through resolution enhancement of local areas and boundary iterative optimization, the accuracy of contour positioning is improved, and the sub-pixel-level segmentation map of the conductor and tower is finally output.
[0020] The pixel-level coarse segmentation specifically refers to:
[0021] Constructing a fuzzy C-mean objective function The calculation of cluster centers and membership degrees is transformed from the pixel level to the gray level, and image classification is completed at the gray level; the fuzzy C-means objective function Express it using the following formula: ;in, Indicates the number of gray levels in an image. This represents the grayscale value of the current pixel. Indicates grayscale level The number of pixels, Indicates the number of cluster centers. Represents grayscale level The pixel belongs to the first Membership degree of each cluster center Represents the fuzzy index. Indicates the first The gray level of each cluster center;
[0022] By iteratively optimizing the membership degree and cluster center until the objective function converges, grayscale classification is finally completed based on the principle of maximum membership degree, achieving pixel-level coarse segmentation of conductors and towers;
[0023] The sub-pixel level refinement stage specifically refers to:
[0024] Regions of Interest (ROIs) are generated based on the area of connected regions, and their pixel-level contours are extracted as initial boundaries. Background noise interference is filtered out by setting an area threshold, the main areas of the conductor and tower are retained, and their outer contours are extracted as initial contours.
[0025] A subpixel processing window is constructed at the contour points, and the image within the window is upsampled at a high magnification to enhance the local detail resolution;
[0026] In the upsampled region, the boundary position is initially located based on the gradient magnitude information, and the initial contour is iteratively adjusted and converged through the boundary evolution optimization method to gradually approach the real boundary and obtain a sub-pixel level boundary point set.
[0027] The optimized subpixel-level boundary point set is mapped back to the original image coordinate system to preserve spatial geometric accuracy. A continuous and smooth contour curve is generated through piecewise curve fitting to obtain the subpixel-level conductor and tower contours, and finally, the subpixel-level segmentation map of the conductor and tower is obtained.
[0028] Step S4 includes the following steps:
[0029] Key features such as conductor sag points and tower corner points are extracted from the sub-pixel segmentation images of conductors and towers, and cross-view feature matching is completed with the help of semantic consistency constraints.
[0030] By using the matched sub-pixel pairs, their three-dimensional spatial positions are recovered through the principle of triangulation, and an initial sparse point cloud is generated.
[0031] Through global optimization and dense matching, a dense point cloud containing details is obtained;
[0032] By performing surface fitting and mesh reconstruction based on the geometric properties of point clouds, a three-dimensional point cloud map for accurate sag calculation is constructed.
[0033] Step S4 is as follows:
[0034] Using sub-pixel-level segmentation maps of conductors and towers as core data support, feature points such as conductor inflection points and tower connection points are extracted, and the pixel positions of feature points are obtained through sub-pixel-level coordinate calibration.
[0035] Within the same category of conductors and towers, cross-view feature point matching is carried out with the help of semantic consistency constraints;
[0036] Based on sub-pixel level matched image points, combined with the camera's calibrated intrinsic and extrinsic parameters, the 3D structure is reconstructed using matching points obtained through triangulation. The projection equation is: ;in, , Represents the depth scale. Represents the camera intrinsic parameter matrix. , and , Let the rotation matrix and translation vector be... , For subpixel level matching of image points, Let be the homogeneous coordinates of a 3D point in world space; rearrange the projection equations, and the resulting linear system is expressed by the following formula: ;in, Camera projection matrix and the coordinates of the observed image points , Linear combination; solving for three-dimensional coordinates using singular value decomposition (SVD). Thus, a preliminary sparse point cloud and camera pose are obtained;
[0037] Global optimization is performed using bundle adjustment (BA), and the objective function is minimized using the following formula: ;in, For camera pose parameters, For three-dimensional point coordinates, For camera frame rate, The number of points in three dimensions; The projection equation is used to calculate the point. 3D coordinates In frame The projected pixel coordinates in the image; For the first Frame observation and The corresponding 2D pixel coordinates; camera parameters and sparse 3D point cloud are obtained through BA optimization;
[0038] A global energy function is constructed, transforming the dense stereo matching problem into a global optimization objective containing data and smoothing terms. Guided by sub-pixel segmentation of the image, the data term calculates the disparity of pixels through precise segmentation boundaries, while the smoothing term employs a semantically adaptive strategy to strengthen disparity continuity constraints within similar objects and preserve the degrees of freedom for disparity jumps at boundaries between different objects. Under the 3D positional constraints provided by sparse reconstruction, an optimal dense disparity map is obtained by minimizing the global energy function. The 3D coordinates are then calculated using the dense disparity map and the obtained camera parameters. This generates a dense point cloud containing details, represented by the following formula: ;in, Focal length Let these be the coordinates of the camera's principal point. Baseline length For parallax, These are the sub-pixel coordinates of the feature point (conductor or tower);
[0039] Based on the generated dense point cloud, an improved moving least squares (MLS) method is used for 3D reconstruction. Local surface fitting is performed in the neighborhood of each point, and then all these local surfaces are seamlessly merged into a smooth overall surface through a weighted average algorithm, finally constructing a 3D point cloud map for accurate sag calculation.
[0040] Step S5 is as follows:
[0041] Based on a smooth 3D model, the spatial geometric information of the conductor and tower is extracted. Taking the catenary equation of a symmetrical tower as the core foundation, the catenary equation is: ;in, It is a constant determined by the boundary conditions. , It refers to the unit mass of the conductor. It is horizontal tension. It is the acceleration due to gravity. The horizontal coordinates of the conductor along the span direction. The longitudinal coordinate of the conductor along the span direction;
[0042] Substituting the boundary conditions into the catenary equation, we obtain the coordinate expression of the traverse along the span direction: ;in, Horizontal spacing The maximum sag is obtained through the final derivation. The calculation expression is: .
[0043] The present invention also provides a system for implementing the UAV line-following flight method based on subpixel segmentation optimized 3D reconstruction, including a data acquisition module, a data processing module, a subpixel segmentation module, a 3D reconstruction module, a data calculation module, and a UAV line-following flight module;
[0044] The data acquisition module uses the onboard sensors of the UAV to acquire image data and uploads the data to the data processing module;
[0045] The data processing module preprocesses the image data based on the received data to obtain filtered image data, and then uploads the data to the subpixel segmentation module.
[0046] The subpixel segmentation module performs subpixel-level segmentation on the conductor and tower based on the received data and the filtered image data, obtains subpixel-level segmentation images of the conductor and tower, and uploads the data to the 3D reconstruction module.
[0047] The 3D reconstruction module uses the received data and the sub-pixel segmentation map of the conductor and tower as a basis to achieve high-precision reconstruction from 2D image to 3D space, obtain a 3D point cloud map, and upload the data to the data calculation module.
[0048] The data calculation module, based on the received data and the spatial geometry of the conductor and tower in the 3D point cloud map, uses the catenary equation as the core foundation to output sag data and upload the data to the UAV line-following flight module.
[0049] The drone's line-following flight module completes the drone's line-following flight based on the received data and the sag data.
[0050] This invention discloses a method and system for UAV line-following flight based on subpixel segmentation-optimized 3D reconstruction. Addressing the insufficient accuracy of traditional 3D reconstruction, it innovatively introduces subpixel-level segmentation technology to achieve refined segmentation of power lines and towers. This invention provides sufficiently accurate data support for high-precision identification of power transmission lines in complex scenarios, centimeter-level dynamic positioning, and equidistant line-following flight of UAVs, thereby improving the perception accuracy and operational reliability of autonomous UAV inspections. Attached Figure Description
[0051] Figure 1 This is a schematic flowchart of the method of the present invention;
[0052] Figure 2 This is a schematic diagram of the system of the present invention. Detailed Implementation
[0053] This invention provides a method for UAV line-following flight based on sub-pixel segmentation optimized 3D reconstruction, the flowchart of which is shown below. Figure 1 As shown, it includes the following steps:
[0054] S1: Acquire image data based on the onboard sensors of the UAV;
[0055] In step S1, the onboard sensors of the UAV include lidar and a visual camera.
[0056] S2: Based on the image data obtained in step S1, preprocess the image data to obtain filtered image data;
[0057] In step S2, the preprocessing is a nonlocal filtering method based on structural similarity measurement to filter the image data and achieve directional protection of the geometric features of the conductor through a feature coupling mechanism.
[0058] The preprocessing specifically includes:
[0059] First, a nonlocal search window is set up with each pixel as the center, and the pixels within the window are used as candidate matching objects. Then, a normalized composite weight function is constructed and normalized within this window. The weights are jointly determined by the similarity of brightness distribution, gradient field consistency, and local structure similarity. Furthermore, an anisotropic constraint factor is introduced to selectively modulate the pixel association, thereby strengthening the high-frequency information matching along the dominant direction of the conductor and weakening the erroneous associations dominated by lateral and noise. Finally, the pixel association model establishes nonlocal association relationships within the nonlocal search window with pixels as nodes and normalized composite weights as edge weights, achieving denoising while suppressing the interference of random noise on the shape of slender targets.
[0060] An adaptive smoothing strategy based on regional saliency discrimination constructs a saliency map by comprehensively analyzing the characteristics of different regions and their neighbors, including differences in brightness distribution, gradient response intensity, and consistency of line orientation. The filtering intensity of each region is then dynamically adjusted based on this saliency map: smoothing is enhanced in low-saliency regions to suppress noise, while smoothing is weakened in high-saliency regions to preserve wire edges and details. Continuous adjustment is used in transition regions to achieve a smooth transition, thus establishing a balance between noise suppression and target edge preservation. This leads to more accurate extraction of wire geometric features and improved wire recognition performance.
[0061] S3: Based on the filtered image data obtained in step S2, perform sub-pixel segmentation on the conductor and tower to obtain a sub-pixel segmentation image of the conductor and tower.
[0062] In step S3, based on the filtered image data obtained in step S2, a two-stage segmentation strategy from coarse to fine is applied to the conductors and towers, specifically as follows:
[0063] In the first stage, pixel-level coarse segmentation is performed to quickly locate the main areas of the conductor and tower. In the second stage, the sub-pixel-level fine segmentation stage is entered. Through resolution enhancement of local areas and boundary iterative optimization, the accuracy of contour positioning is improved, and the sub-pixel-level segmentation map of the conductor and tower is finally output.
[0064] The pixel-level coarse segmentation specifically refers to:
[0065] Constructing a fuzzy C-mean objective function The calculation of cluster centers and membership degrees is transformed from the pixel level to the gray level, and image classification is completed at the gray level; the fuzzy C-means objective function Express it using the following formula: ;in, Indicates the number of gray levels in an image. This represents the grayscale value of the current pixel. Indicates grayscale level The number of pixels, Indicates the number of cluster centers. Represents grayscale level The pixel belongs to the first Membership degree of each cluster center Represents the fuzzy index. Indicates the first The gray level of each cluster center;
[0066] By iteratively optimizing the membership degree and cluster center until the objective function converges, grayscale classification is finally completed based on the principle of maximum membership degree, achieving pixel-level coarse segmentation of conductors and towers;
[0067] The sub-pixel level refinement stage specifically refers to:
[0068] Regions of Interest (ROIs) are generated by filtering connected regions based on their area, and their pixel-level contours are extracted as initial boundaries. Background noise interference is filtered out by setting an area threshold, the main areas of the conductor and tower are retained, and their outer contours are extracted as initial contours, which effectively reduces the processing range and computational complexity.
[0069] A subpixel processing window is constructed at the contour points, and the image within the window is upsampled at a high magnification to enhance the local detail resolution, significantly improve the regional pixel density, and provide sufficient data support for subpixel-level adjustments.
[0070] In the upsampled region, the boundary position is initially located based on the gradient magnitude information, and the initial contour is iteratively adjusted and converged through the boundary evolution optimization method to gradually approach the true boundary and obtain the sub-pixel level boundary point set, thus achieving a substantial leap from the pixel level to the sub-pixel level.
[0071] The optimized subpixel-level boundary point set is mapped back to the original image coordinate system to preserve spatial geometric accuracy. A continuous and smooth contour curve is generated through piecewise curve fitting to obtain the subpixel-level conductor and tower contours, and finally, the subpixel-level segmentation map of the conductor and tower is obtained.
[0072] S4: Based on the sub-pixel segmentation map of the conductor and tower obtained in step S3, high-precision reconstruction from two-dimensional image to three-dimensional space is achieved to obtain a three-dimensional point cloud map.
[0073] Step S4 includes the following steps:
[0074] Key features such as conductor sag points and tower corner points are extracted from the sub-pixel segmentation images of conductors and towers, and cross-view feature matching is completed with the help of semantic consistency constraints.
[0075] By using the matched sub-pixel pairs, their three-dimensional spatial positions are recovered through the principle of triangulation, and an initial sparse point cloud is generated.
[0076] Through global optimization and dense matching, a dense point cloud containing details is obtained;
[0077] By performing surface fitting and mesh reconstruction based on the geometric properties of point clouds, a three-dimensional point cloud map for accurate sag calculation is constructed.
[0078] Step S4 is as follows:
[0079] Using sub-pixel-level segmentation maps of conductors and towers as core data support, feature points such as conductor inflection points and tower connection points are extracted, and the pixel positions of feature points are obtained through sub-pixel-level coordinate calibration.
[0080] Within the same category of conductors and towers, cross-view feature point matching is carried out with the help of semantic consistency constraints;
[0081] Based on sub-pixel level matched image points, combined with the camera's calibrated intrinsic and extrinsic parameters, the 3D structure is reconstructed using matching points obtained through triangulation. The projection equation is: ;in, , Represents the depth scale. Represents the camera intrinsic parameter matrix. , and , Let the rotation matrix and translation vector be... , For subpixel level matching of image points, Let be the homogeneous coordinates of a 3D point in world space; rearrange the projection equations, and the resulting linear system is expressed by the following formula: ;in, Camera projection matrix and the coordinates of the observed image points , Linear combination; solving for three-dimensional coordinates using singular value decomposition (SVD). This yields a preliminary sparse point cloud and camera pose; however, as an initial point cloud, this result is limited by noise during matching, making it difficult to meet overall accuracy requirements.
[0082] Global optimization is performed using bundle adjustment (BA), and the objective function is minimized using the following formula: ;in, For camera pose parameters, For three-dimensional point coordinates, For camera frame rate, The number of points in three dimensions; The projection equation is used to calculate the point. 3D coordinates In frame The projected pixel coordinates in the image; For the first Frame observation and The corresponding 2D pixel coordinates; camera parameters and sparse 3D point cloud are obtained through BA optimization;
[0083] A global energy function is constructed to transform the dense stereo matching problem into a global optimization objective containing data and smoothing terms. Guided by sub-pixel segmentation of the image, the data term calculates the disparity of pixels through precise segmentation boundaries, thereby improving the accuracy of feature localization. The smoothing term employs a semantic adaptive strategy, strengthening the disparity continuity constraint within similar objects while preserving the degree of freedom for disparity jumps at the boundaries of different objects. This global energy function fully utilizes segmentation priors to effectively overcome matching ambiguities in weakly textured regions and complex boundaries. Under the 3D positional constraints provided by sparse reconstruction, an optimal dense disparity map is obtained by minimizing the global energy function. This process is not an independent pixel-level decision but a global collaborative inference guided by a geometric framework composed of sparse feature points. Sparse point clouds, as precise anchor points in space, provide strong geometric constraints for disparity estimation, significantly improving the robustness of matching in weakly textured and occluded regions, thus ensuring that disparity estimation maintains both continuity and consistency globally while effectively suppressing noise interference.
[0084] Calculate 3D coordinates using dense disparity maps and obtained camera parameters This generates a dense point cloud containing details, represented by the following formula: ;in, Focal length Let these be the coordinates of the camera's principal point. Baseline length For parallax, These are the sub-pixel coordinates of the feature point (conductor or tower);
[0085] Based on the generated dense point cloud, an improved moving least squares (MLS) method is used for 3D reconstruction. To overcome the effects of point cloud discreteness and noise, local surface fitting is performed in the neighborhood of each point to robustly recover the underlying geometry. Then, a weighted averaging algorithm is used to seamlessly fuse all these local surfaces into a smooth overall surface, solving the surface fragmentation problem caused by local fitting, and finally constructing a 3D point cloud map for accurate sag calculation.
[0086] S5: Based on the spatial geometric information of the conductor and tower in the 3D point cloud obtained in step S4, output the sag data with the catenary equation as the core basis.
[0087] Step S5 is as follows:
[0088] Based on a smooth 3D model, the spatial geometric information of the conductor and tower is extracted. Taking the catenary equation of a symmetrical tower as the core foundation, the catenary equation is: ;in, It is a constant determined by the boundary conditions. , It refers to the unit mass of the conductor. It is horizontal tension. It is the acceleration due to gravity. The horizontal coordinates of the conductor along the span direction. The longitudinal coordinate of the conductor along the span direction;
[0089] Substituting the boundary conditions into the catenary equation, we obtain the coordinate expression of the traverse along the span direction: ;in, Horizontal spacing The maximum sag is obtained through the final derivation. The calculation expression is: .
[0090] S6. Based on the sag data obtained in step S5, complete the drone's line-following flight.
[0091] The present invention also provides a system for implementing the UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction, the structural schematic diagram of which is shown below. Figure 2 As shown, it includes a data acquisition module, a data processing module, a sub-pixel segmentation module, a 3D reconstruction module, a data calculation module, and a UAV line-following flight module;
[0092] The data acquisition module uses the onboard sensors of the UAV to acquire image data and uploads the data to the data processing module;
[0093] The data processing module preprocesses the image data based on the received data to obtain filtered image data, and then uploads the data to the subpixel segmentation module.
[0094] The subpixel segmentation module performs subpixel-level segmentation on the conductor and tower based on the received data and the filtered image data, obtains subpixel-level segmentation images of the conductor and tower, and uploads the data to the 3D reconstruction module.
[0095] The 3D reconstruction module uses the received data and the sub-pixel segmentation map of the conductor and tower as a basis to achieve high-precision reconstruction from 2D image to 3D space, obtain a 3D point cloud map, and upload the data to the data calculation module.
[0096] The data calculation module, based on the received data and the spatial geometry of the conductor and tower in the 3D point cloud map, uses the catenary equation as the core foundation to output sag data and upload the data to the UAV line-following flight module.
[0097] The drone's line-following flight module completes the drone's line-following flight based on the received data and the sag data.
Claims
1. A method for UAV line-following flight based on sub-pixel segmentation optimized 3D reconstruction, characterized in that, Includes the following steps: S1: Acquire image data based on the onboard sensors of the UAV; S2: Based on the image data obtained in step S1, preprocess the image data to obtain filtered image data; S3: Based on the filtered image data obtained in step S2, perform sub-pixel segmentation on the conductor and tower to obtain a sub-pixel segmentation image of the conductor and tower. S4: Based on the sub-pixel segmentation map of the conductor and tower obtained in step S3, high-precision reconstruction from two-dimensional image to three-dimensional space is achieved to obtain a three-dimensional point cloud map. S5: Based on the spatial geometric information of the conductor and tower in the 3D point cloud obtained in step S4, output the sag data with the catenary equation as the core basis. S6. Based on the sag data obtained in step S5, complete the drone's line-following flight.
2. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 1, characterized in that, In step S2, the preprocessing is a nonlocal filtering method based on structural similarity measurement to filter the image data and achieve directional protection of the geometric features of the conductor through a feature coupling mechanism.
3. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 2, characterized in that, The preprocessing specifically includes: First, a non-local search window is set up with each pixel as the center, and the pixels in the window are used as candidate matching objects. Then, a composite weight function is constructed and normalized in the window. The weight is determined by the similarity of brightness distribution, the consistency of gradient field and the similarity of local structure. Furthermore, an anisotropic constraint factor is introduced to selectively modulate the pixel association, thereby strengthening the high-frequency information matching along the dominant direction of the conductor and weakening the erroneous association dominated by the lateral direction and noise. Ultimately, the pixel association model establishes non-local association relationships within the non-local search window using pixels as nodes and normalized composite weights as edge weights, achieving denoising while suppressing the interference of random noise on the shape of slender targets. Furthermore, an adaptive smoothing strategy based on regional saliency discrimination constructs a saliency map by comprehensively considering the characteristics of different regions and their neighbors, including differences in brightness distribution, gradient response intensity, and consistency of linear direction. Subsequently, the filtering intensity of each region is dynamically adjusted based on the saliency map: smoothing is enhanced in low saliency regions to suppress noise, smoothing is weakened in high saliency regions to preserve the edges and details of the conductors, and continuous adjustment is used in transition regions to achieve a smooth transition. This establishes a balance between noise suppression and target edge preservation, thereby extracting the geometric features of the conductors more accurately and improving the conductor recognition effect.
4. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 1, characterized in that, In step S3, based on the filtered image data obtained in step S2, a two-stage segmentation strategy from coarse to fine is applied to the conductors and towers, specifically as follows: In the first stage, pixel-level coarse segmentation is performed to lock the main areas of the conductor and tower. In the second stage, the sub-pixel-level fine segmentation stage is entered. Through resolution enhancement of local areas and boundary iterative optimization, the accuracy of contour positioning is improved, and the sub-pixel-level segmentation map of the conductor and tower is finally output.
5. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 4, characterized in that, The pixel-level coarse segmentation specifically refers to: Constructing a fuzzy C-mean objective function The calculation of cluster centers and membership degrees is transformed from the pixel level to the gray level, and image classification is completed at the gray level; the fuzzy C-means objective function Express it using the following formula: ;in, Indicates the number of gray levels in an image. This represents the grayscale value of the current pixel. Indicates grayscale level The number of pixels, Indicates the number of cluster centers. Represents grayscale level The pixel belongs to the first Membership degree of each cluster center Represents the fuzzy index. Indicates the first The gray level of each cluster center; By iteratively optimizing the membership degree and cluster center until the objective function converges, grayscale classification is finally completed based on the principle of maximum membership degree, achieving pixel-level coarse segmentation of conductors and towers; The sub-pixel level refinement stage specifically refers to: The ROI region is generated by filtering the area of the connected region and its pixel-level contour is extracted as the initial boundary. Background noise interference is filtered by setting an area threshold, the main area of the conductor and tower is retained, and its outer contour is extracted as the initial contour. A subpixel processing window is constructed at the contour points, and the image within the window is upsampled at a high magnification to enhance the local detail resolution; In the upsampled region, the boundary position is initially located based on the gradient magnitude information, and the initial contour is iteratively adjusted and converged through the boundary evolution optimization method to gradually approach the real boundary and obtain a sub-pixel level boundary point set. The optimized subpixel-level boundary point set is mapped back to the original image coordinate system to preserve spatial geometric accuracy. A continuous and smooth contour curve is generated through piecewise curve fitting to obtain the subpixel-level conductor and tower contours, and finally, the subpixel-level segmentation map of the conductor and tower is obtained.
6. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 1, characterized in that, Step S4 includes the following steps: Features including conductor sag points and tower corner points are extracted from the sub-pixel segmentation images of conductors and towers, and cross-view feature matching is completed with the help of semantic consistency constraints; By using the matched sub-pixel pairs, their three-dimensional spatial positions are recovered through the principle of triangulation, and an initial sparse point cloud is generated. Through global optimization and dense matching, a dense point cloud containing details is obtained; By performing surface fitting and mesh reconstruction based on the geometric properties of point clouds, a three-dimensional point cloud map for accurate sag calculation is constructed.
7. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 6, characterized in that, Step S4 is as follows: Using sub-pixel-level segmentation maps of conductors and towers as core data support, feature points including conductor inflection points and tower connection points are extracted, and the pixel positions of feature points are obtained through sub-pixel-level coordinate calibration. Within the same category of conductors and towers, cross-view feature point matching is carried out with the help of semantic consistency constraints; Based on sub-pixel level matched image points, combined with the camera's calibrated intrinsic and extrinsic parameters, the 3D structure is reconstructed using matching points obtained through triangulation. The projection equation is: ;in, , Represents the depth scale. This represents the camera intrinsic parameter matrix. , and , Let the rotation matrix and translation vector be... , For subpixel level matching of image points, Let be the homogeneous coordinates of a 3D point in world space; rearrange the projection equations, and the resulting linear system is expressed by the following formula: ;in, Camera projection matrix and the coordinates of the observed image points , Linear combination; solving for three-dimensional coordinates through singular value decomposition. Thus, a preliminary sparse point cloud and camera pose are obtained; Global optimization is performed using the bundle adjustment method, and the objective function is minimized using the following formula: ;in, For camera pose parameters, For three-dimensional point coordinates, For camera frame rate, The number of points in three dimensions; The projection equation is used to calculate the point. 3D coordinates In frame The projected pixel coordinates in the image; For the first Frame observation and The corresponding 2D pixel coordinates; camera parameters and sparse 3D point cloud are obtained through BA optimization; A global energy function is constructed, transforming the dense stereo matching problem into a global optimization objective containing data and smoothing terms. Guided by sub-pixel segmentation of the image, the data term calculates the disparity of pixels through precise segmentation boundaries, while the smoothing term employs a semantically adaptive strategy to strengthen disparity continuity constraints within similar objects and preserve the degrees of freedom for disparity jumps at boundaries between different objects. Under the 3D positional constraints provided by sparse reconstruction, an optimal dense disparity map is obtained by minimizing the global energy function. The 3D coordinates are then calculated using the dense disparity map and the obtained camera parameters. This generates a dense point cloud containing details, represented by the following formula: ;in, Focal length Let these be the coordinates of the camera's principal point. Baseline length For parallax, These are the sub-pixel coordinates of the feature point (conductor or tower); Based on the generated dense point cloud, an improved moving least squares method is used for 3D reconstruction. Local surface fitting is performed in the neighborhood of each point, and then all these local surfaces are seamlessly merged into a smooth overall surface through a weighted average algorithm, finally constructing a 3D point cloud map for accurate sag calculation.
8. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 1, characterized in that, Step S5 includes the following steps: Based on a smooth 3D model, information about conductors and towers is extracted, and the equation of the catenary is constructed. Substituting the boundary conditions into the catenary equation, we obtain the coordinate expression of the conductor along the span direction; Based on the catenary equation and the coordinate expression of the conductor along the span direction, the sag data is derived and calculated.
9. The UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction according to claim 8, characterized in that, Step S5 is as follows: Based on a smooth 3D model, the spatial geometric information of the conductor and tower is extracted. Taking the catenary equation of a symmetrical tower as the core foundation, the catenary equation is: ;in, It is a constant determined by the boundary conditions. , It refers to the unit mass of the conductor. It is horizontal tension. It is the acceleration due to gravity. The horizontal coordinates of the conductor along the span direction. The longitudinal coordinate of the conductor along the span direction; Substituting the boundary conditions into the catenary equation, we obtain the coordinate expression of the traverse along the span direction: ;in, Horizontal spacing The maximum sag is obtained through the final derivation. The calculation expression is: .
10. A system for implementing the UAV line-following flight method based on sub-pixel segmentation optimized 3D reconstruction as described in any one of claims 1 to 9, characterized in that, It includes a data acquisition module, a data processing module, a sub-pixel segmentation module, a 3D reconstruction module, a data calculation module, and a drone line-following flight module; The data acquisition module uses the onboard sensors of the UAV to acquire image data and uploads the data to the data processing module; The data processing module preprocesses the image data based on the received data to obtain filtered image data, and then uploads the data to the subpixel segmentation module. The subpixel segmentation module performs subpixel-level segmentation on the conductor and tower based on the received data and the filtered image data, obtains subpixel-level segmentation images of the conductor and tower, and uploads the data to the 3D reconstruction module. The 3D reconstruction module uses the received data and the sub-pixel segmentation map of the conductor and tower as a basis to achieve high-precision reconstruction from 2D image to 3D space, obtain a 3D point cloud map, and upload the data to the data calculation module. The data calculation module, based on the received data and the spatial geometry of the conductor and tower in the 3D point cloud map, uses the catenary equation as the core foundation to output sag data and upload the data to the UAV line-following flight module. The drone's line-following flight module completes the drone's line-following flight based on the received data and the sag data.