Power transmission tower three-dimensional reconstruction method and system based on Vis-MVSNet
The Vis-MVSNet-based method addresses the challenges of high-precision and efficient 3D reconstruction of power towers by combining SuperPoint and AdaLAM for sparse reconstruction, followed by Vis-MVSNet for dense reconstruction, resulting in accurate and efficient 3D models.
Patent Information
- Application Number
- CN202510204418.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-15
AI Technical Summary
In the process of sparse reconstruction and intensive reconstruction, it is difficult to efficiently and accurately generate high-precision three-dimensional model of transmission tower in complex scenarios, especially when facing areas with strong reflections on surface materials and scarce texture characteristics, the traditional method has large matching errors, and the deep learning methods are applicable to different scenarios, resulting in insufficient reconstruction quality and efficiency.
The incremental SfM sparse reconstruction method combined with SuperPoint and AdaLAM is adopted, and the intensive reconstruction method of Vis-MVSNet is combined with the three-dimensional reconstruction method, and a high-precision three-dimensional model of the transmission tower is generated through feature extraction, matching and three-dimensional model construction.
While maintaining high computing efficiency, a high-precision three-dimensional model of transmission tower is generated, suitable for three-dimensional reconstruction tasks in complex scenarios, improving the accuracy and robustness of feature matching, and enhancing the detailed performance and realism of the model.
Smart Images

Figure CN120318453A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and machine learning, and more specifically, to a three-dimensional reconstruction method and system of a transmission tower based on Vis-MVSNet. Background Art
[0002] With the rapid progress of unmanned aerial vehicle (UAV) remote sensing technology, power facility inspection based on UAVs has become a research hotspot in the grid automation inspection solution. A high-precision three-dimensional model of a transmission tower can provide a spatial data foundation for automatic inspection route generation and plan formulation, etc., which is a key and difficult problem to be solved in grid automation inspection. The current three-dimensional reconstruction process of a transmission tower based on UAV images includes steps such as data collection and preparation, sparse reconstruction, dense reconstruction, and post-processing. Among them, sparse reconstruction and dense reconstruction are two core links. Sparse reconstruction is used to restore the accurate imaging parameters of UAV images, and dense reconstruction is responsible for generating a dense three-dimensional point cloud based on the images and the corresponding imaging parameters, and further generating a three-dimensional model with color texture.
[0003] Currently, there are two common framework processes for sparse reconstruction. One is the aerial triangulation framework commonly used in aerial photogrammetry. This framework method takes images, their initial imaging parameters, and UAV camera parameters as inputs, and optimizes the imaging parameters of the images through steps such as feature matching and block adjustment. The other is to use the structure-from-motion (SfM) framework in computer vision. This framework is generally the same as the aerial triangulation method in terms of process, but compared with the aerial triangulation method, it does not require pre-calibrated camera parameters, so it has stronger adaptability. The key basis of the SfM framework is image feature extraction and high-precision matching between features. However, there are currently a wide variety of feature extraction algorithms, such as SIFT, LIFT, SURF, SuperPoint, etc., each with its own advantages and disadvantages. Therefore, it is necessary to sort out the characteristics of UAV images in the power facility inspection scenario, so as to screen out the most suitable feature extraction and matching algorithms for this scenario.
[0004] Dense reconstruction starts from the accurate imaging parameters of the images obtained by sparse reconstruction. Through multi-view stereo matching, a dense point cloud is obtained, and then through a series of steps such as mesh construction and optimization, and texture mapping, a three-dimensional model with color texture is obtained. In this process, multi-view stereo matching, as the core step, still has many challenges in terms of point cloud density and geometric accuracy. Currently, common multi-view stereo matching methods are divided into traditional semi-global matching-based methods (SGM) and recently deep learning network-based methods. However, traditional SGM methods rely on manual features and heuristic rules, and show obvious limitations when dealing with complex scenes and objects with large geometric deformations. For example, in the face of areas with strong surface material reflection and lack of texture features, traditional SGM methods are difficult to accurately find corresponding points, resulting in an increase in matching errors, which in turn affects the density and quality of the point cloud; while deep learning network-based methods, although overcoming some shortcomings of traditional methods to a certain extent, can utilize the powerful feature extraction ability of deep neural networks to automatically learn feature representations in images and achieve good matching results in complex scenes. However, there are currently a wide variety of deep learning methods, and each method is suitable for different scenarios. Therefore, it is also necessary to sort out the three-dimensional characteristics of transmission towers in the power facility inspection scenario, so as to screen out the dense reconstruction network most suitable for this scenario.
[0005] Therefore, a three-dimensional reconstruction method of transmission towers based on Vis-MVSNet is needed. Summary of the Invention
[0006] The present invention proposes a three-dimensional reconstruction method and system of transmission towers based on Vis-MVSNet to solve the problem of how to efficiently perform three-dimensional reconstruction of transmission towers.
[0007] To solve the above problems, according to one aspect of the present invention, a three-dimensional reconstruction method of transmission towers based on Vis-MVSNet is provided, and the method includes:
[0008] Collect the original sequence of drone images covering multiple perspectives of the transmission tower, and preprocess the original data images to obtain processed sequence images;
[0009] Based on the processed sequence images, use an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM for incremental reconstruction to obtain a sparse point cloud model and image imaging parameters;
[0010] Based on the sparse point cloud model and image imaging parameters, use a dense reconstruction method based on Vis-MVSNet for dense reconstruction to obtain an initial three-dimensional model of the transmission tower;
[0011] Perform transformation processing on the initial three-dimensional model of the transmission tower to obtain the final three-dimensional model of the transmission tower.
[0012] Preferably, the preprocessing of the original data image to obtain a processed sequence of images includes:
[0013] Read the GPS-aided information and camera parameter information of the images in the original sequence of images, and perform wallis filtering on all the images to obtain a processed sequence of images with enhanced image quality.
[0014] Preferably, based on the processed sequence of images, an incremental reconstruction is performed using an incremental SfM sparse reconstruction method that combines SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters, including:
[0015] Use the SuperPoint deep network to extract the feature point positions and feature vectors from the processed sequence of images;
[0016] Use the AdaLAM algorithm to perform feature matching on the feature points extracted from any two images to determine the image pairs;
[0017] Select an initial image pair and gradually perform scene reconstruction using the incremental SfM method to obtain a sparse point cloud model and image imaging parameters.
[0018] Preferably, based on the sparse point cloud model and image imaging parameters, a dense reconstruction is performed using a dense reconstruction method based on Vis-MVSNet to obtain a three-dimensional initial model of the transmission tower, including:
[0019] Based on the sparse point cloud model and image imaging parameters, use the Vis-MVSNet dense matching method to perform matching to obtain a dense three-dimensional point cloud;
[0020] Using the dense three-dimensional point cloud as input, connect the points in the point cloud into a TIN grid through the Delaunay triangulation algorithm to perform TIN construction and obtain a TIN structure;
[0021] Perform mesh optimization on the TIN structure to obtain a three-dimensional spatial structure mesh;
[0022] Attach the sequence of images as textures to the surface of the three-dimensional spatial structure mesh to obtain a three-dimensional initial model of the transmission tower.
[0023] Preferably, the conversion process of the three-dimensional initial model of the transmission tower to obtain a three-dimensional final model of the transmission tower includes:
[0024] Convert the local coordinate system of the three-dimensional initial model of the transmission tower to the global coordinate system and then convert it to a three-dimensional file format to obtain a three-dimensional final model of the transmission tower.
[0025] According to another aspect of the present invention, there is provided a three-dimensional reconstruction system of a transmission tower based on Vis-MVSNet, and the system includes:
[0026] An image collection unit, configured to collect the original sequence images of the drone covering multiple perspectives of the transmission tower, and preprocess the original data images to obtain processed sequence images;
[0027] A sparse reconstruction unit, configured to perform incremental reconstruction based on the processed sequence images by using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters;
[0028] A dense reconstruction unit, configured to perform dense reconstruction based on the sparse point cloud model and image imaging parameters by using a dense reconstruction method based on Vis-MVSNet to obtain a three-dimensional initial model of the transmission tower;
[0029] A conversion unit, configured to perform conversion processing on the three-dimensional initial model of the transmission tower to obtain a three-dimensional final model of the transmission tower.
[0030] Preferably, the image collection unit preprocesses the original data images to obtain processed sequence images, including:
[0031] Reading the GPS auxiliary information and camera parameter information of the images in the original sequence images, and performing wallis filtering on all the images to obtain processed sequence images with enhanced image quality.
[0032] Preferably, the sparse reconstruction unit performs incremental reconstruction based on the processed sequence images by using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters, including:
[0033] Using the SuperPoint deep network to extract the feature point positions and feature vectors from the processed sequence images;
[0034] Using the AdaLAM algorithm to perform feature matching on the feature points extracted from any two images to determine the image pairs;
[0035] Selecting an initial image pair, and gradually performing scene reconstruction by using the incremental SfM method to obtain a sparse point cloud model and image imaging parameters.
[0036] Preferably, the dense reconstruction unit performs dense reconstruction based on the sparse point cloud model and image imaging parameters by using a dense reconstruction method based on Vis-MVSNet to obtain a three-dimensional initial model of the transmission tower, including:
[0037] Based on the sparse point cloud model and image imaging parameters, use the Vis-MVSNet dense matching method for matching to obtain a dense three-dimensional point cloud;
[0038] Taking the dense three-dimensional point cloud as the input, connect the points in the point cloud into a TIN grid through the Delaunay triangulation algorithm, perform TIN construction to obtain a TIN structure;
[0039] Perform mesh optimization on the TIN structure to obtain a three-dimensional spatial structure mesh;
[0040] Attach the sequence of images as textures to the surface of the three-dimensional spatial structure mesh to obtain an initial three-dimensional model of the transmission tower.
[0041] Preferably, the conversion unit performs conversion processing on the initial three-dimensional model of the transmission tower to obtain a final three-dimensional model of the transmission tower, including:
[0042] Convert the local coordinate system of the initial three-dimensional model of the transmission tower to the global coordinate system, and then convert it to a three-dimensional file format to obtain a final three-dimensional model of the transmission tower.
[0043] The present invention provides a method and system for three-dimensional reconstruction of a transmission tower based on Vis-MVSNet, including: collecting original sequence images of a drone covering multiple perspectives of the transmission tower, and preprocessing the original data images to obtain processed sequence images; based on the processed sequence images, using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM for incremental reconstruction to obtain a sparse point cloud model and image imaging parameters; based on the sparse point cloud model and image imaging parameters, using a dense reconstruction method based on Vis-MVSNet for dense reconstruction to obtain an initial three-dimensional model of the transmission tower; performing conversion processing on the initial three-dimensional model of the transmission tower to obtain a final three-dimensional model of the transmission tower. The present invention can generate high-precision three-dimensional reconstruction results while maintaining high computational efficiency, and is applicable to three-dimensional reconstruction tasks of complex scenes such as transmission towers and buildings. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] By referring to the following drawings, the exemplary embodiments of the present invention can be more fully understood:
[0045] Figure 1 It is a flowchart of a three-dimensional reconstruction method 100 of a transmission tower based on Vis-MVSNet according to an embodiment of the present invention;
[0046] Figure 2 It is a flowchart of an incremental SfM sparse reconstruction combining SuperPoint and AdaLAM according to an embodiment of the present invention;
[0047] Figure 3 Flowchart of dense reconstruction based on Vis - MVSNet according to an embodiment of the present invention;
[0048] Figure 4 Schematic structural diagram of a three - dimensional reconstruction system 400 of a transmission tower based on Vis - MVSNet according to an embodiment of the present invention. Detailed implementation manners
[0049] Now, exemplary embodiments of the present invention will be introduced with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely, and to fully convey the scope of the present invention to those skilled in the art. The terms in the exemplary embodiments shown in the drawings are not limitations to the present invention. In the drawings, the same units / components are denoted by the same reference numerals.
[0050] Unless otherwise specified, the terms (including scientific and technical terms) used herein have the ordinary meaning understood by those skilled in the art. Additionally, it can be understood that terms defined in a commonly used dictionary should be understood to have a meaning consistent with the context of their related fields, and should not be understood in an idealized or overly formal sense.
[0051] Aiming at the problems of insufficient stability in the existing three - dimensional reconstruction process of transmission towers and low accuracy of the reconstructed three - dimensional models, the present invention provides a three - dimensional reconstruction method of a transmission tower based on Vis - MVSNet. Taking the sequence images of the transmission tower captured by a drone as input, an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM is used to recover high - precision image imaging parameters, then a dense reconstruction method based on Vis - MVSNet is used to obtain the three - dimensional model of the transmission tower, and finally, after post - processing steps such as coordinate transformation and format conversion, a high - precision three - dimensional model of the transmission tower is output.
[0052] Figure 1 Flowchart of a three - dimensional reconstruction method 100 of a transmission tower based on Vis - MVSNet according to an embodiment of the present invention. As Figure 1 shown, the three - dimensional reconstruction method of a transmission tower based on Vis - MVSNet provided by the embodiment of the present invention can maintain a high computational efficiency while generating high - precision three - dimensional reconstruction results, and is applicable to three - dimensional reconstruction of complex scenes such as transmission towers and buildings. Starting from step 101, in step 101, the original sequence images of the drone covering multiple perspectives of the transmission tower are collected, and the original data images are pre - processed to obtain processed sequence images.
[0053] Preferably, the pre - processing of the original data images to obtain processed sequence images includes:
[0054] Read the GPS-aided information and camera parameter information of the images in the original sequence images, and perform Wallis filtering on all the images to obtain the processed sequence images with enhanced image quality.
[0055] In the present invention, it is necessary to collect the UAV sequence images covering multiple perspectives of the transmission tower, and preprocess the sequence images, including: reading the image GPS-aided information, camera parameter information, and performing Wallis filtering on all the images to obtain the sequence images with enhanced image quality.
[0056] The present invention uses the Wallis filtering algorithm to enhance the contrast of all images and reduce the color difference between images, which can effectively improve the contrast and detail performance of the images, especially having a significant effect in scenes with uneven illumination or low contrast. This processing method can effectively enhance the texture details and edge information in the transmission tower images, making the subsequent feature extraction and matching more accurate and reliable.
[0057] In step 102, based on the processed sequence images, an incremental reconstruction is performed using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters.
[0058] Preferably, the incremental reconstruction using the incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM based on the processed sequence images to obtain a sparse point cloud model and image imaging parameters includes:
[0059] Use the SuperPoint deep network to extract the feature point positions and feature vectors from the processed sequence images;
[0060] Use the AdaLAM algorithm to perform feature matching on the feature points extracted from any two images to determine the image pairs;
[0061] Select the initial image pairs, and gradually perform scene reconstruction using the incremental SfM method to obtain a sparse point cloud model and image imaging parameters.
[0062] In the present invention, after obtaining the sequence images with enhanced quality, the sequence images are input into the incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM, and the image imaging parameters are output.
[0063] Combined Figure 2 As shown, the incremental SfM sparse reconstruction method specifically includes the following steps:
[0064] S1. Feature extraction: Use the advanced SuperPoint deep network to extract high-precision feature points and corresponding feature vectors from each input image. The SuperPoint network is a self-supervised feature extraction method based on deep learning, which can automatically detect key points in the image and generate discriminative feature descriptors. Specifically, first input the preprocessed image into the SuperPoint network. The network performs multi-level feature learning on the image through convolutional layers and feature encoders, and finally outputs the positions of significant feature points in the image and their corresponding feature vectors.
[0065] S2. Feature matching: Use the AdaLAM (Adaptive Locally Affine Matching) algorithm to efficiently and robustly match the SuperPoint feature points extracted from pairwise images. AdaLAM is an adaptive feature matching algorithm based on local affine transformation, which can effectively handle problems such as perspective changes, scale differences, and local deformations existing between images. Specifically, the AdaLAM algorithm first uses the feature points and feature vectors extracted by SuperPoint to generate initial matching pairs by calculating the similarity (such as cosine distance or Euclidean distance) between feature vectors. Subsequently, the algorithm verifies and optimizes the initial matching pairs through a local affine model, and uses geometric consistency constraints to eliminate incorrect matching pairs.
[0066] S3. Incremental reconstruction: Select a pair of images with sufficient overlapping regions and a high number of matching points from the sequence of images as the initial image pair. Calculate the relative pose (i.e., image imaging parameters) of the initial image pair through epipolar geometry, and obtain an initial sparse 3D point cloud through triangulation. Select the image with the most matching points with the current reconstruction model from the remaining images as the next image to be added. Use the PnP algorithm to estimate the camera pose and generate new 3D points through triangulation. After each addition, adjust and optimize the reconstruction model through bundle adjustment to minimize the reprojection error. This process is iterated until all images are processed, and finally a globally consistent sparse 3D model is formed, and an accurate sparse point cloud model and image imaging parameters are output.
[0067] The present invention uses the SuperPoint network to extract feature points, which are usually located at the edges, corners or other regions with significant texture changes in the image, and can effectively characterize the structural features of the transmission tower. In addition, the feature vectors extracted by the SuperPoint network have high robustness and can maintain consistency under different perspectives, lighting conditions and scale changes, providing reliable basic data for subsequent multi-view matching and 3D reconstruction. Through this step, the accuracy of feature matching and the precision of the reconstruction model can be significantly improved.
[0068] The present invention uses the AdaLAM algorithm to match feature points. This process can effectively improve the accuracy and robustness of the matching by adaptively analyzing the geometric relationships in the local area. In addition, the AdaLAM algorithm also introduces a multi-scale strategy, which can match feature points in different scale spaces, further enhancing the adaptability to scale changes. Through the matching process of the AdaLAM algorithm, high-quality corresponding relationships of feature points can be obtained, providing reliable basic data for subsequent multi-view geometric reconstruction and depth map generation. This step significantly improves the accuracy and stability of the matching, providing an important guarantee for the overall accuracy of the three-dimensional reconstruction of transmission towers.
[0069] In step 103, based on the sparse point cloud model and image imaging parameters, a dense reconstruction is performed using a dense reconstruction method based on Vis-MVSNet to obtain an initial three-dimensional model of the transmission tower.
[0070] Preferably, the dense reconstruction using a dense reconstruction method based on Vis-MVSNet based on the sparse point cloud model and image imaging parameters to obtain an initial three-dimensional model of the transmission tower includes:
[0071] Based on the sparse point cloud model and image imaging parameters, a dense matching method of Vis-MVSNet is used for matching to obtain a dense three-dimensional point cloud;
[0072] Taking the dense three-dimensional point cloud as the input, the points in the point cloud are connected into a TIN grid through the Delaunay triangulation algorithm to perform TIN construction and obtain a TIN structure;
[0073] Perform mesh optimization on the TIN structure to obtain a three-dimensional space structure mesh;
[0074] Attach the sequence images as textures to the surface of the three-dimensional space structure mesh to obtain an initial three-dimensional model of the transmission tower.
[0075] In the present invention, the sparse point cloud model and image imaging parameters obtained in step 102 are input into a dense reconstruction method based on Vis-MVSNet for dense reconstruction, and a three-dimensional model of the transmission tower is output.
[0076] Combined Figure 3 As shown, the dense reconstruction method based on Vis-MVSNet specifically includes the following steps:
[0077] S1. Dense matching: Use the Vis-MVSNet dense matching method to generate a dense three-dimensional point cloud. First, use the Vis-MVSNet neural network to extract multi-scale feature maps and construct a cost volume; then, optimize feature matching through a visual attention mechanism and an adaptive view selection strategy, generate an initial depth map and optimize it using a probabilistic depth regression method; finally, fuse multi-view depth maps to generate a dense point cloud, and remove noise points through post-processing to output a high-precision dense three-dimensional point cloud.
[0078] S2. Construct TIN: Take the dense three-dimensional point cloud as input, and connect the points in the point cloud into a TIN mesh through the Delaunay triangulation algorithm, ensuring that no other points are contained within the circumcircle of each triangle, thereby generating a TIN structure.
[0079] S3. Mesh optimization: Optimize the TIN according to the transmission tower characteristics or application requirements, such as removing long edges or flat triangles, and retaining important transmission tower feature points to improve the quality and applicability of the TIN, and finally obtain a common three-dimensional space structure mesh. Specifically, first detect and repair defects in the TIN mesh, such as holes, overlapping faces, and self-intersecting faces; then reduce the number of redundant triangles through TIN mesh simplification while retaining important geometric features, or increase the mesh density through subdivision to enhance the detail performance; finally, smooth the mesh to eliminate irregular geometric noise.
[0080] S4. Texture mapping: Attach the sequence of images as textures to the surface of the three-dimensional model mesh to enhance the realism and detail performance of the model. First, transform the texture mapping problem into a global optimization problem by defining an energy function. The energy function includes a data term and a smooth term. The data term measures the matching degree between the texture image and the model surface, and the smooth term ensures natural texture transition in adjacent regions; then use algorithms such as Graph Cut or Belief Propagation to solve the minimum value of the energy function to obtain the optimal texture mapping result; finally, eliminate seams and distortions through post-processing to generate a high-quality and visually consistent texture map.
[0081] The present invention uses the Vis-MVSNet algorithm to generate dense three-dimensional point clouds. By introducing a visual attention mechanism and an adaptive view selection strategy, it can significantly improve the accuracy and robustness of depth estimation. The visual attention mechanism dynamically weights the feature information of different perspectives, highlighting important features and suppressing noise, while the adaptive view selection strategy automatically selects the optimal view combination according to the scene characteristics, effectively handling complex problems such as occlusion, repetitive textures, and illumination changes. In addition, Vis-MVSNet adopts multi-scale feature extraction and probabilistic depth regression methods to further optimize the quality of the depth map, enabling it to generate high-precision three-dimensional reconstruction results while maintaining high computational efficiency, and is applicable to three-dimensional reconstruction tasks of complex scenes such as transmission towers and buildings.
[0082] In step 104, the three-dimensional initial model of the transmission tower is subjected to transformation processing to obtain the three-dimensional final model of the transmission tower.
[0083] Preferably, the transformation processing of the three-dimensional initial model of the transmission tower to obtain the three-dimensional final model of the transmission tower includes:
[0084] The local coordinate system of the three-dimensional initial model of the transmission tower is converted to the global coordinate system and then converted into a three-dimensional file format to obtain the three-dimensional final model of the transmission tower.
[0085] In the present invention, the three-dimensional initial model of the transmission tower obtained in step 103 can output the final three-dimensional model of the transmission tower after post-processing steps such as coordinate transformation and format transformation. The specific process includes: First, the model is converted from the local coordinate system to the global coordinate system (such as the geographic coordinate system) through coordinate transformation to ensure the accurate position and orientation of the model in the actual scene; then, the model is subjected to format transformation and converted into a common three-dimensional file format (such as OBJ, FBX, or PLY) to be compatible with different three-dimensional software and platforms; in addition, the model is subjected to detail optimization, including removing redundant data, repairing geometric defects, and optimizing texture mapping, to ensure that the model reaches high precision and high fidelity in terms of geometric structure and visual effect; the finally generated three-dimensional model of the transmission tower can be directly used for engineering analysis, visualization display, or further three-dimensional applications, providing reliable data support for the detection, maintenance, and planning of transmission towers.
[0086] The following specifically exemplifies the implementation manner of the present invention
[0087] In an embodiment of the present invention, for the 3D reconstruction algorithm of transmission towers based on Vis-MVSNet, first, a multi-view sequence of images covering the transmission tower is collected by an unmanned aerial vehicle (UAV), the image auxiliary information is read, and the images are filtered by Wallis to enhance the image quality; then, an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM is used to calculate the image imaging parameters; on this basis, a dense reconstruction method based on Vis-MVSNet is adopted to obtain the 3D model of the transmission tower; finally, after a series of post-processing operations, the final 3D model of the transmission tower is output. The specific process includes:
[0088] Step 1: Use a UAV equipped with a high-resolution camera to take multi-angle photos around the transmission tower to obtain a multi-view sequence of images of the transmission tower, and preprocess the sequence of images to obtain a sequence of images with enhanced quality. The preprocessing process includes reading GPS auxiliary information, camera parameter information, etc. from the image file. If such information does not exist in the image file, the user inputs this information from the configuration file; there is also a key link in the preprocessing, which is to perform Wallis filtering on all images. Wallis filtering is a local adaptive image enhancement algorithm that can effectively improve the contrast and detail performance of images, especially in scenes with uneven illumination or low contrast. The core idea of this filter is to enhance the local contrast of the image by adjusting the mean and variance within the neighborhood of each pixel, while suppressing the influence of noise. In the specific implementation process, Wallis filtering will perform local statistical analysis on each pixel of the image, calculate the gray mean and variance within its neighborhood, and dynamically adjust the pixel according to the preset target mean and variance values.
[0089] Step 2: The incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM has a process as Figure 2 shown. Traditional sparse reconstruction methods usually rely on manually designed feature extraction algorithms (such as SIFT or ORB) and geometric constraint-based matching strategies. These methods have many deficiencies when dealing with complex scenes such as transmission towers. For example, manually designed feature extraction algorithms are difficult to extract a sufficient number of stable feature points in complex texture or low-contrast areas, and traditional matching methods have poor robustness to view changes, illumination differences, and repetitive textures, and are prone to false matches, affecting the accuracy of sparse reconstruction.
[0090] Aiming at the problems of low matching accuracy, poor computational efficiency and insufficient robustness in traditional sparse reconstruction methods when dealing with complex scenarios such as transmission towers, an incremental SfM sparse reconstruction method based on the combination of SuperPoint and AdaLAM is adopted. The core idea of this method is to utilize the high-quality feature points and their feature vectors extracted by the SuperPoint deep network, and combine with the AdaLAM adaptive local affine matching algorithm to achieve efficient and robust feature matching and incremental reconstruction. Specifically, first, highly discriminative feature points and feature vectors are extracted from multi-view images through the SuperPoint network, and these feature points can accurately represent the structural details and texture information of the transmission tower. Subsequently, the AdaLAM algorithm is used to match the extracted feature points. AdaLAM adaptively optimizes the matching results through the local affine transformation model, and can effectively handle complex situations such as view changes, occlusions and repetitive textures, significantly improving the accuracy and robustness of the matching. After the matching is completed, the incremental SfM method is used to gradually reconstruct the sparse point cloud model of the scene. Incremental reconstruction can effectively reduce the computational complexity while ensuring the reconstruction accuracy by gradually adding images and optimizing the image imaging parameters and the positions of 3D points.
[0091] Step 3: The dense reconstruction method based on Vis-MVSNet has a process as Figure 3 shown. Traditional methods usually rely on manually designed feature extraction and matching algorithms (such as SGM), and perform poorly when dealing with complex textures, occlusions and repetitive structures, resulting in insufficient matching accuracy; secondly, traditional methods have low robustness to light changes and view differences, and are prone to false matches or missing matches; in addition, the computational complexity of traditional methods is relatively high, especially when dealing with large-scale data, the efficiency drops significantly; finally, traditional methods usually lack the ability of global optimization and are difficult to effectively handle the noise and inconsistencies in the scene, resulting in limited quality and integrity of the reconstruction results.
[0092] In contrast, Vis-MVSNet can better overcome these limitations and achieve high-precision and high-efficiency dense reconstruction through deep learning and adaptive strategies. Specifically, multi-scale features are extracted through the Vis-MVSNet network and a cost volume is constructed. The visual attention mechanism and adaptive view selection strategy are used to generate high-precision depth maps, and then multi-view depth maps are fused to generate a dense point cloud. Then, based on the point cloud data, the Delaunay triangulation algorithm is used to construct a TIN to accurately represent the geometric structure of the transmission tower scene. Then, mesh optimization is performed on the TIN, including removing noise points, repairing holes and self-intersecting surfaces, simplifying redundant triangles or subdividing important regions to improve mesh quality and computational efficiency. Finally, the two-dimensional image is used as a texture and attached to the surface of the optimized three-dimensional transmission tower model through texture mapping. The correspondence between the texture and the model is established using UV coordinates, and seams and distortions are eliminated by combining lighting and shadow calculations to generate a three-dimensional transmission tower model with high realism and detailed performance.
[0093] Step 4: After a series of post-processing steps such as coordinate transformation and format conversion, a high-precision and high-fidelity three-dimensional transmission tower model is finally output. First, the model is converted from the local coordinate system to the global coordinate system (such as the geographic coordinate system) through coordinate transformation to ensure the accurate spatial position, orientation, and scale of the model in the actual scene, laying a foundation for subsequent spatial analysis and applications. Then, format conversion is performed on the model to convert it into a widely supported three-dimensional file format (such as OBJ, FBX, PLY, or STL) to ensure that the model is compatible with a variety of three-dimensional software, platforms, and hardware devices to meet the needs of different application scenarios. In addition, detail optimization is performed on the model, including removing redundant data to reduce the file size, repairing geometric defects (such as holes, self-intersecting surfaces, or non-manifold structures) to improve the model integrity, and optimizing texture mapping to enhance the visual effect to ensure that the model meets high-quality standards in both geometric accuracy and visual performance. The finally generated three-dimensional transmission tower model can be directly used for engineering analysis, visualization display, virtual simulation, or further three-dimensional applications (such as digital twins, UAV inspection path planning, etc.), providing comprehensive and reliable data support for the detection, maintenance, planning, and management of transmission towers, and contributing to the efficient operation and intelligent development of the power system.
[0094] Figure 4 FIG. is a schematic structural diagram of a three-dimensional transmission tower reconstruction system 400 based on Vis-MVSNet according to an embodiment of the present invention. As Figure 4 shown, the three-dimensional transmission tower reconstruction system 400 based on Vis-MVSNet provided by the embodiment of the present invention includes: an image collection unit 401, a sparse reconstruction unit 402, a dense reconstruction unit 403, and a conversion unit 404.
[0095] Preferably, the image collection unit 401 is configured to collect the original sequence images of the UAV covering multiple perspectives of the transmission tower, and preprocess the original data images to obtain processed sequence images.
[0096] Preferably, the image collection unit 401 preprocesses the original data images to obtain processed sequence images, including:
[0097] Read the GPS-aided information and camera parameter information of the images in the original sequence images, and perform wallis filtering on all images to obtain processed sequence images with enhanced image quality.
[0098] Preferably, the sparse reconstruction unit 402 is configured to perform incremental reconstruction based on the processed sequence images by using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters.
[0099] Preferably, the sparse reconstruction unit 402 performs incremental reconstruction based on the processed sequence images by using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters, including:
[0100] Use the SuperPoint deep network to extract the feature point positions and feature vectors from the processed sequence images;
[0101] Use the AdaLAM algorithm to perform feature matching on the feature points extracted from any two images to determine the image pairs;
[0102] Select the initial image pairs and gradually perform scene reconstruction by using the incremental SfM method to obtain a sparse point cloud model and image imaging parameters.
[0103] Preferably, the dense reconstruction unit 403 is configured to perform dense reconstruction based on the sparse point cloud model and image imaging parameters by using a dense reconstruction method based on Vis-MVSNet to obtain an initial 3D model of the transmission tower.
[0104] Preferably, the dense reconstruction unit 403 performs dense reconstruction based on the sparse point cloud model and image imaging parameters by using a dense reconstruction method based on Vis-MVSNet to obtain an initial 3D model of the transmission tower, including:
[0105] Based on the sparse point cloud model and image imaging parameters, use the Vis-MVSNet dense matching method to perform matching to obtain a dense three-dimensional point cloud;
[0106] Taking the dense three-dimensional point cloud as the input, connecting the points in the point cloud into a TIN grid through the Delaunay triangulation algorithm, performing TIN construction, and obtaining a TIN structure;
[0107] Performing mesh optimization on the TIN structure to obtain a three-dimensional space structure mesh;
[0108] Attaching the sequence images as textures to the surface of the three-dimensional space structure mesh to obtain an initial three-dimensional model of the transmission tower.
[0109] Preferably, the conversion unit 404 is configured to perform conversion processing on the initial three-dimensional model of the transmission tower to obtain a final three-dimensional model of the transmission tower.
[0110] Preferably, the conversion unit 404 performs conversion processing on the initial three-dimensional model of the transmission tower to obtain a final three-dimensional model of the transmission tower, including:
[0111] Converting the local coordinate system of the initial three-dimensional model of the transmission tower to the global coordinate system and then converting it into a three-dimensional file format to obtain a final three-dimensional model of the transmission tower.
[0112] The transmission tower three-dimensional reconstruction system 400 based on Vis-MVSNet in the embodiment of the present invention corresponds to the transmission tower three-dimensional reconstruction method 100 in another embodiment of the present invention, and will not be elaborated here.
[0113] The present invention has been described by referring to a few embodiments. However, as is well known to those skilled in the art, other embodiments equivalent to those disclosed above of the present invention equally fall within the scope of the present invention.
[0114] Generally, all terms used in the present invention are interpreted according to their ordinary meanings in the technical field, unless otherwise clearly defined therein. All references to "a / the [device, component, etc.]" are open to interpretation as at least one instance of the device, component, etc., unless otherwise clearly stated. The steps of any method disclosed here do not necessarily have to be run in the exact order disclosed, unless clearly stated.
[0115] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the present invention.
Claims
1. A three-dimensional reconstruction method of a transmission tower based on Vis-MVSNet, characterized in that The method includes: Collecting the original sequence images of the UAV covering multiple perspectives of the transmission tower, and preprocessing the original data images to obtain processed sequence images; Based on the processed sequence images, using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM for incremental reconstruction to obtain a sparse point cloud model and image imaging parameters; Based on the sparse point cloud model and image imaging parameters, using a dense reconstruction method based on Vis-MVSNet for dense reconstruction to obtain a three-dimensional initial model of the transmission tower; Performing a transformation process on the three-dimensional initial model of the transmission tower to obtain a three-dimensional final model of the transmission tower.
2. The method according to claim 1, wherein The preprocessing of the original data images to obtain processed sequence images includes: Reading the GPS auxiliary information and camera parameter information of the images in the original sequence images, and performing wallis filtering on all images to obtain processed sequence images with enhanced image quality.
3. The method according to claim 1, characterized in that The incremental reconstruction based on the processed sequence images using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters includes: Using the SuperPoint deep network to extract the feature point positions and feature vectors from the processed sequence images; Using the AdaLAM algorithm to perform feature matching on the feature points extracted from any two images to determine the image pairs; Selecting the initial image pairs and gradually performing scene reconstruction using the incremental SfM method to obtain a sparse point cloud model and image imaging parameters.
4. The method according to claim 1, characterized in that The dense reconstruction based on the sparse point cloud model and image imaging parameters using a dense reconstruction method based on Vis-MVSNet to obtain a three-dimensional initial model of the transmission tower includes: Based on the sparse point cloud model and image imaging parameters, using the Vis-MVSNet dense matching method for matching to obtain a dense three-dimensional point cloud; Taking the dense three-dimensional point cloud as the input, connecting the points in the point cloud into a TIN grid through the Delaunay triangulation algorithm to perform TIN construction and obtain a TIN structure; Performing mesh optimization on the TIN structure to obtain a three-dimensional spatial structure mesh; Attaching the sequence images as textures to the surface of the three-dimensional spatial structure mesh to obtain a three-dimensional initial model of the transmission tower.
5. The method according to claim 1, wherein The transformation process on the three-dimensional initial model of the transmission tower to obtain a three-dimensional final model of the transmission tower includes: Converting the local coordinate system of the three-dimensional initial model of the transmission tower to the global coordinate system and then converting it to a three-dimensional file format to obtain a three-dimensional final model of the transmission tower.
6. A three-dimensional reconstruction system of a transmission tower based on Vis-MVSNet, characterized in that, The system includes: An image collection unit for collecting the original sequence images of the UAV covering multiple perspectives of the transmission tower, and preprocessing the original data images to obtain processed sequence images; A sparse reconstruction unit for performing incremental reconstruction based on the processed sequence images using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters; A dense reconstruction unit for performing dense reconstruction based on the sparse point cloud model and image imaging parameters using a dense reconstruction method based on Vis-MVSNet to obtain an initial 3D model of the transmission tower; A conversion unit for performing conversion processing on the initial 3D model of the transmission tower to obtain a final 3D model of the transmission tower.
7. The system according to claim 6, wherein The image collection unit preprocesses the original data image to obtain a processed sequence of images, including: Reading the GPS-aided information and camera parameter information of the images in the original sequence of images, and performing wallis filtering on all images to obtain a processed sequence of images with enhanced image quality.
8. The system according to claim 6, wherein The sparse reconstruction unit performs incremental reconstruction based on the processed sequence of images using an incremental SfM sparse reconstruction method combining SuperPoint and AdaLAM to obtain a sparse point cloud model and image imaging parameters, including: Using the SuperPoint deep network to extract the feature point positions and feature vectors from the processed sequence of images; Using the AdaLAM algorithm to perform feature matching on the feature points extracted from any two images to determine the image pairs; Selecting an initial image pair and gradually performing scene reconstruction using the incremental SfM method to obtain a sparse point cloud model and image imaging parameters.
9. The system according to claim 6, wherein The dense reconstruction unit performs dense reconstruction based on the sparse point cloud model and image imaging parameters using a dense reconstruction method based on Vis-MVSNet to obtain an initial 3D model of the transmission tower, including: Based on the sparse point cloud model and image imaging parameters, using the Vis-MVSNet dense matching method for matching to obtain a dense 3D point cloud; Using the Delaunay triangulation algorithm to connect the points in the point cloud into a TIN grid with the dense 3D point cloud as the input, performing TIN construction to obtain a TIN structure; Performing mesh optimization on the TIN structure to obtain a 3D spatial structure mesh; Attaching the sequence of images as textures to the surface of the 3D spatial structure mesh to obtain an initial 3D model of the transmission tower.
10. The system according to claim 6, characterized in that, The conversion unit performs conversion processing on the initial 3D model of the transmission tower to obtain a final 3D model of the transmission tower, including: Converting the local coordinate system of the initial 3D model of the transmission tower to the global coordinate system and then converting it to a 3D file format to obtain a final 3D model of the transmission tower.