A method and system for single-view point cloud reconstruction of tree branches in real scenes

By employing a hierarchical cyclic prediction and densification strategy, combined with a cyclic pixel transformation network and point cloud correction, the problem of reconstructing occluded areas of tree branches and trunks was solved, achieving high-precision point cloud reconstruction of tree branches and trunks in real-world scenarios, and supporting more accurate tree morphology analysis and management decisions.

CN119600230BActive Publication Date: 2025-12-05BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411658003.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-12-05
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

In real-world tree branch reconstruction, existing single-view-based point cloud reconstruction methods cannot effectively estimate the depth information of occluded branches due to the occlusion problem of leaves and branches, resulting in poor reconstruction results.

Method used

A hierarchical cyclic prediction strategy for depth information is adopted. By using a cyclic pixel-to-pixel transformation network (R-pix2pix) and a point cloud densification algorithm based on normal and uniform distribution (NUD), combined with a hierarchical storage strategy for 3D point cloud information and a point cloud correction network, the depth information of tree branches and trunks is predicted and densified layer by layer to correct errors and achieve complete reconstruction.

Benefits of technology

It effectively solves the problem of reconstructing parts of tree branches that are obscured, improves reconstruction accuracy and detail, and is superior to traditional single-view and multi-view methods, supporting more accurate tree morphology analysis and intelligent pruning management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600230B_ABST
    Figure CN119600230B_ABST
Patent Text Reader

Abstract

The application discloses a kind of real scene under tree branch single view point cloud reconstruction method, comprising: obtaining real scene under tree branch single RGB image;After single view depth estimation based on cycle prediction is carried out to tree branch single RGB image, obtain multiple depth images;Multiple depth images are converted into point cloud, so as to obtain sparse point cloud;Sparse point cloud is densified to obtain dense point cloud;The fine modification of point cloud coordinates is carried out to dense point cloud to obtain reconstruction result;Real scene under tree branch point cloud is reconstructed based on reconstruction result.The application also discloses corresponding system, electronic equipment and computer readable storage medium, and the problem that the depth information of the occluded branch cannot be effectively estimated is solved based on single view depth estimation based on cycle prediction;The problem of sparse and inaccurate tree branch point cloud obtained initially is solved based on the point cloud densification algorithm based on tree form feature analysis and the point cloud correction method based on point-by-point correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target point cloud reconstruction technology, and in particular to a method and system for reconstructing single-view point clouds of tree branches in a real-world scene. Background Technology

[0002] In recent years, with the rise of digital orchards, the demand for efficient, automated, and digitalized orchard management has been increasing. Among these advancements, tree trunk and branch reconstruction technology has emerged as a crucial tool in digital orchard management. This technology aims to meet the needs of orchard management for tree morphology analysis, growth monitoring, and intelligent pruning, providing orchard managers with more data support. The continuous development of deep learning and computer vision has provided strong technical support for tree trunk and branch reconstruction. Using these technologies, researchers can reconstruct the three-dimensional structure of tree trunks and branches from collected data.

[0003] Currently, some scholars have begun research on tree branch and trunk reconstruction technology. However, in practical scenarios, due to the mutual shading between leaves, branches, and other organs, as well as between branches themselves, some tree branches and trunks become obscured and cannot be observed. This shading problem makes tree branch and trunk reconstruction more complex and difficult. Solving the shading problem is of great significance for the complete reconstruction of tree structure. It helps to achieve more comprehensive tree morphological analysis, allowing orchard managers to better understand the growth status of trees. It also provides strong support for management decisions such as tree growth monitoring and intelligent pruning, promoting the development of orchard management towards a more efficient and intelligent direction.

[0004] The problem of reconstructing occluded tree branches can be categorized into three types based on the raw data used: reconstruction based on LiDAR point clouds, reconstruction based on RGB and depth images, and reconstruction based on RGB images. Among these, RGB image-based reconstruction methods have significant advantages in data acquisition and processing compared to other methods. Specifically, RGB image-based reconstruction methods do not require image data from multiple perspectives, reducing the complexity and cost of data acquisition. Therefore, this invention focuses on in-depth research into tree branch reconstruction methods based on RGB images. LiDAR point cloud-based reconstruction is a method for converting incomplete point clouds to refined point clouds, which is less relevant to this invention; therefore, this invention provides progress in the latter two research areas.

[0005] Reconstruction studies based on RGB and depth images typically use RGBD cameras to acquire data, simultaneously obtaining RGB and depth images. However, due to the complexity of real-world orchard scenes, the acquired depth images are often inaccurate. Therefore, image processing and deep learning techniques are needed to address the inaccuracy of depth information caused by occlusion. For example, Geckeler C et al. used an RGBD camera to acquire RGB and depth images of trees and found that occlusion caused depth discontinuities in the acquired tree depth images. They proposed a U-Net-based network model to optimize the depth images of occluded branches. Experimental results showed that this method effectively improved the reconstruction of occluded areas. Kok E et al. proposed a Unet++-based model to address the problem of branch occlusion in natural orchards. This method can recover the 3D structure of occluded trees from RGB and depth images captured by an RGBD camera. Experimental results showed that the proposed method can effectively reconstruct the spatial structure of occluded branches. Kim CH et al. addressed the problem that pruning robots equipped with RGBD cameras cannot accurately perceive the tree skeleton hidden inside the canopy, thus hindering effective pruning. They proposed a method to estimate the tree skeleton hidden inside the canopy from RGB and depth images. The authors conducted experiments on simulated and real data, demonstrating that the proposed method can effectively estimate the structure of the occluded tree skeleton. Lin G et al. addressed the problem of poor reconstruction results due to depth discontinuities in depth images acquired by RGBD cameras, proposing a 3D reconstruction method. This method first uses a miniature Mask R-CNN for branch detection, converting the detected branches into 3D point clouds. Then, it uses 3D cylinders to simulate the shape of the branches. Finally, it uses a cylinder segment fitting method based on principal component analysis to reconstruct the branches. Experimental results demonstrate that the proposed method can effectively solve the problem of poor branch reconstruction in occluded areas.

[0006] Research on point cloud reconstruction based on RGB images can be divided into point cloud reconstruction methods based on multiple views and point cloud reconstruction methods based on a single view.

[0007] 1) Multi-view point cloud reconstruction methods: namely, Structure from Motion and Multi-View Stereo (SfM-MVS). Gené-Mola J et al. used SfM-MVS technology to reconstruct apple tree branches in 3D. Their results showed that, compared to other reconstruction methods, SfM-MVS better addressed the poor reconstruction results caused by occlusion. Nader KB et al. proposed a multi-view reconstruction method for grape branches, effectively solving the problem of poor results due to occlusion in grape branch reconstruction by controlling different camera shooting angles. He W et al. reconstructed 3D point clouds of soybean plants using multi-view images acquired by smartphones, thereby enabling the detection and segmentation of soybean point clouds. The authors proposed that using smartphones, as portable miniature devices, allows for flexible adjustment of shooting angles to obtain unobstructed soybean plant reconstructed point clouds. G et al. conducted a comprehensive analysis of the processing parameters, image redundancy, and acquisition geometry in the SfM-MVS technique. The authors proposed that adjusting the processing parameters and image acquisition angle in the SfM-MVS technique can effectively improve the reconstruction accuracy of occluded areas. The main reason why multi-view reconstruction can solve the occlusion problem is that it acquires image data of trees from different perspectives by increasing the number of cameras or shooting positions. When a part is occluded, other perspectives can usually capture information about that part, thus compensating for the missing information. Through image data from multiple perspectives, more comprehensive and complete information about the tree structure can be obtained, thereby achieving higher-precision tree branch and trunk reconstruction. However, this type of research requires high data acquisition costs, necessitating the use of multiple cameras or devices for image acquisition, and the data processing and computation are relatively cumbersome. Furthermore, this method places high demands on algorithms and computational power, resulting in a relatively slow reconstruction process.

[0008] 2) Single-view point cloud reconstruction methods: Single-view point cloud reconstruction methods use a single RGB image to reconstruct tree branches and trunks. Typically, a depth estimation network is trained to learn the mapping from RGB images to depth images, and then the depth image is converted into a 3D point cloud to complete the 3D reconstruction. Guénard J et al. proposed a single-view tree reconstruction method. This method takes a single background-free image of a tree as input and, combined with prior knowledge of tree morphology, first roughly estimates the tree's skeleton. Then, by changing the parameter values ​​of the generative model (the main branch structure and leaves of the tree), the 3D structure of the tree is obtained. Experimental results show that the proposed single-view tree reconstruction method can effectively reconstruct the 3D structure of trees. Liu Z et al. used Conditional Generative Adversarial Nets (cGANs) to extract the edge information of trees from images and, in addition, combined with 2D skeleton strokes drawn manually by the user to infer the 3D shape of the tree. Experimental results show that the proposed method has a certain effectiveness in reconstructing 3D tree models from a single image. The innovation of this method lies in its pioneering use of a single image to store the depth information of both the front and back layers of the tree canopy, thus enabling the reconstruction of the envelope shape of the occluded area on the back of the canopy. Single-view point cloud reconstruction research requires only one RGB image as input data, making data acquisition and processing simpler and less costly compared to RGBD camera reconstruction and multi-view reconstruction techniques. Currently, there is relatively little research on tree branch reconstruction using single-view point cloud reconstruction techniques, and solving the occlusion problem remains a significant challenge in the field of point cloud reconstruction.

[0009] In summary, single-view tree branch point cloud reconstruction methods have significant advantages over multi-view point cloud reconstruction methods in terms of data acquisition and processing. The single-view method eliminates the need for image data from multiple perspectives, reducing the complexity and cost of data acquisition. Furthermore, it simplifies the data processing flow and improves processing efficiency. Given these advantages, this invention focuses on research into single-view tree branch reconstruction methods. In tree branch reconstruction technology, single-view methods require only one RGB image as input data, thus offering a simpler data acquisition and processing process compared to multi-view methods. However, current research on single-view tree branch reconstruction primarily focuses on point cloud reconstruction of tree branches in virtual simulation scenarios. This research typically operates in relatively simple scenarios without background interference, failing to adequately consider the complexity of tree growth and occlusion interference in real-world scenarios. In actual scenarios, the mutual occlusion between branches increases the complexity and difficulty of reconstruction. Addressing occlusion issues is crucial for the complete reconstruction of tree branch and trunk structures. This not only facilitates more accurate tree morphology analysis but also supports management decisions such as tree growth monitoring and intelligent pruning. Therefore, researching novel point cloud reconstruction methods to achieve complete reconstruction of tree branches and trunks in real-world scenarios is particularly important. Specifically, given the complex growth states of trees and the presence of occlusion interference, new point cloud reconstruction methods are needed to achieve complete reconstruction of tree branches and trunks in realistic scenarios. Summary of the Invention

[0010] To address the problems existing in the prior art, this paper provides a method and system for reconstructing single-view point clouds of tree branches in real-world scenarios. In this method and system, firstly, a single-view depth estimation method based on cyclic prediction is proposed, which solves the problem that traditional point cloud reconstruction methods cannot effectively estimate the depth information of occluded branches. Secondly, a point cloud density algorithm based on tree morphological feature analysis and a point cloud correction method based on point-by-point correction are proposed, which solve the problems of sparse and inaccurate tree branch point clouds initially obtained in the single-view depth estimation method.

[0011] Tree branch reconstruction can typically be achieved through two methods: single-view point cloud reconstruction and multi-view point cloud reconstruction. Multi-view methods offer more comprehensive and accurate reconstruction results; however, they require higher data acquisition costs and increase the complexity of data processing and computation. Therefore, this invention chooses single-view point cloud reconstruction to reconstruct tree branches. The idea behind single-view point cloud reconstruction is to predict the corresponding depth image from a single RGB image. However, due to viewpoint limitations, only one side of the surface of the same branch can be reconstructed, and some branches may be occluded by other branches. The root cause of this occlusion problem is that a single depth image cannot represent the three-dimensional structural information of the entire tree branch. In this case, the depth image predicted by the depth estimation network can only represent the visible side depth information of the corresponding position in the RGB image, and cannot represent the depth information of the occluded side. To address this problem, this invention considers using multiple depth images to store depth information hierarchically, and then using a cyclic prediction approach to recover the three-dimensional structural information of the occluded portion. Specifically, this invention constructs a depth prediction network that uses a known RGB image and the previous depth image to predict the next depth image, iterating until all depth images are predicted. This hierarchical storage and iterative depth prediction approach effectively solves the problem of poor reconstruction results caused by occlusion.

[0012] The embodiments of this invention employ a hierarchical, cyclic prediction strategy for depth information, effectively addressing the problem of poor prediction performance for occluded portions of tree branches. Secondly, a recurrent pixel-to-pixel network (R-pix2pix) is proposed for cyclically predicting the depth information of tree branches. Next, a Normal and Uniformly Distributed Point Cloud Dense Algorithm (NUD) is proposed, which can achieve a denser effect on sparse tree point clouds. Finally, a point cloud correction network is proposed, which can correct errors caused in the tree branch point cloud reconstruction process.

[0013] This invention proposes a method for reconstructing single-view point clouds of tree branches in real-world scenarios, comprising:

[0014] S1, acquire a single RGB image of tree branches and trunks in a real scene;

[0015] S2, after performing single-view depth estimation based on cyclic prediction on a single RGB image of the tree branches, multiple depth images are obtained;

[0016] S3, convert the multiple depth images into point clouds to obtain sparse point clouds;

[0017] S4, the sparse point cloud is densified to obtain a dense point cloud;

[0018] S5, refine the point cloud coordinates of the dense point cloud to obtain the reconstruction result;

[0019] S6. Based on the reconstruction results, reconstruct a single-view point cloud of tree branches in the real scene.

[0020] Preferably, S2 includes:

[0021] S21, Based on the point cloud 3D information layered storage strategy, the point cloud is divided into grids, and the centroid information in the grids is stored in multiple 2D images in layers;

[0022] S22, a cyclic pixel-to-pixel conversion network is established. The cyclic pixel-to-pixel conversion network is used to implement cyclic prediction. When predicting each layer of depth image, the previous layer of depth image is used as a conditional input, thereby generating a depth image sequence layer by layer. Each neuron in the cyclic pixel-to-pixel conversion network consists of a generator and a discriminator. Each neuron adopts the basic structure of a conditional generative adversarial network, and the generator and discriminator compete with each other and co-evolve.

[0023] S23, input a single RGB image of the tree branches into the cyclic pixel-to-pixel conversion network to cyclically predict the depth information of each layer of the tree branches.

[0024] Preferably, in step S21, a pair of N mapping relationships is set, that is, one RGB image corresponds to N depth images; during network training and testing, the next depth image is predicted using the known RGB image and the previous depth image, iterating until all depth images are predicted, including the following steps:

[0025] Step a, divide the 3D tree into meshes along the depth axis; set the depth axis as the y-axis, set the mesh to M, and the size of each mesh is max(depth) / M, where max(depth) is the maximum value of the coordinates of the points on the y-axis;

[0026] Step b: Calculate the relative centroid position of the points in each grid, and calculate the average relative centroid of all points in each grid by statistically analyzing the point cloud data in each grid, thereby obtaining the representative position of the points in each grid;

[0027] Step c: Encode the average relative centroid position information of M grids into the RGB channel value of each pixel in the depth image; convert the point cloud data in three-dimensional space into pixel values ​​of two-dimensional images, and obtain a total of N depth images.

[0028] Preferably, the specific structure of the cyclic pixel-to-pixel conversion network is as follows:

[0029] (1) Input, output, and training process of the recurrent pixel-to-pixel conversion network:

[0030] The cyclic pixel-to-pixel conversion network receives a single RGB image X RGB As input, it outputs the depth image of the next layer; during training, the depth image of the first layer is predicted through the pix2pix_0 network. Then, X RGB and As conditional input, the pix2pix_1 network is used to predict the depth image of the second layer. This process is repeated iteratively until the depth image of the (N-1)th layer is predicted. This continues until the depth images of all layers are predicted;

[0031] (2) The recurrent pixel-to-pixel transformation network consists of a generator and a discriminator in each neuron. The generator uses a U-Net structure to convert the input RGB image into the corresponding depth image. The discriminator is used to determine the authenticity of the generated image. Through a series of convolutional and pooling layers, it extracts features from the input image. During downsampling, the size of the feature matrix gradually decreases while the number of channels gradually increases, ultimately resulting in a 1*1 feature vector. This feature vector is converted into a score representing the authenticity of the image using the sigmoid function, which is used to optimize the adversarial loss between the generator and the discriminator. The Huber loss function is introduced into the generator.

[0032] Preferably, step S4 includes: based on tree morphological feature analysis, denserizing the sparse point cloud to obtain a dense point cloud, including:

[0033] S41, Establish the NUD algorithm. The NUD algorithm classifies sparse point clouds according to the morphological characteristics of tree branches and trunks. For different categories of sparse point clouds of tree branches and trunks, corresponding densification strategies and densification parameters are used. The densification strategies include normal distribution and uniform distribution. The densification parameters include the number of points in the normal distribution, the mean and the variance, the number of points in the uniform distribution, and one or more of the upper and lower bounds of the distribution.

[0034] S42, based on the NUD algorithm, the sparse point cloud is densified to obtain a dense point cloud.

[0035] Preferably, S42 includes:

[0036] (1) The NUD algorithm classifies sparse point clouds based on the morphological characteristics of tree branches. Specifically, when traversing to a given point, the algorithm determines whether there are other points in the grids before and after that point. Based on the presence of points in the grids before and after that point, the sparse point cloud can be divided into 9 categories, including:

[0037] Category ①: No points exist in either the front or back grid.

[0038] Category ②: A point exists in the previous grid and is located in the first half of the grid, but the point does not exist in the next grid;

[0039] Category ③: A point exists in the previous grid and is located in the middle of the grid, but the point does not exist in the next grid;

[0040] Category 4: There are no points in the previous grid, but there are points in the next grid, and the points are located in the first half of the grid.

[0041] Category 5: A point exists in the previous grid but in the next grid, and its position is close to the second half of the grid it belongs to;

[0042] Category 6: Points exist in both the preceding and following grids, and the points in the preceding grid are located in the first half of their respective grids, while the points in the following grid are located in the first half of their respective grids.

[0043] Category 7: Points exist in both the preceding and following grids, with the points in the preceding grid located near the front half of their respective grid, and the points in the following grid located near the back half of their respective grid.

[0044] Category 8: Points exist in both the preceding and following grids, with the points in the preceding grid located near the latter half of the grid, and the points in the following grid located near the former half of the grid.

[0045] Category 9: Points exist in both the preceding and following grids, and the position of the point in the preceding grid is close to the last 1 / 2 of its own grid, while the position of the point in the following grid is close to the last 1 / 2 of its own grid.

[0046] (2) Based on the classified sparse point cloud, different densification strategies are adopted, including:

[0047] For categories ①, ②, ⑤ and ⑦, a normal distribution densification method is adopted, wherein the mean of the normal distribution is the centroid value of the points in the current grid, while the variance and the number of generated points are dynamically adjusted according to the altitude of the point cloud;

[0048] For category ③, a uniform distribution densification method is adopted. The lower bound of uniform distribution is the centroid value of the point in the previous grid, and the upper bound is the centroid value of the point in the current grid. The number of generated points is also dynamically adjusted according to the altitude of the point cloud.

[0049] For category ④, a uniform distribution densification method is adopted, but the lower and upper bounds of the uniform distribution are adjusted to the centroid values ​​of the points in the current grid and the centroid values ​​of the points in the next grid, respectively.

[0050] For category ⑥, the points in the current grid are densed using a uniform distribution method. The lower bound is the centroid value of the points in the current grid, and the upper bound is the centroid value of the points in the next grid. The number of points generated is still dynamically adjusted according to the altitude of the point cloud.

[0051] For category ⑧, a uniformly distributed densification method is adopted, with the lower bound being the starting position value of the current grid and the upper bound being the ending position value of the current grid. The number of generated points is still dynamically adjusted according to the altitude of the point cloud.

[0052] For category ⑨, a normal distribution density method is used. The mean of the normal distribution is taken from the centroid value of the points in the current grid, while the variance and the number of generated points are dynamically adjusted according to the elevation of the point cloud.

[0053] Preferably, S5 includes:

[0054] S51, establish a point cloud correction network. The point cloud correction network consists of convolutional layers, BN layers, and fully connected layers. It takes the coordinates of the original coarse point cloud as input, learns and adjusts the coordinate positions, and outputs the corrected coordinates at the corresponding positions. This allows for more refined correction of the existing point cloud coordinates, thereby reducing accumulated errors and improving the overall reconstruction accuracy.

[0055] S52, input the coordinates of the original coarse point cloud into the point cloud correction network, and output the corrected coordinates of the corresponding point at the corresponding position.

[0056] A second aspect of the present invention provides a single-view point cloud reconstruction system for tree branches in a real-world scene, for implementing the method of the first aspect, comprising:

[0057] The image acquisition module acquires a single RGB image of tree branches in a real-world scene.

[0058] The single-view depth estimation module based on cyclic prediction is used to obtain multiple depth images by performing single-view depth estimation based on cyclic prediction on a single RGB image of the tree branches.

[0059] The sparse point cloud conversion module is used to convert the multiple depth images into point clouds to obtain sparse point clouds.

[0060] A point cloud densification module based on tree morphology feature analysis is used to densify the sparse point cloud to obtain a dense point cloud.

[0061] The point cloud correction module based on point-by-point correction is used to refine the point cloud coordinates of the dense point cloud to obtain the reconstruction result.

[0062] The point cloud reconstruction module is used to reconstruct a single-view point cloud of tree branches and trunks in a real scene based on the reconstruction results.

[0063] The present invention has the following beneficial technical effects:

[0064] To verify the effectiveness and practicality of the single-view point cloud reconstruction method proposed in this invention, it was compared with traditional single-view point cloud reconstruction methods and multi-view point cloud reconstruction methods. Experimental results show that the single-view point cloud reconstruction method proposed in this invention exhibits superior performance, outperforming traditional methods in reconstruction quality and effectively solving the problem of ineffective reconstruction of occluded areas in traditional methods. Compared with multi-view reconstruction methods, the single-view point cloud reconstruction method proposed in this invention demonstrates significant advantages in the reconstruction of branch details. The single-view point cloud reconstruction technology proposed in this invention brings innovative ideas and practical methods to the research in the field of tree branch reconstruction, specifically:

[0065] (1) A strategy of hierarchical cyclic prediction of depth information was adopted, which effectively solved the problem of poor prediction effect of tree branches and trunks blocking the depth information.

[0066] (2) The proposed Recurrent Pixel-to-Pixel Network (R-pix2pix) is used to cyclically predict the depth information of tree branches, aiming to solve the problem that traditional depth estimation networks cannot effectively estimate the depth information of occluded areas. This network combines a hierarchical storage strategy for point cloud 3D information. First, the 3D information of tree branches is stored hierarchically in multiple 2D images. Then, the improved pixel-to-pixel network cyclically predicts the depth information layer by layer. Through this layer-by-layer prediction process, the network can more accurately capture the structural information of occluded areas, improving the accuracy of depth estimation.

[0067] (3) The Normal and Uniformly Distributed Point Cloud Dense Algorithm (NUD) can achieve the densification effect of sparse tree point clouds, aiming to solve the problem of sparse tree branch point clouds synthesized after the depth image is predicted by the single-view depth estimation method. This algorithm analyzes the morphological features of tree branches and trunks and uses the normal and uniform distribution generation method to densify the sparse point cloud.

[0068] (4) The point cloud correction network can correct errors caused in the single-view point cloud reconstruction process of tree branches, aiming to correct the errors caused by the R-pix2pix network and NUD algorithm. This invention uses a hierarchical cyclic prediction method for depth prediction, which causes errors to gradually accumulate in the network layers, resulting in inaccurate final depth estimation. Furthermore, during point cloud densification, the complex structure of tree branches leads to insufficient detail in the resulting tree branches. To reduce errors caused by depth estimation and densification, this invention proposes a point cloud correction network. This network takes the coordinates of the point cloud as input and outputs the corrected coordinates at the corresponding locations. The correction network's role is to refine the existing point cloud coordinates to reduce accumulated errors and improve the overall reconstruction accuracy. Attached Figure Description

[0069] Figure 1 This is a flowchart of the method for reconstructing tree branches and trunks from a single view in a real-world scenario, as described in this invention.

[0070] Figure 2 This is a schematic diagram illustrating the principle of the single-view point cloud reconstruction method for tree branches in a real-world scenario as described in this invention. The three gray dashed boxes represent one of the innovative aspects of this invention. In the single-view depth estimation module based on cyclic prediction, this invention proposes a hierarchical storage strategy for 3D point cloud information and an R-pix2pix network to address the problem of poor reconstruction of occluded parts in the single-view point cloud reconstruction of tree branches. In the point cloud densification module based on tree morphological feature analysis, the NUD algorithm is proposed to achieve a densification effect for sparse point clouds. In the point cloud optimization module based on point-by-point correction, a point cloud correction network is proposed to achieve fine-grained correction of point cloud coordinates.

[0071] Figure 3 This is a schematic diagram illustrating the layered storage of three-dimensional information of tree branches and trunks as described in this invention, consisting of multiple depth images. Figure 3 The left-middle column shows the overall effect of the 3D point cloud; the upper right dashed box shows the detail of points in the tree branch point cloud along the y-axis; and the lower right dashed box shows the final 8 depth images. These represent the eight depth images obtained from the layering process;

[0072] Figure 4 This is a diagram of the R-pix2pix network architecture described in this invention. Figure 4 ZhongX RGB This represents the RGB image of tree branches input to the R-pix2pix network. Let G(x) represent the depth image at layer t, where t ∈ {0, 1, ..., 7}, and G(x) represent the depth map generated by the generator. The upper dashed box in the figure represents the overall architecture of the R-pix2pix network, which predicts 8 depth maps layer by layer through 8 neurons composed of pix2pix units. The lower dashed box represents the specific architecture of each neuron. The upper area in the figure is the innovation of the R-pix2pix network. This embodiment adopts the idea of ​​cyclic prediction to predict depth images layer by layer, and also introduces the Huber loss function, which significantly improves the prediction performance of the network.

[0073] Figure 5 This is a diagram of the point cloud correction network architecture described in this invention;

[0074] Figure 6 This is a system architecture diagram of a single-view point cloud reconstruction system for tree branches in a real-world scenario, as described in this invention.

[0075] Figure 7 A visual comparison diagram of the depth image prediction performance under MAE, MSE, and the Huber loss function introduced in this invention;

[0076] Figure 8 A schematic diagram showing the quantitative comparison results of eight pix2pix networks under different loss functions (MAE, MSE, and the Huber loss function introduced in this invention) in terms of PSNR and SSIM metrics;

[0077] Figure 9 This is a schematic diagram showing a visual comparison between the method of this invention and traditional single-view point cloud reconstruction in a real-world scenario;

[0078] Figure 10 An example of an RGB image of tree branches taken from a top-down view using a camera mounted on a drone; the area indicated by the arrow in the image is a pole that stretches out the branches, obscuring the branches below.

[0079] Figure 11 This is a visual comparison of the reconstruction details of the occluded region between our method and traditional single-view point cloud reconstruction.

[0080] Figure 12 A visualization showing the comparison between the method of this invention and the multi-view reconstruction method in a real-world scenario;

[0081] Figure 13 A comparison view showing the results of multi-view reconstruction using different numbers of images and the single-view point cloud reconstruction technique proposed in this invention.

[0082] Figure 14 This is a schematic diagram of the electronic device structure described in this invention. Detailed Implementation

[0083] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0084] Example 1

[0085] like Figure 1-2 As shown, this embodiment provides a method for reconstructing single-view point clouds of tree branches in a real-world scene, including:

[0086] S1, acquire a single RGB image of tree branches and trunks in a real scene;

[0087] In single-view point cloud reconstruction methods, the depth estimation network plays a crucial role, responsible for establishing the mapping relationship between RGB images and depth images. This step directly determines the reconstruction quality. However, due to the limitations of single-view point cloud reconstruction, a single RGB image can often only predict a corresponding single depth image. The limitation of a single depth image is that it cannot accurately represent the depth information of occluded areas in tree branches, thus failing to effectively reconstruct the 3D structure of these occluded areas. This becomes an unavoidable problem when using single-view point cloud reconstruction methods to reconstruct tree branches. To solve this problem, this invention stores the 3D structural information of tree branches in multiple depth images in a hierarchical manner. Through this strategy, this invention successfully overcomes the limitation that a single depth image cannot fully express the depth information of occluded areas in tree branches. Then, this invention adopts the idea of ​​iterative prediction, using the known RGB image and the previous depth image to predict the next depth image, and iterates until all depth images are predicted. This iterative prediction idea enables this invention to effectively predict the 3D structure of occluded areas of tree branches.

[0088] S2, after performing single-view depth estimation based on cyclic prediction on a single RGB image of the tree branches, multiple depth images are obtained;

[0089] S3, convert the multiple depth images into point clouds to obtain sparse point clouds;

[0090] S4, the sparse point cloud is densified to obtain a dense point cloud;

[0091] S5, refine the point cloud coordinates of the dense point cloud to obtain the reconstruction result;

[0092] S6. Based on the reconstruction results, reconstruct a single-view point cloud of tree branches in the real scene.

[0093] In a preferred embodiment, S2 includes:

[0094] S21, Based on the point cloud 3D information layered storage strategy, the point cloud is divided into grids, and the centroid information in the grids is stored in multiple 2D images in layers, thereby solving the problem that the 3D information of the occluded area cannot be stored in the 2D image, which leads to the inability to perform depth prediction;

[0095] In traditional single-view depth estimation networks, network training typically employs a one-to-one approach: inputting an RGB image of a tree branch, the network outputs a corresponding depth image of the tree branch from that RGB viewpoint. However, a single depth image cannot represent the 3D structural information of the entire tree branch because 3D information cannot be fully stored in a 2D image. To address this issue, this embodiment proposes a hierarchical storage strategy for point cloud 3D information, storing point cloud information layer by layer across multiple 2D images.

[0096] In this preferred embodiment, a one-to-eight mapping relationship is established, meaning one RGB image corresponds to eight depth images. During network training and testing, the next depth image is predicted using the known RGB image and the previous depth image, iterating until all depth images are predicted. This hierarchical storage strategy effectively solves the problem of poor reconstruction results caused by occlusion.

[0097] Figure 3 This is a schematic diagram illustrating the layered storage of three-dimensional information of tree branches and trunks into eight depth images. Specific step S21 includes the following steps:

[0098] Step a: Mesh the 3D tree along its depth axis. In this embodiment, the depth axis is set to the y-axis, and the grid is set to 24 grids. The size of each grid is max(depth) / 3, where max(depth) is the maximum value of the coordinates of a point on the y-axis.

[0099] Step b: Calculate the relative centroid position of each point in each grid, and calculate the average relative centroid of all points in each grid by statistically analyzing the point cloud data in each grid, thereby obtaining the representative position of each point in each grid.

[0100] This step further simplifies the data while preserving the main features of the point cloud data. In this embodiment, the centroid values ​​in each grid are unified to the range [0, 255] so that they can be stored as pixel values ​​in the image.

[0101] Step c: Encode the average relative centroid position information of the 24 grids into the RGB channel value of each pixel in the depth image.

[0102] This method converts point cloud data in three-dimensional space into pixel values ​​of two-dimensional images, resulting in a total of 8 depth images.

[0103] Through the above steps, this embodiment stores the main three-dimensional information of tree branches and trunks in multiple two-dimensional images in layers, solving the problem that the three-dimensional information of occluded parts cannot be stored in two-dimensional images, thus making depth prediction impossible.

[0104] S22, a recurrent pixel-to-pixel network (R-pix2pix network) is established. This network is used to implement cyclic prediction. When predicting each layer of depth image, the previous layer's depth image is used as a conditional input, thereby generating a sequence of depth images layer by layer. Each neuron in the recurrent pixel-to-pixel network consists of a generator and a discriminator. Each neuron adopts the basic structure of a conditional generative adversarial network, and the generator and discriminator compete with each other and co-evolve.

[0105] In step S21 of this embodiment, the point cloud information of tree branches is stored in eight depth images to solve the problem that a single depth image cannot store depth information of occluded areas in traditional single-view point cloud reconstruction. Through a hierarchical storage strategy for point cloud 3D information, this embodiment obtains a training pair of a single RGB image of tree branches and eight depth images. In traditional single-view point cloud reconstruction schemes, a single RGB image can only predict the corresponding single depth image, failing to complete a one-to-many prediction task. Therefore, this embodiment designs a novel depth estimation network capable of completing a one-to-many prediction task.

[0106] As a preferred embodiment, such as Figure 4 The diagram shows the architecture of a Recurrent Pixel-to-Pixel Network (R-pix2pix). The R-pix2pix network borrows core concepts from Recurrent Neural Networks (RNNs) and Conditional Generative Adversarial Networks (GANs), introducing iterative prediction into depth estimation tasks for multi-layer depth image prediction. When predicting each layer of depth image, the previous layer's depth image is used as a conditional input, thus generating a sequence of depth images layer by layer. This iterative prediction approach effectively improves the prediction accuracy of depth images. Within each neuron, this embodiment of the invention borrows the basic structure of a Conditional Generative Adversarial Network, introducing a mechanism where the generator and discriminator compete and co-evolve, thereby enhancing the prediction capability for depth images.

[0107] As a preferred implementation, the specific structure of the R-pix2pix network is as follows:

[0108] (1) Input, output and training process of R-pix2pix network:

[0109] The R-pix2pix network receives a single RGB image. RGB As input, the next layer's depth image is output.

[0110] During training, the depth image of the first layer is predicted using the pix2pix_0 network. Then, X RGB and As conditional input, the pix2pix_1 network is used to predict the depth image of the second layer. This process is repeated iteratively until the depth image of the seventh layer is predicted. This process continues until the depth images of all layers are predicted.

[0111] (2) Each neuron in the R-pix2pix network consists of a generator and a discriminator.

[0112] The generator uses a U-Net structure to convert the input RGB image into the corresponding depth image, while the discriminator is used to determine the authenticity of the generated image.

[0113] Specifically, in this embodiment, the generator receives two 256*256*3 images as input and performs downsampling and upsampling operations using a U-Net structure. In the downsampling stage, the input images first pass through a series of convolutional and pooling layers, gradually reducing the size of the feature matrix while increasing the number of channels. Specifically, starting from the initial 256*256*3 matrix, it is downsampled sequentially to 128*128*64, 64*64*128, 32*32*256, 16*16*512, 8*8*512, and 4*4*512, until a final feature matrix of size 2*2*512 is obtained. During this process, the network extracts high-dimensional features of the image layer by layer, providing necessary information for subsequent upsampling operations. In the upsampling stage, the generator uses transposed convolutional layers to amplify the feature matrix. First, the 2*2*512 feature matrix is ​​upsampled using transposed convolution, and then concatenated with a feature matrix of the same size from the previous layer. This transposed convolutional and feature concatenation method allows the network to fuse feature information from different dimensions, thereby recovering finer image details. Through layer-by-layer upsampling, the size of the feature matrix gradually increases while the number of channels gradually decreases, ultimately resulting in a 256*256*3 depth image with the same size as the input image.

[0114] The discriminator is used to determine the authenticity of the depth image output by the generator. It receives two 256*256*3 images as input: an RGB image and either the generator's output depth image or a real depth image. The discriminator extracts features from the input image through a series of convolutional and pooling layers. During downsampling, the size of the feature matrix gradually decreases while the number of channels gradually increases, ultimately resulting in a 1*1 feature vector. This feature vector is converted into a score representing the authenticity of the image using the sigmoid function, which is used to optimize the adversarial loss between the generator and the discriminator.

[0115] (3) Loss Function: In this embodiment, the Huber loss function is introduced into the generator, and the specific formula is shown in Equation (1). The generator of the traditional Pix2Pix network uses the MAE (Mean Absolute Error) and MSE (Mean Squared Error) loss functions to achieve the mapping effect from RGB image to depth image. Among them, the MAE loss function has the disadvantage of being sensitive to outliers, which makes it easy to be affected by outliers during training, resulting in the problem of insufficiently refined generated images. The MSE loss function is not sensitive to outliers, but it is prone to gradient explosion, which may lead to unstable network convergence during training. In contrast, the Huber loss function combines the advantages of the MAE and MSE loss functions, has better robustness to outliers, and performs well when the network approaches the optimal solution. It can more effectively balance the network's handling of outliers and accurate values. This invention introduces the Huber loss function to solve the problems caused by the MAE and MSE loss functions of the traditional pix2pix network, making the generation effect of the pix2pix network more refined.

[0116]

[0117] In equation (1), y is the true value, and f(x) is the predicted value output by the network. The hyperparameter δ determines how the network handles outliers. When the residual is greater than δ, the Huber loss uses L1 loss (i.e., linear loss), which is less sensitive to outliers; while when the residual is less than δ, L2 loss (i.e., squared loss) is used for optimization. In this embodiment, the hyperparameter δ is set to 0.5.

[0118] S23, input a single RGB image of the tree branches into the cyclic pixel-to-pixel conversion network to cyclically predict the depth information of each layer of the tree branches.

[0119] By using a recurring prediction process, the network can more accurately capture the structural information of the occluded areas, thereby recovering the three-dimensional structural information of the tree branches.

[0120] In a preferred embodiment, S4 includes: based on tree morphological feature analysis, the sparse point cloud is densified to obtain a dense point cloud.

[0121] Tree branches and trunks, serving as the skeleton and support of a tree, exhibit rich and diverse morphological characteristics. They not only reflect the tree's growth habits and environmental adaptability but also form a crucial foundation for the overall shape of the crown. The main trunk is typically upright and robust, while the central trunk extends upwards from the main trunk to the top of the crown. It is usually more slender than the main trunk, providing support and balance for the various parts of the crown. Main branches are large branches that grow directly from the central trunk, forming the main framework of the crown. Lateral branches are smaller branches that grow on the main branches, further subdividing the crown and forming more complex structures. However, due to potential occlusion of tree branches, the complexity of real-world environments, and the limitations of the data acquisition equipment, the collected tree branch point clouds are often sparse. This sparsity is mainly manifested in a smaller number of points and a lower point cloud density. This results in insufficient detail and richness in the branch morphology within the point cloud data, leading to missing information. Furthermore, sparse branch point clouds often fail to accurately capture the subtle features and details of branches, such as branch forking and bending. This limits the use of point cloud data in applications such as morphological analysis and 3D reconstruction. Furthermore, it results in an incomplete overall spatial structure of trees, especially at branch junctions and crown edges. This incompleteness can hinder an accurate understanding of tree morphology and structure.

[0122] To address the various problems that may arise from the sparsity of point clouds containing tree branches, point cloud densification is a necessary and important task. Point cloud densification aims to increase the density of point cloud data through algorithms and techniques to more accurately reflect the morphology and structure of tree branches. Traditional point cloud densification algorithms mainly include methods based on spatial patch diffusion and methods based on contour constraints. These methods typically rely on certain spatial geometric rules and constraints to densify existing sparse point cloud information through derivation and calculation. However, when dealing with complex scenes, such as point clouds containing tree branches with complex morphological characteristics, these methods cannot fully consider their morphological and structural information, resulting in unsatisfactory densification effects.

[0123] Therefore, this embodiment proposes a point cloud densification algorithm, NUD, based on tree morphological feature analysis, which deeply analyzes the morphological characteristics of tree branches and trunks. By preprocessing and classifying the point cloud data, the algorithm employs different densification methods and parameters for different types of points based on their location and density variations. This classification approach allows the algorithm to better adapt to the complexity and diversity of tree branch and trunk morphology. This helps improve the accuracy and reliability of the densification results, making the densified point cloud data closer to the actual morphology of tree branches and trunks.

[0124] In a preferred embodiment, S4 includes:

[0125] S41, establish the NUD algorithm. The NUD algorithm classifies sparse point clouds based on the morphological characteristics of tree branches and trunks. For different categories of sparse point clouds of tree branches and trunks, corresponding densification strategies (normal and uniform distribution) and densification parameters (number of points in the normal distribution, mean and variance; number of points in the uniform distribution, upper and lower bounds of the distribution, etc.) are used. By fully considering the morphological characteristics of tree branches and trunks, the NUD algorithm can achieve a denser effect for tree branch and trunk point clouds.

[0126] S42, based on the NUD algorithm, the sparse point cloud is densified to obtain a dense point cloud.

[0127] In this embodiment, S42 includes:

[0128] (1) The NUD algorithm classifies sparse point clouds based on the morphological characteristics of tree branches. Specifically, when traversing to a point, the algorithm determines whether there are other points in the grids before and after that point.

[0129] In this embodiment, based on the presence of points in the preceding and following grids, the current sparse point cloud can be divided into the following 9 categories, which correspond to the 9 classifications in Table 1:

[0130] Category ①: No points exist in either the front or back grid.

[0131] Category ②: A point exists in the previous grid and is located in the first half of the grid, but the point does not exist in the next grid;

[0132] Category ③: A point exists in the previous grid and is located in the middle of the grid, but the point does not exist in the next grid;

[0133] Category 4: There are no points in the previous grid, but there are points in the next grid, and the points are located in the first half of the grid.

[0134] Category 5: A point exists in the previous grid but in the next grid, and its position is close to the second half of the grid it belongs to;

[0135] Category 6: Points exist in both the preceding and following grids, and the points in the preceding grid are located in the first half of their respective grids, while the points in the following grid are located in the first half of their respective grids.

[0136] Category 7: Points exist in both the preceding and following grids, with the points in the preceding grid located near the front half of the grid and the points in the following grid located near the back half of the grid.

[0137] Category 8: Points exist in both the preceding and following grids, with the points in the preceding grid located near the latter half of the grid, and the points in the following grid located near the former half of the grid.

[0138] Category 9: Points exist in both the preceding and following grids, and the position of the point in the preceding grid is close to the last 1 / 2 of its own grid, while the position of the point in the following grid is close to the last 1 / 2 of its own grid.

[0139] (2) Based on the sparse point cloud after classification, different densification strategies are adopted.

[0140] In this embodiment, the densification strategies corresponding to specific categories are shown in Table 1. Details are as follows:

[0141] For categories ①, ②, ⑤, and ⑦, these categories correspond to the situation in real-world scenarios where tree branch point clouds exist only in the current grid and not in the preceding or following neighboring grids. Therefore, the algorithm employs a normal distribution densification method. The mean of the normal distribution is taken from the centroid value of the points in the current grid, while the variance and the number of generated points are dynamically adjusted based on the elevation of the point cloud.

[0142] For category ③, which corresponds to the case where the tree branch point cloud exists in the current grid and the previous neighboring grid, but not in the subsequent neighboring grid, the algorithm adopts a uniform distribution densification method. The lower bound of uniform distribution is the centroid value of the points in the previous grid, and the upper bound is the centroid value of the points in the current grid. The number of points generated is also dynamically adjusted according to the elevation of the point cloud.

[0143] For category ④, the point cloud of tree branches exists in the current grid and the next neighboring grid, but not in the previous neighboring grid. The densification strategy is similar to that of category ③, but the lower and upper bounds of the uniform distribution are adjusted to the centroid values ​​of the points in the current grid and the centroid values ​​of the points in the next grid, respectively.

[0144] For category ⑥, the point cloud corresponding to tree branches exists in the current grid, the previous neighbor grid, and the next neighbor grid. In this case, there are two branches, one located in the previous neighbor grid and the other located in both the current and next neighbor grids. In this case, the algorithm uses a uniform distribution to densify the points in the current grid, with the lower bound being the centroid value of the points in the current grid and the upper bound being the centroid value of the points in the next grid. The number of points generated is still dynamically adjusted according to the elevation of the point cloud.

[0145] For category ⑧, the point cloud corresponding to tree branches exists in the current grid, the previous neighboring grid, and the next neighboring grid. However, in this case, there is only one branch that spans the previous neighboring grid, the current grid, and the next neighboring grid. In this case, a uniformly distributed densification method is adopted. The lower bound is the starting position value of the current grid and the ending position value of the current grid. The number of generated points is still dynamically adjusted according to the elevation of the point cloud.

[0146] For category ⑨, the point cloud corresponding to tree branches exists in the current grid, the previous neighbor grid, and the next neighbor grid. In this case, there are three branches, located in the previous neighbor grid, the current grid, and the next neighbor grid, respectively. In this case, a normal distribution densification method is used. The mean of the normal distribution is taken from the centroid value of the points in the current grid, while the variance and the number of generated points are dynamically adjusted according to the elevation of the point cloud.

[0147] Table 1. Classification of each point in the tree branch point cloud by the NUD algorithm and the corresponding densification strategy adopted.

[0148]

[0149]

[0150] In a preferred embodiment, S5 includes:

[0151] S51, establish a point cloud correction network. The point cloud correction network consists of convolutional layers, BN layers, and fully connected layers. It takes the coordinates of the original coarse point cloud as input, learns and adjusts the coordinate positions, and outputs the corrected coordinates at the corresponding positions. This allows for more refined correction of the existing point cloud coordinates, thereby reducing accumulated errors and improving the overall reconstruction accuracy.

[0152] S52, input the coordinates of the original coarse point cloud into the point cloud correction network, and output the corrected coordinates of the corresponding point at the corresponding position.

[0153] In this invention, when using the R-pix2pix network for depth estimation, the layered, iterative training method leads to a gradual accumulation of errors within each layer, resulting in inaccurate depth estimation. Simultaneously, during point cloud densification, the tree branch point cloud obtained by the NUD algorithm is not refined enough. These two issues jointly affect the overall accuracy of point cloud reconstruction, specifically manifesting as poor reconstruction of details in the tree branch point cloud and deviations in the point cloud coordinates of these details. Therefore, an effective method is needed to correct the point cloud coordinates.

[0154] Traditional point cloud coordinate correction methods primarily rely on registration algorithms, aiming to adjust point cloud coordinates through mathematical transformations or model learning to eliminate deviations caused by factors such as acquisition equipment errors and environmental noise. Registration algorithms calculate the similarities and differences between different point clouds, using transformations such as rotation and translation to achieve optimal alignment in space, thereby correcting coordinate deviations. However, due to the complex and varied shapes and structures of tree branch point cloud data, registration algorithms often struggle to find accurate matching points when dealing with complex point cloud formations, resulting in unsatisfactory correction effects. Furthermore, these methods typically only adjust and optimize point cloud coordinates holistically, failing to fine-tune local details, thus affecting the accuracy and effectiveness of the correction. Therefore, while traditional point cloud coordinate correction methods can correct coordinate deviations to some extent, their effectiveness in handling detailed tree branch areas still needs improvement. This invention aims to explore more suitable methods and technologies to improve the accuracy and efficiency of point cloud coordinate correction.

[0155] To address the aforementioned problems, this invention proposes a novel point cloud correction network. The core idea of ​​this network is to perform point-by-point correction on existing point cloud coordinates to reduce accumulated errors and improve the accuracy of the overall reconstruction.

[0156] Specifically, this invention employs a deep learning network that takes the coordinates of the original coarse point cloud as input, learns and adjusts the coordinate positions, and outputs the corrected coordinates at the corresponding locations. In this way, fine-tuning can be performed on the tree branch point cloud obtained by single-view point cloud reconstruction methods, correcting errors and imprecision. The advantage of this point cloud correction network lies in its ability to make fine adjustments to each point, rather than uniformly transforming or optimizing the entire point cloud. This allows the network to better handle local details and complex structures, improving the accuracy of point cloud reconstruction. A detailed description of this network follows.

[0157] Figure 5This is the architecture diagram of the network model. The network consists of convolutional layers, batch normalization (BN) layers, and fully connected layers. A 1x4 data vector passes through Conv-32, Conv-64, Conv-128, and Conv-256 layers sequentially, interspersed with BN layers to prevent gradient vanishing. It then passes through three fully connected layers: FC-512, FC-256, and FC-1, finally outputting a 1x1 value, which is the corrected y-coordinate. The core of the network consists of multiple convolutional layers that progressively extract deeper features from the point cloud data. Specifically, the 1x4 data vector first passes through the Conv-32 layer to initially capture local features. Subsequently, the data passes through Conv-64, Conv-128, and Conv-256, learning more abstract and higher-level feature representations. Batch normalization (BN) layers are interspersed between the convolutional layers. The main function of the Batch Normalization (BN) layer is to normalize each batch of data, thus addressing the vanishing gradient problem that may occur during training, thereby accelerating the network's convergence speed and improving its stability. After a series of convolutional and BN layers, the data is passed to the fully connected layers for further processing. The role of the fully connected layers is to globally integrate the previously extracted features and map them to the output space. This network uses three fully connected layers, FC-512, FC-256, and FC-1, to progressively map the feature vectors to the final output space through layer-by-layer dimensionality reduction. After processing by the fully connected layers, the network outputs a 256*4 feature vector. Finally, this feature vector outputs a 1*1 value, which is the corrected y-coordinate.

[0158] Example 2

[0159] like Figure 6 As shown, this embodiment provides a single-view point cloud reconstruction system for tree branches in a real-world scene, used to implement the method of Embodiment 1, including:

[0160] Image acquisition module 101 acquires a single RGB image of tree branches in a real scene;

[0161] The single-view depth estimation module 102 based on cyclic prediction is used to obtain multiple depth images by performing single-view depth estimation based on cyclic prediction on a single RGB image of the tree branches.

[0162] The sparse point cloud conversion module 103 is used to convert the multiple depth images into point clouds to obtain sparse point clouds.

[0163] The point cloud densification module 104 based on tree morphology feature analysis is used to densify the sparse point cloud to obtain a dense point cloud.

[0164] The point cloud correction module 105 based on point-by-point correction is used to refine the point cloud coordinates of the dense point cloud to obtain the reconstruction result.

[0165] The point cloud reconstruction module 106 is used to reconstruct a single-view point cloud of tree branches in a real scene based on the reconstruction results.

[0166] Experimental Results and Analysis: The proposed single-view point cloud reconstruction method for tree branches in real-world scenarios was evaluated, including experimental results evaluating the performance of the R-pix2pix network and the reconstruction results.

[0167] (I) Experimental Environment and Parameter Settings

[0168] The computer equipment used in the experiments of this invention embodiment has the following configuration: Intel(R) Core(TM) i7-12700F@2.1GHz, NVIDIA GeForce GTX 4090 graphics card, Python 3.8.0 (64-bit) and PyTorch 1.7.1.

[0169] Furthermore, for the single-view depth estimation method based on cyclic prediction, this embodiment sets the batch size to 32 to balance training efficiency and resource consumption. Simultaneously, to address the training requirements of image input network structures with different characteristics, this embodiment sets the learning rate of the discriminator to 1e-6 and the learning rate of the generator to 1e-5. This learning rate configuration aims to ensure the stability and convergence of the training process.

[0170] (II) Evaluation Indicators

[0171] In evaluating the method of this invention, this embodiment employs multiple evaluation metrics to ensure a comprehensive and objective measurement of its performance. These metrics cover the performance of the R-pix2pix network and the overall quality of the reconstruction results. The specific reasons for their selection and their advantages are as follows:

[0172] 1. First, for performance evaluation of the R-pix2pix network, this invention selects Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM). PSNR is an important indicator of image quality, reflecting the accuracy of reconstruction by comparing the pixel differences between the original image and the reconstructed image. The PSNR value ranges from [0, +∞). A higher PSNR value indicates less distortion and better performance in the point cloud correction process. SSIM, on the other hand, focuses on the structural information of the image and can evaluate the network's ability to maintain the integrity of the image structure. The SSIM value ranges from [0, 1]. The closer the SSIM value is to 1, the better the network performance. By combining PSNR and SSIM, this invention can comprehensively evaluate the performance of the R-pix2pix network in the tree branch depth prediction task.

[0173] 2. Secondly, to comprehensively evaluate the overall reconstruction results, this embodiment of the invention introduces Hausdorff distance and Chamfer distance. Hausdorff distance is an important indicator for measuring the similarity between two point sets; it reflects the maximum mismatch between the reconstructed point cloud of tree branches and the real point cloud. By calculating Hausdorff distance, this invention can determine the magnitude of the deviation between the reconstructed point cloud and the real point cloud, thereby evaluating the accuracy of the tree branch reconstruction results. Chamfer distance, on the other hand, focuses on the overall distribution of the reconstructed point cloud of tree branches and can measure the average distance between point pairs between the reconstructed point cloud and the real point cloud. The smaller the chamfer distance, the closer the reconstruction result is to the real tree branch point cloud overall.

[0174] In addition to the quantitative indicators mentioned above, this invention also visualizes the reconstruction results to more intuitively observe the differences between this method and other methods in the reconstruction results.

[0175] (III) Experimental Results and Analysis

[0176] 1. Performance of R-pix2pix network under different loss functions

[0177] To verify the performance of the proposed deep estimation network, this invention evaluates the R-pix2pix network in real-world scenarios. The embodiments of this invention experimentally test the performance of the R-pix2pix network under different loss functions, and use two commonly used metrics, PSNR and SSIM, along with visualization results, to measure the network's performance.

[0178] Experimental results are as follows Figure 7As shown in the figure, each row represents the depth image prediction performance at different layers, and each column represents the prediction performance under different loss functions and the difference from the ground truth. The cleaner the difference image, the better the performance of that loss function in the depth estimation task. The results show that the three loss functions perform differently in depth image prediction tasks at different layers. Near the edge layers (i.e., layers one and two, and layers seven and eight), the performance of MAE, MSE, and Huber loss functions is relatively similar, with no obvious advantage or disadvantage. However, when focusing on depth image prediction near the middle layers, the differences among the three begin to emerge. Specifically, the Huber loss function performs best, generating more detailed and accurate depth images; the MAE and MSE loss functions perform poorly, and their sensitivity to outliers results in less detailed images in some areas. This is because the Huber loss function cleverly combines the advantages of MAE and MSE. When the error is small, it behaves like MSE, which can better approximate the optimal solution; when the error is large, it behaves like MAE, which has better robustness to outliers. This characteristic allows the Huber loss function to more effectively balance the network's handling of outliers and accurate values ​​in depth image prediction tasks, resulting in more accurate and stable prediction results.

[0179] like Figure 8The figure shows a quantitative comparison of PSNR and SSIM metrics for eight pix-pix networks under different loss functions, including MAE, MSE, and the Huber loss function introduced in this invention. The horizontal axis represents the pix-pix network at each layer, and the horizontal axis represents the PSNR and SSIM values. Experimental results show that the R-pix2pix network exhibits a trend of first decreasing and then increasing in both PSNR and SSIM metrics. This trend can be attributed to the distribution characteristics of depth information at different layers in the depth image. Depth images near the edge layers store less depth information, making prediction tasks relatively easy for the R-pix2pix network at these layers, resulting in higher performance. However, as the depth increases towards the middle layers, the amount of depth information stored in the depth image gradually increases, the network needs to process more information, and the prediction task becomes more complex and difficult, leading to a decrease in performance metrics. The comparison results of the three loss functions show that, overall, the Huber loss function outperforms MAE and MSE. Further analysis of the performance of neurons at different layers reveals that the MSE loss function performs better in neurons near the edge layers. This is because depth images near the edges often contain less depth information and have more white space, while the MSE loss function performs better on images with more pronounced edge features. MSE evaluates the loss by calculating the sum of the squares of the differences between the predicted and true values, and is more sensitive to subtle changes in edges, thus achieving better results in this case. However, in neurons near the middle layers, the depth image stores more depth information and has less white space. In this case, the MAE loss function performs better. MAE evaluates the loss by calculating the absolute value of the difference between the predicted and true values, and performs better on images with rich foreground features. Because the middle layers contain more depth information, MAE can better capture these features, thus achieving more accurate depth estimation.

[0180] 2. Comparison of reconstruction effects in obscured areas

[0181] To comprehensively evaluate the occlusion region reconstruction effect of the proposed method, this embodiment employs a single-view point cloud reconstruction method as a comparative algorithm for in-depth comparative analysis. This comparative algorithm is a general single-view point cloud reconstruction method proposed by Hassner T et al., which constructs a mapping network from RGB images to depth images to achieve depth prediction of RGB images, thereby reconstructing the 3D structure of objects. Experiments have demonstrated the effectiveness of this method. Because this method abandons complex parameterized networks, it exhibits greater flexibility and can effectively adapt to objects of different shapes and structures. Simultaneously, this method demonstrates high accuracy in depth estimation. Therefore, to fully verify the effectiveness of the proposed single-view point cloud reconstruction method in tree branch reconstruction, this preferred embodiment selects this method as a comparative algorithm and conducts comparative experiments with the method of this invention to further highlight the advantages of the proposed method. Figure 9 The results are compared with those of traditional single-view point cloud reconstruction.

[0182] like Figure 9As shown, the analysis of the visualization results clearly demonstrates the differences between the traditional single-view point cloud reconstruction method and the method proposed in this invention in terms of point cloud reconstruction. Specifically, from the xoy perspective, the reconstruction results obtained by both the traditional single-view point cloud reconstruction method and the method proposed in this invention show only minor differences in morphology compared to the true values. In contrast, from the xoz and yoz perspectives, the point cloud reconstructed by the method proposed in this invention is very close to the true values, with almost no significant difference, while the results obtained by the traditional single-view point cloud reconstruction method show significant differences compared to the true values, with the structure of both main branches and secondary branches differing somewhat from the true values. This result is due to two reasons: First, it may be because the RGB images of tree branches in this data were taken from the xoy perspective. From the xoy perspective, the method of this invention, due to its cyclic depth prediction strategy, provides more accurate depth information compared to the traditional single-view point cloud reconstruction method, which predicts all depth information at once. Secondly, the high similarity between tree branches and the background environment in real RGB images makes it difficult for traditional single-view point cloud reconstruction methods to distinguish between foreground and background, resulting in discrepancies in the reconstruction results for some structures. This difference is particularly noticeable from an xoy perspective, further highlighting the limitations of traditional methods in handling such complex scenes. For example, in practical applications of distinguishing foreground and background, taking peach trees with specific planting methods as an example, a pole is usually placed in the center of the tree, and the branches are stretched with ropes. This approach aims to ensure the tree's growth form and stability. However, since the RGB images used in this invention were captured by a drone with a camera, the perspective is primarily overhead. From this perspective, the central pole inevitably obscures some branches of the tree, such as... Figure 10 As shown.

[0183] Figure 11 This paper presents a detailed visual comparison of the point cloud reconstruction methods proposed in this invention and traditional point cloud reconstruction methods in terms of reconstruction effects in occluded areas. The first column of the figure shows the results of the traditional point cloud reconstruction algorithm comparison, the second column shows the reconstruction results of this method, and the third column shows the ground truth results. Each row represents the reconstruction results of the detailed parts of the occluded area for each algorithm. The black boxes indicate the comparison results of the occluded area reconstruction. When using traditional single-view point cloud reconstruction methods, this invention's embodiments found that tree branches occluded by poles cannot be effectively reconstructed. This limitation mainly stems from the shortcomings of traditional methods in handling occlusion problems. In contrast, the single-view point cloud reconstruction method proposed in this invention can overcome this problem. Figure 11As can be seen, even branches obscured by poles can be completely reconstructed using the method of this invention. More importantly, the reconstructed branch structures are almost identical to the corresponding branch structures in the ground truth, thus ensuring high accuracy and realism of the reconstruction results. Traditional single-view point cloud reconstruction methods, on the other hand, cannot reconstruct tree branches in obscured areas. Figure 11 As can be seen, in the reconstruction results of the traditional single-view point cloud reconstruction method, the tree branches in the occluded areas appear incomplete.

[0184] As can be seen from the visualization results, the single-view point cloud reconstruction method proposed in this invention has significant advantages in handling occlusion problems and can more accurately restore the branch structure of the occluded part.

[0185] 3. Comparison of global reconstruction results

[0186] Figure 12 A visualization showing the comparison between the multi-view point cloud reconstruction method and the single-view point cloud reconstruction method proposed in this invention is presented. idThis is used as an identification code for trees. This comparison aims to explore the differences in reconstruction results between different methods. In multi-view reconstruction, this invention selected 5, 100, and 200 RGB images as input data. These RGB images were randomly selected from all available images, reflecting the impact of different input quantities on the reconstruction results. From the visualization results, this invention can observe the following obvious trends. First, when using 5 RGB images for multi-view reconstruction, the reconstructed point cloud is very sparse, and even the structure of the main branches has a large error compared to the true value. This indicates that a small number of input images leads to insufficient reconstruction information. Second, as the number of input images increases, the reconstruction effect gradually improves. When using 100 RGB images for reconstruction, the main branches of the point cloud gradually become clearer and closer to the true branch structure, but the overall structure is still relatively sparse. Finally, when the number of input images reaches 200, although the structure of the main branches is very clear, the overall point cloud still appears relatively sparse, and the secondary branches are almost impossible to reconstruct effectively. This phenomenon may be due to the inherent characteristics of multi-view reconstruction technology itself. In SfM-MVS, feature extraction is typically performed using algorithms such as SIFT. However, in the real-world scenario used in this invention, the foreground and background in the RGB images of tree branches are too similar, making it difficult to effectively extract features from secondary branch regions, thus affecting the restoration of secondary branch structures during reconstruction. In contrast, the single-view point cloud reconstruction method proposed in this invention overcomes this problem. By using the image-to-point cloud angle registration method proposed in this section, even subtle regions like secondary branches can have their depth information features captured. Therefore, the reconstruction results of the method in this invention are denser and more refined, enabling a more accurate restoration of the detailed structures in the real-world scene.

[0187] Figure 13 This paper presents a comparison of the results of multi-view reconstruction using different numbers of images with the single-view point cloud reconstruction technique proposed in this invention. Through this comparison, the present invention aims to explore how many RGB images are equivalent to in multi-view reconstruction for the effectiveness of the proposed single-view point cloud reconstruction technique.

[0188] For quantitative comparison, this invention uses Hausdorff distance as the evaluation metric. Hausdorff distance, rather than chamfer distance, was chosen because it focuses on the maximum distance between the nearest points in two point sets, better reflecting the similarity of the overall structure. Specifically, Hausdorff distance is calculated by taking the distance between each point in one point set and all points in the other, then selecting the minimum distance value, and finally choosing the maximum of these minimum distance values ​​as the final Hausdorff distance. Chamfer distance, on the other hand, is a weighted distance metric that focuses more on measuring the local similarity between two point clouds or shapes. Considering that the goal of this invention is to evaluate the performance of single-view point cloud reconstruction technology from the perspective of overall structure, Hausdorff distance was chosen as the evaluation metric.

[0189] according to Figure 13 The experimental results show that the single-view point cloud reconstruction technology proposed in this invention is equivalent to multi-view reconstruction using approximately 200 RGB images. This result demonstrates that the single-view point cloud reconstruction technology proposed in this invention achieves a considerably high performance level, effectively approaching or even surpassing the effect of multi-view reconstruction using a large number of images.

[0190] The visualization and quantitative comparison results above demonstrate that the single-view point cloud reconstruction technology proposed in this invention exhibits excellent performance, outperforming traditional single-view point cloud reconstruction methods and effectively solving the problem of ineffective reconstruction of occluded areas in traditional single-view point cloud reconstruction. A comparison with multi-view reconstruction shows that the proposed single-view point cloud reconstruction technology demonstrates superior performance; specifically, its reconstruction effect is equivalent to multi-view reconstruction using approximately 200 images, showcasing outstanding performance. The single-view point cloud reconstruction technology proposed in this invention provides new ideas and methods for research in the field of tree branch reconstruction, and is expected to promote technological progress and application expansion in the future.

[0191] (IV) Conclusion

[0192] To fully verify the effectiveness and practicality of the proposed single-view point cloud reconstruction method, experiments were conducted on a real peach tree dataset. The proposed method was also comprehensively compared with traditional single-view point cloud reconstruction and multi-view reconstruction techniques. Experimental results show that the proposed single-view point cloud reconstruction technique exhibits superior performance. In terms of reconstruction quality, the proposed method significantly outperforms traditional single-view point cloud reconstruction methods, especially in handling occluded areas, effectively addressing the problem of traditional methods' inability to reconstruct effectively. Furthermore, compared to multi-view reconstruction techniques, the proposed single-view point cloud reconstruction method achieves the equivalent reconstruction effect of approximately 200 images used in multi-view reconstruction, and also surpasses multi-view reconstruction techniques in reconstructing details of tree branches, fully demonstrating its superior performance and practicality. The proposed single-view point cloud reconstruction technique provides new ideas and methods for research in the field of tree branch reconstruction, and is expected to promote technological progress and application expansion in the future.

[0193] The present invention also provides a memory that stores a plurality of instructions for implementing the method as described in Embodiment 1.

[0194] like Figure 14 As shown, the present invention also provides an electronic device, including a processor 301 and a memory 302 connected to the processor 301. The memory 302 stores a plurality of instructions, which can be loaded and executed by the processor to enable the processor to perform the method as described in Embodiment 1.

[0195] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A method for single-view point cloud reconstruction of tree branches in real scenes, characterized in that, The method comprises the following steps: S1, acquiring a single RGB image of a tree branch in a real scene; S2, obtaining a plurality of depth images by performing single-view depth estimation based on cycle prediction on the single RGB image of the tree branch; S3, converting the plurality of depth images into a sparse point cloud to obtain a sparse point cloud; S4, densifying the sparse point cloud to obtain a dense point cloud; S5, performing fine correction on the coordinates of the dense point cloud to obtain a reconstruction result; S6, reconstructing a single-view point cloud of the tree branch in the real scene based on the reconstruction result. The S2 comprises the following steps: S21, dividing the point cloud into grids based on a point cloud three-dimensional information hierarchical storage strategy, and storing the gravity information in the grids into a plurality of two-dimensional images in a hierarchical manner; S22, establishing a cycle pixel-to-pixel conversion network, wherein the cycle pixel-to-pixel conversion network is used to implement cycle prediction, and in the prediction of each layer of depth image, the previous layer of depth image is taken as a conditional input, so as to generate a sequence of depth images layer by layer; each neuron of the cycle pixel-to-pixel conversion network is composed of a generator and a discriminator, wherein the basic structure of the conditional generative adversarial network is adopted, and the generator and the discriminator compete with each other and co-evolve; S23, inputting the single RGB image of the tree branch into the cycle pixel-to-pixel conversion network to cyclically predict the depth information of each layer of the tree branch; The specific structure of the cycle pixel-to-pixel conversion network is as follows: (1) input, output and training process of the cycle pixel-to-pixel conversion network: A recurrent pixel-to-pixel translation network receives a single RGB image As input, output the next layer of depth images; in the training process, the first layer of depth images is predicted by the pix2pix_0 network ; then, and As conditional input, use the pix2pix_1 network to predict the second layer of depth images ; such an iterative cycle until the Nth layer of depth images is predicted , so as to predict all layers of depth images (2) each neuron of the cycle pixel-to-pixel conversion network is composed of a generator and a discriminator; wherein the generator adopts a U-Net structure and is used to convert the input RGB image into a corresponding depth image, and the discriminator is used to judge the authenticity of the generated image; through a series of convolutional layers and pooling layers, the features of the input image are extracted; in the down-sampling process, the size of the feature matrix gradually decreases and the number of channels gradually increases, and finally a 1*1 feature vector is obtained; the feature vector is converted into a score representing the authenticity of the image through a sigmoid function, which is used to optimize the adversarial loss between the generator and the discriminator; the Huber loss function is introduced in the generator.

2. The method of claim 1, wherein, In the S21, a mapping relationship of one RGB image corresponding to N depth images is set; in the network training and testing process, the next depth image is predicted by using the known RGB image and the previous depth image, and the iteration is performed until all the depth images are predicted, comprising the following steps: Step a, grid division is performed on the depth axis of the three-dimensional tree; the depth axis is set as the y axis, the grid is set as M, and the size of each grid is wherein is the maximum value of the coordinates of the points on the y axis; Step b, calculating the relative gravity position of the points in each grid, and calculating the average relative gravity of all points in each grid by counting the point cloud data in each grid, so as to obtain the representative position of the points in each grid; Step c, encoding the average relative gravity position information of the M grids into the RGB channel value of each pixel point in the depth image; the point cloud data in the three-dimensional space is converted into the pixel value of the two-dimensional image, and N depth images are obtained.

3. The method of claim 1, wherein, The S4 comprises: based on tree morphology feature analysis, the sparse point cloud is densified to obtain a dense point cloud, comprising: S41, establishing a NUD algorithm, the NUD algorithm classifies the sparse point cloud according to the morphology of the tree branches, and uses corresponding densification strategies and densification parameters for different categories of tree branch sparse point clouds; wherein the densification strategy includes normal and uniform distribution; the densification parameters include one or more of the number of normal distribution mean points, mean and variance, the number of uniform distribution mean points, and the upper and lower bounds of the uniform distribution; S42, based on the NUD algorithm, the sparse point cloud is densified to obtain a dense point cloud.

4. The method of claim 3, wherein, The S42 comprises: (1) the NUD algorithm classifies the sparse point cloud according to the morphology of the tree branches, specifically, when traversing to the current point, the algorithm will judge whether there are other points in the front and back grids of the point; according to the existence of the points in the front and back grids, the current sparse point cloud is divided into 9 categories, including: Category ①: no points in the front and back grids; Category ②: there are points in the front grid, and the position is close to the front 1 / 2 of the grid, and there are no points in the back grid; Category ③: there are points in the front grid, and the position is close to the back 1 / 2 of the grid, and there are no points in the back grid; Category ④: there are no points in the front grid, and there are points in the back grid, and the position is close to the front 1 / 2 of the grid; Category ⑤: there are no points in the front grid, and there are points in the back grid, and the position is close to the back 1 / 2 of the grid; Category ⑥: there are points in the front and back grids, and the position of the point in the front grid is close to the front 1 / 2 of the grid, and the position of the point in the back grid is close to the front 1 / 2 of the grid; Category ⑦: there are points in the front and back grids, and the position of the point in the front grid is close to the front 1 / 2 of the grid, and the position of the point in the back grid is close to the back 1 / 2 of the grid; Category ⑧: there are points in the front and back grids, and the position of the point in the front grid is close to the back 1 / 2 of the grid, and the position of the point in the back grid is close to the front 1 / 2 of the grid; Category ⑨: there are points in the front and back grids, and the position of the point in the front grid is close to the back 1 / 2 of the grid, and the position of the point in the back grid is close to the back 1 / 2 of the grid; (2) based on the classified sparse point cloud, different densification strategies are adopted, including: For categories ①, ②, ⑤ and ⑦, the normal distribution densification method is adopted, wherein the mean of the normal distribution adopts the center of gravity value of the current grid, and the variance and the number of generated points are dynamically adjusted according to the altitude of the point cloud; For category ③, the uniform distribution densification method is adopted, the lower bound of the uniform distribution is the center of gravity value of the point in the front grid, the upper bound is the center of gravity value of the point in the current grid, and the number of generated points is also dynamically adjusted according to the altitude of the point cloud; For category ④, the uniform distribution densification method is adopted, but the lower bound and the upper bound of the uniform distribution are adjusted to the center of gravity value of the point in the current grid and the center of gravity value of the point in the back grid, respectively; For category ⑥, a uniform distribution densification method is adopted for the points in the current grid, the lower bound is the barycenter value of the points in the current grid, the upper bound is the barycenter value of the points in the next grid, and the generated point number is still dynamically adjusted according to the elevation of the point cloud; For category ⑧, a uniform distribution densification method is adopted, the lower bound is the start position value of the current grid, the upper bound is the end position value of the current grid, and the generated point number is still dynamically adjusted according to the elevation of the point cloud; For category ⑨, a normal distribution densification method is adopted, the mean value of the normal distribution adopts the barycenter value of the points in the current grid, and the variance and the generated point number are dynamically adjusted according to the elevation of the point cloud.

5. The method of claim 4, wherein, The S5 comprises: S51, establishing a point cloud correction network composed of convolution layers, BN layers and fully connected layers, taking the coordinates of the original rough point cloud as input, adjusting the coordinate position through learning, and outputting the corrected coordinates at the corresponding position, so as to make more fine correction on the existing point cloud coordinates, reduce the cumulative error and improve the accuracy of the overall reconstruction; S52, inputting the coordinates of the original rough point cloud into the point cloud correction network, and outputting the coordinates of the corresponding points after correction.

6. A system for single-view point cloud reconstruction of tree branches in real scenes, for implementing the method according to any one of claims 1 to 5, characterized in that, Comprise: An image acquisition module (101) acquires a single RGB image of a tree branch in a real scene; A single-view depth estimation module (102) based on cycle prediction is used to perform single-view depth estimation based on cycle prediction on the single RGB image of the tree branch to obtain multiple depth images; A sparse point cloud conversion module (103) is used to convert the multiple depth images into point clouds to obtain a sparse point cloud; A point cloud densification module (104) based on tree morphology feature analysis is used to densify the sparse point cloud to obtain a dense point cloud; A point cloud correction module (105) based on point-by-point correction is used to finely correct the coordinates of the dense point cloud to obtain a reconstruction result; A point cloud reconstruction module (106) is used to reconstruct the point cloud of the tree branch in the real scene based on the reconstruction result.

7. An electronic device, comprising a processor and a memory, the memory storing a plurality of instructions, wherein, The processor is configured to read the instructions and perform the method of any one of claims 1-5.

8. A computer-readable storage medium, the computer-readable storage medium storing a plurality of instructions, characterized in that, The plurality of instructions can be read and executed by the processor to perform the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Rapid three-dimensional reconstruction method for mine inspection scene of unmanned aerial vehicle

    CN112288875A

  • Tree branch skeleton extraction method and system for point cloud data collected by unmanned aerial vehicle

    CN116843693A