Multi-platform laser point cloud interactive individual tree extraction method based on deep learning
By constructing a Gaussian guided feature map through a point cloud interactive method and combining it with a deep neural network and a spatial neighborhood voting strategy, the accuracy and robustness issues of single tree extraction in multi-platform laser point cloud data are solved, and efficient and accurate single tree instance segmentation is achieved, which is suitable for forestry resource surveys and ecological monitoring.
Patent Information
- Application Number
- CN202510762230.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies have difficulty in efficiently and accurately extracting individual tree instances from multi-platform laser point cloud data in complex forest environments. In particular, problems such as missed detection, multiple detections, and blurred boundaries exist in natural forest areas with dense overlapping crowns or complex tree species structures. In addition, deep learning models lack portability and robustness, and lack user interactive control capabilities.
Through interactive point cloud visualization, the target area is selected, and positive and negative sample points are obtained to construct a Gaussian guided feature map. After fusion with the original point cloud data, it is input into a deep neural network for segmentation. The label is completed by combining the spatial neighborhood majority voting strategy to generate a structurally complete and semantically consistent single tree instance mask result.
It achieves high-precision, low-interaction-cost single-tree extraction in complex forest environments. It is applicable to multi-platform point cloud data, improves the structural integrity and label consistency of segmentation results, and has strong practicality and promotion value.
Smart Images

Figure CN120747730A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a multi-platform laser point cloud interactive single tree extraction method based on deep learning. Background Art
[0002] Accurately extracting individual trees from a forest is fundamental for forestry applications such as tree species identification, biomass estimation, and ecological modeling. With the development of multi-platform acquisition methods such as unmanned aerial vehicle (UAV)-LiDAR and mobile laser scanning (MLS), the acquisition of forest point cloud data has become more efficient and dense. However, due to complex forest structure, significant crown overlap, and frequent trunk occlusion, efficient and accurate extraction of individual tree instances from massive 3D point clouds remains a significant challenge in forest information modeling.
[0003] The current mainstream single tree extraction methods mainly include single tree instance segmentation based on traditional algorithms or deep learning-based methods. The former relies on geometric structure analysis, density clustering or trunk / crown positioning for segmentation, and usually requires manual setting of parameters, such as crown width threshold or density radius, and special adaptation for different forest types and different point cloud platforms. This type of method is effective in regular woodlands, but in natural forest areas with dense overlapping crowns or complex tree species structures, problems such as missed detection, multiple detections and blurred boundaries often occur. In particular, segmentation methods developed based on ground-based lidar (TLS) often assume that the trunk is completely visible. When applied to data obtained by UAV or MLS, their performance is greatly reduced due to view occlusion and their generalization ability is limited.
[0004] In recent years, with the advancement of artificial intelligence, deep neural networks such as YOLO, Faster-RCNN, and Point Transformer have been gradually introduced to forest point cloud segmentation tasks, and some models have achieved high accuracy on existing datasets. However, due to the strong dependence of deep models on training samples, the forestry field generally faces problems such as high sample acquisition costs and difficulty ensuring label quality, resulting in insufficient model transferability and robustness. Furthermore, most existing methods lack user interactive control capabilities, making it difficult to quickly correct deviations in segmentation results, limiting their widespread application and practical application in real-world forestry scenarios.
[0005] In addition, point cloud data acquired from different platforms exhibit significant differences in density, noise level, distribution characteristics, and scanning trajectories, further exacerbating the difficulty of cross-domain model application. Even models that have currently achieved relatively high accuracy still face the problem of a sharp drop in accuracy when applied to cross-platform data. Therefore, there is an urgent need for a new method that combines user interaction prompts with deep neural network reasoning. This method is both flexible and controllable, and has both high accuracy and high efficiency. It can adapt to the requirements of complex forest structures, multi-platform point cloud formats, and diverse forest types, and promote the development of single tree instance extraction towards practicality and intelligence. Summary of the Invention
[0006] In view of this, the present invention provides a multi-platform laser point cloud interactive single tree extraction method based on deep learning to at least solve the above technical problems.
[0007] According to a first aspect of an embodiment of the present invention, a multi-platform laser point cloud interactive single tree extraction method based on deep learning is provided, including: S1, interactively selecting a target area through a point cloud visualization view, and obtaining positive and negative sample points in the target area for marking the target foreground and background areas; S2, constructing a Gaussian guided feature map based on the positive and negative sample points, and fusing it with the obtained original point cloud data to obtain fused data; S3, inputting the fused data into a deep neural network model for binary classification segmentation of foreground and background to obtain a preliminary tree mask result; S4, post-processing the unassigned points in the preliminary tree mask result, and using a majority voting strategy based on spatial neighborhood to complete the labels to generate a structurally complete and semantically consistent single tree instance mask result.
[0008] Furthermore, step S1 specifically includes: interacting in the three-dimensional point cloud view by manually drawing a frame or zooming the view to set an area to select an area of interest and extract the target tree and local point cloud data; determining positive sample points and negative sample points in the area of interest, respectively, the positive sample points are located in the target tree area, and the negative sample points are located in the non-target tree area; generating corresponding spherical guidance areas according to the spatial coordinates of the positive and negative sample points, which are used to construct initial label templates of the foreground guidance map and the background guidance map.
[0009] Furthermore, step S2 specifically includes: calculating the Euclidean distance between each sample point and all points in the local point cloud data according to the spatial coordinates of the positive and negative sample points, thereby forming a distance distribution matrix between the foreground and background; and performing an attenuation conversion on the distance value using a Gaussian function based on the distance distribution matrix, thereby generating a foreground guide map and a background guide map, respectively. The calculation formula of the foreground guide map is as follows:
[0010]
[0011] Among them, N posis the number of positive sample points, σ is the standard deviation of the Gaussian kernel, and the background guide map G neg (x) is constructed in the same way;
[0012] The foreground guide image G pos (x), background guide map G neg (x) is spliced with the geometric coordinates (x, y, z) and additional attributes of the corresponding local point cloud data in the channel dimension to construct a multi-channel tensor structure that meets the neural network input format requirements as the fused data. The spliced multi-channel tensor structure is as follows:
[0013] I(x)=[x,y,z,G pos (x),G neg (x),a1(x),a2(x),…,a k (x)]
[0014] Among them, a k (x) represents the value of the kth additional feature channel.
[0015] Furthermore, the step S3 specifically includes: inputting the multi-channel tensor structure I(x) into the point cloud segmentation neural network constructed based on the PointTransformer V3 architecture as the initial input feature of the deep neural network model; extracting the multi-channel tensor structure layer by layer through the multi-layer attention mechanism module set inside the network to extract the multi-scale spatial structure features; performing foreground and background probability prediction on each point in the multi-scale spatial structure features in the output layer of the neural network, and outputting a binary semantic mask map The preliminary tree mask result is used to identify the instance-level division of tree points and non-tree points. During the training process, the deep neural network model uses a joint loss function for parameter optimization. The joint loss function consists of point-level cross entropy loss and Lovász-Softmax region consistency loss, and is expressed as follows:
[0016]
[0017] in, represents the cross entropy classification loss between foreground and background, is the structural loss for the IoU indicator, and λ1 and λ2 are weighted coefficients.
[0018] Furthermore, step S4 specifically includes: performing multiple random sampling on the point cloud data blocks in the preliminary tree mask result to generate multiple specified number of point sets and inputting the multiple specified number of point sets into the deep neural network model for foreground and background prediction, thereby obtaining multiple sets of semantic mask results as multiple prediction results; fusing the multiple prediction results into the original point cloud coordinate system, identifying points that are not effectively assigned labels as unassigned points, and forming a point set to be completed X unlabeled , specifically defined as follows:
[0019]
[0020] in, represents the complete point cloud set, y(x) is the predicted label of point x;
[0021] Construct each unassigned point x∈X unlabeled spatial neighborhood And count the spatial neighbors The label distribution of the labeled points in the matrix is obtained, and the majority voting strategy is used to determine the final label of the labeled point:
[0022]
[0023] Here, mode(·) represents the label with the highest frequency in the set. If the spatial neighborhood of the current unassigned point does not contain any valid labeled points, the nearest neighbor strategy is used to assign the label of the nearest labeled point to the current unassigned point. After completing the label completion of all unassigned points, a structurally complete and semantically consistent single-tree instance mask result is output.
[0024] According to a second aspect of an embodiment of the present invention, a multi-platform laser point cloud interactive single tree extraction system based on deep learning is provided, including: an interactive module for interactively selecting a target area through a point cloud visualization view, and obtaining positive and negative sample points in the target area for marking the target foreground and background areas; a construction module for constructing a Gaussian guided feature map based on the positive and negative sample points, and fusing it with the acquired original point cloud data to obtain fused data; a segmentation module for inputting the fused data into a deep neural network model for binary classification segmentation of foreground and background to obtain a preliminary tree mask result; a post-processing module for post-processing unassigned points in the preliminary tree mask result, and using a majority voting strategy based on spatial neighborhood to complete labels to generate a structurally complete and semantically consistent single tree instance mask result.
[0025] According to a third aspect of an embodiment of the present invention, there is provided an electronic device comprising a processor and a memory storing a program, wherein the program comprises instructions that, when executed by the processor, cause the processor to perform the steps of the method according to the first aspect.
[0026] According to a fourth aspect of an embodiment of the present invention, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method of the first aspect is implemented.
[0027] In summary, the present invention can achieve high-precision, low-interaction-cost single-tree extraction from forestry scene point clouds in complex forest environments. The method constructs a Gaussian guidance map through positive and negative sample clicks, guides the neural network to perform point-level semantic prediction, and combines the attention mechanism to enhance the modeling ability of local structure and global context. Through multiple sampling fusion and spatial majority voting strategies, the label-missing areas in the semantic mask are supplemented, which effectively improves the structural integrity and label consistency of the segmentation results. This method is applicable to laser point cloud data acquired by drones or mobile handheld platforms, and has strong practicality and promotion value. It is particularly suitable for application scenarios such as forestry resource surveys and ecological monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0029] Figure 1 Flow chart of the steps of the method of the present invention.
[0030] Figure 2 Schematic diagram of the Point Transformer V3 network structure.
[0031] Figure 3 Schematic diagram of the comparison between before and after segmentation and error mask.
[0032] Figure 4 A schematic diagram comparing the effects before and after post-processing.
[0033] Figure 5 Schematic diagram of the impact of the number of interactive clicks on segmentation accuracy and inference time. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0035] This paper discloses a multi-platform interactive tree extraction method based on deep learning, suitable for high-precision tree extraction and structure recognition in three-dimensional laser point cloud data for forestry. This method requires no prior ground points or noise removal, is applicable to data collected by various LiDAR platforms, and achieves high-precision and efficient instance-level tree segmentation with low interaction costs. It is suitable for a variety of scenarios, including forest resource surveys and digital forestry.
[0036] See also Figure 1 , a multi-platform laser point cloud interactive single tree extraction method based on deep learning in an embodiment of the present invention includes:
[0037] S1. Interactively select the target area through the point cloud visualization view, and obtain positive and negative sample points in the target area to mark the target foreground and background areas;
[0038] S2, constructing a Gaussian guided feature map based on the positive and negative sample points, and fusing it with the acquired original point cloud data to obtain fused data;
[0039] S3. Input the fused data into the deep neural network model for binary segmentation of foreground and background to obtain the preliminary tree mask result;
[0040] S4. Post-process the unassigned points in the preliminary tree mask results and use the majority voting strategy based on spatial neighborhood to complete the labels and generate a structurally complete and semantically consistent single tree instance mask result.
[0041] Furthermore, the step S1 specifically includes:
[0042] S11, interactively selecting an area of interest and extracting target tree and local point cloud data by manually drawing a frame or zooming in on the 3D point cloud view;
[0043] S12, determining positive sample points and negative sample points in the region of interest, wherein the positive sample points are located in the target tree region and the negative sample points are located in the non-target tree region, that is, clicking the positive sample points in the target tree region and the negative sample points in the non-target tree region in the region of interest;
[0044] S13. Generate corresponding spherical guidance areas according to the spatial coordinates of the positive and negative sample points, and use them to construct initial label templates for the foreground guidance map and the background guidance map.
[0045] Furthermore, the step S2 specifically includes:
[0046] S21. Calculate the Euclidean distance between each sample point and all points in the local point cloud data based on the spatial coordinates of the positive and negative sample points to form a distance distribution matrix between the foreground and background.
[0047] S22. Based on the distance distribution matrix, use a Gaussian function to perform attenuation conversion on the distance value to generate a foreground guidance map and a background guidance map respectively to reflect the distribution of the guidance intensity of the sample point to the surrounding points. The calculation formula of the foreground guidance map is as follows:
[0048]
[0049] Among them, N pos is the number of positive sample points, σ is the standard deviation of the Gaussian kernel, and the background guide map G neg (x) is constructed in the same way;
[0050] S23, the foreground guide image G pos (x), background guide map G neg (x) is spliced with the geometric coordinates (x, y, z) and additional attributes (including color, intensity or normal vector, etc.) of the corresponding local point cloud data in the channel dimension to construct a multi-channel tensor structure that meets the neural network input format requirements as the fused data. The spliced multi-channel tensor structure is as follows:
[0051] I(x)=[x,y,z,G pos (x),G neg (x),a1(x),a2(x),…,a k (x)]
[0052] Among them, a k (x) represents the value of the kth additional feature channel.
[0053] Further, see Figure 2 This is a diagram of the Point Transformer V3 (PTV3) network structure (segmentation network). Figure 3 This is a schematic diagram comparing the error mask before and after segmentation. Figure 3 The segmentation effect of the proposed method on multiple real forest point cloud scenes is demonstrated, including the actual manual annotation results, model prediction results and error visualization. Figure 3 Columns. Each instance of the tree is presented in the same color. Gray represents background points, green represents correctly predicted points, and red represents incorrectly predicted points.
[0054] The step S3 specifically includes:
[0055] S31, inputting the multi-channel tensor structure I(x) into a point cloud segmentation neural network built based on the Point Transformer V3 (PTV3) architecture as the initial input feature of the deep neural network model;
[0056] S32. Perform layer-by-layer feature extraction on the input multi-channel tensor structure through the multi-layer attention mechanism module set inside the network to extract multi-scale spatial structure features;
[0057] S33. In the output layer of the neural network, the probability of foreground and background is predicted for each point in the multi-scale spatial structure feature, and a binary classification semantic mask is output. As the preliminary tree mask result, an instance-level partition is used to identify tree points and non-tree points;
[0058] S34. During the training process, the deep neural network model uses a joint loss function for parameter optimization. The joint loss function is composed of point-level cross entropy loss and Lovász-Softmax regional consistency loss, and the expression is as follows:
[0059]
[0060] in, represents the cross entropy classification loss between foreground and background, is the structural loss for the IoU indicator, and λ1 and λ2 are the weighted coefficients of the two.
[0061] Further, see Figure 4 This is a schematic diagram comparing the effects before and after post-processing. Figure 4 (a) Before automatic refinement, (b) After automatic refinement, red points represent unclassified points, and gray points represent non-tree points. Figure 5 This is a diagram showing the impact of the number of interactive clicks on segmentation accuracy and inference time. Figure 5 The middle left figure shows that the mean segmentation accuracy (mIoU) increases significantly with the number of interactive clicks. Figure 5 The middle right figure shows that the inference time remains basically stable.
[0062] The step S4 specifically includes:
[0063] S41, performing multiple random sampling on the point cloud data blocks in the preliminary tree mask result to generate multiple specified number of point sets, and inputting the multiple specified number of point sets into the deep neural network model for foreground and background prediction, thereby obtaining multiple sets of semantic mask results as multiple prediction results;
[0064] S42: Fusing the multiple prediction results into the original point cloud coordinate system, identifying points that are not effectively assigned labels as unassigned points, and forming a point set to be completed X. unlabeled , specifically defined as follows:
[0065]
[0066] in, represents the complete point cloud set, y(x) is the predicted label of point x;
[0067] S43. Construct each unassigned point x∈X unlabeled spatial neighborhood And count the spatial neighbors The label distribution of the labeled points in the matrix is obtained, and the majority voting strategy is used to determine the final label of the labeled point:
[0068]
[0069] Among them, mode(·) represents the label with the highest frequency in the set;
[0070] S44. If the spatial neighborhood of the current unassigned point does not contain any valid label point, the nearest neighbor strategy is adopted to assign the label of the nearest labeled point to the current unassigned point;
[0071] S45. After completing the label completion of all unassigned points, output the single tree instance mask result with complete structure and consistent semantics.
[0072] In summary, the method of the present invention can achieve high-precision, low-interaction-cost tree extraction in complex forest environments. The method constructs a Gaussian guidance map through positive and negative sample clicks, guides the neural network to perform point-level semantic prediction, and combines the attention mechanism to enhance the modeling ability of local structure and global context. Through multiple sampling fusion and spatial majority voting strategies, the label missing areas in the semantic mask are supplemented, which effectively improves the structural integrity and label consistency of the segmentation results. This method is applicable to laser point cloud data obtained by drones or mobile handheld platforms, has strong practicality and promotion value, and is particularly suitable for application scenarios such as forestry resource surveys and ecological monitoring.
[0073] The embodiment of the present invention further provides a multi-platform laser point cloud interactive single tree extraction system based on deep learning, comprising:
[0074] The interactive module is used to interactively select the target area through the point cloud visualization view and obtain positive and negative sample points in the target area to mark the target foreground and background areas;
[0075] A construction module is used to construct a Gaussian guided feature map based on positive and negative sample points, and fuse it with the acquired original point cloud data to obtain fused data;
[0076] The segmentation module is used to input the fused data into the deep neural network model for binary segmentation of foreground and background to obtain the preliminary tree mask result;
[0077] The post-processing module is used to post-process the unassigned points in the preliminary tree mask results, and use the majority voting strategy based on spatial neighborhood to complete the labels to generate structurally complete and semantically consistent single tree instance mask results.
[0078] It should be understood that the system of this embodiment is used to implement the corresponding methods in the aforementioned multiple method embodiments and has the beneficial effects of the corresponding method embodiments.
[0079] As another example, the present invention also provides an electronic device, which will now be described as an electronic device that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0080] The server is a cloud server, which includes a central server and an edge node server.
[0081] The electronic device may include: a processor, a communication interface, a memory, and a communication bus.
[0082] The processor, communication interface and memory communicate with each other through a communication bus. The communication interface is used to communicate with other electronic devices or servers.
[0083] The processor is used to execute programs, and specifically can execute the relevant steps in the above method embodiments.
[0084] Specifically, the program may include program codes including computer operation instructions.
[0085] The processor may be a CPU, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.
[0086] The memory is used to store programs and may include high-speed RAM memory or non-volatile memory, such as at least one disk storage.
[0087] When executed by a processor, the program is used to enable an electronic device to perform a multi-platform laser point cloud interactive single tree extraction method based on deep learning of the present invention.
[0088] In addition, the specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiments, and will not be repeated here.
[0089] An exemplary embodiment of the present invention further provides a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, the methods of the various embodiments of the present invention are implemented. The corresponding process descriptions in the aforementioned method embodiments can be referred to and will not be repeated here.
[0090] The method according to the embodiment of the present invention described above can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.
[0091] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.
[0092] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
[0093] Finally, it should be noted that the above implementation methods are only used to illustrate the embodiments of the present invention, and are not limitations on the embodiments of the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the scope of patent protection of the embodiments of the present invention should be defined by the claims.
Claims
1. A multi-platform laser point cloud interactive single tree extraction method based on deep learning, characterized in that: include: S1. Interactively select the target area through the point cloud visualization view, and obtain positive and negative sample points in the target area to mark the target foreground and background areas; S2, constructing a Gaussian guided feature map based on the positive and negative sample points, and fusing it with the acquired original point cloud data to obtain fused data; S3. Input the fused data into the deep neural network model for binary segmentation of foreground and background to obtain the preliminary tree mask result; S4. Post-process the unassigned points in the preliminary tree mask results and use the majority voting strategy based on spatial neighborhood to complete the labels and generate a structurally complete and semantically consistent single tree instance mask result.
2. The method according to claim 1, characterized in that The step S1 specifically includes: In the 3D point cloud view, you can interact by manually drawing a frame or zooming in and out to select the area of interest and extract the target tree and local point cloud data. Determine positive sample points and negative sample points in the region of interest, wherein the positive sample points are located in the target tree area and the negative sample points are located in the non-target tree area; According to the spatial coordinates of the positive and negative sample points, corresponding spherical guidance areas are generated to construct the initial label templates of the foreground guidance map and the background guidance map.
3. The method according to claim 2, characterized in that The step S2 specifically includes: According to the spatial coordinates of the positive and negative sample points, the Euclidean distance between each sample point and all points in the local point cloud data is calculated to form the distance distribution matrix between the foreground and background; Based on the distance distribution matrix, a Gaussian function is used to perform attenuation conversion on the distance value to generate a foreground guide map and a background guide map respectively. The calculation formula of the foreground guide map is as follows: Among them, N pos is the number of positive sample points, σ is the standard deviation of the Gaussian kernel, and the background guide map G neg (x) is constructed in the same way; The foreground guide image G pos (x), background guide map G neg (x) is spliced with the geometric coordinates (x, y, z) and additional attributes of the corresponding local point cloud data in the channel dimension to construct a multi-channel tensor structure that meets the neural network input format requirements as the fused data. The spliced multi-channel tensor structure is as follows: I(x)=[x,y,z,G pos (x),G neg (x),a1(x),a2(x),…,a k (x)] Among them, a k (x) represents the value of the kth additional feature channel.
4. The method according to claim 3, characterized in that The step S3 specifically includes: Input the multi-channel tensor structure I(x) into the point cloud segmentation neural network built based on the Point Transformer V3 architecture as the initial input feature of the deep neural network model; The multi-layer attention mechanism module set up inside the network extracts features of the input multi-channel tensor structure layer by layer to extract multi-scale spatial structure features; At the output layer of the neural network, the probability of foreground and background is predicted for each point in the multi-scale spatial structure feature, and a binary classification semantic mask is output. As the preliminary tree mask result, an instance-level partition is used to identify tree points and non-tree points; During the training process, the deep neural network model uses a joint loss function for parameter optimization. The joint loss function consists of point-level cross entropy loss and Lovász-Softmax region consistency loss, and is expressed as follows: in, represents the cross entropy classification loss between foreground and background, is the structural loss for the IoU indicator, and λ1 and λ2 are weighted coefficients.
5. The method according to claim 1, characterized in that The step S4 specifically includes: Perform multiple random sampling on the point cloud data blocks in the preliminary tree mask result to generate multiple specified number of point sets, and input the multiple specified number of point sets into the deep neural network model for foreground and background prediction, and obtain multiple sets of semantic mask results as multiple prediction results; The multiple prediction results are integrated into the original point cloud coordinate system, and the points that are not effectively assigned labels are identified as unassigned points to form the point set to be completed X unlabeled , specifically defined as follows: in, represents the complete point cloud set, y(x) is the predicted label of point x; Construct each unassigned point x∈X unlabeled spatial neighborhood And count the spatial neighbors The label distribution of the labeled points in the matrix is obtained, and the majority voting strategy is used to determine the final label of the labeled point: Among them, mode(·) represents the label with the highest frequency in the set; If the spatial neighborhood of the current unassigned point does not contain any valid label points, the nearest neighbor strategy is used to assign the label of the nearest labeled point to the current unassigned point; After completing the label completion of all unassigned points, the output is a single tree instance mask result with complete structure and consistent semantics.
6. A multi-platform laser point cloud interactive single tree extraction system based on deep learning, characterized by: include: The interactive module is used to interactively select the target area through the point cloud visualization view and obtain positive and negative sample points in the target area to mark the target foreground and background areas; A construction module is used to construct a Gaussian guided feature map based on positive and negative sample points, and fuse it with the acquired original point cloud data to obtain fused data; The segmentation module is used to input the fused data into the deep neural network model for binary segmentation of foreground and background to obtain the preliminary tree mask result; The post-processing module is used to post-process the unassigned points in the preliminary tree mask results, and use the majority voting strategy based on spatial neighborhood to complete the labels to generate structurally complete and semantically consistent single tree instance mask results.
7. An electronic device, characterized in that: include: processor; Memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the steps of the method according to any one of claims 1 to 5.
8. A computer storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.