Farm three-dimensional point cloud semantic segmentation method and device for unmanned aerial vehicle path planning
The three-dimensional point cloud data of farmland is obtained through drones and semantic segmentation is used for deep learning models, which solves the problem of building high-precision semantic maps in farmland scenarios, realizes high-precision land object recognition and semantic segmentation, and supports the precise operation of unmanned agricultural machinery.
Patent Information
- Application Number
- CN202510218286.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The prior art is difficult to meet the needs of high-precision semantic map construction in farmland scenarios, especially in unstructured, multi-scale and complex environments. Traditional methods are difficult to balance the fusion of detail retention and global information, resulting in insufficient semantic consistency, feature capture and geographic classification accuracy.
The farmland image data is obtained through the drone, converted into three-dimensional point cloud data, and input the farm three-dimensional point cloud semantic segmentation model for classification to obtain the landform type corresponding to different point clouds. This model extracts the local geometric features and global semantic information of each sampled point in the point cloud data subset, and trains it in combination with the deep learning model to obtain high-precision semantic segmentation results.
It realizes high-precision semantic segmentation of farmland land objects, improves the ability to identify complex land objects, improves the accuracy and robustness of land objects recognition in farmland scenarios, and meets the needs of accurate navigation, path planning and task allocation of unmanned agricultural machinery.
Smart Images

Figure CN120147634A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a three-dimensional point cloud semantic segmentation method and device for farmland oriented to UAV path planning. Background Art
[0002] With the development of agricultural intelligence, the application of unmanned agricultural machinery in precise navigation, path planning, and task allocation is becoming increasingly widespread. The realization of these functions requires high-precision three-dimensional semantic information of farmland ground objects. However, due to the characteristics of unstructured, multi-scale, and complex environments in farmland scenes, traditional ground object information extraction methods are difficult to meet actual needs, especially when constructing high-precision semantic maps in large-scale scenes, facing many challenges.
[0003] Currently, the acquisition of farmland ground object information mainly relies on geometric maps, topological maps, or grid maps. However, these methods still have great limitations when facing the complexity, diversity, and large-scale feature extraction of farmland environments, and cannot achieve a balance between detail retention and global information fusion, resulting in deficiencies in semantic consistency, feature capture, and ground object classification accuracy. It is difficult to guarantee the accuracy of ground object information acquisition, thereby reducing the efficiency and reliability of automated operations (such as driving path planning, task allocation, and navigation) in farmland scenes. Summary of the Invention
[0004] The present invention provides a three-dimensional point cloud semantic segmentation method and device for farmland oriented to UAV path planning to solve the technical problem of low ground object segmentation accuracy in the prior art.
[0005] In a first aspect, the present invention provides a three-dimensional point cloud semantic segmentation method for farmland oriented to UAV path planning, including the following steps.
[0006] Acquire image data of farmland by a UAV and convert the image data into three-dimensional point cloud data; Input the three-dimensional point cloud data into a three-dimensional point cloud semantic segmentation model for farmland to perform classification and obtain the ground object types corresponding to different point clouds; the three-dimensional point cloud semantic segmentation model for farmland is trained through the following steps: Obtain multiple subsets of point cloud data based on the three-dimensional point cloud data; Extract the local geometric features and global semantic information of each sampling point in the subset of point cloud data; Determine the predicted category of the sampling point based on the local geometric features and the global semantic information; Train a preset deep learning model based on the predicted category and the true category of the sampling point to obtain a three-dimensional point cloud semantic segmentation model for farmland.
[0007] In some embodiments, obtaining a plurality of subsets of point cloud data based on the three-dimensional point cloud data includes: Processing the three-dimensional point cloud data based on a voxel grid downsampling algorithm to obtain sampled point cloud data; Dividing the sampled point cloud data into a plurality of point cloud data subsets of the same size.
[0008] In some embodiments, extracting the local geometric features of each sampling point in the point cloud data subset includes: Obtaining all neighborhood points of each sampling point based on the K-nearest neighbor algorithm; Obtaining the local spatial features of the sampling point based on the sampling point and the neighborhood points of the sampling point; the local spatial features include a distance attenuation exponent and a local geometric bounding box; Encoding the local spatial features based on a multi-layer perceptron to obtain the local geometric features of the sampling point.
[0009] In some embodiments, extracting the global semantic information of each sampling point in the point cloud data subset includes: Calculating the volume ratio of the local neighborhood space of each sampling point to the global space; the local neighborhood space is the space where all neighborhood points of the sampling point are located; the global space is the space corresponding to the point cloud data subset; Determining the global semantic information of the sampling point based on the volume ratio.
[0010] In some embodiments, determining the predicted category of the sampling point based on the local geometric features and the global semantic information includes: Performing feature fusion on the local geometric features and the global semantic information to obtain the aggregated features of the sampling point; Determining the predicted category of the sampling point based on the aggregated features of the sampling point and the aggregated features of the neighborhood points of the sampling point.
[0011] In some embodiments, training a preset deep learning model based on the predicted category and the true category of the sampling point includes: Determining the cross-entropy loss based on the predicted category and the true category of the sampling point; Adjusting the model parameters of the deep learning model based on the cross-entropy loss.
[0012] In some embodiments, the method further includes: Generating a farmland semantic map based on the ground object types corresponding to different point clouds; the farmland semantic map is used to display the three-dimensional distribution information of different ground object types in the farmland.
[0013] Second aspect, the present invention provides a three-dimensional point cloud semantic segmentation device for farmland for UAV path planning, including the following modules.
[0014] An acquisition module, configured to acquire image data of farmland through a UAV and convert the image data into three-dimensional point cloud data; A semantic segmentation module, configured to input the three-dimensional point cloud data into a three-dimensional point cloud semantic segmentation model for a farm to perform classification to obtain the ground object types corresponding to different point clouds; the three-dimensional point cloud semantic segmentation model for a farm is obtained through the following steps: Obtain multiple subsets of point cloud data based on the three-dimensional point cloud data; Extract the local geometric features and global semantic information of each sampling point in the subset of point cloud data; Determine the predicted category of the sampling point based on the local geometric features and the global semantic information; Train a preset deep learning model based on the predicted category and the true category of the sampling point to obtain a three-dimensional point cloud semantic segmentation model for a farm.
[0015] Third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the three-dimensional point cloud semantic segmentation method for farmland for UAV path planning as described in any one of the above.
[0016] Fourth aspect, a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the three-dimensional point cloud semantic segmentation method for farmland for UAV path planning as described in any one of the above.
[0017] Fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the three-dimensional point cloud semantic segmentation method for farmland for UAV path planning as described in any one of the above.
[0018] The farm three-dimensional point cloud semantic segmentation method and device for unmanned aerial vehicle path planning provided by the present invention obtain image data of farmland through an unmanned aerial vehicle, convert the image data into three-dimensional point cloud data, input the three-dimensional point cloud data into a farm three-dimensional point cloud semantic segmentation model for classification to obtain the ground object types corresponding to different point clouds. The farm three-dimensional point cloud semantic segmentation model is trained through the following steps: obtaining multiple subsets of point cloud data based on the three-dimensional point cloud data, extracting the local geometric features and global semantic information of each sampling point therein, determining the predicted category of the sampling point based on the local geometric features and global semantic information, and training a deep learning model based on the predicted category and the true category to obtain the farm three-dimensional point cloud semantic segmentation model. It effectively solves the problem of feature extraction of multi-scale targets in the farmland scene, improves the recognition ability of complex ground objects, and enhances the accuracy of ground object recognition in the farmland scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic flow chart of the farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning provided by the present invention.
[0021] Figure 2 It is a schematic flow chart of the farm point cloud construction based on unmanned aerial vehicle aerial survey provided by the present invention.
[0022] Figure 3 It is a schematic structural diagram of the farm three-dimensional point cloud semantic segmentation model provided by the present invention.
[0023] Figure 4 It is a schematic structural diagram of the farm three-dimensional point cloud semantic segmentation device for unmanned aerial vehicle path planning provided by the present invention.
[0024] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Currently, there are many deficiencies in the field of three-dimensional point cloud semantic segmentation of farmland scenes, which are difficult to meet the requirements of high-precision semantic information for the precise navigation and path planning of unmanned agricultural machinery.
[0026] The methods based on geometric maps, topological maps, and grid maps have limitations: Although geometric maps can capture the three-dimensional structural information of the environment, they rely on simple geometric features and are difficult to reflect the detailed information in complex farmland scenes. For example, common irregular structures in farmland (such as irrigation facilities and paddy field boundaries) often cannot be accurately expressed by traditional geometric descriptions, resulting in insufficient semantic understanding ability of farmland features. In addition, these methods usually rely on artificial rules or predefined paths, with low operation efficiency and difficulty in adapting to unstructured farmland environments. Although topological maps and grid maps are applied in path planning and task allocation, they cannot effectively combine environmental semantic information, are difficult to provide refined navigation support, and have low computational efficiency in large-scale scenes.
[0027] Projection methods and voxel methods simplify the data structure by converting three-dimensional point cloud data into two-dimensional images or voxel grids. However, this conversion inevitably leads to a reduction in point cloud resolution and loss of detailed information, especially when dealing with small targets in farmland scenes (such as field roads and irrigation pipelines). In addition, the computational complexity of such methods is relatively high, making it difficult to efficiently process large-scale point cloud data.
[0028] Point-by-point processing methods such as PointNet operate directly on the original point cloud, can better retain geometric and semantic information, and are suitable for processing complex scenes. However, PointNet and subsequent improved methods (such as RandLA-Net) still face scalability problems when dealing with large-scale point clouds. For example, their ability to extract features of multi-scale targets in farmland scenes is limited, and it is difficult to balance local details and global semantic consistency, resulting in a decline in classification accuracy and segmentation effect. At the same time, these methods have weak ability to handle class imbalance, and small-class targets (such as hard ground or irrigation facilities) are often ignored due to sparse point clouds.
[0029] Generally speaking, current segmentation methods can usually only handle regular scenes such as paved roads, and it is difficult to meet the requirements of large-scale, multi-scale, and high semantic consistency in the semantic segmentation of farmland three-dimensional point clouds. These deficiencies limit the precise operation and path planning capabilities of unmanned agricultural machinery in complex farmland environments such as irregular farmland scenes like ridges, paddy fields, and irrigation facilities. There is an urgent need for a new type of point cloud segmentation technology that can efficiently extract local and global features and adapt to multi-scale targets to promote the development of agricultural automation.
[0030] Based on the above technical problems, the present invention proposes a farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning. The method obtains image data of farmland by an unmanned aerial vehicle, converts the image data into three-dimensional point cloud data, and inputs the three-dimensional point cloud data into a farm three-dimensional point cloud semantic segmentation model for classification to obtain the ground object types corresponding to different point clouds. The farm three-dimensional point cloud semantic segmentation model is trained through the following steps: obtaining a plurality of subsets of point cloud data based on the three-dimensional point cloud data, extracting the local geometric features and global semantic information of each sampling point therein, determining the predicted category of the sampling point based on the local geometric features and global semantic information, and training a deep learning model based on the predicted category and the true category to obtain the farm three-dimensional point cloud semantic segmentation model. It can efficiently extract the local geometric features and global semantic information in the farmland scene, effectively solve the deficiencies of the prior art in dealing with unstructured farmland, multi-scale targets, and semantic consistency, achieve high-precision ground object semantic segmentation, and provide technical support for the precise navigation, path planning, and task allocation of unmanned agricultural machinery. The present invention not only improves the accuracy and robustness of semantic segmentation, but also can significantly improve the classification consistency and small target recognition ability in complex farmland scenes, thus meeting the requirements of modern agricultural automation and intelligence.
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] Figure 1 is a schematic flowchart of the farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning provided by the present invention. As Figure 1 shown, the present invention provides a farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning. The method includes: Step 101, obtaining image data of farmland by an unmanned aerial vehicle, and converting the image data into three-dimensional point cloud data.
[0033] Specifically, currently, the point cloud collection and segmentation for autonomous driving are both based on vehicle-mounted devices, while in the embodiment of the present application, an unmanned aerial vehicle is used to collect point clouds to guide the autonomous driving of ground vehicles. Specifically, a high-precision imaging sensor (for example, DJI Phantom 4 RTK) can be carried by the unmanned aerial vehicle to perform full-coverage scanning on the target farmland area, obtain high-definition image data of the farmland, and identify the farmland ground objects, so as to obtain a farmland map for assisting autonomous driving and the like.
[0034] In some embodiments, the flight altitude of the drone is between 25 - 50 meters, the flight speed is 2 m / s, and the overlap rates in the vertical direction and the heading direction are both 70% to ensure complete coverage and data quality.
[0035] Then, the image data can be converted into 3D point cloud data through professional software (such as DJI Terra).
[0036] In some embodiments, after converting the image data into 3D point cloud data, a check is also performed on whether the point cloud data is qualified.
[0037] For example, Figure 2 is a schematic diagram of the farm point cloud construction process based on UAV aerial survey provided by the present invention. As Figure 2 shown, first, checkpoints are arranged in the farmland, then the UAV is used to perform photogrammetry on the farmland at each checkpoint, and point clouds are generated through in - house processing. Finally, a quality check is performed on the point cloud data. If the check is qualified, the process ends; if the check is unqualified, the UAV performs photogrammetry on the farmland again and generates point clouds until qualified point cloud data is obtained.
[0038] Step 102: Input the 3D point cloud data into the farm 3D point cloud semantic segmentation model for classification to obtain the ground object types corresponding to different point clouds.
[0039] Specifically, the pre - trained farm 3D point cloud semantic segmentation model is used to perform semantic segmentation on the 3D point cloud data, that is, to classify and identify the ground object types of the 3D point cloud data, and obtain the ground object types corresponding to different point clouds.
[0040] Among them, the farm 3D point cloud semantic segmentation model is trained through the following steps: Obtain multiple subsets of point cloud data based on the 3D point cloud data; Extract the local geometric features and global semantic information of each sampling point in the subset of point cloud data; Determine the predicted category of the sampling point based on the local geometric features and the global semantic information; Train a preset deep learning model based on the predicted category and the true category of the sampling point to obtain the farm 3D point cloud semantic segmentation model.
[0041] Specifically, the farm 3D point cloud semantic segmentation model adopts a structure of random sampling - encoding - decoding to perform multi - scale feature extraction. In the sampling module, multiple subsets of point cloud data are obtained; in the encoding module, the local geometric features and global semantic information of each sampling point in the subset of point cloud data are extracted; in the decoding module, the predicted ground object category of the sampling point is identified based on the extracted features.
[0042] In the embodiments of the present application, a custom dataset and a public dataset (such as Semantic3D) can be used for joint training to improve the generalization ability of the model.
[0043] The method for farm three-dimensional point cloud semantic segmentation for UAV path planning provided by the embodiments of the present application can efficiently extract local geometric features and global semantic information in the farmland scene, effectively solve the deficiencies of the prior art in dealing with unstructured farmland, multi-scale targets and semantic consistency, realize high-precision ground object semantic segmentation, and provide technical support for the precise navigation, path planning and task allocation of unmanned agricultural machinery. It not only improves the accuracy and robustness of semantic segmentation, but also can significantly improve the classification consistency and small target recognition ability in complex farmland scenes, so as to meet the needs of modern agricultural automation and intelligence. Moreover, it has high operational simplicity and application flexibility, adapts to the needs of various farmland scenes, and performs well in energy conservation and environmental protection, demonstrating good technical advantages and broad application prospects.
[0044] In some embodiments, obtaining multiple subsets of point cloud data based on the three-dimensional point cloud data includes: Processing the three-dimensional point cloud data based on the voxel grid downsampling algorithm to obtain sample point cloud data; Dividing the sample point cloud data into multiple point cloud data subsets of the same size.
[0045] Specifically, the three-dimensional point cloud data is processed based on the voxel grid downsampling algorithm to reduce the number of points and obtain sample point cloud data. Then, using the batch data generation algorithm, the sample point cloud data is divided into multiple point cloud data subsets of the same size to ensure that each batch of data can be evenly input into the model for training.
[0046] The method for farm three-dimensional point cloud semantic segmentation for UAV path planning provided by the embodiments of the present application reduces redundant data by performing voxel downsampling on the point cloud data, reduces the data scale while trying to retain the original geometric information, divides the point cloud into multiple point cloud data subsets of the same size, generates a standardized dataset suitable for model input, and ensures the consistency and uniformity of the data input into the model; in addition, it also reduces the time for processing large-scale point cloud data.
[0047] In some embodiments, extracting the local geometric features of each sampling point in the subset of point cloud data includes: Obtaining all neighborhood points of each sampling point based on the K-nearest neighbor algorithm; Obtaining the local spatial features of the sampling point based on the sampling point and the neighborhood points of the sampling point; the local spatial features include a distance attenuation index and a local geometric bounding box; Encode the local spatial features based on a multi-layer perceptron to obtain the local geometric features of the sampling points.
[0048] Specifically, use the K-nearest neighbor (KNN) algorithm to obtain all the neighborhood points of each sampling point. All the neighborhood points of each sampling point constitute the local spatial neighborhood of the sampling point. Extract the local spatial features of the sampling point based on the local spatial neighborhood of the sampling point. The local spatial features include the distance decay exponent and the local geometric bounding box, etc.
[0049] Then encode the local spatial features based on a multi-layer perceptron (MLP), and combine the self-attention mechanism to strengthen the representation ability of the local geometric features of the point cloud, and generate local geometric features with higher semantic understanding ability.
[0050] The farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning provided by the embodiments of the present application captures the microscopic structure information of complex ground objects by extracting the local geometric features of the point cloud, improves the recognition ability of complex ground objects, and is beneficial to improving the accuracy of ground object recognition in the farmland scene.
[0051] In some embodiments, extracting the global semantic information of each sampling point in the subset of the point cloud data includes: Calculate the volume ratio of the local neighborhood space of each sampling point to the global space; the local neighborhood space is the space where all the neighborhood points of the sampling point are located; the global space is the space corresponding to the subset of the point cloud data. Determine the global semantic information of the sampling point based on the volume ratio.
[0052] Specifically, calculate the volume ratio of the local neighborhood space of each sampling point to the global space, and determine the global semantic information of the sampling point based on the volume ratio.
[0053] Among them, the local neighborhood space is the space constituted by all the neighborhood points of the sampling point. The global space is the space corresponding to the subset of the point cloud data where the sampling point is located.
[0054] The farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning provided by the embodiments of the present application captures the global semantic information through the volume feature of the volume ratio of the local neighborhood space to the global space, improves the recognition ability of complex ground objects, and improves the accuracy of ground object recognition in the farmland scene.
[0055] In some embodiments, the determining the predicted category of the sampling point based on the local geometric features and the global semantic information includes: Fuse the local geometric features and the global semantic information to obtain the aggregated features of the sampling point; Determine the predicted category of the sampling point based on the aggregated features of the sampling point and the aggregated features of the domain points of the sampling point.
[0056] Specifically, fuse or aggregate the local geometric features of the sampling point and the global semantic information to obtain the aggregated features of the sampling point.
[0057] Then, in the point cloud classification stage, combine the neighborhood consistency constraint, aggregate the aggregated features of the sampling point and its neighborhood points, and determine the predicted category of the sampling point based on this aggregated feature and the aggregated features of the domain points of the sampling point.
[0058] The three-dimensional point cloud semantic segmentation method for farmland oriented to UAV path planning provided by the embodiments of the present application enhances the segmentation performance of the model for multi-scale targets, the recognition ability for large-scale targets, especially the accurate recognition ability for complex structures by aggregating local and global features. Classification is performed by aggregating the aggregated features of the sampling point and its neighborhood points to expand the receptive field, improve the classification accuracy and consistency. Combining residual connections can prevent feature loss and enhance the classification consistency of the model for large target areas, effectively reducing classification errors in large target areas and significantly improving the recognition ability for small category targets such as irrigation facilities and hard ground.
[0059] In some embodiments, training the preset deep learning model based on the predicted category and the true category of the sampling point includes: Determine the cross-entropy loss based on the predicted category and the true category of the sampling point; Adjust the model parameters of the deep learning model based on the cross-entropy loss.
[0060] Specifically, determine the cross-entropy loss based on the predicted category and the true category of the sampling point, and adjust the model parameters of the deep learning model based on the cross-entropy loss to obtain a trained three-dimensional point cloud semantic segmentation model for farmland.
[0061] In the embodiments of the present application, for the optimization of the model, the Adam optimizer can be used, with an initial learning rate of 0.01, and training for 100 rounds to ensure stable convergence of the model. By comparing the performance contributions of different modules, hyperparameter tuning of the model is performed, including setting the number of neighborhood points K (the optimal value is 16), so as to achieve a balance between local feature capture and noise suppression.
[0062] The three-dimensional point cloud semantic segmentation method for farmland oriented to UAV path planning provided by the embodiments of the present application can use the cross-entropy weight adjustment strategy to solve the problem of class imbalance and improve the recognition ability for small category targets (such as irrigation facilities and hard ground).
[0063] Figure 3 It is a schematic structural diagram of the three-dimensional point cloud semantic segmentation model for farmland provided by the present invention, asFigure 3 As shown in Figure 3 , in the encoding module (i.e., the encoder Encoder) of the farm three-dimensional point cloud semantic segmentation model, a random sampling module is used to perform voxel downsampling and chunking on the point cloud data to obtain subsets of point cloud data with the same size; a feature fusion module is used to aggregate the local geometric features and global features of the sampled points to obtain aggregated features. In the decoding module (i.e., the decoder Decoder) of the farm three-dimensional point cloud semantic segmentation model, a nearest neighbor search such as the K-nearest neighbor algorithm is used to obtain the local neighborhood set of the point cloud, construct the local space and extract the local space features, a multi-layer perceptron (MLP) is used to encode the local space features to obtain local geometric features, a classifier is used to aggregate the aggregated features of the sampled points and their neighboring points, and the ground object category of the sampled points is identified based on this feature. The navigation accuracy requirement for agricultural machinery autonomous driving is ±2.5 cm. Therefore, the accuracy requirements for the point cloud data and the ground object segmentation results are extremely high. The design of the model structure proposed in the embodiments of this application improves the model accuracy and robustness, and can achieve higher-precision autonomous driving navigation, etc.
[0064] In some embodiments, the method further includes: Generating a farmland semantic map based on the ground object types corresponding to different point clouds; the farmland semantic map is used to display the three-dimensional distribution information of different ground object types in the farmland.
[0065] Specifically, the point cloud data after semantic segmentation is visually displayed, and a farmland semantic map is generated according to the classification results (i.e., the ground object types corresponding to different point clouds). The farmland semantic map is used to display the three-dimensional distribution information of different ground object types in the farmland, including the three-dimensional distribution information of farmland targets such as paddy fields, entrances and exits, roads, buildings, irrigation facilities, etc.
[0066] The farm three-dimensional point cloud semantic segmentation method for UAV path planning provided by the embodiments of this application displays the semantic segmentation results and converts them into a farmland semantic map, so that it can be applied to scenarios such as path planning, navigation, and task allocation of unmanned agricultural machinery, provides reliable ground object information support for unmanned agricultural machinery, meets the precise requirements of path planning and task allocation, significantly improves the efficiency and reliability of agricultural automation, and reduces the need for manual intervention.
[0067] Figure 4 is a schematic structural diagram of a farm three-dimensional point cloud semantic segmentation device for UAV path planning provided by the present invention. As Figure 4 shown, the present invention provides a farm three-dimensional point cloud semantic segmentation device for UAV path planning, including an acquisition module 401 and a semantic segmentation module 402.
[0068] The acquisition module 401 is used to acquire the image data of the farmland by means of a UAV and convert the image data into three-dimensional point cloud data.
[0069] The semantic segmentation module 402 is used to input the three-dimensional point cloud data into the farm three-dimensional point cloud semantic segmentation model for classification to obtain the ground object types corresponding to different point clouds; the farm three-dimensional point cloud semantic segmentation model is trained through the following steps: Obtain multiple subsets of point cloud data based on the three-dimensional point cloud data; Extract the local geometric features and global semantic information of each sampling point in the subset of point cloud data; Determine the predicted category of the sampling point based on the local geometric features and the global semantic information; Train a preset deep learning model based on the predicted category and the true category of the sampling point to obtain the farm three-dimensional point cloud semantic segmentation model.
[0070] In some embodiments, the obtaining multiple subsets of point cloud data based on the three-dimensional point cloud data includes: Process the three-dimensional point cloud data based on the voxel grid downsampling algorithm to obtain the sample point cloud data; Divide the sample point cloud data into multiple point cloud data subsets of the same size.
[0071] In some embodiments, the extracting the local geometric features of each sampling point in the subset of point cloud data includes: Obtain all the neighborhood points of each sampling point based on the K-nearest neighbor algorithm; Obtain the local spatial features of the sampling point based on the sampling point and the neighborhood points of the sampling point; the local spatial features include the distance attenuation index and the local geometric bounding box; Encode the local spatial features based on a multi-layer perceptron to obtain the local geometric features of the sampling point.
[0072] In some embodiments, the extracting the global semantic information of each sampling point in the subset of point cloud data includes: Calculate the volume ratio of the local neighborhood space of each sampling point to the global space; the local neighborhood space is the space where all the neighborhood points of the sampling point are located; the global space is the space corresponding to the subset of point cloud data; Determine the global semantic information of the sampling point based on the volume ratio.
[0073] In some embodiments, the determining the predicted category of the sampling point based on the local geometric features and the global semantic information includes: Fuse the local geometric features and the global semantic information to obtain the aggregated features of the sampling point; Determine the predicted category of the sampling point based on the aggregated features of the sampling point and the aggregated features of the neighborhood points of the sampling point.
[0074] In some embodiments, training a preset deep learning model based on the predicted category and the true category of the sampling point includes: Determining a cross-entropy loss based on the predicted category and the true category of the sampling point; Adjusting the model parameters of the deep learning model based on the cross-entropy loss.
[0075] In some embodiments, it further includes: A generation module, configured to generate a farm semantic map based on the ground object types corresponding to different point clouds; the farm semantic map is used to display the three-dimensional distribution information of different ground object types in the farm.
[0076] Specifically, the above-mentioned farm three-dimensional point cloud semantic segmentation device for unmanned aerial vehicle path planning provided by the present invention can implement all the method steps implemented by the above-mentioned method embodiments of the farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described herein.
[0077] It should be noted that the division of units / modules in the above-mentioned embodiments of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0078] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 5 shown, the electronic device may include: a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 complete mutual communication through the communication bus 504. The processor 501 can call the logical instructions in the memory 503 to execute the above-mentioned method for farm three-dimensional point cloud semantic segmentation for unmanned aerial vehicle path planning provided by each method embodiment. The method includes: Obtaining image data of a farm by an unmanned aerial vehicle and converting the image data into three-dimensional point cloud data; Inputting the three-dimensional point cloud data into a farm three-dimensional point cloud semantic segmentation model for classification to obtain the ground object types corresponding to different point clouds; the farm three-dimensional point cloud semantic segmentation model is trained through the following steps: Obtain multiple subsets of point cloud data based on three-dimensional point cloud data; Extract the local geometric features and global semantic information of each sampling point in the subset of point cloud data; Determine the predicted category of the sampling point based on the local geometric features and the global semantic information; Train a preset deep learning model based on the predicted category and the true category of the sampling point to obtain a three-dimensional point cloud semantic segmentation model for the farm.
[0079] In some embodiments, the obtaining multiple subsets of point cloud data based on three-dimensional point cloud data includes: Process the three-dimensional point cloud data based on the voxel grid downsampling algorithm to obtain sample point cloud data; Divide the sample point cloud data into multiple point cloud data subsets of the same size.
[0080] In some embodiments, the extracting the local geometric features of each sampling point in the subset of point cloud data includes: Obtain all the neighborhood points of each sampling point based on the K-nearest neighbor algorithm; Obtain the local spatial features of the sampling point based on the sampling point and its neighborhood points; the local spatial features include the distance attenuation exponent and the local geometric bounding box; Encode the local spatial features based on a multi-layer perceptron to obtain the local geometric features of the sampling point.
[0081] In some embodiments, the extracting the global semantic information of each sampling point in the subset of point cloud data includes: Calculate the volume ratio of the local neighborhood space of each sampling point to the global space; the local neighborhood space is the space where all the neighborhood points of the sampling point are located; the global space is the space corresponding to the subset of point cloud data; Determine the global semantic information of the sampling point based on the volume ratio.
[0082] In some embodiments, the determining the predicted category of the sampling point based on the local geometric features and the global semantic information includes: Fuse the local geometric features and the global semantic information to obtain the aggregated features of the sampling point; Determine the predicted category of the sampling point based on the aggregated features of the sampling point and the aggregated features of the domain points of the sampling point.
[0083] In some embodiments, the training a preset deep learning model based on the predicted category and the true category of the sampling point includes: Determine the cross-entropy loss based on the predicted category and the true category of the sampling point; Adjust the model parameters of the deep learning model based on the cross-entropy loss.
[0084] In some embodiments, the method further includes: Generate a farm semantic map based on the ground object types corresponding to different point clouds; the farm semantic map is used to display the three-dimensional distribution information of different ground object types in the farm.
[0085] Specifically, the processor 501 can be a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or a Complex Programmable Logic Device (CPLD). The processor can also adopt a multi-core architecture.
[0086] When the logical instructions in the memory 503 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disks, or optical discs that can store program codes.
[0087] In some embodiments, a computer program product is further provided. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning provided in the above method embodiments.
[0088] Specifically, the computer program product provided in the embodiments of the present application can implement all the method steps implemented in the above method embodiments, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described herein again.
[0089] In some embodiments, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and the computer program is used to cause a computer to execute the farm three-dimensional point cloud semantic segmentation method for unmanned aerial vehicle path planning provided in the above method embodiments.
[0090] Specifically, the computer-readable storage medium provided by the present invention can implement all the method steps implemented in the above method embodiments, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described herein again.
[0091] It should be noted that: the computer-readable storage medium can be any available medium or data storage device accessible by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSDs)), etc.
[0092] The term "a plurality of" in the present invention refers to two or more, and other quantifiers are similar thereto.
[0093] "Determining B based on A" in the present invention means that the factor A should be considered when determining B. It is not limited to "determining B only based on A", but should also include: "determining B based on A and C", "determining B based on A, C, and E", "determining C based on A, and further determining B based on C", etc. Additionally, it can also include using A as a condition for determining B. For example, "when A meets the first condition, use the first method to determine B"; for another example, "when A meets the second condition, determine B"; for yet another example, "when A meets the third condition, determine B based on the first parameter", etc. Of course, it can also be a condition for using A as a factor for determining B. For example, "when A meets the first condition, use the first method to determine C, and further determine B based on C", etc.
[0094] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0095] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer-executable instructions. These computer-executable instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0096] These processor-executable instructions can also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the processor-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0097] These processor-executable instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0098] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A farm 3D point cloud semantic segmentation method for UAV path planning, characterized by: include: Acquire image data of farmland through a drone, and convert the image data into three-dimensional point cloud data; The three-dimensional point cloud data is input into the farm three-dimensional point cloud semantic segmentation model for classification to obtain the types of objects corresponding to different point clouds; the farm three-dimensional point cloud semantic segmentation model is trained by the following steps: Acquire multiple point cloud data subsets based on three-dimensional point cloud data; Extracting local geometric features and global semantic information of each sampling point in the point cloud data subset; Determining a prediction category of the sampling point based on the local geometric features and the global semantic information; A preset deep learning model is trained based on the predicted categories and the real categories of the sampling points to obtain a semantic segmentation model of the farm three-dimensional point cloud.
2. The farm 3D point cloud semantic segmentation method for UAV path planning according to claim 1 is characterized in that: The step of acquiring a plurality of point cloud data subsets based on the three-dimensional point cloud data comprises: Processing the three-dimensional point cloud data based on a voxel grid downsampling algorithm to obtain sample point cloud data; The sample point cloud data is divided into a plurality of point cloud data subsets of the same size.
3. The farm 3D point cloud semantic segmentation method for UAV path planning according to claim 1 is characterized in that: The extracting of local geometric features of each sampling point in the point cloud data subset comprises: Obtain all neighboring points of each sampling point based on the K-nearest neighbor algorithm; Acquire the local spatial features of the sampling point based on the sampling point and the neighborhood points of the sampling point; the local spatial features include a distance decay index and a local geometric bounding box; The local spatial features are encoded based on a multi-layer perceptron to obtain local geometric features of the sampling points.
4. The farm 3D point cloud semantic segmentation method for UAV path planning according to claim 1 is characterized in that: Extracting global semantic information of each sampling point in the point cloud data subset includes: Calculate the volume ratio of the local neighborhood space of each sampling point to the global space; the local neighborhood space is the space where all neighborhood points of the sampling point are located; the global space is the space corresponding to the point cloud data subset; Global semantic information of the sampling point is determined based on the volume ratio.
5. The farm 3D point cloud semantic segmentation method for UAV path planning according to claim 1 is characterized in that: The determining the prediction category of the sampling point based on the local geometric features and the global semantic information includes: Performing feature fusion on the local geometric features and the global semantic information to obtain aggregate features of the sampling points; Based on the aggregate features of the sampling points and the aggregate features of the domain points of the sampling points, a predicted category of the sampling points is determined.
6. The farm 3D point cloud semantic segmentation method for UAV path planning according to claim 1 is characterized in that: The preset deep learning model is trained based on the predicted category and the real category of the sampling point, comprising: Determine a cross entropy loss based on the predicted category and the true category of the sampling point; Model parameters of the deep learning model are adjusted based on the cross entropy loss.
7. The farm 3D point cloud semantic segmentation method for UAV path planning according to claim 1 is characterized in that: The method further comprises: A farmland semantic map is generated based on the types of land objects corresponding to different point clouds; the farmland semantic map is used to display the three-dimensional distribution information of different types of land objects in the farmland.
8. A farm 3D point cloud semantic segmentation device for drone path planning, characterized in that: include: An acquisition module, used to acquire image data of farmland through a drone and convert the image data into three-dimensional point cloud data; The semantic segmentation module is used to input the 3D point cloud data into the farm 3D point cloud semantic segmentation model for classification to obtain the types of objects corresponding to different point clouds; the farm 3D point cloud semantic segmentation model is trained by the following steps: Acquire multiple point cloud data subsets based on three-dimensional point cloud data; Extracting local geometric features and global semantic information of each sampling point in the point cloud data subset; Determining a prediction category of the sampling point based on the local geometric features and the global semantic information; A preset deep learning model is trained based on the predicted categories and the real categories of the sampling points to obtain a semantic segmentation model of the farm three-dimensional point cloud.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the farm three-dimensional point cloud semantic segmentation method for drone path planning as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements the farm three-dimensional point cloud semantic segmentation method for drone path planning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Top coal caving quantity measuring system based on laser radar and coal caving control method
CN113267124A
Real-time track obstacle detection method based on three-dimensional point cloud
CN113378647A
Complementary pseudo-multi-mode characteristic system and method
CN116109777A
Three-dimensional target detection method based on class enhancement and geometric enhancement
CN117315646A
Farmland feature semantic segmentation method based on laser radar and SegNet
CN119360384A