Point cloud semantic segmentation method for unstructured scenes in natural orchards
By combining the local feature geometric differentiation scaling strategy and the global feature attention mechanism, the problems of local feature extraction difficulties and category imbalance in natural orchard environments are solved, and the accuracy and performance of point cloud semantic segmentation are significantly improved.
Patent Information
- Application Number
- CN202411518656.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-10-29
AI Technical Summary
The prior art is difficult to effectively extract local features in natural orchard environments, and due to category imbalance, the segmentation accuracy is low, making it difficult to meet the needs of complex agricultural scenarios.
The point cloud semantic segmentation method combining the local feature geometric differentiation scaling strategy GDS and the global feature attention mechanism GFM is adopted. The local feature aggregation module LFA and the global feature mapping module GFM are enhanced, and the class imbalance problem is alleviated through the improved composite loss function ωCEloss.
It significantly improves the accuracy of point cloud semantic segmentation in natural orchard scenes, enhances the model's understanding of local and global features, slows down training degradation and potential errors, and achieves higher semantic segmentation performance.
Smart Images

Figure CN119418054B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and in particular relates to a point cloud semantic segmentation method for unstructured scenes in natural orchards. Background Art
[0002] With the rapid development of intelligent agriculture, the role of fruit recognition and segmentation in agricultural production has become increasingly prominent. At present, fruit recognition and segmentation methods have been widely used in robot picking, orchard yield prediction and other fields. The versatility of computer vision makes it a technical tool suitable for many fields, including precision agriculture. In the field of precision agriculture and smart farms, deep learning technology can more effectively solve problems such as lack of robustness and generalization compared with other machine vision technologies. Semantic segmentation is a classic research field in computer vision. It is of great significance to accurately understand information at the pixel or voxel level and convert it into a region of interest with highlighted display. In the field of fruit planting in agricultural scenes, accurately parsing semantic information such as fruits, branches and leaves plays a very important role in tasks such as fruit growth detection, phenotype extraction, automatic picking and yield estimation.
[0003] Two-dimensional semantic segmentation has been widely used in fruit identification and phenotyping tasks. Many methods identify apple fruits by directly processing two-dimensional images. Although these methods have achieved good overall performance in terms of recognition accuracy, their semantic segmentation average intersection-over-union ratio is relatively poor. The reason for this phenomenon is that environmental factors such as light and weather can cause interference, and fruits are easily occluded in agricultural natural orchard environments. Due to the lack of depth information, the three-dimensional space cannot be well described. For highly unstructured agricultural scenes, the occlusion of objects such as branches and leaves and the cluttered background will cause a significant decrease in accuracy. Therefore, conventional two-dimensional semantic segmentation methods can no longer meet the needs of many scenes, and researchers have turned to the semantic segmentation of three-dimensional point clouds. However, unlike urban street scenes or indoor scenes, the unstructured and unevenly distributed three-dimensional point clouds of natural orchards in agricultural scenes are particularly prominent. It is still very challenging to accurately and efficiently semantically segment three-dimensional point clouds in agricultural scenes. Existing methods have problems such as ineffective local feature extraction and insufficient attention to important information. Summary of the invention
[0004] In view of the technical problems in the above-mentioned prior art such as lack of depth information, low segmentation accuracy, and category imbalance in the natural orchard environment, the present invention provides a point cloud semantic segmentation method for unstructured scenes in natural orchards. By combining advanced strategies such as local differentiation scaling strategy GDS and attention mechanism mapping GFM, the segmentation accuracy of the model in the natural orchard scene is improved.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] The point cloud semantic segmentation method for unstructured scenes in natural orchards includes the following steps:
[0007] S1. Effectively process the natural orchard 3D point cloud data through spatial voxelization grid subdivision, data enhancement and spatial sampling, and organize and construct the natural orchard point cloud dataset;
[0008] S2. Based on the Encoder architecture, a local feature aggregation module LFA is designed to effectively geometrically differentiate and scale the local neighborhood features GDS, balance the feature differences and scaling scales of different neighborhood groups, enhance the feature extraction capability of the network, and use MLP and Pooling for efficient feature aggregation;
[0009] S3. Based on the decoder architecture, a global feature mapping module (GFM) is designed to explore the semantic dependencies of different spatial locations, fully extract the contextual semantic relationship between global and local features, and enhance the network's perception of global spatial locations.
[0010] S4. The improved composite loss function is used to solve the category imbalance problem of 3D point clouds in natural orchards, and to mitigate the degradation and potential errors of model training.
[0011] S5. Input the natural orchard 3D point cloud data into the model network to obtain the training results, and establish evaluation indicators for performance evaluation of semantic segmentation accuracy, and compare the results with other SOTA models.
[0012] The dataset in S1 is: a public natural orchard 3D point cloud dataset PFuji-Size, which consists of a set of 3D point clouds of Fuji apple trees; 3D point clouds of 6 complete Fuji apple trees are generated using motion structure and multi-view stereo technology, containing a total of 615 apples; these point clouds are scanned at different stages of maturity, with two different modes of XYZ and RGB, and also contain corresponding spatial position annotation information and fruit diameter information for apple fruit detection and size estimation.
[0013] The data processing method in S1 is: spatial voxelization grid subdivision, data enhancement and spatial sampling. The data enhancement includes data enhancement and expansion operations such as random translation, random scaling, adding Gaussian noise, random point removal, random color inactivation, etc. to simulate the impact of environmental factors such as weather, lighting, and occlusion on model performance.
[0014] The specific construction method of the local feature aggregation module LFA in S2 is: based on the Encoder architecture, the KNN algorithm is used to group the input features into local neighborhoods; for local neighborhood features, the geometric differentiation scaling GDS strategy is used to balance the feature differences of different local neighborhoods; and local features are extracted and aggregated through efficient MLP and Pooling operations.
[0015] The specific construction method of the global feature mapping module GFM in S3 is: based on the Decoder decoder architecture, the global feature map is extracted through MLP; and the semantic importance and spatial position distribution dependency of the three-dimensional point cloud are mapped through the efficient PESA attention mechanism.
[0016] The composite loss function in S4 is: the initial CEloss is improved to ωCEloss, different weights are assigned to apple fruits and branches and leaves, the problem of category imbalance is alleviated, the model training degradation is prevented, and the potential errors of the model are reduced. The composite loss function formula is:
[0017]
[0018] Among them, ωCEloss represents the composite loss function, C represents different categories, such as apple fruit and branch background area; N represents the number of points; ω i is the weight factor for each category; y ij represents the true label of the jth point in the i-th category; P ij is the probability that the model predicts that the jth point belongs to the i-th class.
[0019] The model training method in S5 is as follows: initialize the trainable parameters in the model network and input the training data into the network in batches; construct a loss function based on the predicted labels and the true labels and calculate the loss in the iterative process, use the optimization algorithm to back-propagate and update the network parameters until the loss no longer decreases within a certain range, at which time the network parameters are saved as the final model; then use the same training strategy and the same data set to train other SOTA models and compare them based on the evaluation indicators.
[0020] The model evaluation indicators in S5 are: mAcc average accuracy, mIoU average intersection over union, and the formulas are as follows:
[0021]
[0022] Among them, mAcc average accuracy, mIoU average intersection over union, C represents the number of categories in the semantic scene segmentation task, n i Indicates the number of correctly segmented points in this category, N i Indicates all the numbers under this category; Y irepresents the true set of points of the i-th category, B i Represents the set of points that the model predicts for the i-th category.
[0023] Other SOTA models in S5 include: DGCNN, PointNet++ (SSG), PointNet++ (MSG), PointTransformer and PointMLP.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] The present invention constructs the OrchardNet natural orchard semantic segmentation model by combining advanced structures such as local feature geometric differentiation scaling (GDS) and global feature attention mechanism (GFM), thereby improving the segmentation accuracy in complex natural orchard environments; and alleviates the category imbalance problem through an improved loss function, slowing down training degradation and reducing potential model errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0027] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.
[0028] Figure 1 This is a flow chart for acquiring and processing natural orchard data used in the present invention.
[0029] Figure 2 This is a model structure diagram of the point cloud semantic segmentation method for unstructured scenes in natural orchards according to the present invention.
[0030] Figure 3 Schematic diagram of the local feature extraction module LFA in the model constructed by the present invention.
[0031] Figure 4 Schematic diagram of the global feature mapping module GFM in the model constructed by the present invention.
[0032] Figure 5This is an example diagram of the semantic segmentation results of the natural orchard scene point cloud of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than to limit the claims of the present invention. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0034] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0035] like Figure 1-5 As shown, this embodiment is implemented under the pytorch deep learning framework. This embodiment proposes a point cloud semantic segmentation method for unstructured scenes in natural orchards. By combining the local feature extraction module LFA with local feature geometric differential scaling and the global feature mapping module GFM with the global feature attention mechanism, the model's ability to understand local neighborhood features is enhanced, the spatial dependency of global features is obtained, and the semantic segmentation accuracy in natural orchard scenes is improved; and the category imbalance problem is alleviated through the improved composite loss function ωCEloss.
[0036] Step 1: Data preparation
[0037] Natural orchard 3D point cloud dataset: PFuji-Size, a public natural orchard 3D point cloud dataset, consists of a collection of 3D point clouds of Fuji apple trees. 3D point clouds of 6 complete Fuji apple trees are generated using motion structure and multi-view stereo technology, containing a total of 615 apples. These point clouds are scanned at different stages of maturity, with two different modalities, XYZ and RGB, and also contain corresponding spatial position annotation information and fruit diameter information for apple fruit detection and size estimation.
[0038] The above datasets are preprocessed to convert the different annotation information of the detection dataset into the format data required for training. Then the data space is voxelized and gridded, data is enhanced, and spatial sampling is performed. The processing flow is as follows: Figure 1The data enhancement includes data enhancement and expansion operations such as random translation, random scaling, adding Gaussian noise, random point removal, and random color inactivation to simulate the impact of environmental factors such as weather, lighting, and occlusion on model performance. Spatial sampling is performed, and then the data set is divided into training set, validation set, and test set in a ratio of 8:1:1 for subsequent model training.
[0039] Among them, the data space segmentation strategy will greatly affect the model performance. This embodiment conducted an exploratory experiment, and the results are shown in Table 1.
[0040] Table 1 Data space segmentation strategy experimental results
[0041] test Size Rate Time mAcc QUR 1 0.3m 24.2% 22.6s 94.4% 87.1% 2 0.4m 28.6% 15.8s 91.9% 84.1% 3 0.5m 32.0% 11.7s 89.9% 82.3%
[0042] The Size and Rate in Table 1 refer to the scale of the Voxel Grid spatial subdivision algorithm and the ratio of the subdivision subblocks involved in model training in all subdivision subblocks, respectively. Time refers to the time it takes for the benchmark model to infer all subdivision scene subblocks. From the comparison of tests 1-3, it can be seen that a smaller Size will segment out scene subdivision subblocks with smaller scales and clearer semantic details, so the model segmentation accuracy is higher. However, the ratio of its subdivided scene subblocks participating in model training evaluation is also smaller (because the finer division results in the absence or few positive sample points in most subdivision subblocks), and detailed reasoning will increase the total processing time. While ensuring the accuracy of model segmentation, in order to make full use of data and take into account processing efficiency, this paper uses Size = 0.4m as a typical value with relatively balanced indicators.
[0043] Step 2: Building an Apple Semantic Segmentation Model
[0044] The constructed OrchardNet model is based on the 3D semantic segmentation network PointNet++. The specific network model structure is as follows: Figure 2 As shown in the figure. PointNet++ is based on an encoder-decoder structure. The encoder network consists of a point set abstraction layer (SA), which realizes the extraction of local features of multi-scale point clouds based on downsampling operations. SA Block includes a sampling layer and a feature extraction layer. The sampling layer uniformly samples a fixed number of points in the input point set by iterating the farthest point sampling FPS. The feature extraction layer first divides the local neighborhood; then maps and extracts the local neighborhood features. The model uses a hierarchical structure. The number of sampling points (N) and sampling radius (R) of each layer of SA Block are different. As the level deepens, the number of sampling points and sampling radius will increase accordingly. Different extraction scales can extract multi-scale features.
[0045] In order to enhance the feature extraction capability of the original PointNet++ model, this embodiment adds the local feature extraction module in the PointNet++ model network to the geometric differentiation scaling strategy GDS. GDS improves the consistency of features in different local areas by balancing the feature differences and scale differences in different local neighborhoods, and then extracts local features through efficient MLP and Pooling. Figure 3 As shown in Figure 1. In the PointNet++ model network, the global feature mapping module adds a global attention mechanism PESA, extracts features through MLP, and the PESA attention mechanism is similar to the position encoding-based attention mechanism of Point Transformer. It incorporates spatial position encoding and attention mechanism, and can map the semantic importance and spatial position distribution dependency of three-dimensional point clouds. Figure 4 shown.
[0046] Step 3: Model feature extraction and mapping strategy verification
[0047] This embodiment fully experimentally verifies the geometric differential scaling GDS in the proposed local feature aggregation module LFA. The results are shown in Table 2. Test 1 refers to direct feature extraction of multimodal information without scaling; GDS (SN) refers to differential adjustment of the local receptive field of a single neighborhood, and GDS (AN) refers to differential adjustment of the local receptive field of all neighborhoods for the xyz position offset, taking into account the feature differences of different local neighborhood groups. GDS (R) refers to the scale scaling of the local receptive field query radius for the xyz position offset; GDS (ALL) refers to describing the feature differences and scale differences of all local neighborhood groups.
[0048] Table 2 Local feature extraction LFA experimental results
[0049] test LFA mAcc QUR 1 XYZ RGB 91.9% 84.1% 2 GDS(SN) 92.4% 84.9% 3 GDS(AN) 93.4% 86.5% 4 GDS(R) 93.3% 85.9% 5 GDS(ALL) 93.5% 86.8%
[0050] GDS(ALL) achieves the best results, with mAcc reaching 93.5% and mIoU reaching 86.8%. It can accelerate model convergence and improve semantic segmentation accuracy. Its differentiated adjustment balances the feature differences of all local neighborhood groups, and the scaling of the query radius alleviates the optimization difficulties caused by the absolute offset.
[0051] This embodiment fully experimentally verifies the global attention mechanism in the proposed global feature mapping module GFM. The results are shown in Table 3.
[0052] Table 3 Global feature map GFM experimental results
[0053] test GFM mAcc QUR 1 MLPs 91.9% 84.1% 2 SA 91.8% 84.6% 3 CA 91.7% 84.2% 4 CCI 92.3% 84.3% 5 ECA 92.3% 84.7% 6 SE 92.3% 84.9% 7 PESA (k=16) 92.6% 85.3% 8 PESA (k=64) 92.7% 85.7%
[0054] Among them, MLPs represents the basic feature mapping operation; SA represents the addition of self-attention mechanism mapping to feature channels; CA represents the addition of attention weighting to different feature channels; CCI represents the cross-channel interaction adjustment of features; ECA represents ECA-Net attention mechanism mapping; SE represents SE-Net attention mechanism mapping; PESA represents self-attention mapping of features through the position encoding self-attention module (Transformer Block), and k refers to the KNN neighborhood range.
[0055] The results in Table 3 show that the attention mechanisms based on channel adjustment such as SA and ECA are generally ineffective, and they cannot fully understand the dependencies between different spatial position coordinates. The PESA module can achieve better results. PESA models the spatial position semantic importance of global features, increases the ability of relative position information, and can learn a universal and consistent global feature representation to capture the long-range dependencies of global semantic position importance. In the best case, mAcc reaches 92.7% and mIoU reaches 85.7%.
[0056] Step 4: Model overall performance evaluation
[0057] Initialize the trainable parameters in the network and input the training set into the network in batches; construct a composite loss function and calculate the model loss based on the predicted values and true labels, and use the optimization algorithm to backpropagate and update the network parameters until the loss no longer decreases within a certain range. At this time, save the network parameters as the final model.
[0058] In order to verify the optimization effect of the two newly added modules on the model, this embodiment conducted an ablation experiment, and the results are shown in Table 4.
[0059] Table 4. Ablation experiment results of semantic segmentation model
[0060] test LFA GFM mAcc QUR 1 91.9% 84.1% 2 √ 93.5% 86.8% 3 √ 92.7% 85.7% 4 √ √ 93.7% 87.7%
[0061] As can be seen from Table 4, in Test 2, after the local feature extraction module was improved (LFA), the mAc value of the model increased by 1.6% and the mIoU increased by 2.7%; in Test 3, after the global feature mapping module was improved (GFM), the mAc value increased by 0.8% and the mIoU increased by 1.6%; in Test 4, after improving LFA and GFM at the same time, mAc increased to 93.7% and mIoU increased to 87.7%. In summary, after adding LFA or GFM to the baseline model, the semantic segmentation performance is significantly improved. In this embodiment, the model of Test 4 is named OrchardNet. Compared with the original PointNet++ model, its semantic segmentation performance is significantly improved, which lays a better foundation for apple picking and fruit phenotype extraction.
[0062] Then, we use the same dataset and the same training strategy to train other SOTA models (including DGCNN, PointNet++ (SSG), PointNet++ (MSG), Point Transformer, and PointMLP), and compare the parameter results of these models. The results are shown in Table 5.
[0063] Table 5 Performance comparison of OrchardNet model with other SOTA models
[0064] Model QUR PointNet++(SSG) 81.5% PointNet++(MSG) 84.1% DGCNN 79.0% Point Transformer 78.1% PointMLP 85.4% OrchardNet 87.7%
[0065] As can be seen from Table 5, the OrchardNet proposed in this embodiment enhances the ability to correctly segment apple semantics in a complex natural orchard environment, laying a better foundation for apple picking and fruit phenotype extraction. Figure 5 shown.
[0066] Step 5: Alleviate the class imbalance problem
[0067] The positive and negative ratios of the three-dimensional point cloud in the natural orchard are quite different, and the class imbalance problem is more prominent. In order to prevent the model training from degrading and reduce the potential deviation of the model, this embodiment improves the loss function of the model, improves the model performance and ensures the segmentation accuracy through the improved loss function.
[0068] This example uses the ωCEloss weighted loss function to alleviate the class imbalance problem. In classification or segmentation problems, ωCEloss can quantify the difference between probability distributions of different categories through class weight factors, and is relatively stable. ωCEloss is defined as follows:
[0069]
[0070] Among them, C represents different categories, such as apple fruit and branch background area; N represents the number of points; ω i is the weight factor for each category; y ij represents the true label of the jth point in the i-th category; P ij is the probability that the model predicts that the jth point belongs to the ith class. Among them, CE represents the cross entropy loss function, CE+DL represents the composite form of the cross entropy loss function and the Dice loss function, ωCE represents the weighted cross entropy loss function, and the weight ratio refers to the weighted weight ratio of negative samples and positive samples. This embodiment conducts loss function related experiments on the benchmark model PointNet++, and the results are shown in Table 6.
[0071] Table 6 Loss function experimental results
[0072] test Loss Function Weight Ratio mAcc QUR 1 CE — 91.9% 84.1% 2 CE+DL — 92.7% 84.4% 3 ωCE 1.0:1.25 93.5% 84.5% 4 ωCE 1.0:1.5 94.5% 84.4% 5 ωCE 1.0:1.75 94.9% 84.0% 6 ωCE 1.0:2.0 95.5% 84.0%
[0073] The results of tests 1-6 show that CE loss and CE+DL composite loss perform poorly; on the basis of cross entropy loss CE, the category weight can better alleviate the category imbalance problem of the natural orchard 3D point cloud scene, improve the model segmentation accuracy, and accelerate the model convergence speed. As shown in test2-test5, mAcc has been improved, and as the category weight ratio increases, the mAcc value shows an upward trend.
[0074] The point cloud semantic segmentation method for the natural orchard unstructured scene proposed in this embodiment enhances the model's ability to understand local neighborhood features, obtains the spatial dependency of global features, and improves the semantic segmentation accuracy in the natural orchard scene by combining the local feature extraction module LFA with the local feature geometric differential scaling and the global feature mapping module GFM with the global feature attention mechanism; and alleviates the category imbalance problem through the improved composite loss function ωCEloss. It can be seen from the results that compared with other SOTA models, the semantic segmentation method proposed in this embodiment has a higher accuracy, laying a solid foundation for fruit picking and fruit phenotype extraction.
[0075] Only the preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention, and various changes should be included in the protection scope of the present invention.
Claims
1. A point cloud semantic segmentation method for unstructured scenes in natural orchards, characterized by: The following steps are involved: S1. Effectively process the 3D point cloud data of natural orchards through spatial voxelization grid subdivision, data enhancement and spatial sampling, and organize and construct a natural orchard point cloud dataset; the dataset is: a public natural orchard 3D point cloud dataset PFuji-Size, which consists of a collection of 3D point clouds of Fuji apple trees; 3D point clouds of 6 complete Fuji apple trees are generated using motion structure and multi-view stereo technology, containing a total of 615 apples; these point clouds are scanned at different maturity stages, with two different modalities of XYZ and RGB, and also contain corresponding spatial position annotation information and fruit diameter information, which are used for apple fruit detection and size estimation; S2. Based on the Encoder architecture, a local feature aggregation module LFA is designed to effectively geometrically differentiate and scale the local neighborhood features GDS, balance the feature differences and scaling scales of different neighborhood groups, enhance the feature extraction capability of the network, and use MLP and Pooling for efficient feature aggregation; S3. Based on the Decoder architecture, a global feature mapping module GFM is designed to explore the semantic dependencies of points at different spatial locations, fully extract the contextual semantic relationships between global and local features, and enhance the network's perception of global spatial locations. The specific construction method of the global feature mapping module GFM is as follows: based on the Decoder architecture, the global feature map is extracted through MLP; and the semantic importance of the three-dimensional point cloud and the spatial location distribution dependency are mapped through the efficient PESA attention mechanism. S4. The improved composite loss function is used to solve the category imbalance problem of 3D point clouds in natural orchards, and to mitigate the degradation and potential errors of model training. S5. Input the natural orchard 3D point cloud data into the model network to obtain the training results, and establish evaluation indicators for performance evaluation of semantic segmentation accuracy, and compare the results with other SOTA models.
2. The point cloud semantic segmentation method for natural orchard unstructured scenes according to claim 1 is characterized in that: The data processing method in S1 is: spatial voxelization grid subdivision, data enhancement and spatial sampling. The data enhancement includes random translation, random scaling, adding Gaussian noise, random point removal, and random color inactivation data enhancement and expansion operations to simulate the impact of weather, lighting, and occlusion environmental factors on model performance.
3. The point cloud semantic segmentation method for natural orchard unstructured scenes according to claim 1 is characterized in that: The specific construction method of the local feature aggregation module LFA in S2 is: based on the Encoder architecture, the KNN algorithm is used to group the input features into local neighborhoods; for local neighborhood features, the geometric differentiation scaling GDS strategy is used to balance the feature differences of different local neighborhoods; and local features are extracted and aggregated through efficient MLP and Pooling operations.
4. The point cloud semantic segmentation method for natural orchard unstructured scenes according to claim 1 is characterized in that: The composite loss function in S4 is: the initial CEloss is improved to ωCEloss, different weights are assigned to apple fruits and branches and leaves, the problem of category imbalance is alleviated, the model training degradation is prevented, and the potential errors of the model are reduced. The composite loss function formula is: Among them, ωCEloss represents the composite loss function, C represents different categories, apple fruit and branch background area; N represents the number of points; ω i is the weight factor for each category; y ij represents the true label of the jth point in the i-th category; P ij is the probability that the model predicts that the jth point belongs to the i-th class.
5. The point cloud semantic segmentation method for natural orchard unstructured scenes according to claim 1 is characterized in that: The model training method in S5 is as follows: initialize the trainable parameters in the model network and input the training data into the network in batches; construct a loss function based on the predicted labels and the true labels and calculate the loss in the iterative process, use the optimization algorithm to back-propagate and update the network parameters until the loss no longer decreases within a certain range, at which time the network parameters are saved as the final model; then use the same training strategy and the same data set to train other SOTA models and compare them based on the evaluation indicators.
6. The point cloud semantic segmentation method for natural orchard unstructured scenes according to claim 1 is characterized in that: The model evaluation indicators in S5 are: mAcc average accuracy, mIoU average intersection over union, and the formulas are as follows: Among them, mAcc average accuracy, mIoU average intersection over union, C represents the number of categories in the semantic scene segmentation task, n i Indicates the number of correctly segmented points in this category, N i Indicates all the numbers under this category; Y i represents the true set of points of the i-th category, B i Represents the set of points that the model predicts for the i-th category.
7. The point cloud semantic segmentation method for natural orchard unstructured scenes according to claim 1 is characterized in that: Other SOTA models in S5 include: DGCNN, PointNet++ (SSG), PointNet++ (MSG), PointTransformer and PointMLP.
Citation Information
Patent Citations
Three-dimensional point cloud semantic segmentation method based on deep learning
CN111489358A
Urban street point cloud semantic segmentation method based on self-attention global feature enhancement
CN115147601A