Deep Learning-Based Point Cloud Segmentation Method

Through the point cloud segmentation method based on deep learning, the problem of point cloud segmentation difficulty in complex scenarios is solved, and the precise segmentation and material classification of material piles in the material field are realized, supporting intelligent management and efficient transportation of the material field.

CN119832013BActive Publication Date: 2025-07-22CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510300445.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-22
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing point cloud segmentation algorithm has poor segmentation effect in complex scenarios, and cannot handle semantic segmentation, instance segmentation and panoramic segmentation at the same time, and is not suitable for environments such as material fields, resulting in large bulk equipment not being able to identify material piles and being unable to realize automatic material collection and intelligent management of material field data.

Method used

A point cloud segmentation method based on deep learning is designed. By acquiring material field point cloud data, denoising and annotating, a local neighborhood graph is constructed, curvature features are introduced, point features are extracted using 3D sparse convolution, and super points are generated. Transformer decoder is used for multi-scale feature interaction, combining multi-task loss function and de-entanglement matching strategy to achieve end-to-end training.

Benefits of technology

It realizes the completion of three segmentation tasks of semantics, examples and panoramicity in one training cycle, accurately segment the material stack point cloud, identify the shape, location and boundaries of the material stack, and classify material categories, support material field inventory management and material scheduling, improve transportation efficiency, and reduce manual measurement errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832013B_ABST
    Figure CN119832013B_ABST
Patent Text Reader

Abstract

The present invention discloses a point cloud segmentation method based on deep learning, which relates to fields such as 3D vision and machine learning. Aiming at the problems that the point cloud in the stockyard cannot be accurately segmented during the intelligentization process of bulk material equipment, and the bulk material equipment cannot accurately perceive the stockpile information, etc., first, data collection and preprocessing are carried out; secondly, point features are aggregated to form superpoints; then, a query strategy and a Transformer decoder layer are designed to predict instance masks and semantic categories; finally, a multi-task loss function and a matching strategy are designed. Compared with the prior art, through the training of an instance segmentation dataset, the present invention completes three segmentation tasks, and can accurately identify the positions and boundaries of different stockpiles in the stockyard and classify their material categories, which can significantly improve the warehouse management efficiency and intelligent level of the stockyard, provide technical support for the efficient operation of bulk material equipment, and can be widely applied to fields such as bulk material equipment management, mineral collection, and dock handling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of three-dimensional vision, machine learning, etc., and specifically to a point cloud segmentation method based on deep learning. Background Art

[0002] Large bulk material equipment plays an important role in industrial production, and its intelligent upgrade is an inevitable trend in future development. Intelligent bulk material equipment can achieve continuous operation, meet the needs of large-scale bulk material storage and reclaiming, support multiple stacking methods, and can handle various bulk materials. This upgrade can not only significantly reduce labor costs, but also reduce equipment downtime, improve production efficiency and safety.

[0003] In recent years, with the introduction of technologies such as the Internet of Things, big data, and artificial intelligence, the intelligent level of bulk material equipment has been continuously improved. Three-dimensional point cloud segmentation is one of the key technologies to realize the intelligence of large bulk material equipment. By effectively segmenting point cloud data, it can provide an accurate basis for subsequent data processing and analysis, such as calculating the volume and weight of the stockpile, optimizing the reclaiming path, etc. However, point cloud segmentation faces many challenges in complex scenarios, especially when dealing with bulk material stockpiles, it is necessary to accurately segment the stockpile from the cluttered point cloud.

[0004] At present, significant progress has been made in point cloud segmentation algorithms. PointNet++ has improved local feature extraction, and the Graph Convolutional Network (GCN) can handle local and global features of point clouds. However, there are still problems such as high computational complexity, poor segmentation effect in complex scenarios, inability to handle three segmentation tasks of semantic segmentation, instance segmentation, and panoramic segmentation simultaneously, difficulty in capturing complex features, which limits its application in actual complex environments; at the same time, existing algorithms are not suitable for application in environments such as stockyards and mines. Due to problems such as many dynamic objects and high similarity of point clouds of different types of materials, it is difficult to segment point clouds, resulting in the inability of large bulk material equipment to identify stockpiles for automatic reclaiming, and also unable to calculate the volume and weight of stockpiles, unable to realize the intelligent management of stockyard data, and unable to meet the further development needs of the intelligence of large bulk material equipment.

[0005] In view of the above problems and actual needs, the present invention designs a point cloud segmentation method based on deep learning to solve the problem that the point clouds of stockpiles in the stockyard of large bulk material equipment cannot be accurately segmented. By training on a dataset with stockpile instance labels, it can achieve three segmentation tasks of semantics, instance, and panorama, avoiding spending three times the time and training volume to complete the three segmentation tasks, finally accurately segmenting each group of stockpile point clouds, and predicting the material category of each group of stockpile point clouds by obtaining the instance mask. Summary of the Invention

[0006] The present invention relates to a point cloud segmentation method based on deep learning. After obtaining the point cloud data of the stockpile in the stockyard of large bulk material equipment, operations such as noise reduction, outlier removal, and annotation are performed. Then, point - to - point weights are designed, a local neighborhood graph is constructed through geometric and color similarity, curvature features are introduced, 3D sparse convolution is used to extract point features, superpoints are generated through clustering, and dynamic weighted average pooling is used to aggregate superpoint features. Then, an initial query vector is generated through task - conditioned query, multi - scale feature interaction is realized using a Transformer decoder, convolutional kernels are dynamically generated, and joint output of multi - tasks including semantic segmentation, instance segmentation, and panoramic segmentation is supported. Finally, the multi - task loss is jointly optimized, curvature loss is introduced for geometric regularization, and the matching between predictions and real instances is simplified through a disentanglement matching strategy to optimize the end - to - end training efficiency.

[0007] The specific implementation steps are as follows:

[0008] Step 1: Obtain the environmental point cloud data set of the operation area of large bulk material equipment and pre - process it.

[0009] Step 1.1: Obtain the point cloud data set of the stockpile to be segmented. The original parameters of the point cloud need to include coordinate parameters, which are converted into x, y, z and carry RGB color information.

[0010] Step 1.2: Perform noise reduction and outlier removal on the original point cloud, and label the stockpile point cloud, including an instance segmentation data set with different categories and different individual labels.

[0011] Step 1.3: Data augmentation. Through operations such as random rotation and random flipping, data diversity is increased and the generalization ability of the model is improved.

[0012] Step 2: Since the number of point clouds is huge, directly training with the original point clouds will consume a large amount of resources, and the curvature of the stockpile point cloud is an important feature. Therefore, this method strengthens the effect of point cloud segmentation by adding curvature features to the original point clouds, and then forms superpoints from the point clouds, greatly reducing the number of point clouds entering the pooling layer and the resources consumed in training.

[0013] Step 2.1: Construct a neighborhood. For each point p i , an edge weight is used to divide points with similar geometric features and colors into the same neighborhood. The specific design steps are as follows:

[0014] N(p i ) = {p i |w ij > T},

[0015] where T represents the minimum weight threshold; w ij represents the weight between point p i and point p jThe weight value between them has the following specific expression:

[0016] w ij = α·S geo (p i , p j ) + (1 - α)·S color (c i , c j ),

[0017] where α is a balance parameter, S geo (p i , p j ) is the geometric similarity between point p i and point p j , and S color (c i , c j ) is the color similarity between point p i and point p j . The similarity is converted into a weight through a Gaussian kernel function:

[0018]

[0019] where ||p i - p j || 2 is the square of the Euclidean distance, c i = (r i , g i , b i ) is the color vector of point p i , and σ and τ are normalization parameters used to control the range of distance attenuation.

[0020] Step 2.2: Calculate the curvature. The specific steps are as follows:

[0021] Step 1: Calculate the covariance matrix C of the neighborhood points from the neighborhood constructed above:

[0022]

[0023] where k is the total number of points in the neighborhood of point p j , is the centroid of the neighborhood points, and ω j is the distance weight between the neighborhood points and the center point, and its expression is:

[0024]

[0025] where σ is a scale parameter.

[0026] Step 2: Curvature calculation. Perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues λ1 ≥ λ2 ≥ λ3. The curvature k is calculated using the smallest eigenvalue λ3:

[0027]

[0028] Step 3: Preprocessing of curvature features. Normalize the curvature values to the range [-1, 1] to avoid numerical instability. For points with larger curvatures, assign higher weights to highlight their importance;

[0029] Step 4: Attach a curvature value to each point. Suppose there are N input points, denoted as The point cloud features change from N×6 (x, y, z, r, g, b) to N×7 (x, y, z, k, r, g, b), obtaining where x, y, and z are coordinates, k is the curvature, and r, g, and b are red, green, and blue respectively, with a range of 0 to 255.

[0030] Step 2.3: Superpoint generation and feature extraction. First, use 3D sparse convolution on P′ ∈ R N×7 Extract point features to obtain Then, perform superpoint aggregation using connectivity-based clustering. The purpose is to merge points with similar features into the same superpoint and generate a superpoint set S i ∈ {S1, S2,..., S M}, where the number of superpoints M << N, and the features of the superpoints are obtained by dynamically weighted average pooling of pre-computed superpoints . The formula for dynamically weighted average pooling is:

[0031]

[0032] where f(p j ) is the feature vector of p j , K is the number of point clouds in the superpoint region, w j is the weight of point p j . Calculate the weight by combining the curvature feature k j :

[0033] w j = Softmax(β·k j + MLP(f(p j ))),

[0034] where β is a weight parameter used to control the influence of curvature, k j is the curvature of this superpoint, f(p j ) is the feature vector of p j , and MLP is a multi-layer perceptron.

[0035] Step 3: Design a query decoder that takes two sets of queries as the input of the query decoder, with the numbers being K ins and K sem , where K ins represents the number of instance queries, and K sem represents the number of semantic queries. The core function of the query decoder is to convert these queries into an equal number of kernels, generating a dynamically adjusted network behavior for cross-scale feature interaction through task-conditioned queries, without designing independent branches for different tasks.

[0036] Step 3.1: Task-conditioned query initialization. The input of the query decoder is a task-specific query vector, and its initialization process combines task semantics and the geometric features of the point cloud:

[0037] Step1: Assign a learnable task token as the conditional information of the task, where d1 is the number of task types;

[0038] Step2: Generate an initial query vector through the interaction between the task token and the global feature:

[0039]

[0040] where, is the global feature obtained by pooling the features of the highest layer of the sparse convolutional encoder, where C is the dimension of the superpoint feature, is the concatenation operation, and MLP is the multi-layer perceptron.

[0041] Step 3.2: The query decoder interacts with the multi-scale feature maps through the Transformer decoder layer to dynamically fuse local and global information:

[0042] Step1: Obtain the multi-resolution feature maps {F1, F2,......, F L} of the encoder, where N l is the number of points in the l-th layer, and C is the dimension of the projection matrix;

[0043] Step2: Use the cross-attention mechanism to interact the query vector Q with each scale feature map F l to extract task-related features:

[0044]

[0045] where, is the learnable projection matrix, and d is the dimension of the projection matrix;

[0046] Step 3: Feature fusion. The results of multi-scale cross-attention are fused to generate the final query vector:

[0047]

[0048] Among them, γ l is a learnable scale weight, calculated through the attention mechanism.

[0049] Step 3.3: Dynamically generate convolutional kernels through query features to directly predict the segmentation results. Each query Q final generates the corresponding convolutional kernel (k is the convolutional kernel size, and d is the dimension of the above projection matrix), and its expression is:

[0050] W kernel = MLP(Q final ),

[0051] Then apply the convolutional kernel to the multi-scale feature map to generate a mask:

[0052]

[0053] Among them, * represents the convolution operation, and the output M task is the task-specific segmentation result. The convolution operation enables the kernel to capture useful information in the superpoint features, and then generates two sets of outputs: K ins instances and K sem semantic masks. Instances represent independent objects or entities in the data, which are segmented from the original point cloud through the processing of the decoder, while the semantic masks are used to identify the semantic categories of these instances and determine which type of object they belong to.

[0054] Step 4: This design defines a cost function between the query and the ground truth object, designs a matching strategy to minimize this cost function, and formulates a loss function applied to the matching pairs to implement an end-to-end training method based on the transformer.

[0055] Step 4.1: Design the cost function C ik , which is used to measure the similarity between the i-th prediction and the k-th ground truth. The specific formula is as follows:

[0056]

[0057] Among them, is used to measure the probability that the i-th prediction belongs to the C k -th semantic category. λ is an adjustable weight parameter, and the overlapping point mask matching cost is the sum of binary cross-entropy (BCE) and Dice loss with Laplace smoothing:

[0058]

[0059] Among them, are respectively the predicted true value mask and the actual true value mask of an overlapping point.

[0060] Step 4.2: Disentangled matching. Establish a correspondence between an actual true value object, a superpoint, an instance query, and a guessed instance derived from this instance query. Skip the intermediate correspondences and directly match the predicted instance to a real instance, thus disentangling the guessed value and the actual value. However, the number of predictions still exceeds the number of actual true value instances. Therefore, it is necessary to filter out the predictions that do not correspond to the actual true value object. The disentangled matching technique simplifies the cost function optimization by setting the largest weight in the cost matrix to infinity:

[0061]

[0062] Step 4.3: Design the total loss function L, which is:

[0063] L = λ1L semantic + λ2(L BCE + L Dice ) + λ3L curvature ,

[0064] Among them, λ1, λ2, λ3 are weight parameters used to balance the importance of each part of the loss. The loss function L semantic for semantic segmentation is the cross-entropy loss:

[0065]

[0066] where S is the total number of superpoints, C is the total number of categories, y i,c is the true label of the i-th point belonging to category c, is the predicted probability that the i-th point is classified as category c; the loss function for instance segmentation is the sum of the binary cross-entropy loss L BCE and the Dice loss. The binary cross-entropy loss is used to predict whether each point belongs to a certain instance, and its expression is:

[0067]

[0068] where y i is the instance label of the i-th point, is the instance probability predicted by the model, and S is the number of superpoints; the Dice loss function is used to calculate the overlap degree of the instance masks to improve the segmentation accuracy:

[0069]

[0070] where yi is the instance label of the i-th point, is the instance probability predicted by the model, and S is the number of superpoints; then the curvature L curvature loss is used for geometric regularization to enhance the model's perception ability of the surface morphology. This loss is to ensure the curvature consistency of the model during the optimization process and reduce misclassifications in the boundary regions:

[0071]

[0072] where is the curvature predicted by the model, and k i is the true curvature, and S is the number of superpoints.

[0073] The present invention specifically relates to a point cloud segmentation method based on deep learning, which is used to solve the problem that the point cloud of the stockpile in the stockyard of large bulk material equipment cannot be accurately and quickly segmented. First, by constructing a special neighborhood of the original point cloud, calculating the curvature features of each point in the neighborhood, and then forming the features of superpoints through dynamically weighted average pooling, the computational complexity is greatly reduced; then, through task-conditioned query to generate a dynamic adjustment network behavior for cross-scale feature interaction, there is no need to design independent branches for different tasks, and a cost function, a loss function and a matching mechanism are designed for it. Finally, each group of stockpile point clouds is accurately segmented, and the material category of each group of stockpile point clouds is predicted by obtaining an instance mask. The technical solution of the present invention has the following beneficial technical effects:

[0074] 1. Through the training of one training cycle and one instance segmentation data set, three segmentation tasks of semantics, instance and panorama are completed, greatly reducing the model training time and computational cost, and reducing costs;

[0075] 2. The curvature features highlight the geometric structures of edges and corners, and can accurately identify the shapes, positions and boundaries of different stockpiles in the stockyard, and classify their material categories (such as coal, ore, etc.), providing accurate data support for inventory management and material scheduling;

[0076] 3. Each group of stockpile point clouds is accurately segmented, and the material type of each group of point clouds is predicted by the finally obtained instance mask. The weight of the stockpile is calculated through volume and density, which helps to realize the automated management of the stockyard inventory, reduce manual measurement errors, and improve the accuracy and timeliness of inventory data;

[0077] 4. By accurately segmenting the stockpile, the scheduling and transportation plans of materials can be optimized, the no-load rate of equipment can be reduced, the transportation efficiency can be improved, and the operating cost can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 is the overall flowchart of the present invention;

[0079] Figure 2 This is the network structure diagram of the present invention. Specific embodiments

[0080] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0081] Appendix Figure 1 This is the overall flowchart of the present invention. This embodiment relates to a point cloud segmentation method based on deep learning. Taking the point cloud of a large bulk material equipment storage yard as an example, the specific process includes the following: data preprocessing, generation and feature extraction of superpoints, design of a query decoder, and design of a loss function and matching mechanism.

[0082] Appendix Figure 2 This is the network structure diagram of the present invention. Its structure is that the input data undergoes data preprocessing, a sparse 3D convolutional layer, an average pooling layer to output superpoint features, and the values, keys, and query selections of the output superpoint features are input into the Transformer decoder layer. Then, through a disentangled matching mechanism, the instance kernel, instance score, semantic kernel of the point cloud are output and compared with the superpoint features, and finally, three segmentation results of semantics, panorama, and instance are output.

[0083] The specific implementation is described as follows:

[0084] Step 1: Obtain the environmental point cloud dataset of the operation area of large bulk material equipment and preprocess it.

[0085] Step 1.1: Use a 3D lidar and a depth camera to collect the point cloud data of the bulk material pile in the silo, ensure that the data contains the three-dimensional coordinates (x, y, z) and RGB color information (0 - 255) of each point, and deploy a real-time scanning device and set a regular scanning task.

[0086] Step 1.2: Use statistical filtering (Statistical Outlier Removal) to remove outliers, then smooth the noise through Gaussian filtering, and then use the CloudCompare tool to assign a unique instance label and semantic category (such as coal, ore, sand and gravel, etc.) to each pile instance. Finally, divide the dataset (70% training set, 15% validation set, 15% test set).

[0087] Step 1.3: Data augmentation, perform random rotation (±10°), random translation (±0.5 m), random scaling (0.9 - 1.1 times), and random color perturbation (±20% brightness) on the point cloud to increase data diversity.

[0088] Step 2: Since the number of point clouds is huge, directly training with the original point clouds will consume a large amount of resources, and the curvature of the stockpile point cloud is an important feature. Therefore, this method strengthens the effect of point cloud segmentation by adding curvature features to the original point clouds, and then forms superpoints from the point clouds, greatly reducing the number of point clouds entering the pooling layer and the resources consumed in training.

[0089] Step 2.1: Construct a neighborhood. For each point p i , use an edge weight to divide points with similar geometric features and colors into the same neighborhood. The specific design steps are as follows:

[0090] N(p i ) = {p i |w ij > T},

[0091] where T represents the minimum weight threshold; w ij represents the weight value between point p i and point p j , and its specific expression is:

[0092] w ij = α · S geo (p i , p j ) + (1 - α) · S color (c i , c j ),

[0093] where α is a balance parameter, S geo (p i , p j ) is the geometric similarity between point p i and point p j , S color (c i , c j ) is the color similarity between point p i and point p j . The similarity is converted into a weight through a Gaussian kernel function:

[0094]

[0095] where, ||p i - p j || 2 is the square of the Euclidean distance, c i = (r i , g i , b i ) is the color vector of point p i , and σ and τ are normalization parameters used to control the range of distance attenuation.

[0096] Step 2.2: Calculate the curvature. The specific steps are as follows:

[0097] Step1: From the neighborhood constructed above, calculate the covariance matrix C of the neighborhood points:

[0098]

[0099] where k is the total number of points in the neighborhood of point p j is the centroid of the neighborhood points, and ω is the distance weight between the neighborhood points and the center point, and its expression is: j where σ is the scale parameter.

[0100]

[0101] where σ is the scale parameter.

[0102] Step2: Curvature calculation. Perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues λ1≥λ2≥λ3. The curvature k is calculated through the minimum eigenvalue λ3:

[0103]

[0104] Step3: Preprocessing of curvature features. Normalize the curvature values to the range [-1, 1] to avoid numerical instability. For points with larger curvatures, assign higher weights to highlight their importance;

[0105] Step4: Attach a curvature value to each point. If N points are input, expressed as the point cloud features change from N×6(x, y, z, r, g, b) to N×7(x, y, z, k, r, g, b), obtaining where x, y, and z are coordinates, k is the curvature, and r, g, and b are red, green, and blue respectively, with a range of 0 to 255.

[0106] Step 2.3: Superpoint generation and feature extraction. First, use 3D sparse convolution on P′∈R N×7 to extract point features to obtain Then, use Euclidean clustering for superpoint aggregation. The purpose is to merge points with similar features into the same superpoint and generate a superpoint set S i ∈{S1, S2,..., S M}, where the number of superpoints M << N, and the features of the superpoint are obtained by dynamically weighted average pooling of the pre-computed superpoint The formula for dynamically weighted average pooling is:

[0107]

[0108] where f(p j ) is the feature vector of p j , K is the number of point clouds in the superpoint region, w j is the weight of point p j . Combining the curvature feature k j , the weight is calculated as follows:

[0109] w j = Softmax(β·k j + MLP(f(p j ))),

[0110] where β is the weight parameter used to control the influence of curvature, k j is the curvature of this superpoint, f(p j ) is the feature vector of p j , and MLP is the multi-layer perceptron.

[0111] Step 3: Design a query decoder that takes two sets of queries as the input of the query decoder, and their numbers are K ins and K sem , where K ins represents the number of instance queries, and K sem represents the number of semantic queries. The core function of the query decoder is to convert these queries into an equal number of kernels, generating a dynamically adjusted network behavior for cross-scale feature interaction through task-conditioned queries, without designing independent branches for different tasks.

[0112] Step 3.1: Task-conditioned query initialization. The input of the query decoder is a task-specific query vector, and its initialization process combines task semantics and the geometric features of the point cloud:

[0113] Step 1: Each task (semantic segmentation, instance segmentation, panoramic segmentation) is assigned a learnable task token as the conditional information of the task, where d1 is the number of task types;

[0114] Step 2: Generate the initial query vector through the interaction between the task token and the global feature:

[0115]

[0116] where, is the global feature obtained by pooling the features of the highest layer of the sparse convolutional encoder, where C is the superpoint feature dimension, is the concatenation operation, and MLP is the multi-layer perceptron.

[0117] Step 3.2: The query decoder interacts with the multi-scale feature maps through the Transformer decoder layer to dynamically fuse local and global information:

[0118] Step1: Obtain the multi-resolution feature maps {F1, F2,......, F L}, where N l is the number of points in the l-th layer, and C is the dimension of the projection matrix;

[0119] Step2: Use the cross-attention mechanism to interact the query vector Q with each scale feature map F l to extract task-related features:

[0120]

[0121] where, is a learnable projection matrix, and d is the dimension of the projection matrix;

[0122] Step3: Feature fusion, fuse the results of multi-scale cross-attention to generate the final query vector:

[0123]

[0124] where, γ l is a learnable scale weight, calculated through the attention mechanism.

[0125] Step 3.3: Dynamically generate convolution kernels through query features to directly predict the segmentation results. Each query Q final generates a corresponding convolution kernel (k is the convolution kernel size, and d is the dimension of the above projection matrix), and its expression is:

[0126] W kernel = MLP(Q final ),

[0127] Then apply the convolution kernel to the multi-scale feature maps to generate a mask:

[0128]

[0129] where, * represents the convolution operation, and the output M task is the task-specific segmentation result. The convolution operation enables the kernel to capture useful information in the superpoint features, and then generates two sets of outputs: K ins instances and K sem semantic masks. Instances represent independent objects or entities in the data, which are segmented from the original point cloud through the processing of the decoder, while the semantic masks are used to identify the semantic categories of these instances and determine which type of object they belong to.

[0130] Step 4: In this design, a cost function is defined between the query and the actual ground truth object, a matching strategy for minimizing this cost function is designed, and a loss function applied to the matching pairs is formulated to implement an end-to-end training method based on transformers.

[0131] Step 4.1: Design the cost function C ik , which is used to measure the similarity between the i-th prediction and the k-th actual ground truth. The specific formula is as follows:

[0132]

[0133] where is used to measure the probability that the i-th prediction belongs to the C k -th semantic category, λ is an adjustable weight parameter, and the overlap point mask matching cost is the sum of binary cross-entropy (BCE) and Dice loss with Laplace smoothing:

[0134]

[0135] where are the predicted ground truth mask and the actual ground truth mask of an overlap point respectively.

[0136] Step 4.2: Disentangled matching is performed to establish a correspondence between an actual ground truth object, a superpoint, an instance query, and a guessed instance derived from this instance query. By skipping the intermediate correspondences and directly matching the predicted instance to a real instance, the entanglement between the guessed value and the actual value is removed. However, the number of predictions still exceeds the number of actual ground truth instances. Therefore, predictions that do not correspond to the actual ground truth object need to be filtered out. The disentangled matching technique simplifies the cost function optimization by setting the largest weight in the cost matrix to infinity:

[0137]

[0138] Step 4.3: Design the total loss function L, which is:

[0139] L = λ1L semantic + λ2(L BCE + L Dice ) + λ3L curvature ,

[0140] where λ1, λ2, λ3 are weight parameters used to balance the importance of each part of the loss. The loss function L semantic for semantic segmentation is the cross-entropy loss:

[0141]

[0142] Where S is the total number of superpoints, C is the total number of categories, and y i,c is the true label that the i-th point belongs to category c, and is the predicted probability that the i-th point is classified as category c; the loss function for instance segmentation is the sum of binary cross-entropy loss and Dice loss. The binary cross-entropy loss is used to predict whether each point belongs to a certain instance, and its expression is:

[0143]

[0144] where y i is the instance label of the i-th point, is the instance probability predicted by the model, and S is the number of superpoints; the Dice loss function is used to calculate the overlap degree of the instance mask to improve the segmentation accuracy:

[0145]

[0146] where y i is the instance label of the i-th point, is the instance probability predicted by the model, and S is the number of superpoints; then the curvature L curvature loss is used for geometric regularization to enhance the model's perception ability of the surface morphology. This loss is to ensure the curvature consistency of the model during the optimization process and reduce misclassification in the boundary region:

[0147]

[0148] where is the curvature predicted by the model, k i is the true curvature, and S is the number of superpoints.

[0149] Through the above steps, a neural network model that can simultaneously perform semantic segmentation, instance segmentation, and panoramic segmentation on 3D point clouds can be built. Then, the divided data is put into the model for training, and finally, each pile of stockpile point clouds can be accurately segmented, and the material types of each pile of stockpile point clouds can be predicted through the finally obtained instance mask.

Claims

1. A point cloud segmentation method based on deep learning, characterized in that Including the following steps: Step 1: Obtain the environmental point cloud data set of the large bulk material equipment operation area, including collecting the point cloud data of the stockyard with three-dimensional coordinates and RGB color information, completing noise reduction, removing abnormal points, labeling the point cloud of the stockpile, and then enhancing the data to improve the generalization ability of the model; Step 2: Construct a local neighborhood for the processed point cloud data, and calculate the curvature of the points in the neighborhood and the superpoint features weighted by curvature; Step 3: Initialize task-conditioned query. The input to the query decoder is a task-specific query vector, and its initialization process combines task semantics and the geometric features of the point cloud. The query decoder interacts with the multi-scale feature maps through the Transformer decoder layer, dynamically fusing local and global information. Finally, convolution kernels are dynamically generated based on the query features, and K ins instances and K sem semantic masks are output. The convolution operation enables the kernel to capture useful information in the superpoint features; Step 4: Calculate the mask matching cost by predicting the true value mask and the actual true value mask, and design the cost function C ik , and design the disentangled matching and total loss function L.

2. The point cloud segmentation method based on deep learning according to claim 1, wherein The steps of constructing a local neighborhood for the processed point cloud data and calculating the curvature of the points in the neighborhood and the superpoint features weighted by curvature described in Step 2 are as follows: Step 2.1: Construct the neighborhood N(p i ), for each point p i : N(p i ) = {p i | w ij > T}, where T represents the minimum weight threshold; w ij represents the weight value between point p i and point p j and its specific expression is: w ij = α·S geo (p i , p j ) + (1 - α)·S color (c i , c j ), where α is a balance parameter, S geo (p i , p j ) is the geometric similarity between point p i and point p j , and S color (c i , c j ) is the color similarity between point p i and point p j . The similarity is converted into weights through a Gaussian kernel function: where ||p i -p j || 2 is the square of the Euclidean distance, c i =(r i , g i , b i ) is the color vector of point p i , and σ and τ are normalization parameters; Step 2.2: According to the constructed neighborhood above, calculate the curvature k of each point through the covariance matrix and add it to the features of the point. The input point cloud is represented as where N is the number of points in the point cloud, and the number of features of the point cloud is 6, namely x, y, z, r, g, b. x, y, z are coordinates, and r, g, b are red, green, and blue respectively, with a range of 0 to 255. After adding the curvature k, we get Step 2.3: Superpoint generation and feature extraction. First, use 3D sparse convolution on P′ ∈ R N×7 Extract point features to obtain Then, use a connectivity-based clustering algorithm to aggregate superpoints to form a superpoint domain. Merge points with similar features into the same superpoint and generate a superpoint set S i ∈ {S1, S2,..., S M}, where the number of superpoints M << N, and the features of the superpoints are obtained by dynamically weighted average pooling of the point cloud in each superpoint domain. The formula for dynamically weighted average pooling is: where f(p j ) is the eigenvector of p j , K is the number of point clouds in the superpoint region, w j is the weight of point p j . Combining the curvature feature k j to calculate the weight: w j = Softmax(β·k j + MLP(f(p j ))), where β is the weight parameter, and k j is the curvature of the point cloud, f(p j ) is the eigenvector of p j , and MLP is the multi-layer perceptron.

3. The point cloud segmentation method based on deep learning according to claim 1, characterized in that The task-conditioned query initialization described in step 3. The input to the query decoder is a task-specific query vector, and its initialization process combines task semantics and the geometric features of the point cloud. The query decoder interacts with the multi-scale feature maps through the Transformer decoder layer, dynamically fusing local and global information, and finally dynamically generates convolution kernels through the query features and outputs K ins instances and K sem semantic masks. The convolution operation enables the kernel to capture useful information in the superpoint features. The steps are as follows: Step 3.1: Initialize the task-conditioned query. The input of the query decoder is a task-specific query vector, and its initialization process combines the task semantics and the geometric features of the point cloud: Step1: Assign a task token to each task As the conditional information of the task, where d1 is the number of task types; Step 2: Generate an initial query vector through the interaction between the task token and the global feature: Among them, is the global feature obtained by pooling the features of the highest layer of the sparse convolutional encoder, where C is the dimension of the superpoint feature, is the concatenation operation, and MLP is the multi-layer perceptron; Step 3.2: Design the query decoder: Step1: Obtain the multi-resolution feature maps {F1, F2,......, F L}, where N l is the number of points in the l-th layer, and C is the dimension of the superpoint features; Step 2: Use the cross-attention mechanism to interact the query vector Q with each scale feature map F l for interaction: Among them, is the projection matrix, and d is the dimension of the projection matrix; Step 3: Expression of the final query vector: Among them, γ l is the scale weight; Step 3.3: Each query vector Q final generates a corresponding convolution kernel where k is the convolution kernel size and d is the dimension of the above projection matrix, and its expression is: W kernel = MLP(Q final ) Then, the convolution kernel is convolved with the multi-scale feature map to obtain the mask M task : Among them, * represents the convolution operation, and W kernel is the convolution kernel, and the output M task is the task-specific segmentation result. The convolution operation enables the kernel to capture useful information in the superpoint features.

4. The point cloud segmentation method based on deep learning according to claim 1, wherein, Calculating the mask matching cost by predicting the true value mask and the actual true value mask, and designing the cost function C ik , designing the disentangled matching and the total loss function L, the steps are as follows: Step 4.1: Design a cost function C ik , which is used to measure the similarity between the i-th prediction and the k-th actual true value. The specific formula is as follows: Among them, is used to measure the probability that the i-th prediction belongs to the C k -th semantic category, λ is an adjustable weight parameter, and the overlapping point mask matching cost is the sum of binary cross-entropy and Dice loss with Laplace smoothing: where m i , are respectively the predicted segmentation mask and the actual ground truth mask of an overlapping point. Step 4.2: Disentangled matching When the i-th point belongs to the k-th object, the cost is C ik , otherwise it is infinite: Step 4.3: Design the total loss function L: L = λ1L semantic + λ2(L BCE + L Dice ) + λ3L curvature , Among them, λ1, λ2, λ3 are weight parameters, and the loss function L of semantic segmentation semantic is the cross-entropy loss: where S is the total number of superpoints, C is the total number of categories, and y i,c is the true label that the i-th point belongs to category c, is the predicted probability that the i-th point is classified into category c; the loss function for instance segmentation is the sum of the binary cross-entropy loss L BCE and the Dice loss, and the expression of the binary cross-entropy loss is: where y i is the instance label of the i-th point, is the instance probability predicted by the model, and S is the number of superpoints; Dice loss function: where y i is the instance label of the i-th point, is the instance probability predicted by the model, and S is the number of superpoints; then the curvature L curvature loss is used for geometric regularization: Among them is the curvature predicted by the model, k i is the true curvature, and S is the number of superpoints.

Citation Information

Patent Citations

  • CBCT and laser scanning point cloud data tooth registration method based on supervoxels

    CN112200843A

  • Method, device and equipment for segmenting point cloud data and computer readable medium

    CN114549838A