Non-diffuse reflection three-dimensional noise point cloud elimination method based on improved LightGBM
Through the improved LightGBM algorithm and point cloud cluster feature extraction method, the positioning error and non-diffuse noise point cloud detection problems of LiDAR technology in dynamic and degraded scenarios are solved, and a more efficient and reliable noise point cloud removal effect is achieved.
Patent Information
- Application Number
- CN202510060367.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The existing LiDAR technology has positioning errors when dealing with dynamic and degraded scenarios, and the noise point cloud caused by non-diffuse reflection scenarios is difficult to effectively detect and eliminate, affecting the application reliability of the mapping results.
The improved LightGBM algorithm is used to construct a point cloud cluster classification feature system based on the geometry, attributes and mirror features of point cloud clusters, and through feature dimensionality reduction, transformation and screening, the optimal point cloud cluster classification feature subset is constructed to realize the detection and elimination of non-diffuse reflection noise point clouds.
It improves the application reliability of LiDAR mapping results, reduces the dependence of heuristic thresholds and prior premise assumptions, enhances detection efficiency and accuracy, and is suitable for a variety of non-diffuse reflection scenarios.
Smart Images

Figure CN119991481A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of simultaneous positioning and mapping based on laser radar (Light Detection and Ranging Simultaneous Localization and Mapping, LiDAR SLAM), and specifically relates to a method for removing non-diffuse three-dimensional noise point cloud based on improved LightGBM. Background Art
[0002] In recent years, the rapid development of technologies such as artificial intelligence, digital twins, and three-dimensional reconstruction has not only improved people's daily life patterns, but also prompted various types of remote sensing data and geographic data to become important basic and strategic resources. Among them, due to the advantages of high measurement accuracy, all-weather operation, and strong penetration of laser detection and ranging (LiDAR) technology, its original point cloud and mapping results are widely used in three-dimensional reconstruction, resource exploration, urban mapping, autonomous driving and other fields. However, as the application fields become more and more extensive, the application limitations of original LiDAR point clouds in some complex scenes are gradually exposed. For example, in dynamic scenes, there are large inter-frame pose changes in moving targets, which directly affect the accuracy of point cloud registration. In degraded scenes, due to the small number of effective features in a single-frame point cloud or the large number of similar features in adjacent frame point clouds, inter-frame registration is usually abnormal. The problems caused by the above two situations will be directly reflected by large positioning errors, so that they can be discovered and blocked in time in practical applications. In addition to the explicit problems that can be directly reflected above, non-diffuse reflection scenes also bring certain challenges, and their impact is usually not directly reflected by the positioning results. It is an implicit problem, which is mainly related to the propagation law of lasers. During the working process of LiDAR, after the laser beam broadcast by the laser transmitter hits an object with a rough surface, diffuse reflection occurs at the reflective surface, and the reflected light is finally captured by the laser receiver, thereby forming a point cloud cluster corresponding to the object in the actual collected point cloud. However, when the laser hits objects with smooth surfaces such as mirrors, water surfaces, smooth floor tiles, and smooth glass, mirror reflection will occur, resulting in the existence of false point cloud clusters formed by the reflected light hitting the object behind the corresponding reflective surface in the actual point cloud. Since the absolute position of the false point cloud cluster does not change with the change of the LiDAR scanning angle, it has little effect on the accuracy of the point cloud frame registration, but directly leads to large noise in the mapping results, which in turn affects the actual application effect. In addition, when the laser hits translucent objects such as glass and water, light transmission will occur, resulting in almost no laser points on the surface of the translucent object in the actual point cloud, while the transmitted light hits the object behind it to form a transmission point cloud cluster. The transmission point cloud cluster also has little effect on the accuracy of point cloud frame registration, but when the point cloud scanned by the translucent object does not exist in the point cloud map, it is considered that the space on both sides of the translucent object is connected, resulting in the mapping result being unable to provide effective application value. In summary, detecting non-diffuse reflection objects or directly eliminating the above two types of non-diffuse reflection noise point clouds is of great significance and value to improving the application reliability of LiDAR mapping results, so it is used as the research focus of this patent.
[0003] For the detection or elimination of non-diffuse reflective objects in LiDAR, the commonly used methods are mainly divided into methods assisted by other sensors and detection methods that rely only on point clouds. Among them, the former mainly uses heterogeneous sensors such as sonar to assist in the detection of non-diffuse reflective objects. Although the detection effect is good, the overall cost of the equipment is high, and the time and space benchmarks between the sensors need to be calibrated, so it is gradually replaced by the latter. At present, the detection methods that rely only on point clouds are mainly divided into three categories, including: detection methods based on single-echo point clouds, detection methods based on double-echo point clouds, and detection methods based on multi-echo point clouds. The detection methods based on single-echo point clouds are further divided into methods assisted by laser intensity and methods assisted by two-dimensional grids. The laser intensity assisted method constructs a mirror detector and a reflection classifier by analyzing the law of laser intensity changes. This method can simultaneously detect the position and material of the reflective surface, but it is necessary to pre-determine the reflection model parameters of various objects to be detected, and it requires a large amount of work and storage. With the help of two-dimensional grid-assisted method, an online heuristic algorithm is used based on LiDAR measurement information to obtain the quasi-diffuse reflection angle subset of each reflective surface, and a multi-channel or single-channel detection method is used to release the dynamic grid while building a two-dimensional grid map, so as to ensure that non-diffuse reflection objects participate in the mapping. The mapping results of this method can be directly applied to fields such as robot path planning, but it is mostly used in two-dimensional space, and necessary heuristic operations are required for different application scenarios. The detection method based on double echo point cloud is divided into a detection method based on reflection symmetry and a method assisted by laser intensity. The detection method based on reflection symmetry determines the approximate position of the surface to be detected according to the mirror symmetry and uses the extended Kalman filter (Extended Kalman Filter, EKF) to integrate the spatiotemporal information, thereby updating the direction and contour of the detection surface in real time. This method has high accuracy and real-time performance, but it is only suitable for detecting plane mirrors with obvious boundaries. With the help of laser intensity-assisted method, the double echo plane containing the reflective surface is determined according to the difference between the strongest and the last echo points, and then the intensity peak plane is screened according to the intensity change, and the boundary of the actual reflective surface is eroded and searched. This method can detect various smooth reflective surfaces, but it needs to be based on the premise that the surface to be detected has a clear boundary, and the detection accuracy decreases rapidly as the scene range increases. The detection method based on multi-echo point cloud is divided into detection method based on reflection symmetry and detection method based on learning. The detection method based on reflection symmetry divides the point cloud patches through spherical projection, and uses mixed K Gaussian distribution to model each patch and screen glass patches according to the characteristics of multi-echo point cloud, and calculates the symmetry and similarity between the point clouds on both sides of the main reflective glass, so as to detect the virtual points behind it. This method does not require pre-assumptions and training, and has high detection efficiency, but can only detect the corresponding virtual points of the main reflective surface of the current frame.In view of the above problems, the patch screening process is optimized and the reliability of reflection trajectory is added to the virtual point detection, so that the virtual point detection in the case of multiple reflecting surfaces is realized. However, the efficiency decreases with the increase of the number of surfaces, and this type of method is sensitive to the stability of multi-echoes, and the detection effect depends on the quality of the heuristic threshold. The learning-based detection method divides the two-dimensional grid and inputs its multi-echo pulse distribution into the probability estimation network, and inputs the estimation result into the three-dimensional feature similarity evaluation network to distinguish the real points and virtual points behind the glass. This method does not rely on prior assumptions and thresholds, but is only applicable to the detection of flat glass and is also affected by the stability of multi-echo point cloud acquisition. In summary, each method has its advantages and limitations, and it is necessary to further study how to balance the effect and efficiency of non-diffuse noise point cloud detection or removal.
[0004] In order to reduce the influence of heuristic threshold, prior assumption and multi-echo stability on the detection or removal of non-diffuse noise point cloud, this patent extracts and quantifies multiple features for each independent point cloud cluster based on the cluster segmentation results of single echo point cloud, so as to form a point cloud cluster classification feature system, and adopts the improved Light Gradient Boosting Machine (LightGBM) algorithm to directly detect whether each point cloud cluster is a non-diffuse noise point cloud cluster, so as to achieve the removal of noise point cloud that is insensitive to the type, shape and number of non-diffuse objects. Among the many machine learning algorithms currently available, LightGBM has the advantages of fast training speed, small memory usage, high accuracy and support for distributed computing, and many experts and scholars have optimized it from the aspects of dynamic parameter adjustment and comprehensive feature screening. Therefore, this algorithm is widely used to solve various classification and regression problems, but it is rarely used to process LiDAR point clouds, and it is almost never used to detect non-diffuse noise point clouds, so it is used as the core algorithm for the removal of non-diffuse three-dimensional noise point clouds. Based on the clustering results of independent point cloud clusters, this patent designs and extracts 7 geometric features, 9 attribute features and 3 mirror features to construct a point cloud cluster classification feature system, and then constructs a LightGBM classification decision tree for direct detection and elimination of non-diffuse three-dimensional noise point clouds. In order to improve the training efficiency of LightGBM, this patent calculates the Spearman's Rank Correlation Coefficient (SRCC) between each feature before bundling mutually exclusive features and reduces the dimension of the feature subsets with correlation. In addition, it is proposed to use the density-based spatial clustering of applications with noise (DBSCAN) algorithm to reasonably convert continuous numerical features into discrete category features, thereby assisting LightGBM in constructing feature histograms. In order to take into account the classification effect of LightGBM, this patent proposes a cosine similarity-based feature bidirectional screening (CS-FBS) algorithm. For the point cloud cluster classification feature system after feature bundling, the cosine similarity (CS) between each feature is calculated as a benchmark, and the optimal point cloud cluster classification feature subset is constructed through bidirectional feature screening (FBS), so as to obtain the corresponding optimal LightGBM classification model. Compared with other learning-based non-diffuse noise point detection methods, this patented method does not require a complex network structure, and has stronger applicability and stability.Although it is also necessary to build a training set and train a classification model, by improving the feature dimension reduction, transformation, screening and other steps in LightGBM, the training efficiency and classification effect of the model are better taken into account.
[0005] Light Gradient Boosting Machine (LightGBM) is a Gradient Boosting Machine (GBM) algorithm used to solve classification and regression problems. This patent mainly uses the above algorithm to detect whether each independent point cloud cluster is non-diffuse noise based on the constructed optimal point cloud cluster classification feature subset.
[0006] Speaman's rank Corelation Coefficient (SCC) is a statistical analysis indicator that reflects the degree of rank correlation. It is a statistic obtained by arranging the sample values of two elements in order of data size and replacing the actual data with the rank of the sample values of each element. This patent mainly uses the above coefficient to measure the correlation between the features of each point cloud cluster, thereby reducing the dimension of the point cloud cluster classification feature system.
[0007] Density-Based Spatial Clustering of Applications with Noise (DBSCAN) is a representative density-based clustering algorithm that defines a cluster as the largest set of density-connected points. It can divide areas with sufficiently high density into clusters and can find clusters of any shape in a noisy spatial database. This patent mainly uses the above algorithm to segment the sample sets corresponding to each continuous numerical feature, thereby converting it into discrete categorical features.
[0008] Cosine Similarity (CS) is a similarity evaluation index between vectors, also known as cosine similarity, which evaluates the similarity between two vectors by calculating the cosine value of the angle between them. This patent mainly uses the above index to judge the similarity between features, so as to construct the optimal feature subset for point cloud cluster classification. Summary of the invention
[0009] In view of the problems that the current non-diffuse three-dimensional noise point cloud detection methods require heuristic thresholds, rely on prior assumptions, and are affected by multi-echo stability, the present invention proposes a non-diffuse three-dimensional noise point cloud removal method based on improved LightGBM, which realizes a comprehensive judgment of whether an independent point cloud cluster is a non-diffuse noise point cloud cluster based on its basic characteristics.
[0010] A non-diffuse 3D noise point cloud removal method based on improved LightGBM includes the following steps:
[0011] Step 1. After LiDAR point cloud preprocessing and point cloud cluster segmentation, 7 geometric features, 9 attribute features and 3 mirror features are designed and extracted based on the independent point cloud cluster clustering results to construct a point cloud cluster classification feature system, and then a Light Gradient Boosting Machine (LightGBM) classification decision tree is constructed;
[0012] Step 2: To improve the training efficiency of LightGBM, the Spearman's Rank Correlation Coefficient (SRCC) between the features is calculated and the feature subset with correlation is reduced in dimension. In addition, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is used to reasonably transform continuous numerical features into discrete categorical features, thereby assisting LightGBM in constructing feature histograms.
[0013] Step 3: In order to take into account the classification effect of LightGBM, a cosine similarity-based feature bidirectional screening (CS-FBS) algorithm is designed to obtain the best LightGBM classification model corresponding to the optimal classification feature subset;
[0014] Step 4: Based on the best LightGBM classification model in step 3, detect and remove non-diffuse 3D noise point clouds.
[0015] After the LiDAR point cloud preprocessing and point cloud cluster segmentation described in step 1, 7 geometric features, 9 attribute features and 3 mirror features are designed and extracted based on the independent point cloud cluster clustering results to construct a point cloud cluster classification feature system, and then construct a LightGBM classification decision tree;
[0016] The specific steps are as follows:
[0017] Step 1-1: Aiming at the four key problems of LiDAR point cloud, namely, equipment system error, point cloud motion distortion, a large number of noise points, and over-dense data, an unsupervised LiDAR point cloud calibration method is used to compensate for the equipment system error, and the V-ICP (Velocity updating-Iterative Closest Point) method is used to remove point cloud motion distortion; on this basis, the voxel filtering method is used to remove discrete points and noise points;
[0018] Step 1-2: Based on the LiDAR point cloud processed in step 1-1, ground segmentation and point cloud clustering are performed in sequence;
[0019] Steps 1-3, geometric features refer to the features that describe the basic geometric structures such as the span, projection area, and volume of the point cloud cluster. A total of 7 geometric features were extracted, including: X-axis span L X , Y axis span L Y , Z axis span L Z 、The minimum circumscribed rectangular area S of the XY projection surface XY 、The minimum circumscribed rectangular area S of the XZ projection surface XZ , the minimum circumscribed rectangular area S of the YZ projection surface YZ , the volume V of the minimum circumscribed parallelepiped, the specific quantification methods of the above characteristics are as follows;
[0020] Taking the i-th frame point cloud as an example, the maximum and minimum values of the X, Y, and Z coordinates of all laser points in the k-th point cloud cluster are recorded as Calculate the corresponding three-axis span
[0021]
[0022] based on Quantify the results and calculate the area of the minimum circumscribed rectangle of the point cloud cluster on the three projection planes XY, YZ, and XZ
[0023]
[0024] Also based on Quantify the results and calculate the volume V of the smallest circumscribed parallelepiped of the point cloud cluster in three-dimensional space k :
[0025]
[0026] Where k represents the point cloud cluster number;
[0027] Steps 1-4: Attribute features refer to the features that describe the basic attributes of the point cloud cluster, such as the internal structure and intensity distribution. A total of 9 attribute features were extracted, including: the number of laser points N, the laser point density P, the average laser intensity I mean , laser intensity variance I var , the average laser intensity I of the three additional point cloud clusters closest to the geometric center of the current point cloud cluster ext , the first element N of the principal plane normal vector one , the second element of the principal plane normal vector N two , the third element of the principal plane normal vector N thr , the average error S of the principal plane fitting, the specific quantification methods of the above characteristics are as follows;
[0028] Taking the kth point cloud cluster in the i-th frame point cloud as an example, we directly count the number of laser points N k , and based on V k Quantitative results to calculate the laser point density P k :
[0029] P k =N k / V k (4)
[0030] While counting the number of laser points, the corresponding laser intensity of each laser point is accumulated and recorded as
[0031]
[0032] Among them, j represents the laser point number, intensity j represents the laser intensity corresponding to the jth laser point;
[0033] based on and N k Quantify the results and calculate the mean laser intensity
[0034]
[0035] based on and N k Quantify the results and calculate the laser intensity variance
[0036]
[0037] Mark the geometric center point of the current point cloud cluster as (X k ,Y k ,Z k), by calculating the Euclidean distance between the geometric center points of other point cloud clusters, the three additional point cloud clusters closest to each other are selected, and the overall laser intensity mean is calculated.
[0038]
[0039] in, They represent the mean laser intensity of the screened point cloud clusters;
[0040] For the current point cloud cluster, the Random Sample Consensus (RANSAC) algorithm is used to fit the main plane; the number of laser points in the fitting plane is recorded as And calculate the coordinate mean vector c k :
[0041]
[0042] in, Respectively represent vector c k The component element of j represents the serial number of the laser point in the main plane point cloud set. They represent the three-axis coordinate values corresponding to the jth laser point; the normal vector of the main plane is recorded as n k , construct the projection function of all laser points in the plane set:
[0043]
[0044] in, represents the residual sum of the projections of all laser points on the principal plane, p j represents the three-dimensional coordinate vector of the jth laser point; based on the principal component analysis (PCA) algorithm, the vertical direction of the concentrated direction of the laser point projection distribution is equivalent to the direction with the minimum variance of the laser point projection; for the convenience of solving, the formula (10) is simplified as:
[0045]
[0046] in, Calculate M k The eigenvalues and eigenvectors of the matrix are normalized for the eigenvector corresponding to the minimum eigenvalue, which is used as the normal vector of the current principal plane, and the three elements in the vector are assigned to
[0047] Based on the principal plane fitting results, calculate the average error S of the principal plane fitting k :
[0048]
[0049] Among them, ||*|| represents the modulus length of the orientation vector;
[0050] Step 1-5: When the laser hits a non-diffuse reflecting object at a nearly vertical angle, a point cloud will exist on the surface of the object in the actual point cloud. In order to identify the pseudo-diffuse reflecting noise point cloud generated by the above situation, three mirror features are extracted, including: the degree of conformity between the overall point cloud intensity and the cone model C, the ratio of the concave hull polygon area of the XZ projection surface to the minimum circumscribed rectangle area S cxz , the ratio of the area of the concave hull polygon of the YZ projection surface to the area of the minimum circumscribed rectangle S cyz ,The specific quantification methods of the above features are as follows;
[0051] For the case where the laser hits a non-diffuse reflecting object at an approximately vertical angle, the overall laser intensity of this type of point cloud cluster usually shows a trend of being high in the middle and gradually decreasing towards the surroundings; therefore, the point cloud clusters are projected onto the XZ plane and the YZ plane respectively, and the laser intensity is used as the coordinate value of the third dimension respectively. By fitting the three-dimensional cone surface and calculating the fitting error, it is determined whether the corresponding laser intensity of the point cloud cluster conforms to the above trend; taking the projection onto the XZ plane as an example, the jth laser point in the kth point cloud cluster is assigned a new three-dimensional coordinate, recorded as (X j ,Z j ,intensity j ); based on the updated coordinates of the laser point (X j ,Y j ,Z j ), define the distance function d from the jth point to the cone surface j :
[0052]
[0053] in, Represents the three-axis projection coordinates of the geometric center of the current point cloud cluster on the central axis of the cone, a k 、b k 、c k Represents the three component elements of the cone center axis direction vector, α k Represents the vertex angle of the cone normal cross section, r k represents the radius of the cone base circle;
[0054] Construct a scoring function based on the distance function of all laser points in the point cloud cluster
[0055]
[0056] For formula (14), the Gauss-Newton Iteration Method (GNIM) algorithm is used to solve it. When the value is the smallest, it corresponds to the best fitting conical surface;
[0057] Similarly, when the point cloud cluster is projected onto the YZ plane, the score corresponding to the best fitting conical surface can be obtained. based on and The quantitative results of the calculation of the overall point cloud strength and the degree of conformity C of the cone model k :
[0058]
[0059] Taking the projection of the point cloud cluster to the XZ plane as an example, the AlphaShape (AS) algorithm is used to extract the concave hull contour of the projected two-dimensional point cloud set. The contour points are numbered in the order of extraction and recorded as Calculate the area of the concave hull polygon in the XZ plane by segmenting the image based on the contour point numbers and corresponding coordinates
[0060]
[0061] in, represents the X-axis coordinate and Z-axis coordinate of the j-th contour point, and |*| represents taking the absolute value;
[0062] based on and Quantification result, calculate the ratio of the concave hull polygon area of the XZ projection surface to the minimum circumscribed rectangle area
[0063]
[0064] Similarly, calculate the ratio of the area of the concave hull polygon of the YZ projection surface to the area of the minimum circumscribed rectangle
[0065]
[0066] By analyzing and quantifying the data basis and information transmission relationship required for the above 19 features, a point cloud traversal is performed and a multi-threaded processing mode is adopted to construct a classification feature system for each independent point cloud cluster in a hierarchical manner, thereby ensuring the timeliness of data processing and transmission.
[0067] In order to improve the training efficiency of LightGBM as described in step 2, the SRCC between each feature is calculated and the dimensionality of the feature subset with correlation is reduced; in addition, the DBSCAN algorithm is used to reasonably convert continuous numerical features into discrete category features, thereby assisting LightGBM in constructing a feature histogram;
[0068] The specific steps are as follows:
[0069] Step 2-1. Since the feature bundling in LightGBM mainly reduces the dimension of mutually exclusive features, but the extracted features are almost not sparse, the number of mutually exclusive features is small, and feature dimensionality reduction cannot be effectively achieved; therefore, in order to further reasonably compress the feature categories, the highly correlated features are reduced in dimension by calculating the SRCC between the features before feature bundling; for the nth feature and the mth feature whose correlation is to be calculated, their data are sorted respectively, and the corresponding ranking is recorded as the rank rank, and then the corresponding rank difference d of the jth row of data is calculated j :
[0070] d j =rank nj -rank mj (19)
[0071] Among them, rank nj Indicates the rank corresponding to the j-th row of data of the n-th feature, rank mj Indicates the rank corresponding to the j-th row of data of the m-th feature;
[0072] Calculate SRCCP based on the corresponding rank difference of each row of data:
[0073]
[0074] Among them, l represents the number of data contained in the feature;
[0075] Formula (20) is used to calculate the correlation coefficients between all features. The features with a calculated result greater than 0.8 are respectively used to form high-correlation feature subsets, and only the features with the highest information content in each subset are retained to achieve feature dimensionality reduction while retaining the original feature information.
[0076] Step 2-2, since the classification feature system after dimensionality reduction only contains continuous numerical features, whether the feature binning can be performed accurately and efficiently directly determines the reliability of subsequent steps and the final classification accuracy; in the traditional LightGBM classification model, continuous numerical feature binning usually adopts simple equal-width binning or equal-amount binning, and then splits and merges the preliminary binning results, which is prone to uneven segmentation; in addition, the above binning method requires the maximum number of bins to be set in advance, which has certain limitations, and for features with different data distribution forms, a fixed number of bins usually leads to different degrees of over-segmentation or under-segmentation problems; to address the above problems, the DBSCAN algorithm is used to cluster each feature in advance, and the category label is directly assigned to the corresponding data according to the clustering results, so as to realize the conversion of continuous numerical features into discrete category features, and directly use it as the basis for constructing feature histograms; the above process makes the conversion results of discrete category features more reasonable and reliable, and does not need to pre-set the number of categories, avoiding the introduction of human errors, and realizing feature binning according to the data distribution form of each feature;
[0077] Step 2-3: Based on the classification feature system after dimensionality reduction and transformation, the exclusive feature bundling (EFB) algorithm is used to screen and bundle mutually exclusive features.
[0078] In step 3, a CS-FBS algorithm is designed to take into account the classification effect of LightGBM, so as to obtain the best LightGBM classification model corresponding to the optimal classification feature subset;
[0079] The specific steps are as follows:
[0080] Step 3-1, after feature dimensionality reduction, transformation, and bundling, the newly generated features may be correlated again; however, since the bundling process accumulates regional differences, the SRCC between the features cannot reflect the correct correlation at this time. Therefore, the CS between the corresponding data vectors of each feature is calculated to reflect the degree of correlation; taking the nth feature and the mth feature as an example, the CS between the two is recorded as C nm :
[0081]
[0082] Among them, η n represents the data vector corresponding to the nth feature, η m represents the data vector corresponding to the mth feature, ||*|| represents the modulus of the vector to be queried, N represents the number of data in the feature, μ nj represents the jth data in the nth feature, μ mj Represents the jth data in the mth feature;
[0083] Step 3-2: Based on the CS calculation results between features, the FBS strategy is used to construct the optimal classification feature subset; all features in the current classification feature system are sorted from large to small according to the importance of the model output and recorded as sequence F ps , and record its reverse sequence as F io ; F ps The first feature in the classification feature subset is added to the classification feature subset. Then, based on the CS calculation results between other features and this feature, the features with absolute values greater than 0.8 are also added to the classification feature subset to build a classification decision tree and calculate the current classification accuracy as the accuracy benchmark. Press F io The corresponding features are removed from the subset in the order of feature arrangement in the feature tree, and a decision tree is constructed and the classification accuracy is calculated. If the accuracy is lower than the accuracy benchmark, the removal operation is withdrawn, otherwise the removal operation is retained and the accuracy benchmark is updated. This feature removal process lasts until F io Stop after being completely traversed; continue to iterate the above process of adding and removing features from the subset. After each feature addition operation, the corresponding classification accuracy is also compared with the current accuracy benchmark. If the accuracy increases, the addition operation is retained and the accuracy benchmark is updated, and the feature removal process is performed at the same time. Otherwise, no operation is performed and the next iteration is entered. This iterative process continues until F ps Stop after emptying;
[0084] Step 3-3: Finally, based on the specific composition of the bundled features, determine the optimal point cloud cluster classification feature subset and the corresponding LightGBM classification model.
[0085] In step 4, based on the best LightGBM classification model in step 3, non-diffuse reflection three-dimensional noise point cloud is detected and eliminated;
[0086] The specific steps are as follows:
[0087] Step 4-1: In order to ensure that the constructed non-diffuse reflection three-dimensional noise point cloud detection model has universal applicability, LiDAR point clouds are collected statically in multiple real indoor and outdoor scenes, and the corresponding coordinate system is recorded as the L system. At the same time, the contour point coordinates of the non-diffuse reflection objects in the corresponding environment and the environmental feature point coordinates are collected by the total station, and the corresponding coordinate system is recorded as the W system. In the above process, if the boundary of the non-diffuse reflection object is regular and clear, the contour corner point coordinates are directly collected; if the boundary of the non-diffuse reflection object is smooth or blurred, the contour point coordinates are collected in appropriate interval segments; based on the collected LiDAR point cloud and non-diffuse reflection object contour coordinates, non-diffuse reflection noise point cloud cluster labels are assigned to each point cloud cluster in a single frame point cloud in batches;
[0088] Taking the i-th frame point cloud as an example, the laser points matched with the environmental feature points collected by the total station are extracted from the point cloud. Based on multiple pairs of coordinates of the same-name points, the Bursa seven-parameter model is used to solve Li Department and W i The conversion relationship between the systems is denoted as Then, using Convert all point clouds to W i system, and based on the property of laser propagation in a straight line, connect the LiDAR center and all non-diffuse reflection object contour points and extend them; assign the point cloud cluster whose geometric center is in the above three-dimensional extended space the non-diffuse reflection noise point cloud cluster label "1", and the remaining point cloud clusters are assigned the label "0";
[0089] Step 4-2: Train the best LightGBM classification model based on the quantized optimal point cloud cluster classification feature subset and labels, and use the model to detect whether each independent point cloud cluster is a non-diffuse noise point cloud cluster, and directly eliminate the non-diffuse 3D noise point cloud according to the detection results.
[0090] Beneficial effects of the present invention:
[0091] 1. In order to ensure that the constructed point cloud cluster classification feature system is comprehensive and reliable, the present invention comprehensively considers the structure and properties of each point cloud cluster, extracts and quantifies 19 features, covering multi-dimensional information such as point cloud cluster span, projection area, internal composition, intensity distribution, etc.; at the same time, based on the LiDAR point cloud and the environmental point coordinates collected by the total station and the coordinates of the contour points of non-diffuse reflection objects, each point cloud cluster in a single-frame point cloud is assigned a non-diffuse reflection noise point cloud cluster label in batches, so as to reduce human errors and reduce workload; based on the above-mentioned point cloud cluster classification feature system and non-diffuse reflection noise point cloud cluster labels, an experimental platform is built to collect data in 12 common real indoor and outdoor scenes, thereby constructing a training data set for the LightGBM classification model.
[0092] 2. Since the LiDAR data acquisition frequency is high and the research object of the present invention is an independent point cloud cluster, the training set and the test set both contain a large number of data samples, and the point cloud cluster classification feature system contains more classification features; therefore, in order to better balance the detection effect and training efficiency of the LightGBM classification model, the present invention calculates the SRCC auxiliary feature dimension reduction, and proposes to use the DBSCAN algorithm to assist in the construction of feature histograms, and also proposes the CS-FBS algorithm to construct the optimal point cloud cluster classification feature subset; relevant experimental results show that considering the classification effect and training efficiency, the improved LightGBM of the present invention is better than the comparison method and has a wider applicability, and the corresponding average F1-Measure is improved by 23.16%. Although the average training time is increased by 20.54%, the ratio of the accuracy improvement rate to the efficiency reduction rate is 1.1276, which means that the improvement of the present invention sacrifices a small part of the computational efficiency but obtains a greater accuracy improvement.
[0093] 3. In order to test the detection effect of the method of the present invention on noise point clouds in real indoor and outdoor non-diffuse reflection scenes, two suitable methods in the same field were selected for comparison according to actual conditions; the results show that the correct detection rate of non-diffuse reflection noise point clouds of the method of the present invention is 0.9340, the error detection rate of non-noise point clouds is 0.0617, and the time required to process a single frame of point clouds is 0.0271s; compared with the other two comparison methods, the correct detection rate of non-diffuse reflection noise point clouds is increased by 2.59% and 2.90% respectively, and the error detection rate of non-noise point clouds is reduced by The time required for processing a single frame of point cloud is shortened by 7.50% and 24.66% respectively; the time required for processing a single frame of point cloud is shortened by 28.12% and 73.35% respectively; the method of the present invention does not require a preset reflection parameter model or heuristic calculation, and can directly determine whether it belongs to non-diffuse reflection noise based on the structure and properties of the point cloud cluster, and constructs an optimal point cloud cluster classification feature subset and training parameter set, so the corresponding detection effect and efficiency have significant advantages; in addition, based on the point cloud processed by the method of the present invention, the quality of the corresponding point cloud map is significantly improved compared with the original point cloud map.
[0094] 4. Compared with the existing LiDAR non-diffuse noise point cloud detection method based on single echo point cloud, the method of the present invention does not need to build a parameter model for the object to be detected in advance, nor does it rely on the global pose accuracy of the point cloud and heuristic calculations. It can directly determine whether each point cloud cluster in a single-frame point cloud is non-diffuse noise; in addition, the constructed detection model takes into account both detection effect and training efficiency, and has strong universal applicability; the method of the present invention can not only be used in LiDAR-based SLAM (LiDAR Simultaneous Localization and Mapping, LiDAR SLAM), robot path planning and other aspects, but also can play a role in spatial database construction, agricultural and forestry census, geological survey and other fields to varying degrees; at the same time, the improved LightGBM of the present invention is also applicable to other fields, which can improve the effect and efficiency of solving corresponding problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 This is a flow chart of a method for removing non-diffuse three-dimensional noise point cloud based on improved LightGBM in the present invention;
[0096] Figure 2 A specific flow chart of step 1 of an embodiment of the present invention;
[0097] Figure 3 A specific flow chart of step 2 of an embodiment of the present invention;
[0098] Figure 4 A specific flow chart of step 3 of an embodiment of the present invention;
[0099] Figure 5 A specific flow chart of step 4 of an embodiment of the present invention;
[0100] Figure 6 A summary flow chart of one embodiment of the present invention;
[0101] Figure 7 Radar chart for evaluating the effectiveness of improving LightGBM in the present invention;
[0102] Figure 8 This is a sequence diagram of the non-diffuse noise point cloud detection effect of the method of the present invention and two comparative methods in three test scenes changing over time. DETAILED DESCRIPTION
[0103] An embodiment of the present invention is further described below with reference to the accompanying drawings.
[0104] In the embodiment of the present invention, a non-diffuse three-dimensional noise point cloud removal method based on improved LightGBM is used. Figure 1 As shown, the following steps are included:
[0105] Step 1: After LiDAR point cloud preprocessing and point cloud cluster segmentation, 7 geometric features, 9 attribute features and 3 mirror features are designed and extracted based on the independent point cloud cluster clustering results to construct a point cloud cluster classification feature system, and then construct a LightGBM classification decision tree;
[0106] Step 1-1: Aiming at the four key problems of LiDAR point cloud, namely, equipment system error, point cloud motion distortion, many noise points, and over-dense data, an unsupervised LiDAR point cloud calibration method is used to compensate for the equipment system error, and the V-ICP method is used to remove the point cloud motion distortion; on this basis, the voxel filtering method is used to remove discrete points and noise points;
[0107] Step 1-2: Based on the LiDAR point cloud processed in step 1-1, ground segmentation and point cloud clustering are performed in sequence;
[0108] Steps 1-3, geometric features refer to the features that describe the basic geometric structures such as the span, projection area, and volume of the point cloud cluster. A total of 7 geometric features were extracted, including: X-axis span L X , Y axis span L Y , Z axis span L Z 、The minimum circumscribed rectangular area S of the XY projection surface XY 、The minimum circumscribed rectangular area S of the XZ projection surface XZ , the minimum circumscribed rectangular area S of the YZ projection surface YZ , the volume V of the minimum circumscribed parallelepiped, the specific quantification methods of the above characteristics are as follows;
[0109] Taking the i-th frame point cloud as an example, the maximum and minimum values of the X, Y, and Z coordinates of all laser points in the k-th point cloud cluster are recorded as Calculate the corresponding three-axis span
[0110]
[0111] based on Quantify the results and calculate the area of the minimum circumscribed rectangle of the point cloud cluster on the three projection planes XY, YZ, and XZ
[0112]
[0113] Also based on Quantify the results and calculate the volume V of the smallest circumscribed parallelepiped of the point cloud cluster in three-dimensional space k :
[0114]
[0115] Where k represents the point cloud cluster number;
[0116] Steps 1-4: Attribute features refer to the features that describe the basic attributes of the point cloud cluster, such as the internal structure and intensity distribution. A total of 9 attribute features were extracted, including: the number of laser points N, the laser point density P, the average laser intensity I mean , laser intensity variance I var , the average laser intensity I of the three additional point cloud clusters closest to the geometric center of the current point cloud cluster ext , the first element N of the principal plane normal vector one , the second element of the principal plane normal vector N two , the third element of the principal plane normal vector N thr , the average error S of the principal plane fitting, the specific quantification methods of the above characteristics are as follows;
[0117] Taking the kth point cloud cluster in the i-th frame point cloud as an example, we directly count the number of laser points N k , and based on V k Quantitative results to calculate the laser point density P k :
[0118] P k =N k / V k (4)
[0119] While counting the number of laser points, the corresponding laser intensity of each laser point is accumulated and recorded as
[0120]
[0121] Among them, j represents the laser point number, intensity j represents the laser intensity corresponding to the jth laser point;
[0122] based on and N k Quantify the results and calculate the mean laser intensity
[0123]
[0124] based on and N k Quantify the results and calculate the laser intensity variance
[0125]
[0126] Mark the geometric center point of the current point cloud cluster as (X k ,Y k ,Z k ), by calculating the Euclidean distance between the geometric center points of other point cloud clusters, the three additional point cloud clusters closest to each other are selected, and the overall laser intensity mean is calculated.
[0127]
[0128] in, They represent the mean laser intensity of the screened point cloud clusters;
[0129] For the current point cloud cluster, the RANSAC algorithm is used to fit the main plane; the number of laser points in the fitting plane is recorded as And calculate the coordinate mean vector c k :
[0130]
[0131] in, Respectively represent vector c k The component element of j represents the serial number of the laser point in the main plane point cloud set. They represent the three-axis coordinate values corresponding to the jth laser point; the normal vector of the main plane is recorded as n k , construct the projection function of all laser points in the plane set:
[0132]
[0133] in, represents the residual sum of the projections of all laser points on the principal plane, p jrepresents the three-dimensional coordinate vector of the jth laser point. Based on the PCA algorithm, the vertical direction of the concentrated direction of the laser point projection distribution is equivalent to the direction with the smallest laser point projection variance. In order to facilitate the solution, equation (10) is simplified as follows:
[0134]
[0135] in, Calculate M k The eigenvalues and eigenvectors of the matrix are normalized for the eigenvector corresponding to the minimum eigenvalue, which is used as the normal vector of the current principal plane, and the three elements in the vector are assigned to
[0136] Based on the principal plane fitting results, calculate the average error S of the principal plane fitting k :
[0137]
[0138] Among them, ||*|| represents the modulus length of the orientation vector;
[0139] Step 1-5: When the laser hits a non-diffuse reflecting object at a nearly vertical angle, a point cloud will exist on the surface of the object in the actual point cloud. In order to identify the pseudo-diffuse reflecting noise point cloud generated by the above situation, three mirror features are extracted, including: the degree of conformity between the overall point cloud intensity and the cone model C, the ratio of the concave hull polygon area of the XZ projection surface to the minimum circumscribed rectangle area S cxz , the ratio of the area of the concave hull polygon of the YZ projection surface to the area of the minimum circumscribed rectangle S cyz ,The specific quantification methods of the above features are as follows;
[0140] For the case where the laser hits a non-diffuse reflecting object at an approximately vertical angle, the overall laser intensity of this type of point cloud cluster usually shows a trend of being high in the middle and gradually decreasing towards the surroundings; therefore, the point cloud clusters are projected onto the XZ plane and the YZ plane respectively, and the laser intensity is used as the coordinate value of the third dimension respectively. By fitting the three-dimensional cone surface and calculating the fitting error, it is determined whether the corresponding laser intensity of the point cloud cluster conforms to the above trend; taking the projection onto the XZ plane as an example, the jth laser point in the kth point cloud cluster is assigned a new three-dimensional coordinate, recorded as (X j ,Z j ,intensity j ); based on the updated coordinates of the laser point (X j ,Y j ,Z j ), define the distance function d from the jth point to the cone surface j :
[0141]
[0142] in, Represents the three-axis projection coordinates of the geometric center of the current point cloud cluster on the central axis of the cone, a k 、b k 、c k Represents the three component elements of the cone center axis direction vector, α k Represents the vertex angle of the cone normal cross section, r k represents the radius of the cone base circle;
[0143] Construct a scoring function based on the distance function of all laser points in the point cloud cluster
[0144]
[0145] For formula (14), the GNIM algorithm is used to solve it. When the value is the smallest, it corresponds to the best fitting conical surface;
[0146] Similarly, when the point cloud cluster is projected onto the YZ plane, the score corresponding to the best fitting conical surface can be obtained. based on and The quantitative results of the calculation of the overall point cloud strength and the degree of conformity C of the cone model k :
[0147]
[0148] Taking the projection of the point cloud cluster to the XZ plane as an example, the AS algorithm is used to extract the concave hull contour of the projected two-dimensional point cloud set, and The contour points are numbered in the order of extraction and recorded as Calculate the area of the concave hull polygon in the XZ plane by segmenting the image based on the contour point numbers and corresponding coordinates
[0149]
[0150] in, represents the X-axis coordinate and Z-axis coordinate of the j-th contour point, and |*| represents taking the absolute value;
[0151] based on and Quantification result, calculate the ratio of the concave hull polygon area of the XZ projection surface to the minimum circumscribed rectangle area
[0152]
[0153] Similarly, calculate the ratio of the area of the concave hull polygon of the YZ projection surface to the area of the minimum circumscribed rectangle
[0154]
[0155] By analyzing and quantifying the data basis and information transmission relationship required for the above 19 features, a point cloud traversal is performed and a multi-threaded processing mode is adopted to construct a classification feature system for each independent point cloud cluster in a hierarchical manner, thereby ensuring the timeliness of data processing and transmission.
[0156] Step 2: To improve the training efficiency of LightGBM, the SRCC between each feature is calculated and the dimensionality of the feature subset with correlation is reduced; in addition, the DBSCAN algorithm is used to reasonably convert continuous numerical features into discrete category features, thereby assisting LightGBM in constructing a feature histogram;
[0157] Step 2-1. Since the feature bundling in LightGBM mainly reduces the dimension of mutually exclusive features, but the extracted features are almost not sparse, the number of mutually exclusive features is small, and feature dimensionality reduction cannot be effectively achieved; therefore, in order to further reasonably compress the feature categories, the highly correlated features are reduced in dimension by calculating the SRCC between the features before feature bundling; for the nth feature and the mth feature whose correlation is to be calculated, their data are sorted respectively, and the corresponding ranking is recorded as the rank rank, and then the corresponding rank difference d of the jth row of data is calculated j :
[0158] d j =rank nj -rank mj (19)
[0159] Among them, rank nj Indicates the rank corresponding to the j-th row of data of the n-th feature, rank mj Indicates the rank corresponding to the j-th row of data of the m-th feature;
[0160] Calculate SRCCP based on the corresponding rank difference of each row of data:
[0161]
[0162] Among them, l represents the number of data contained in the feature;
[0163] Formula (20) is used to calculate the correlation coefficients between all features. The features with a calculated result greater than 0.8 are respectively used to form high-correlation feature subsets, and only the features with the highest information content in each subset are retained to achieve feature dimensionality reduction while retaining the original feature information.
[0164] Step 2-2, since the classification feature system after dimensionality reduction only contains continuous numerical features, whether the feature binning can be performed accurately and efficiently directly determines the reliability of subsequent steps and the final classification accuracy; in the traditional LightGBM classification model, continuous numerical feature binning usually adopts simple equal-width binning or equal-amount binning, and then splits and merges the preliminary binning results, which is prone to uneven segmentation; in addition, the above binning method requires the maximum number of bins to be set in advance, which has certain limitations, and for features with different data distribution forms, a fixed number of bins usually leads to different degrees of over-segmentation or under-segmentation problems; to address the above problems, the DBSCAN algorithm is used to cluster each feature in advance, and the category label is directly assigned to the corresponding data according to the clustering results, so as to realize the conversion of continuous numerical features into discrete category features, and directly use it as the basis for constructing feature histograms; the above process makes the conversion results of discrete category features more reasonable and reliable, and does not need to pre-set the number of categories, avoiding the introduction of human errors, and realizing feature binning according to the data distribution form of each feature;
[0165] Step 2-3: Based on the classification feature system after dimensionality reduction and transformation, the EFB algorithm is used to screen and bundle mutually exclusive features.
[0166] Step 3: In order to take into account the classification effect of LightGBM, a CS-FBS algorithm is designed to obtain the best LightGBM classification model corresponding to the optimal classification feature subset;
[0167] Step 3-1, after feature dimensionality reduction, transformation, and bundling, the newly generated features may be correlated again; however, since the bundling process accumulates regional differences, the SRCC between the features cannot reflect the correct correlation at this time. Therefore, the CS between the corresponding data vectors of each feature is calculated to reflect the degree of correlation; taking the nth feature and the mth feature as an example, the CS between the two is recorded as C nm :
[0168]
[0169] Among them, η n represents the data vector corresponding to the nth feature, η m represents the data vector corresponding to the mth feature, ||*|| represents the modulus of the vector to be queried, N represents the number of data in the feature, μ nj represents the jth data in the nth feature, μ mj Represents the jth data in the mth feature;
[0170] Step 3-2: Based on the CS calculation results between features, the FBS strategy is used to construct the optimal classification feature subset; all features in the current classification feature system are sorted from large to small according to the importance of the model output and recorded as sequence F ps , and record its reverse sequence as F io ; F ps The first feature in the classification feature subset is added to the classification feature subset. Then, based on the CS calculation results between other features and this feature, the features with absolute values greater than 0.8 are also added to the classification feature subset to build a classification decision tree and calculate the current classification accuracy as the accuracy benchmark. Press F io The corresponding features are removed from the subset in the order of feature arrangement in the feature tree, and a decision tree is constructed and the classification accuracy is calculated. If the accuracy is lower than the accuracy benchmark, the removal operation is withdrawn, otherwise the removal operation is retained and the accuracy benchmark is updated. This feature removal process lasts until F io Stop after being completely traversed; continue to iterate the above process of adding and removing features from the subset. After each feature addition operation, the corresponding classification accuracy is also compared with the current accuracy benchmark. If the accuracy increases, the addition operation is retained and the accuracy benchmark is updated, and the feature removal process is performed at the same time. Otherwise, no operation is performed and the next iteration is entered. This iterative process continues until F ps Stop after emptying;
[0171] Step 3-3: Finally, based on the specific composition of the bundled features, determine the optimal point cloud cluster classification feature subset and the corresponding LightGBM classification model.
[0172] Step 4: Based on the best LightGBM classification model in step 3, detect and remove non-diffuse 3D noise point clouds;
[0173] Step 4-1: In order to ensure that the constructed non-diffuse reflection three-dimensional noise point cloud detection model has universal applicability, LiDAR point clouds are collected statically in multiple real indoor and outdoor scenes, and the corresponding coordinate system is recorded as the L system. At the same time, the contour point coordinates of the non-diffuse reflection objects in the corresponding environment and the environmental feature point coordinates are collected by the total station, and the corresponding coordinate system is recorded as the W system. In the above process, if the boundary of the non-diffuse reflection object is regular and clear, the contour corner point coordinates are directly collected; if the boundary of the non-diffuse reflection object is smooth or blurred, the contour point coordinates are collected in appropriate interval segments; based on the collected LiDAR point cloud and non-diffuse reflection object contour coordinates, non-diffuse reflection noise point cloud cluster labels are assigned to each point cloud cluster in a single frame point cloud in batches;
[0174] Taking the i-th frame point cloud as an example, the laser points matched with the environmental feature points collected by the total station are extracted from the point cloud. Based on multiple pairs of coordinates of the same-name points, the Bursa seven-parameter model is used to solve L i Department and W i The conversion relationship between the systems is denoted as Then, using Convert all point clouds to W i system, and based on the property of laser propagation in a straight line, connect the LiDAR center and all non-diffuse reflection object contour points and extend them; assign the point cloud cluster whose geometric center is in the above three-dimensional extended space the non-diffuse reflection noise point cloud cluster label "1", and the remaining point cloud clusters are assigned the label "0";
[0175] Step 4-2: Train the best LightGBM classification model based on the quantized optimal point cloud cluster classification feature subset and labels, and use the model to detect whether each independent point cloud cluster is a non-diffuse noise point cloud cluster, and directly eliminate the non-diffuse 3D noise point cloud according to the detection results.
[0176] In the example of the present invention, LiDAR data is collected in a real environment by building a combined positioning mobile platform. The LiDAR uses RS-LiDAR-32 produced by RoboSense, which has a horizontal field of view of 360°, a resolution of [0.1°, 0.4°], a vertical field of view of [-25°, 15°], a resolution of 0.33°, and is set to collect point clouds at a frequency of 10Hz; the total station model is Leica TS50, which is equipped with a Leica full reflection prism, and the three-dimensional coordinates of environmental feature points and non-diffuse reflection object contour points are collected through point measurement mode. In order to verify whether the improvement of the present invention on LightGBM better takes into account the classification effect and training efficiency, three data sets are selected from the UCI machine learning database (http: / / archive.ics.uci.edu / ) to form a test set. Among them, there are 1 Iris data set containing only continuous numerical features, 1 Car Evaluation data set containing only discrete category features, and 1 Abalone data set containing both of the above two types of features. When constructing the training data set of the non-diffuse reflection three-dimensional noise point cloud detection model, in order to ensure that the model will not be overfitted or underfitted, the number of data items of the non-diffuse reflection point cloud cluster in the data set cannot differ by more than 2 times from the number of other data items. Based on the constructed experimental platform, training data are collected in 12 indoor and outdoor scenes, including: indoor hall scene with multiple mirrors, corridor scene with multiple windows on one side, foyer scene with multiple glass, corridor scene with smooth marble, indoor hall scene with smooth marble, park shore scene near water, underground garage scene, commercial street scene, traffic artery scene, campus playground scene, indoor shopping mall scene, and residential community scene. Three additional common indoor and outdoor scenes are selected to collect test data, including: indoor hall scene with both glass and mirrors, hall scene with both smooth marble floor and glass, and commercial pedestrian street scene, to verify the removal effect of the method of the present invention on non-diffuse reflection three-dimensional noise point cloud. Experimental scene 1 is an indoor hall scene with both glass and mirrors. Although the floor material is marble, the surface is relatively rough and can be regarded as a diffuse reflection surface. There are non-diffuse reflection objects such as glass doors, glass windows, large plane mirrors, and diffuse reflection objects such as sculptures, people, and door and window frames in the environment. Experimental scene 2 is a hall scene with both smooth marble floor and glass. The floor material is smooth marble and is regarded as a non-diffuse reflection surface. There are non-diffuse reflection objects such as glass doors and glass windows in the environment, as well as diffuse reflection objects such as painted load-bearing columns, fire hydrants, and trash cans. Experimental scene 3 is an urban commercial pedestrian street scene. The floor material is limestone, which is a diffuse reflection surface. There are many non-diffuse reflection objects such as glass display windows, glass doors, one-way mirrors, and diffuse reflection objects such as cars, stone piers, trash cans, and lamp posts in the environment. The trajectory lengths in scenes 1, 2, and 3 are approximately 25m, 45m, and 145m, respectively, and the acquisition time is approximately 20s, 30s, and 100s, respectively.
[0177] like Figure 7 As shown in the figure, the larger the area surrounded by the same color line, the better the corresponding algorithm can balance the detection effect and training efficiency. In each sub-figure, the area surrounded by the red line corresponding to the improved LightGBM of the present invention is significantly larger than the area surrounded by the blue line corresponding to the traditional LightGBM. Therefore, the improved LightGBM of the present invention can balance better classification effect and good training efficiency, successfully verifying the effectiveness of the improved LightGBM.
[0178] like Figure 8As shown in the figure, because the detection effect of the comparative method 1 depends on the construction of the corresponding reflection parameter model of various objects to be detected in advance, but the construction process of each parameter model is very complicated and requires a large storage space, the default model of the algorithm is directly used to process each test data set. The default model only contains the reflection parameters of glass, mirrors, and individual metals. Therefore, as shown in sub-figure (c), the blue line trend in the figure is relatively stable, but significantly lower than the other two methods. This is mainly because the smooth marble floor in scene 2 cannot be effectively recognized by the default model, resulting in most of the non-diffuse reflection point clouds generated by its reflection being missed, but because the parameters in the detection model are always fixed, it can maintain a relatively stable detection effect. In comparison, as shown in sub-figures (a) and (e), the blue line trend in the figure is stable and generally between the other two methods. This is mainly because the non-diffuse reflection objects in scenes 1 and 3 are almost all glass, mirrors and common metals, so the default model can take into account both higher detection accuracy and more stable detection effects. In addition, as shown in sub-figures (b), (d), and (f), the blue line trend in the figure is also stable and in most cases better than the other two methods. This is mainly because the detection model constructed by this method is highly targeted. For objects other than the preset categories to be detected, this method is extremely unlikely to misdetect them, so the corresponding non-noise point cloud error detection rate is generally low. Comparative method 2 unifies the coordinate system between the point cloud and the two-dimensional grid map based on the global pose of the single-frame point cloud, and then detects repeatability through a multi-channel mode or detects local information through a single-channel mode, thereby indirectly screening out non-diffuse noise point cloud clusters. The above detection mode causes the detection effect of this method to be too dependent on the accuracy of the global pose of the point cloud. Therefore, as shown in the 6 sub-figures, the trend of the green broken line in the figure is unstable, and the detection accuracy fluctuates greatly, especially when the carrier turns or passes through a complex environment, the corresponding detection accuracy is the lowest. In addition, there are more dynamic targets in scene 3 and the distribution of objects in the environment is complex. The occlusion between targets is more common and random, resulting in the instability of the reproducibility of the point cloud clusters between frames and the local information of the point cloud clusters, which in turn affects the detection effect and the degree of influence gradually accumulates over time. Therefore, as shown in sub-figures (e) and (f), the green broken line in the figure fluctuates and also shows an overall trend of gradually decreasing and gradually increasing, respectively. Since the method performs necessary heuristic operations for different application scenarios to ensure the applicability of the detection model, when the influence of external environmental factors is small, as shown in sub-images (a), (b), (c), and (d), the differences between the main distribution areas of the green broken lines in the corresponding sub-images are small, indicating that the method is generally applicable to various environments, but is sensitive to environmental noise. As shown in the six sub-images, in the three test scenarios, the correct detection rate of the non-diffuse noise point cloud of the method of the present invention is generally higher than 0.9, and the false detection rate of the non-noise point cloud is generally lower than 0.1, indicating that the method of the present invention has strong universal applicability to different environments and is insensitive to external noise.Among them, as shown in sub-figures (e) and (f), the detection effect of the method of the present invention in scene 3 is slightly worse than that in the other two scenes. This is mainly because there are many glass windows of shops in this scene. When there are posters or stains on the glass, it can be directly approximated as a diffuse reflection object, resulting in the inability to accurately distinguish the above diffuse reflection objects from other real diffuse reflection objects based only on the constructed classification feature system, so there are some non-diffuse reflection objects that are missed. In addition, in outdoor scenes, the detection results of distant point cloud clusters have greater randomness, which also leads to the omission of some non-diffuse reflection noise point clouds or misdetection of other point clouds. However, since the classification feature system constructed by the present invention contains 19 features, and each feature plays a role in the classification process, the occlusion phenomenon between targets has little effect on the detection effect. As shown in sub-figures (a), (b), (c), and (d), the detection effects of the method of the present invention in scenes 1 and 2 are stable and relatively similar, further illustrating that the method of the present invention is generally applicable to reflective objects of different categories or materials. Through quantitative comparison, compared with the two comparison methods, the average non-diffuse noise point cloud correct detection rate of the method of the present invention increased by 2.59% and 2.90% respectively, the average non-noise point cloud error detection rate decreased by 7.50% and 24.66% respectively, the average detection efficiency increased by 28.12% and 73.35% respectively, the non-diffuse noise point cloud correct detection rate was as high as 0.9340, and the non-noise point cloud error detection rate was only 0.0617, and the average time required to process a single frame of point cloud was 0.0271s. In summary, the method of the present invention can efficiently and accurately detect non-diffuse three-dimensional noise point clouds, and effectively improve the quality of LiDAR point clouds in non-diffuse scenes by eliminating the above noise points.
[0179] The above is only the most basic specific implementation of the present invention, but the protection scope of the present invention is not limited thereto. Any replacement that can be understood by a person skilled in the art within the technical scope disclosed by the present invention should be included in the scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A non-diffuse three-dimensional noise point cloud removal method based on improved LightGBM, characterized in that: The following steps are involved: Step 1. After LiDAR point cloud preprocessing and point cloud cluster segmentation, 7 geometric features, 9 attribute features and 3 mirror features are designed and extracted based on the independent point cloud cluster clustering results to construct a point cloud cluster classification feature system, and then a Light Gradient Boosting Machine (LightGBM) classification decision tree is constructed; Step 2: To improve the training efficiency of LightGBM, the Spearman's Rank Correlation Coefficient (SRCC) between the features is calculated and the feature subset with correlation is reduced in dimension. In addition, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is used to reasonably transform continuous numerical features into discrete categorical features, thereby assisting LightGBM in constructing feature histograms. Step 3: In order to take into account the classification effect of LightGBM, a cosine similarity-based feature bidirectional screening (CS-FBS) algorithm is designed to obtain the best LightGBM classification model corresponding to the optimal classification feature subset; Step 4: Based on the best LightGBM classification model in step 3, detect and remove non-diffuse 3D noise point clouds.
2. According to the method for removing non-diffuse three-dimensional noise point cloud based on improved LightGBM in claim 1, it is characterized in that: After the LiDAR point cloud preprocessing and point cloud cluster segmentation described in step 1, 7 geometric features, 9 attribute features and 3 mirror features are designed and extracted based on the independent point cloud cluster clustering results to construct a point cloud cluster classification feature system, and then construct a LightGBM classification decision tree; The specific steps are as follows: Step 1-1: Aiming at the four key problems of LiDAR point cloud, namely, equipment system error, point cloud motion distortion, a large number of noise points, and over-dense data, an unsupervised LiDAR point cloud calibration method is used to compensate for the equipment system error, and the V-ICP (Velocity updating-Iterative Closest Point) method is used to remove point cloud motion distortion; on this basis, the voxel filtering method is used to remove discrete points and noise points; Step 1-2: Based on the LiDAR point cloud processed in step 1-1, ground segmentation and point cloud clustering are performed in sequence; Steps 1-3, geometric features refer to the features that describe the basic geometric structures such as the span, projection area, and volume of the point cloud cluster. A total of 7 geometric features were extracted, including: X-axis span L X , Y axis span L Y , Z axis span L Z 、The minimum circumscribed rectangular area S of the XY projection surface XY 、The minimum circumscribed rectangular area S of the XZ projection surface XZ , the minimum circumscribed rectangular area S of the YZ projection surface YZ , the volume V of the minimum circumscribed parallelepiped, the specific quantification methods of the above characteristics are as follows; Taking the i-th frame point cloud as an example, the maximum and minimum values of the X, Y, and Z coordinates of all laser points in the k-th point cloud cluster are recorded as Calculate the corresponding three-axis span based on Quantify the results and calculate the area of the minimum circumscribed rectangle of the point cloud cluster on the three projection planes XY, YZ, and XZ Also based on Quantify the results and calculate the volume V of the smallest circumscribed parallelepiped of the point cloud cluster in three-dimensional space k : Among them, k represents the point cloud cluster number; Steps 1-4: Attribute features refer to the features that describe the basic attributes of the point cloud cluster, such as the internal structure and intensity distribution. A total of 9 attribute features were extracted, including: the number of laser points N, the laser point density P, the average laser intensity I mean , laser intensity variance I var , the average laser intensity I of the three additional point cloud clusters closest to the geometric center of the current point cloud cluster ext , the first element N of the principal plane normal vector one , the second element of the principal plane normal vector N two , the third element of the principal plane normal vector N thr , the average error S of the principal plane fitting, the specific quantification methods of the above characteristics are as follows; Taking the kth point cloud cluster in the i-th frame point cloud as an example, we directly count the number of laser points N k , and based on V k Quantitative results to calculate the laser point density P k : P k =N k / V k (4) While counting the number of laser points, the corresponding laser intensity of each laser point is accumulated and recorded as Among them, j represents the laser point number, intensity j represents the laser intensity corresponding to the jth laser point; based on and N k Quantify the results and calculate the mean laser intensity based on and N k Quantify the results and calculate the laser intensity variance Mark the geometric center point of the current point cloud cluster as (X k ,Y k ,Z k ), by calculating the Euclidean distance between the geometric center points of other point cloud clusters, the three additional point cloud clusters closest to each other are selected, and the overall laser intensity mean is calculated. in, They represent the mean laser intensity of the screened point cloud clusters; For the current point cloud cluster, the Random Sample Consensus (RANSAC) algorithm is used to fit the main plane; the number of laser points in the fitting plane is recorded as And calculate the coordinate mean vector c k : in, Respectively represent vector c k The component element of j represents the serial number of the laser point in the main plane point cloud set. They represent the three-axis coordinate values corresponding to the jth laser point; the normal vector of the main plane is recorded as n k , construct the projection function of all laser points in the plane set: in, represents the residual sum of the projections of all laser points on the principal plane, p j represents the three-dimensional coordinate vector of the jth laser point; based on the principal component analysis (PCA) algorithm, the vertical direction of the concentrated direction of the laser point projection distribution is equivalent to the direction with the minimum variance of the laser point projection; for the convenience of solving, the formula (10) is simplified as: in, Calculate M k The eigenvalues and eigenvectors of the matrix are normalized for the eigenvector corresponding to the minimum eigenvalue, which is used as the normal vector of the current principal plane, and the three elements in the vector are assigned to Based on the principal plane fitting results, calculate the average error S of the principal plane fitting k : Among them, ||*|| represents the modulus length of the orientation vector; Step 1-5: When the laser hits a non-diffuse reflecting object at a nearly vertical angle, a point cloud will exist on the surface of the object in the actual point cloud. In order to identify the pseudo-diffuse reflecting noise point cloud generated by the above situation, three mirror features are extracted, including: the degree of conformity between the overall point cloud intensity and the cone model C, the ratio of the concave hull polygon area of the XZ projection surface to the minimum circumscribed rectangle area S cxz , the ratio of the area of the concave hull polygon of the YZ projection surface to the area of the minimum circumscribed rectangle S cyz ,The specific quantification methods of the above features are as follows; For the case where the laser hits a non-diffuse reflecting object at an approximately vertical angle, the overall laser intensity of this type of point cloud cluster usually shows a trend of being high in the middle and gradually decreasing towards the surroundings; therefore, the point cloud clusters are projected onto the XZ plane and the YZ plane respectively, and the laser intensity is used as the coordinate value of the third dimension respectively. By fitting the three-dimensional cone surface and calculating the fitting error, it is determined whether the corresponding laser intensity of the point cloud cluster conforms to the above trend; taking the projection onto the XZ plane as an example, the jth laser point in the kth point cloud cluster is assigned a new three-dimensional coordinate, recorded as (X j ,Z j ,intensity j ); based on the updated coordinates of the laser point (X j ,Y j ,Z j ), define the distance function d from the jth point to the cone surface j : in, Represents the three-axis projection coordinates of the geometric center of the current point cloud cluster on the central axis of the cone, a k , b k 、c k Represents the three component elements of the cone center axis direction vector, α k Represents the vertex angle of the cone normal cross section, r k represents the radius of the cone base circle; Construct a scoring function based on the distance function of all laser points in the point cloud cluster For formula (14), the Gauss-Newton Iteration Method (GNIM) algorithm is used to solve it. When the value is the smallest, it corresponds to the best fitting conical surface; Similarly, when the point cloud cluster is projected onto the YZ plane, the score corresponding to the best fitting conical surface can be obtained. based on and The quantitative results of the calculation of the overall point cloud strength and the degree of conformity C of the cone model k : Taking the projection of the point cloud cluster to the XZ plane as an example, the AlphaShape (AS) algorithm is used to extract the concave hull contour of the projected two-dimensional point cloud set. The contour points are numbered in the order of extraction and recorded as Calculate the area of the concave hull polygon in the XZ plane by segmenting the image based on the contour point numbers and corresponding coordinates in, represents the X-axis coordinate and Z-axis coordinate of the j-th contour point, and |*| represents taking the absolute value; based on and Quantification result, calculate the ratio of the concave hull polygon area of the XZ projection surface to the minimum circumscribed rectangle area Similarly, calculate the ratio of the area of the concave hull polygon of the YZ projection surface to the area of the minimum circumscribed rectangle By analyzing and quantifying the data basis and information transmission relationship required for the above 19 features, a point cloud traversal is performed and a multi-threaded processing mode is adopted to construct a classification feature system for each independent point cloud cluster in a hierarchical manner, thereby ensuring the timeliness of data processing and transmission.
3. According to the method for removing non-diffuse three-dimensional noise point cloud based on improved LightGBM in claim 1, it is characterized in that: In order to improve the training efficiency of LightGBM as described in step 2, the SRCC between each feature is calculated and the dimensionality of the feature subset with correlation is reduced; in addition, the DBSCAN algorithm is used to reasonably convert continuous numerical features into discrete category features, thereby assisting LightGBM in constructing a feature histogram; The specific steps are as follows: Step 2-1. Since the feature bundling in LightGBM mainly reduces the dimension of mutually exclusive features, but the extracted features are almost not sparse, the number of mutually exclusive features is small, and feature dimensionality reduction cannot be effectively achieved; therefore, in order to further reasonably compress the feature categories, the highly correlated features are reduced in dimension by calculating the SRCC between the features before feature bundling; for the nth feature and the mth feature whose correlation is to be calculated, their data are sorted respectively, and the corresponding ranking is recorded as the rank rank, and then the corresponding rank difference d of the jth row of data is calculated j : d j =rank nj -rank mj (19) Among them, rank nj Indicates the rank corresponding to the j-th row of data of the n-th feature, rank mj Indicates the rank corresponding to the j-th row of data of the m-th feature; Calculate SRCCP based on the corresponding rank difference of each row of data: Among them, l represents the number of data contained in the feature; Formula (20) is used to calculate the correlation coefficients between all features. The features with a calculated result greater than 0.8 are respectively used to form high-correlation feature subsets, and only the features with the highest information content in each subset are retained to achieve feature dimensionality reduction while retaining the original feature information. Step 2-2, since the classification feature system after dimensionality reduction only contains continuous numerical features, whether the feature binning can be performed accurately and efficiently directly determines the reliability of subsequent steps and the final classification accuracy; in the traditional LightGBM classification model, continuous numerical feature binning usually adopts simple equal-width binning or equal-amount binning, and then splits and merges the preliminary binning results, which is prone to uneven segmentation; in addition, the above binning method requires the maximum number of bins to be set in advance, which has certain limitations, and for features with different data distribution forms, a fixed number of bins usually leads to different degrees of over-segmentation or under-segmentation problems; to address the above problems, the DBSCAN algorithm is used to cluster each feature in advance, and the category label is directly assigned to the corresponding data according to the clustering results, so as to realize the conversion of continuous numerical features into discrete category features, and directly use it as the basis for constructing feature histograms; the above process makes the conversion results of discrete category features more reasonable and reliable, and does not need to pre-set the number of categories, avoiding the introduction of human errors, and realizing feature binning according to the data distribution form of each feature; Step 2-3: Based on the classification feature system after dimensionality reduction and transformation, the exclusive feature bundling (EFB) algorithm is used to screen and bundle mutually exclusive features.
4. According to the method for removing non-diffuse three-dimensional noise point cloud based on improved LightGBM in claim 1, it is characterized in that: In step 3, a CS-FBS algorithm is designed to take into account the classification effect of LightGBM, so as to obtain the best LightGBM classification model corresponding to the optimal classification feature subset; The specific steps are as follows: Step 3-1, after feature dimensionality reduction, transformation, and bundling, the newly generated features may be correlated again; however, since the bundling process accumulates regional differences, the SRCC between the features cannot reflect the correct correlation at this time. Therefore, the CS between the corresponding data vectors of each feature is calculated to reflect the degree of correlation; taking the nth feature and the mth feature as an example, the CS between the two is recorded as C nm : Among them, η n represents the data vector corresponding to the nth feature, η m represents the data vector corresponding to the mth feature, ||*|| represents the modulus of the vector to be queried, N represents the number of data in the feature, μ nj represents the jth data in the nth feature, μ mj Represents the jth data in the mth feature; Step 3-2: Based on the CS calculation results between features, the FBS strategy is used to construct the optimal classification feature subset; all features in the current classification feature system are sorted from large to small according to the importance of the model output and recorded as sequence F ps , and record its reverse sequence as F io ; F ps The first feature in the classification feature subset is added to the classification feature subset. Then, based on the CS calculation results between other features and this feature, the features with absolute values greater than 0.8 are also added to the classification feature subset to build a classification decision tree and calculate the current classification accuracy as the accuracy benchmark. Press F io The corresponding features are removed from the subset in the order of feature arrangement in the feature tree, and a decision tree is constructed and the classification accuracy is calculated. If the accuracy is lower than the accuracy benchmark, the removal operation is withdrawn, otherwise the removal operation is retained and the accuracy benchmark is updated. This feature removal process lasts until F io Stop after being completely traversed; continue to iterate the above process of adding and removing features from the subset. After each feature addition operation, the corresponding classification accuracy is also compared with the current accuracy benchmark. If the accuracy increases, the addition operation is retained and the accuracy benchmark is updated, and the feature removal process is performed at the same time. Otherwise, no operation is performed and the next iteration is entered. This iterative process continues until F ps Stop after emptying; Step 3-3: Finally, based on the specific composition of the bundled features, determine the optimal point cloud cluster classification feature subset and the corresponding LightGBM classification model.
5. According to the method for removing non-diffuse three-dimensional noise point cloud based on improved LightGBM in claim 1, it is characterized in that: In step 4, based on the best LightGBM classification model in step 3, non-diffuse reflection three-dimensional noise point cloud is detected and eliminated; The specific steps are as follows: Step 4-1: In order to ensure that the constructed non-diffuse reflection three-dimensional noise point cloud detection model has universal applicability, LiDAR point clouds are collected statically in multiple real indoor and outdoor scenes, and the corresponding coordinate system is recorded as the L system. At the same time, the contour point coordinates of the non-diffuse reflection objects in the corresponding environment and the environmental feature point coordinates are collected by the total station, and the corresponding coordinate system is recorded as the W system. In the above process, if the boundary of the non-diffuse reflection object is regular and clear, the contour corner point coordinates are directly collected; if the boundary of the non-diffuse reflection object is smooth or blurred, the contour point coordinates are collected in appropriate interval segments; based on the collected LiDAR point cloud and non-diffuse reflection object contour coordinates, non-diffuse reflection noise point cloud cluster labels are assigned to each point cloud cluster in a single frame point cloud in batches; Taking the i-th frame point cloud as an example, the laser points matched with the environmental feature points collected by the total station are extracted from the point cloud. Based on multiple pairs of coordinates of the same-name points, the Bursa seven-parameter model is used to solve L i Department and W i The conversion relationship between the systems is denoted as Then, using Convert all point clouds to W i system, and based on the property of laser propagation in a straight line, connect the LiDAR center and all non-diffuse reflection object contour points and extend them; assign the point cloud cluster whose geometric center is in the above three-dimensional extended space the non-diffuse reflection noise point cloud cluster label "1", and the remaining point cloud clusters are assigned the label "0"; Step 4-2: Train the best LightGBM classification model based on the quantized optimal point cloud cluster classification feature subset and labels, and use the model to detect whether each independent point cloud cluster is a non-diffuse noise point cloud cluster, and directly eliminate the non-diffuse 3D noise point cloud according to the detection results.
Citation Information
Cited By
Building structure intelligent reverse modeling and analysis system based on BIM
CN120910977A
BIM-based building structure intelligent reverse modeling and analysis system
CN120910977B