A Dynamic Point Cloud Segmentation and Prediction Label Correction Method for an Inspection Robot

Through dynamic point cloud segmentation and prediction tag correction methods, the problem of low positioning accuracy of 3D laser SLAM system in dynamic environment is solved, accurate identification and segmentation of dynamic point clouds is realized, and the efficiency and accuracy of robot patrol are improved.

CN119671908BActive Publication Date: 2025-06-20SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411652127.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-06-20
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

The existing 3D laser SLAM system has low robustness to dynamic objects in a shared human-machine maintenance channel environment, resulting in a reduced positioning accuracy and affecting the robot patrol efficiency and accuracy.

Method used

Dynamic point cloud segmentation and predictive tag correction methods are adopted to extract the appearance and motion characteristics of point clouds through sliding time windows, use the dual-branch network of appearance-motion bridge to process noise and occlusion, and correct the wrong tags through density-based spatial clustering.

Benefits of technology

Effectively identify and segment dynamic point clouds, improve the accuracy of dynamic point recognition, reduce the interference of dynamic point clouds to the SLAM system, and thus improve the working efficiency and positioning accuracy of under-vehicle patrol robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119671908B_ABST
    Figure CN119671908B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for dynamic point cloud segmentation and prediction label correction of an inspection robot, the method comprising: the method for dynamic point cloud segmentation and prediction label correction is mainly divided into two parts: dynamic point cloud recognition and prediction point cloud label correction; through-filtering the collected point cloud data, and performing dimension reduction projection on the region of interest to optimize processing efficiency, using a sliding time window to extract the appearance and motion features of the pre-processed point cloud, and using a dual-branch network of appearance-motion bridging to solve the erroneous residuals that may be caused by noise and occlusion, and finally generating a predicted moving object label; in the prediction process, for possible mislabeling problems, a density-based spatial clustering method is used to cluster the mislabels and recalibrate them. The method can separate dynamic point clouds from input data, enhance the segmentation effect through label correction, and improve the accuracy of dynamic point recognition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical Field

[0002] The present invention relates to the technical field of robot data processing, and particularly to a method for dynamic point cloud segmentation and prediction label correction of an inspection robot. Background Art

[0003] Train underbody inspection robots have received wide attention due to their potential value in enhancing rail transit safety and maintenance efficiency. The main task of a train underbody inspection robot is to regularly inspect the underbody of a subway train in a maintenance trench to ensure operational safety. The robot moves in the maintenance trench, stops at preset inspection points, and collects images and features of the components at the bottom of the train through a camera at the end of its robotic arm. Based on the collected data, the robot can evaluate the maintenance requirements of the train and provide maintenance suggestions. During the inspection process, the robot uses a hybrid solid-state LiDAR (Light Detection and Ranging) for Simultaneous Localization and Mapping (SLAM).

[0004] The environmental features in the maintenance trench are highly repetitive and usually in a long and narrow shape, belonging to a degenerate scene. In the absence of artificial external aids such as QR codes or magnetic strips, a SLAM system based on 3D LiDAR is required to achieve robot positioning. Currently, the robustness of 3D laser SLAM systems to dynamic object interference is significantly lower than that of 2D laser SLAM systems, which poses difficulties for actual robot operations in a shared human-robot maintenance trench environment. During the deployment teaching or autonomous inspection of the robot, the deployment personnel or the maintenance personnel in human-robot collaboration must leave the maintenance pit when the robot needs to move to the next position. After the robot stops, the personnel re-enter the maintenance pit for robotic arm teaching or collaborative maintenance to avoid contamination of the robot's point cloud map by dynamic point clouds generated by "personnel following the robot", thereby affecting the positioning accuracy. Repeatedly entering and leaving the maintenance trench affects the efficiency of deployment teaching and reduces the flexibility of human-robot collaboration. Summary of the Invention

[0005] In order to overcome the deficiencies of the prior art, the present invention provides a method for dynamic point cloud segmentation and prediction label correction of an inspection robot. It can be added as a data preprocessing module for the SLAM front end to solve the impact of dynamic point clouds on the SLAM system of the train underbody inspection robot.

[0006] A method for dynamic point cloud segmentation and prediction label correction of an inspection robot installs a hybrid solid-state 3D lidar on the robot body to achieve the acquisition of three-dimensional point cloud data, autonomous mapping, and path planning. The industrial control computer is responsible for running the core algorithm to achieve the functions of dynamic point cloud segmentation and prediction label correction.

[0007] The dynamic point cloud segmentation and predicted label correction method mainly consists of two parts: dynamic point cloud recognition and predicted point cloud label correction. Due to the large amount of point cloud data and heavy computational burden, this method performs a pass-through filter on the collected point cloud data and conducts a dimensionality reduction projection on the region of interest to optimize the processing efficiency. The appearance and motion features of the preprocessed point cloud are extracted using a sliding time window, and a dual-branch network with appearance-motion bridging is used to address the errors in residuals that may be caused by noise and occlusion, ultimately generating predicted labels for moving objects. During the prediction process, to address the possible problem of mislabels, a density-based spatial clustering method is used to cluster and analyze the mislabels and recalibrate them to improve the accuracy of the labels.

[0008] The implementation steps of the dynamic point cloud segmentation and predicted label correction method are as follows:

[0009] Step S1, LiDAR input preprocessing: Define the point cloud data as: To prevent interference from objects outside the inspection channel, a pass-through filter is used to crop the point cloud, and only the point cloud within the inspection channel area is extracted. For solid-state LiDAR with a small field of view and non-repeating scans, a bird's-eye view (BEV) projection method is used to project the 3D point cloud into a polar coordinate system to balance the uneven spatial distribution of the point cloud. Define F (u,v),i as the set of all points within the Grid th th cell defined by the (u, v) coordinates in the polar coordinate BEV image, and its expression is:

[0010]

[0011] where F (u,v),i represents the set of all points belonging to the point cloud frame F i in the grid cell defined by the (u, v) coordinates in the polar coordinate top-down image. p j represents a point in the point cloud frame F i , located at the polar coordinates (θ j , p j ) of the point cloud. This inequality represents the range that the angle θ j of the point p j satisfies. θ max and θ min are the maximum and minimum values of the angle respectively. h represents the height of the polar coordinate top-down image, and the angle is divided into h units accordingly. v represents the row index in the polar coordinate image. This condition restricts that ρ j must be within the angle range of the v-th row. This inequality represents the range that the distance ρ j of the point p j satisfies. ρ max and ρmin They are the maximum and minimum values of the distance respectively. w represents the width of the polar coordinate top - view image, and the distance is divided into w units accordingly. u represents the column index in the polar coordinate image. This condition restricts that p j must be within the distance range of the u - th column.

[0012] Step S2, Dynamic Object Recognition Module: When processing the current frame F i , integrate the data of the current frame and its previous and next frames, and construct a sliding window for comprehensive analysis. First, use the laser odometer to determine the transformation matrix that describes the spatial evolution between the current frame and the previous n frames, expressed as:

[0013]

[0014] where represents the overall transformation matrix from time i - n to time i. This matrix transforms the point cloud data at time i - n to the coordinate system at time i. represents the single - step transformation matrix from time i - n to time i - k + 1. This matrix can be regarded as the pose transformation of the (i - k + 1)-th frame relative to the i - n - th frame. means multiplying item - by - item from k = 1 to k = n to obtain the overall transformation This represents the step - by - step accumulation of multiple frames.

[0015] Subsequently, use the transformation matrix to calculate the corresponding point cloud set:

[0016]

[0017] where represents the new point cloud frame obtained by transforming the point cloud frame F i-n from time i - n to the coordinate system at time i. This set contains all the points in F i-n that, after being transformed by T i i-n , are located in the coordinate system at time i. represents the overall transformation matrix from time i - n to time i. This matrix performs coordinate transformation on each point p j to transform the point cloud of F i-n from the coordinate system at time i - n to the coordinate system at time i. p j represents a point in the point cloud frame F i-n . The point p j has coordinates in the coordinate system at time i - n.

[0018] Motion features required to identify dynamic objects are extracted by examining height differences in the BEV image. To address issues caused by low point cloud density, point clouds extracted from two consecutive time windows W1 and W2 are used. Motion features are generated by evaluating the height variance of corresponding grid cells over two time periods.

[0019] First, self-motion compensation is performed through relative pose transformation to align W1 and W2 to the current local coordinate system. Define H Grid,a as the height occupied by the grid in Wa for each pixel value, and its expression is:

[0020] H Grid,a = Max{Z Grid,a}-Min{Z Grid,a}, a ∈ {1, 2}

[0021] where H Grid,a represents the height difference of the unit grid Grid in the a-th sliding window. This value is represented by the difference between the maximum height and the minimum height within the window. Z Grid,a represents all the height values of the unit grid Grid in the a-th sliding window, that is, the height of each point within the grid area. Max{Z Grid,a} represents the maximum height of all points in the unit grid Grid in the a-th sliding window. Min{Z Grid,a} represents the minimum height of all points in the unit grid Grid in the a-th sliding window.

[0022] Among them,

[0023] Z Grid,a = {z j ∈ p j | p j ∈ W Grid,a,}

[0024] where Z Grid,a represents the set of height values of the unit grid Grid in the a-th sliding window. z j represents the Z coordinate of a specific point, that is, the value of point p j on the Z axis. p j represents a point in grid a. W Grid,a, represents the Grid grid area in the a-th sliding window.

[0025] Define as the motion feature in the n-th channel of the i-th frame. The value of n is determined according to the position of the frame within the time window. When a new frame arrives, both time windows move to the next position. The residual is obtained by subtracting H1 and H2:

[0026]

[0027] Among them, These are the motion features of each channel of the grid cell Grid at different time frames. Represents the motion feature of the 0th channel of the i-th frame. Represents the motion feature of the 1st channel of the (i - 1)-th frame. Represents the motion feature of the 2nd channel of the (i - 2)-th frame. H Grid,1 and H Grid,2 respectively represent the height difference of the grid cell Grid in the first sliding window and the height difference of the second sliding window on the grid cell Grid, that is, the difference between the maximum and minimum heights in each area. H Grid,1 -H Grid,2 This is the height difference between the first sliding window unit and the second sliding window. By calculating the height difference, the residual is obtained, that is, the height change amount between the two sliding windows.

[0028]

[0029] Among them, These are the motion features of each channel of the grid cell Grid at different time frames. Represents the motion feature of the 3rd channel of the (i - 3)-th frame. Represents the motion feature of the 4th channel of the (i - 4)-th frame. Represents the motion feature of the 5th channel of the (i - 5)-th frame. H Grid,1 and H Grid,2 respectively represent the height difference of the first sliding window on the grid cell Grid and the height difference of the second sliding window on the grid cell Grid, that is, the difference between the maximum and minimum heights in each area. H Grid,2 -H Grid,1 This is the height difference between the second sliding window unit and the first sliding window. By calculating the height difference, the residual is obtained, that is, the height change amount between the two sliding windows.

[0030] Step S3, Network architecture: Based on the dual-branch structure of PolarNet, combined with the appearance-motion co-attention mechanism. The maximum pooling operation is used to capture the distribution characteristics of the BEV projected point cloud within the vertical range of the grid cell. To enhance the cross-modal interaction between the appearance and motion features, an appearance-motion co-attention module (AMCM) is introduced. This module consists of two important components: a co-attention gate for aligning the salient features of the two modalities, and a motion-guided attention module based on motion cues for refining the attention focus according to the motion cues.

[0031] Step S4. Segmentation correction based on spatial clustering: This dynamic point cloud recognition method may still encounter misrecognition problems, and some dynamic point clouds are not correctly labeled. To solve this problem, a combination of density-based spatial clustering (DBSCAN) and cloth simulation filtering (CSF) is adopted. DBSCAN is a density-based spatial clustering algorithm that can identify and segment high-density regions in a dataset into multiple clusters while identifying isolated points as noise. The core of this algorithm lies in identifying core points within the neighborhood and performing clustering based on this. Specifically, DBSCAN requires two parameters, namely the neighborhood radius ε (eps) and the minimum number of points minPts required to form a dense region. The ε-neighborhood of the starting point is retrieved. If the number of points within the neighborhood exceeds minPts, then this point is defined as the core point of the cluster, and a new cluster is initiated; otherwise, this point is labeled as noise. Once a point is recognized as a dense part of the cluster, the points within its ε-neighborhood will also be included in the same cluster. If these points also meet the density condition, then the points within their ε-neighborhoods are continuously added to this cluster. This process is iterated until all density-connected points in the cluster are recognized. After completely determining the current cluster, the algorithm reselects new unvisited points for clustering recognition or noise labeling until all data points are processed.

[0032] Considering the existence of ground point clouds, there is a correlation between spatial point clouds. DBSCAN clustering requires that the point clouds in the target cluster meet the density reachability condition, so it is easy to cause a large number of ground and non-ground point clouds to gather into the same cluster. For this reason, CSF filtering is applied to effectively separate the ground point clouds, enabling the non-ground point clouds to meet the requirements of DBSCAN clustering, thereby more accurately identifying and segmenting the dynamic point clouds.

[0033] The beneficial technical effects of the present invention are as follows:

[0034] In view of the working scenario characteristics of the underbody inspection robot, the present invention proposes a low-cost hybrid solid-state Li DAR dynamic point cloud segmentation method. This method can separate dynamic point clouds from the input data and enhance the segmentation effect through label correction, improving the accuracy of dynamic point recognition. This method can effectively identify and segment dynamic point clouds, minimize the interference of dynamic point clouds on SLAM, and further improve the overall working efficiency and accuracy of the underbody inspection robot. Description of the Drawings

[0035] Figure 1 is a schematic diagram of the robot working scenario structure;

[0036] Figure 2 is the overall technical flow chart of the present invention;

[0037] Figure 3 is the simulated inspection channel diagram of the present invention;

[0038] Figure 4 It is a diagram of the dynamic point cloud recognition module of the present invention;

[0039] Figure 5 It is a flowchart of the dynamic point cloud label clustering and correction of the present invention;

[0040] Figure 6 It is a diagram of the result of the dynamic point cloud label clustering and correction of the present invention; Specific implementation manner

[0041] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following further describes this application in detail with reference to the drawings and specific embodiments.

[0042] As Figure 1 shown, a method for dynamic point cloud segmentation and predicted label correction of an inspection robot is installed with: a Livox-Mid 70 solid-state lidar, a Zivid 2 camera, a UR5e robotic arm, and an industrial computer, and dynamic object point cloud recognition and segmentation are performed in the simulated inspection channel built. The steps are as follows:

[0043] Combined with Figure 2 shown, it specifically includes the following steps.

[0044] Step S1, LiDAR input preprocessing: Define the point cloud data as: To prevent interference from objects outside the simulated inspection channel, a pass-through filter is used to crop the point cloud, and only the point cloud within the simulated inspection channel area is extracted. Here, the point cloud with a height of 2m is extracted. For the solid-state LiDAR with a small field of view and non-repeating scanning, the bird's-eye view (BEV) projection method is used to project the 3D point cloud into the polar coordinate system to balance the uneven spatial distribution of the point cloud. Define F (u,v),i as the set of all points in the Grid th th cell defined by the (u, v) coordinates in the polar coordinate BEV image, and its expression is:

[0045]

[0046] where F (u,v),i represents the set of all points belonging to the point cloud frame F i in the grid cell defined by the (u, v) coordinates in the polar coordinate top-down image. p j represents a point in the point cloud frame F i , located at the polar coordinates (θ j , p j ) of the point cloud. This inequality represents the range satisfied by the angle θ j of the point p j θmax and θ min are the maximum and minimum values of the angle respectively. h represents the height of the polar coordinate top - view image, and the angle is divided into h units based on this. v represents the row index in the polar coordinate image. This condition restricts ρ j to be within the angle range of the v - th row. This inequality represents the distance ρ j of the point p j satisfied range. ρ max and ρ min are the maximum and minimum values of the distance respectively. w represents the width of the polar coordinate top - view image, and the distance is divided into w units based on this. u represents the column index in the polar coordinate image. This condition restricts p j to be within the distance range of the u - th column.

[0047] Step S2, Dynamic Object Recognition Module: When processing the current frame F i , integrate the data of the current frame and its previous and next frames, and construct a sliding window for comprehensive analysis. First, use the laser odometer to determine the transformation matrix that describes the spatial evolution between the current frame and the previous n frames. Here, n = 5, which is expressed as:

[0048]

[0049] Among them, represents the overall transformation matrix from time i - n to time i. This matrix transforms the point cloud data at time i - n to the coordinate system at time i. represents the single - step transformation matrix from time i - n to time i - k + 1. This matrix can be regarded as the pose transformation of the (i - k + 1)-th frame relative to the (i - n)-th frame. represents multiplying term by term from k = 1 to k = n to obtain the overall transformation This represents the step - by - step accumulation of multiple frames.

[0050] Subsequently, use the transformation matrix to calculate the corresponding point cloud set:

[0051]

[0052] Among them, represents the new point cloud frame obtained by transforming the point cloud frame F i-n at time i - n to the coordinate system at time i. This set contains all the points in F i-n that, after being transformed by , are located in the coordinate system at time i. represents the overall transformation matrix from time i - n to time i. This matrix performs coordinate transformation on each point p j and transforms F i-nThe point cloud is transformed from the coordinate system at time i - n to the coordinate system at time i. p j represents a point cloud frame F i-n and a point in it. Point p j is the coordinate in the coordinate system at time i - n.

[0053] Motion features required to identify dynamic objects are extracted by examining the height differences in the BEV image. To address the problem caused by the low point cloud density, point clouds extracted from two consecutive time windows W1 and W2 are used. Motion features are generated by evaluating the height variance of corresponding grid cells within two time periods.

[0054] First, self - motion compensation is performed through relative pose transformation to align W1 and W2 to the current local coordinate system. Define H Grid,a as the height occupied by each pixel value in the grid in Wa, and its expression is:

[0055] H Grid,a = Max{Z Grid,a}-Min{Z Grid,a}, a ∈ {1, 2}

[0056] where H Grid,a represents the height difference of the unit grid Grid in the a - th sliding window. This value is represented by the difference between the maximum height and the minimum height within the window. Z Grid,a represents all the height values of the unit grid Grid in the a - th sliding window, that is, the height of each point within the grid area. Max{Z Grid,a} represents the maximum height of all points in the unit grid Grid in the a - th sliding window. Min{Z Grid,a} represents the minimum height of all points in the unit grid Grid in the a - th sliding window.

[0057] Z Grid,a = {z j ∈ p j | p j ∈ W Grid,a,}

[0058] where Z Grid,a represents the set of height values of the unit grid Grid in the a - th sliding window. z j represents the Z - coordinate of a specific point, that is, the value of point p j on the Z - axis. p j represents a point in grid a. W Grid,a, represents the Grid grid area in the a - th sliding window.

[0059] Define Represents the motion feature within the n-th channel of the i-th frame, where the value of n is determined according to the position of the frame within the time window. When a new frame arrives, both time windows shift to the next position. The residual is obtained by subtracting H1 and H2:

[0060]

[0061] Among them, These are the motion features of each channel of the grid cell Grid at different time frames. Represents the motion feature of the 0-th channel of the i-th frame. Represents the motion feature of the 1-st channel of the (i - 1)-th frame. Represents the motion feature of the 2-nd channel of the (i - 2)-th frame. H Grid,1 and H Grid,2 respectively represent the height difference of the grid cell Grid in the first sliding window and the height difference of the second sliding window on the grid cell Grid, that is, the difference between the maximum and minimum heights in each region. H Grid,1 -H Grid,2 This is the height difference between the first sliding window unit and the second sliding window. By calculating the height difference, the residual is obtained, that is, the height change amount between the two sliding windows.

[0062]

[0063] Among them, These are the motion features of each channel of the grid cell Grid at different time frames. Represents the motion feature of the 3-rd channel of the (i - 3)-th frame. Represents the motion feature of the 4-th channel of the (i - 4)-th frame. Represents the motion feature of the 5-th channel of the (i - 5)-th frame. H Grid,1 and H Grid,2 respectively represent the height difference of the first sliding window on the grid cell Grid and the height difference of the second sliding window on the grid cell Grid, that is, the difference between the maximum and minimum heights in each region. H Grid,2 -H Grid,1 This is the height difference between the second sliding window unit and the first sliding window. By calculating the height difference, the residual is obtained, that is, the height change amount between the two sliding windows.

[0064] Step S3, Network Architecture: Based on the dual-branch structure of PolarNet, combined with the appearance-motion co-attention mechanism. The maximum pooling operation is used to capture the distribution characteristics of the BEV projected point cloud within the vertical range of the grid cells. To enhance the cross-modal interaction between appearance and motion features, an appearance-motion co-attention module (AMCM) is introduced. This module consists of two important components: a co-attention gate for aligning the salient features of the two modalities, and a motion-guided attention module based on motion cues for refining the attention focus according to motion cues.

[0065] Step S4, Segmentation Correction Based on Spatial Clustering: This dynamic point cloud recognition method may still encounter misrecognition problems, and some dynamic point clouds are not correctly labeled. To solve this problem, a combination of density-based spatial clustering (DBSCAN) and cloth simulation filtering (CSF) is adopted. DBSCAN is a density-based spatial clustering algorithm that can identify and segment high-density regions in the dataset into multiple clusters, while identifying isolated points as noise. The core of this algorithm lies in identifying the core points within the neighborhood and clustering based on this. Specifically, DBSCAN requires two parameters, namely the neighborhood radius ε (eps) and the minimum number of points minPts required to form a dense region. The ε-neighborhood of the starting point is retrieved. If the number of points within the neighborhood exceeds minPts, this point is defined as the core point of the cluster, and a new cluster is started; otherwise, this point is labeled as noise. Once a point is identified as a dense part of the cluster, the points within its ε-neighborhood will also be included in the same cluster. If these points also meet the density condition, the points within their ε-neighborhoods will continue to be added to this cluster. This process continues iteratively until all density-connected points in the cluster are identified. After completely determining the current cluster, the algorithm reselects a new unvisited point for cluster identification or noise labeling until all data points are processed. Here, ε (eps) = 0.5 and minPts = 10 are taken.

[0066] Considering the existence of ground point clouds, there is a correlation between spatial point clouds. DBSCAN clustering requires the point clouds in the target cluster to meet the density reachability condition, so it is easy to cause a large number of ground and non-ground point clouds to gather into the same cluster. For this reason, CSF filtering is applied to effectively separate the ground point clouds, enabling the non-ground point clouds to meet the requirements of DBSCAN clustering, thereby more accurately identifying and segmenting dynamic point clouds.

[0067] As Figure 3 shown, it presents a comparison between the real inspection channel and the simulated channel built with cardboard boxes. The real channel map shows the inspection scenario where the inspection robot actually operates, while the simulated inspection channel uses materials such as cardboard boxes to construct the basic structure of the channel, providing a controlled experimental environment.

[0068] Combined with Figure 4The dynamic point cloud recognition module shown in the figure conducts experiments in the simulated train inspection trench built. Considering the actual application environment of the algorithm design, only people are classified as dynamic objects in the dataset, while other objects such as walls are classified as static background point clouds. The specific steps are as follows:

[0069] Step Q1, time window processing: Push the current scan data into the time window and convert all the scan data within the window into point cloud data in the current perspective.

[0070] Step Q2, motion feature extraction: Divide the 6-frame point cloud data into two sliding time windows for processing again, so as to extract motion features for each frame of point cloud data. Specifically, after 6 motion iterations, motion features of 6 channels are extracted, and these features represent the motion information of the point cloud changing over time. More specifically, Figure 4 The black, white, and gray circles respectively simulate the existence states of the point cloud of pedestrians in different frames during the movement process after the point cloud of each frame is converted to the current frame in the sliding window. This process can effectively capture the motion trajectory of dynamic objects and provide a key basis for subsequent dynamic point cloud recognition and positioning.

[0071] Step Q3, appearance feature learning: Pop the scan data from the time window and use the PointNet algorithm to learn the appearance features of the pedestrian point cloud and extract the geometric feature information of the point cloud.

[0072] Step Q4, spatio-temporal information fusion: Input the appearance features and motion features into the dual-branch network bridged by AMCM for joint learning and fusion of spatio-temporal information.

[0073] In terms of the dataset format, we refer to Semantic KITTI and normalize the dataset coordinates. Table 1 presents the results of comparison with the PointNet and PointNet++ algorithms, focusing on two metrics: ACC (accuracy) and mIoU (mean intersection over union). To better understand the content in the table, the following is a detailed introduction to these two algorithms and these two evaluation metrics. PointNet is a network model that directly processes 3D point clouds without converting the point cloud data into meshes or voxels. It aggregates the features of each point to handle the disorder of the point cloud. PointNet uses a shared multi-layer perceptron (MLP) to extract features point by point and finally summarizes them into global features. Its advantages are simple structure and high computational efficiency, making it suitable for quickly processing simple geometric structures. However, its limitation is that it is not good at capturing local geometric features and may perform poorly in processing complex or detail-rich point clouds. PointNet++ enhances the ability to capture local structures based on PointNet and is suitable for point clouds with multi-scale and complex geometries. Through hierarchical sampling, it extracts and aggregates multi-scale local features and summarizes the features within the local area through neighborhood definitions (such as spherical neighborhoods), thus better adapting to complex geometric environments. Its advantage is that it is good at capturing multi-scale geometric details and is applicable to scenes with rich details. However, its computational cost is high and the processing speed is relatively slow. ACC represents the overall accuracy of the model in classifying or segmenting point cloud data, and the specific definition is:

[0074]

[0075] The higher the Acc value, the higher the overall accuracy of the model in recognition and classification tasks. mIoU (Mean Intersection over Union) is a commonly used metric for segmentation tasks and evaluates the segmentation performance by comparing the overlapping area between the model prediction and the actual label. mIoU is the average of the IoU for all classes, and the calculation formula for IoU is:

[0076]

[0077] mIoU is the average of the IoU for all classes and is used to evaluate the overall accuracy in multi-class segmentation tasks:

[0078]

[0079] Among them, the higher the mIoU value, the stronger the ability of the model to distinguish all classes in the segmentation task; (detailed introduction and explanation of the algorithms in the table and what ACC represents) Figure 6 Shows the results in the case of using the label correction method. Figure 6Among them, for two randomly selected frames of point cloud data, it can be seen that the point cloud predicted by the point cloud recognition module is not completely correctly labeled, and there are parts of the pedestrian's point cloud data that are not correctly recognized. Combining Figure 5 with the dynamic point cloud label clustering and correction method, the specific process includes:

[0080] Step R1, point cloud segmentation and clustering: First, for the original point cloud data, the CSF filtering algorithm is used to divide the original point cloud into ground point cloud and non-ground point cloud. Subsequently, the DBSCAN clustering algorithm is used for clustering the non-ground point cloud part.

[0081] Step R2, label correction: According to the clustering results, the labels are corrected. Specifically, if all the point clouds in a certain cluster are marked with the same label in the clustering results, but in the prediction results, some of the point clouds belonging to this cluster are mislabeled as other labels, then these mislabeled point clouds in the prediction are uniformly corrected to a unified label. This correction strategy assumes that the point clouds of the same type of object have a relatively high density in space and have similar attributes, so the clustering label can effectively be used as the basis for correcting the prediction label.

[0082] Finally, it can be seen that Figure 6 in the misidentified pedestrian prediction point cloud labels are unified into the same category of point cloud labels.

[0083] Table 1 Results of comparing the PointNet and PointNet++ algorithms

[0084] Method mIoU Acc PointNet 84.04% 91.28% PointNet++ 87.28% 93.17% The method proposed in this paper 89.21% 95.62%

[0085] After inspection, this dynamic point cloud segmentation and prediction label correction design method can be applied to robots equipped with low-cost hybrid solid-state lidar. This method effectively segments dynamic point clouds using appearance and motion features in 2D BEV. By integrating the spatial density clustering results, the segmentation accuracy and dynamic point recognition ability are improved. Experiments in a simulated channel environment verify the effectiveness of the method.

[0086] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "set", "install", "connected", "connected", "fixed" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0087] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various equivalent changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for dynamic point cloud segmentation and prediction label correction of an inspection robot, characterized in that , the method comprises: Step S1, laser radar input preprocessing: define point cloud data as p j =[x j ,y j ,z j ] T , where F i represents the point cloud data of the i-th frame; p j represents the jth point cloud point, which belongs to the point cloud data of the i-th frame; M represents the total number of point clouds in the i-th frame point cloud data; R 3 Represents three-dimensional space, each point of the point cloud data is in three-dimensional space; j =[x j ,y j ,z j ] T represents the three-dimensional coordinates of the jth point in the point cloud; x j ,y j 、z j They represent the coordinate values ​​of the point in the x, y, and z axes respectively. A straight-through filter is used to crop the point cloud and only extract the point cloud in the inspection channel area. For solid-state laser radars with a small field of view and non-repetitive scanning, a bird's-eye view projection method is used to project the 3D point cloud into a polar coordinate system to balance the uneven spatial distribution of the point cloud. Define F (u,v),i is the grid defined by the (u,v) coordinates in the polar coordinate bird's-eye view image th The set of all points within a cell; Step S2, dynamic object recognition module: when processing the current frame F i When , the data of the current frame and its previous and next frames are integrated to construct a sliding window for comprehensive analysis; Step S3, network architecture: Based on the dual-branch structure of PolarNet, combined with the appearance-motion co-attention mechanism, the maximum pooling operation is used to capture the distribution characteristics of the bird's-eye view projection point cloud within the vertical range of the grid unit, and the appearance-motion co-attention module is introduced to enhance the cross-modal interaction between appearance and motion features; Step S4, segmentation correction based on spatial clustering: A combination of density-based spatial clustering algorithm and cloth simulation filtering is used to identify misidentification and some dynamic point clouds are not correctly labeled. The density-based spatial clustering algorithm is used to identify and segment the high-density areas in the data set into multiple clusters, and identify isolated points as noise.

2. The method for dynamic point cloud segmentation and predicted label correction of an inspection robot according to claim 1, characterized in that: The F (u,v),i , the expression is: Among them, F (u,v),i Represents the polar coordinate top view image in the grid cell defined by the (u,v) coordinates, belonging to the point cloud frame F i The set of all points, p j Represents the point cloud frame F i A point in the point cloud, located at the polar coordinates (θ j ,p j )Down, This inequality means that the point p j The angle θ j The range of satisfaction, θ max and θ min are the maximum and minimum values ​​of the angle, respectively. h represents the height of the polar coordinate top view image, which divides the angle into h units. v represents the row index in the polar coordinate image. This condition constrains ρ j must be within the angle range of row v, This inequality means that the point p j The distance ρ j The range satisfied, ρ max and ρ min are the maximum and minimum values ​​of the distance, respectively. w represents the width of the polar coordinate top view image, which divides the distance into w units. u represents the column index in the polar coordinate image. This condition constrains p j Must be within the u-th column distance range.

3. The method for dynamic point cloud segmentation and predicted label correction of an inspection robot according to claim 1, characterized in that: The construction of the sliding window for comprehensive analysis includes: The laser odometry is used to determine the transformation matrix describing the spatial evolution between the current frame and the previous n frames, which is expressed as: in, Represents the overall transformation matrix from time in to time i, which converts the point cloud data at time in to the coordinate system at time i. Represents the single-step transformation matrix from time in to time i-k+1. This matrix is ​​the pose transformation of the i-k+1th frame relative to the inth frame. It means that from k = 1 to k = n Multiply each term to get the overall transformation This represents a gradual accumulation of multiple frames; Use the transformation matrix to calculate the corresponding point cloud set: in, It means that after transformation, the point cloud frame F at time in i-n The new point cloud frame converted to the coordinate system of time i contains all the points in F i-n point in After transformation, it is located in the coordinate system of time i. Represents the overall transformation matrix from time in to time i. This matrix is ​​for each point p j Perform coordinate transformation and convert F i-n The point cloud is transformed from the coordinate system of time in to the coordinate system of time i, p j Represents the point cloud frame F i-n A point in j is the coordinate in the coordinate system at time in; The motion features required for identifying dynamic objects are extracted by examining the height differences in the bird's-eye view images. The point clouds extracted from two consecutive time windows W1 and W2 are used to generate motion features by evaluating the height variance of the corresponding grid cells in the two time periods. Compensate for self-motion by relative posture transformation, align W1 and W2 to the current local coordinate system, and define H Grid,a Each pixel value represents the height of the grid in Wa, and its expression is: H Grid,a =Max{Z Grid,a }-Min{Z Grid,a },a∈{1,2} Among them, H Grid,a Indicates the height difference of the cell grid Grid in the ath sliding window. This value is represented by the difference between the maximum height and the minimum height in the window. Grid,a Represents all the height values ​​of the unit grid Grid in the a-th sliding window, that is, the height of each point in the grid area, Max{Z Grid,a } represents the maximum height of all points in the unit grid Grid in the a-th sliding window, Min{Z Grid,a } represents the minimum height of all points of the cell grid Grid in the a-th sliding window, WITH Grid,a ={z j ∈p j |p j ∈W Grid,a, } Among them, Z Grid,a Represents the set of height values ​​of the cell grid Grid in the a-th sliding window, z j Represents the Z coordinate of a specific point, that is, point p j The value on the Z axis; p j represents a point in the grid a, W Grid,a, Represents the Grid area in the a-th sliding window; definition Represents the motion features in the nth channel of the i-th frame. The value of n is determined by the position of the frame in the time window. When a new frame arrives, both time windows move to the next position. The residual is obtained by subtracting H1 and H2: in, These are the motion features of each channel of the grid cell Grid at different time frames. represents the motion features of the 0th channel of the i-th frame, represents the motion features of the 1st channel of the i-1th frame, represents the motion feature of the second channel of the i-2th frame, H Grid,1 and H Grid,2 They represent the height difference of the grid cell Grid in the first sliding window and the height difference of the second sliding window on the grid cell Grid, that is, the difference between the maximum and minimum heights in each area, H Grid,1 -H Grid,2 This is the height difference between the first sliding window unit and the second sliding window. By calculating the height difference, we get the residual, which is the height change between the two sliding windows. in, These are the motion features of each channel of the grid cell Grid at different time frames. represents the motion features of the 3rd channel of the i-3th frame, represents the motion features of the 4th channel of the i-4th frame, represents the motion feature of the 5th channel of the i-5th frame, H Grid,1 and H Grid,2 They represent the height difference of the first sliding window on the grid unit Grid and the height difference of the second sliding window on the grid unit Grid, that is, the difference between the maximum and minimum heights in each area, H Grid,2 -H Grid,1 This is the height difference between the second sliding window unit and the first sliding window. By calculating the height difference, we get the residual, which is the height change between the two sliding windows.

4. The method for dynamic point cloud segmentation and predicted label correction of an inspection robot according to claim 1, characterized in that: The appearance-motion co-attention module consists of two parts: a co-attention gate for aligning the salient features of the two modalities, and a motion-guided attention module based on motion cues for refining the focus points according to the motion cues.

5. The method for dynamic point cloud segmentation and predicted label correction of an inspection robot according to claim 1, characterized in that: The density-based spatial clustering algorithm includes two parameters, the neighborhood radius ε and the minimum number of points minPts required to form a dense area. The ε neighborhood of the starting point is retrieved. If the number of points in the neighborhood exceeds minPts, the point is defined as the core point of the cluster and a new cluster is started; otherwise, the point is marked as noise; once a point is identified as a dense part of a cluster, the points in its ε neighborhood will also be included in the same cluster; If these points also meet the dense condition, continue to add the points in their ε neighborhood to the cluster; This process continues to iterate until all density-connected points in the cluster are identified. After the current cluster is fully determined, the algorithm reselects new unvisited points for cluster identification or noise marking until all data points are processed.

6. The method for dynamic point cloud segmentation and predicted label correction of an inspection robot according to claim 1, characterized in that: This method installs a hybrid solid-state 3D laser radar on the robot body to collect three-dimensional point cloud data, autonomously build maps and plan paths; an industrial computer is used to run the core algorithm to perform dynamic point cloud segmentation and correction of predicted labels.

Citation Information

Patent Citations

  • Dynamic object recognition method based on clustering algorithm

    CN113344112A

  • Automatic generation method of point cloud label and dynamic segmentation method based on convolutional neural network

    CN118334657A