Point cloud traversal analysis network based on self-supervised learning

By introducing a point cloud traversability analysis network based on self-supervised learning into the point cloud traversability analysis technology, qualitative and quantitative analysis is performed using point cloud intraframe and global information, the problem of insufficient robustness and scalability in the existing technology is solved, and efficient and highly adaptable point cloud traversability analysis is achieved.

CN120030396APending Publication Date: 2025-05-23TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311549991.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing point cloud traversability analysis technology is not robust and scalable, and relies on a large amount of manual annotation data and manual rules, making it difficult to adapt to different types of robots and environments.

Method used

A point cloud traversibility analysis network based on self-supervised learning is proposed. Using the information of the point cloud intra-frame and global point cloud, the qualitative and quantitative analysis of the point cloud traversability is realized through the SLAM module, the backbone network module and the traversability analysis module, and the traversibility analysis module is reduced.

Benefits of technology

It realizes efficient point cloud traversability analysis without the need for a large amount of manual data labeling, which improves the generalization ability and robustness of the algorithm, is highly adaptable, and can adapt to a variety of different environments and application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030396A_ABST
    Figure CN120030396A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud traversal analysis network based on self-supervised learning, and aims to solve the problems of poor robustness and poor expandability of the existing point cloud traversal analysis technology. According to the network, a point cloud frame sequence is utilized, and traversal analysis is carried out through three main modules, namely an SLAM module, a backbone network module and a traversal analysis module. The SLAM module is responsible for processing each frame of point cloud, calculating odometer information of the robot and maintaining a point cloud map. The backbone network module receives point cloud frames and speedometer information as input, uses global point cloud features as reference, provides intra-frame context information and global context information for each point in the point cloud of the current frame, and outputs aggregation features. And the trafficability analysis module is used for receiving the aggregated features and carrying out qualitative and quantitative trafficability analysis. According to the network structure, multi-level feature information of the global point cloud is considered, the context information of the current frame point cloud is enriched, direct reprocessing of the historical global point cloud is avoided, and the time cost is reduced. Effective local feature aggregation is realized through local space coding, attention pooling and residual block expansion. Meanwhile, the qualitative analysis subnet and the quantitative analysis subnet are used to realize the qualitative and quantitative analysis of the traversal of the point cloud. The invention provides a network architecture which does not depend on manual labeling data and can accurately analyze the traversal of the point cloud, and provides powerful technical support for navigation and path planning of the autonomous mobile robot in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the environmental perception technology of an autonomous mobile robot, and mainly to the field of point cloud processing and self-supervised learning. Specifically, the present invention provides a point cloud traversability analysis network based on self-supervised learning. Background Art

[0002] After entering the 21st century, robotics has been rapidly expanded in multiple application scenarios through deep integration with artificial intelligence, deep learning, and the Internet of Things. Modern robots have not only achieved automation and optimization in industrial production, but have also found wide applications in fields such as home, medical care, exploration, and entertainment. In recent years, the development of fields such as autonomous mobile robots and driverless cars has also made significant progress, opening up new possibilities for the future application of robotics.

[0003] Autonomous mobile robots are able to perform tasks independently without external intervention and have the ability to make autonomous decisions. With the expansion of application areas, robots are facing the challenge of operating in more complex and diverse environments. The environmental perception ability of autonomous mobile robots depends greatly on the sensor technology they are equipped with. Among them, 3D LiDAR can accurately measure the distance between the robot and the surrounding environment by emitting laser pulses and capturing reflected signals, and construct a three-dimensional point cloud model of the environment. Compared with RGB cameras and ultrasonic sensors, 3D LiDAR has significant advantages in high accuracy, robustness in lighting conditions, and capturing large-scale three-dimensional environmental information. After the raw point cloud of 3D LiDAR is collected, it needs to be further processed to achieve effective environmental perception. The traversability analysis and obstacle detection of the point cloud are key steps to achieve safe and efficient navigation of robots.

[0004] Point cloud traversability analysis is a key technology to ensure that the robot can fully understand the difficulty of traversing various areas of the environment. It helps to avoid the robot from entering areas that may cause jamming or damage. However, existing point cloud traversability analysis solutions, whether traditional traversability determination methods based on manual rules or supervised deep learning methods, all have a certain degree of human intervention. Traditional methods usually rely on manually set traversability determination rules, which are often based on the personal experience of the rule setter. Therefore, the effectiveness of such methods is largely subject to the accuracy of rule setting. Since manually set rules are usually too simple or too focused on specific scenarios, the algorithm may lack generalization ability and robustness in diverse environments. On the other hand, the method of using supervised deep learning for point cloud traversability analysis requires a lot of manual resources for data annotation. The efficiency and effectiveness of this method are largely limited by the quality and quantity of the labeled data. Furthermore, when the environment in which the robot is located is significantly different from the environment of the training data set, the model may need to be retrained to adapt to the new environment, which not only increases the cost of training, but also may not be realistic in the absence of sufficient labeled data. Therefore, finding a point cloud traversability analysis method that can reduce human intervention and has high generalization ability and robustness has become an important research direction in this field.

[0005] Since Langer et al. first proposed the concept of traversability to describe whether a robot can pass through a certain area in 1994, research in this field has made some progress. Langer et al. used stereo cameras as sensors to obtain three-dimensional information of the terrain surface and stored it in the elevation map. By calculating the height difference and slope in each elevation cell, each cell is divided into two states: non-traversable and traversable. Then in 1999, Gennery studied the relationship between ground roughness and traversability. Since roughness is an indicator of small-scale changes in the terrain surface, it is obvious that the roughness of the area of ​​interest is related to traversability. Gennery used the covariance matrix to represent the accuracy of each data point, and then fitted the data points within the unit range through the least squares plane, and calculated the residual value to represent the roughness. Subsequently, he further used a manually constructed cost function to calculate the cost value, which represented the probability of traversability of the area represented by a certain data. This method was applied to the lunar probe. With the development and application of 3D lidar, a single acquisition of data results in a frame of point cloud, which is called a point cloud frame. As the robot moves with the laser radar, the point cloud frames of multiple scans will have overlapping scanning areas, and the traversability of the overlapping areas should be stable. Shan et al. applied the generalized Bayesian theorem to traversability analysis. The study took the laser radar point cloud frame sequence as input and output the height estimate of the corresponding area. Since the regression model is used for all units, the traversability of the overlapping areas of multiple scans can be directly calculated. Molino et al. further subdivided traversability into overall cost and single-step cost, where the overall cost is used to measure the difficulty of the robot exploring all parts of a certain area, and the single-step cost measures the difficulty of the robot moving between two points in the area. Their research has actually noticed that the traversability of different robots passing through the same area is different. It is obviously unreasonable to use exactly the same traversability for all types of robots. The present invention also takes this issue into consideration. However, most of the above-mentioned studies and other related studies rely on artificially formulated rules, and there are doubts about whether they can truly reflect the traversability of mobile robots passing through specific areas. When these methods are applied to other robots or other environments, the rules may need to be adjusted because a single rule cannot adapt to different types of environments. This makes the traversability analysis method based on artificial rules have obvious defects in generalization ability and robustness. Therefore, seeking a point cloud traversability analysis method that reduces manual intervention and has high generalization ability and robustness has become an important research direction in this field.

[0006] In recent years, with the rapid progress of artificial intelligence technology, more and more researchers have begun to use deep learning technology to solve the problem of point cloud traversability analysis. In these deep learning-based methods, the traversability of a certain area is usually simplified into two states: traversable and non-traversable (i.e., obstacle area). In fact, this type of method has many similarities with the task of point cloud semantic segmentation. Both methods use neural networks to assign binary labels of traversable or non-traversable to each point in the point cloud. Therefore, exploring the point cloud semantic segmentation method based on deep learning is of great significance for understanding and solving the problem of traversability analysis. The SegCloud method proposed by Tchapmi et al. is a typical example. It first processes the point cloud data into 3D voxels and then uses a 3D fully convolutional network to process the voxels. The network solves the common problem of coarse granularity in voxelization methods by applying trilinear interpolation. Although Tchapmi et al. experimentally evaluated SegCloud using the Semantic3D dataset, the method does not support real-time processing and has not been verified in unstructured environments. On the other hand, the exploration of directly using neural networks to process point cloud data has also made some important progress. PointNet and its subsequent version PointNet++ proposed by Qi et al. are key works in this direction. They directly apply convolutional neural network (CNN) technology to point cloud data processing. By relying on the inherent symmetry of the maximum pooling layer to backpropagate signals to train the network, PointNet++ even expands small "neighborhoods" in the point cloud, allowing the network to learn features from different scales. Based on the work of the PointNet series, a variety of successful point cloud semantic segmentation networks have been constructed, such as Semantic3D proposed by Hackel et al. and ScanNet benchmark dataset proposed by Dai et al. These network models are mainly used for semantic segmentation tasks in large scenes. For the point cloud traversability analysis task, Sock et al. used a supervised deep learning network to perform traversability analysis on the data generated by the lidar and camera respectively, and fused the two independently constructed probability maps through the Bayesian rule, thereby improving the detection performance. However, these supervised deep learning-based methods rely on a large amount of manually annotated data for training, and do not fully consider the problem that the traversability of different robots may be different when traversing the same area. These problems show that although some progress has been made in point cloud traversability analysis through deep learning technology, there are still many challenges and problems to be solved, especially in reducing dependence on manual annotation, improving algorithm generalization ability and handling traversability analysis of different types of robots. Summary of the invention

[0007] The present invention aims to solve the deficiencies of existing point cloud traversability analysis technologies in terms of robustness and scalability, and is committed to realizing traversability analysis of point cloud data without the need for manual data annotation. To this end, a point cloud traversability analysis network based on self-supervised learning is proposed. In the point cloud feature extraction stage, the present invention utilizes two different information sources, namely, the point cloud frame and the global point cloud, and fully considers the connection between single point cloud frame data and historical point cloud data and their respective characteristics. In terms of traversability analysis function, the present invention not only qualitatively analyzes the traversability of the point cloud, but also performs quantitative analysis, thereby providing a comprehensive and accurate traversability assessment. In this way, the present invention provides strong technical support for the safe and efficient navigation of autonomous mobile robots in different environments.

[0008] The technical solution of the present invention is as follows:

[0009] SLAM module processing: For the input point cloud frame sequence, each point cloud frame first passes through the first part of the present invention, namely the SLAM (Simultaneous Localization And Mapping) module. This module is mainly responsible for calculating the robot's odometer information based on the current point cloud frame, and maintaining the point cloud map. Although the SLAM module usually belongs to the autonomous positioning function part of the autonomous mobile robot and is downstream of environmental perception, the network of the present invention has already utilized the processing results of the current point cloud frame by the SLAM module in the environmental perception stage. This design is based on the framework of the present invention, which aims to utilize the contextual information of the current point cloud in the historical point cloud to enrich the information of the current frame, thereby enabling the perception architecture to obtain more accurate analysis results, and avoids the complete re-processing of the merged global point cloud, reducing the waste of computing resources.

[0010] Feature enhancement and backbone network module: The second part of the network of the present invention is the backbone network module. Before entering the backbone network, the features of the current point cloud frame are first enhanced, including adding a timestamp and a vector from the laser radar position to the point at the time of acquisition as additional features to further enhance the information of the current point cloud frame. The timestamp can provide time information for point cloud acquisition, which helps to understand the temporal changes of the point cloud. At the same time, the position of the laser radar in the world coordinate system can be calculated from the odometer information output by the SLAM module. The vector value enables the continuity of some single point cloud frames to be retained after the current point cloud frame is merged into the global point cloud, which helps to understand the geometric changes of the point cloud. The backbone network module uses the odometer information, global point cloud information and global point cloud historical features output by the SLAM module to extract the features of the current point cloud frame. First, based on the effective range of the laser radar and the position information of the current point cloud frame in the global point cloud, points related to the current point cloud frame in the global point cloud are selected. Subsequently, when extracting features from points in the current point cloud frame, neighboring points in the same frame and in the global point cloud are found at the same time, and feature enhancement is performed based on these two parts of neighboring points. This design enables the backbone network module to effectively utilize the contextual information of historical point cloud data and analysis results, so that the network has a more comprehensive and in-depth understanding of the current point cloud frame, while avoiding the high time cost of directly processing the global point cloud.

[0011] Traversability analysis module: The third part is the traversability analysis module. After the backbone obtains the features of the current point cloud frame, the module performs qualitative and quantitative analysis on the traversability of the point cloud. The qualitative analysis stage of traversability aims to predict the traversability of the area represented by each point in the point cloud. Specifically, an anomaly detection method is used to determine whether the feature of a point is located in a defined traversable hypersphere in the feature space, thereby realizing the qualitative analysis of the traversability of the point cloud. In order to prevent the collapse of the hypersphere, the present invention adopts a self-labeling method to maintain the stability and effectiveness of the hypersphere. After completing the qualitative analysis, the present invention realizes the quantitative analysis of the traversable point cloud by predicting the motion change factor. The motion change factor is a parameter innovatively proposed in this study, which is used to evaluate the ability of a certain point to affect the robot's motion state. Through this parameter, not only can the traversability of the point cloud be evaluated more finely, but also the network's ability to understand the environment can be enhanced. Similarly, the self-labeling method is also applied in this stage, which helps to further enhance the network's understanding ability and prediction accuracy.

[0012] The present invention has the following beneficial effects:

[0013] By adopting a self-supervised learning method, the present invention can autonomously learn and optimize the point cloud traversability analysis model without the need for a large amount of manually labeled data. This greatly reduces the labor cost of model training and improves the system's adaptability and practicality.

[0014] The present invention makes comprehensive use of the information in the point cloud frame and the global point cloud, and fully considers the connection and characteristics between the single point cloud frame data and the historical point cloud data when extracting point cloud features. This helps to build a more accurate and efficient environmental perception model, thereby improving the navigation and decision-making capabilities of autonomous mobile robots in complex environments. At the same time, it avoids the need to completely reprocess the merged global point cloud, thereby reducing the consumption of computing resources and improving processing efficiency.

[0015] The present invention can perform qualitative and quantitative analysis on the traversability of point clouds, providing the robot with more comprehensive and accurate environmental information. This not only helps to improve the robot's environmental perception ability, but also provides important technical support for achieving safe and efficient autonomous navigation. By introducing the self-labeling method and the design of motion change factors, the present invention shows good generalization ability, can adapt to a variety of different environments and application scenarios, and has high practical value and promotion potential.

[0016] The design and implementation of the present invention provides a useful demonstration and reference for further research and development of point cloud processing technology, especially in point cloud traversability analysis and self-supervised learning applications, and makes a positive contribution to the technological progress in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Point cloud traversability analysis network architecture diagram based on self-supervised learning;

[0018] Figure 2 The backbone network architecture diagram of the network of the present invention;

[0019] Figure 3 The structure diagram of each unit of local feature aggregation in the backbone network of the network of the present invention;

[0020] Figure 4 A diagram of the network architecture for traversability analysis in the network of the present invention; DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0022] This experiment was conducted on an Ubuntu 20.04 system, with an Intel Core i9-10900K CPU, an NVIDIA RTX 3090 GPU (24GB video memory), and 32GB memory. The model was implemented based on the TensorFlow framework (TensorFlow 2.3 version, CUDA 11 version), and the Adam optimizer was used to optimize the model training process. The parameters used for training were a batch_size of 128, 50 epochs, and an initial learning rate of 0.01. During the training process, a strategy of dynamically adjusting the learning rate was adopted. When the model performance (validation loss) no longer decreased, the learning rate was automatically reduced with a strategy of learning_rate*0.8. It took 6.5 hours to complete the training task of the entire model.

[0023] This example demonstrates a point cloud traversability analysis network based on self-supervised learning. Figure 1 As shown in the figure, in the initialization phase, each frame in the point cloud frame sequence is first processed by the SLAM module to obtain the position of the current point cloud frame in the global point cloud coordinate system and the posture of the robot, and then output the corresponding odometer information. Subsequently, the backbone network receives the point cloud frame and odometer information as input, and based on the global point cloud features as a reference, provides two kinds of contextual information including intra-frame and global for each point in the current frame point cloud, and finally outputs the aggregated features. After that, the data flows to the traversability analysis module to perform qualitative and quantitative traversability analysis respectively. Finally, the output result is updated to the global result to complete the processing flow of a frame of point cloud.

[0024] In the backbone network part, the workflow is as follows: Figure 2 As shown in the figure. For each point cloud frame output, the front part of the backbone network first enhances the features of the point cloud frame, which is achieved by embedding the position of the point cloud frame in the global point cloud and recording the timestamp of the current point cloud frame, aiming to provide the network with contextual information of the current point cloud frame in time and space. Then, the data passes through a fully connected layer, followed by four random downsampling and local feature aggregation processes to gradually extract the features of the current point cloud frame. Subsequently, high-dimensional features are processed by multi-layer perceptrons and upsampling to finally generate aggregated features. The backbone network fully considers the multi-level feature information of the global point cloud, enriches the contextual information of the current frame point cloud, and avoids directly reprocessing the historical global point cloud, reducing the high time cost that may be incurred in processing the global point cloud.

[0025] The local feature aggregation function of the backbone network is composed of three main neural units, such as Figure 3As shown, they are the local space encoding unit, the attention pooling unit and the dilated residual block. Among them, the local space encoding unit is the core component of the local feature aggregation function, and its input includes the coordinates of the points in the point cloud frame, the features of the points, and the historical point cloud and its corresponding features. The local space encoding unit clearly embeds the three-dimensional (3D) coordinates of the relevant adjacent points by spatially encoding the adjacent points in the current point cloud and the adjacent points in the global point cloud respectively, ensuring that the features of the corresponding points are accurately represented. This structure gradually expands the receptive field as the depth of the network increases, so that the local space encoding unit can consider local geometric features from multiple levels, thereby helping the entire network to effectively learn complex point cloud structures. The above calculations are only for the points in the current point cloud frame, and the points in the historical point cloud are not processed. The points in the historical point cloud and their corresponding features are backups of the previous calculations. Since the features of the points correspond one-to-one to the coordinates of the points, they are very easy to maintain. The attention pooling unit is responsible for calculating the attention score. The present invention uses a shared function to learn a unique attention score for each feature, thereby achieving more effective feature aggregation.

[0026] The traversability analysis network can be divided into qualitative analysis subnet and quantitative analysis subnet according to its function. Its specific structure is as follows: Figure 4 As shown in the figure. The upper part is the qualitative analysis branch of traversability, which first contains a channel attention module, the purpose of which is to strengthen the feature channels that are more important for the qualitative analysis of traversability. The qualitative analysis network of traversability is trained using the Deep SVDD method. During training, the feature set corresponding to the positive sample points and the feature set corresponding to the unlabeled points need to be processed separately. The positive sample points are points around the robot trajectory, and the rest are unlabeled points. During the training process, only the feature embedding space of the positive sample points is used, and the goal is to minimize the volume of the hypersphere containing the positive sample data. When the qualitative evaluation of traversability of a certain point is performed, the qualitative traversability score of the point is the similarity between the embedded classification feature of the point and the center point of the hypersphere. The quantitative analysis network branch of traversability only processes the positive sample points (traversable points) during training. The features corresponding to the positive sample points are passed through a channel attention module, so that the feature channels that are more important for the quantitative analysis of traversability are strengthened. The goal of the subsequent stage is to learn a model that is responsible for estimating the point-by-point regression of the motion change factor of each positive sample point. The motion change factor is given by the average of the environmental force modulus in the K IMU data related to the point.

[0027] Obviously, the described embodiment is only one possible embodiment of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of the present invention.

Claims

1. A point cloud traversability analysis network based on self-supervised learning, It is characterized in that The network includes: ●A SLAM (Simultaneous Localization And Mapping) module, which receives each frame in the point cloud frame sequence, calculates the coordinates of the current point cloud frame in the global point cloud coordinate system and the robot's position and pose, and outputs the obtained odometer information; ●A backbone network module, which receives the point cloud frame and odometer information from the SLAM module, and uses the global point cloud features as a reference to provide two kinds of context information for each point in the current frame point cloud, namely, intra-frame and global context information, and finally outputs the aggregated features; ●A traversability analysis module that receives aggregated features from the backbone network module and performs qualitative and quantitative analysis of the traversability of the point cloud, respectively.

2. The point cloud traversability analysis network based on self-supervised learning according to claim 1, It is characterized in that The traversability analysis module includes a qualitative analysis subnet and a quantitative analysis subnet. The qualitative analysis subnet is used to predict whether the area represented by the points in the point cloud is traversable, and the quantitative analysis subnet is used to perform quantitative analysis on the traversable point cloud.

3. The point cloud traversability analysis network based on self-supervised learning according to claim 1 or 2, It is characterized in that The network can perform traversability analysis on point cloud data without manually annotating data, thereby achieving qualitative and quantitative analysis of the traversability of the point cloud.

4. The point cloud traversability analysis network based on self-supervised learning according to claim 1 or 2, It is characterized in that The network uses two information sources, namely, the point cloud frame information and the global point cloud information, when extracting point cloud features, and fully considers the connection and respective characteristics between single point cloud frame data and historical point cloud data.

5. The point cloud traversability analysis network based on self-supervised learning according to claim 1, It is characterized in that The backbone network module includes: ● A pre-processing unit for enhancing the features of the input current point cloud frame, including by embedding the position of the point cloud frame in the global point cloud and the time when the current point cloud frame was acquired, so as to understand the context information of the current point cloud frame in time and space; ●A fully connected layer, which receives the output of the pre-processing unit and gradually extracts the features of the current point cloud frame through random downsampling and local feature aggregation; ●A multi-layer perceptron that receives the output of the fully connected layer and processes the high-dimensional features through upsampling to generate aggregated features.

6. The point cloud traversability analysis network based on self-supervised learning according to claim 5, It is characterized in that The backbone network module further includes a local feature aggregation functional unit, which is composed of three neural units: local spatial encoding, attention pooling and dilated residual block. It is used to simultaneously find neighboring points in the same frame and neighboring points in the global point cloud when extracting feature values ​​of points in the current point cloud frame, and perform feature enhancement based on the two parts of neighboring points.

7. The point cloud traversability analysis network based on self-supervised learning according to claim 6, It is characterized in that The local spatial encoding function performs spatial encoding on the neighboring points of the point in the current point cloud and the neighboring points in the global point cloud respectively, thereby explicitly embedding the 3D coordinates of the related neighboring points, so that the features of the corresponding points are explicitly represented.

8. The point cloud traversability analysis network based on self-supervised learning according to claim 7, It is characterized in that The attention pooling unit calculates the attention score and uses a shared function to learn a unique attention score for each feature for a given set of adjacent features, thereby achieving more effective feature aggregation.

9. The point cloud traversability analysis network based on self-supervised learning according to claim 1, It is characterized in that The implementation configuration of the network includes: ●Central Processing Unit: Intel Core i9-10900K or its equivalent processor, used to perform computing tasks of the network; Graphics processor: NVIDIA RTX 3090 or its equivalent with 24GB of video memory to accelerate the training and inference process of the network; Memory: 32GB or more for storing and processing required data and model parameters; ● Operating system: Ubuntu 20.04 or its equivalent operating system for managing hardware resources and running the network; Deep Learning Framework: TensorFlow framework, version 2.3 or higher, for implementing and running the network; CUDA version: 11 or higher, for running the network on a graphics processor; ●Optimizer: Adam optimizer, used to optimize the model training process.

10. The point cloud traversability analysis network based on self-supervised learning according to claim 9, It is characterized in that The training parameters of the network include but are not limited to: ●Batch size: 128 or other appropriate batch size; ● Training rounds: 50 rounds or other appropriate training rounds; ●Starting learning rate: 0.01 or other appropriate starting learning rate; ●Strategy for dynamically adjusting the learning rate: When the model performance (validation loss) no longer decreases, the learning rate is automatically reduced with a strategy of learning_rate*0.8.