A Method for Perceiving the Navigable Area for a Robot
Through the deep learning model combined with the pavement segmentation model and the plane detection method, the robot can be output to output the driving area, solving the performance degradation of traditional positioning algorithms in complex scenarios, and improving the stability and safety of the robot system.
Patent Information
- Application Number
- CN202110677857.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-06-18
AI Technical Summary
Traditional positioning algorithms have degraded performance in occlusion and multipath scenarios, resulting in the inspection robots having problems such as positioning loss and positioning drift, which in turn affects the stability and security of the robot system.
The deep learning model is used to input RGB images and depth maps, and the travelable area is output. The road segmentation model and plane detection method are combined to combine multi-frame data and the final travelable area is obtained through weight accumulation.
It improves the accuracy of the robot's perceived driving area in complex scenarios, enhances the stability and safety of the system, and avoids the robot's misoperation caused by positioning loss and positioning drift.
Smart Images

Figure CN113343875B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and particularly to a method for perceiving drivable areas for robots. Background Art
[0002] Traditional positioning algorithms usually perform combined navigation or SLAM positioning based on GNSS signals, lidar, etc. However, GNSS signals are usually only applicable to open scenarios, and the positioning performance will drop sharply in occluded and multipath scenarios. The applicable scenarios of lidar are relatively broad, but it is easy to lose positioning in scenarios such as open spaces, dynamic environments, and long corridors.
[0003] The positioning system is the basis for inspection robots to correctly execute other tasks. A stable and reliable drivable area perception technology can provide a solid guarantee for the entire robot system to correctly and effectively execute commands. However, when there are problems such as positioning loss and positioning drift in the positioning system, it will cause the inspection robot to travel along the wrong trajectory, resulting in problems in all subsequent logical processing. In severe cases, it may lead to problems such as the robot tipping over or hitting people. Summary of the Invention
[0004] Object of the Invention: In view of the above deficiencies, the present invention proposes a method for perceiving drivable areas for robots, which inputs RGB images and depth maps based on a deep learning model and outputs drivable areas.
[0005] Technical Solution:
[0006] A method for perceiving drivable areas for robots includes the steps of:
[0007] Step 1, the robot collects a road surface depth image and an RGB image, and aligns the depth image and the RGB image;
[0008] Step 2, the RGB image is processed by a pre-trained road surface segmentation model to obtain a drivable area, and then the depth image is subjected to plane detection to obtain a drivable area, and the two drivable areas are aligned through Step 1;
[0009] Step 3, the drivable areas are fused by using a method of multi-frame superposition and weight accumulation to obtain a final drivable area.
[0010] In the said Step 1, the robot collects a road surface depth image and an RGB image through an RGB-D camera; or collects an RGB image through an RGB camera, and then collects a road surface depth image through a depth camera or a lidar.
[0011] Before step 3, the intersection over union of two drivable regions is used to determine whether the two drivable regions overlap. If they overlap, go to step 3; if not, calculate the average distance between all interior points in the drivable region obtained by plane detection and the plane, and normalize it to obtain the plane detection error value. Calculate the road surface segmentation error value according to the confidence level output by the road surface segmentation model, and compare the two. Use the drivable region with the smaller error value as the basis for the fusion in step 3.
[0012] The training of the road surface segmentation model in step 2 is as follows:
[0013] (11) The robot uses cameras installed on it to collect RGB images under different working conditions, at different time periods, or at different positions, and mark the road surfaces in the RGB images.
[0014] (12) Use a neural network to train the images obtained in step (11) through the neural network to obtain a labeling model, and use this labeling model to label the collected RGB images.
[0015] (13) Design a road surface segmentation model, and use the method of transfer learning to use the parameters of the labeling model obtained in step (12) as the initial parameters of the road surface segmentation model, and use the distillation learning method to train the labeled RGB images to obtain the road surface segmentation model.
[0016] The designed road surface segmentation model is specifically: adopt depth_wise convolution, point_wise convolution, channelshuffle or Inception convolution structure, and use a small 3*3 convolution kernel; in the detail branch of the network structure, use wide channels and shallow convolution, and in the semantic branch, use narrow channels and deep convolution. In the fusion layer, use an attention mechanism to fuse information of two different dimensions.
[0017] In step 2, the process of using the pre-trained road surface segmentation model to process the RGB image to obtain the drivable region is specifically as follows:
[0018] (21) Input the RGB image into the trained road surface segmentation model to obtain the road surface mask of the drivable region.
[0019] (22) Downsample the road surface mask obtained in step (21) to obtain a number of discrete points on the contour of the road surface mask.
[0020] (23) Use the Kalman filtering method to filter the number of discrete points on the contour of the road surface mask obtained after downsampling in step (22) to obtain the 2D coordinates of all points, and then obtain the drivable region.
[0021] After obtaining the drivable area in the second step, the area change of the drivable area in the front and rear frames is detected in real time and compared with a preset area change threshold; when the area change of the drivable area in the front and rear frames is greater than the preset area change threshold, it is determined that there is a misdetection, and the current frame image is input into the previously trained road surface segmentation model for training.
[0022] In the second step, the specific process of obtaining the drivable area by performing a plane detection on the depth image is as follows:
[0023] (31) Convert the depth image into a 3D point cloud, assuming the plane model is Ax + By + Cz + D = 0;
[0024] (32) Randomly select three points from the 3D point cloud to construct a plane model, and calculate the parameters of the plane model and the plane normal vector;
[0025] (33) Substitute all the points in the 3D point cloud into the plane model constructed in step (32) to calculate the number of inliers; among them, the criterion for judging whether a point is an inlier is to determine whether the distance from the point to the plane is less than a set threshold;
[0026] (34) Determine whether the number of inliers of the plane model reaches a set quantity. If not, go to step (35); if it has reached, use it as the final plane model; among them, the set quantity satisfies the condition that the number of inliers accounts for 80% of all the points in the 3D point cloud;
[0027] (35) Repeat steps (31) - (32), compare the number of inliers of the current plane model and the plane model at the previous moment, use the plane model with the largest number of inliers as the candidate plane model, and record its parameters and the number of inliers;
[0028] (36) Repeat step (34) until all the points in the 3D point cloud are traversed and the iteration ends. Use the final candidate plane model as the final plane model, and then obtain the drivable area.
[0029] The specific process of the third step is as follows:
[0030] (41) Multi-frame superposition:
[0031] (411) The slam system feeds back the translation and rotation relationship TF of the current frame image relative to the initial frame image;
[0032] (412) For each frame of image, obtain the driving area, transform it to the vehicle body coordinate system through the installation parameters of the RGB-D camera, and splice it through TF to obtain the spliced driving area;
[0033] (42) Weight accumulation:
[0034] (421) The entire map is virtualized into a checkerboard map with a grid resolution of 0.05m * 0.05m. When the grid corresponding to a certain area is determined by the robot to be a drivable area, the weight of the corresponding grid position is incremented by 1. Conversely, when it is determined to be an undrivable area, the weight of the corresponding grid position is decremented by 1;
[0035] (422) After driving for a set time, the grids with weights greater than or equal to the set threshold of 3 are drivable areas, and the grids with weights less than the set threshold of 3 are undrivable areas;
[0036] (43) Finally, the fused drivable area is obtained.
[0037] It also includes an alignment step:
[0038] (51) Install the calibration board at a certain point on the map, and collect the information of the calibration points on the calibration board through the RGB-D camera and the lidar respectively;
[0039] (52) Calculate the rotation matrix:
[0040] R * ·N = M
[0041] R * ·(N·N′) = M·N′
[0042] R * =(M·N′)·(N·N′) -1 (if rank(N·N′)=3)
[0043] In the above formula, R is the rotation matrix, N is the plane normal vector of the plane where the calibration points on the calibration board are located in the camera coordinate system, and M is the plane normal vector of the plane where the calibration points on the calibration board are located in the lidar coordinate system;
[0044] (53) Calculate the translation vector:
[0045]
[0046] In the above formula, α, β, γ are the magnitudes of the Euler angles of three degrees of freedom, x, y, z are the translation distances of three degrees of freedom, T is the translation vector, p i is the coordinate of the marker point in the camera coordinate system, q i,j is the coordinate of the feature point captured by the lidar in the lidar coordinate system, n i is the normal vector of the plane where the marker point is located in the camera coordinate system;
[0047] (54) Detect a feature point on the plane where the calibration board is located in the laser coordinate system, and at the same time select the corresponding feature point on the plane where the calibration board is located in the camera coordinate system and transform it to the lidar coordinate system according to steps (52) and (53);
[0048] (55) In the lidar coordinate system, the line connecting the two aforementioned feature points must be perpendicular to the normal vector of the plane where the calibration plate is located, that is, the dot product of the line and the normal vector is 0; repeat step (54) to select several feature points, obtain a system of equations and solve for RT;
[0049] (56) Calculate the drivable area in the slam system coordinate system through the RT obtained in step (55).
[0050] Beneficial effects: Different from the commonly used integrated navigation and SLAM systems, the present invention inputs RGB images and depth maps based on a deep learning model and outputs the drivable area. Under normal circumstances, the drivable area perception technology works in cooperation with the positioning system to enhance the system stability. When encountering some special situations, such as: 1) When dynamic factors such as people or vehicles appear, the module will be responsible for detecting the dynamic area; 2) When the positioning is lost, the module will be responsible for detecting the drivable area of the road surface to prevent the robot from tipping over; 3) When the positioning deviation occurs on a narrow road surface, the module will be responsible for detecting the road edge to prevent the robot from falling. Description of the Drawings
[0051] Figure 1 It is the training flowchart of the CNN model of the present invention;
[0052] Figure 2 It is the deployment and operation algorithm flowchart of the present invention;
[0053] Figure 3 It is the post - processing flowchart of the present invention.
[0054] Figure 4 It is the schematic diagram of the network structure of the road surface segmentation model of the present invention. Detailed Embodiments
[0055] The following further clarifies the present invention in conjunction with the drawings and specific embodiments.
[0056] The method for perceiving the drivable area of the present invention for a robot includes the steps:
[0057] Step 1: Training of the road surface segmentation model;
[0058] The road surface segmentation model adopts a CNN model. The input for training the CNN model is the labeled RGB image. Considering that obtaining high - quality labeled data is usually costly and time - consuming, semi - automatic annotation is introduced to reduce the annotation cost, and the CNN model is iteratively upgraded through feedback annotation technology. Then, combined with model lightweight design and model compression technology, the size of the CNN model is controlled to reduce the consumption of computing resources during the actual deployment and operation of the CNN model, so as to ensure that the CNN model can output the road surface mask in real time, as Figure 1 shown.
[0059] The specific steps are as follows:
[0060] (11) Data collection: The robot captures RGB images through the cameras installed on it; to ensure the diversity of data samples, relevant data is collected under different working conditions, at different time periods, or in different locations; such as in different weather conditions like sunny, cloudy, rainy, snowy, or foggy days, at different time periods like morning, noon, or evening, and substations of different years, locations, and sizes;
[0061] (12) Data annotation: Manual annotation is carried out using the annotation tool cvat, and the road surface in the captured RGB images is marked through masks. After annotation, it is used for the training of the road surface segmentation model;
[0062] When the number of RGB images to be annotated reaches a certain quantity, a large model will be trained for semi-automatic annotation to reduce the annotation cost;
[0063] Among them, the large model is a deep learning model, specifically a CNN model. Using the aforementioned RGB images as the model input, the model is trained, and finally a large model for annotating RGB images is obtained;
[0064] (13) Lightweight design of the road surface segmentation model:
[0065] Since the deep learning model needs to be deployed on the robot side, but the computing power of the robot side is usually limited, and there are high requirements for power consumption, running time, and resource consumption, it is necessary to design a road surface segmentation model with low computing power requirements and high deployment performance;
[0066] Among them, in terms of the selection of the convolution structure of the road surface segmentation model, lightweight convolution methods are adopted, such as depth_wise convolution, point_wise convolution, channel shuffle, and Inception structure. Small 3*3 convolution kernels are used and cross-channel connections are minimized as much as possible.
[0067] In terms of the network structure of the road surface segmentation model, as Figure 4 shown, the network of the road surface segmentation model of the present invention adopts two branches. The detail branch uses wide channels and shallow convolutions to obtain the detail features and rich feature expressions of the low_level of the image; the semantic branch uses narrow channels and deep convolutions. A larger receptive field can obtain high-level semantic context information such as contours, boundaries, and corners. Then, an attention mechanism is used in the fusion layer to fuse the information of the two layers with different dimensions, obtaining a high segmentation accuracy while ensuring low computing power;
[0068] (14) Road surface segmentation model training and compression: Using the method of transfer learning, the initial parameters of the road surface segmentation model are obtained based on the large model that has been preliminarily trained in step (12). In order to deploy on the robot side, the model needs to be further compressed. Therefore, the method of distillation learning is used to train and obtain the compressed road surface segmentation model; specifically as follows:
[0069] (141) Perform data augmentation on the labeled images. During training, operations such as random cropping, random masking, adding Gaussian noise, random affine transformation, and resizing are performed on the data to increase the diversity of the data;
[0070] (142) Based on the large model that has been preliminarily trained in step (12), use the method of transfer learning to obtain the initial parameters of the road surface segmentation model to achieve faster training speed and effect;
[0071] (143) Perform CNN training based on the initial parameters of the road surface segmentation model obtained in step (142) to obtain a large teacher model (without considering the model computing power and size) to ensure good segmentation effect (the edge of the drivable area is clear, and the segmentation MAP reaches 90%), and obtain the fixed parameters of the teacher model;
[0072] (144) Obtain the student model according to the road surface segmentation model designed in step (13) and the initial parameters in step (142). Add the intermediate layer output and the final output of the teacher model obtained in step (143) to the loss layer of the student model for training, so that the student model is easier to train and obtain good segmentation effect. Finally, obtain the compressed road surface segmentation model; after the road surface segmentation model training is completed, it will be converted from the original float32 type to int8 type to reduce the running time and memory consumption of the road surface segmentation model;
[0073] (15) Model transplantation: The framework used during the training of the road surface segmentation model is pytorch, and the feedforward framework actually deployed on the robot side is openvino. The model needs to be first converted to onnx and then to the bin file supported by openvino;
[0074] Step two, as Figure 2 shown, drivable area acquisition;
[0075] (21) The robot collects the road surface depth image and RGB image through the RGB-D camera installed on it, and aligns the depth image and RGB image; in the present invention, the RGB image can also be collected through the RGB camera, and the road surface depth image can be collected through the depth camera or lidar;
[0076] (22) AsFigure 3 As shown, input the RGB image into the road surface segmentation model trained in Step 1 to obtain the road surface mask of the drivable area;
[0077] (23) Post - processing algorithm processing:
[0078] (231) Contour fitting: Considering reducing the computational complexity of the drivable area fusion and occasional mis - detection problems, it is necessary to down - sample the obtained road surface mask to obtain several discrete points on the contour of the road surface mask;
[0079] (232) Use the Kalman filtering method to filter the several discrete points on the contour of the road surface mask obtained after down - sampling in Step (231) to obtain the 2D coordinates of all points, and then obtain the drivable area;
[0080] In the present invention, after obtaining the drivable area, the area change of the drivable area between the front and rear frames is detected in real - time. First, a threshold for area change is preset. When the area change of the drivable area between the front and rear frames is greater than the area change threshold, it is determined that there is a mis - detection. Save the image being processed at this time, and input this image into the previously trained road surface segmentation model for training, which is used for the above - mentioned feedback annotation, and re - collect and recognize and segment the road surface depth image and RGB image;
[0081] (24) Plane detection for the depth image: First, assume that the drivable area must be a plane, and this assumption conforms to the actual scenario of the substation; therefore, the drivable area can be obtained through a plane detection algorithm, that is, convert the depth image into a 3D point cloud, and perform plane detection on the 3D point cloud through the RANSAC algorithm to obtain the drivable area; specifically as follows:
[0082] (241) Assume the plane model is Ax + By + Cz + D = 0; randomly select three points from the 3D point cloud to construct a plane model, and calculate the parameters of this plane model and the plane normal vector;
[0083] (242) Substitute all points in the 3D point cloud into the plane model constructed in Step (241) to calculate the number of inliers; among them, the criterion for judging whether a point is an inlier can be set artificially. In the present invention, it is whether the distance of the point from the plane is less than the set threshold; in the present invention, the size of the set threshold is set to 0.5mm;
[0084] (243) Judge whether the number of inliers of this plane model reaches the set threshold. If not, go to Step (244); if it has reached, take it as the final plane model; among them, the set threshold meets the condition that the number of inliers accounts for 80% of all points in the 3D point cloud;
[0085] (244) Repeat steps (241) to (242), compare the number of inliers of the current plane model and the plane model at the previous moment, use the plane model with the largest number of inliers as the candidate plane model, and record its parameters and the number of inliers;
[0086] (245) Repeat step (243) until all points in the 3D point cloud are traversed and the iteration ends. Use the final candidate plane model as the final plane model, and then obtain the drivable area;
[0087] (25) Align the drivable area obtained in step (23) and the drivable area obtained in step (24) according to the internal parameters of the RGB-D camera;
[0088] (26) Through the two aligned drivable areas obtained in the previous steps, in most cases the two drivable areas coincide, but not all planes are drivable and there is a certain probability of false detection in the road surface segmentation model;
[0089] (261) Determine whether the two drivable areas coincide by comparing the intersection over union (IoU) of the two drivable areas with the set IoU threshold, where the set IoU threshold is 0.6 - 1; if they coincide, go to step three with the drivable area obtained in step (245) as the benchmark; if they do not coincide, go to step (262);
[0090] (262) Calculate the average distance between all inliers in the drivable area detected by plane detection and the plane, normalize it through the standard threshold to obtain the plane detection error value, calculate the road surface segmentation error value according to the confidence level output by the road surface segmentation model in step (22), and compare the two. Use the drivable area with the smaller error value as the benchmark; where the standard threshold is the average distance of all points in the plane to the plane set in advance;
[0091] Step Three: Drivable Area Fusion:
[0092] (31) Multi-frame Superposition: Multi-frame superposition requires the slam system to feedback the translation and rotation relationship of the current frame relative to the initial frame, simply referred to as TF; for each detected drivable area, transform it to the vehicle body coordinate system through the calibrated external parameters, and then splice the drivable areas detected in multiple frames through TF;
[0093] (32) Weight Accumulation: Weight accumulation can effectively avoid false detection in a single frame. First, virtualize the entire map into a checkerboard graph with a grid resolution of 0.05m * 0.05m. When a grid corresponding to a certain area is determined by the robot to be a drivable area, the weight of the corresponding grid position is increased by 1, and vice versa, when it is determined to be a non-drivable area, the weight of the corresponding grid position is decreased by 1; after driving for a set time, the set threshold is 3, and the grids with weights greater than or equal to 3 are drivable areas, and the grids with weights less than 3 are non-drivable areas;
[0094] (33) Thus, the finally fused drivable area is obtained;
[0095] Step 4: Calibration of RGB-D camera and lidar: After obtaining the fused drivable area, it is necessary to align it with the existing slam system coordinate system in order to use the detected drivable area for the positioning system. Therefore, it is necessary to calibrate the external parameters RT of the RGB-D camera and the lidar;
[0096] (41) Before calibration, a calibration board needs to be prepared in advance for the detection of planar feature points and the calculation of planar normal vectors; the calibration board is installed at a certain point on the map, and the information of the marked points on the calibration board is collected by the RGB-D camera and the lidar respectively;
[0097] (42) Calculate the rotation matrix:
[0098] R * ·N = M
[0099] R * ·(N·N′) = M·N′
[0100] R * = (M·N′)·(N·N′) -1 (if rank(N·N′) = 3)
[0101] In the above formula, R is the rotation matrix, N is the planar normal vector of the plane where the marked points on the calibration board are located in the camera coordinate system, and M is the planar normal vector of the plane where the marked points on the calibration board are located in the lidar coordinate system;
[0102] (43) Calculate the translation vector:
[0103]
[0104] In the above formula, α, β, γ are the magnitudes of the Euler angles of three degrees of freedom, x, y, z are the translation distances of three degrees of freedom, T is the translation vector, p i is the coordinate of the marked point in the camera coordinate system, q i,j is the coordinate of the feature point captured by the lidar in the lidar coordinate system, n i is the normal vector of the plane where the marked point is located in the camera coordinate system.
[0105] (44) In the lidar coordinate system, the line connecting two feature points on a plane must be perpendicular to the normal vector of the plane. First, a feature point on the plane where the calibration board is located is detected in the lidar coordinate system. At the same time, a feature point on the plane where the calibration board is located is selected in the camera coordinate system and transformed to the lidar coordinate system through the extrinsic rotation and translation transformation. Theoretically, if the extrinsic parameters are accurate, the line connecting the two feature points must be perpendicular to the normal vector of the plane where the calibration board is located, and the dot product of the line and the normal vector is 0. Therefore, only by selecting multiple groups of feature points to minimize the above formula can the extrinsic parameters RT be calculated.
[0106] (45) Calculate the drivable area in the slam system coordinate system through the extrinsic parameters RT obtained in step (44).
[0107] Traditional plane detection methods do not have road surface semantic information and cannot be directly used for drivable area detection. And CNN models usually cannot be directly applied in engineering due to the diversity of training data samples and the lack of generalization of the model itself. The main idea of the drivable area perception technology is to fuse the CNN deep learning method and the traditional plane detection method. Obtain semantic information through the CNN model, and then introduce the plane detection method to combine with the CNN model to eliminate occasional false detections to improve the robustness of the detection. In the post-processing module, an RGB image of the false detections of the CNN model is obtained through an adaptive threshold, and then the false detection image is put back into the model for enhanced training to form a feedback annotation closed loop to ensure the iterative upgrade ability of the method.
[0108] The preferred embodiments of the present invention have been described in detail above, but the present invention is not limited to the specific details in the above embodiments. Within the technical concept scope of the present invention, various equivalent transformations (such as quantity, shape, position, etc.) can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A method for perceiving a drivable area for a robot, characterized in that: It includes the steps: Step 1: The robot collects a road surface depth image and an RGB image, and aligns the depth image and the RGB image; Step 2: Process the RGB image through a pre-trained road surface segmentation model to obtain a drivable area, then perform a plane detection on the depth image to obtain a drivable area, and align the two drivable areas through Step 1; Step 3: Use the method of multi-frame superposition and weight accumulation to fuse the drivable areas to obtain the final drivable area, including: 1) Multi-frame superposition: 11) The slam system feeds back the translation and rotation relationship TF of the current frame image relative to the initial frame image; 12) Obtain the driving area for each frame of image, transform it to the vehicle body coordinate system through the installation parameters of the RGB-D camera, and splice it through TF to obtain the spliced driving area; 2) Weight accumulation: 21) Virtualize the entire map into a checkerboard map with a grid resolution of 0.05m * 0.05m. When a grid corresponding to a certain area is determined by the robot to be a drivable area, the weight of the corresponding grid position is increased by 1. Conversely, when it is determined to be a non-drivable area, the weight of the corresponding grid position is decreased by 1; 22) After driving for a set time, the grids with weights greater than or equal to the set threshold 3 are drivable areas, and the grids with weights less than the set threshold 3 are non-drivable areas; 3) Combine 1) and 2) to obtain the fused drivable area.
2. The method for perceiving a drivable area according to claim 1, characterized in that: In the said Step 1, the robot collects a road surface depth image and an RGB image through an RGB-D camera; or collects an RGB image through an RGB camera, and then collects a road surface depth image through a depth camera or a lidar.
3. The method for perceiving a drivable area according to claim 1, characterized in that: Before the said Step 3, judge whether the two drivable areas overlap through the intersection over union of the two drivable areas. If they overlap, go to Step 3; if they do not overlap, calculate the average distance between all the inliers in the drivable area obtained by plane detection and the plane and normalize it to obtain a plane detection error value, calculate a road surface segmentation error value according to the confidence level output by the road surface segmentation model, compare the two, and use the drivable area with a smaller error value as the basis for the fusion in Step 3.
4. The method for perceiving a drivable area according to claim 1, characterized in that: The training of the road surface segmentation model in the said Step 2 is as follows: (11) The robot respectively collects RGB images under different working conditions, different time periods or different positions through the cameras installed on it and marks the road surfaces in the RGB images; (12) Use a neural network to train the images obtained in step (11) through the neural network to obtain a marking model, and use this marking model to mark the collected RGB images; (13) Design a road surface segmentation model, and use the method of transfer learning to use the parameters of the marking model obtained in step (12) as the initial parameters for obtaining the road surface segmentation model, and use the distillation learning method to train the marked RGB images to obtain the road surface segmentation model.
5. The drivable area perception method according to claim 4, characterized in that: The designed road surface segmentation model is specifically: adopting depth_wise convolution, point_wise convolution, channel shuffle or Inception convolution structure, and using a small convolution kernel of 3*3; in the detail branch of the network structure, wide channels and shallow convolution are adopted, and in the semantic branch, narrow channels and deep convolution are adopted, and an attention mechanism is adopted in the fusion layer to fuse information of two layers with different dimensions.
6. The drivable area perception method according to claim 1, characterized in that: In the second step, the specific process of obtaining the drivable area by processing the RGB image through the pre-trained road surface segmentation model is as follows: (21) Input the RGB image into the trained road surface segmentation model to obtain the road surface mask of the drivable area; (22) Downsample the road surface mask obtained in step (21) to obtain a number of discrete points on the contour of the road surface mask; (23) Use the Kalman filtering method to filter the number of discrete points on the contour of the road surface mask obtained after downsampling in step (22) to obtain the 2D coordinates of all points, and then obtain the drivable area.
7. The drivable area perception method according to claim 1, characterized in that: After obtaining the drivable area in the second step, the area change of the drivable area in the front and rear frames is detected in real time and compared with the preset area change threshold; when the area change of the drivable area in the front and rear frames is greater than the preset area change threshold, it is determined that there is a misdetection, and the current frame image is input into the previously trained road surface segmentation model for training.
8. The drivable area perception method according to claim 1, characterized in that: In the second step, the specific process of obtaining the drivable area by performing plane detection on the depth image is as follows: (31) Convert the depth image into a 3D point cloud, and assume the plane model is Ax + By + Cz + D = 0; (32) Randomly select three points from the 3D point cloud to construct a plane model, and calculate the parameters of the plane model and the plane normal vector; (33) Substitute all points in the 3D point cloud into the plane model constructed in step (32) to calculate the number of inliers; among them, the basis for judging whether it is an inlier is to judge whether the distance of the point from the plane is less than the set threshold; (34) Judge whether the number of inliers of the plane model reaches the set quantity. If not, go to step (35); if it has reached, use it as the final plane model; among them, the set quantity meets the condition that the number of inliers accounts for 80% of all points in the 3D point cloud; (35) Repeat steps (31) to (32), compare the number of inliers of the current plane model and the plane model at the previous moment, and use the plane model with the largest number of inliers as the candidate plane model, and record its parameters and the number of inliers; (36) Repeat step (34) until all points in the 3D point cloud are traversed and the iteration ends, and use the final candidate plane model as the final plane model, and then obtain the drivable area.
9. The drivable area perception method according to claim 1, characterized in that: It further includes an alignment step: (51) Install the calibration board at a certain point on the map, and collect the information of the calibration points on the calibration board through the RGB-D camera and the lidar respectively; (52) Calculate the rotation matrix: R * .N = M R * .(N·N′) = M.N′ R * = (M.N′)·(N.N′) -1 , if rank(N·N′) = 3 In the above formula, R is the rotation matrix, N is the plane normal vector of the plane where the calibration points on the calibration board are located in the camera coordinate system, and M is the plane normal vector of the plane where the calibration points on the calibration board are located in the lidar coordinate system; (53) Calculate the translation vector: In the above formula, α, β, and γ are the magnitudes of the Euler angles of the three degrees of freedom, x, y, and z are the translation distances of the three degrees of freedom, T is the translation vector, p i is the coordinate of the marked point in the camera coordinate system, q i,j is the coordinate of the feature point captured by the lidar in the lidar coordinate system, n i is the normal vector of the plane where the marked point is located in the camera coordinate system; (54) Detect a feature point on the plane where the calibration board is located in the laser coordinate system, and at the same time select the corresponding feature point on the plane where the calibration board is located in the camera coordinate system and transform it to the lidar coordinate system according to steps (52) and (53); (55) In the lidar coordinate system, the connection line of the above two feature points must be perpendicular to the normal vector of the plane where the calibration board is located, that is, the dot product of the connection line and the normal vector is 0; repeat step (54) to select several feature points, obtain a system of equations and solve for RT; (56) Calculate the drivable area in the slam system coordinate system through the RT obtained in step (55).
Citation Information
Patent Citations
Road drivable area identification method based on binocular stereoscopic vision
CN110008848A